Head of Central Quality and Project Enablement

HumanSignalSan Francisco, CaliforniaOn-siteFull-timePrincipal, 12–15+ yearsListed 2 hours ago

Apply now

About this role

About HumanSignal

Real-world data is the competitive edge in AI.

HumanSignal is a human data partner for companies building AI models and products. Our customers ship better AI, faster, because we partner with their researchers from real-world data creation to annotation to delivery.

We design and create datasets from scratch, recruit and manage the domain experts who evaluate model output, and run everything through our own platform, Label Studio, the open-source standard for data labeling and evaluation, used by over 1 million practitioners worldwide.

We specialize in the operationally complex: real-world data collection, multimodal pipelines, and multi-step workflows. Advanced ML and AI teams use our enterprise platform to run their own data factories, and our services team to extend their reach where in-house capacity runs out.

If you want to do work that materially shapes how the next generation of AI products gets built, we'd love to talk.

Head of Quality & Project Enablement, Data Services

Location: San Francisco or Austin, TX
Reports to: Head of Data Services
Compensation: $140,000 - $180,000

About the role

Every Data Services engagement succeeds or fails on two things: whether the team doing the work understands the spec, and whether we can prove the output meets it. This role owns both.

As Head of Quality & Project Enablement, you'll take each customer spec and turn it into the materials that get annotators and experts productive. You'll then design the quality pipeline and sampling methodology that verifies the work before it ships. When quality slips, you'll trace it back to its source, whether that's a gap in the spec, the training, or the review process, and fix it.

What you'll own

Project enablement

- Turn customer spec documents into project-ready enablement packages: annotator guidelines, decision trees, edge-case libraries with worked examples, and onboarding walkthroughs.

- Build qualification tests and gold-standard sets that confirm annotators understand the spec before they touch production work.

- Run pilot and calibration rounds at project kickoff to surface spec ambiguities early. Resolve those ambiguities with the customer and delivery lead before scaling up.

- Maintain versioned guidelines throughout the project. Roll out updates as new edge cases appear, and confirm the workforce has absorbed the changes.

Quality pipeline design

- Design the quality plan for every project. Choose the review structure (gold tasks, overlap/consensus, multi-stage review, expert adjudication) based on the task type, risk, and budget.

- Set the sampling methodology. Size samples to hit target confidence levels, stratify by class and difficulty, and adjust review rates for each annotator based on performance.

- Select the right metrics for each task, such as accuracy against gold, agreement measures like κ or α, or per-class error rates. Set acceptance thresholds that fit how subjective the task is.

- Configure these workflows in our labeling platform, and turn what works into reusable quality playbooks by task type.

Delivery quality control

- Own final quality sign-off. No delivery ships without meeting its agreed acceptance criteria.

- Monitor quality throughout each project. Catch drift early, run root-cause analysis on defects, and drive corrective action for individuals and for guidelines.

- Produce a clear quality report for every delivery that shows the methodology, the results, and any known limitations.

What you'll bring

- 6+ years in quality assurance or quality operations for data labeling, human data, or ML training/evaluation data, including experience leading a quality function or team

- A proven ability to translate complex or ambiguous specs into guidelines and training that produce consistent results

- Strong applied statistics for QA: sampling design, confidence intervals, and agreement metrics, plus the ability to explain your choices to a customer

- Hands-on experience configuring review and QA workflows in an annotation platform

- Proficiency in Claude Code for quality analysis

- Crisp written communication. Your guidelines and quality reports are the product.

Nice to have: Experience with LLM, RLHF, or preference-data projects; experience with expert or domain-specialist workforces; a background in instructional design; familiarity with Label Studio; experience with model-assisted QA.

What success looks like

- 90 days: A standard enablement package and quality plan template is in use on every new project. You own sign-off on all active deliveries.

- 6 months: Annotators ramp to target accuracy faster, mid-project guideline churn drops, and rework and customer-reported defects are measurably down.

- 12 months: Quality reports and methodology are a selling point for Data Services, and the playbooks are in place so the function can scale beyond you.