Live roles / clera
Compatibility brief · Discovered by NoBoards 1d ago

Lead Research Engineer, Data Quality

clera · San Francisco
Sponsorship not statedMode not stated

ABOUT THE ROLE

This is a senior individual contributor and team lead position at an early-stage AI infrastructure startup building the tooling and data pipelines that power frontier RL-based agent training. You'll own the strategy and execution of data quality systems end-to-end — from evaluation frameworks to synthetic data validation — and help shape internal research culture around what makes agent training data genuinely useful.

WHAT YOU'LL DO

  • Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
  • Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
  • Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
  • Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
  • Turn qualitative research insights into production systems — internal tools, dashboards, validation pipelines, and feedback loops.
  • Build internal research taste around what makes agent training data realistic, learnable, diverse, and reliable — not just superficially correct.
  • Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed.

WHAT WE'RE LOOKING FOR

  • 5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation.
  • Demonstrated track record leading technical teams or projects on ambiguous problems from definition through to iteration.
  • Advanced proficiency in Python, Docker, and Linux environments.
  • Experience building QC systems, evals, benchmarks, synthetic data pipelines, or model evaluation infrastructure.
  • Strong intuition for the characteristics of high-quality training data for AI agents — realistic, learnable, diverse, reliable, and useful.
  • Ability to design metrics, experiments, and QA/QC processes, not just execute them.
  • Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
  • Strong written communication skills with the ability to explain methodology clearly to technical and non-technical audiences.
  • Comfort navigating complex systems involving domain experts, vendors, model outputs, graders, and infrastructure.
  • Prior experience in an early-stage startup environment; able to work independently and move quickly.

COMPENSATION & BENEFITS

Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available.

LOCATION

On-site in San Francisco, CA, United States.