Compatibility brief · Discovered by NoBoards 1d ago
Lead Research Engineer, Data Quality
clera · San Francisco
Sponsorship not statedMode not stated
ABOUT THE ROLE
This is a senior individual contributor and team lead position at an early-stage AI infrastructure startup building the tooling and data pipelines that power frontier RL-based agent training. You'll own the strategy and execution of data quality systems end-to-end — from evaluation frameworks to synthetic data validation — and help shape internal research culture around what makes agent training data genuinely useful.
WHAT YOU'LL DO
- Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
- Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
- Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
- Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
- Turn qualitative research insights into production systems — internal tools, dashboards, validation pipelines, and feedback loops.
- Build internal research taste around what makes agent training data realistic, learnable, diverse, and reliable — not just superficially correct.
- Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
WHAT WE'RE LOOKING FOR
- 5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation.
- Demonstrated track record leading technical teams or projects on ambiguous problems from definition through to iteration.
- Advanced proficiency in Python, Docker, and Linux environments.
- Experience building QC systems, evals, benchmarks, synthetic data pipelines, or model evaluation infrastructure.
- Strong intuition for the characteristics of high-quality training data for AI agents — realistic, learnable, diverse, reliable, and useful.
- Ability to design metrics, experiments, and QA/QC processes, not just execute them.
- Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
- Strong written communication skills with the ability to explain methodology clearly to technical and non-technical audiences.
- Comfort navigating complex systems involving domain experts, vendors, model outputs, graders, and infrastructure.
- Prior experience in an early-stage startup environment; able to work independently and move quickly.
COMPENSATION & BENEFITS
Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available.
LOCATION
On-site in San Francisco, CA, United States.