Lead Research Engineer, Data Quality
Visa sponsorship available.
This is an on-site role based in San Francisco, CA.
Lead the data quality team to build systems evaluating and scaling training data for frontier AI agents.
ABOUT THE ROLE
Join an early-stage AI/ML startup building the infrastructure layer for reinforcement learning environments and post-training data at the frontier. As Lead Research Engineer, Data Quality, you will own the strategy and systems that measure, improve, and scale training data for frontier agents. This is a high-impact, hands-on leadership role where you'll shape both the technical direction and internal research culture around what makes agent training data truly useful.
The company is a well-funded, rapidly growing AI infrastructure platform (Series A/B stage) focused on RL environment tooling, synthetic data generation, and model evaluation — working directly with AI labs and research teams.
WHAT YOU'LL DO
- Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
- Define the data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
- Develop new methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
- Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
- Turn qualitative research insights into production systems — internal tools, dashboards, validation pipelines, and feedback loops.
- Help build internal research intuition around what makes agent training data realistic, learnable, diverse, reliable, and useful — not just superficially correct.
- Mentor other research engineers, maintaining a high bar for technical rigor, clarity, and execution speed.
WHAT WE'RE LOOKING FOR
Required
- 5+ years of relevant engineering or research experience.
- Proven track record leading technical teams on ambiguous projects from problem definition through implementation and iteration.
- Advanced proficiency in Python, Docker, and Linux environments.
- Hands-on experience building QC systems, evals, benchmarks, synthetic data pipelines, validation workflows, or model evaluation infrastructure.
- Deep intuition for data quality — what makes training tasks realistic, learnable, diverse, reliable, and useful.
- Comfort designing metrics, experiments, and QA/QC processes, not just executing them.
- Strong written communication; ability to explain methodology clearly to researchers, engineers, and external audiences.
- Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
- Early-stage startup experience with demonstrated ability to work independently in fast-paced environments.
- Detail-oriented mindset with an eye for subtle inconsistencies or edge cases in data.
Nice to Have
- Background in reinforcement learning, reward modeling, or agent evaluation.
- Experience shipping production research infrastructure (not just prototypes).
- Familiarity with large-scale task execution systems or distributed evaluation pipelines.
COMPENSATION & BENEFITS
- Salary: $150,000 – $250,000 USD annually, commensurate with experience.
- Equity participation in an early-stage, high-growth AI company.
- Visa sponsorship available.
LOCATION
This is an on-site role based in San Francisco, CA. Candidates should be willing and able to work from the office. Relocation support may be available for strong candidates.