Live roles / clera
Compatibility brief · Discovered by NoBoards 9h ago

Research Engineer, Synthetic Data

clera · San Francisco
Sponsorship statedOn-site2+ years requestedIndustry: ai
Direct source excerpts
- Visa sponsorship available
This role is on-site in San Francisco, CA. Candidates based in or willing to relocate to San Francisco are strongly preferred. The team also has a presence in Singapore.
Extracted role summary · not eligibility evidence

Research Engineer role focused on building synthetic data pipelines for AI agents at a ~15-person startup in San Francisco.

PythonDockerLinux

ABOUT THE ROLE

We're a ~15-person engineering team — made up of Olympiad medalists and published researchers — building infrastructure that aligns AI to real-world workflows through reinforcement learning environments and post-training data. We're hiring Research Engineers to own the synthetic data pipeline: transforming domain-specific workflows into scalable, high-quality training tasks for AI agents.

This is a high-ownership, low-bureaucracy role. You'll be working in genuinely unstructured problem spaces where the roadmap is yours to define. Visa sponsorship is available.

WHAT YOU'LL DO

  • Build and maintain the end-to-end synthetic data pipeline, converting domain-specific workflows into realistic, structured, and challenging training tasks for AI agents.
  • Collaborate with subject-matter experts to generate synthetic tasks across professional and technical domains.
  • Design synthetic task generation methods that produce diverse, realistic, and learnable outputs.
  • Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
  • Analyze model and agent performance on synthetic tasks to understand what they teach and where they break down.
  • Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.

WHAT WE'RE LOOKING FOR

Required

  • 2–4 years of experience in software engineering, ML engineering, or AI research — with a track record of shipping data pipelines, ML infrastructure, or synthetic data systems.
  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python; comfortable working in Linux environments with containerization tools such as Docker.
  • Demonstrated understanding of synthetic data quality criteria and evaluation metrics (diversity, realism, learnability) and their limitations — from production or research work.
  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
  • Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements.

Nice to Have

  • Experience detecting edge cases, inconsistencies, or quality issues in synthetic or algorithmically generated datasets.
  • Experience creating synthetic tasks, data, or evaluations across multiple distinct professional or technical domains.
  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines.

You'll thrive here if you

  • Reason from first principles about task design, scoring, and failure modes.
  • Are detail-oriented and naturally spot subtle inconsistencies in data and systems.
  • Are energised by early-stage, ambiguous environments rather than frustrated by them.
  • Communicate clearly and collaborate effectively across time zones.

COMPENSATION & BENEFITS

  • Salary: $150,000 – $250,000 USD annually
  • Visa sponsorship available
  • Equity participation (early-stage startup)

LOCATION

This role is on-site in San Francisco, CA. Candidates based in or willing to relocate to San Francisco are strongly preferred. The team also has a presence in Singapore.