Live roles / clera
Compatibility brief · Discovered by NoBoards 9h ago

Research Engineer, Benchmarks

clera · San Francisco
Sponsorship statedMode not stated2+ years requestedIndustry: ai
Direct source excerpts
- Visa sponsorship available.
Extracted role summary · not eligibility evidence

Research Engineer, Benchmarks at clera, designing and implementing AI benchmark evaluations for frontier agents, on-site in San Francisco.

PythonDockerLinuxAIML

ABOUT THE ROLE

Join a small, highly technical team of researchers and engineers — including International Olympiad medalists and published AI researchers — at an early-stage startup building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks, you'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to measure real-world agent performance. This is a critical, high-ownership role at the intersection of research rigor and engineering execution.

The company operates in the AI/ML evaluation and reinforcement learning infrastructure space, providing a platform for building, running, and scaling RL environments and post-training datasets. The team is based in San Francisco, CA and works on-site. Visa sponsorship is available.

WHAT YOU'LL DO

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.

WHAT WE'RE LOOKING FOR

Required

  • 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments.
  • Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
  • Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure.
  • Experience building and operating infrastructure to reliably run AI models or agents against benchmark or evaluation tasks at scale.
  • Experience developing metrics, statistical analyses, or validation studies to assess benchmark difficulty, reliability, and real-world correlation.
  • Experience collaborating with subject-matter experts to translate domain workflows into benchmark tasks and evaluation criteria.
  • Experience analyzing workflows across diverse technical or business domains to inform task design.
  • Strong technical writing skills — able to produce benchmark reports and documentation for research and engineering audiences.

Nice to Have

  • Published papers or technical blog posts on AI benchmarking, model evaluation, or model failure modes.
  • Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation.
  • Background at frontier AI labs, research institutions, or involvement in widely used public benchmark projects.

Traits We Value

  • Deep curiosity about how workflows operate across varied domains.
  • Sharp attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design.
  • Ability to reason from first principles about task design, scoring, and failure modes.
  • Comfort thriving in unstructured problem spaces and working independently in a fast-paced, early-stage environment.
  • Excellent communication skills for collaborating across time zones and with technical teams.

COMPENSATION & BENEFITS

  • Salary: $150,000 – $250,000 USD annually, depending on experience.
  • Early-stage equity participation.
  • Visa sponsorship available.

LOCATION

This is an on-site role based in San Francisco, CA, United States. Candidates must be willing and able to work from the office. Fully remote arrangements are not available for this position.