Live roles / pareto-ai
Compatibility brief · Discovered by NoBoards 2d ago

Applied AI Engineer

pareto-ai · US Remote
Sponsorship not statedRemote listed8+ years requestedIndustry: ai
Extracted role summary · not eligibility evidence

Build and scale internal AI data pipelines, HITL workflows, and agentic systems for leading frontier labs.

PythonKubernetesDockerAirflowSynthetic DataHuman-in-the-Loop (HITL)

ABOUT PARETO

Humanity is in a virtuous cycle: human insight improves AI, and better AI expands what people can do. Sustaining it depends on the one input that can't be automated: expert human judgment https://pareto.ai/blog/debating-persuasive-llms-truthful-answers.

At Pareto, we build the platform that turns that judgment into the data https://pareto.ai/blog/community-driven-knowledge-resource-ai, evals https://pareto.ai/blog/introducing-attunebench, and RL environments frontier models https://pareto.ai/blog/llm-metacognition-shared-and-shallow learn from. We work with leading frontier labs like Anthropic and GDM, and we give skilled people everywhere a way to shape the future of AI and share in what it creates.

This RL environment and human-data infrastructure is already in production. Our job now is to scale it.

About This Role

Applied AI Engineers at Pareto build the internal systems that let the company scale. The work is unglamorous in the best way: data pipelines, training workflows, HITL infrastructure. Your job is to get ahead of the complexity before it slows things down. Design the agentic workflows and automation pipelines that let Pareto operate faster and break less.

You work directly with research, data, and product teams. You'll scope hard problems, decide what's worth building, and own what ships. The work feeds into training pipelines at Anthropic, GDM, and other frontier labs.

What You'll Do

  • Design and build pipelines that generate synthetic tasks and eval environments for model training. Factory floor of AI development, not the models themselves.
  • Architect HITL workflows. Decide what gets automated vs. what needs human review, how state survives handoffs, and what breaks first under load.
  • Lead system design discussions. Write one-page scoping docs that surface hidden risks before development starts.
  • Assess whether a technical idea is worth building. Get early signal, align stakeholders, then decide fast.
  • Work across research, ops, and data teams. Translate ambiguous requirements into concrete architecture.
  • Build frameworks and guidelines the rest of the team can actually use.

What You'll Need

  • 8+ years of software engineering experience with end-to-end ownership of complex systems
  • Software engineering foundation first. You think in systems and architecture.
  • Production experience with agentic workflows, HITL pipelines, and LLM-powered applications. RAG, vector stores, multi-model stacks. Shipped and monitored, not demoed.
  • Solid context engineering practices. You know where AI breaks and how to build around it.
  • Experience with distributed systems applied to AI or data platforms. It needs to be reliable and observable at scale.
  • Daily use of agentic coding tools (Claude Code, Cursor, or equivalent).
  • Strong written and verbal English communication. Design discussions, stakeholder pushback, architecture docs. None of it can be a bottleneck.

You'll Stand Out If You Have

  • Experience at an AI data company (Scale, Surge, Snorkel, Labelbox, or similar), especially on synthetic data pipelines, eval environments, or task generation systems
  • Experience building annotation workflows, data labeling interfaces, or data collection pipelines
  • Familiarity with preference data and reward models: RLHF, RLVR, or similar
  • Proficiency with our stack: Python, TypeScript, AWS, GCP, Terraform, Temporal Cloud, containerization, LLM gateways, and data pipeline tooling

You Probably Aren't the Fit If

  • You need requirements locked before you start. What Pareto builds is new, and the specs evolve with the research.
  • You want to stay close to one system. This role spans architecture, delivery, and cross-team work in the same week.
  • You'd rather be research-adjacent than production-first. This work ends up in training pipelines, not decks.

Why Pareto

The pipelines you build here train models at Anthropic and GDM. Not adjacent to that work. Inside it. What you ship determines what the frontier labs can build next. Equity is part of the package.