Live roles / clera
Compatibility brief · Discovered by NoBoards 3h ago

Founding AI Engineer

clera · Palo Alto
Sponsorship unavailableOn-siteHiring geography needs source reviewTimezone: PST ±2h1+ years requestedIndustry: ai
Direct source excerpts
Visa sponsorship: Not available
Willingness to work on-site 5 days/week in San Francisco, CA.
Extracted role summary · not eligibility evidence

Founding AI Engineer to build and ship production multimodal agentic systems on industrial smart glasses.

PythonPyTorchvLLMTritonRay ServeONNXTensorRTHugging Face TransformersLangChainDocker

ABOUT THE ROLE

We're a seed-stage wearable AI company building an AI co-pilot for skilled field technicians — delivered through industrial smart glasses — that helps workers in high-stakes industries like data centers and energy infrastructure operate at an expert level. Our stack spans edge inference, real-time voice/video, and agentic visual reasoning running on real hardware in demanding environments.

We're looking for a Founding AI Engineer with 1–5 years of experience who has shipped production multimodal and agentic AI systems to real users. This is a hands-on, end-to-end ownership role at the core of our product — not a research or prototyping position.

WHAT YOU'LL DO

  • Build and ship the production agentic Vision-Language Model (VLM) pipeline running on industrial smart glasses — multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service).
  • Own model orchestration and runtime optimization for edge inference, balancing model quality against latency with graceful degradation across variable connectivity conditions.
  • Design and build the evaluation harness and data flywheel from scratch — failure-mode capture, customer-data fine-tuning loops, and measurable quality improvements each release cycle.
  • Ship real-time voice and video AI interfaces tailored to different end-user profiles: video-heavy, conversational speech, and proactive alerting.
  • Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data.
  • Drive multimodal model training for on-premise deployments: open-source model SFT, RL post-training, and quantization.

WHAT WE'RE LOOKING FOR

Required (Dealbreakers)

  • Demonstrable track record shipping production multimodal and computer vision systems in the VLM era, owning the model layer end-to-end — with hands-on expertise in visual-language and/or video-language VLMs/VLAs.
  • Bachelor's degree in Computer Science, Machine Learning, Engineering, or equivalent — graduated 2018 or later.
  • Willingness to work on-site 5 days/week in San Francisco, CA.

Also Required

  • Experience with applied agentic AI or model orchestration in production settings.
  • Experience building production AI products at a startup or high-ownership AI team, or relevant big-company experience (AR/smart glasses, real-time video/streaming, on-device/edge ML) paired with a strong builder signal (e.g., early startup, side projects, open-source contributions).
  • Experience with rigorous evaluation methodologies — ground-truth evals, trajectory evals, tool-call accuracy, and regression testing for comparing models and orchestration stacks.
  • Strong foundation in CS, ML, or engineering, or a demonstrated equivalent through shipping history.

Nice to Have

  • Experience with production AR or wearable AI (e.g., AR headsets, mixed reality platforms) or autonomous driving computer vision.
  • On-prem / self-hosted model deployment, including serving and optimizing open-weight models on customer hardware, or hands-on fine-tuning and deployment of vLLMs.
  • Exposure to industrial domains such as data centers, energy grids, aerospace, or manufacturing.
  • Experience with in-context grounding or RAG against a knowledge base, including tool and knowledge base wiring.
  • Master's degree with a vision or multimodal research component.

Tech Stack Includes: Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, LangChain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker.

COMPENSATION & BENEFITS

  • Salary: $180,000 – $240,000 USD annually
  • Early-stage equity as a founding team member
  • Visa sponsorship: Not available

LOCATION

On-site, 5 days/week in San Francisco, CA. This is not a remote or hybrid role.