Live roles / trener-robotics
Compatibility brief · Discovered by NoBoards 9h ago

Senior Cloud Infrastructure Engineer

trener-robotics · San Jose (US-HQ)
Sponsorship not statedMode not statedIndustry: robotics
Extracted role summary · not eligibility evidence

Senior Cloud Infrastructure Engineer at Trener, a robotics software company, responsible for multi-cloud infrastructure (GCP and AWS), including networking, infrastructure-as-code, Kubernetes, and security controls.

GCPAWSTerraformOpenTofuKubernetesIAMVPCDNS

About Trener

Trener is building the software foundation for a new generation of industrial robotics.

We combine advanced AI, intuitive programming, and pre-trained skill models to make robots easier to deploy,

operate, and adapt to real-world industrial work. With teams in San Jose and Trondheim, we build software

that connects cloud infrastructure, developer platforms, and deployed robotic systems.

About the Role

We run a production cloud platform on Google Cloud today and are expanding our AWS footprint. We need a senior infrastructure engineer to help establish a deliberate multi-cloud foundation and define how GCP and AWS should coexist as the platform grows.

You will own the Cloud Foundations layer: cloud account and project structure, infrastructure-as-code, state strategy, networking, Kubernetes foundations, and implementation of cloud access controls. You will help decide where Trener should standardize across providers and where provider-specific architecture is the right boundary.

This is a senior individual contributor role. You will not be managing people, and you will not be inheriting a greenfield.

What You Will Own

  • Cloud foundations across GCP and AWS - project and account topology, landing zones, organization policies, baseline guardrails, and enrollment of new environments.
  • Cloud and cross-cloud networking - VPCs, routing, peering, DNS, load balancing, firewall policy, IP addressing, and private connectivity between cloud environments.
  • Infrastructure as code - reusable Terraform/OpenTofu modules, environment composition, lifecycle management, and patterns that keep infrastructure understandable as the estate grows.
  • State strategy and blast-radius boundaries - how infrastructure state is partitioned, how dependencies are expressed, and how those patterns evolve across providers.
  • Cloud identity implementation - implement cloud IAM, workload identity, federation, and access controls in partnership with Security.
  • Kubernetes foundations - cluster lifecycle, baseline configuration, upgrades, and shared infrastructure services beneath application workloads.
  • Infrastructure cost visibility - implement the tagging, labeling, budgets, alerts, and reporting needed to make cloud spend understandable and surface obvious infrastructure waste.
  • Infrastructure controls supporting our SOC 2 program - resource labeling, access boundaries, configuration standards, and evidence-producing infrastructure practices.

Requirements

  • Proven experience operating production infrastructure as a cloud, infrastructure, platform, or site reliability engineer.
  • Terraform or OpenTofu at module-authoring depth - writing and versioning reusable modules, managing state across environments, and handling the lifecycle of real infrastructure over time.
  • Cloud foundation design experience - multi-account, multi-project, landing-zone, guardrail, IAM, or network architecture beyond a handful of isolated workloads.
  • Experience with infrastructure composition or orchestration - Atmos, Terragrunt, Terraspace, or an equivalent approach to managing reusable infrastructure across environments.
  • Strong networking depth - VPCs, routing, peering, DNS, TLS, load balancing, firewall policy, IP addressing, and connectivity troubleshooting.
  • Kubernetes working knowledge - cluster operations, RBAC, networking, shared services, and troubleshooting workloads that will not start or communicate correctly.
  • Helm chart authoring, not just chart installation.
  • Production cloud experience across at least two major providers, or deep experience with one plus substantial hands- on involvement extending an organization into another.
  • Python and Bash for automation, CLIs, diagnostics, and glue.
  • Linux fluency.
  • Experience with controlled infrastructure environments - SOC 2, HIPAA, HITRUST, FedRAMP, PCI, or similar security/compliance expectations.
  • Fluency in English and strong written communication. Architectural decisions here are expected to be documented and reviewed in writing

Experience That Will Help You Succeed

  • GitOps workflows across multiple environments and comfort operating infrastructure through reviewed, auditable changes.
  • Externalized secrets management and Kubernetes-native secret delivery patterns.
  • Operating Prometheus/Grafana/Loki or comparable observability tooling as an infrastructure consumer and operator.
  • Experience introducing or formalizing a second cloud, including a clear account of what you would repeat and what you would change.
  • Cloud identity federation and workload identity across providers.
  • Disaster-recovery, regional resilience, or high-availability infrastructure work.
  • Private cloud networking or overlay connectivity.

What We Offer

The opportunity to work at the intersection of robotics, AI, and cloud infrastructure. A senior technical role with meaningful influence over the foundations on which the company builds.

A growing international engineering organization with teams in the United States and Norway