Lead Site Reliability Engineer
We're a remote-first company with a team spanning ten countries.
Lead Site Reliability Engineer to lead the team ensuring AppSignal runs reliably 24/7, with a player-coach role owning reliability and security strategy across a bare-metal infrastructure managed with Ansible, data stack in Rust/Kafka, and Rails UI.
Lead Site Reliability Engineer We're looking for a Lead Site Reliability Engineer (SRE) to lead the team that's responsible for making AppSignal run reliably and securely 24/7. Full-time Remote, within 2 hours of CET About AppSignal AppSignal helps thousands of teams in 60+ countries monitor and improve the performance of their applications. We're a remote-first company with a team spanning multiple countries, built on values like impact, transparency, and continuous improvement.
We build developer tools that engineers actually enjoy using - with a strong focus on clarity, reliability, and customer experience. We're looking for somebody within 2-4 hours of the CET time zone for this role.
About the role AppSignal is looking for a Lead Site Reliability Engineer (SRE) to lead the team that's responsible for making AppSignal run reliably 24/7. This is a player-coach role: you'll set the technical direction and own the reliability and security strategy for our platform, while staying hands-on with the systems your team runs. AppSignal's infrastructure was developed by a very small team. We value resilience and peace of mind highly, so there's a long-standing tradition of fixing underlying issues that lead to us getting alerted.
As we enter the next phase of our business's growth, we're expanding this team, and we're looking for the person to lead it. Our stack We currently run a mostly bare-metal infrastructure, managed by Ansible We have a data ingestion and processing stack that's written in Rust, running on top of Kafka There's a Rails app that serves the UI our customers interact with We store data in MongoDB, Clickhouse and ElasticSearch As the Lead Site Reliability Engineer, you'll own the reliability and scalability of this system end-to-end.
You'll get to work on all the layers; there's no black box underneath. You'll shape the long-term infrastructure roadmap and security approach, decide where we invest, and grow the team that gets us there. You'll also stay close to the work: you'll be part of the on-call rotation, and we do a lot of our infrastructure tuning in our Rust codebase. How we work We run infrastructure by a written operating model, grounded in established SRE practice. These are the three principles we strive for: - Make changes boringly.
Every production change is safe, visible, and reversible: the repo is the source of truth, staging absorbs mistakes before customers feel them, and rollouts are progressive. Our tooling announces the work, checks for drift, and records what ran so the safe path is also the easy path. - Talk while you work. When something breaks, we switch into incident mode explicitly a lightweight take on the Incident Command System, with one named coordinator, one written source of truth, and narrated state changes. We mitigate first and diagnose after.
In high-stakes moments, communication is the work. - Every miss teaches. Every customer-facing incident and every near-miss ends in a blameless postmortem with owned, tracked action items. A manual fix needed twice becomes automation; recurring toil is a bug. As the lead, you'll be the keeper of this model: you'll uphold it under pressure, notice when we're drifting from it, and evolve it as the team grows.
Your responsibilities Lead the SRE team: set priorities, unblock the team, and grow each engineer through mentoring and feedback Own the reliability strategy and long-term infrastructure roadmap to get us to the next level of scalability Be part of the on-call rotation, and continuously improve the rotation itself so that we can all have amazing nights' rest Act as incident coordinator when it counts: keep communication flowing, mitigate first, and stand incidents down explicitly Lead blameless postmortems for incidents and near-misses, and make sure every action item has an owner and root causes get fixed, not worked around Guide strategic projects, such as building new infrastructure in AWS Stay hands-on: optimize and tune our Rust-based processing code, and make impactful infrastructure automation improvements Help hire as the team expands, and work with leadership on capacity planning and infrastructure budgets Work cross-functionally to respond to security report requests, ISO certification renewals Coordinate penetration tests, document results, and coordinate development and deployment of any responses Handle incoming reports by security researchers.
Assess impact and get fixes shipped promptly. What you bring (and what helps you thrive) We are looking for candidates with 8+ years of experience who...
Have proven experience with keeping large systems running reliably, particularly in Linux environments Have led an SRE, platform, or infrastructure team, either as a formal manager or as a technical lead Are both detail-oriented and interested in the big picture of a complex system Have experience working with development teams to help influence product development approaches, including database performance tuning, refactoring needs, and identifying alternatives or limitations given current architecture Are a competent software developer, with experience in multiple languages/frameworks Have experience with Rust, either personally or professionally Have experience using Ansible, and/or similar configuration management systems Have experience defining and running incident response and post-mortem processes Can translate business growth into an infrastructure strategy, and communicate it clearly to both engineers and leadership Get energized by solving problems and working collaboratively with a thoughtful, low-ego team Are proactive, organized, and comfortable managing your own schedule in a remote environment Are comfortable being in an on-call rotation Bonus: Have corporate experience with AWS, professional experience deploying and supporting Kubernetes and Docker containers What we offer Competitive salary tailored to your experience and location Remote-first work culture with support for co-working if needed Eligibility to participate in employee stock option program Flexible and generous PTO (Paid Time Off) policy Personal development budget for books, courses, or conferences Who we are We're a team of kind, curious people from different backgrounds, each bringing unique strengths (and yes, a few quirks too).
We'd love for you to add yours. We welcome candidates of all backgrounds, genders, orientations, ethnicities, ages, and abilities. If you're looking for a place to do your best work and know your contributions are valued, you'll feel right at home here. How to apply Do you want to join our team as our new Lead Site Reliability Engineer? Then we'd love to hear about you! Apply now AppSignal helps thousands of teams in 60+ countries monitor their web apps. We're a remote-first company with a team spanning ten countries. Our website Hiring with Homerun Pri