Senior Site Reliability Engineer
Employ is transforming how hiring gets done. With our three ATS solutions - Jobvite, Lever, and JazzHR - plus cutting-edge AI Companions, we free recruiters from admin work and give them more time for what matters: connecting with people. More than 26,000 global customers rely on Employ to make millions of candidate connections each year, helping them hire smarter, faster, and at scale. From growing startups to the world’s most recognized brands, we’re redefining what’s possible in talent acquisition.
We’re a fast-moving, remote-first team of builders and innovators who live our people-first philosophy every day. We back it up with flexible work scheduling and paid time off, comprehensive benefits, and career development opportunities that help you thrive. At Employ, you won’t just grow your career - you’ll help millions of others grow theirs.
Come join our team where we have each other’s backs, champion our customers, hold ourselves accountable, and shape what’s next in talent acquisition.
Candidate Safety Notice Please note, employ takes candidate safety seriously. Employ recruiters and employees communicate only via official @employinc.com email addresses. We will never request payment or personal financial information during the hiring process. If you receive suspicious outreach claiming to be from Employ, please report it to [email protected].
Job Title: Senior Site Reliability Engineer Location: Bangalore, India - Hybrid Job Type: Full-time Department: Engineering
Position Overview We are looking for a Senior Site Reliability Engineer (SRE) to join our team and play a key role in building highly reliable, scalable, and efficient systems. You will be instrumental in driving modern engineering practices that blend software engineering and infrastructure expertise. As an SRE, you’ll help maintain the health of our production environment, uphold service reliability standards. and implement tooling and automation that empowers engineering teams to move fast without compromising system stability.
You’ll work cross-functionally with developers and security teams to build observability, manage incidents, and proactively reduce toil through automation—creating systems that are not just available, but resilient and maintainable in an AI-native SDLC.
Key Responsibilities Use AI as a core tool in daily SRE work, embedding architectural context and change history directly into code and systems so they remain legible to both engineers and AI agents over time Deep dive into application codebases and directly contribute to improvements that drive the SRE mission, leaving code better than you found with each issue you tackle.
Participate in on-call rotations and command incident response with effective practices that improve time to recovery, preservation of evidence, and root cause analysis that prevents future occurrence and implements lessons learned into runbook improvements. Own the design and effectiveness of observability systems that empower all engineers with alerting and visibility into the applications we are supporting.
Develop and manage sustainable Infrastructure as Code (IaC) automation using tools such as Terraform, Ansible, or similar along with CI/CD and orchestrators such as Kubernetes and ArgoCD to bake SRE into the SDLC and eliminate toilsome operational work. Partner with roadmap delivery teams to implement and promote Site Reliability Engineering best practices within their workflows such as the systematic implementation of SLIs/SLOs, production readiness assessments, etc.
Collaborate with Security teams to ensure systems align with ISO 27001, SOC 2, and other compliance standards Stay current with industry trends and emerging technologies to continually improve our SRE capabilities
Minimum Qualifications 5+ years of experience in Site Reliability Engineering, Software Engineer or a similar role Proficiency in one or more programming/scripting languages such as Typescript, Java, Python, Go, PHP, or Ruby Strong experience with Unix/Linux systems administration and internals Solid understanding of system design, distributed computing, and SRE principles Expertise with containerization and orchestration technologies such as Docker and Kubernetes Experience with one or more cloud platforms: AWS, Azure, or Google Cloud Platform (GCP) Hands-on experience with any of the CI/CD tools such as GitHub Actions CI/CD, Argo CD etc..
Experience with relational databases like PostgreSQL, MySQL, or SQL Server Excellent troubleshooting, problem-solving, and analytical skills
Preferred Qualifications (Good to Have) Proficiency in scripting and automation using tools such as NodeJS Familiarity with NoSQL solutions such as MongoDB, Redis, DynamoDB Ability to monitor, optimize, and troubleshoot performance in large scale high-availability environments Experience with backup, replication, and data recovery strategies Excellent communication and collaboration abilities
Why Join Us
Continuous learning culture with opportunities to explore emerging technologies Will be engaged in developing and maintaining applications across multiple ecosystems, each with distinct architectural patterns and design considerations Collaborate with talented engineers who value innovation and ownership Flexible work environment with a focus on outcomes and autonomy