Site Reliability Engineer
Our client's Cloud Operations team is expanding its SRE function. Site Reliability Engineers keep all user-facing services and production systems running smoothly. SREs here are a blend of pragmatic operators and software craftspeople who apply sound engineering principles, operational discipline and mature automation to the environment and the codebase. The team specialises in systems — networking, the Linux kernel, and scaling, algorithms and distributed systems.
js Collaborate and communicate asynchronously, and document so nothing is learned twice Have a go-for-it attitude: when you see something broken, you fix it Have experience with Nginx, HAProxy, Docker, Kubernetes, Terraform or similar technologies Projects you could work on Coding infrastructure automation with Ansible and Terraform Improving Prometheus monitoring or building new metrics Helping release managers deploy and fix new versions of application software Planning and executing the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS) Developing a relationship with a product group and defining their SRE KPIs — the SRE practice here is early in its journey Créé par