Site Reliability Engineer
Evlo AI
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
As a Site Reliability Engineer (SRE), you will own the reliability of critical production services—covering availability, latency, performance, efficiency, change management, and emergency response. You will work with software engineers to design, build, and operate resilient distributed systems that scale under high load.
Key Responsibilities
- Design, provision, and maintain cloud infrastructure using Terraform, Kubernetes, and AWS or GCP.
- Build robust CI/CD pipelines using GitHub Actions or ArgoCD for rapid, automated, and safe deployments.
- Establish observability frameworks with Prometheus, Grafana, and Datadog (metrics, logging, distributed tracing).
- Participate in on-call rotation to troubleshoot and resolve production incidents, including post-mortem analysis.
- Automate operational tasks and infrastructure provisioning to reduce toil and enforce reliability standards.
- Collaborate with engineering teams on capacity planning, architectural reviews, and load testing.
Requirements
- 4–7 years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering in high-scale environments.
- Deep hands-on experience with Kubernetes, containerization, and service mesh technologies in production.
- Strong infrastructure-as-code skills (e.g., Terraform, and/or CloudFormation or Pulumi).
- Proficiency in a scripting/programming language: Python, Go, or Bash.
- Solid understanding of networking fundamentals: TCP/IP, DNS, TLS, and load balancing architectures.
Bonus
- Experience with chaos engineering.
- Experience managing multi-region cloud deployments.
About Evlo AI
Evlo AI appears to develop AI-focused products and services that run in production environments requiring reliable distributed systems. The role centers on ensuring availability, performance, and safe operations of critical services, suggesting a software-driven, high-scale engineering organization.
Scraped 8/6/2026