Site Reliability Engineer
Evlo AI
midpermanentdevopsbackend Austin, TX Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)AWSTerraformKubernetesPrometheusGrafanaDatadogCI/CDIncident ResponseResilience Engineering
About the role
Role Overview
Site Reliability Engineer responsible for maintaining high availability, low latency, and robust security for large-scale distributed systems serving millions of active users. You will partner with software engineers to automate infrastructure provisioning, strengthen observability, and lead incident response.
Responsibilities
- Design and manage AWS cloud infrastructure using Terraform and Kubernetes; ensure high availability and zero-downtime deployments
- Build and maintain monitoring, logging, and alerting with Prometheus, Grafana, and Datadog
- Automate deployment and release processes using CI/CD tools such as GitHub Actions and Argo CD
- Lead incident response, perform post-mortems, and implement preventative measures
- Optimize infrastructure costs, resource utilization, and database performance across production environments
- Participate in an on-call rotation and collaborate on production readiness for new services
Requirements
- 3–6 years experience in SRE, DevOps, or systems administration in cloud-native environments
- Strong Linux systems administration skills
- Networking fundamentals: TCP/IP, DNS, TLS
- Containerization: Docker and Kubernetes
- Infrastructure as Code: Terraform (and CloudFormation experience)
- Configuration management tools (unspecified)
- Scripting/programming proficiency in Python, Go, or Bash
- Solid understanding of distributed systems, microservices, and resilience engineering
Bonus
- Bachelor’s degree in Computer Science
- CKA (Certified Kubernetes Administrator)
- Experience with service mesh (e.g., Istio)
Scraped 8/1/2026