Site Reliability Engineer
Scale.jobs
midpermanentdevops Atlanta, GA 24 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability EngineeringAWSGCPKubernetesTerraformCI/CDObservabilityPrometheusGrafanaPython
About the role
Role overview
As a Site Reliability Engineer, you will build and maintain foundational infrastructure, CI/CD pipelines, and observability frameworks that enable high-throughput, low-latency microservices. You’ll partner with backend engineers to design resilient systems, automate operations, and prevent bottlenecks before they affect customers.
Responsibilities
- Design, provision, and operate multi-region cloud infrastructure on AWS or GCP using Terraform (IaC)
- Maintain and optimize Kubernetes clusters (EKS/GKE), including ingress, service meshes, and autoscaling policies
- Build and support CI/CD pipelines using tools such as GitHub Actions, GitLab CI, or ArgoCD
- Develop observability strategies using Prometheus, Grafana, Datadog, and/or OpenTelemetry
- Participate in a blameless on-call rotation, lead incident response, perform root-cause analysis (RCA), and implement permanent fixes
- Automate recurring operational tasks using Python, Go, or Bash to reduce toil and improve developer velocity
Requirements
- 3–6 years experience in SRE, DevOps, or systems engineering managing production cloud infrastructure
- Deep expertise in AWS or GCP and Kubernetes container orchestration
- Strong proficiency with Terraform and configuration management/IaC
- Solid scripting/system engineering skills in Go, Python, or Ruby
- Experience configuring and tuning logging, metrics, and tracing for microservices
Bonus
- Experience with service meshes (Istio/Linkerd)
- Database administration at scale (PostgreSQL / NoSQL)
- Security/compliance frameworks (SOC2 / ISO27001)
About Scale.jobs
Scale.jobs is hiring for infrastructure-focused engineering roles centered on running reliable, high-throughput microservices. The company’s work focuses on building production-grade cloud platforms with strong automation, CI/CD, and observability.
Scraped 7/2/2026