xelys jobs xelys jobs

Site Reliability Engineer

Scale.jobs

midpermanentdevops Atlanta, GA 24 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringAWSGCPKubernetesTerraformCI/CDObservabilityPrometheusGrafanaPython

About the role

Role overview

As a Site Reliability Engineer, you will build and maintain foundational infrastructure, CI/CD pipelines, and observability frameworks that enable high-throughput, low-latency microservices. You’ll partner with backend engineers to design resilient systems, automate operations, and prevent bottlenecks before they affect customers.

Responsibilities

  • Design, provision, and operate multi-region cloud infrastructure on AWS or GCP using Terraform (IaC)
  • Maintain and optimize Kubernetes clusters (EKS/GKE), including ingress, service meshes, and autoscaling policies
  • Build and support CI/CD pipelines using tools such as GitHub Actions, GitLab CI, or ArgoCD
  • Develop observability strategies using Prometheus, Grafana, Datadog, and/or OpenTelemetry
  • Participate in a blameless on-call rotation, lead incident response, perform root-cause analysis (RCA), and implement permanent fixes
  • Automate recurring operational tasks using Python, Go, or Bash to reduce toil and improve developer velocity

Requirements

  • 3–6 years experience in SRE, DevOps, or systems engineering managing production cloud infrastructure
  • Deep expertise in AWS or GCP and Kubernetes container orchestration
  • Strong proficiency with Terraform and configuration management/IaC
  • Solid scripting/system engineering skills in Go, Python, or Ruby
  • Experience configuring and tuning logging, metrics, and tracing for microservices

Bonus

  • Experience with service meshes (Istio/Linkerd)
  • Database administration at scale (PostgreSQL / NoSQL)
  • Security/compliance frameworks (SOC2 / ISO27001)

About Scale.jobs

Scale.jobs is hiring for infrastructure-focused engineering roles centered on running reliable, high-throughput microservices. The company’s work focuses on building production-grade cloud platforms with strong automation, CI/CD, and observability.

Scraped 7/2/2026