xelys jobs xelys jobs

Site Reliability Engineer

Evlo AI

midpermanentdevopsbackend Minneapolis, MN 2 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)TerraformAWSGCPKubernetesEKSDockerPrometheusGrafanaCI/CD

About the role

Role Overview

Site Reliability Engineer responsible for the reliability, scalability, and performance of critical cloud-native production platforms. You will build automation to reduce operational toil, define and monitor SLOs/SLIs, and architect resilient systems in collaboration with product development teams.

Responsibilities

  • Design, provision, and maintain secure, multi-region cloud infrastructure using Terraform and cloud services (AWS-heavy)
  • Build and optimize container orchestration using Kubernetes (EKS), including ingress controllers, service meshes, and autoscaling policies
  • Establish observability pipelines with Prometheus, Grafana, Jaeger, and the ELK stack to monitor health and latency
  • Lead incident response and run blameless post-mortems; implement durable mitigations for systemic failures
  • Develop and maintain automated CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins for continuous, zero-downtime deployments
  • Automate operational tasks (toil) by writing infrastructure tooling in Go, Python, or Bash

Requirements

  • 3–6 years in SRE/DevOps/systems engineering managing high-traffic production environments
  • Strong Infrastructure as Code (IaC) experience, especially Terraform, plus cloud platforms like AWS or GCP
  • Deep expertise in containerization and production-grade Kubernetes administration (Docker/Kubernetes)
  • Proficiency in at least one development language for automation, preferably Go, Python, or Ruby
  • Solid networking knowledge (DNS, TCP/IP, HTTP/S, VPCs, load balancing) and Linux operating system internals

Nice to Have

  • Service mesh experience (Istio/Linkerd)
  • GitOps workflows (e.g., ArgoCD)
  • Multi-cloud or cloud certifications (AWS Certified Solutions Architect, CKA)

About Evlo AI

Evlo AI builds and operates cloud-native platforms that require high reliability, scalability, and performance. The role focuses on ensuring production systems can handle large-scale transaction volumes through modern DevOps and SRE practices.

Scraped 7/23/2026