xelys jobs xelys jobs

Site Reliability Engineer

Evlo AI

midpermanentdevops Minneapolis, MN Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringAWSKubernetesTerraformCloudFormationHelmPrometheusGrafanaDatadogCI/CDGitHub ActionsIAMObservabilityOn-callService MeshPythonGoBashRoot Cause Analysis

About the role

Role Overview

Own the availability, latency, performance, efficiency, and capacity management of core infrastructure used by millions of global users. You’ll build and maintain platform automation, CI/CD pipelines, and observability to help engineering teams ship rapidly and safely.

Responsibilities

  • Design, build, and maintain production infrastructure on AWS using Terraform, Kubernetes, and Helm charts
  • Manage and scale observability pipelines using Prometheus, Grafana, and Datadog for alerting and metrics
  • Automate deployment workflows and CI/CD pipelines with GitHub Actions to reduce downtime
  • Participate in an on-call rotation; troubleshoot incidents and perform thorough root cause analyses
  • Implement security hardening, including IAM policies and cloud compliance standards
  • Write infrastructure as code and internal tooling in Go or Python to reduce operational toil

Requirements

  • 3–6 years of experience in SRE, DevOps, or systems engineering in high-scale environments
  • Deep expertise in Kubernetes administration, containerization, and service mesh architectures
  • Strong infrastructure-as-code skills, specifically Terraform and CloudFormation
  • Solid programming skills in Python, Go, or Bash for automation and tooling
  • Extensive AWS experience (networking, IAM, and compute fundamentals)

Bonus

  • Experience migrating legacy systems to cloud-native architectures
  • Contributions to open-source infrastructure projects

About Evlo AI

Evlo AI builds AI-driven products and services that operate at scale for global users. The team focuses on reliability and performance of core infrastructure, including automation, CI/CD, and observability systems.

Scraped 8/2/2026