xelys jobs xelys jobs

Site Reliability Engineer

Empower

midpermanentdevopsbackend United States Today via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringKubernetesAWSEKSTerraformInfrastructure as CodeObservabilityCI/CDGitOpsIncident Management

About the role

Role Overview

The Site Reliability Engineer (SRE) ensures the reliability, scalability, and performance of Empower’s financial services platform. You’ll support production systems for millions of customers, improve operational excellence, and partner with development teams to maintain high availability and strong observability in a regulated environment.

Responsibilities

  • Own operational excellence for assigned systems and services; support cross-team projects
  • Participate in on-call rotations, respond to incidents, troubleshoot complex system/deployment issues, and drive resolution
  • Run postmortems, perform root cause analysis, and implement preventive measures
  • Define service level indicators (SLIs), build proactive monitoring and alerting, and manage observability for Kubernetes (including EKS)
  • Build, maintain, and optimize infrastructure as code across AWS environments
  • Manage/optimize EKS clusters for availability, resilience, and scalability of containerized applications
  • Partner with developers on releases and implement scalable/resilient services using GitOps and progressive delivery
  • Maintain and improve CI/CD pipelines and automation to reduce toil and improve operational efficiency
  • Perform capacity planning and right-sizing for performance and reliability
  • Document systems, runbooks, and architecture decisions; mentor entry-level SREs

Requirements

  • Bachelor’s degree in CS/IT or equivalent practical experience
  • 2–4 years in SRE, DevOps, or Systems Engineering
  • Hands-on production experience with high availability and resiliency on AWS, including EKS, EC2, RDS, S3, VPC
  • Production Kubernetes experience and containerization (e.g., Docker)
  • Infrastructure as code proficiency: Terraform and/or CloudFormation
  • Observability/incident detection and response experience
  • CI/CD principles; experience with GitLab CI, Jenkins, or equivalent
  • Networking fundamentals and high-availability architecture patterns
  • Familiarity with GitOps, incident management, and on-call practices

Nice to Have

  • Experience in financial services or other highly regulated industries
  • Familiarity with compliance frameworks such as SOC 2 or PCI DSS
  • Observability/APM tools such as Datadog, AppDynamics, New Relic, or similar
  • Strong programming skills in Shell, Go, Python, or similar
  • Experience supporting Java Spring Boot applications
  • Additional Kubernetes production experience (posting cuts off mid-sentence)

About Empower

Empower is a financial services company focused on helping customers achieve financial freedom. The role supports the reliability and scalability of Empower’s financial services platform in a regulated environment, emphasizing operational excellence and internal mobility.

Scraped 7/31/2026