xelys jobs xelys jobs

Sr Site Reliability Engineer

Commence

seniorpermanentbackend Virginia, United States 141 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringAWSKubernetesObservabilitySLOs/SLIsIncident ResponseTerraformCI/CDHIPAASOC 2

About the role

Sr Site Reliability Engineer (SRE)

Commence is looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational health of its mission-critical healthcare data platform.

Responsibilities

  • Own reliability across the lifecycle: embed reliability as a first-class concern from architecture through deployment.
  • Design and operate observability: implement metrics, logging, tracing, and alerting for distributed systems.
  • Define reliability targets: establish and enforce SLOs/SLIs and error budgets with product and engineering teams.
  • Lead incident response: triage, coordinate remediation, run blameless post-mortems, and drive systemic fixes.
  • Build CI/CD pipelines to enable rapid and safe production releases.
  • Partner on infrastructure changes and contribute to existing Infrastructure-as-Code (Terraform or CloudFormation).
  • Design and operate resilient systems: highly available, fault-tolerant architecture with auto-scaling, failover, and disaster recovery.
  • Reduce operational toil via automation and elimination of manual processes.
  • Review architectures for operational risk and help establish reliability-first design patterns.
  • Operate container orchestration at scale (Kubernetes/container platforms).
  • Ensure compliance and security requirements for healthcare data (e.g., HIPAA, SOC 2).
  • Mentor engineers on reliability practices.
  • Participate in on-call rotation with a continuous focus on reducing the need for it.

Requirements

  • 7+ years in SRE, platform engineering, or DevOps.
  • Proven ability to solve complex, high-stakes system failures under pressure.
  • Deep hands-on experience with AWS services: EC2, EKS/ECS, Lambda, RDS, S3, CloudWatch.
  • Familiar with Terraform or CloudFormation to contribute to existing IaC.
  • Experience designing/operating distributed systems with strict availability/latency requirements.
  • Proficiency in at least one scripting/systems language: Python, Go, Bash (or similar).
  • Production experience with container orchestration (Kubernetes, ECS).
  • Observability tooling experience: OpenSearch, Prometheus/Grafana (or equivalents).
  • CI/CD experience with platforms such as GitHub Actions, Jenkins, CircleCI.
  • Ability to define and operationalize SLOs and error budgets.
  • Experience with relational and NoSQL databases (performance tuning, replication, backup strategies).
  • Strong networking fundamentals (DNS, load balancing, VPCs, TLS).
  • Excellent communication skills to translate technical risk into business impact.

Nice-to-have / Additional Requirements

  • AWS Certifications (exact requirements appear cut off in the provided text).

About Commence

Commence is a healthcare-focused technology company building data-centric solutions to improve health outcomes. It combines clinical expertise with quality, data-driven technology to support more efficient, value-based care through mission-critical healthcare data platforms.

Scraped 5/12/2026