Senior Site Reliability Engineer
Akuity
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
As a Senior Site Reliability Engineer (SRE) at Akuity, you’ll help keep the Akuity platform running at the reliability level enterprise customers expect. This is a high-ownership role where you don’t just respond to incidents—you help define and defend reliability across the platform, working closely with engineering, infrastructure, and product.
What You’ll Own
- Platform Reliability & SLAs
- Define and improve SLI/SLO/SLA targets for the Akuity SaaS platform
- Design, instrument, and maintain observability systems (metrics, logs, traces) across multi-region infrastructure
- Identify reliability gaps, lead blameless post-mortems, and drive permanent fixes
- Partner with engineering to build reliability into new features before production
- On-Call & Incident Response
- Participate in an on-call rotation and act as incident commander for high-severity events
- Maintain runbooks, escalation paths, and incident playbooks to reduce MTTR
- Improve alerting fidelity to reduce noise and eliminate toil
- Lead post-incident reviews with clear timelines, root cause analysis, and action item follow-through
Requirements
- 5+ years in SRE, platform engineering, or production operations in a SaaS environment
- Deep, hands-on Kubernetes expertise (scheduler, networking, storage, autoscaling)
- Strong AWS fundamentals:
- Compute: EC2, EKS
- Networking: VPC, NLB, Route53
- Storage: S3, RDS
- Security: IAM
- Proven ability to define and operate SLOs in production (including error budgets)
- Observability tooling proficiency (e.g., Prometheus, Grafana, OpenTelemetry, Datadog)
- Strong scripting/automation skills (Go, Python, Bash, or similar)
- Excellent written communication (runbooks, incident reports, post-mortems)
- Must live in US time zones (Pacific through Eastern) (including Canada/other regions)
Nice to Have
- Experience with Argo CD, Kargo, or GitOps delivery workflows
- Familiarity with multi-region, multi-cluster Kubernetes deployments
- Experience with compliance-adjacent infrastructure (SOC 2, ISO 27001, HIPAA, PCI DSS)
- Background operating infrastructure for other platform/developer tooling companies
About Akuity
Akuity builds Argo CD and a GitOps platform for enterprises to simplify Kubernetes-based software delivery. Their platform helps teams manage development and deployment across hundreds or thousands of Kubernetes clusters from a single control plane, improving reliability and speed for DevOps and Platform Engineering teams.
Scraped 7/3/2026