xelys jobs xelys jobs

Senior Site Reliability Engineer

Akuity

seniorpermanentdevops Sunnyvale, CA 83 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesAWSSRESLOsObservabilityPrometheusGrafanaOpenTelemetryGitOpsArgo CD

About the role

Role Overview

As a Senior Site Reliability Engineer (SRE) at Akuity, you’ll help keep the Akuity platform running at the reliability level enterprise customers expect. This is a high-ownership role where you don’t just respond to incidents—you help define and defend reliability across the platform, working closely with engineering, infrastructure, and product.

What You’ll Own

  • Platform Reliability & SLAs
    • Define and improve SLI/SLO/SLA targets for the Akuity SaaS platform
    • Design, instrument, and maintain observability systems (metrics, logs, traces) across multi-region infrastructure
    • Identify reliability gaps, lead blameless post-mortems, and drive permanent fixes
    • Partner with engineering to build reliability into new features before production
  • On-Call & Incident Response
    • Participate in an on-call rotation and act as incident commander for high-severity events
    • Maintain runbooks, escalation paths, and incident playbooks to reduce MTTR
    • Improve alerting fidelity to reduce noise and eliminate toil
    • Lead post-incident reviews with clear timelines, root cause analysis, and action item follow-through

Requirements

  • 5+ years in SRE, platform engineering, or production operations in a SaaS environment
  • Deep, hands-on Kubernetes expertise (scheduler, networking, storage, autoscaling)
  • Strong AWS fundamentals:
    • Compute: EC2, EKS
    • Networking: VPC, NLB, Route53
    • Storage: S3, RDS
    • Security: IAM
  • Proven ability to define and operate SLOs in production (including error budgets)
  • Observability tooling proficiency (e.g., Prometheus, Grafana, OpenTelemetry, Datadog)
  • Strong scripting/automation skills (Go, Python, Bash, or similar)
  • Excellent written communication (runbooks, incident reports, post-mortems)
  • Must live in US time zones (Pacific through Eastern) (including Canada/other regions)

Nice to Have

  • Experience with Argo CD, Kargo, or GitOps delivery workflows
  • Familiarity with multi-region, multi-cluster Kubernetes deployments
  • Experience with compliance-adjacent infrastructure (SOC 2, ISO 27001, HIPAA, PCI DSS)
  • Background operating infrastructure for other platform/developer tooling companies

About Akuity

Akuity builds Argo CD and a GitOps platform for enterprises to simplify Kubernetes-based software delivery. Their platform helps teams manage development and deployment across hundreds or thousands of Kubernetes clusters from a single control plane, improving reliability and speed for DevOps and Platform Engineering teams.

Scraped 7/3/2026