Senior Site Reliability Engineer
Doghouse Recruitment
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role overview
Senior/Staff Site Reliability Engineer to own production reliability end-to-end in a bare-metal Linux data center environment (100% remote in the US). You will define reliability targets, drive error-budget discussions, and ship changes that reduce incidents and improve latency (p95/p99).
Responsibilities
- Own production reliability end-to-end (from defining goals to shipping improvements)
- Define SLIs/SLOs and run error budget conversations
- Reduce incidents and improve latency (p95/p99)
- Build automation to eliminate toil
- Improve deployment safety with canary/rollback approaches
- Turn observability into actionable signal (reduce noise)
- Work close to the metal across:
- Kubernetes internals (scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane)
- Linux performance (CPU/memory/IO contention and low-level behaviors)
- Network debugging (DNS/TCP/TLS, latency, packet loss, congestion)
- Participate in on-call; success is measured by how much you reduce it
Requirements (must-have)
- Extensive production engineering experience running bare metal/on-prem/data center infrastructure (not public cloud only)
- Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO)
- Strong networking understanding: DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting (including underload)
- Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane
- Experience with Terraform, Docker, Helm, and modern CI/CD practices
- Strong coding skills (Go and/or Python) beyond automation scripting
- Experience in low-latency environments
Compensation
- Total compensation: OTE up to $350k (base + variable) depending on experience
About Doghouse Recruitment
The client is building a cloud platform for high-throughput, compute-heavy workloads, operating large-scale infrastructure where reliability must be engineered end-to-end. The environment is largely bare-metal/on-prem in data centers, with emphasis on Linux performance, networking troubleshooting, and Kubernetes internals.
Scraped 6/16/2026