xelys jobs xelys jobs

Senior Site Reliability Engineer

Doghouse Recruitment

full-remoteseniorpermanentdevopsbackend United States 52 days ago via LinkedIn
350,000 - 350,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSRELinuxKubernetesTerraformDockerHelmCI/CDNetworkingLow Latency

About the role

Role overview

Senior/Staff Site Reliability Engineer to own production reliability end-to-end in a bare-metal Linux data center environment (100% remote in the US). You will define reliability targets, drive error-budget discussions, and ship changes that reduce incidents and improve latency (p95/p99).

Responsibilities

  • Own production reliability end-to-end (from defining goals to shipping improvements)
  • Define SLIs/SLOs and run error budget conversations
  • Reduce incidents and improve latency (p95/p99)
  • Build automation to eliminate toil
  • Improve deployment safety with canary/rollback approaches
  • Turn observability into actionable signal (reduce noise)
  • Work close to the metal across:
    • Kubernetes internals (scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane)
    • Linux performance (CPU/memory/IO contention and low-level behaviors)
    • Network debugging (DNS/TCP/TLS, latency, packet loss, congestion)
  • Participate in on-call; success is measured by how much you reduce it

Requirements (must-have)

  • Extensive production engineering experience running bare metal/on-prem/data center infrastructure (not public cloud only)
  • Deep hands-on expertise in Linux systems debugging and performance (CPU, memory, IO)
  • Strong networking understanding: DNS/TCP/TLS, latency, packet loss, congestion, troubleshooting (including underload)
  • Strong Kubernetes experience beyond manifests: scheduler behavior, autoscaling edge cases, kubelet pressure/evictions, etcd/control plane
  • Experience with Terraform, Docker, Helm, and modern CI/CD practices
  • Strong coding skills (Go and/or Python) beyond automation scripting
  • Experience in low-latency environments

Compensation

  • Total compensation: OTE up to $350k (base + variable) depending on experience

About Doghouse Recruitment

The client is building a cloud platform for high-throughput, compute-heavy workloads, operating large-scale infrastructure where reliability must be engineered end-to-end. The environment is largely bare-metal/on-prem in data centers, with emphasis on Linux performance, networking troubleshooting, and Kubernetes internals.

Scraped 6/16/2026