Infrastructure Site Reliability Engineer
Nebius
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Nebius is hiring an Infrastructure Site Reliability Engineer (Network SRE / NetSRE) to help build and run the network layer of its AI cloud platform. This is an engineering-first SRE role focused on defining reliability targets, automating improvements, and making the network safer to operate as Nebius scales.
Responsibilities
- Define and own network reliability goals for network services and critical paths using SLIs/SLOs, availability targets, and error budgets.
- Drive reliability improvements across the whole network, including:
- service reliability
- site readiness
- inter-site connectivity (DCI)
- operational standards
- Own incident response for your areas:
- lead investigations and postmortems
- convert failures into durable fixes (reduce repeat incidents)
- Build and evolve observability with actionable metrics/logs/traces, alerting, and faster debug loops.
- Design safer change workflows using automation, CI/CD, test/staging, canarying, rollbacks, and auditability for network changes.
- Partner with network engineers and platform teams to embed operability into designs and keep operations fast and practical.
Requirements
- Strong production Linux fundamentals and a structured approach to debugging complex systems.
- Solid networking fundamentals and understanding how real networks fail (e.g., control plane vs data plane, latency/loss, failure domains).
- Hands-on experience operating high-availability systems and improving them over time.
- Ability to write and maintain software/automation (Go is common; Python is welcome).
- Experience with modern infrastructure tooling such as IaC, CI/CD, and container platforms, plus comfort automating operational workflows.
Nice to Have
- High-throughput traffic processing experience (e.g., load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems).
- Low-level networking/performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals).
- Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection).
- Large-scale network observability/telemetry experience (routing/flow telemetry, regression detection at scale).
Compensation
- Base compensation range: $179,500–$224,300 USD (plus benefits).
About Nebius
Nebius builds a full-stack AI cloud infrastructure platform for the global AI economy, spanning compute, storage, networking, and inference optimization. The company supports developers and enterprises from data/model training through production deployment, enabling teams to avoid the cost and complexity of running large in-house AI/ML infrastructure.
Scraped 8/4/2026