xelys jobs xelys jobs

Infrastructure Site Reliability Engineer

Nebius

midpermanentdevopsbackend United States 5 days ago via LinkedIn
179,500 - 224,300 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)Network ReliabilitySLIs/SLOsIncident ResponsePostmortemsObservabilityCI/CDInfrastructure as Code (IaC)LinuxGo

About the role

Role Overview

Nebius is hiring an Infrastructure Site Reliability Engineer (Network SRE / NetSRE) to help build and run the network layer of its AI cloud platform. This is an engineering-first SRE role focused on defining reliability targets, automating improvements, and making the network safer to operate as Nebius scales.

Responsibilities

  • Define and own network reliability goals for network services and critical paths using SLIs/SLOs, availability targets, and error budgets.
  • Drive reliability improvements across the whole network, including:
    • service reliability
    • site readiness
    • inter-site connectivity (DCI)
    • operational standards
  • Own incident response for your areas:
    • lead investigations and postmortems
    • convert failures into durable fixes (reduce repeat incidents)
  • Build and evolve observability with actionable metrics/logs/traces, alerting, and faster debug loops.
  • Design safer change workflows using automation, CI/CD, test/staging, canarying, rollbacks, and auditability for network changes.
  • Partner with network engineers and platform teams to embed operability into designs and keep operations fast and practical.

Requirements

  • Strong production Linux fundamentals and a structured approach to debugging complex systems.
  • Solid networking fundamentals and understanding how real networks fail (e.g., control plane vs data plane, latency/loss, failure domains).
  • Hands-on experience operating high-availability systems and improving them over time.
  • Ability to write and maintain software/automation (Go is common; Python is welcome).
  • Experience with modern infrastructure tooling such as IaC, CI/CD, and container platforms, plus comfort automating operational workflows.

Nice to Have

  • High-throughput traffic processing experience (e.g., load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems).
  • Low-level networking/performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals).
  • Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection).
  • Large-scale network observability/telemetry experience (routing/flow telemetry, regression detection at scale).

Compensation

  • Base compensation range: $179,500–$224,300 USD (plus benefits).

About Nebius

Nebius builds a full-stack AI cloud infrastructure platform for the global AI economy, spanning compute, storage, networking, and inference optimization. The company supports developers and enterprises from data/model training through production deployment, enabling teams to avoid the cost and complexity of running large in-house AI/ML infrastructure.

Scraped 8/4/2026