xelys jobs xelys jobs

Site Reliability Engineer II

RemoteHunter

full-remotemidpermanentbackenddevops United States 175 days ago via LinkedIn
95,000 - 171,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSREPythonGoKubernetesLinuxPrometheusGrafanaCI/CDTerraform

About the role

Role Overview

Site Reliability Engineer II focused on automating, monitoring, and maintaining the reliability of AI inference workloads on a cloud platform. The role aims to reduce operational toil, improve stability, and support continuous deployment while ensuring AI application availability and performance.

Responsibilities

  • Build and maintain dashboards, alerts, and monitoring for inference workloads using the existing observability stack
  • Develop reliability automation and tooling in Python or Go to reduce manual effort
  • Create and improve runbooks for inference-specific operations
  • Support SLO tracking and reporting to identify reliability trends and improvement opportunities
  • Maintain CI/CD pipelines, deployment safety checks, and rollback processes
  • Collaborate with product engineering teams to troubleshoot complex issues across the stack
  • Participate in on-call rotations, respond to production incidents, and run blameless post-mortems

Requirements

  • 2+ years of Site Reliability Engineering experience and a Bachelor’s degree (or equivalent)
  • Proficiency in Python or Go for automation scripting
  • Experience with Linux systems administration and infrastructure troubleshooting
  • Familiarity with Kubernetes and containerization
  • Experience with observability tools such as Prometheus or Grafana
  • Exposure to CI/CD pipelines and infrastructure-as-code tools like Terraform or SaltStack
  • Curiosity and willingness to learn about AI infrastructure and distributed systems

Compensation & Benefits

  • US base salary range: $95,000–$171,000 per year (depending on experience, skills, certifications, and location)
  • Additional incentives may include annual bonuses, equity awards, and an ESPP
  • Healthcare, 401K, holidays and PTO, sick leave, parental leave, and employee assistance program
  • Flexible work arrangements: remote or office within the advertised country

About RemoteHunter

The employer operates in cloud computing and AI infrastructure, helping customers run AI inference models and enabling developers to build AI applications. It designs, implements, deploys, and operates AI platforms using scalable serverless inference workloads, GPU infrastructure, and Kubernetes to deliver reliable AI services at scale.

Scraped 4/1/2026