Site Reliability Engineer II
RemoteHunter
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Site Reliability Engineer II focused on automating, monitoring, and maintaining the reliability of AI inference workloads on a cloud platform. The role aims to reduce operational toil, improve stability, and support continuous deployment while ensuring AI application availability and performance.
Responsibilities
- Build and maintain dashboards, alerts, and monitoring for inference workloads using the existing observability stack
- Develop reliability automation and tooling in Python or Go to reduce manual effort
- Create and improve runbooks for inference-specific operations
- Support SLO tracking and reporting to identify reliability trends and improvement opportunities
- Maintain CI/CD pipelines, deployment safety checks, and rollback processes
- Collaborate with product engineering teams to troubleshoot complex issues across the stack
- Participate in on-call rotations, respond to production incidents, and run blameless post-mortems
Requirements
- 2+ years of Site Reliability Engineering experience and a Bachelor’s degree (or equivalent)
- Proficiency in Python or Go for automation scripting
- Experience with Linux systems administration and infrastructure troubleshooting
- Familiarity with Kubernetes and containerization
- Experience with observability tools such as Prometheus or Grafana
- Exposure to CI/CD pipelines and infrastructure-as-code tools like Terraform or SaltStack
- Curiosity and willingness to learn about AI infrastructure and distributed systems
Compensation & Benefits
- US base salary range: $95,000–$171,000 per year (depending on experience, skills, certifications, and location)
- Additional incentives may include annual bonuses, equity awards, and an ESPP
- Healthcare, 401K, holidays and PTO, sick leave, parental leave, and employee assistance program
- Flexible work arrangements: remote or office within the advertised country
About RemoteHunter
The employer operates in cloud computing and AI infrastructure, helping customers run AI inference models and enabling developers to build AI applications. It designs, implements, deploys, and operates AI platforms using scalable serverless inference workloads, GPU infrastructure, and Kubernetes to deliver reliable AI services at scale.
Scraped 4/1/2026