Site Reliability Engineer (HPC)
Microsoft
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Site Reliability Engineer (HPC) for Microsoft’s High Performance Computing (HPC) infrastructure team within the Microsoft AI Superintelligence organization. You will combine software and systems engineering to keep large-scale, distributed AI infrastructure reliable, efficient, and highly available.
Responsibilities
- Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters used for MAI model training and inference.
- Observability: Design and maintain monitoring, alerting, and logging for HPC systems across GPUs, clusters, storage, and networking.
- Automation & Tooling: Automate deployments, incident response, scaling, and failover for CPU + GPU environments.
- Incident Management: Participate in and lead on-call rotations, troubleshoot production issues, run blameless postmortems, and drive continuous improvements.
- Security & Compliance: Ensure secure operations, data privacy, and compliance across model training and serving.
- Collaboration: Partner with ML engineers and platform teams to improve developer experience and speed research-to-production workflows.
Required Qualifications
- Master’s degree in Computer Science/IT (or related) plus 2+ years technical experience in SRE/DevOps/Infrastructure Engineering, OR
- Bachelor’s degree in Computer Science/IT (or related) plus 4+ years technical experience in SRE/DevOps/Infrastructure Engineering.
Additional Information (Work Location / Policy)
- Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days per week if living within 50 miles (U.S.) or 25 miles (non-U.S.) of that office (subject to local law).
About Microsoft
Microsoft is a global technology company building software, cloud infrastructure, and AI platforms for consumers and businesses. This role sits within Microsoft AI’s Superintelligence organization, focused on developing advanced, controllable, safety-aligned AI systems and enabling large-scale AI infrastructure to run reliably in production.
Scraped 6/15/2026