AI Infrastructure & Platform Operations Engineer
Mirantis
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
As an AI Infrastructure & Platform Operations Engineer, you’ll join a dedicated team responsible for operating and supporting production AI infrastructure platforms. The work focuses on NVIDIA GPU-accelerated environments, high-performance networking, and Kubernetes-based platforms, ensuring reliability and continuous improvement of operations.
Key Missions
- Monitor, operate, and support production AI infrastructure platforms, including NVIDIA GPU infrastructure and Kubernetes environments.
- Investigate and resolve incidents across infrastructure, networking, hardware, and platform layers.
- Collaborate with engineering teams and vendors to drive effective troubleshooting and resolution.
- Improve monitoring and observability, automation, and operational processes.
- Maintain operational documentation and contribute to operational runbooks and best practices.
Requirements
- Strong Linux administration and troubleshooting skills.
- Demonstrated analytical and problem-solving abilities.
- Experience supporting production infrastructure and services.
- Familiarity with structured operational and incident management processes.
- Good understanding of networking concepts, with ability to diagnose infrastructure issues.
- Excellent communication and collaboration skills.
- 3+ years in infrastructure operations, platform operations, network operations, SRE, cloud operations, datacenter operations, or related roles.
- Ability to work in a shift-based operational environment.
- Working knowledge of Kubernetes in production.
Nice-to-Haves / Additional Expertise
- NVIDIA GPU infrastructure and accelerated computing platforms.
- InfiniBand networking and NVIDIA UFM.
- Observability tools/platforms such as Grafana, Prometheus, ELK, OpenTelemetry.
- Experience with AI infrastructure and/or HPC environments.
- SRE/Platform Engineering background.
- Infrastructure automation and Infrastructure-as-Code (IaC) practices.
- Experience with large-scale distributed systems and production platforms.
About Mirantis
Mirantis is a technology company focused on enterprise cloud infrastructure and Kubernetes-based platforms. It helps organizations build, run, and operate large-scale cloud and container environments, including mission-critical infrastructure for modern workloads. The role here is centered on operating and improving production AI infrastructure ecosystems.
Scraped 7/26/2026