Senior AI Infrastructure & Platform Operations Engineer
Mirantis
full-remoteseniorpermanentdevopssecuritybackend Full remote - Barcelona, ES 55 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
KubernetesLinuxNetworkingIncident ManagementRoot Cause AnalysisSite Reliability Engineering (SRE)ObservabilityPrometheusOpenTelemetryInfrastructure as Code (IaC)
About the role
Role overview
Join Mirantis’ European AI Infrastructure & Platform Operations team as a Senior AI Infrastructure & Platform Operations Engineer. You will be a technical leader responsible for operational excellence in complex production environments, leading incident response, and collaborating with engineering and operations partners.
Key missions / responsibilities
- Lead investigation and resolution of complex incidents across infrastructure, networking, and platform layers.
- Act as a senior escalation point during critical, service-impacting events.
- Drive operational excellence: shape operational standards, reliability practices, and automation initiatives; help evolve AI-powered operational services.
- Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve technical challenges.
- Participate in major incident management and service restoration activities.
- Conduct operational readiness reviews and evaluate emerging technologies to improve service delivery.
- Mentor and support other engineers; provide technical leadership during investigations and improvements.
Requirements
- Production Kubernetes operations experience.
- 7+ years in infrastructure operations, platform operations, SRE, network operations, cloud operations, datacenter operations, or related technical roles.
- Proven ability to lead technical investigations and manage complex incidents.
- Strong networking expertise (diagnosing performance, connectivity, and reliability issues).
- Strong root cause analysis skills and ability to drive long-term operational improvements.
- Solid understanding of observability, monitoring, and service reliability practices.
- Expert-level Linux administration and troubleshooting.
- Experience supporting large-scale production infrastructure and distributed systems.
- Excellent communication, collaboration, and stakeholder management.
Nice-to-haves
- NVIDIA GPU infrastructure / accelerated computing platforms.
- InfiniBand networking and NVIDIA UFM.
- AI infrastructure and/or HPC environments.
- Platform Engineering and/or SRE experience.
- Large-scale Kubernetes operations.
- Infrastructure automation and Infrastructure-as-Code practices.
- Observability tools: Grafana, Prometheus, ELK, OpenTelemetry.
- Performance analysis/optimization for distributed infrastructure.
- Prior technical leadership, mentoring, or team lead responsibilities.
About Mirantis
Mirantis is a technology company focused on enterprise cloud and infrastructure solutions. The role is within their European AI Infrastructure & Platform Operations team, supporting production AI and platform services through reliability and operational excellence.
Scraped 7/30/2026