xelys jobs xelys jobs

Senior AI Infrastructure & Platform Operations Engineer

Mirantis

full-remoteseniorpermanentdevopssecuritybackend Full remote - Barcelona, ES 55 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesLinuxNetworkingIncident ManagementRoot Cause AnalysisSite Reliability Engineering (SRE)ObservabilityPrometheusOpenTelemetryInfrastructure as Code (IaC)

About the role

Role overview

Join Mirantis’ European AI Infrastructure & Platform Operations team as a Senior AI Infrastructure & Platform Operations Engineer. You will be a technical leader responsible for operational excellence in complex production environments, leading incident response, and collaborating with engineering and operations partners.

Key missions / responsibilities

  • Lead investigation and resolution of complex incidents across infrastructure, networking, and platform layers.
  • Act as a senior escalation point during critical, service-impacting events.
  • Drive operational excellence: shape operational standards, reliability practices, and automation initiatives; help evolve AI-powered operational services.
  • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve technical challenges.
  • Participate in major incident management and service restoration activities.
  • Conduct operational readiness reviews and evaluate emerging technologies to improve service delivery.
  • Mentor and support other engineers; provide technical leadership during investigations and improvements.

Requirements

  • Production Kubernetes operations experience.
  • 7+ years in infrastructure operations, platform operations, SRE, network operations, cloud operations, datacenter operations, or related technical roles.
  • Proven ability to lead technical investigations and manage complex incidents.
  • Strong networking expertise (diagnosing performance, connectivity, and reliability issues).
  • Strong root cause analysis skills and ability to drive long-term operational improvements.
  • Solid understanding of observability, monitoring, and service reliability practices.
  • Expert-level Linux administration and troubleshooting.
  • Experience supporting large-scale production infrastructure and distributed systems.
  • Excellent communication, collaboration, and stakeholder management.

Nice-to-haves

  • NVIDIA GPU infrastructure / accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • AI infrastructure and/or HPC environments.
  • Platform Engineering and/or SRE experience.
  • Large-scale Kubernetes operations.
  • Infrastructure automation and Infrastructure-as-Code practices.
  • Observability tools: Grafana, Prometheus, ELK, OpenTelemetry.
  • Performance analysis/optimization for distributed infrastructure.
  • Prior technical leadership, mentoring, or team lead responsibilities.

About Mirantis

Mirantis is a technology company focused on enterprise cloud and infrastructure solutions. The role is within their European AI Infrastructure & Platform Operations team, supporting production AI and platform services through reliability and operational excellence.

Scraped 7/30/2026