xelys jobs xelys jobs

AI Infrastructure & Platform Operations Engineer

Mirantis

full-remotemidpermanentdevopsbackend Full remote Yesterday via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

LinuxKubernetesNVIDIA GPUsInfiniBandPrometheusGrafanaELKOpenTelemetryObservabilityInfrastructure as Code (IaC)

About the role

Role Overview

As an AI Infrastructure & Platform Operations Engineer, you’ll join a dedicated team responsible for operating and supporting production AI infrastructure platforms. The work focuses on NVIDIA GPU-accelerated environments, high-performance networking, and Kubernetes-based platforms, ensuring reliability and continuous improvement of operations.

Key Missions

  • Monitor, operate, and support production AI infrastructure platforms, including NVIDIA GPU infrastructure and Kubernetes environments.
  • Investigate and resolve incidents across infrastructure, networking, hardware, and platform layers.
  • Collaborate with engineering teams and vendors to drive effective troubleshooting and resolution.
  • Improve monitoring and observability, automation, and operational processes.
  • Maintain operational documentation and contribute to operational runbooks and best practices.

Requirements

  • Strong Linux administration and troubleshooting skills.
  • Demonstrated analytical and problem-solving abilities.
  • Experience supporting production infrastructure and services.
  • Familiarity with structured operational and incident management processes.
  • Good understanding of networking concepts, with ability to diagnose infrastructure issues.
  • Excellent communication and collaboration skills.
  • 3+ years in infrastructure operations, platform operations, network operations, SRE, cloud operations, datacenter operations, or related roles.
  • Ability to work in a shift-based operational environment.
  • Working knowledge of Kubernetes in production.

Nice-to-Haves / Additional Expertise

  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Observability tools/platforms such as Grafana, Prometheus, ELK, OpenTelemetry.
  • Experience with AI infrastructure and/or HPC environments.
  • SRE/Platform Engineering background.
  • Infrastructure automation and Infrastructure-as-Code (IaC) practices.
  • Experience with large-scale distributed systems and production platforms.

About Mirantis

Mirantis is a technology company focused on enterprise cloud infrastructure and Kubernetes-based platforms. It helps organizations build, run, and operate large-scale cloud and container environments, including mission-critical infrastructure for modern workloads. The role here is centered on operating and improving production AI infrastructure ecosystems.

Scraped 7/26/2026