xelys jobs xelys jobs

Engineering Manager, Machine Learning Platform

Jobgether

leadengineering-managementbackend United States Today via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine Learning InfrastructureDistributed SystemsEngineering ManagementModel TrainingModel ServingGPU InfrastructureLow-Latency SystemsTransformer WorkloadsData Quality & ReproducibilityPlatform Reliability

About the role

Role Overview

Lead the engineering team building and scaling infrastructure for advanced machine learning capabilities. You’ll blend people leadership with hands-on technical judgment to deliver reliable ML training, deployment, GPU infrastructure, and low-latency serving systems.

Responsibilities

  • Own and execute the roadmap for ML training and serving platforms (training systems, deployment workflows, GPU infrastructure, low-latency serving).
  • Lead, mentor, and develop platform engineers while staying engaged in technical strategy and implementation decisions.
  • Balance reliability, scalability, developer experience, performance, and infrastructure costs to drive operational excellence.
  • Build and improve platforms that help ML teams develop, deploy, and operate models efficiently.
  • Evaluate and adopt modern ML infrastructure technologies, including support for deep learning, transformer-based workloads, and large-scale compute.
  • Partner with ML, product, and infrastructure teams on critical AI initiatives.
  • Collaborate with senior engineers on architecture, trade-offs, and long-term platform strategy.
  • Recruit, retain, and grow high-performing engineering talent across career levels.
  • Establish best practices to improve platform quality, reliability, and engineering effectiveness.

Requirements

  • 7+ years of software and/or ML engineering experience, including 2+ years managing engineering teams.
  • Strong background in ML infrastructure, distributed systems, and team development.
  • Demonstrated success building production-grade platforms and guiding technical decisions through complex challenges.
  • Hands-on experience building and operating production ML platforms and/or distributed systems infrastructure.
  • Experience in one or more areas such as model training, model serving, deployment workflows, GPU infrastructure, or large-scale compute.
  • Solid understanding of ML data requirements (datasets, data quality, reproducibility, evaluation).
  • Familiarity with modern ML technologies: deep learning, transformer architectures, and large-scale model workloads.
  • Strong systems thinking and collaboration skills with senior engineers on architecture.
  • Proven ability to deliver platforms that improve engineering productivity and ML team impact.
  • Experience recruiting/coaching/developing engineers across multiple experience levels.

Nice-to-Haves

  • Not explicitly stated beyond the requirements.

Scraped 7/28/2026