Engineering Manager, Machine Learning Platform
Jobgether
leadengineering-managementbackend United States Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine Learning InfrastructureDistributed SystemsEngineering ManagementModel TrainingModel ServingGPU InfrastructureLow-Latency SystemsTransformer WorkloadsData Quality & ReproducibilityPlatform Reliability
About the role
Role Overview
Lead the engineering team building and scaling infrastructure for advanced machine learning capabilities. You’ll blend people leadership with hands-on technical judgment to deliver reliable ML training, deployment, GPU infrastructure, and low-latency serving systems.
Responsibilities
- Own and execute the roadmap for ML training and serving platforms (training systems, deployment workflows, GPU infrastructure, low-latency serving).
- Lead, mentor, and develop platform engineers while staying engaged in technical strategy and implementation decisions.
- Balance reliability, scalability, developer experience, performance, and infrastructure costs to drive operational excellence.
- Build and improve platforms that help ML teams develop, deploy, and operate models efficiently.
- Evaluate and adopt modern ML infrastructure technologies, including support for deep learning, transformer-based workloads, and large-scale compute.
- Partner with ML, product, and infrastructure teams on critical AI initiatives.
- Collaborate with senior engineers on architecture, trade-offs, and long-term platform strategy.
- Recruit, retain, and grow high-performing engineering talent across career levels.
- Establish best practices to improve platform quality, reliability, and engineering effectiveness.
Requirements
- 7+ years of software and/or ML engineering experience, including 2+ years managing engineering teams.
- Strong background in ML infrastructure, distributed systems, and team development.
- Demonstrated success building production-grade platforms and guiding technical decisions through complex challenges.
- Hands-on experience building and operating production ML platforms and/or distributed systems infrastructure.
- Experience in one or more areas such as model training, model serving, deployment workflows, GPU infrastructure, or large-scale compute.
- Solid understanding of ML data requirements (datasets, data quality, reproducibility, evaluation).
- Familiarity with modern ML technologies: deep learning, transformer architectures, and large-scale model workloads.
- Strong systems thinking and collaboration skills with senior engineers on architecture.
- Proven ability to deliver platforms that improve engineering productivity and ML team impact.
- Experience recruiting/coaching/developing engineers across multiple experience levels.
Nice-to-Haves
- Not explicitly stated beyond the requirements.
Scraped 7/28/2026