Software Engineer 5 – Model Runtime, AI Platform
Netflix
seniorpermanentbackenddata United States 84 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
PyTorchDistributed TrainingFSDPCUDAGPU ProfilingReinforcement LearningDPOGRPOAWSInference Optimization
About the role
Role Overview
Software Engineer (Level 5) on Netflix’s Model Runtime team, owning systems that train, align, and serve critical ML models. The team is small and autonomous, building infrastructure at the intersection of systems engineering and ML.
Responsibilities
- Alignment & post-training infrastructure: design infrastructure for reinforcement learning and preference optimization, including GRPO, DPO, PPO, reward modeling, and preference optimization for recommendation model training.
- Next-gen GenAI workloads: build infrastructure for multimodal and diffusion models, including distributed training, disaggregated serving, and real-time / near-real-time / batch inference with asynchronous GPU pipelines.
- Scale distributed training: develop fault-tolerant training systems using FSDP and tensor/pipeline/context parallelism, plus mixed-precision across clusters of hundreds of GPUs.
- Full-stack optimization: profile and tune from PyTorch operators down to GPU kernels, improve utilization, and build cost models to guide infrastructure strategy.
- Evaluate emerging hardware/frameworks: assess specialized accelerators and the open-source ecosystem (including next-gen NVIDIA silicon) to maintain efficiency leadership.
- Promote operational best practices such as observability, logging, reporting, and on-call processes.
Requirements (Minimum)
- Experience in ML systems engineering, building infrastructure for training, fine-tuning, or inference of pre-LLM and post-LLM era models at scale.
- Strong systems programming skills across multiple layers of the stack, from ML frameworks to GPU kernels and memory management.
- Hands-on experience with PyTorch internals, large-scale distributed training, and system-model co-design.
- Ability to work through ambiguity and deliver 0-to-1 and 1-to-100 initiatives across business and technical domains.
- Operations & reliability mindset: observability/logging/on-call practices.
- Cloud experience, preferably AWS.
- Strong written and verbal communication across distributed time zones.
Nice to Have (Preferred)
- Deep experience with distributed training at scale (e.g., FSDP, parallelism, checkpointing) and/or LLM post-training (SFT, RLHF, DPO/GRPO).
- Inference optimization: vLLM, TensorRT, quantization, continuous batching, KV-cache management.
- GPU performance profiling/tuning: CUDA, NCCL, Nsight, PyTorch profiler.
- Experience with multimodal or diffusion model generation pipelines.
- Track record building reusable ML libraries or contributing to open-source ML.
About Netflix
Netflix is a global streaming entertainment company operating TV series, feature films, and games for over 300 million members across more than 190 countries. It uses machine learning and artificial intelligence to power personalization, content discovery, and other key business workflows, investing heavily in scalable ML infrastructure.
Scraped 7/4/2026