Senior Machine Learning Systems Engineer
Cohere
full-remoteseniorpermanentbackenddata Full remote 73 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine LearningDistributed TrainingHPCCUDANCCLPerformance EngineeringData PipelinesDockerKubernetesJAX Internals
About the role
Role overview
You will build, maintain, and evolve the training framework for frontier-scale language models. The focus is on designing core training components that are fast, reliable, and scalable, with end-to-end ownership across critical parts of the training stack.
Key missions
- Design and maintain essential components enabling rapid, reliable, and scalable model training.
- Collaborate with infrastructure teams to ensure clusters, container environments, and hardware configurations support high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack (from compute to networking, IO, and data pipelines).
Requirements
- Strong engineering judgment around trade-offs (e.g., performance vs. complexity, research velocity vs. maintainability).
- Ability to debug performance issues across CUDA/NCCL, networking, IO, and ML data pipelines.
- Proven experience in large-scale distributed training or HPC systems.
- Comfort working on performance engineering, profiling, and low-level systems.
- Experience with containerized environments (Docker and Singularity/Apptainer).
- Strong collaboration skills across infrastructure, research, and deployment teams.
Nice to have
- Contributions to ML frameworks (e.g., PyTorch, JAX, DeepSpeed, Megatron, xFormers).
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
- Experience optimizing data pipelines (e.g., sharded datasets, caching strategies).
- Familiarity with evaluation/serving frameworks (e.g., vLLM, TensorRT-LLM, KV caches).
- Experience with multi-node cluster orchestration (e.g., Slurm, Ray, Kubernetes).
- A publication track record in top venues (NeurIPS, ICML, ICLR, etc.).
- Experience training LLMs and large transformer architectures.
Location
Full remote.
About Cohere
Cohere is an AI company focused on building frontier-scale language models. The company works on training and deploying large transformer systems, partnering closely across infrastructure and research to deliver high-performance ML capabilities.
Scraped 5/13/2026