xelys jobs xelys jobs

Senior Machine Learning Systems Engineer

Cohere

full-remoteseniorpermanentbackenddata Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine LearningDistributed TrainingHPCCUDANCCLPerformance EngineeringData PipelinesDockerKubernetesJAX Internals

About the role

Role overview

You will build, maintain, and evolve the training framework for frontier-scale language models. The focus is on designing core training components that are fast, reliable, and scalable, with end-to-end ownership across critical parts of the training stack.

Key missions

  • Design and maintain essential components enabling rapid, reliable, and scalable model training.
  • Collaborate with infrastructure teams to ensure clusters, container environments, and hardware configurations support high-performance training.
  • Investigate and resolve performance bottlenecks across the ML systems stack (from compute to networking, IO, and data pipelines).

Requirements

  • Strong engineering judgment around trade-offs (e.g., performance vs. complexity, research velocity vs. maintainability).
  • Ability to debug performance issues across CUDA/NCCL, networking, IO, and ML data pipelines.
  • Proven experience in large-scale distributed training or HPC systems.
  • Comfort working on performance engineering, profiling, and low-level systems.
  • Experience with containerized environments (Docker and Singularity/Apptainer).
  • Strong collaboration skills across infrastructure, research, and deployment teams.

Nice to have

  • Contributions to ML frameworks (e.g., PyTorch, JAX, DeepSpeed, Megatron, xFormers).
  • Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
  • Experience optimizing data pipelines (e.g., sharded datasets, caching strategies).
  • Familiarity with evaluation/serving frameworks (e.g., vLLM, TensorRT-LLM, KV caches).
  • Experience with multi-node cluster orchestration (e.g., Slurm, Ray, Kubernetes).
  • A publication track record in top venues (NeurIPS, ICML, ICLR, etc.).
  • Experience training LLMs and large transformer architectures.

Location

Full remote.

About Cohere

Cohere is an AI company focused on building frontier-scale language models. The company works on training and deploying large transformer systems, partnering closely across infrastructure and research to deliver high-performance ML capabilities.

Scraped 5/13/2026