xelys jobs xelys jobs

MLOps Engineer | Remote | $70 –$110/hr

Call For Referral

full-remotemidcontractdevopsbackend United States 85 days ago via LinkedIn
108,000 - 168,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsML InfrastructureJAXPyTorchTritonPallasGPU KernelsDistributed TrainingGenAITechnical Documentation

About the role

Role Overview

MLOps Engineer supporting frontier GenAI initiatives by improving large-scale ML infrastructure, training performance, and distributed training systems.

Responsibilities

  • Support AI research and engineering teams in improving ML infrastructure and training systems
  • Design advanced MLOps and ML system solutions with structured technical deliverables
  • Evaluate ML system outputs and provide detailed technical feedback
  • Develop evaluation rubrics and frameworks for distributed systems, training pipelines, and kernel-level optimization
  • Collaborate with domain experts to maintain consistency and quality across AI training workflows
  • Improve large-scale model training performance and infrastructure reliability

Requirements

  • 2+ years of professional experience in ML infrastructure, MLOps, or ML systems engineering
  • Hands-on production experience with JAX and/or PyTorch at scale
  • Experience writing or optimizing GPU kernels using Pallas or Triton
  • Strong understanding of ML training systems and distributed infrastructure
  • Demonstrated engineering career progression
  • Ability to commit to full-time 40 hours/week on weekdays
  • Strong written communication and technical documentation skills

Nice to Haves

  • (Not explicitly stated; focus is on the kernel-level optimization and distributed training/evaluation expertise.)

About Call For Referral

Call For Referral is recruiting for an MLOps Engineer role supporting next-generation AI systems. The work centers on large-scale ML infrastructure, training optimization, and framework-level engineering for GenAI initiatives.

Scraped 7/1/2026