xelys jobs xelys jobs

Machine Learning & Operations Engineer

Sundayy

full-remotemidpermanentbackenddevops United States 65 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsPythonPyTorchTensorFlowCI/CDDockerGPU WorkloadsDistributed TrainingAWSData Pipelines

About the role

Role Overview

You’ll join a remote team as a Machine Learning & Operations Engineer, designing, automating, and scaling MLOps systems for reliable deployment and management of machine learning models. You’ll work at the intersection of ML engineering and infrastructure to automate data validation, orchestrate large-scale experiments, and deploy high-performance algorithms.

Responsibilities

  • Design and maintain automated ML training pipelines to streamline model development
  • Build and optimize infrastructure for large-scale distributed experimentation
  • Create ML-focused CI/CD workflows for continuous deployment
  • Orchestrate data ingestion, preprocessing, validation, and model versioning
  • Implement experiment tracking, hyperparameter tuning automation, and reproducibility systems
  • Optimize GPU/compute resource utilization across cloud and on-premises
  • Deploy, monitor, and maintain production ML models with high availability and performance
  • Establish and enforce MLOps best practices (model registry, artifact management, observability)
  • Improve reliability, security, and performance through continuous iteration
  • Collaborate with ML researchers to productionize new algorithms
  • Support DevOps-related software development tasks as needed

Requirements

  • 3+ years in MLOps, ML infrastructure, or related roles (or equivalent degree experience)
  • Python proficiency
  • Experience with ML frameworks: PyTorch and/or TensorFlow
  • Hands-on CI/CD experience using tools such as GitHub Actions, GitLab CI, Jenkins
  • Docker and container orchestration experience
  • Experience managing GPU workloads and distributed training systems
  • Familiarity with cloud platforms: AWS, GCP, or Azure
  • Strong understanding of automation, infrastructure reliability, and data pipelines
  • Ability to collaborate effectively with international teams across regions (e.g., US and Europe)

Nice to Have

  • Not explicitly stated, but exposure to production ML observability, model registries, and artifact management is implied by responsibilities.

About Sundayy

OptiTrack (Sundayy is listed as the company) is a global provider of motion capture technology and precision tracking solutions. Its hardware and software are used in animation, robotics, virtual production, biomechanics, and industrial applications, enabling clients to capture and analyze motion with high accuracy. The company supports production and research teams worldwide with innovative, reliable motion capture systems.

Scraped 5/21/2026