xelys jobs xelys jobs

Staff Engineer - ML Operations - USA Remote

Danaher

full-remoteseniorpermanentengineering-managementbackend New York, NY 10 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine Learning Operations (MLOps)MLflowKubeflowWeights & BiasesExperiment TrackingModel Registry/Versioning/LineageModel Serving (Batch & Low-Latency)KubernetesObservability (OpenTelemetry/Prometheus/Grafana)GPU Compute Optimization

About the role

Role Overview

Staff Engineer - ML Operations for Danaher’s AI-driven research platform. Own significant parts of the machine learning lifecycle, taking models from experimentation to reliable, scalable production. Fully remote and reporting to the Senior Director, Data and AI Platform within the Chief Scientific Officer (CSO) Office.

Responsibilities

  • Own end-to-end ML lifecycle and deployment, including:
    • experiment tracking, model registry, versioning, lineage, and reproducibility (e.g., MLflow, Weights & Biases, Kubeflow)
  • Design and operate model serving for batch and low-latency online inference:
    • autoscaling, GPU efficiency, performance optimization (batching, quantization, caching)
    • ensure models in production are traceable, auditable, and performant
  • Productionize large-scale protein design and structure-prediction workflows with bioinformatics/computational biology teams:
    • build scalable, repeatable, high-throughput pipelines using Airflow, Dagster, Prefect, Nextflow
    • run containerized, reproducible execution supporting many concurrent researchers
  • Implement CI/CD, continuous training, and ML observability:
    • automate path from model code to validated production via testing/evaluation gates
    • safe deployment patterns (blue/green, canary)
    • monitor performance, data/prediction drift, latency, and cost
    • instrument monitoring/alerts with OpenTelemetry/Prometheus/Grafana
  • Drive GPU and accelerated compute efficiency:
    • scheduling, quota/utilization management, driver/CUDA image hygiene
  • Build self-service ML tooling and provide technical leadership:
    • create “golden paths” so researchers can train/track/serve/monitor without deep infrastructure expertise
    • set MLOps standards while staying hands-on with architecture and delivery

Requirements

  • Degree in Computer Science, Engineering, Computational Biology, or related field (or equivalent practical experience)
  • 5+ years of software, ML, or infrastructure engineering experience
  • Hands-on MLOps experience with a demonstrated track record taking ML models to production at scale
  • Strong ML lifecycle tooling experience (experiment tracking, observability, model registry, versioning, lineage, reproducibility)
    • e.g., MLflow, Kubeflow, Weights & Biases
  • Strong containerization and orchestration experience:
    • Docker, Kubernetes, including scaling GPU workloads

Nice-to-haves

  • Experience with automated retraining, alerting, and drift monitoring in production
  • Experience optimizing GPU workloads (scheduling, utilization, performance/cost optimization)
  • Experience building high-throughput scientific pipelines for concurrent users

About Danaher

Danaher is a leading science and technology company operating across life sciences, diagnostics, and biotechnology. It delivers innovations that help save lives, with a global workforce and a focus on continuous improvement and scalable impact.

Scraped 7/18/2026