Staff Engineer - ML Operations - USA Remote
Danaher
full-remoteseniorpermanentengineering-managementbackend New York, NY 10 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine Learning Operations (MLOps)MLflowKubeflowWeights & BiasesExperiment TrackingModel Registry/Versioning/LineageModel Serving (Batch & Low-Latency)KubernetesObservability (OpenTelemetry/Prometheus/Grafana)GPU Compute Optimization
About the role
Role Overview
Staff Engineer - ML Operations for Danaher’s AI-driven research platform. Own significant parts of the machine learning lifecycle, taking models from experimentation to reliable, scalable production. Fully remote and reporting to the Senior Director, Data and AI Platform within the Chief Scientific Officer (CSO) Office.
Responsibilities
- Own end-to-end ML lifecycle and deployment, including:
- experiment tracking, model registry, versioning, lineage, and reproducibility (e.g., MLflow, Weights & Biases, Kubeflow)
- Design and operate model serving for batch and low-latency online inference:
- autoscaling, GPU efficiency, performance optimization (batching, quantization, caching)
- ensure models in production are traceable, auditable, and performant
- Productionize large-scale protein design and structure-prediction workflows with bioinformatics/computational biology teams:
- build scalable, repeatable, high-throughput pipelines using Airflow, Dagster, Prefect, Nextflow
- run containerized, reproducible execution supporting many concurrent researchers
- Implement CI/CD, continuous training, and ML observability:
- automate path from model code to validated production via testing/evaluation gates
- safe deployment patterns (blue/green, canary)
- monitor performance, data/prediction drift, latency, and cost
- instrument monitoring/alerts with OpenTelemetry/Prometheus/Grafana
- Drive GPU and accelerated compute efficiency:
- scheduling, quota/utilization management, driver/CUDA image hygiene
- Build self-service ML tooling and provide technical leadership:
- create “golden paths” so researchers can train/track/serve/monitor without deep infrastructure expertise
- set MLOps standards while staying hands-on with architecture and delivery
Requirements
- Degree in Computer Science, Engineering, Computational Biology, or related field (or equivalent practical experience)
- 5+ years of software, ML, or infrastructure engineering experience
- Hands-on MLOps experience with a demonstrated track record taking ML models to production at scale
- Strong ML lifecycle tooling experience (experiment tracking, observability, model registry, versioning, lineage, reproducibility)
- e.g., MLflow, Kubeflow, Weights & Biases
- Strong containerization and orchestration experience:
- Docker, Kubernetes, including scaling GPU workloads
Nice-to-haves
- Experience with automated retraining, alerting, and drift monitoring in production
- Experience optimizing GPU workloads (scheduling, utilization, performance/cost optimization)
- Experience building high-throughput scientific pipelines for concurrent users
About Danaher
Danaher is a leading science and technology company operating across life sciences, diagnostics, and biotechnology. It delivers innovations that help save lives, with a global workforce and a focus on continuous improvement and scalable impact.
Scraped 7/18/2026