ML Ops Engineer
Circadia Health
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
As an ML Ops Engineer at Circadia Health, you will own the infrastructure and operational lifecycle of the company’s machine learning systems used in a clinical monitoring platform. You will build production ML pipelines, deployment infrastructure, and monitoring systems—working across ML, backend, data, and clinical teams—to ensure models are reliably trained, versioned, deployed, and monitored in both cloud and edge environments.
Key Responsibilities
- ML pipeline orchestration: Own and extend production ML pipelines using Apache Airflow for training, evaluation, and deployment workflows.
- Automated retraining & promotion: Build pipelines for model retraining, validation, and promotion across dev/staging/production.
- Reliability & monitoring: Implement monitoring, alerting, and failure recovery to prevent silent failures and improve operational robustness.
- Reproducibility for experimentation: Design architectures that enable rapid experimentation while enforcing production-grade reproducibility.
- Deployment & rollout safety:
- Deploy and manage models on AWS (e.g., AWS Batch for batch inference).
- Support deployment to edge devices in clinical monitoring hardware.
- Implement safe rollout strategies like shadow deployments and canary releases.
- Model management with MLflow:
- Manage model versioning, promotion, and rollback via MLflow Model Registry.
- Maintain MLflow experiment tracking/registry infrastructure.
- Define conventions for experiment logging, artifact storage, metadata, and lineage tracking.
- Dataset versioning & lineage:
- Establish training data versioning and dataset management for reproducibility.
- Track dataset lineage, labeling provenance, and feature dependencies.
- Collaborate to formalize dataset release and validation workflows.
- Production performance monitoring: Build monitoring for model performance, including data drift detection, prediction quality tracking, and degradation alerting; create operational dashboards for pipeline health and deployment status.
- Incident readiness: Develop incident response procedures and runbooks for ML system failures.
- Cloud resource optimization: Manage and optimize AWS compute resources (Batch/EC2 or similar) and drive cost optimization.
- Infrastructure as Code: Build reproducible ML environments using Infrastructure-as-Code.
- Data integrations: Support Snowflake integrations for feature generation and training data pipelines.
- Best practices & enablement:
- Champion ML engineering best practices: CI/CD for models, automated testing for ML pipelines, reproducible training workflows.
- Build internal tooling/templates to streamline the path from experimentation to production.
- Document processes, architecture decisions, and onboarding materials.
- Participate in architecture discussions and planning.
- Healthcare security compliance: Ensure ML pipelines and infrastructure meet healthcare security requirements (text truncated in posting).
Requirements / Skills
- Strong ownership of end-to-end ML operations for production ML systems.
- Experience orchestrating production ML workflows (explicitly: Apache Airflow).
- Hands-on work with AWS for training/inference and operational deployment.
- Use of MLflow for experiment tracking and model registry.
- Capability to implement monitoring/alerting for ML pipelines and production model health.
- Experience with dataset versioning/lineage to ensure training reproducibility.
- Ability to support/coordinate deployments to edge hardware environments.
- Familiarity with Infrastructure-as-Code and cost optimization in cloud ML.
Nice-to-Haves (Implied)
- Experience integrating with Snowflake for data/feature pipelines.
- Experience implementing rollout safety patterns (canary/shadow) for model releases.
- Experience with CI/CD and automated testing for ML pipelines.
- Experience with healthcare/security-sensitive environments.
About Circadia Health
Circadia Health builds a clinical monitoring platform that uses predictive machine learning models to detect early signs of patient deterioration. The company operates at the intersection of healthcare, data, and production ML to improve patient outcomes through reliable and continuously improving model operations.
Scraped 7/24/2026