MLOps Platform Engineer
dv01
seniorpermanentdevopsbackend United States 116 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
MLOpsKubernetesTerraformCI/CDObservabilityAnomaly DetectionDistributed SystemsAgentic SystemsLLMGenAIOps
About the role
Role Overview
MLOps Platform Engineer at dv01. You will design, build, and operate cloud-native infrastructure and platform tooling that helps teams develop, deploy, and run AI-powered services safely and efficiently in production.
Responsibilities
- Build and operate an AI infrastructure platform to accelerate AI development across the company.
- Own the DevOps/infrastructure side of MLOps and agentic systems, including:
- CI/CD for AI workloads
- Scalable inference infrastructure
- Observability, cost management, and reliability
- Shared services and repeatable patterns to reduce friction for AI application teams
- Enable AI services, agents, and runtime platforms, such as:
- LLM-backed APIs
- Model Context Protocol (MCP) servers
- Agentic systems used by production applications
- Secure tool access, runtime orchestration, and isolation boundaries
- Integrate MLOps into platform operations using techniques like:
- AI-driven monitoring and alerting
- anomaly detection and incident response
- Establish governance, security, and operational guardrails for AI systems:
- access controls, deployment policies, auditability, secure-by-default patterns
- partner with security/compliance to meet risk and regulatory needs
- Provide technical leadership and enablement by influencing architecture and mentoring engineers while partnering with product, data, and application teams.
Requirements
- 8+ years experience in cloud infrastructure, DevOps, or platform engineering; deep expertise operating distributed systems in production.
- 5+ years MLOps experience (required) with direct exposure to ML/GenAIOps practices such as monitoring, anomaly detection, predictive alerting, or automated remediation.
- Strong cloud-native infrastructure skills, including:
- Kubernetes
- containerized workloads
- Infrastructure as Code (e.g., Terraform)
- Hands-on experience supporting platforms that run AI workloads.
Nice-to-haves
- Experience operating AI-enabled platforms at scale with advanced observability and incident response patterns.
- Familiarity with AI runtime governance and secure-by-default deployment practices (access controls, auditability, policy enforcement).
About dv01
dv01 is a data-first analytics company focused on structured finance. It provides transparency into investment performance and risk using large-scale coverage of loans across mortgages, personal loans, auto, BNPL, small business, and student loans, serving 400+ major financial institutions.
Scraped 4/1/2026