xelys jobs xelys jobs

MLOps Platform Engineer

dv01

seniorpermanentdevopsbackend United States 116 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsKubernetesTerraformCI/CDObservabilityAnomaly DetectionDistributed SystemsAgentic SystemsLLMGenAIOps

About the role

Role Overview

MLOps Platform Engineer at dv01. You will design, build, and operate cloud-native infrastructure and platform tooling that helps teams develop, deploy, and run AI-powered services safely and efficiently in production.

Responsibilities

  • Build and operate an AI infrastructure platform to accelerate AI development across the company.
  • Own the DevOps/infrastructure side of MLOps and agentic systems, including:
    • CI/CD for AI workloads
    • Scalable inference infrastructure
    • Observability, cost management, and reliability
    • Shared services and repeatable patterns to reduce friction for AI application teams
  • Enable AI services, agents, and runtime platforms, such as:
    • LLM-backed APIs
    • Model Context Protocol (MCP) servers
    • Agentic systems used by production applications
    • Secure tool access, runtime orchestration, and isolation boundaries
  • Integrate MLOps into platform operations using techniques like:
    • AI-driven monitoring and alerting
    • anomaly detection and incident response
  • Establish governance, security, and operational guardrails for AI systems:
    • access controls, deployment policies, auditability, secure-by-default patterns
    • partner with security/compliance to meet risk and regulatory needs
  • Provide technical leadership and enablement by influencing architecture and mentoring engineers while partnering with product, data, and application teams.

Requirements

  • 8+ years experience in cloud infrastructure, DevOps, or platform engineering; deep expertise operating distributed systems in production.
  • 5+ years MLOps experience (required) with direct exposure to ML/GenAIOps practices such as monitoring, anomaly detection, predictive alerting, or automated remediation.
  • Strong cloud-native infrastructure skills, including:
    • Kubernetes
    • containerized workloads
    • Infrastructure as Code (e.g., Terraform)
  • Hands-on experience supporting platforms that run AI workloads.

Nice-to-haves

  • Experience operating AI-enabled platforms at scale with advanced observability and incident response patterns.
  • Familiarity with AI runtime governance and secure-by-default deployment practices (access controls, auditability, policy enforcement).

About dv01

dv01 is a data-first analytics company focused on structured finance. It provides transparency into investment performance and risk using large-scale coverage of loans across mortgages, personal loans, auto, BNPL, small business, and student loans, serving 400+ major financial institutions.

Scraped 4/1/2026