xelys jobs xelys jobs

MLOps Engineer

Talener

full-remoteseniorpermanentbackenddevops New York, NY 47 days ago via LinkedIn
140,000 - 140,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsAWS SageMakerPythonCI/CDPyTorchTensorFlowCI/CDContainersMonitoringML Inference

About the role

Role overview

You’ll be an MLOps Engineer responsible for production operations of machine learning systems—not for data science or model building. Your goal is to ensure ML inference pipelines run reliably, cost-effectively, and at scale, with safe releases across Dev, QA, and Prod.

Responsibilities

  • Own the full production lifecycle of ML systems: deploying, scaling, monitoring, and governing inference pipelines
  • Ensure production ML systems meet SLAs and operate safely across environments
  • Build and standardize deployment patterns and operational frameworks, including:
    • Containerization strategies
    • Environment isolation
    • Versioned rollouts and rollback mechanisms
    • Monitoring frameworks for ongoing operations
  • Maintain production stability and cost control for media pipelines handling hundreds of thousands of queries per day
  • Establish operational reliability practices in shared ownership with ML, Data Science, DevOps, and Platform teams

Required skills

  • 5+ years professional experience
  • Experience working with large amounts of data, specifically text, images, and videos
  • Experience in a media firm or large real-time environment where uptime is paramount
  • Hands-on production experience deploying and operating ML inference systems
  • Strong AWS SageMaker experience (pipelines, endpoints, monitoring, multi-environment deployments)
  • Python for pipeline and tooling work
  • PyTorch and TensorFlow from an ops/serving perspective (not modeling)
  • Production experience with BERT/transformer-based NLP models
  • CI/CD experience (Jenkins and/or GitLab)
  • Containerized inference and autoscaling (deployment and orchestration)
  • Compute optimization for production ML (CPU/GPU selection, benchmarking, cost/performance)
  • Monitoring, alerting, drift detection, and A/B testing frameworks for ML
  • Comfortable with shared ownership across multiple teams

Nice to have

  • Computer vision or ranking/reranking systems experience
  • Familiarity with ANN methods (e.g., HNSW)
  • Experience running ML workloads over large-scale media datasets (text, image, video)

About Talener

Talener is a technology team within a global newswire and media organization focused on large-scale media processing. The organization provides high-volume daily content delivery and is investing heavily in modern machine learning infrastructure for text, image, and video pipelines.

Scraped 6/18/2026