xelys jobs xelys jobs

Lead DevOps Engineer

Paramount

leadpermanentdevopsbackend New York, NY 7 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesCI/CDTerraformHelmPrometheusOpenTelemetryGrafanaGoogle Cloud Platform (GCP)AWSMachine Learning Operations

About the role

Role overview

Lead DevOps Engineer for Paramount’s Applied Intelligence Personalization team. You will build and maintain scalable, low-latency infrastructure that powers personalization and engagement across Paramount streaming platforms, including real-time ML inference and high-traffic services.

Responsibilities

  • Design, implement, and manage scalable Kubernetes-based infrastructure for personalization services
  • Build and own CI/CD pipelines for safe production releases (canary, rollback, progressive delivery) including deployment of ML models
  • Set up observability and monitoring (Prometheus, New Relic, OpenTelemetry, Grafana), define SLIs/SLOs, and apply error-budget discipline
  • Ensure production high availability, security, and performance for APIs and streaming data pipelines
  • Partner with application, data, and ML engineers to integrate their workloads into the platform
  • Implement autoscaling strategies (HPA, KEDA, traffic-driven) for bursty traffic
  • Manage Pub/Sub and event-driven architectures for real-time messaging and engagement analytics
  • Optimize hot-path services with caching (Redis, Memcached, etc.)
  • Debug and resolve production issues related to latency, scaling, and reliability

Key projects

  • Real-time serving infrastructure for personalization and engagement (ML inference workloads)
  • Scalable, secure CI/CD pipelines for services and ML models
  • Platform-wide log aggregation and monitoring
  • Kubernetes serving optimization for minimal latency and efficient resource use
  • Improved A/B testing infrastructure for personalized experience measurement
  • Enhanced streaming data pipelines for real-time data consumers

Requirements (Basic qualifications)

  • 4+ years of production DevOps/SRE/Cloud Infrastructure experience
  • Strong Kubernetes and container orchestration experience
  • CI/CD experience with GitHub Actions, Jenkins, and ArgoCD
  • Deep cloud infrastructure knowledge (GCP, AWS, or Azure)
  • Infrastructure as Code (IaC) with Terraform and Helm
  • Experience with event-driven architectures/message queues (Pub/Sub, Kafka, etc.)
  • Monitoring/logging experience (New Relic, Prometheus, OpenTelemetry, etc.)
  • Scripting skills in Python, Bash, or Go for automation
  • Proven ownership of production reliability (SLIs/SLOs, incident response, postmortems)

Nice to have

  • Experience deploying and operating ML models in production (TF Serving, Triton, TorchServe, Ray Serve)
  • GPU/accelerator scheduling and Kubernetes node-pool management
  • Knowledge of load balancing, API gateways, and caching at scale
  • Experience with A/B testing frameworks and experimentation infrastructure

About Paramount

Paramount is a global entertainment company operating major media and streaming brands. It builds and delivers content and experiences for audiences through technology platforms, including personalization and engagement systems.

Scraped 8/6/2026