xelys jobs xelys jobs

Engineering Manager (Ads ML Efficiency)

Reddit

full-remoteleadpermanentengineering-managementbackend Full remote 28 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine LearningML EngineeringDistributed SystemsPyTorchGPU OptimizationTraining OptimizationInference OptimizationModel PerformanceLoad TestingEngineering Management

About the role

Role overview

Join Reddit as an Engineering Manager for Ads ML Efficiency. You will lead a team responsible for improving the efficiency of Ads ML systems across training, inference, GPU enablement, load/performance testing, and model performance tooling.

Responsibilities

  • Lead and manage an Ads ML efficiency team (optimization, training efficiency, GPU enablement, load testing, and performance tooling).
  • Define and execute roadmaps for:
    • training optimization
    • inference optimization
    • launch-readiness tooling
    • reusable efficiency primitives across Ads ML
  • Drive measurable reductions in:
    • model training time
    • online latency
    • serving cost
    • infra-driven launch risk
  • Partner with model owners and platform teams to deliver improvements.

Requirements

  • Deep ML engineering experience, including hands-on understanding of training, serving, debugging, and optimization.
  • Strong people leadership: building/leads teams, coaching engineers, delivery management, and prioritization under ambiguity.
  • Fluency in distributed systems and production-scale ML tradeoffs (reliability, speed, cost, scale).
  • Ability to act with customer/platform instincts: support modeling teams while building reusable systems.
  • Clear technical communication with engineers, PMs, and senior stakeholders.
  • Hands-on optimization background in training loops, serving systems, profiling workflows, model/inference efficiency, or GPU utilization.

Nice to have

  • Ads or recommender/marketplace ML experience (ads ranking, recommender systems, production ML adjacent domains).
  • Experience with GPU training/serving migrations.
  • PyTorch and distributed training frameworks; familiarity with kernel/performance optimization.
  • Experience building efficiency benchmarking or launch certification frameworks.
  • Experience working where ML platform and applied modeling are split across multiple teams.

Location / Remote

  • Full remote

About Reddit

Reddit is a global online platform built around community-driven discussion. It operates large-scale services and uses data and machine learning to power products such as advertising and personalized experiences.

Scraped 6/27/2026