xelys jobs xelys jobs

Senior Manager of Cloud Platform & Site Reliability

Baseten

full-remoteseniorpermanentengineering-managementdevops Full remote 65 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesSite Reliability Engineering (SRE)TerraformPulumiCI/CDGitOpsObservabilitySLOs/SLIsIncident ManagementMulti-cloud

About the role

Role overview

Baseten is hiring a Senior Manager, Cloud Platform & Site Reliability (full remote). You will lead and grow the organization responsible for the infrastructure powering Baseten’s machine learning platform—setting technical direction, reliability standards, and scaling the SRE/platform engineering practice alongside product growth.

Key missions

  • Lead and grow the infrastructure organization that powers Baseten’s ML platform.
  • Set org-level technical direction and roadmap for infrastructure, reliability, and platform engineering.
  • Own end-to-end reliability posture for the platform by defining and enforcing standards for:
    • SLOs/SLIs
    • Incident response
    • Observability-as-code
    • Runbooks
    • Post-incident reviews
  • Drive cross-functional collaboration and ensure best practices are adopted and maintained across the organization.

Responsibilities

  • Own the end-to-end health of cloud infrastructure and the SRE practice.
  • Engage credibly in architectural decisions while still setting organization-wide direction.
  • Lead incident management, including executive-level communication during high-severity incidents and disciplined post-incident follow-through.
  • Lead complex, multi-stakeholder technical initiatives from scoping through execution.
  • Manage and grow multiple infrastructure/platform/SRE teams, including managers.

Requirements

  • Strong ability to zoom out for org-level strategy while staying technically credible across:
    • Kubernetes and multi-cloud infrastructure
    • Reliability engineering and distributed systems
  • Excellent communication with executive presence; can explain technical work to technical and non-technical audiences.
  • Experience owning incident management and enterprise SLAs at scale.
  • Proven experience managing managers and leading multiple high-performing infrastructure/platform/SRE teams in a fast-paced, high-growth environment.
  • Hands-on experience with Infrastructure-as-Code (e.g., Terraform, Pulumi) and CI/CD (e.g., GitHub Actions, GitLab CI, Jenkins).
  • Familiarity with GitOps workflows (e.g., Flux CD, Argo CD, Helm).
  • Deep expertise in Kubernetes (e.g., EKS, GKE) and distributed systems.
  • Strong observability background and experience raising reliability standards via SLOs/SLIs and observability-as-code.
  • Familiarity with observability stack/tools including Prometheus, VictoriaMetrics, Loki, ELK, Grafana, and OpenTelemetry.
  • Openness to learning ML infrastructure and model serving (no prior ML experience required).
  • Interest/experience with high-performance AI workloads, including troubleshooting ML pipelines and serving.
  • Experience with GPU infrastructure, including fractional GPU provisioning and multi-node model serving (e.g., H100/B200).
  • Experience with incident management platforms (e.g., incident.io, PagerDuty) and building AI-assisted incident triage.
  • Experience scaling an SRE practice: runbook standards, self-healing automations, and systematic mitigations.

Nice to have

  • Experience with GPU/multi-node model serving specifically.
  • Experience building AI-assisted tooling for incident response.

About Baseten

Baseten builds a machine learning platform, including the infrastructure and reliability needed to run ML workloads in production. The role focuses on leading cloud platform and Site Reliability Engineering to scale Baseten’s ML infrastructure with strong reliability standards.

Scraped 7/4/2026