Member of the Technical Staff - Platform
Andromeda
hybridseniorpermanentbackenddevops San Francisco, CA Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
KubernetesPostgreSQLLinuxCI/CDCloud InfrastructureFleet ManagementCapacity PlanningKubernetes OperatorsOn-CallDatabase Operations
About the role
Role Overview
Member of the Technical Staff – Platform, working on the control plane that runs and manages Andromeda’s compute fleet.
Responsibilities
- Build and operate automated systems that take clusters from bare machines to customer-ready environments
- Manage machine lifecycle between tenants, including join, wipe, verify, rejoin
- Operate Kubernetes and Postgres across the fleet
- Develop and contribute to custom Kubernetes operators
- Scale clusters from tens of nodes to thousands
- Participate in on-call rotations
Requirements
- 2+ years of on-call experience for critical production services
- Strong Kubernetes expertise
- Solid Linux fundamentals (kernel, cgroups, containers, networking, storage)
- Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale
- Familiarity with fleet management and capacity planning
Nice to Have
- Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware
About Andromeda
Andromeda builds AI infrastructure that makes compute “liquid,” deploying into external datacenters and turning available hardware into clusters for AI labs. The company operates compute for 80+ customers across 30+ capacity providers, managing tens of thousands of GPUs and billions of GPU-hours.
Scraped 7/31/2026