Infrastructure Engineer - Kubernetes
Alexander Chapman
on-sitemidpermanentdevopsbackend San Francisco Bay Area Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
KubernetesAWSTerraformPulumiPythonLinux NetworkingML/LLM InfrastructureGPU SchedulingCiliumIstio
About the role
Role Overview
Infrastructure Engineer (Kubernetes) to build and operate the infrastructure behind large-scale LLM inference. You’ll own the systems that enable reliable, low-latency AI workloads used by millions of users.
Responsibilities
- Design, deploy, and operate Kubernetes clusters across multiple regions and cloud providers
- Build and scale infrastructure for global AI inference workloads
- Own platform components including:
- Networking and load balancing
- Service mesh
- Observability (monitoring, logging, tracing)
- Cloud infrastructure for production systems
- Manage GPU infrastructure and autoscaling for ML workloads
- Build infrastructure using Infrastructure as Code tools (e.g., Python, Pulumi, Terraform)
- Improve platform reliability and operational excellence
Requirements
- Strong production Kubernetes experience
- Deep AWS and cloud infrastructure knowledge
- Experience with Infrastructure as Code (Pulumi, Terraform, etc.)
- Solid Linux networking fundamentals
- Python programming experience
- Experience supporting high-scale production systems
Nice to Have
- Experience with ML/LLM infrastructure, GPU scheduling, and/or Kubernetes networking (Cilium, Istio, Envoy)
- Experience with multi-cloud environments
About Alexander Chapman
Alexander Chapman is partnering with a frontier AI startup building large-scale AI systems focused on reliability, reasoning, and autonomous decision-making. The company is backed by nearly $40M in Seed funding and is developing infrastructure that powers AI products at global scale.
Scraped 7/31/2026