Data Center Engineer — Kubernetes
CaelumenAI
full-remotemidpermanentdevopsbackend San Francisco Bay Area Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
KubernetesAWS EKSLinuxTerraformAnsibleHelmInfrastructure as CodeObservabilityIncident ResponseHybrid Cloud Networking
About the role
Role overview
Design, build, and operate the infrastructure that connects physical data-center machines, AWS EKS, Kubernetes clusters, and PERCO’s worker system into a reliable deployment platform.
Responsibilities
- Design data-center deployment solutions across compute, storage, networking, power, access, and operational constraints
- Provision, commission, inventory, and troubleshoot physical Linux servers and their network paths
- Build and operate Kubernetes clusters across AWS EKS and self-managed environments
- Integrate clusters and machines with PERCO’s worker system for controlled scheduling, execution, status reporting, and lifecycle management
- Automate provisioning, upgrades, scaling, recovery, and compliance checks using Infrastructure as Code
- Establish observability for cluster health, workloads, capacity, networking, and deployment evidence
- Diagnose production failures across hardware, Linux, networking, Kubernetes, cloud services, and application workloads
- Write runbooks and ensure incidents end with root-cause fixes or clearly owned follow-ups
Requirements (signal)
- Proven production experience operating Kubernetes and Linux infrastructure
- Hands-on experience with AWS and EKS (networking, identity, storage, cluster lifecycle)
- Ability to provision and troubleshoot physical or virtual machines, networking, DNS, certificates, and container runtimes
- Experience automating infrastructure with Terraform, Ansible, Helm, or equivalent tools
- Strong incident-debugging skills using logs, metrics, events, and system state
- Ability to communicate architecture decisions, operating risks, and recovery evidence clearly
- Willingness and ability to travel to U.S. data centers or customer sites as needed
- Authorization to work in the U.S. when employment begins (OPT allowed; H-1B sponsorship supported for eligible candidates)
Bonus (nice to have)
- Bare-metal Kubernetes, Cluster API, Talos Linux, PXE/image-based provisioning, or immutable host systems
- GPU infrastructure, workload scheduling, or distributed worker systems
- Hybrid-cloud networking (VPN/private connectivity) and multi-cluster operations
- Designing secure customer-managed or regulated deployment environments
Location / work model
- Remote-first (based in the United States)
- Option to work from the San Francisco, CA area
- Planned travel required for physical deployment, commissioning, or incident response.
About CaelumenAI
CaelumenAI is building a governed production layer for agent-native software. The company focuses on infrastructure that can securely deploy and operate workloads across public cloud and customer-controlled data centers while maintaining security, observability, and human control.
Scraped 8/5/2026