xelys jobs xelys jobs

Data Center Engineer — Kubernetes

CaelumenAI

full-remotemidpermanentdevopsbackend San Francisco Bay Area Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesAWS EKSLinuxTerraformAnsibleHelmInfrastructure as CodeObservabilityIncident ResponseHybrid Cloud Networking

About the role

Role overview

Design, build, and operate the infrastructure that connects physical data-center machines, AWS EKS, Kubernetes clusters, and PERCO’s worker system into a reliable deployment platform.

Responsibilities

  • Design data-center deployment solutions across compute, storage, networking, power, access, and operational constraints
  • Provision, commission, inventory, and troubleshoot physical Linux servers and their network paths
  • Build and operate Kubernetes clusters across AWS EKS and self-managed environments
  • Integrate clusters and machines with PERCO’s worker system for controlled scheduling, execution, status reporting, and lifecycle management
  • Automate provisioning, upgrades, scaling, recovery, and compliance checks using Infrastructure as Code
  • Establish observability for cluster health, workloads, capacity, networking, and deployment evidence
  • Diagnose production failures across hardware, Linux, networking, Kubernetes, cloud services, and application workloads
  • Write runbooks and ensure incidents end with root-cause fixes or clearly owned follow-ups

Requirements (signal)

  • Proven production experience operating Kubernetes and Linux infrastructure
  • Hands-on experience with AWS and EKS (networking, identity, storage, cluster lifecycle)
  • Ability to provision and troubleshoot physical or virtual machines, networking, DNS, certificates, and container runtimes
  • Experience automating infrastructure with Terraform, Ansible, Helm, or equivalent tools
  • Strong incident-debugging skills using logs, metrics, events, and system state
  • Ability to communicate architecture decisions, operating risks, and recovery evidence clearly
  • Willingness and ability to travel to U.S. data centers or customer sites as needed
  • Authorization to work in the U.S. when employment begins (OPT allowed; H-1B sponsorship supported for eligible candidates)

Bonus (nice to have)

  • Bare-metal Kubernetes, Cluster API, Talos Linux, PXE/image-based provisioning, or immutable host systems
  • GPU infrastructure, workload scheduling, or distributed worker systems
  • Hybrid-cloud networking (VPN/private connectivity) and multi-cluster operations
  • Designing secure customer-managed or regulated deployment environments

Location / work model

  • Remote-first (based in the United States)
  • Option to work from the San Francisco, CA area
  • Planned travel required for physical deployment, commissioning, or incident response.

About CaelumenAI

CaelumenAI is building a governed production layer for agent-native software. The company focuses on infrastructure that can securely deploy and operate workloads across public cloud and customer-controlled data centers while maintaining security, observability, and human control.

Scraped 8/5/2026