xelys jobs xelys jobs

Platform Engineer

Harrison Clarke

seniorbackenddevops San Francisco Bay Area 107 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

KubernetesInfrastructure as CodeGitOpsGPUArgo CDTerraformPulumiHelmObservabilityNetworking

About the role

Role overview

You will be a Senior Platform Engineer responsible for owning the company’s core platform for production AI systems. This role goes beyond traditional DevOps: it covers GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability.

Responsibilities

  • Design & manage multi-region Kubernetes clusters across cloud and GPU-focused providers using infrastructure-as-code
  • Own the deployment lifecycle using GitOps practices (e.g., Helm, Kustomize, automated releases, continuous delivery)
  • Manage GPU infrastructure, including:
    • scheduling efficiency
    • workload placement
    • cold-start optimization
  • Oversee networking systems such as ingress, gateways, load balancing, and cross-region connectivity
  • Build and maintain observability across metrics, logs, traces, and performance profiling
  • Ensure infrastructure security around identity, secrets, and encryption
  • Support CI/CD workflows for a monorepo of services and deployment artifacts
  • Partner with ML engineers to optimize model serving and GPU utilization

Requirements

  • Strong experience operating Kubernetes in production (troubleshooting, autoscaling, upgrades)
  • Proven experience with infrastructure-as-code (e.g., Terraform, Pulumi)
  • Hands-on experience running GPU workloads on Kubernetes and optimizing resources
  • Familiarity with GitOps tooling such as Argo CD or Flux, and Helm-based deployments
  • Experience with Redis or other in-memory/distributed systems and distributed architectures
  • Strong understanding of observability tooling and practices
  • Solid networking fundamentals for low-latency/distributed systems
  • Ability to work with broad ownership across infrastructure

Preferred background

  • Exposure to GPU cloud providers beyond major hyperscalers
  • Experience with real-time/streaming infrastructure
  • Proficiency in Go or Python
  • Familiarity with ML model deployment and optimization
  • Experience managing infrastructure costs, especially for GPU-heavy workloads

About Harrison Clarke

Harrison Clarke is recruiting for an early-stage company building advanced AI systems. The company operates production AI workloads on Kubernetes across multiple clusters, regions, and hardware types, and is expanding to additional cloud providers. The platform team focuses on scaling, stabilizing, and securing the infrastructure that powers model serving.

Scraped 6/11/2026