xelys jobs xelys jobs

Site Reliability Engineer

BridgeSource Utilities Solutions

hybridseniorpermanentdevopsbackend United States Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSRELinuxKubernetesAWSAzureGCPTerraformCI/CDObservability

About the role

Role Overview

BridgeSource Utilities Solutions is hiring for Site Reliability Engineer, Senior Site Reliability Engineer, and Principal Site Reliability Engineer roles (levels based on experience and leadership). You will help build, operate, and continuously improve large-scale AI, cloud, and infrastructure platforms supporting mission-critical workloads.

Responsibilities

  • Improve reliability and operational excellence across distributed production environments
  • Build and maintain automation, observability, and incident response capabilities
  • Own incident management, on-call practices, and participate in root cause analysis (RCA)
  • Support and troubleshoot complex production issues
  • Apply SLIs/SLOs to drive service quality and continuous improvement
  • For senior/principal levels: provide technical leadership, architecture/design, and mentorship

Requirements

  • Experience in Site Reliability Engineering, Platform Engineering, Systems Engineering, Software Engineering, or Infrastructure Engineering
  • Strong knowledge of Linux, Kubernetes, cloud platforms (AWS, Azure, or GCP), networking, and distributed systems
  • Proficiency in Python, Go, or a similar language for automation/tooling
  • Experience with Infrastructure as Code (e.g., Terraform, Ansible)
  • Experience with CI/CD, monitoring, and observability tools
  • Hands-on production support and complex troubleshooting with RCA
  • Understanding of SLIs, SLOs, incident management, and on-call best practices
  • Strong communication and collaboration across engineering teams

Nice to Have

  • Experience building HPC clusters (1,000+ GPUs) for AI/ML workloads with hyperscale clients
  • Experience with AI infrastructure, GPU environments, HPC, InfiniBand, or RDMA

Location / Work Policy

  • Remote, with a possibility of on-site as programs and responsibilities grow
  • Preference for candidates living in the Seattle, New York, San Francisco, and Houston Metro areas

About BridgeSource Utilities Solutions

BridgeSource Utilities Solutions builds and operates large-scale AI, cloud, and infrastructure platforms for mission-critical workloads. The work centers on reliability, automation, observability, and operational excellence across distributed production environments.

Scraped 7/29/2026