xelys jobs xelys jobs

Senior Site Reliability Engineer

Jobgether

full-remoteseniorpermanentdevops United States 74 days ago via LinkedIn
149,100 - 157,800 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

AWSKubernetesTerraformGitOpsObservabilitySLIs/SLOsIncident ResponseFinOpsInternal Developer PlatformsPython

About the role

Role Overview

Senior Site Reliability Engineer (fully remote within the United States) for an AI-driven platform. You will lead reliability initiatives across complex cloud infrastructure, AI workloads, and internal developer enablement systems, partnering with platform, software engineering, and AI engineering teams.

Responsibilities

  • Own platform reliability initiatives, defining and managing SLIs, SLOs, and error budgets.
  • Design resilient infrastructure patterns for AI pipelines, including observability, failure detection, graceful degradation, and workload isolation.
  • Lead incident response, disaster recovery planning, and post-incident reviews to drive long-term improvements.
  • Establish reliability standards, deployment best practices, and scalable CI/CD workflows with engineering teams.
  • Build and maintain observability using monitoring, tracing, logging, and telemetry.
  • Manage Infrastructure as Code (IaC), cloud cost optimization (FinOps), and automation strategies.
  • Build/enhance Internal Developer Platforms (IDP), service catalogs, and self-service tooling.
  • Mentor junior and intermediate engineers and support technical growth and knowledge sharing.

Requirements

  • Bachelor’s degree (CS/Engineering or related) or equivalent practical experience.
  • 6–8 years in Site Reliability Engineering, Platform Engineering, or DevOps with technical leadership responsibilities.
  • Deep expertise in AWS, Kubernetes, Docker, Terraform, and GitOps.
  • Strong observability and operational tooling experience (including distributed tracing, monitoring).
  • Proficiency in Python and/or Bash scripting; experience with microservices and CI/CD pipelines.
  • Experience with Internal Developer Platforms (e.g., Backstage) is highly desirable.
  • Strong analytical, communication, mentoring, and problem-solving skills.

Nice-to-haves

  • Experience supporting AI/ML infrastructure, LLM integrations, or agentic systems.
  • Experience with disaster recovery, policy-as-code, regulated environments, or FinOps.

About Jobgether

Jobgether is an employment platform that uses an AI-powered matching process to connect candidates with partner companies. The role described is for a partner company building and operating an AI-driven platform, with a focus on reliability, observability, and developer enablement.

Scraped 5/13/2026