Senior Site Reliability Engineer
Jobgether
full-remoteseniorpermanentdevops United States 74 days ago via LinkedIn
149,100 - 157,800 USD/annual
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
AWSKubernetesTerraformGitOpsObservabilitySLIs/SLOsIncident ResponseFinOpsInternal Developer PlatformsPython
About the role
Role Overview
Senior Site Reliability Engineer (fully remote within the United States) for an AI-driven platform. You will lead reliability initiatives across complex cloud infrastructure, AI workloads, and internal developer enablement systems, partnering with platform, software engineering, and AI engineering teams.
Responsibilities
- Own platform reliability initiatives, defining and managing SLIs, SLOs, and error budgets.
- Design resilient infrastructure patterns for AI pipelines, including observability, failure detection, graceful degradation, and workload isolation.
- Lead incident response, disaster recovery planning, and post-incident reviews to drive long-term improvements.
- Establish reliability standards, deployment best practices, and scalable CI/CD workflows with engineering teams.
- Build and maintain observability using monitoring, tracing, logging, and telemetry.
- Manage Infrastructure as Code (IaC), cloud cost optimization (FinOps), and automation strategies.
- Build/enhance Internal Developer Platforms (IDP), service catalogs, and self-service tooling.
- Mentor junior and intermediate engineers and support technical growth and knowledge sharing.
Requirements
- Bachelor’s degree (CS/Engineering or related) or equivalent practical experience.
- 6–8 years in Site Reliability Engineering, Platform Engineering, or DevOps with technical leadership responsibilities.
- Deep expertise in AWS, Kubernetes, Docker, Terraform, and GitOps.
- Strong observability and operational tooling experience (including distributed tracing, monitoring).
- Proficiency in Python and/or Bash scripting; experience with microservices and CI/CD pipelines.
- Experience with Internal Developer Platforms (e.g., Backstage) is highly desirable.
- Strong analytical, communication, mentoring, and problem-solving skills.
Nice-to-haves
- Experience supporting AI/ML infrastructure, LLM integrations, or agentic systems.
- Experience with disaster recovery, policy-as-code, regulated environments, or FinOps.
About Jobgether
Jobgether is an employment platform that uses an AI-powered matching process to connect candidates with partner companies. The role described is for a partner company building and operating an AI-driven platform, with a focus on reliability, observability, and developer enablement.
Scraped 5/13/2026