Senior Site Reliability Engineer
Jobgether
seniorpermanentdevops United States 4 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)KubernetesDockercontainerdCI/CDAWSGCPAzurePrometheusHelm
About the role
Role Overview
Senior Site Reliability Engineer (SRE) responsible for designing, automating, and maintaining scalable, highly available infrastructure that supports healthcare technology solutions. You’ll work closely with engineering and data teams to improve operational efficiency, developer experience, and reliability across complex, cloud-native environments.
Responsibilities
- Design and implement systems for application and infrastructure lifecycle management, including CI/CD, continuous deployment, and Kubernetes cluster operations
- Develop automation to reduce manual work, remove operational bottlenecks, and improve engineering efficiency
- Monitor, troubleshoot, and resolve infrastructure issues while minimizing downtime
- Support and optimize containerized workloads (Kubernetes and cloud-native tooling)
- Contribute to evolution of SRE practices, standards, and team objectives
- Improve infrastructure processes (deployment workflows, database change workflows, monitoring, operational tooling)
- Collaborate across engineering and data teams to improve performance, scalability, and reliability
- Participate in incident response and alert management to maintain production availability
- Evaluate and implement improvements to infrastructure security, automation, and operational maturity
Requirements
- 5+ years of programming experience with Python, Go, or Shell scripting
- Strong experience with containerization and orchestration: Docker, containerd, Kubernetes
- Hands-on experience with cloud platforms: AWS, GCP, or Azure
- Knowledge of cloud-native tooling such as Helm, gRPC, Prometheus and related CNCF solutions
- Solid networking knowledge (TCP/IP, UDP, DNS, firewalls, routing, load balancing)
- Linux system administration and architecture principles
- Strong SRE foundation (monitoring, automation, performance optimization, reliability engineering)
- Experience building/maintaining CI/CD pipelines and modern deployment workflows
- Ability to work independently, identify problems proactively, and drive solutions
- Strong communication and cross-functional collaboration
Nice-to-haves
- Adaptability and willingness to learn new technologies (explicitly mentioned, though not otherwise specified)
Scraped 7/26/2026