Senior Site Reliability Engineer
Jobgether
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Senior Site Reliability Engineer responsible for designing, building, and optimizing reliable infrastructure that supports healthcare technology solutions. You will work across software engineering, cloud infrastructure, automation, and operational excellence to improve scalability, reliability, and performance across distributed systems.
Responsibilities
- Design, maintain, and improve scalable infrastructure platforms with a focus on automation and operational excellence.
- Implement declarative application/infrastructure lifecycle management (CI/CD, continuous deployment, service inventory management).
- Build and automate infrastructure processes to reduce manual effort and operational complexity.
- Manage and support containerized workloads, including Kubernetes-based environments and cloud infrastructure platforms.
- Monitor, troubleshoot, and resolve infrastructure issues while minimizing downtime.
- Develop automation tools/workflows to improve deployment efficiency, system performance, and operational consistency.
- Improve monitoring, alerting, performance tuning, and incident response processes.
- Support reliability across compute, storage, and networking environments.
- Contribute to the strategic direction of SRE practices and align technical initiatives with business objectives.
- Collaborate with engineering teams, technical leads, and data professionals; promote a culture of knowledge sharing and continuous improvement.
Requirements
- 5+ years of programming experience (Python, Go, and/or Shell scripting).
- Strong containerization experience (Docker, containerd, Kubernetes).
- Hands-on CNCF technologies including Helm, gRPC, and Prometheus.
- Experience with public cloud environments (AWS, GCP, or Azure).
- Strong networking fundamentals (TCP/IP, UDP, DNS, routing, firewalls, load balancing).
- Solid Linux system administration experience and Linux architecture knowledge.
- Understanding of SRE concepts: monitoring, automation, performance optimization, incident management.
- Proven experience building and maintaining reliable production infrastructure.
- Ability to work independently, proactively identify challenges, and drive solutions.
- Strong communication and collaboration skills.
- Adaptable and willing to learn new technologies quickly.
Nice to Have
- Not explicitly stated, beyond the listed CNCF and cloud ecosystem experience.
About Jobgether
The role is listed on behalf of a partner company that provides and manages applications supporting innovative healthcare technology solutions. The work focuses on building reliable, scalable infrastructure and improving operational excellence for production systems.
Scraped 7/28/2026