Senior Site Reliability Engineer (Cloud Platform)
Jobgether
full-remoteseniorpermanentdevopsbackend United States 72 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)KubernetesAWSObservabilityPrometheusGrafanaELK StackTerraformAnsibleLinux
About the role
Role Overview
Senior Site Reliability Engineer (Cloud Platform) responsible for building and evolving highly available, mission-critical cloud platforms at scale. You’ll embed SRE principles across the development lifecycle, with a focus on reliability, automation, observability, and continuous improvement in a fully remote environment.
Accountabilities
- Maintain reliability, performance, and availability of production and pre-production cloud environments.
- Design and optimize observability using metrics, logging, and tracing.
- Respond to incidents, perform root cause analysis, and implement preventive improvements.
- Collaborate with engineering teams to improve application reliability and integrate SRE best practices into workflows.
- Automate operational processes to reduce manual work.
- Create and maintain documentation, runbooks, troubleshooting guides, and incident response procedures.
- Participate in on-call rotations and continuously improve incident management.
- Promote a culture of reliability, automation, knowledge sharing, and continuous improvement.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
- Hands-on experience managing Kubernetes and containerized production environments.
- Proven experience supporting large-scale cloud infrastructures and mission-critical services.
- Strong expertise in AWS and cloud-native architectures.
- Observability/tooling experience with Prometheus, Grafana, and ELK Stack.
- Strong scripting/automation skills using Python, Bash, Go, or similar.
- Linux system administration experience and infrastructure automation tools such as Terraform or Ansible.
- Solid networking understanding (TCP/IP, DNS, routing, load balancing).
- Strong troubleshooting, communication, and collaboration skills with an automation-first mindset.
Nice-to-haves
- SIP/VoIP experience.
- Relational databases: MySQL/PostgreSQL.
- NoSQL: Redis.
Benefits / Work Style
- Fully remote and flexible.
- Professional development support and continuous learning opportunities.
- Mentorship, knowledge sharing, and continuous improvement culture.
Scraped 7/16/2026