Site Reliability Engineer
Jobgether
seniorpermanentdevops United States Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)AWSOracle CloudLinuxWindows ServerKubernetesTerraformPrometheusGrafanaAnsible
About the role
Role Overview
Site Reliability Engineer supporting mission-critical, high-availability platforms across cloud and on-premises environments. You’ll work with cross-functional teams to improve reliability through automation, monitoring, and operational excellence.
Responsibilities
- Design, implement, maintain, and optimize highly available infrastructure for critical applications and services
- Monitor production, analyze performance, and proactively improve stability, scalability, and operational efficiency
- Troubleshoot and resolve critical incidents across infrastructure, networking, hardware, and software
- Build and maintain monitoring, alerting, backup/recovery, and disaster recovery procedures to maximize uptime
- Manage cloud platforms, virtualization, network infrastructure, and remote monitoring for secure operations
- Implement infrastructure automation using configuration management, scripting, and Infrastructure-as-Code
- Participate in post-incident reviews, document improvements, and support continuous reliability and security enhancements
- Collaborate to strengthen CI/CD pipelines and infrastructure resilience and best practices
Requirements
- Bachelor’s degree in Computer Science/IT or related field (certifications a plus)
- 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, Systems Administrator, or similar
- Strong cloud experience (AWS or Oracle Cloud)
- Proficient with Linux and Windows server administration, virtualization, and enterprise infrastructure management
- Experience with: Docker, Kubernetes, Terraform, Git, GitLab CI/CD, ELK Stack, Prometheus, and Grafana
- Knowledge of MySQL and PostgreSQL, networking concepts (LAN/WAN, HTTP, TCP/IP), security, and backup/recovery
- Automation/scripting: Ansible, Bash, Rundeck, or Puppet
- Strong troubleshooting and analytical skills; ability to work independently and collaboratively
- Ability to respond to critical production incidents outside standard business hours when required
Nice to Have
- Nginx, PHP-FPM, SSL, DNS, and Cloudflare
About Jobgether
Jobgether is an AI-powered recruiting platform that matches candidates to roles at partner companies. It uses an automated screening and ranking process to present shortlists to hiring employers for the final selection steps.
Scraped 8/5/2026