xelys jobs xelys jobs

Senior Site Reliability Engineer (Cloud Platform)

Jobgether

full-remoteseniorpermanentdevopsbackend United States 72 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)KubernetesAWSObservabilityPrometheusGrafanaELK StackTerraformAnsibleLinux

About the role

Role Overview

Senior Site Reliability Engineer (Cloud Platform) responsible for building and evolving highly available, mission-critical cloud platforms at scale. You’ll embed SRE principles across the development lifecycle, with a focus on reliability, automation, observability, and continuous improvement in a fully remote environment.

Accountabilities

  • Maintain reliability, performance, and availability of production and pre-production cloud environments.
  • Design and optimize observability using metrics, logging, and tracing.
  • Respond to incidents, perform root cause analysis, and implement preventive improvements.
  • Collaborate with engineering teams to improve application reliability and integrate SRE best practices into workflows.
  • Automate operational processes to reduce manual work.
  • Create and maintain documentation, runbooks, troubleshooting guides, and incident response procedures.
  • Participate in on-call rotations and continuously improve incident management.
  • Promote a culture of reliability, automation, knowledge sharing, and continuous improvement.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • Hands-on experience managing Kubernetes and containerized production environments.
  • Proven experience supporting large-scale cloud infrastructures and mission-critical services.
  • Strong expertise in AWS and cloud-native architectures.
  • Observability/tooling experience with Prometheus, Grafana, and ELK Stack.
  • Strong scripting/automation skills using Python, Bash, Go, or similar.
  • Linux system administration experience and infrastructure automation tools such as Terraform or Ansible.
  • Solid networking understanding (TCP/IP, DNS, routing, load balancing).
  • Strong troubleshooting, communication, and collaboration skills with an automation-first mindset.

Nice-to-haves

  • SIP/VoIP experience.
  • Relational databases: MySQL/PostgreSQL.
  • NoSQL: Redis.

Benefits / Work Style

  • Fully remote and flexible.
  • Professional development support and continuous learning opportunities.
  • Mentorship, knowledge sharing, and continuous improvement culture.

Scraped 7/16/2026