xelys jobs xelys jobs

Site Reliability Engineer

Jobgether

seniorpermanentdevops United States Today via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringObservabilityIncident ManagementRoot Cause Analysis (RCA)New RelicDatadogAWSAzureCI/CDInfrastructure as Code (IaC)

About the role

Role Overview

Site Reliability Engineer (SRE) focused on improving platform reliability, scalability, and operational performance across cloud-based SaaS environments. You will reduce incident impact, accelerate resolution times, and build proactive solutions to prevent future disruptions.

Responsibilities

  • Enhance platform resilience and improve incident management processes
  • Lead post-incident investigations and perform root cause analysis (RCA)
  • Produce clear, actionable RCA documentation for internal teams and customer delivery
  • Develop and implement preventative reliability strategies
  • Monitor and improve reliability metrics (e.g., time to resolution, incident response effectiveness)
  • Configure and maintain observability tooling for monitoring, alerting, and performance visibility
  • Build client-focused dashboards and alerts to proactively identify issues
  • Collaborate with Engineering and Cloud Operations to improve scalability and stability
  • Provide guidance and knowledge sharing on observability usage during investigations
  • Create feedback loops to improve engineering practices and operational patterns
  • Contribute to automation initiatives that streamline incident response and workflows
  • Troubleshoot application and infrastructure issues across cloud environments

Requirements

  • 5+ years of professional SRE experience
  • Strong understanding of SRE principles, reliability practices, and incident management
  • Experience with observability platforms (e.g., New Relic, Datadog, Sumo Logic, or similar)
  • Ability to read and write code using JavaScript, .NET, and SQL
  • Experience troubleshooting C#/.NET web applications and performance issues
  • Familiarity with cloud platforms (AWS or Azure); AWS strongly preferred
  • Experience operating in public cloud environments and SaaS platforms
  • Strong understanding of cloud architecture patterns and operational best practices

Nice to Have

  • Knowledge of CI/CD pipelines and Infrastructure as Code (IaC)
  • Experience troubleshooting Windows environments and SQL Server

Scraped 7/30/2026