xelys jobs xelys jobs

Site Reliability Engineer

Origami Risk

hybridseniorpermanentdevopsbackend United States 74 days ago via LinkedIn
100,000 - 120,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)Incident ManagementObservabilityNew RelicDatadogSumo LogicAWSAzureCI/CDInfrastructure as Code (IaC)

About the role

Role overview

As a Site Reliability Engineer, you will improve Origami’s time to resolution and drive site reliability, scalability, and stability. You’ll lead post-incident investigations, identify root causes, and implement preventive measures. You’ll also analyze client performance issues and help track key performance metrics across client implementations.

Responsibilities

  • Lead post-incident investigations for the Site Reliability team
  • Perform root-cause analyses (RCAs) and develop preventive strategies
  • Draft clear, insightful RCAs for customer delivery
  • Cross-train others on using observability tools during incident and performance investigations
  • Provide stakeholder visibility across the SRE process
  • Collaborate cross-functionally to implement system enhancements for scalability and stability
  • Build client-focused dashboards/alerts to proactively detect performance challenges
  • Monitor and continuously improve time-to-resolution metrics
  • Maintain and configure observability tooling for incident response and performance investigations
  • Create an actionable feedback loop to Observability and Engineering teams to improve MELT and development patterns
  • Contribute to automation tools to streamline incident response
  • Proactively prevent incidents and reduce platform impact
  • Partner with Cloud Operations, SRE, Engineering, and business teams to advance SaaS platforms

Requirements

  • 5+ years in a Site Reliability Engineering role
  • Strong knowledge of SRE best practices and incident management protocols
  • Deep experience using and/or configuring observability tools such as:
    • New Relic, Datadog, Sumo Logic (or similar)
  • Proficiency reading/writing code (e.g., JavaScript, .NET, SQL)
  • Familiarity with cloud platforms (e.g., AWS, Azure) and architectural patterns
  • Data-driven, strong problem-solving skills for incident analysis
  • Experience operating in public cloud environments (AWS strongly preferred)
  • Ability to troubleshoot C#/.NET web applications for bugs and performance issues
  • Solid knowledge of SaaS operations
  • Comfort working with ambiguity across varying levels of operational maturity
  • Strong written and verbal communication skills

Nice to have

  • Windows and SQL Server troubleshooting experience
  • Experience with CI/CD pipelines
  • Experience with Infrastructure as Code (IaC) environments

About Origami Risk

Origami Risk provides SaaS platforms focused on risk-related workflows. The company operates public-cloud services at scale and emphasizes site reliability, observability, and fast incident recovery to support customer delivery.

Scraped 5/13/2026