Site Reliability Engineer
Origami Risk
hybridseniorpermanentdevopsbackend United States 74 days ago via LinkedIn
100,000 - 120,000 USD/annual
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)Incident ManagementObservabilityNew RelicDatadogSumo LogicAWSAzureCI/CDInfrastructure as Code (IaC)
About the role
Role overview
As a Site Reliability Engineer, you will improve Origami’s time to resolution and drive site reliability, scalability, and stability. You’ll lead post-incident investigations, identify root causes, and implement preventive measures. You’ll also analyze client performance issues and help track key performance metrics across client implementations.
Responsibilities
- Lead post-incident investigations for the Site Reliability team
- Perform root-cause analyses (RCAs) and develop preventive strategies
- Draft clear, insightful RCAs for customer delivery
- Cross-train others on using observability tools during incident and performance investigations
- Provide stakeholder visibility across the SRE process
- Collaborate cross-functionally to implement system enhancements for scalability and stability
- Build client-focused dashboards/alerts to proactively detect performance challenges
- Monitor and continuously improve time-to-resolution metrics
- Maintain and configure observability tooling for incident response and performance investigations
- Create an actionable feedback loop to Observability and Engineering teams to improve MELT and development patterns
- Contribute to automation tools to streamline incident response
- Proactively prevent incidents and reduce platform impact
- Partner with Cloud Operations, SRE, Engineering, and business teams to advance SaaS platforms
Requirements
- 5+ years in a Site Reliability Engineering role
- Strong knowledge of SRE best practices and incident management protocols
- Deep experience using and/or configuring observability tools such as:
- New Relic, Datadog, Sumo Logic (or similar)
- Proficiency reading/writing code (e.g., JavaScript, .NET, SQL)
- Familiarity with cloud platforms (e.g., AWS, Azure) and architectural patterns
- Data-driven, strong problem-solving skills for incident analysis
- Experience operating in public cloud environments (AWS strongly preferred)
- Ability to troubleshoot C#/.NET web applications for bugs and performance issues
- Solid knowledge of SaaS operations
- Comfort working with ambiguity across varying levels of operational maturity
- Strong written and verbal communication skills
Nice to have
- Windows and SQL Server troubleshooting experience
- Experience with CI/CD pipelines
- Experience with Infrastructure as Code (IaC) environments
About Origami Risk
Origami Risk provides SaaS platforms focused on risk-related workflows. The company operates public-cloud services at scale and emphasizes site reliability, observability, and fast incident recovery to support customer delivery.
Scraped 5/13/2026