xelys jobs xelys jobs

Senior Site Reliability Engineer

Mondo

full-remoteseniorcontractdevopsbackend Horsham, PA 79 days ago via LinkedIn
728,000 - 1,044,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)AWSDistributed SystemsObservabilitySLIs/SLOsIncident ManagementCI/CDAutomationInfrastructure as CodeServiceNow

About the role

Role Overview

Senior Site Reliability Engineer (Senior SRE) working 100% remotely (contract-to-hire, start ASAP). You will improve reliability and operational excellence across distributed systems using automation, observability, and engineering best practices.

Responsibilities

  • Lead reliability and recovery planning for critical systems and services
  • Define and maintain SLIs, SLOs, and error budget practices
  • Drive incident response and lead complex outage investigations
  • Perform root cause analysis and implement corrective actions
  • Build automation to reduce operational toil
  • Improve CI/CD pipelines, deployment workflows, and rollback strategies
  • Develop and maintain observability tooling (metrics, logs, tracing)
  • Partner with release/change management teams on release readiness
  • Mentor engineers and promote operational excellence across teams
  • Influence system architecture to improve reliability and scalability

Requirements

Must-haves

  • Bachelor’s degree in CS, Engineering, or related field
  • 6–10 years experience in SRE, DevOps, platform engineering, or software engineering
  • Strong experience with AWS and distributed systems
  • Root cause analysis and incident management experience
  • Experience documenting SRE systems and operational processes
  • Strong programming skills across multiple languages
  • Observability/monitoring experience (metrics, logs, tracing)
  • CI/CD pipelines and deployment automation experience
  • Familiarity with containerization and orchestration technologies
  • Strong troubleshooting and problem-solving skills
  • Ability to work in regulated/enterprise environments

Nice-to-haves

  • Infrastructure as Code (e.g., CloudFormation)
  • Experience with Spring/Spring Boot, React, or Angular
  • Experience with Tomcat, Netty, Node.js, or Next.js
  • Relational and non-relational databases
  • Agile/Scrum experience
  • ITSM tools such as ServiceNow
  • ITIL-based release/change management familiarity
  • Security compliance framework knowledge (e.g., ISO 27001, SOC 2)

About Mondo

Mondo is a technology and services company providing solutions for enterprise customers, with work spanning distributed systems, operations, and engineering best practices. The role focuses on reliability, scalability, and operational excellence in cloud-based environments.

Scraped 5/12/2026