xelys jobs xelys jobs

Site Reliability Engineer

OneStream Software

full-remoteseniorpermanentdevopsbackend United States 48 days ago via LinkedIn
114,000 - 148,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSREAzureKubernetesAKSObservabilityDynatraceTerraformInfrastructure as CodeFedRAMP

About the role

Role Overview

As a Site Reliability Engineer (SRE) in OneStream Software’s Cloud Services, you will ensure the platform and services customers depend on are reliable, performant, and highly available. You’ll design, implement, and monitor scalable and secure cloud services, while partnering with Product and Engineering to improve reliability through automation and observability.

Primary Responsibilities

  • Implement application and infrastructure observability solutions to meet availability, reliability, and performance goals.
  • Participate in on-call rotations and communicate incident learnings via post-mortems and regular review meetings.
  • Partner with Product and Engineering to design, deploy, and maintain reliable systems and services.
  • Influence and create architectures, standards, and methods for large-scale systems.
  • Sustain reliability for key services and automated systems.
  • Automate processes to improve reliability, performance, and availability.
  • Maintain technical documentation and workflow/knowledge-base articles.
  • Provide feedback in pull requests and peer coding reviews.
  • Build codified automated integrations between Dynatrace, Azure DevOps, and Jira.
  • Mentor others across technical areas.
  • Apply practical SOC / FedRAMP controls to support Compliance and Security.

Required Qualifications

  • BS/BA in computer science/engineering/technology (or equivalent experience).
  • Proven experience as an SRE (or similar role).
  • 6+ years cloud infrastructure and software development experience.
  • 2+ years hands-on with Azure Kubernetes Services (AKS) for container-based deployments (or similar: OpenShift, GKE, EKS).
  • Advanced observability/APM experience with tools such as Dynatrace, Application Insights, Datadog, Log Analytics, New Relic, Prometheus, Grafana.
  • Advanced Infrastructure as Code (IaC) using tools such as Terraform, CloudFormation, Bicep, or ARM templates on Azure/AWS/GCP.
  • Deep configuration management/orchestration experience (e.g., Ansible, PowerShell DSC, Chef, Puppet).
  • Strong knowledge of cloud concepts including elasticity, security, and identity management.
  • Familiarity with Agile methodologies using Jira or Azure DevOps Boards.
  • 6+ years hands-on automation with PowerShell, Bash, CLI, REST APIs, Python, and ARM templates (or other scripting languages).
  • Experience with Git and Azure DevOps or GitHub.
  • Knowledge of container orchestration platforms (Kubernetes, OpenShift, AKS, GKE, helm).
  • Experience with Microsoft Azure, AWS, or GCP.

Preferred Qualifications

  • Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider.
  • 6+ years of relevant Azure experience deploying and managing infrastructure/services.

About OneStream Software

OneStream Software provides enterprise performance management (EPM) and financial consolidation solutions for organizations. The role sits within OneStream’s Cloud Services, focused on operating and scaling the cloud platform that supports customer-facing applications.

Scraped 6/17/2026