xelys jobs xelys jobs

Senior Site Reliability Engineer - Observability

Optum

seniordevopssecurity United States 3 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

ObservabilitySite Reliability Engineering (SRE)DynatraceElasticElasticsearchOpenTelemetry (OTEL)PrometheusGrafanaTerraformAnsible

About the role

Role Overview

Senior Site Reliability Engineer (Observability) on the OptumServe Enterprise Monitoring team.

You will help maintain the reliability, scalability, and availability of the organization’s log management solution and metrics/observability platform, with strong emphasis on automation and performance/SLO management.

Responsibilities

  • Maintain and deploy monitoring and alerting systems
  • Design, configure, and maintain log aggregation at large scale
  • Set up and manage ingestion pipelines and data transformations
  • Adopt an “automate any task” mindset
  • Build and maintain robust monitoring/alerting to detect issues early and trigger timely alerts
  • Maintain documentation aligned to audit and certification requirements
  • Participate in troubleshooting, capacity planning, and performance analysis
  • Research new monitoring requirements and write code to implement them
  • Define and maintain monitoring policies/rules/templates
  • Develop scripts to meet monitoring requirements
  • Maintain performance KPIs and define SLOs
  • Medium-to-expert work creating AI rules for tools such as Dynatrace (DavisAI) and/or Elastic GenAI

Required Qualifications

  • 2+ years working directly with monitoring tools as an admin/SME/architect, preferably Dynatrace and/or Elasticsearch
  • 2+ years with Dynatrace (managed, cloud, and offline) including best practices and setup (ActiveGate, cloud, on-prem, custom workflows)
    • or 2+ years with Elastic (on-prem and cloud) with best practices
  • 1+ years designing data pipelines using Filebeat, Logstash, and/or Fluent Bit/Fluentd
  • 1+ years of AI expertise related to observability to reduce workload and improve reliability
  • 1+ years scripting to automate tasks using Python and Bash or PowerShell
  • 1+ years working with Linux OS
  • United States Citizenship

Preferred Qualifications

  • BS/MS in Computer Science/Engineering (or equivalent), or 5+ years experience
  • 1+ years with Terraform and Ansible (including complex configuration in multi commercial and Gov clouds)
  • 1+ years scripting (JavaScript, Java, PowerShell, or others)
  • Familiarity with SNMP, TCP dump, and tracing
  • Proven knowledge of AIOps platforms

About Optum

Optum is part of the Optum family of business within UnitedHealth Group, focused on simplifying complex healthcare-related logistics and workforce health programs. Logistics Health Incorporated (LHI), a part of Optum, provides cost-effective solutions and distribution processes, supporting on-prem, hybrid, and cloud-based environments.

Scraped 7/30/2026