xelys jobs xelys jobs

Site Reliability Engineer

NationsBenefits

full-remotemidpermanentdevops United States Today via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringDatadogKubernetesDockerCI/CDMonitoringIncident ManagementHIPAASOC 2Terraform

About the role

Role Overview

Site Reliability Engineer II (SRE) responsible for reliability, availability, and performance of production platforms. You will monitor system health, respond to incidents, troubleshoot Kubernetes-based environments, and collaborate with Development, DevSecOps, and Engineering teams.

Responsibilities

  • Production Support & Incident Management
    • Act as first line of response for production incidents
    • Monitor, triage, troubleshoot, and resolve production issues
    • Perform initial root-cause analysis and escalate appropriately
    • Communicate incident updates to stakeholders
  • Monitoring & Platform Reliability
    • Use Datadog (or similar) observability tools to monitor infrastructure/application health
    • Optimize alerting to reduce false positives
    • Troubleshoot Kubernetes workloads (pods, deployments, logs, rollbacks)
    • Maintain high availability and performance
  • Collaboration
    • Partner with Development, DevSecOps, Infrastructure, and Engineering for cross-functional troubleshooting
    • Recommend improvements to monitoring, tooling, and operational processes
    • Collaborate with global teams across time zones
  • Automation & Continuous Improvement
    • Build automation scripts and operational tools using Python, PowerShell, Bash, C#, Java
    • Support CI/CD pipeline monitoring and deployment reliability
    • Contribute to self-healing and automated recovery
  • Documentation & Compliance
    • Maintain incident documentation and post-mortems
    • Follow security and compliance standards including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST
  • On-Call & Support Rotation
    • Participate in a weekday production support rotation in a follow-the-sun model

Requirements

  • 3–5 years of experience in SRE, DevOps, Production Support, or Platform Engineering
  • Hands-on experience with production incident management and troubleshooting
  • Experience with Datadog (or similar observability tools)
  • Strong support experience for Kubernetes and Docker
  • Experience with SQL, MySQL, or NoSQL databases
  • Familiarity with Azure, AWS, or GCP
  • Ability to work weekday shifts in a global follow-the-sun support model
  • Strong troubleshooting, analytical, and communication skills

Preferred Qualifications

  • Experience with CI/CD pipelines and deployment automation
  • Knowledge of Helm Charts
  • Familiarity with ITIL processes and Agile methodologies
  • Scripting/programming experience using Python, PowerShell, Bash, C#, Java
  • Familiarity with security and compliance standards in Healthcare or FinTech

About NationsBenefits

NationsBenefits is a fast-growing Healthcare FinTech company that delivers supplemental benefits, flex card solutions, and member engagement platforms for managed care organizations. Its technology helps health plans improve member outcomes, reduce healthcare costs, and address social determinants of health through secure, scalable, compliance-driven solutions.

Scraped 7/27/2026