xelys jobs xelys jobs

Site Reliability Engineer

NationsBenefits

full-remotemidpermanentbackend United States 9 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringKubernetesDockerDatadogObservabilityCI/CDIncident ManagementAutomationHIPAASOC 2

About the role

Role Overview

Site Reliability Engineer II (SRE) for NationsBenefits’ Site Reliability Engineering team. You’ll help ensure reliability, availability, and performance of production platforms by monitoring system health, responding to incidents, troubleshooting Kubernetes-based environments, and partnering with Development, DevSecOps, and Engineering teams.

Responsibilities

  • Production Support & Incident Management
    • Act as first line of response for production incidents
    • Monitor, triage, troubleshoot, and resolve production issues
    • Perform initial root cause analysis and escalate when needed
    • Communicate incident updates to stakeholders
  • Monitoring & Platform Reliability
    • Monitor infrastructure/application health using Datadog (or similar)
    • Optimize alerts to reduce false positives
    • Troubleshoot Kubernetes workloads (pods, deployments, logs, rollbacks)
    • Maintain high availability and performance
  • Collaboration
    • Work with Development, DevSecOps, Infrastructure, and Engineering on incident resolution
    • Participate in cross-functional troubleshooting
    • Recommend improvements to monitoring, tooling, and operational processes
    • Collaborate with global teams across multiple time zones
  • Automation & Continuous Improvement
    • Build automation scripts/tools using Python, PowerShell, Bash, C#, Java
    • Support CI/CD pipeline monitoring and deployment reliability
    • Contribute to self-healing and automated recovery
  • Documentation & Compliance
    • Maintain incident documentation and post-mortems
    • Follow security/compliance standards including HIPAA, PCI DSS, SOC 2, ISO 27001, HITRUST
  • On-Call & Support Rotation
    • Participate in weekday production support in a global “follow-the-sun” model
    • Join on-call rotation for critical systems as needed

Required Qualifications

  • 3–5 years experience in SRE, DevOps, Production Support, or Platform Engineering
  • Hands-on production incident management and troubleshooting
  • Experience with Datadog or similar observability tools
  • Strong experience supporting Kubernetes and Docker
  • Experience with SQL/MySQL or other NoSQL databases
  • Familiarity with cloud platforms: Azure, AWS, or GCP
  • Experience with high-availability production environments
  • Strong troubleshooting/analytical and communication skills
  • Ability to work weekday shifts in a global follow-the-sun support model

Preferred Qualifications

  • Experience with CI/CD pipelines and deployment automation
  • Knowledge of Helm Charts
  • Familiarity with ITIL and Agile methodologies
  • Scripting/programming experience: Python, PowerShell, Bash, C#, Java
  • Familiarity with security and compliance standards in Healthcare/FinTech

About NationsBenefits

NationsBenefits is a healthcare FinTech company that delivers supplemental benefits, flex card solutions, and member engagement platforms for managed care organizations. It helps health plans improve member outcomes and reduce healthcare costs through secure, scalable, and compliance-driven technology delivered by teams across the US, South America, and India.

Scraped 7/17/2026