xelys jobs xelys jobs

Site Reliability Engineer II

NationsBenefits

full-remotemidpermanentdevops Plantation, FL 3 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSREDatadogKubernetesDockerCI/CDIncident ManagementAutomationHIPAASOC 2

About the role

Role: Site Reliability Engineer II (SRE)

You’ll join the Site Reliability Engineering team to help ensure the availability, reliability, and performance of NationsBenefits production platforms. You’ll monitor systems, respond to incidents, troubleshoot infrastructure issues, and drive automation—in collaboration with Development, DevSecOps, and Engineering teams.

Responsibilities

  • Incident Management
    • Act as first responder for production incidents (identify, triage, resolve)
    • Monitor and respond to alerts from Datadog and other monitoring tools
    • Perform initial root-cause analysis and escalate per SLAs
    • Provide incident status updates to stakeholders
  • Monitoring & Platform Reliability
    • Continuously monitor application health, infrastructure performance, and uptime
    • Configure/optimize dashboards and alert thresholds
    • Troubleshoot Kubernetes issues (pod failures, deployment rollbacks, log analysis)
    • Support containerized applications in Kubernetes and Docker environments
  • Production Support / On-call
    • Participate in a weekday Follow-the-Sun support model with global engineering teams
    • Join an on-call rotation for critical production systems
    • Help maintain high availability and system uptime
  • Automation & Continuous Improvement
    • Build automation scripts and operational tooling using one or more of:
      • Python, PowerShell, Bash, C#, Java
    • Support CI/CD pipeline monitoring and deployment reliability
    • Contribute to self-healing solutions and automation to reduce manual work
  • Collaboration
    • Work with Software Engineering, DevSecOps, Infrastructure, and Platform teams
    • Recommend improvements to monitoring, tooling, and operational processes
  • Documentation & Compliance
    • Maintain documentation for incidents, troubleshooting procedures, and post-incident reviews
    • Ensure alignment with security/compliance standards including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST

Requirements

  • 3–5 years of experience in Site Reliability / production support (required; posting cuts off after this section)

Nice-to-haves / Strong signals (from the posting)

  • Experience with Datadog monitoring
  • Hands-on troubleshooting in Kubernetes/Docker
  • Ability to automate with scripting/programming languages listed above
  • Familiarity with CI/CD reliability practices and operational documentation
  • Experience operating in environments governed by HIPAA/PCI/SOC2/ISO27001/HITRUST

About NationsBenefits

NationsBenefits is a healthcare fintech provider offering supplemental benefits, flex cards, and member engagement solutions. The company partners with managed care organizations to improve outcomes, reduce costs, and deliver compliant, high-quality healthcare benefits using fintech payment platforms and cloud-native infrastructure.

Scraped 7/30/2026