Site Reliability Engineer
NationsBenefits
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Site Reliability Engineer II (SRE) responsible for reliability, availability, and performance of production platforms. You will monitor system health, respond to incidents, troubleshoot Kubernetes-based environments, and collaborate with Development, DevSecOps, and Engineering teams.
Responsibilities
- Production Support & Incident Management
- Act as first line of response for production incidents
- Monitor, triage, troubleshoot, and resolve production issues
- Perform initial root-cause analysis and escalate appropriately
- Communicate incident updates to stakeholders
- Monitoring & Platform Reliability
- Use Datadog (or similar) observability tools to monitor infrastructure/application health
- Optimize alerting to reduce false positives
- Troubleshoot Kubernetes workloads (pods, deployments, logs, rollbacks)
- Maintain high availability and performance
- Collaboration
- Partner with Development, DevSecOps, Infrastructure, and Engineering for cross-functional troubleshooting
- Recommend improvements to monitoring, tooling, and operational processes
- Collaborate with global teams across time zones
- Automation & Continuous Improvement
- Build automation scripts and operational tools using Python, PowerShell, Bash, C#, Java
- Support CI/CD pipeline monitoring and deployment reliability
- Contribute to self-healing and automated recovery
- Documentation & Compliance
- Maintain incident documentation and post-mortems
- Follow security and compliance standards including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST
- On-Call & Support Rotation
- Participate in a weekday production support rotation in a follow-the-sun model
Requirements
- 3–5 years of experience in SRE, DevOps, Production Support, or Platform Engineering
- Hands-on experience with production incident management and troubleshooting
- Experience with Datadog (or similar observability tools)
- Strong support experience for Kubernetes and Docker
- Experience with SQL, MySQL, or NoSQL databases
- Familiarity with Azure, AWS, or GCP
- Ability to work weekday shifts in a global follow-the-sun support model
- Strong troubleshooting, analytical, and communication skills
Preferred Qualifications
- Experience with CI/CD pipelines and deployment automation
- Knowledge of Helm Charts
- Familiarity with ITIL processes and Agile methodologies
- Scripting/programming experience using Python, PowerShell, Bash, C#, Java
- Familiarity with security and compliance standards in Healthcare or FinTech
About NationsBenefits
NationsBenefits is a fast-growing Healthcare FinTech company that delivers supplemental benefits, flex card solutions, and member engagement platforms for managed care organizations. Its technology helps health plans improve member outcomes, reduce healthcare costs, and address social determinants of health through secure, scalable, compliance-driven solutions.
Scraped 7/27/2026