xelys jobs xelys jobs

Site Reliability Engineer ll

Cohere Health

full-remoteseniordevopsbackend United States 8 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)AWS LambdaPySparkNode.jsPythonTerraformHIPAASOC 2HITRUSTObservability

About the role

Role Overview

Operational-focused Site Reliability Engineer (SRE) for a remote-first organization. You will bridge AWS cloud infrastructure, MERN stack applications, and large-scale data workflows, maximizing availability, performance, and resilience.

You’ll spend approximately:

  • 60% on live operations: incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning
  • 40% on engineering automation to reduce operational toil

Responsibilities

  • Production Operations: Maintain uptime, scalability, and security of AWS-hosted MERN applications and backend data architectures.
  • Serverless Execution: Run and optimize event-driven architectures on AWS Lambda (cold-start mitigation, memory allocation, execution timeouts).
  • Data Pipeline Execution: Monitor and operate scheduled PySpark workflows; triage and rerun/patch failed data processing jobs.
  • Incident Management: Participate in an on-call rotation to triage, debug, and mitigate application and data flow outages.
  • Healthcare Compliance: Maintain HIPAA, SOC2, HITRUST compliance across runtime environments, storage, and PHI-handling pipelines.
  • Toil Elimination: Automate repetitive tasks (e.g., manual data seeding, infrastructure provisioning, PySpark pipeline recovery steps).
  • Observability Engineering: Build dashboards/alerts for Node.js event loops, PySpark stages, memory leak indicators, and pipeline throughput anomalies.
  • Post-Mortem Culture: Lead blameless post-mortems and convert failures into permanent structural fixes.

Requirements

  • 3+ years operating multi-tenant, cloud-hosted/cloud-native SaaS platforms at scale.
  • AWS expertise across: Lambda, ECS/EKS, EMR or Glue (Spark), EC2, VPC, IAM, and CloudWatch.
  • Strong automation and scripting in Python (incl. PySpark APIs) and Node.js.
  • Experience with distributed orchestration/ETL pipelines and tools such as message queues (AWS SQS/SNS, RabbitMQ) or stream processing frameworks.
  • Deep operational understanding of JavaScript/TypeScript applications (memory management, async runtimes, Node.js clustering).
  • Production experience with MySQL and Athena (RDS or self-hosted).
  • Infrastructure as Code using Terraform or OpenTofu (immutable infrastructure).
  • 1+ year working in HIPAA-regulated environments; securing data-at-rest and data-in-transit with PHI is preferred.
  • 4+ years overall software/systems experience; 1–2+ years focused on live cloud operations and distributed data workflow management preferred.
  • Ability to stay calm during outages; methodical troubleshooting (e.g., PySpark driver OOM, data corruption) and effective communication with stakeholders.

Additional Notes

  • Remote-first role with potential travel to Boston, MA for onboarding and occasional in-person meetings/events.

About Cohere Health

Cohere Health is a healthcare technology company focused on improving healthcare outcomes through technology-enabled care coordination and data-driven operations. The role centers on running and securing production systems that support healthcare workflows and Protected Health Information (PHI).

Scraped 7/22/2026