Site Reliability Engineer ll
Cohere Health
full-remoteseniordevopsbackend United States 8 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)AWS LambdaPySparkNode.jsPythonTerraformHIPAASOC 2HITRUSTObservability
About the role
Role Overview
Operational-focused Site Reliability Engineer (SRE) for a remote-first organization. You will bridge AWS cloud infrastructure, MERN stack applications, and large-scale data workflows, maximizing availability, performance, and resilience.
You’ll spend approximately:
- 60% on live operations: incident remediation, data pipeline operations, and Node.js/Python infrastructure tuning
- 40% on engineering automation to reduce operational toil
Responsibilities
- Production Operations: Maintain uptime, scalability, and security of AWS-hosted MERN applications and backend data architectures.
- Serverless Execution: Run and optimize event-driven architectures on AWS Lambda (cold-start mitigation, memory allocation, execution timeouts).
- Data Pipeline Execution: Monitor and operate scheduled PySpark workflows; triage and rerun/patch failed data processing jobs.
- Incident Management: Participate in an on-call rotation to triage, debug, and mitigate application and data flow outages.
- Healthcare Compliance: Maintain HIPAA, SOC2, HITRUST compliance across runtime environments, storage, and PHI-handling pipelines.
- Toil Elimination: Automate repetitive tasks (e.g., manual data seeding, infrastructure provisioning, PySpark pipeline recovery steps).
- Observability Engineering: Build dashboards/alerts for Node.js event loops, PySpark stages, memory leak indicators, and pipeline throughput anomalies.
- Post-Mortem Culture: Lead blameless post-mortems and convert failures into permanent structural fixes.
Requirements
- 3+ years operating multi-tenant, cloud-hosted/cloud-native SaaS platforms at scale.
- AWS expertise across: Lambda, ECS/EKS, EMR or Glue (Spark), EC2, VPC, IAM, and CloudWatch.
- Strong automation and scripting in Python (incl. PySpark APIs) and Node.js.
- Experience with distributed orchestration/ETL pipelines and tools such as message queues (AWS SQS/SNS, RabbitMQ) or stream processing frameworks.
- Deep operational understanding of JavaScript/TypeScript applications (memory management, async runtimes, Node.js clustering).
- Production experience with MySQL and Athena (RDS or self-hosted).
- Infrastructure as Code using Terraform or OpenTofu (immutable infrastructure).
- 1+ year working in HIPAA-regulated environments; securing data-at-rest and data-in-transit with PHI is preferred.
- 4+ years overall software/systems experience; 1–2+ years focused on live cloud operations and distributed data workflow management preferred.
- Ability to stay calm during outages; methodical troubleshooting (e.g., PySpark driver OOM, data corruption) and effective communication with stakeholders.
Additional Notes
- Remote-first role with potential travel to Boston, MA for onboarding and occasional in-person meetings/events.
About Cohere Health
Cohere Health is a healthcare technology company focused on improving healthcare outcomes through technology-enabled care coordination and data-driven operations. The role centers on running and securing production systems that support healthcare workflows and Protected Health Information (PHI).
Scraped 7/22/2026