Site Reliability Engineer
MyFitnessPal
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
MyFitnessPal is hiring a Software Engineer III – Site Reliability for the PEAS (Productivity Engineering: Automation & Self-Service) team within Technology Operations (TechOps). The team owns automation, CI/CD, and self-service platforms that help product teams build, test, and ship quickly and safely. As an SRE, you’ll ensure production services are fast, available, and safe to change.
Responsibilities
- Own and evolve reliability: define and improve SLI/SLO and error-budget frameworks to guide prioritization and product decisions.
- Incident management: lead incident response, drive postmortems, and convert findings into systemic fixes.
- Observability: build and maintain metrics, logs, and traces observability using Datadog, improving signal quality and reducing alert fatigue.
- Resilient infrastructure: design and operate scalable infrastructure using Infrastructure as Code (e.g., Terraform).
- Production Kubernetes operations: manage Kubernetes/container workloads, including capacity planning and cloud-cost optimization.
- CI/CD and deployment safety: own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback).
- Security in the delivery pipeline: implement and tune SAST/DAST/SCA scanning (e.g., in GitHub Actions) so issues are surfaced while code is still under review.
- Policy-as-code: implement admission-time guardrails using OPA/Rego, Kyverno, or Conftest.
- Vulnerability SLAs: drive vulnerability triage and remediation for pipeline/infrastructure findings based on real risk.
- On-call excellence: improve on-call sustainability via runbooks and automation.
- Mentorship: coach team members and engineers on reliability patterns and operational best practices.
Requirements
- 5+ years in site reliability, platform, or infrastructure engineering with senior-level ownership of production systems.
- Strong programming skills for automation/tooling (Go, Python, TypeScript, or similar), building software/custom tooling (not just scripts).
- Deep hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus).
- Proven experience leading incident response and building SLO-driven reliability practices.
- Fluency with observability tooling (Datadog is a plus).
- Practical experience integrating security into CI/CD (SAST/DAST/SCA, dependency scanning, and/or policy-as-code).
- Strong cloud security fundamentals (e.g., IAM, least-privilege, policy/guardrails).
Nice-to-haves / Additional Signals
- Experience with GitHub Actions security scanning.
- Familiarity with OPA/Rego, Kyverno, or Conftest for policy-as-code.
About MyFitnessPal
MyFitnessPal provides tools, resources, and support to help users reach their health goals. It operates a digital fitness and nutrition ecosystem that relies on reliable, secure software delivery and operations at scale.
Scraped 7/27/2026