Senior Site Reliability Engineer
Juul Labs
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role overview
As a Senior Site Reliability Engineer (SRE), you will own operational stability and performance for Juul’s hybrid cloud infrastructure across Nutanix and AWS/GCP. You will lead reliability-focused automation, architect for resilience, and serve as the final escalation point for critical incidents to keep the platform scalable and efficient.
Responsibilities
- Lead automation and reliability architecture for hybrid cloud operations (Nutanix + AWS/GCP)
- Act as the L3 escalation point for critical incidents and troubleshoot complex issues
- Design and maintain enterprise-scale Nutanix environments:
- Deploy and manage Nutanix AHV clusters and Prism Central for multi-cluster management
- Use Nutanix CLI (nCLI/aCLI) for advanced operations, troubleshooting, and automation
- Build infrastructure automation using:
- Nutanix REST APIs, Python SDK, PowerShell, and Terraform (IaC)
- Create/manage VM templates, golden images, and standardized provisioning catalogs
- Design and implement disaster recovery solutions (Nutanix Leap, protection domains, cross-cluster replication, metro clustering)
- Implement network and security controls:
- Micro-segmentation using Nutanix Flow
- Configure RBAC, encryption, and security hardening
- Manage performance and availability for mission-critical workloads:
- Configure HA, VM affinity, QoS policies, and optimize performance
- Manage AHV networking (OVS bridges, VLANs, bonds, LACP, workload/resource balancing)
- Architect multi-cloud infrastructure across Nutanix HCI, AWS, and GCP with HA and DR
- AWS/GCP platform engineering:
- Multi-region/multi-account strategy for AWS and GCP
- Organization hierarchy, centralized governance, and IAM policy design
- Centralized logging and security integrations (CloudWatch/CloudTrail, GCP Cloud Logging, S3 aggregation, Splunk via HEC)
- Build and manage advanced load balancing:
- AWS ALB/NLB/ELB and GCP Cloud Load Balancing (SSL termination, auto-scaling)
- Deliver repeatable IaC and CI/CD:
- Terraform, CloudFormation, CDK
- Identity & access management:
- AWS SSO, cross-account IAM roles, GCP Workload Identity, federated access
- Hybrid connectivity:
- AWS Transit Gateway / PrivateLink; GCP Shared VPC / VPC peering
- Container workload operations:
- EKS, GKE, ECS, Cloud Run with service mesh, observability, and security best practices
- Disaster recovery on AWS/GCP:
- AWS Backup, cross-region replication, GCP snapshots, multi-region failover strategies
Requirements
- Expert proficiency in Nutanix platform management (AHV, Prism Central) and advanced Nutanix CLI operations
- Strong automation and infrastructure engineering experience across AWS and GCP
- Hands-on ability to architect and implement HA, scalability, and disaster recovery
- Advanced troubleshooting skills and L3 incident ownership
Nice-to-haves (implied by the posting)
- Experience integrating security/logging stacks, including Splunk and SIEM correlation
- Experience with container platforms and service mesh patterns
- Strong IaC/automation depth across multi-cloud tooling
About Juul Labs
Juul Labs’ mission is to transition the world’s adult smokers away from combustible cigarettes, eliminate their use, and help combat underage usage of its products. The company emphasizes quality, research, design, and innovation, backed by leading technology investors. It operates across technology and product functions to deliver reliable, high-performing systems and experiences.
Scraped 6/13/2026