xelys jobs xelys jobs

Senior Site Reliability Engineer

Filevine

full-remoteseniorpermanentdevopsbackendsecurity Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)AWSKubernetesCI/CDMonitoringIncident ResponseAutomationPythonCloudWatchPerformance Optimization

About the role

Role Overview

Senior Site Reliability Engineer (SRE) at Filevine, embedded with a cross-functional team. You’ll lead reliability efforts for mission-critical parts of the platform, driving improvements across the SDLC and supporting production through an on-call rotation.

Responsibilities

  • Lead and mentor: Provide reliability engineering leadership, judgment, and guidance to the team (including mentoring junior engineers).
  • Design and maintain autonomous systems for:
    • building, deploying, testing, and operating Filevine products
    • minimizing human intervention through automation
  • Own reliability across the SDLC as the authoritative voice for reliability engineering.
  • Monitoring and incident support:
    • monitor software and infrastructure events
    • proactively identify and resolve gaps in availability, performance, and security
    • participate in a 24/7 on-call rotation for production support
  • CI/CD improvements: Enhance and maintain CI/CD pipelines to improve reliability and delivery outcomes.
  • Operational rigor: Document processes, collaborate with the team, and support reliability best practices (e.g., incident response, capacity planning, and performance optimization).

Requirements

  • 8+ years hands-on technical experience in software engineering, infrastructure, or operations, including at least 4 years in SRE.
  • Hands-on AWS expertise, including at least some of: EC2, Kubernetes/EKS, CloudWatch, Lambda, S3, IAM.
  • Python and scripting proficiency: Python, Bash, PowerShell, and other SRE automation/tooling.
  • Proven ability to independently drive reliability improvements, reduce toil via automation, and deliver high-availability, scalable production systems.
  • Expert-level ability to design/build/maintain autonomous systems spanning build → deploy → test → monitor → operate.

Nice-to-haves

  • Certifications (e.g., AWS or comparable cloud certifications like Google Cloud Professional) or equivalent education/experience.
  • Demonstrated continuous learning, self-motivation, and curiosity for process and system improvements.

About Filevine

Filevine is a technology company that builds software products for legal professionals, helping teams manage case workflows and related operations. The role sits within a reliability engineering function focused on keeping these products highly available, secure, and performant in production environments.

Scraped 5/12/2026