Site Reliability Engineer
Technatomy Corporation
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role Overview
Technatomy is seeking a Site Reliability Engineer to support the Technical Director’s team. The role focuses on reliability engineering, cloud operations, automation, and resilient service delivery for Department of Veterans Affairs enterprise healthcare platforms and applications. You will work with senior engineers, platform/operations teams, and VA stakeholders to maintain availability, performance, and operational stability.
Responsibilities
- Support day-to-day SRE activities across platform services, hosted applications, and cloud environments.
- Maintain service reliability, availability, and performance using runbooks and engineering standards.
- Review operational metrics, alerts, logs, and system health to identify issues and drive improvements.
- Own/maintain monitoring, logging, alerting, and dashboards for infrastructure and application visibility.
- Participate in incident response, including service restoration, escalation, and post-incident follow-up.
- Document incidents, recurring issues, procedures, configuration details, and troubleshooting guidance.
- Build simple scripts and automation to reduce manual effort and handle recurring operational tasks.
- Support CI/CD processes and environment maintenance across development, test, and production.
- Assist with Infrastructure as Code, configuration changes, and environment updates using approved tools/templates.
- Perform routine operational checks for AWS and container-based platforms.
- Maintain service inventory and related configuration/operational documentation artifacts.
- Assist with release validation, testing, deployment readiness, and operational acceptance.
- Follow security, access, change, and operational procedures supporting Federal compliance.
- Collaborate with engineering, infrastructure/platform, monitoring, incident-management, and support teams to improve reliability.
Requirements
- 1–3 years of experience in Site Reliability Engineering, DevOps, systems administration, cloud operations, platform support, software engineering, or related.
- Foundational understanding of Linux, cloud infrastructure concepts, enterprise application support, and basic networking.
- Scripting/programming exposure with Python, Bash, PowerShell, or similar.
- Familiarity with monitoring/logging/alerting, troubleshooting, incident response, and service restoration.
- Basic knowledge of CI/CD, version control, automation, configuration management, and/or Infrastructure as Code concepts.
- Ability to follow technical procedures, document accurately, analyze operational information, and escalate appropriately.
Nice-to-Haves
- Experience supporting AWS and container-based platforms.
- Demonstrated ability to improve monitoring/automation and reduce operational toil through scripting.
About Technatomy Corporation
Technatomy Corporation provides innovative solutions for government agencies and entities, including the Department of Veterans Affairs and Department of Defense. The company focuses on delivering mission-critical services in support of customer success, emphasizing reliability, security, and strong operational practices.
Scraped 7/16/2026