xelys jobs xelys jobs

DevOps Engineer

Careflow

full-remoteseniorpermanentdevops Miami, FL 48 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Google Cloud Platform (GCP)DevOpsSite Reliability EngineeringCI/CDObservabilityMonitoringLoggingIncident ResponseSecurityDisaster Recovery

About the role

Role Overview

Careflow is hiring an experienced DevOps Engineer to own and improve cloud infrastructure, security, observability, and operational reliability. You’ll help ensure the platform stays secure, scalable, performant, and highly available as it grows, with opportunities to troubleshoot production issues and occasionally dive into the application codebase.

Responsibilities

Cloud Infrastructure & Operations

  • Manage and maintain the Google Cloud Platform (GCP) environment
  • Design and improve infrastructure for scalability, reliability, and cost efficiency
  • Oversee networking, compute, databases, storage, and other cloud services
  • Monitor system health and proactively address performance bottlenecks

Monitoring, Logging & Observability

  • Build and maintain centralized logging and monitoring
  • Create dashboards and alerts for system health, application performance, and business-critical workflows
  • Establish operational metrics and usage tracking
  • Lead incident response and perform root cause analysis
  • Monitor and manage platform spend

Security & Compliance

  • Implement and maintain security best practices for infrastructure and applications
  • Manage identity/access controls, secrets management, and environment security
  • Conduct security reviews and support vulnerability remediation
  • Assist with compliance initiatives and audit readiness

CI/CD & Automation

  • Improve deployment pipelines and release processes
  • Automate infrastructure provisioning and operational workflows
  • Enhance deployment reliability for development environments
  • Reduce manual operational work via automation

Reliability Engineering

  • Improve uptime, resiliency, backups, and disaster recovery
  • Establish service-level objectives (SLOs) and operational standards
  • Drive improvements in platform stability and performance

Cross-Functional Support

  • Partner with engineering, product, and leadership
  • Provide technical guidance on infrastructure and operational considerations
  • Participate in an on-call / operational support rotation

Bonus (Application-Level)

  • Troubleshoot and fix application-level issues when needed
  • Contribute code improvements and bug fixes
  • Assist with performance optimization and debugging

Requirements

  • 5+ years in DevOps, Site Reliability Engineering, Cloud Engineering, or related experience
  • Strong hands-on experience with GCP
  • Experience building and maintaining CI/CD pipelines
  • Strong understanding of monitoring, logging, and alerting in cloud environments

Schedule / Availability

  • Fully remote
  • Flexible schedule with availability for Saturday coverage and taking another weekday off in exchange

Reporting Line & Structure

  • Reports to: Lead Architect

About Careflow

Careflow is a software company building a cloud-based platform that requires secure, scalable, and highly available operations. The role focuses on maintaining platform reliability and operational excellence, including infrastructure, security, observability, and incident response across the stack.

Scraped 6/16/2026