Software Engineer (Infrastructure & Reliability)
CrewAI
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreAbout the role
Join CrewAI as a Software Engineer focused on Infrastructure & Reliability. In this role, you will build and operate the platform infrastructure for CrewAI's cloud and enterprise deployments. You will work across multiple hyperscalers, including AWS, Azure, and GCP, and will be responsible for improving systems, designing deployment paths, and hardening production. This is not a pure DevOps support role; you will write code, improve systems, and build the internal platform that allows CrewAI to scale. Key missions: Construire et exploiter l'infrastructure de la plateforme derrière les déploiements cloud et d'entreprise de CrewAI.. Écrire du code, améliorer les systèmes, concevoir des chemins de déploiement, renforcer la production et construire la plateforme interne qui permet à CrewAI de se développer.. Améliorer la fiabilité des déploiements cloud et d'entreprise : vérifications de santé, alertes, réponse aux incidents, planification de la capacité, chemins de récupération et manuels opérationnels. Profile: - Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access - Strong debugging instincts across app, infra, network, deploy, and dependency layers - Ability to write reliable automation in Python, Ruby, Go, Bash, or similar - Calm, rigorous approach to incidents, rollbacks, migrations, and production change management - Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services - Experience with ECS and/or Kubernetes; Helm experience is a strong plus - Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production - Strong infrastructure/platform engineering experience in production SaaS environments - Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems - Experience supporting enterprise/self-hosted deployments - Terraform or other IaC experience - Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability - SRE background: SLOs, incident review, capacity planning, load testing
Scraped 9/25/2026