AI Infrastructure / MLOps Engineer — NYC
LaStellar Group
midpermanentdevopsbackendproduct-management New York City Metropolitan Area 2 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
PythonTerraformAzure Key VaultGCP Secret ManagerCI/CDDockerKubernetesOpenTelemetryPolicy-as-CodePrefect
About the role
Role Overview
AI Infrastructure / MLOps Engineer (NYC) You will build and operate the production infrastructure that keeps live agentic AI and data pipelines reliable, scalable, and production-grade. This is an engineering and platform ownership role—not a research or modeling position.
What You’ll Own
AI Platform & Agent Operations
- Operate and scale live agentic AI systems across Azure and GCP (high availability, performance, resilience under load)
- Build and maintain observability for agent execution (logging, tracing, alerting, performance monitoring)
- Support agent integration with data platforms and Model Context Protocol (MCP) servers
- Implement auto-scaling for containerized agent workloads across:
- Azure Container Apps
- GCP Cloud Run
- GKE
- Contribute to evaluation frameworks and production quality standards for AI agents
MLOps & Python Engineering
- Write and improve production Python powering data pipelines, agent workflows, and platform tooling
- Own the full lifecycle of Python services: containerization, deployment, versioning, runtime behavior
- Orchestrate workflows with Prefect (scheduling, error handling, retries, human-in-the-loop patterns)
- Build shared Python tooling/internal packages to help data science teams deploy faster
Cloud Infrastructure & CI/CD
- Write and maintain Terraform across Azure and GCP (container registries, managed identities, Key Vault, Secret Manager, storage backends, VNet configurations)
- Build CI/CD pipelines and release management workflows across repositories
- Enforce coding standards, security policies, and compliance controls in the pipeline
- Ensure production systems have documentation including runbooks and data lineage
Observability & Reliability
- Build and own the observability stack: metrics, logging, distributed tracing, alerting
- Drive SLO/SLI frameworks and incident response as the platform matures
- Troubleshoot production issues end-to-end (application logic through infrastructure)
Required Qualifications
- 3–5 years of software engineering / DevOps / MLOps / platform engineering with clear production ownership
- Strong Python engineering (production-grade code, packaging, containerization, dependency management)
- Hands-on Docker and container orchestration experience on Azure and/or GCP
- Terraform across cloud providers (designed it, not just configured it)
- Secrets management experience: Azure Key Vault and GCP Secret Manager (runtime injection patterns)
- CI/CD and Git-based release management
- Systems thinking for end-to-end troubleshooting
- Curiosity about AI and agentic systems (platform concepts)
- Policy-as-code enforcement in CI/CD: OPA or equivalent
- Observability depth: familiarity with OpenTelemetry as a protocol/specification
- Experience with:
- Azure Container Apps / ACI / ACR / Managed Identities / VNets
- GCP Cloud Run / GKE / Vertex AI / IAM / Secret Manager
- Familiarity with agentic frameworks: MCP, LangChain, or similar
- AI observability platforms: Langfuse, MLflow, or similar
- Data transformation/warehousing tools: dbt, Snowflake, or similar
Nice to Have
- Prefect (or similar) workflow orchestration in production
- Multi-cloud networking and identity management experience
- Financial services/fintech domain exposure
About LaStellar Group
LaStellar Group is a fast-growing fintech and investment platform operating at the intersection of AI and financial markets. It runs production AI systems and a scaling data platform, requiring reliable, production-grade AI infrastructure and MLOps capabilities.
Scraped 7/30/2026