Senior Forward Deployed Platform Engineer
Recare
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreAbout the role
Join our team as a Senior Forward Deployed Platform Engineer, where you will play a crucial role in ensuring the reliability and performance of our AI portfolio. You will be responsible for monitoring, capacity planning, and supporting scaling initiatives, as well as assisting AI engineers with productionalizing AI workloads. This is a cross-functional role that requires strong communication skills and the ability to collaborate with other teams. You will work with a modern cloud-native stack in a regulated healthcare environment, with the flexibility to work remotely or from our office in Berlin. Key missions: As a Senior Forward Deployed Platform Engineer, your main responsibilities will include ensuring the reliability and performance of the AI portfolio, planning capacity and supporting scaling initiatives, and continuously delivering SageMaker models while following AWS best practices. Profile: - Experience deploying SageMaker models, setting up custom inference containers (HuggingFace and/or OSS model) and endpoints (provisioning, autoscaling, and blue/green rollouts) - This is a cross-functional role bridging the gap between applications and the platform. Ability to both explain complex concepts clearly to non-domain experts, and extract information through well-formed questions during verbal communication, is a must - Knowing how to set up ML/LLM GPU inference on EKS (GPU-backed instances, Nvidia driver/plugin) - Experience with Kubernetes (EKS) in production (kustomize, HPA and KEDA based autoscaling) - Experience with AWS (CloudFormation, IAM, ECR, SageMaker, Bedrock, S3, SQS, DynamoDB, RDS, KMS) - Ability to own technical decisions and collaborate directly with other teams - Familiarity with observability & LLMOps best practices and solutions (Datadog, Langfuse, LLM/GenAI tracing, OTel, CloudWatch) - Preference for writing over talking, resulting in clear, concise, yet complete documentation and asynchronous communications - Proficiency in Python, specifically with FastAPI and async task management - Troubleshooting, RCA, and resolving problems. When things get stressful, you stay cool headed and work through issues step by step, asking for support from SMEs when needed - Healthcare or regulated-domain experience (e.g. ISO 27001, C5) - Self-hosting/home lab - Nvidia Triton/vLLM or similar - MLflow, Kubeflow, ClearML, DVC, Vertex AI - Slack and GitHub automations (bots, workflows, apps, etc.)
Scraped 9/25/2026