xelys jobs xelys jobs

Platform Engineer, Model Shaping

Together AI

midpermanentbackenddevops San Francisco, CA 32 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

PythonGoLinuxKubernetesTerraformAnsiblePrometheusGrafanaCI/CDGPU Workloads

About the role

Role Overview

Platform Engineer on the Model Shaping team at Together AI. You’ll build the backend and infrastructure layers that enable model customization and evaluation workflows for both production users and internal research experiments.

Responsibilities

  • Design, develop, and operate systems and infrastructure for model customization (user-facing and internal improvements)
  • Improve platform reliability, participate in an on-call rotation, and enhance incident response processes
  • Build and improve internal tooling for deployment, CI/CD, and observability
  • Create a job orchestration platform spanning multiple datacenters and heterogeneous hardware
  • Collaborate with engineering and research teams to co-design internal services and integrate them into broader Together systems (e.g., Data Engineering, Cloud Infrastructure, Commerce)

Requirements

  • 3+ years building infrastructure or backend components for production services
  • Strong experience designing/operating/troubleshooting production Linux environments and Kubernetes-based platforms
  • Strong software engineering background in Python or Go
  • Experience with Terraform/Ansible, Prometheus/Grafana, and CI/CD (GitHub Actions, ArgoCD)
  • Cloud administration experience with AWS/GCP/Azure (preferably with hybrid bare-metal/cloud)
  • Strong communication skills; comfortable documenting systems and collaborating cross-functionally
  • Ability to operate across the stack, from cluster/infrastructure automation to backend service development

Nice to Have

  • Experience building large-scale, high-reliability production systems
  • Pipeline orchestration frameworks: Kubeflow, Argo Workflows, Flyte
  • Operating GPU workloads on HPC clusters; familiarity with NVIDIA networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA)
  • Networking fundamentals (TCP/IP, DNS, routing, load balancing, TLS) and network debugging
  • Experience maintaining or contributing to open-source projects

About Together AI

Together AI is a research-driven artificial intelligence company building open and transparent AI systems. The company aims to significantly lower the cost of modern AI by co-designing software, hardware, algorithms, and models, and has contributed to leading open-source research, models, and datasets.

Scraped 6/24/2026