ML Infrastructure Engineer
Mach9
midpermanentbackenddata San Francisco, CA Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine Learning InfrastructureData VersioningDataset LineageML Pipeline OrchestrationCI/CDModel ServingInference OptimizationPyTorchPythonAWS
About the role
Role Overview
Mach9 is seeking an ML Infrastructure Engineer to build and maintain the systems powering production AI models for civil engineering and surveying. You will work across the ML lifecycle—training pipelines, data/lineage management, and real-time inference—including integration with CAD software.
Responsibilities
- Design and build a centralized system for versioning training data, generated datasets, and model artifacts, including end-to-end lineage tracking.
- Develop and maintain reliable, reproducible ML training and data generation pipelines.
- Refactor and harden existing training/data generation scripts into composable, testable, maintainable components.
- Create CI/CD workflows to validate data pipelines and training runs with automated correctness checks and regression detection.
- Build tooling to let ML engineers launch, monitor, and debug training jobs with minimal friction.
- Optimize and scale real-time inference services (profiling, batching strategies, resource-efficient serving) to meet latency/throughput targets.
- Own deployment from trained artifacts to production endpoints, ensuring rollouts, rollbacks, and monitoring.
Requirements
- 3+ years of relevant work experience.
- BS/MS in Computer Science, Engineering, or equivalent experience.
- Strong communication skills and ability to partner with ML researchers/engineers to translate workflows into robust systems.
- Experience with data versioning, artifact management, and dataset lineage (e.g., DVC, LakeFS, Weights & Biases, or similar).
- Hands-on experience with ML pipeline orchestration (e.g., Airflow, Prefect, Metaflow, or similar).
- Experience with model serving and inference optimization (latency profiling, memory reduction, scaling for real-time constraints).
- Ability to read and refactor ML training code to make pipelines reliable (not required to design model architectures).
- Proficient in Python and PyTorch.
Bonus Qualifications
- Familiarity with AWS.
- Experience with containerized ML workflows and GPU-accelerated training environments.
- Familiarity with model optimization (e.g., quantization, TensorRT, ONNX Runtime, distillation).
- Experience with infrastructure-as-code (e.g., AWS CDK, Terraform).
- Experience operating ML systems with large unstructured datasets (imagery, 3D, sensor data).
About Mach9
Mach9 builds and maintains production AI systems for civil engineering and surveying. Its ML infrastructure supports training and real-time inference using large-scale labeled survey data, including image segmentation and 3D prediction models.
Scraped 8/1/2026