Staff AI Systems Engineer
Flock Safety
full-remoteleadpermanentbackenddata Full remote 134 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine LearningLLM EvaluationAgentic AIObservabilityKubernetesNVIDIA TritonPyTorchTensorRTLangChainAWS
About the role
Role overview
Join Flock Safety as a Staff AI Systems Engineer and help develop Night Shift, an AI copilot that improves investigators’ efficiency. You’ll work closely with the Machine Learning team and engineering partners to design agentic AI systems and make them measurable, dependable, and fast.
Key missions
- Contribute to system design and architecture for agentic AI.
- Own the AI evaluation framework.
- Deliver the MVP of the evaluation framework to generate initial metrics, enable debugging, and run regression tests.
- Productionalize the evaluation and observability platform so it becomes the source of truth for quality and safety.
Responsibilities (what you’ll drive)
- Improve measurable outcomes such as lead accuracy and speed for law enforcement officers.
- Enable robust evaluation, debugging, and regression workflows for LLM/agent performance.
- Ensure production readiness with strong observability, quality, and safety practices.
Requirements
- 5+ years building and shipping ML/LLM systems to production.
- Data & storage: ClickHouse, Postgres, Redis.
- Observability: Prometheus, Grafana, OpenTelemetry, LangSmith/Langfuse.
- ML inference: PyTorch, TensorRT, NVIDIA Triton (preferably multimodal: text/image/video).
- Web services: Express/FastAPI, REST, SSE, JWTs.
- Backend JS familiarity (Node.js); TypeScript and Python familiarity welcomed.
- Compute orchestration: Kubernetes, Prefect, Ray.
- LLM inference & tooling: LangChain/LangGraph, vLLM, and LLM APIs (OpenAI/Gemini/Anthropic).
- Agent design: tool use (via MCP), retrieval, memory, grounding/attribution, guardrails.
- Safety & robustness: security, compliance, red-teaming, regression testing.
- Ability to reason about cost/performance/latency trade-offs.
Nice to have / additional experience
- LLM evaluations at scale: offline/online eval harnesses; familiarity with eval methodologies and metrics.
- RAG with vector/hybrid search (e.g., pgvector, Chroma, turbopuffer) and re-rankers (e.g., Cohere, JinaAI).
- Agentic success metrics: trajectory quality, preference learning (SFT, DPO, RLHF, LLM-as-judge).
- Architectural patterns for multi-agent planning, hand-off, and context management.
Location
- Full remote
About Flock Safety
Flock Safety builds AI-powered solutions for public safety and law enforcement investigations. The company develops technology such as an AI copilot to help investigators work more efficiently by improving accuracy, speed, and overall system quality.
Scraped 5/12/2026