Senior Software Engineer - Agent Evaluation - Freelance/Remote 100+ openings
Braintrust
full-remoteseniorfreelancebackend United States Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
AI agent evaluationPythonFastAPIReactTypeScriptDockerPostgreSQLKafkaIntegration TestingEnglish B2+
About the role
Role Overview
You’ll contribute to building a dataset used to evaluate AI coding agents—measuring how well a model handles real-world developer tasks. Work is project-based (not permanent employment).
Responsibilities
- Build realistic developer environments: create a simulated “virtual company” with a codebase, infrastructure, and context (tickets, docs, conversations).
- Design evaluation tasks from intermediate states: craft prompts, define what it means for the agent to be successful, and ensure tasks are solvable by an AI agent.
- Write tests to verify solutions: accept all valid approaches and correctly reject incorrect ones (balanced strictness).
- Iterate with QA feedback: review agent solutions, analyze failures, and refine tasks/tests to make evaluations fair and robust.
What This Is NOT
- Not data labeling
- Not prompt engineering
- Not writing code from scratch (the agent writes most of the code; you guide and evaluate)
Requirements
- 5+ years software development experience
- Core stack experience: Python (FastAPI), JavaScript/TypeScript (React), Docker, PostgreSQL, Kafka, Redis
- Experience writing functional and integration tests
- English proficiency: B2+
Nice-to-haves / Additional Expectations
- Ability to complete identity verification and a technical assessment
- Willingness to use Discord for communication and updates
- Reliable internet and effective communication in a remote setting
Logistics
- Estimated effort: ~20 hours per task (varies by complexity)
- You choose when/how to work, but tasks have deadlines and must meet acceptance criteria.
About Braintrust
Braintrust is a talent network connecting specialists to project-based opportunities with leading technology companies. The work described here focuses on evaluating and improving AI coding agents by building realistic test datasets.
Scraped 7/27/2026