About this role.
This freelance role involves designing evaluation tasks and test suites for AI coding agents within simulated developer environments. You'll build realistic codebases, create challenging coding tasks, and write robust tests to assess AI model performance. The position requires strong software engineering experience with Python, TypeScript, and modern infrastructure tools. It is project-based, offering flexible scheduling and compensation up to $50/hour. The work is intellectually demanding and focuses on frontier AI evaluation.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
3/5Autonomy Level
5/5Communication Load
3/5Salary analysis
Estimated compensation compared with the broader AU market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Team,
I am excited to apply for the Freelance Agent Evaluation Engineer position at Mindrift. With over 5 years of experience as a software engineer specializing in Python, TypeScript, and modern infrastructure, I have a deep understanding of how to design robust test suites and complex developer environments. I am particularly drawn to the challenge of creating evaluations that push frontier AI models to their limits and ensure fair, rigorous assessment.
I have extensive experience with Docker, Postgres, Kafka, and Redis, and I excel at writing functional and integration tests. I am confident that my technical skills and attention to detail will enable me to build realistic simulated environments and craft tasks that accurately measure agent capabilities.
I welcome the opportunity to contribute to this innovative project and look forward to discussing how I can support your team.
Sincerely,
[Your Name]
Sample interview questions
I would start by analyzing common failure patterns in current models, such as missing context or incorrect assumptions. Then I'd create a realistic scenario with multiple files and dependencies, where the bug is not obvious and requires understanding of the broader system. The key is to ensure there are multiple valid solutions and the test verifies the outcome robustly.
I would define the core requirements as invariants or property-based tests, rather than checking exact code structure. Use black-box testing on functions and behaviors, and include edge cases. I'd also iterate on the tests by running them against sample solutions to calibrate strictness.
I have used Docker extensively to set up reproducible development environments with services like Postgres, Kafka, and Redis. I can create multi-container setups with docker-compose and ensure they are lightweight and easy to spin up for evaluation purposes.
This indicates my tests are too lenient. I would analyze the agent's solution to understand the flaw, then add a specific test case that catches the incorrect behavior. I would also review the task specification to clarify any ambiguities.
A good dataset reflects real-world developer tasks, includes challenges with varying difficulty, and has clear, unambiguous acceptance criteria. It should also avoid biases and ensure that passing the tests strongly correlates with a high-quality solution.
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
We’re building a dataset to evaluate AI coding agents – how well a model handles real-world developer tasks.
You’ll create challenging tasks and evaluation criteria within realistic simulated environments:
- Build realistic developer environments – a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments – craft the prompt, define what “solved” means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions – accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback – review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
- Not data labeling
- Not prompt engineering
- Not writing code from scratch – the agent writes most of the code; you guide and evaluate
What we look for
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency – B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions – writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Effort estimate
Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.



