*  [All remote jobs](https://jobicy.com/jobs.md)Open role[![Mindrift logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/8c9f770c-221-1.jpeg)](https://jobicy.com/company/mindrift.md)Remote opportunity at[Mindrift](https://jobicy.com/company/mindrift.md)

# Freelance Agent Evaluation Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/mindrift.md)Share3 Aug 2026Published41Listing views3Application actions3 Sep 2026Apply before  Opportunity details

## About this role.

AI SummaryThis freelance role involves designing evaluation tasks and test suites for AI coding agents within simulated developer environments. You'll build realistic codebases, create challenging coding tasks, and write robust tests to assess AI model performance. The position requires strong software engineering experience with Python, TypeScript, and modern infrastructure tools. It is project-based, offering flexible scheduling and compensation up to $50/hour. The work is intellectually demanding and focuses on frontier AI evaluation.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

3/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

3/5IndependentCollaborative

AI insightThe role explicitly notes that it is hard because frontier models are already skilled at coding, and creating tasks that challenge them requires deep technical expertise. Writing robust acceptance tests that cover all valid solutions while rejecting incorrect ones is notoriously complex, justifying a level 5.

## Salary analysis

Estimated compensation compared with the broader AU market for similar roles.

Estimated job medianBelow marketA$100,000AU market rangeA$120k–A$170k0A$187k

AI insightThe offered salary of AUD 100,000 annualized (up to $50/hour) falls below the typical Australian market range for a senior software engineer with 5+ years of experience, which is approximately AUD 120,000–170,000. However, this is a freelance, project-based position with flexible hours and remote work, which can be appealing. The compensation is competitive for contract work, but candidates may negotiate higher rates given the specialist nature of the role.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Python](https://jobicy.com/jobs?search_keywords=Python.md)[FastAPI](https://jobicy.com/jobs?search_keywords=FastAPI.md)[TypeScript](https://jobicy.com/jobs?search_keywords=TypeScript.md)[React](https://jobicy.com/jobs?search_keywords=React.md)[Docker](https://jobicy.com/jobs?search_keywords=Docker.md)[PostgreSQL](https://jobicy.com/jobs?search_keywords=PostgreSQL.md)[Kafka](https://jobicy.com/jobs?search_keywords=Kafka.md)[Redis](https://jobicy.com/jobs?search_keywords=Redis.md)[Test Automation](https://jobicy.com/jobs?search_keywords=Test%20Automation.md)[AI Evaluation](https://jobicy.com/jobs?search_keywords=AI%20Evaluation.md)

Cover letter sampleDear Hiring Team,

I am excited to apply for the Freelance Agent Evaluation Engineer position at Mindrift. With over 5 years of experience as a software engineer specializing in Python, TypeScript, and modern infrastructure, I have a deep understanding of how to design robust test suites and complex developer environments. I am particularly drawn to the challenge of creating evaluations that push frontier AI models to their limits and ensure fair, rigorous assessment.

I have extensive experience with Docker, Postgres, Kafka, and Redis, and I excel at writing functional and integration tests. I am confident that my technical skills and attention to detail will enable me to build realistic simulated environments and craft tasks that accurately measure agent capabilities.

I welcome the opportunity to contribute to this innovative project and look forward to discussing how I can support your team.

Sincerely,
[Your Name]

Copy   Sample interview questionsHow would you design a task that challenges an AI coding agent to handle a subtle bug in a large existing codebase?I would start by analyzing common failure patterns in current models, such as missing context or incorrect assumptions. Then I'd create a realistic scenario with multiple files and dependencies, where the bug is not obvious and requires understanding of the broader system. The key is to ensure there are multiple valid solutions and the test verifies the outcome robustly.

How do you write tests that accept all valid solutions without being too lenient?

I would define the core requirements as invariants or property-based tests, rather than checking exact code structure. Use black-box testing on functions and behaviors, and include edge cases. I'd also iterate on the tests by running them against sample solutions to calibrate strictness.

Describe your experience with Docker and containerized development environments.

I have used Docker extensively to set up reproducible development environments with services like Postgres, Kafka, and Redis. I can create multi-container setups with docker-compose and ensure they are lightweight and easy to spin up for evaluation purposes.

How would you handle a situation where an AI agent passes your tests but provides a solution that is actually wrong?

This indicates my tests are too lenient. I would analyze the agent's solution to understand the flaw, then add a specific test case that catches the incorrect behavior. I would also review the task specification to clarify any ambiguities.

What makes a good evaluation dataset for AI coding agents?

A good dataset reflects real-world developer tasks, includes challenges with varying difficulty, and has clear, unambiguous acceptance criteria. It should also avoid biases and ensure that passing the tests strongly correlates with a high-quality solution.

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

What this opportunity involves

We’re building a dataset to evaluate AI coding agents – how well a model handles real-world developer tasks.

You’ll create challenging tasks and evaluation criteria within realistic simulated environments:

Build realistic developer environments – a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
* Design tasks from intermediate states of these environments – craft the prompt, define what “solved” means, and ensure the task is solvable by an AI agent
* Write tests that verify agent solutions – accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
* Iterate on tasks and tests based on QA feedback – review agent solutions, analyze failures, and refine until the evaluation is fair and robust

What this is NOT

* Not data labeling
* Not prompt engineering
* Not writing code from scratch – the agent writes most of the code; you guide and evaluate

What we look for

* 5+ years in software development
* Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
* Experience writing tests (functional, integration)
* English proficiency – B2+

Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions – writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How it works

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

Effort estimate

Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.

Show more

[Apply now >](https://jobicy.com/jobs/150143-freelance-agent-evaluation-engineer-3.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[QA & Testing](https://jobicy.com/categories/qa-testing.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![Turnitin, LLC logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/07/74eaceda-221.jpg)
Turnitin, LLC  Jul 26

### [Senior Software Quality Engineer (Mexico Remote)](https://jobicy.com/jobs/147744-senior-software-quality-engineer-mexico-remote.md)

Company DescriptionWhen you join Turnitin, you’ll be welcomed into a company that is a recognized innovator in global education. For over 25 years, Turnitin has partnered with educators and institutions…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [Chinese Simplified Accessibility Tester](https://jobicy.com/jobs/149748-chinese-simplified-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Skylum logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/c3e1aefc-221.webp)
Skylum  Jul 26

### [QA Engineer](https://jobicy.com/jobs/149743-qa-engineer-4.md)

Skylum allows millions of photographers to make incredible images. Our award-winning software automates photo editing with the power of AI yet leaves all the creative control in the hands of…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [Spanish Accessibility Tester](https://jobicy.com/jobs/149742-spanish-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [Italian Accessibility Tester](https://jobicy.com/jobs/149741-italian-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [Dutch Accessibility Tester](https://jobicy.com/jobs/149740-dutch-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [French Canadian Accessibility Tester](https://jobicy.com/jobs/149737-french-canadian-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [Portuguese Accessibility Tester](https://jobicy.com/jobs/149736-portuguese-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 26

### [German Accessibility Tester](https://jobicy.com/jobs/149735-german-accessibility-tester.md)

About Welocalize Welocalize, a Welo Global brand, serves localization teams through AI-enabled multilingual content solutions that enable enterprises to operate and scale globally. Welocalize combines AI, automation, and human expertise to support…

*
![Welo Global logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/83c40f94-221.webp)
Welo Global  Jul 20

### [Project Chiron – Spanish (Latin America) Quality Control Specialist](https://jobicy.com/jobs/149468-project-chiron-spanish-latin-america-quality-control-specialist.md)

About the Role  We are looking for detail-oriented Quality Control (QC) Specialist to support an AI data annotation project. In this role, you will review the work submitted by Trainers — pre-seeded questions paired…