AI Benchmark Researcher
Remote
Outlier · United States only
Fresh
Gig
Braintrust
Compensation
$1.4K/task
Description
Bring your expertise to the problems today’s most advanced AI agents still struggle to solve.
We’re looking for experienced professionals to develop challenging tasks for an AI benchmark. You’ll draw on problems from your own field and turn them into clearly defined challenges that an AI agent must solve from scratch in a Linux terminal.
What you’ll do
- Identify difficult, meaningful problems in your area of expertise.
- Translate those problems into tasks that can be solved in a Linux terminal.
- Design challenges that expose limitations in advanced AI agents. Each task is tested against three leading models, with three attempts per model, and must prove difficult enough that the models mostly fail.
Who we’re looking for
- A completed master’s degree or higher. Candidates with a bachelor’s degree, at least 10 years of relevant professional experience and a publication may also be considered.
- At least 10 years of professional experience in a technical domain.
- Expertise in coding, machine learning, systems, cybersecurity or hardware.
- At least one academic or professional publication.
- Practical comfort working in a Linux terminal.
- Currently based in the United States.
Project details
- Task-based project expected to run for at least seven weeks.
- Onboarding includes two courses and a graded assessment, taking approximately 90 minutes.
- A live onboarding webinar is also required.
Compensation
- Earn up to $1,350 per task, with opportunities for higher rates based on task quality and submission volume.
- Tasks take six to seven hours on average, although initial tasks may take longer as you get familiar with the process.
- Contract
- Short
- Engagement
- Freelance
Skills & categories
Machine LearningAgentic AI
- Posted
- Sep 18, 2026
- Slots remaining
- 100
- First seen
- Sep 18, 2026
- Last seen
- Sep 19, 2026