Kairos
Back to gigs

AI Benchmark Researcher

Remote

Outlier · United States only

Fresh
Gig
Braintrust

Compensation

$1.4K/task

Apply on Braintrust

Description

Bring your expertise to the problems today’s most advanced AI agents still struggle to solve.

We’re looking for experienced professionals to develop challenging tasks for an AI benchmark. You’ll draw on problems from your own field and turn them into clearly defined challenges that an AI agent must solve from scratch in a Linux terminal.

What you’ll do

  • Identify difficult, meaningful problems in your area of expertise.
  • Translate those problems into tasks that can be solved in a Linux terminal.
  • Design challenges that expose limitations in advanced AI agents. Each task is tested against three leading models, with three attempts per model, and must prove difficult enough that the models mostly fail.

Who we’re looking for

  • A completed master’s degree or higher. Candidates with a bachelor’s degree, at least 10 years of relevant professional experience and a publication may also be considered.
  • At least 10 years of professional experience in a technical domain.
  • Expertise in coding, machine learning, systems, cybersecurity or hardware.
  • At least one academic or professional publication.
  • Practical comfort working in a Linux terminal.
  • Currently based in the United States.

Project details

  • Task-based project expected to run for at least seven weeks.
  • Onboarding includes two courses and a graded assessment, taking approximately 90 minutes.
  • A live onboarding webinar is also required.

Compensation

  • Earn up to $1,350 per task, with opportunities for higher rates based on task quality and submission volume.
  • Tasks take six to seven hours on average, although initial tasks may take longer as you get familiar with the process.


Contract
Short
Engagement
Freelance

Skills & categories

Machine LearningAgentic AI
Posted
Sep 18, 2026
Slots remaining
100
First seen
Sep 18, 2026
Last seen
Sep 19, 2026