Compensation
$402,048-$687,759/yrDescription
About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
Overview
METR has started embedding researchers inside frontier labs to investigate incidents, stress-test labs’ internal agent monitoring systems, and assess loss-of-control risks from internal deployment. As agent capabilities increase, we expect this to be one of the most important sources of independent information the world has about catastrophic risks from advanced AI.
As the source of information becomes more important, we'll need many more talented researchers and engineers who can conduct embedded exercises. We expect these assessors to have deep access, and for their work to be a large part of METR's impact in the next year. We want to build on the momentum from previous exercises to further develop our risk assessments.
What the job looks like
You'll be embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology, talk to other researchers, build tooling to make future exercises go better, help us hire and scale, write up results, and plan/coordinate future exercises.
We are looking for embedded researchers and engineers across multiple current and potential future exercises: AI R&D acceleration assessment, compute allocation, monitorability red-teaming, incident investigations, and more. We expect this role to evolve significantly as we develop and prototype this new form of risk assessment.
Ideal candidate
We are looking for candidates with at least two of the following skills, although the ideal candidate will have most or all of them:
You’re good at prompting LLMs, are familiar with research relevant to the fields above (or can get up to speed quickly).
You have some security experience.
You can get spun up on large codebases quickly.
You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel and you’ll need to figure a lot of stuff out on the fly largely by yourself.
You are excellent at loss-of-control threat modeling and breaking down safety cases.
You’re good at verbal communication, writing, and stakeholder management (e.g., navigating complex relationships with frontier labs).
You’re trustworthy and have a good reputation.
Stack
- Posted
- Jul 19, 2026
- Last seen
- Aug 27, 2026
- First seen
- Aug 27, 2026




