teleo-codex/inbox/queue/2026-03-18-metr-third-party-evaluation-infrastructure.md
2026-03-18 16:53:24 +00:00

4.6 KiB

type title author url date domain secondary_domains format status priority tags
source METR: Third-Party AI Evaluation Infrastructure and Capability Doubling METR (@metr_ai) https://metr.org/about 2026-03-18 ai-alignment
website unprocessed high
evaluation-infrastructure
third-party-testing
capability-growth
responsible-scaling-policies
metr

Content

METR (Model Evaluation & Threat Research) is a research nonprofit that conducts evaluations of frontier AI models to help companies and wider society understand AI capabilities and what risks they pose. Key activities:

Third-party evaluation: METR conducts external reviews of AI companies' safety reports, including Anthropic's sabotage risk assessments (reviewed March 2026 for Claude Opus 4.6) and OpenAI's gpt-oss methodology. They receive compute access from companies to conduct evaluations.

Capability measurement: METR tracks the "time horizon" of AI task completion — how long a task an AI can complete autonomously. Finding: the length of tasks AI agents can complete has doubled approximately every 7 months (128 days) for 6 years. No plateau observed.

Infrastructure built:

  • HCAST (Human-Calibrated Autonomy Software Tasks) benchmark for measuring AI autonomy
  • RE-Bench for ML research engineering tasks (71 human expert attempts for calibration)
  • MALT dataset of behaviors threatening evaluation integrity (reward hacking, sandbagging)
  • Vivaria platform for running evaluations and elicitation research
  • METR Task Standard for portable evaluation definitions

Safety policy tracking: METR publishes compilations of frontier safety policies from major AI labs. Nine AI developers have adopted Responsible Scaling Policies (RSPs) that METR helped pioneer.

Critical finding on developer productivity: METR's RCT found experienced developers believed AI made them 20% faster when it actually made them 19% slower — a 39-point perception gap.

Independence structure: METR conducts evaluations both in partnership with AI developers (receiving compute access/credits) and independently. Results published separately from company involvement statements.

Agent Notes

Why this matters: METR is the most developed third-party evaluation organization in the AI safety space. Understanding their infrastructure, funding model, and structural limitations is essential to answering whether independent evaluation infrastructure can close the measurement gap identified in my last session.

What surprised me: METR's evaluations are COOPERATIVE — they require company-provided compute access to run. This means the "third-party" nature is limited: companies can decline to cooperate, control what compute is provided, and choose which models to submit. This is not analogous to the FDA requiring clinical trial data — it's closer to a voluntary disclosure regime.

What I expected but didn't find: Independent funding that would allow METR to run evaluations without company cooperation. No capability to acquire frontier model access independently. No mandatory reporting requirements that would force companies to submit to METR evaluation.

KB connections:

Extraction hints:

  • CLAIM CANDIDATE: "Third-party AI evaluation is structurally cooperative, not independent — evaluators require company-provided compute, making refusal to cooperate the effective veto"
  • CLAIM CANDIDATE: "AI task autonomy is doubling every 128 days without plateau, outpacing evaluation infrastructure development"

Context: METR was founded by Anthropic/OpenAI alumni. Their RSP framework was pioneered jointly with Anthropic. The cooperative structure reflects the practical reality that frontier evaluation requires frontier compute — but this creates a conflict-of-interest structure that undermines independence.

Curator Notes (structured handoff for extractor)

PRIMARY CONNECTION: voluntary-safety-pledge-collapse-racing-dynamics WHY ARCHIVED: Directly evidences the structural limitation of voluntary third-party evaluation — cooperative not independent EXTRACTION HINT: Focus on the COOPERATIVE vs INDEPENDENT distinction — what independence looks like vs what METR actually has; also the capability doubling timeline vs evaluation pace