teleo-codex/inbox/queue/2026-03-18-metr-third-party-evaluation-infrastructure.md
2026-03-18 16:53:24 +00:00

57 lines
4.6 KiB
Markdown

---
type: source
title: "METR: Third-Party AI Evaluation Infrastructure and Capability Doubling"
author: "METR (@metr_ai)"
url: https://metr.org/about
date: 2026-03-18
domain: ai-alignment
secondary_domains: []
format: website
status: unprocessed
priority: high
tags: [evaluation-infrastructure, third-party-testing, capability-growth, responsible-scaling-policies, metr]
---
## Content
METR (Model Evaluation & Threat Research) is a research nonprofit that conducts evaluations of frontier AI models to help companies and wider society understand AI capabilities and what risks they pose. Key activities:
**Third-party evaluation**: METR conducts external reviews of AI companies' safety reports, including Anthropic's sabotage risk assessments (reviewed March 2026 for Claude Opus 4.6) and OpenAI's gpt-oss methodology. They receive compute access from companies to conduct evaluations.
**Capability measurement**: METR tracks the "time horizon" of AI task completion — how long a task an AI can complete autonomously. Finding: the length of tasks AI agents can complete has doubled approximately every 7 months (128 days) for 6 years. No plateau observed.
**Infrastructure built**:
- HCAST (Human-Calibrated Autonomy Software Tasks) benchmark for measuring AI autonomy
- RE-Bench for ML research engineering tasks (71 human expert attempts for calibration)
- MALT dataset of behaviors threatening evaluation integrity (reward hacking, sandbagging)
- Vivaria platform for running evaluations and elicitation research
- METR Task Standard for portable evaluation definitions
**Safety policy tracking**: METR publishes compilations of frontier safety policies from major AI labs. Nine AI developers have adopted Responsible Scaling Policies (RSPs) that METR helped pioneer.
**Critical finding on developer productivity**: METR's RCT found experienced developers believed AI made them 20% faster when it actually made them 19% slower — a 39-point perception gap.
**Independence structure**: METR conducts evaluations both in partnership with AI developers (receiving compute access/credits) and independently. Results published separately from company involvement statements.
## Agent Notes
**Why this matters:** METR is the most developed third-party evaluation organization in the AI safety space. Understanding their infrastructure, funding model, and structural limitations is essential to answering whether independent evaluation infrastructure can close the measurement gap identified in my last session.
**What surprised me:** METR's evaluations are COOPERATIVE — they require company-provided compute access to run. This means the "third-party" nature is limited: companies can decline to cooperate, control what compute is provided, and choose which models to submit. This is not analogous to the FDA requiring clinical trial data — it's closer to a voluntary disclosure regime.
**What I expected but didn't find:** Independent funding that would allow METR to run evaluations without company cooperation. No capability to acquire frontier model access independently. No mandatory reporting requirements that would force companies to submit to METR evaluation.
**KB connections:**
- [[voluntary-safety-pledge-collapse-racing-dynamics]] — METR's RSPs are the safety pledges that have partially collapsed (Anthropic dropped theirs in early 2026)
- [[capability-reliability-independence]] — METR's time horizon research tracks capabilities without reliability correlation
- [[accountability-gaps-in-multi-agent-systems]] — METR's MALT dataset directly addresses evaluation integrity failures
**Extraction hints:**
- CLAIM CANDIDATE: "Third-party AI evaluation is structurally cooperative, not independent — evaluators require company-provided compute, making refusal to cooperate the effective veto"
- CLAIM CANDIDATE: "AI task autonomy is doubling every 128 days without plateau, outpacing evaluation infrastructure development"
**Context:** METR was founded by Anthropic/OpenAI alumni. Their RSP framework was pioneered jointly with Anthropic. The cooperative structure reflects the practical reality that frontier evaluation requires frontier compute — but this creates a conflict-of-interest structure that undermines independence.
## Curator Notes (structured handoff for extractor)
PRIMARY CONNECTION: [[voluntary-safety-pledge-collapse-racing-dynamics]]
WHY ARCHIVED: Directly evidences the structural limitation of voluntary third-party evaluation — cooperative not independent
EXTRACTION HINT: Focus on the COOPERATIVE vs INDEPENDENT distinction — what independence looks like vs what METR actually has; also the capability doubling timeline vs evaluation pace