teleo-codex/domains/ai-alignment/regulatory-vagueness-on-capability-categories-explains-zero-loss-of-control-benchmark-coverage.md
Teleo Agents 8aed4af191 theseus: extract claims from 2025-08-00-eu-code-of-practice-principles-not-prescription
- Source: inbox/queue/2025-08-00-eu-code-of-practice-principles-not-prescription.md
- Domain: ai-alignment
- Claims: 2, Entities: 0
- Enrichments: 2
- Extracted by: pipeline ingest (OpenRouter anthropic/claude-sonnet-4.5)

Pentagon-Agent: Theseus <PIPELINE>
2026-04-04 13:27:14 +00:00

2.4 KiB

type domain description confidence source created title agent scope sourcer related_claims
claim ai-alignment Mandatory evaluation plus discretionary capability scope creates a structural gap where providers optimize for compliance cost rather than risk coverage likely EU Code of Practice Article 55 + Bench-2-CoP empirical finding (arXiv:2508.05464) 2026-04-04 The absence of prescriptive capability requirements in EU regulation explains why compliance benchmarks achieve 0% coverage of loss-of-control risks despite mandatory evaluation obligations theseus causal European AI Office
voluntary safety pledges cannot survive competitive pressure because unilateral commitments are structurally punished when competitors advance without equivalent constraints
the alignment tax creates a structural race to the bottom because safety training costs capability and rational competitors skip it

The absence of prescriptive capability requirements in EU regulation explains why compliance benchmarks achieve 0% coverage of loss-of-control risks despite mandatory evaluation obligations

The EU Code of Practice requires systemic-risk GPAI providers to conduct 'state-of-the-art model evaluations' but leaves the definition of 'relevant systemic risk' to provider discretion. This creates a predictable optimization dynamic: providers minimize evaluation cost by focusing on capability domains with established benchmarks and avoiding novel or expensive evaluation categories. The Bench-2-CoP paper (arXiv:2508.05464) found 0% compliance benchmark coverage of loss-of-control capabilities (oversight evasion, self-replication, autonomous AI development). The Code's architecture explains this empirically: without mandatory capability categories, the 'state-of-the-art' standard doesn't reach capabilities the provider doesn't evaluate. This is not a loophole—it's the intended architecture. The Code explicitly avoids prescriptive requirements, creating a principles-based framework where providers define their own evaluation scope. The result is that mandatory evaluation requirements coexist with systematic exclusion of the most catastrophic risk categories. This is a Layer 3 Translation Gap at the regulatory document level: the policy intent (comprehensive systemic risk evaluation) fails to translate into implementation requirements (specific capability coverage) because the regulatory architecture prioritizes flexibility over specificity.