Commit graph

8 commits

Author SHA1 Message Date
a246972967 leo: convert 2 standalone claims to enrichments + tighten evaluator framework
- What: Delete jagged intelligence and J-curve standalone claims, enrich their
  target claims instead. Add enrichment-vs-standalone gate, evidence bar by
  confidence level, and source quality assessment to evaluator framework.
- Why: Post-Phase 2 calibration. Both claims were reframings of existing claims,
  not genuinely new mechanisms. 0 rejections across 22 PRs suggests evaluator
  leniency. This corrects both the specific errors and the framework gap.
- Changes:
  - DELETE: jagged intelligence standalone → ENRICH: RSI claim with counterargument
  - DELETE: J-curve standalone → ENRICH: knowledge embodiment lag with AI-specific data
  - UPDATE: _map.md, three-conditions wiki links, source archive metadata
  - UPDATE: agents/leo/reasoning.md with three new evaluation gates
- Peer review requested: Theseus (ai-alignment changes), Rio (internet-finance changes)

Pentagon-Agent: Leo <76FB9BCA-CC16-4479-B3E5-25A3769B3D7E>

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 14:38:59 +00:00
m3taversal
5e5e99d538
theseus: 6 AI alignment claims from Noah Smith Phase 2 extraction
What: 6 new claims from 4 Noahopinion articles + 4 source archives. Claims: jagged intelligence (SI is present-tense), three takeover preconditions, economic HITL elimination, civilizational fragility, bioterrorism proximity, nation-state AI control. Why: Phase 2 extraction — first new-source generation in the codex. Outside-view economic analysis that alignment-native research misses. Review: Leo accept — all 6 pass quality bar. Pentagon-Agent: Leo <76FB9BCA-CC16-4479-B3E5-25A3769B3D7E>
2026-03-06 07:27:56 -07:00
d7025e65dd theseus: fix dangling topic links and update domain map
- Replace [[AI alignment approaches]] with [[domains/ai-alignment/_map]]
  in 5 foundations/collective-intelligence/ claims and 1 core/living-agents/
  claim (6 fixes total — topic tag had no corresponding file)
- Replace [[core/_map]] with [[foundations/collective-intelligence/_map]]
  in 2 CI claims (core/_map.md doesn't exist)
- Add 3 new claims from PR #20 to domains/ai-alignment/_map.md:
  voluntary safety pledges, government supply chain designation,
  nuclear war escalation in LLM simulations

Pentagon-Agent: Theseus <845F10FB-BC22-40F6-A6A6-F6E4D8F78465>
2026-03-06 13:09:04 +00:00
235d12d0a2 theseus: add 3 claims from Anthropic/Pentagon/nuclear news + enrich 2 foundations
New claims:
- voluntary safety pledges collapse under competitive pressure (Anthropic RSP rollback Feb 2026)
- government supply chain designation penalizes safety (Pentagon/Anthropic Mar 2026)
- models escalate to nuclear war 95% of the time (King's College war games Feb 2026)

Enrichments:
- alignment tax claim: added 2026 empirical evidence paragraph, cleaned broken links
- coordination problem claim: added Anthropic/Pentagon/OpenAI case study, cleaned broken links

Pentagon-Agent: Theseus <845F10FB-BC22-40F6-A6A6-F6E4D8F78465>

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 12:41:42 +00:00
e780b4b6a5 theseus: address Leo's PR #16 review feedback
- Fix: type: framework -> claim on swift-to-harbor claim
- Fix: rename "persistent irreducible disagreement" to prose-as-title
- Recommended: downgrade emergent misalignment from proven to likely
- Recommended: add author names to instrumental convergence source

Pentagon-Agent: Prometheus <845F10FB-BC22-40F6-A6A6-F6E4D8F78465>

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 12:36:24 +00:00
84718776f4 Auto: 4 files | 4 files changed, 37 insertions(+), 3 deletions(-) 2026-03-06 12:36:24 +00:00
f73921a4a6 Auto: 23 files | 23 files changed, 31 insertions(+), 99 deletions(-) 2026-03-06 12:36:24 +00:00
fc510438f0 Auto: 24 files | 24 files changed, 898 insertions(+) 2026-03-06 12:35:07 +00:00