teleo-codex/domains/ai-alignment/alignment-framing-as-manhattan-project-assumes-five-properties-that-alignment-lacks.md
Teleo Agents 318190eb24 theseus: extract claims from 2026-00-00-friederich-against-manhattan-project-alignment.md
- Source: inbox/archive/2026-00-00-friederich-against-manhattan-project-alignment.md
- Domain: ai-alignment
- Extracted by: headless extraction cron (worker 3)

Pentagon-Agent: Theseus <HEADLESS>
2026-03-11 04:07:31 +00:00

4.2 KiB

type domain description confidence source created depends_on
claim ai-alignment The Manhattan Project metaphor for AI alignment encodes five philosophical assumptions—binary achievement, natural kinds, technical solvability, one-shot solutions, and operationalizability—that mischaracterize alignment's actual nature experimental Friederich & Dung (2026), Mind & Language 2026-03-11
AI alignment is a coordination problem not a technical problem.md
the specification trap means any values encoded at training time become structurally unstable as deployment contexts diverge from training conditions.md
some disagreements are permanently irreducible.md

The Manhattan Project framing of alignment encodes five philosophical assumptions that mischaracterize the problem

Friederich and Dung (2026) argue that AI companies and researchers frame alignment as a "Manhattan Project"—a clear, well-delineated, unified scientific problem solvable within years—but this framing encodes five philosophical assumptions that fail on analysis:

The Five Assumptions

  1. Binary achievement: The framing assumes alignment is a yes/no state that can be achieved and verified, rather than a continuous spectrum or context-dependent property that admits degrees and variations.

  2. Natural kind: It treats alignment as a single unified phenomenon with a discoverable essence, rather than a heterogeneous collection of distinct problems that may not share a common structure.

  3. Technical-scientific solvability: It assumes alignment is primarily a technical problem solvable through scientific methods and engineering, excluding irreducible social, political, and value-pluralistic dimensions.

  4. One-shot achievability: It presumes alignment can be solved once and then implemented, rather than requiring ongoing adaptation, renegotiation, and course correction as contexts and values evolve.

  5. Operationalizability: It assumes alignment can be operationalized—defined in implementable, measurable terms—such that "solving the alignment problem and implementing the solution would be sufficient to rule out AI takeover." The authors argue this is "probably impossible" because alignment may not admit the kind of complete formal specification this assumption requires.

The Harm of the Framing

The paper argues this framing "may bias societal discourse and decision-making towards faster AI development and deployment than is responsible" by making alignment appear more tractable, bounded, and solvable than it actually is. This creates pressure for premature deployment and underestimates the ongoing governance challenges.

Philosophical Grounding

The argument draws on philosophy of science (natural kinds, operationalization, problem specification) rather than AI safety or governance literatures. The operationalizability claim is the strongest: not just that alignment is hard to operationalize, but that it's "probably impossible" to define it such that solving the defined problem would be sufficient to prevent takeover. This suggests a category error—alignment may not be the kind of thing that admits complete formal specification in the way the Manhattan Project framing assumes.

Limitations

The full text is paywalled, so the specific arguments supporting each of the five points cannot be evaluated in depth. The claim rests on abstract and summary descriptions. The "probably impossible" language on operationalizability is strong but not proven—it's a philosophical argument about the nature of the problem, not an empirical demonstration.