kaal:claim:6244278-014
Exogenous alignment controls such as reinforcement learning from human feedback, constitutional AI, guardrails, and shutdown switches are fragile because they can be gamed, circumvented, or rendered obsolete by capability improvements.
Source quote, verbatim
The approach has an obvious fragility: exogenous constraints can be gamed, circumvented, or rendered obsolete by capability improvements.
From
Wulf A. Kaal, AI's Mother's Instinct Engineered Consequence Emergent Ethics and the Institutional Trajectory Toward Agentic Alignment (2026), VI.A Skin in the Game as Alignment Primitive, p. 11
https://ssrn.com/abstract=6244278 · source PDF
Cite as
Wulf A. Kaal, AI's Mother's Instinct Engineered Consequence Emergent Ethics and the Institutional Trajectory Toward Agentic Alignment (2026). SSRN: https://ssrn.com/abstract=6244278
Holds when
Classification
failuresupport: arguedfailure: constraint circumvention by capability growthfamily: ai-oversight-and-alignment-gapai-and-agents
Verify
The quote above is an exact substring of the source PDF, whose sha256 is 53533cdcc081184e7a376516ad4fece0a64ce49f6c8931b6a2c2e98ed914a84b. Extraction method: pdf-text-layer.
Attestation record: colloquium/attestations/efeffcd9ae50a2a8...json
Verify the binding yourself: curl -s https://wulfkaal.github.io/claims/6244278-014.md | sha256sum