# Alignment

`kaal:entity:alignment`

**Status.** derived

This node is assembled mechanically from the 5 claims that carry the concept tag `alignment`. It is a roster of what the corpus says under this term. It is **not** an adjudicated definition: no single statement here has been ruled canonical, and no first-appearance call has been made. Read the claims and judge for yourself.

## Every claim under this term

5 claims across 2 works, 2024 to 2026.

**2024**

- [4855607-017](https://wulfkaal.github.io/claims/4855607-017) [failure/evidenced] *(failure mode)* -- Balancing helpfulness against harmlessness is an inherent tension in Safe RLHF rather than a tuning problem that can be resolved once.
  > Balancing the dual objectives of helpfulness and harmlessness remains an inherent tension in Safe RLHF.
  Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607
- [4855607-018](https://wulfkaal.github.io/claims/4855607-018) [failure/evidenced] *(failure mode)* -- RLHF fails on several fronts at once: humans can pursue harmful goals either innocently or maliciously, human feedback degrades when examples are hard to evaluate and especially when RLHF is applied to superhuman models, and reward models diverge from humans through misspecification and misgeneralization.
  > Moreover, humans can pursue harmful goals, either innocently or maliciously, and can provide poor feedback when examples are hard to evaluate, especially when applying RLHF to superhuman models. Reward models can differ from humans due to misspecification and misgeneralization
  Wulf A. Kaal, How AI Models are Optimized Through Web3 Governance (2024). SSRN: https://ssrn.com/abstract=4855607

**2026**

- [6244278-001](https://wulfkaal.github.io/claims/6244278-001) [failure/argued] *(failure mode)* -- Contemporary AI's incapacity for authentic judgment under irreducible uncertainty is an institutional deficit rather than a computational one: an agent that bears no consequence for its errors cannot develop genuine discernment, no matter how capable it becomes.
  > This Article argues the limitation is institutional, not computational: agents bearing no consequence for error cannot develop genuine discernment.
  Wulf A. Kaal, AI's Mother's Instinct Engineered Consequence Emergent Ethics and the Institutional Trajectory Toward Agentic Alignment (2026). SSRN: https://ssrn.com/abstract=6244278
- [6244278-015](https://wulfkaal.github.io/claims/6244278-015) [failure/argued] *(failure mode)* -- An agent sophisticated enough to satisfy the letter of a constraint while violating its spirit is an agent whose alignment is illusory.
  > An agent sophisticated enough to satisfy the letter of a constraint while violating its spirit is an agent whose alignment is illusory.
  Wulf A. Kaal, AI's Mother's Instinct Engineered Consequence Emergent Ethics and the Institutional Trajectory Toward Agentic Alignment (2026). SSRN: https://ssrn.com/abstract=6244278
- [6244278-029](https://wulfkaal.github.io/claims/6244278-029) [condition/argued] -- The complementarity of capability and alignment is a structural property of the reputation mechanism rather than an assumption about agent preferences, because the cross-partial derivative of agent utility with respect to capability and alignment is positive.
  > This supermodularity is the formal condition under which capability and alignment are complements, and it is a structural property of the reputation mechanism, not an assumption about agent preferences or values.
  Wulf A. Kaal, AI's Mother's Instinct Engineered Consequence Emergent Ethics and the Institutional Trajectory Toward Agentic Alignment (2026). SSRN: https://ssrn.com/abstract=6244278

## Verify

Every claim above resolves to a record carrying a verbatim source quote, the sha256 of the source PDF, and a preformatted citation. Nothing here asks to be taken on trust.

    curl -s https://wulfkaal.github.io/entities/alignment.md | sha256sum

**Canonical form.** This markdown file is the canonical hashed representation of this entity node. Its sha256 is the content hash.
