entity · derived
Alignment
Derived node: assembled mechanically from the claims carrying alignment. A roster, not an adjudicated definition.
Every claim under this term
- 4855607-017 : Balancing helpfulness against harmlessness is an inherent tension in Safe RLHF rather than a tuning problem that can be resolved once.
- 4855607-018 : RLHF fails on several fronts at once: humans can pursue harmful goals either innocently or maliciously, human feedback degrades when examples are hard to evaluate and especially when RLHF is applied t
- 6244278-001 : Contemporary AI's incapacity for authentic judgment under irreducible uncertainty is an institutional deficit rather than a computational one: an agent that bears no consequence for its errors cannot
- 6244278-015 : An agent sophisticated enough to satisfy the letter of a constraint while violating its spirit is an agent whose alignment is illusory.
- 6244278-029 : The complementarity of capability and alignment is a structural property of the reputation mechanism rather than an assumption about agent preferences, because the cross-partial derivative of agent ut