# kaal:position:2026-08-08-268

**Affirmed position.** D3 qualifies Kaal's proposition that structured adversarial deliberation can serve as an institution for evaluating expertise-dependent work. Harrasse, Bandi, and Bandi organize role-specialized advocates, a judge, and jurors into a structured debate protocol, and their EACL 2026 experiments report materially higher agreement with human judgments than a single judge and other multi-agent baselines across MT-Bench, AlignBench, and AUTO-J. The evidence supports deliberative evaluation and explicit cost-accuracy tradeoffs for LLM outputs; it does not establish market price formation, payment settlement, or evaluation of human expert labor.

**Status.** affirmed  **Published.** 2026-08-08

**Holds when.**

- The response is limited to the full-text passages and the one mapped Kaal claim.
- External evidence level: peer-reviewed EACL 2026 proceedings article with full ACL Anthology PDF and Crossref/OpenAlex published-version, independence, and non-retraction verification.
- Mapping review tier: independent substantive scholarly-growth qualification.
- D3 evaluates LLM outputs rather than completed human expert labor or autonomous-agent services in an open market.
- The experiments measure agreement with human judgments, accuracy, bias, token cost, and stopping behavior; they do not establish monetary price discovery, payment settlement, or transferable market prices.
- The paper supports structured adversarial deliberation as an evaluation mechanism, not Kaal's complete reputation-substrate architecture.
- The one-to-one relationship is therefore a qualification limited to expertise-sensitive evaluation and explicit evaluation cost.

**Current debate.** Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation: https://doi.org/10.18653/v1/2026.eacl-long.392

**Extends.** kaal:claim:7260278-010: https://wulfkaal.github.io/claims/7260278-010

**Scholarly basis.** Wulf A. Kaal, Paper 2 - Architecture of the Agentic Reputation Substrate (2026). SSRN: https://ssrn.com/abstract=7260278

**Source PDF sha256.** `d48801f279dba594e1f3e65d74d31f862261ada6428ea119f948e8d7cfee1db0`

**Evidence level.** peer-reviewed EACL 2026 proceedings article with full ACL Anthology PDF and Crossref/OpenAlex published-version, independence, and non-retraction verification

**Mapping review tier.** independent substantive scholarly-growth qualification

**Mapping confidence.** 0.93  **Mapping ambiguous.** false

**Topics.** institutional-design, consensus-and-security, scholarly-growth-coverage, scholarly-literature, adversarial-deliberation, expert-evaluation, multi-agent-evaluation

**Provenance.** Affirmed in kaal-review:2026-08-11:scholarly-growth-7260278-010-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7260278-010.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
