# kaal:position:2026-08-08-323

**Affirmed position.** Structured assessment changes a decision process only if the assessment arrives before judgment. That sequence matters. Harrasse, Bandi, and Bandi route paired LLM evaluations through advocates, a criteria-based judge, and an independent jury that renders the final verdict. D3 compares that deliberative architecture with a single judge that selects directly. On MT-Bench, D3-MORE reports 85.1 percent accuracy, 12.6 percentage points above the single-judge baseline. This result qualifies Kaal's treatment-control design. It supports separating an assessment-exchange treatment from a no-exchange adjudication baseline. Yet the study does not establish a matched control with identical token budgets or Kaal's validator incentives, binding rules, and reputation consequences. The correspondence is limited to the comparative design and its measured evaluation setting. Deliberation has evidentiary value only when the baseline preserves the decision it replaces.

**Status.** affirmed  **Published.** 2026-08-08

**Holds when.**

- The response is limited to the exact protocol, baseline, and reported-result passages and the one mapped Kaal claim.
- External evidence level: complete public EACL manuscript bound through ACL Anthology ID, DOI, title, three authors, conference record, PDF hash, HTML identity, extracted text, comparative baseline, reported result, and printed-page locators.
- Mapping review tier: independent substantive scholarly-growth qualification.
- D3 evaluates paired LLM responses rather than Kaal's validation substrate.
- The single-judge baseline does not match D3's agent count, token budget, or deliberation cost.
- The source does not test Kaal's validator eligibility, incentive rules, reputation consequences, or binding implementation.
- The reported comparison supports the treatment-control architecture but does not establish the effects registered by Kaal.
- Semantic Scholar search returned HTTP 429, while the ACL Anthology, Crossref, OpenAlex, and the complete public manuscript supplied stable identity and evidence records.

**Current debate.** Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation: https://aclanthology.org/2026.eacl-long.392/

**Extends.** kaal:claim:7261018-026: https://wulfkaal.github.io/claims/7261018-026

**Scholarly basis.** Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018

**Source PDF sha256.** `1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8`

**Evidence level.** complete public EACL manuscript bound through ACL Anthology ID, DOI, title, three authors, conference record, PDF hash, HTML identity, extracted text, comparative baseline, reported result, and printed-page locators

**Mapping review tier.** independent substantive scholarly-growth qualification

**Mapping confidence.** 0.97  **Mapping ambiguous.** false

**Topics.** ai-and-agents, consensus-and-security, research-methods, scholarly-growth-coverage, scholarly-literature, multi-agent-deliberation, llm-evaluation, controlled-comparison

**Provenance.** Affirmed in kaal-review:2026-08-12:scholarly-growth-7261018-026-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7261018-026.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
