Qualification: Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation

Record: kaal:position:2026-08-08-323 · 2026-08-08

Structured assessment changes a decision process only if the assessment arrives before judgment. That sequence matters. Harrasse, Bandi, and Bandi route paired LLM evaluations through advocates, a criteria-based judge, and an independent jury that renders the final verdict. D3 compares that deliberative architecture with a single judge that selects directly. On MT-Bench, D3-MORE reports 85.1 percent accuracy, 12.6 percentage points above the single-judge baseline. This result qualifies Kaal's treatment-control design. It supports separating an assessment-exchange treatment from a no-exchange adjudication baseline. Yet the study does not establish a matched control with identical token budgets or Kaal's validator incentives, binding rules, and reputation consequences. The correspondence is limited to the comparative design and its measured evaluation setting. Deliberation has evidentiary value only when the baseline preserves the decision it replaces.

Affirmed commentary position. This record extends a source-bound scholarly claim but is not a verbatim paper claim.
Holds when
Current debate

Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation

Scholarly basis

kaal:claim:7261018-026
Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018
Source PDF sha256: 1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8

Evidence and mapping

Evidence: complete public EACL manuscript bound through ACL Anthology ID, DOI, title, three authors, conference record, PDF hash, HTML identity, extracted text, comparative baseline, reported result, and printed-page locators
Review tier: independent substantive scholarly-growth qualification
Mapping confidence: 0.97
Mapping ambiguous: false

Topics

ai-and-agentsconsensus-and-securityresearch-methodsscholarly-growth-coveragescholarly-literaturemulti-agent-deliberationllm-evaluationcontrolled-comparison

Provenance

Affirmed in kaal-review:2026-08-12:scholarly-growth-7261018-026-reviewed-v1 on 2026-08-08. Review record.

Verify

Canonical markdown sha256: 3311234caf6a97e23137a22db145e4e149f54cd594a88ffbfc74630f455f10e7
curl -s https://wulfkaal.github.io/positions/2026-08-08-323.md | sha256sum