# kaal:position:2026-08-26-012

**Affirmed position.** Scale is an amortization problem before it is a throughput problem. Chen and her coauthors make this distinction explicit in AgentSlimming. Their starting workflow is a graph of agents and communication edges. The method prunes redundant nodes, substitutes lower-cost models, and retains changes only when baseline performance survives. Across eight benchmarks, the resulting workflows reduced recurring token and API expense while preserving or improving selected task scores. The authors then separate the one-time search and evaluation burden from per-query inference savings and calculate time to break even.

This evidence extends Kaal's claim in a narrower and testable form. A benchmark that reports accuracy and inference cost after workflow selection measures the performance of an arrangement. It does not disclose what it cost to discover, validate, and maintain that arrangement. Scale cannot be inferred from per-query cost alone. The design cost must be amortized over actual use, and a topology change must be charged for the validation required to preserve acceptable behavior.

The limits are material. AgentSlimming studies cloud-model workflow graphs, public reasoning benchmarks, and task-level optimization. It does not examine sovereign local runtimes, latency, institutional coordination, security review, or legal arrangements. It also finds cost reductions for a compression method, not a universal constraint on scaling. Sovereign runtime benchmarks should therefore report design and search effort, validation cost, recurring inference cost, break-even volume, topology changes, and performance loss as separate quantities. The cost of arranging the system belongs in the scaling result.

**Status.** affirmed  **Published.** 2026-08-26

**Holds when.**

- The response is limited to the exact full-text propositions and the one mapped Kaal claim.
- External evidence level: peer-reviewed conference paper with complete official proceedings full text.
- Mapping review tier: independent substantive scholarly-growth extension.
- The source studies cloud-model workflow graphs rather than sovereign local agent runtimes or organizational arrangements.
- Its experiments use public reasoning benchmarks and task-level optimization, not production runtime governance.
- The source measures optimization search cost and recurring API expense, not latency, security review, maintenance, or legal arrangement cost.
- The evidence establishes a useful amortization method for the tested workflows. It does not prove that coordination cost is universally the binding constraint on scale.
- Break-even volume depends on model prices, workflow use, benchmark choice, and acceptable performance loss.

**Current debate.** AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems: https://aclanthology.org/2026.acl-long.1387/

**Extends.** kaal:claim:7314479-012: https://wulfkaal.github.io/claims/7314479-012

**Scholarly basis.** Wulf A. Kaal, Institutional Requirements for Sovereign Local Agent Runtimes (2026). SSRN: https://ssrn.com/abstract=7314479

**Source PDF sha256.** `debace24a155ae924a155b1fafe98856d98cf83689feff2f87a32f1c06171ce6`

**Evidence level.** peer-reviewed conference paper with complete official proceedings full text

**Mapping review tier.** independent substantive scholarly-growth extension

**Mapping confidence.** 0.97  **Mapping ambiguous.** false

**Topics.** economics, research-methods, ai-and-agents, benchmarking, coordination-costs, workflow-optimization, systems-integration, performance-measurement

**Provenance.** Affirmed in kaal-review:2026-08-26:scholarly-growth-7314479-012-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7314479-012.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
