# kaal:position:2026-08-08-328

**Affirmed position.** Scientific evidence is apparatus-bound. Kanewala and Bieman's systematic review explains why. The review reports that scientific software generates evidence for research publications, while also finding that testing is often limited to the initial scientific problem addressed by the code. Reliability on a different problem therefore cannot be guaranteed. This supports treating the completed E1 and E2a evidence as evidence about the historical research apparatus that produced it.

The relationship is a qualification. The review does not examine Kaal's apparatus, autonomous agents, either experiment, or the current reference implementation, and it cannot establish which code, models, or institutional conditions produced Kaal's results. Those facts remain source-bound to Kaal's Article. The external evidence instead supplies the methodological reason for keeping the completed evidence tied to its tested apparatus. A later or modified system requires separate verification before the earlier results can be transferred to it.

**Status.** affirmed  **Published.** 2026-08-08

**Holds when.**

- The response is limited to the exact PMC full-text passages and the one mapped Kaal claim.
- External evidence level: complete peer-reviewed scholarly full text from PMC BioC XML with concordant Crossref, OpenAlex, and Semantic Scholar identities.
- Mapping review tier: independent substantive scholarly-growth qualification.
- The review does not examine Kaal's historical apparatus, autonomous agents, E1, E2a, or the current reference implementation.
- It cannot verify which code, models, or institutional conditions produced Kaal's results.
- The external relationship is methodological and does not independently establish Kaal's source-bound provenance statement.
- Crossref Spanish, Semantic Scholar search, arXiv, and one mapped Semantic Scholar record returned HTTP 429. No rate-limited response was promoted.

**Current debate.** Testing Scientific Software: A Systematic Literature Review: https://doi.org/10.1016/j.infsof.2014.05.006

**Extends.** kaal:claim:7261018-031: https://wulfkaal.github.io/claims/7261018-031

**Scholarly basis.** Wulf A. Kaal, Empirical Evaluation of the Agentic Reputation Substrate: Deliberation, the Composition of Error, and the Registered Measurement of Agency Costs in a Controlled Multi-Model Cohort (2026). SSRN: https://ssrn.com/abstract=7261018

**Source PDF sha256.** `1d6cbe544bd0055133f7cf8ff308be4fde955867bd8dc764992b8d516af15fa8`

**Evidence level.** complete peer-reviewed scholarly full text from PMC BioC XML with concordant Crossref, OpenAlex, and Semantic Scholar identities

**Mapping review tier.** independent substantive scholarly-growth qualification

**Mapping confidence.** 0.97  **Mapping ambiguous.** false

**Topics.** research-methods, institutional-design, scholarly-growth-coverage, scholarly-literature, scientific-software, software-testing, evidence-provenance, reproducibility

**Provenance.** Affirmed in kaal-review:2026-08-12:scholarly-growth-7261018-031-reviewed-v1 at https://wulfkaal.github.io/positions/by-claim/7261018-031.html.

**Record type.** This is a dated commentary position that extends a scholarly corpus claim. It is not a verbatim claim extracted from the paper.

**Canonical form.** This markdown file is the canonical hashed representation of the position.
