Qualification: A survey of data provenance in e-science

Record: kaal:position:2026-08-26-022 · 2026-08-26

Provenance is not established by recording activity during execution. Simmhan, Plale, and Gannon classify provenance systems by representation, storage, and dissemination. Their classification matters because each function answers a different institutional question. Representation determines what a record says. Storage determines whether the record survives. Dissemination determines whether another system or person can use it. The survey treats persistence as a design problem. Provenance may be embedded with the data or stored separately. The authors explain that a system must decide whether provenance is immutable, updated, or versioned. They also identify archival retention as an answer to storage cost. These choices are not equivalent. An execution trace held only by a running service has not resolved any of them. It may describe what the service observed, but its availability still depends on the service and the session that produced it. Portability and interpretation impose separate conditions. The survey describes derivation graphs, search, retrieval interfaces, and replication at another site as methods for disseminating provenance. It also explains why syntax alone is insufficient. Semantic information, including ontologies that define concepts and relationships, supplies context for interpreting a lineage record. A detailed history permits users to assess whether data is acceptable only when the record preserves enough meaning for that assessment. A transferable file without stable identifiers, defined relations, and contextual terms may be portable as bytes while remaining unusable as evidence. The source is a peer-reviewed survey of scientific data systems published in 2005. It does not study autonomous agents, sovereign runtimes, legal admissibility, or adversarial deletion. It reports a taxonomy and surveyed design choices rather than a controlled comparison. The authors also found no metadata standard that worked across disciplines. The evidence therefore does not establish that one format guarantees durable or universally intelligible provenance. The institutional requirement is narrower. Runtime telemetry should be treated as an input to provenance, not as completed provenance. Before the originating session ends, the relevant record should be sealed into a versioned and retrievable artifact. Its identifiers, relations, time references, and governing vocabulary should travel with it. A party that did not observe the execution must be able to recover the record, verify its integrity, and understand the causal account without depending on the original service.

Affirmed commentary position. This record extends a source-bound scholarly claim but is not a verbatim paper claim.
Holds when
Current debate

A survey of data provenance in e-science

Scholarly basis

kaal:claim:7314479-022
Wulf A. Kaal, Institutional Requirements for Sovereign Local Agent Runtimes (2026). SSRN: https://ssrn.com/abstract=7314479
Source PDF sha256: debace24a155ae924a155b1fafe98856d98cf83689feff2f87a32f1c06171ce6

Evidence and mapping

Evidence: peer-reviewed survey article with complete publisher full text and provenance-system taxonomy
Review tier: independent substantive scholarly-growth qualification
Mapping confidence: 0.99
Mapping ambiguous: false

Topics

institutional-designgovernance-designprovenanceauditabilityinteroperabilitydata-lineagesemantic-contextagent-runtime

Provenance

Affirmed in kaal-review:2026-08-26:scholarly-growth-7314479-022-reviewed-v1 on 2026-08-26. Review record.

Verify

Canonical markdown sha256: 489ef10e15b5d6697601d3ac7f28ba38421fd31aebfaa180e3ec1bebb15c55b9
curl -s https://wulfkaal.github.io/positions/2026-08-26-022.md | sha256sum