Reference

Source Provenance and Claim Traceability in AI Information Systems

Abstract. Source provenance records where information came from and how it was transformed. Claim traceability connects a statement shown to a reader back to the specific evidence that supports it. The two are related but not identical: a document can have excellent provenance and contain a false claim, while a correct claim can be displayed with a weak or fabricated citation. AI-generated summaries, search answers, and news digests need both a transformation history and a claim-to-evidence map.

Coverage note: sources checked through August 2026.

1. Four questions that should not be collapsed

When an AI system presents a factual statement, readers may ask:

  1. Origin: Where did the source material come from?
  2. Integrity: Has the source or record been altered since a trusted step?
  3. Support: Does the cited passage actually support this claim?
  4. Truth: Is the claim accurate in the world?

These questions require different evidence.

Question Typical evidence What it cannot prove alone
Origin Publisher, URL, identifier, signature That the content is true
Integrity Hash, signature, immutable version That the signer was honest
Support Claim-to-passage mapping, entailment check That the source itself is correct
Truth Corroboration, primary evidence, expert method Complete transformation history

An official press release has clear origin and may still contain a selective claim. A cryptographically signed image can be authentic to its editing history and still depict a staged event. A citation can exist while failing to support the sentence beside it.

Trustworthy systems keep these layers separate. The Support layer has its own formal treatment: the Attributable to Identified Sources (AIS) framework defines a generated statement as attributable to a source when a generic hearer, given the source, would affirm "According to that source, this is so" — a pragmatic test that admits ordinary inference rather than the stricter entailment relation used elsewhere in this article — and demonstrates that human raters can apply it reliably — which makes "supported" an evaluable property rather than a stylistic impression. 7

2. Provenance as a data model

The W3C PROV family provides a general vocabulary for describing provenance. Its core distinguishes: 1, 2

  • entities — documents, datasets, chunks, prompts, summaries, or answers;
  • activities — ingestion, editing, extraction, translation, retrieval, or generation;
  • agents — people, organizations, services, or software responsible for activities.

Relationships record that an entity was generated by an activity, derived from another entity, attributed to an agent, or used by a process.

For an AI summary, a minimal provenance chain might be:

source article (entity)
      ↓ used by
claim extraction (activity)
      ↓ generated
claim records (entities)
      ↓ used by
summary generation (activity)
      ↓ generated
summary version (entity)
      ↓ attributed to
model + editorial workflow (agents)

This is more useful than a flat “Sources” list because it preserves transformation.

3. Claim traceability

Claim traceability works at a finer level.

A claim record should include:

  • claim text or a stable claim identifier;
  • type: factual, interpretive, prediction, quotation, or recommendation;
  • supporting source identifier;
  • exact supporting span or structured data rows;
  • publication and retrieval version;
  • support status;
  • contradictions or alternative sources;
  • confidence and review status;
  • which displayed sentences depend on it.

The goal is not to force every sentence to contain a citation marker. It is to ensure that every material factual assertion can be traced internally and that the interface exposes enough of that trace to the reader.

This supports correction. When a source changes or a claim is challenged, the system can find affected summaries instead of searching prose manually.

4. Citation presence is not citation quality

The gap between the two is measured, not conjectural. When Liu, Zhang, and Liang audited commercial generative search engines against the AIS standard, only about half of generated sentences were fully supported by their citations, and about a quarter of citations did not support their associated sentence — while the answers read as fluent and authoritative throughout. 8 The ALCE benchmark formalized the evaluation for LLMs directly, scoring citation recall (is every claim supported?) and citation precision (does every citation earn its place?) alongside fluency and correctness, and found the shortfall concentrated in citation quality: even the best systems lacked complete citation support about half the time on one of its datasets, while fluency was not where the headroom was. 9

AI systems can produce citations that fail in several ways:

  • the URL does not exist;
  • the source exists but the cited passage does not;
  • the passage concerns a related topic but not the claim;
  • the source supports only part of a compound sentence;
  • the source is secondary when the prose implies primary evidence;
  • the citation points to a later summary of the same claim;
  • one source is reused for claims it does not support;
  • the citation marker moves during rewriting.

Citation evaluation therefore needs at least:

Validity. Does the referenced source and location exist?

Entailment or support. Does the evidence support the claim as written?

Completeness. Are material claims supported?

Source quality. Is the source appropriate for the claim?

Attribution. Is credit assigned correctly?

Freshness. Is the cited version current enough?

Benchmarks such as FACTS Grounding evaluate whether long-form responses are accurate with respect to a supplied document. The benchmark’s design reflects how difficult it is to evaluate long-form grounding at scale; it uses multiple automated judges and a disqualification stage for responses that fail the task. 3 Grounding to a document is important, but it still does not establish the document’s real-world truth.

5. Provenance for retrieved generation

In Retrieval-Augmented Generation, the final answer depends on a pipeline:

  1. source ingestion;
  2. parsing or OCR;
  3. chunking;
  4. metadata attachment;
  5. embedding or indexing;
  6. query transformation;
  7. retrieval;
  8. re-ranking;
  9. context selection;
  10. generation;
  11. citation rendering.

A failure at any stage can produce a wrong answer with a plausible citation.

A provenance-aware RAG system retains:

  • stable document and version IDs;
  • content hashes;
  • canonical URL and publisher;
  • license or access conditions;
  • parser and chunk boundaries;
  • index and embedding versions;
  • query and retrieved candidates;
  • selected passages;
  • prompt and model version;
  • output claims and evidence links;
  • validation result.

The visible answer may show only title, link, and highlighted passage. The deeper trace supports debugging and audit.

6. Provenance for news and media

The Coalition for Content Provenance and Authenticity (C2PA) defines Content Credentials: signed manifests that can record an asset’s origin and editing history. The specification is designed for interoperable, tamper-evident provenance across media workflows. 4

C2PA’s own explainer states an essential limitation: provenance can provide evidence about origin, history, and authenticity, but provenance alone cannot tell whether content is true or factually accurate. 5

This is a model for broader AI information systems:

  • cryptographic provenance answers whether a record and its asserted history validate;
  • editorial verification answers whether the represented event or claim is credible;
  • claim traceability answers what the system’s prose relied on.

The three reinforce each other without becoming the same thing.

7. Source hierarchy and independence

Several links do not necessarily mean several independent sources.

Ten articles may reproduce:

  • one company announcement;
  • one wire-service report;
  • one preprint;
  • one social-media post;
  • one anonymous claim.

A curation system should model derivation and independence. Useful source roles include:

  • primary research or technical artifact;
  • direct observation or official record;
  • original reporting;
  • independent analysis;
  • commentary;
  • aggregation;
  • copied or syndicated coverage.

The system should identify the earliest available source and whether later sources add independent evidence.

This avoids “citation counting” as a substitute for corroboration.

8. The provenance of the model, not only of the sources

Everything above traces the sources a system retrieves. There is a second corpus in every such system whose provenance is usually undocumented: the data the model was trained on. It matters here for a specific reason. When a model answers from parametric knowledge rather than from a retrieved passage, the provenance question has not disappeared; it has become unanswerable.

Two lines of work make this tractable rather than merely lamentable.

Dataset documentation. Gebru and colleagues proposed datasheets for datasets by analogy to the electronics industry, where every component ships with a datasheet describing operating characteristics, test results, and recommended uses. Their proposal is that every dataset carry a document recording its motivation, composition, collection process, and recommended uses, with the explicit aim of improving communication between dataset creators and consumers. 11 The proposal dates to 2018 and the practice is still uneven: a 2024 audit of Hugging Face found 24,065 dataset repositories, 14,011 of them carrying a dataset-card file, and 6,578 of those cards empty — so “only 30.9% (7,433 out of 24,065)” had a non-empty card at all, and those 7,433 are what the analysis covers. Within them, completion varied sharply with a dataset's popularity: the well-used datasets documented and the long tail largely not, which is itself informative about where the incentives sit. The 69% with nothing to analyse is the more striking number. 19

Dataset lineage at scale. The Data Provenance Initiative audited more than 1,800 text datasets with a multi-disciplinary team of legal and machine-learning researchers, tracing each dataset's source, creators, license conditions, properties, and subsequent use. Two findings are directly relevant to an information system built on such models. Licenses are frequently miscategorized on aggregator sites relative to what the original source states, meaning the license a practitioner reads is often not the license that applies. And the composition of commercially open and closed datasets has diverged sharply, with closed collections dominating lower-resource languages, more creative tasks, richer topic variety, and newer or synthetic data. 12

For the four questions in section 1, this adds a fifth that is easy to skip:

Where did the model's own prior come from, and can that be stated?

Three practical positions follow:

  • Distinguish retrieved answers from parametric ones in the trace. A claim supported by a fetched document has provenance; a claim the model produced from training has, at best, a model identifier and a version. Recording which mode produced a sentence is cheap at generation time and expensive to reconstruct later: a replay from the retained prompt, model version and passages (section 5) can recover it in principle, but nobody has measured how often that recovery is right. Section 4's numbers are about a different question — whether the citations a system already displays support the sentences they are attached to — and they are not reassuring about the neighbouring problem even though they do not measure it. Record the mode at generation time; the reconstruction is unvalidated, not merely expensive.
  • Model version is a provenance field. The same prompt against two checkpoints is two different sources. A trace that records the retrieved documents but not the model build is missing the component most likely to have changed.
  • Aggregator metadata is a claim, not a fact. A license field copied through an intermediary and then treated as authoritative is its own failure mode, distinct from section 14's broken-source-identity case: there the identifier is unstable, here the identifier is fine and the value attached to it was never checked against the source that issued it.

9. Versioning and time

Fast-moving AI claims change:

  • a preprint is revised;
  • a benchmark fixes its scorer;
  • a leaderboard updates;
  • a model card adds limitations;
  • a news report is corrected;
  • a repository changes its defaults.

A URL is not a version.

Provenance should preserve:

  • publication date;
  • event or measurement date;
  • retrieval date;
  • document version or commit;
  • effective date;
  • correction date;
  • the date the summary was reviewed.

A correction should create a new summary version linked to the old one. Silent overwrite destroys the evidence needed to understand what readers previously saw.

10. A claim-centered editorial workflow

A practical pipeline can use claim records as the interface between research and prose.

Ingest

Capture source identity, version, publisher, date, and content hash.

Extract candidate claims

Separate factual claims from interpretation and prediction.

Attach evidence

Link each factual claim to exact passages, tables, code, or data.

Check support

Use deterministic validation for links and spans, model-assisted entailment for triage, and human review for important claims. The claim-against-evidence verdict structure — supported, refuted, or not enough information, judged against retrieved evidence — is the design the FEVER fact-verification tradition standardized at dataset scale, and it transfers to editorial pipelines more or less intact. 10

Detect contradiction

Retrieve sources that disagree or bound the claim differently.

Draft

Generate prose from the claim set rather than from an unstructured pile of documents.

Validate again

Re-extract claims from the draft and compare them with the approved claim records. Rewriting can introduce stronger language than the evidence.

Publish with a trace

Show clear citations and preserve the internal graph.

Monitor

When a source changes, identify dependent claims and articles.

This workflow treats citations as data rather than decorative text.

11. Claim types need different evidence

Not every sentence should be evaluated with the same rule.

Direct factual claim

The benchmark contains a stated number of tasks.

Best evidence: the benchmark paper, dataset card, or repository version that defines the release.

Comparative claim

System A outperformed system B.

Best evidence: the result table plus the settings, uncertainty, and whether the comparison was controlled.

Causal claim

The architecture change caused the performance gain.

Best evidence: ablation, experiment, or other design that isolates the mechanism. A before-and-after product announcement is weaker.

Interpretive claim

The project’s deeper contribution is benchmark lifecycle design.

This is analysis. Sources can support the underlying facts, but the interpretation should be labeled as the article’s synthesis rather than attributed to a paper that did not make it.

Prediction

This pattern will become common.

The source can establish trends and constraints; the forecast remains uncertain and date-bound.

Normative claim

The system should preserve a common, non-personalized edition.

This requires an argument about values and tradeoffs. A user-engagement study cannot prove the “should.”

Typing claims prevents citation laundering across categories.

12. Automated checks and their limits

Some checks can be deterministic:

  • URL parses and resolves;
  • document ID exists;
  • quoted text matches the recorded span;
  • citation marker refers to an allowed source ID;
  • source hash matches the indexed version;
  • every factual claim has at least one evidence link;
  • a revision did not drop citations silently.

Other checks are model-assisted:

  • does the passage entail the claim?
  • is the claim stronger than the evidence?
  • do two sources contradict?
  • is the source primary?
  • did the summary omit a material qualifier?

Model-assisted checks should produce review candidates, not invisible truth labels. The verifier may share the generator’s blind spots or be persuaded by the same fluent wording.

TREC’s RAG evaluation work reflects this layered view by evaluating response completeness and attribution alongside retrieval and answer quality, with explicit citation processes in the track design. 6

For important content, a human reviewer should see the claim, exact source span, source context, and any contradiction—not only a pass/fail score.

13. Correction propagation

Suppose a technical report revises one benchmark number.

A provenance graph can identify:

  1. the changed source entity;
  2. extracted claims derived from that version;
  3. summaries using those claims;
  4. cards, digests, or wiki entries displaying them;
  5. downstream recommendations that used the old classification.

The correction workflow can then:

  • mark affected claims stale;
  • suppress automatic republication;
  • generate a candidate correction;
  • require review;
  • publish a new version;
  • link old and new;
  • record when the visible correction reached each surface.

Without dependency links, corrections rely on memory and search. The system may fix one page while leaving the same error in a digest and model context.

14. Failure modes

Provenance theatre

The system stores elaborate metadata but does not check whether the displayed claim is supported.

Signed falsehood

A valid signature is presented as a truth verdict.

Citation laundering

An AI answer cites a reputable source for wording the source never asserted.

Compound-claim blur

One sentence contains several claims and one citation supports only the easiest one.

Stale dependency

The source is updated, but derived summaries are not flagged.

Broken source identity

Documents are indexed by mutable URL or title, so revisions cannot be distinguished.

Missing negative evidence

The system records support but not contradiction, retraction, or failed replication.

Private-source exposure

Provenance metadata reveals a confidential source, user, or document path. Traceability must respect access control and source protection.

Model-generated authority

The system cites its own earlier summary, creating a circular chain with no primary evidence.

15. Watermarking is not provenance

Watermarking is frequently discussed alongside content credentials, and the two are easily conflated — not least because the C2PA specification itself now combines them. The distinction is worth stating precisely, because a watermark that points back to a manifest and a watermark that stands alone as an origin signal are doing different jobs.

Provenance metadata, as in section 6, is an assertion attached to an asset: signed, verifiable, and removable. The specification knows this and answers it with what it calls durable Content Credentials: a hard binding by cryptographic hash plus a soft binding by watermark or fingerprint, so that a manifest stripped from a file can still be rediscovered from the content itself. 5 That description reads against the 2.2 explainer, which is two releases behind this article's coverage date: 2.3 landed in December 2025 and 2.4 in April 2026, the latter adding asset-format and assertion support the durable-credential story here does not depend on but a reader building against the spec should start from. 20 Used that way, a watermark is a lookup key back to provenance. Used on its own, it is something weaker — a statistical signal inside the content, designed to survive stripping and to say only which generator the content came from — and that weaker use is what the rest of this section is about. Kirchenbauer and colleagues' scheme for language models is the reference construction: a randomized set of "green" tokens is selected before each word is generated and softly promoted during sampling, with a statistical test yielding interpretable p-values, and detection possible from a short span of tokens without access to the model's parameters or API. 13

SynthID-Text demonstrated the production version. It modifies only the sampling procedure, leaves training untouched, integrates with speculative decoding so it is viable at serving speed, and was evaluated in a live experiment covering feedback from nearly 20 million Gemini responses without measurable change in text quality on standard benchmarks or human side-by-side ratings. 14 That is a genuine engineering result: the objection that watermarking must degrade output has been answered empirically.

The objection that has not been answered is robustness, and two results bound what may be claimed:

  • Paraphrase attacks. Sadasivan and colleagues stress-tested a range of detectors (watermarking, neural detectors, zero-shot classifiers, and retrieval-based schemes) with a recursive paraphrasing attack on passages of roughly 300 tokens, and found detection rates reduced substantially while text quality degraded only slightly in many cases. 15
  • An impossibility result. Zhang and colleagues prove that under stated and natural assumptions, strong watermarking is impossible: a scheme a computationally bounded attacker cannot erase without significant quality loss cannot be built, including in the private-key setting where the attacker knows neither the key nor which scheme is in use. Their attack requires only access to a quality oracle and a perturbation oracle, and does not require knowing the scheme. 16

The correct reading is neither "watermarking is useless" nor "watermarking solves attribution":

  • It is asymmetric evidence, and neither direction is proof. A missing watermark is nearly no evidence at all, because removal is cheap. A detected watermark is the stronger of the two signals, but it is not conclusive either: the same paper that ran the paraphrase attack also demonstrated spoofing — inferring a scheme's hidden signature without white-box access and stamping it onto human-written text, so that a detector attributes to the generator something it never produced. 15 Detection raises the probability of that origin; it does not settle it.
  • It addresses accidental provenance loss, and that is the case worth designing for. Text that lost its attribution by being copied, quoted and re-hosted is the population a watermark can still speak for; text an adversary worked on is not. Which of those two is the larger share in the wild is not something this article can source, and the design argument does not need it: the adversarial share is the one strong watermarking provably cannot cover, so it is the share that has to be handled another way regardless of its size. The proof is narrower than the slogan and the narrowness matters: it assumes an attacker with a quality oracle and a perturbation oracle “which induces an efficiently mixing random walk on high-quality outputs”, and the authors set weak schemes aside explicitly — “We do not investigate weak watermarking schemes in this paper” — while noting they “can still be useful”. A watermark that resists a specified set of transformations is not covered by the impossibility, and is also not what a determined adversary is up against.
  • It does not establish truth, and neither does provenance. The limitation section 6 quotes from C2PA applies unchanged: knowing where something came from says nothing about whether it is correct.

16. Evaluation

A provenance system should be tested with deliberately difficult cases:

  • moved or deleted sources;
  • changed documents at the same URL;
  • real citations with non-supporting passages;
  • compound claims with partial evidence;
  • syndicated coverage;
  • contradictory sources;
  • retracted or corrected material;
  • private sources with redacted public traces;
  • screenshots or copied text without reliable origin;
  • model-generated citations not in the retrieval set.

Metrics include:

  • valid-link rate;
  • passage-location accuracy;
  • claim-support precision and recall;
  • primary-source rate;
  • independent-source count;
  • version completeness;
  • correction propagation time;
  • unsupported-claim rate;
  • reader success at opening and understanding the evidence.

17. Retraction, and knowing that a source moved

Section 13 describes propagating a correction once the system knows one is needed. The prior question, how the system finds out, is a separate problem with its own literature, and that literature is not reassuring.

Retraction is the strongest correction signal the scholarly record produces, and it propagates badly. Bar-Ilan and Halevi queried ScienceDirect in October 2014 for retracted articles — 987 of them, retracted between 1995 and 2014 — then selected the 15 that went on to receive more than ten citations between January 2015 and March 2016, and manually coded 238 citing documents by the context of each citation — a case study of heavily-cited retractions rather than a census. The result: the vast majority of post-retraction citations were positive, citing the retracted work as valid support, despite a clear retraction notice on the publisher's platform, and regardless of the reason for retraction, including retractions for data fabrication and false reporting. 17 The study tests no pipeline and offers no human-versus-automation comparison, so the obvious next sentence — that an automated pipeline which fetched the PDF once and never revisited it would do worse — is a prediction, not a result. It is the prediction this section's design advice is built on, and it is worth saying that it rests on the mechanism (nothing in a fetch-once pipeline looks again) rather than on a measurement.

The infrastructure response is a standardized signal rather than a per-publisher convention. NISO's recommended practice on the Communication of Retractions, Removals, and Expressions of Concern specifies how such notices should be expressed and distributed so that downstream systems can consume them consistently. 18 Whether an individual pipeline consumes that signal is a build decision, and it is one that separates a provenance system that works from one that merely records.

One further asymmetry is worth naming, because it decides where the engineering effort goes. The two failures this section has to defend against are not alike. A retraction notice is published: it exists somewhere and it is addressable, which is already more than a silent edit offers — an edit that removes or alters a claim without announcing itself leaves no signal at all. What it is not is guaranteed to be machine-findable. The format is recommended, not required — NISO RP-45-2024 is a Recommended Practice whose “primary aim … is to establish best practices for metadata creation, transfer, and display”, and a practice a publisher may decline to follow is not a property a consuming system can assume. So the first problem is a matter of subscribing to the right feeds and matching identifiers where the publisher emits them, and of a fallback where it does not — a step this article's sections do not otherwise cover, since section 12's checks are per-fetch and section 13 begins after the system already knows a correction exists. The second needs the system to have kept something to compare against — a content hash of what was actually fetched is the cheapest such thing, and section 5's retained passages are another. A system that builds the first and skips the second is defended against the failure that announces itself and undefended against the one that does not. Which of the two is more frequent in a given corpus is an empirical question this article has no number for; the asymmetry in detectability is the reason to build both regardless.

Four consequences for the claim-centered workflow in section 10:

  • A source reference is a subscription, not a snapshot. Storing a DOI without ever re-resolving it means the system's belief about the source is frozen at fetch time. Re-checking status is a scheduled job, not an event handler.
  • Retraction is a claim-level event, not a document-level one. A retracted paper may have supported five extracted claims across three surfaces. Section 13's dependency graph is what turns one upstream notice into the right five invalidations.
  • Silent removal is worse than retraction. A withdrawn preprint, a deleted post, or a quietly edited page produces no notice at all. Content hashing at fetch time is the cheapest way to detect that the thing now at the URL is not the thing that was cited; retaining the fetched text itself (section 5) is the other.
  • Record the retraction, do not just delete the claim. A claim withdrawn because its source was retracted is different from a claim that was never made, and section 14's missing-negative-evidence mode is what happens when the system erases rather than annotates.

18. Bottom line

Provenance does not make a claim true, and a citation does not make a sentence supported.

A trustworthy AI information system needs a chain:

displayed claim
    → supporting evidence span
    → versioned source
    → transformation history
    → responsible process

That chain lets readers inspect, editors correct, and engineers debug. It turns “trust us” into “here is what this statement depends on.”

Knowledge systems: Retrieval-Augmented Generation, Incremental Indexing, Information Overload and AI Curation, Evaluation Harness

Trust: Trust Calibration in Human-AI Systems, Benchmark Contamination, Distribution Shift, Factuality Evaluation

Standards: W3C PROV, C2PA Content Credentials, Data Lineage, Model Cards