The problem: the right number in the wrong sentence
An assistant says an operating mode reduces a machine’s annual energy use by 15%, linking a technical report. The report indeed contains “15%”, but it concerns reduced downtime over thirty days. The citation exists, the document concerns the correct machine, and the number was copied correctly. Yet the conclusion is unsupported: both the measured quantity and time horizon changed.
How can we distinguish a documented answer from one merely decorated with sources? We will construct a four-document synthetic archive and compare three answers. Separating retrieval, claim support and completeness reveals errors hidden by a single percentage. We are not measuring a commercial product: documents, answers and annotations are constructed to make the reasoning inspectable.
Three steps that do not guarantee one another
A RAG system, retrieval-augmented generation, selects passages and provides them to a generator. An agent may search again, open a full document or ask a more precise question. Flexibility does not erase the distinction between finding a source and using it correctly. A passage similar to the question may be obsolete, concern another model or support only half the generated sentence.
Call relevance the relationship to the topic; support the logical relationship between a source and a specific claim. A report relevant to Eco mode does not establish annual energy savings. Completeness is the presence of information needed to answer the question. A perfectly supported one-sentence answer may still omit most of what was asked. These are different questions requiring different checks.
The test archive: small enough to inspect completely
The reader asks: “For current-version P7, what is the continuous temperature limit, what does the Eco pilot document, and how much annual energy does it save?”. P7 is fictional. D1 is current manual v3 and gives an 80 °C continuous limit. D2 is obsolete manual v2, which gave 100 °C. D3 is an Eco report: 30 days observed, 15% downtime reduction, no annual energy estimate and no sample size stated. D4 is a catalogue mentioning Eco without quantified performance.
These four texts are the exercise’s entire information universe. D3 is deliberately incomplete: it does not permit causal benefit estimation or generalisation to other machines. We study whether an answer faithfully reports what is written, not whether the fictional report proves an industrial result. No hidden source authorises 15% energy savings or a thousand-machine sample.
First check: did we retrieve the available evidence?
Assume a manually fixed ranking: D2, D4, D1, D3. It is not output from an executed search engine; it isolates the counting. Valid documents for the requested facts are D1 and D3. Define document recall at k as the fraction of those two in the first k results. At k=2 it is zero despite two topical texts; at k=3 it is 1/2; at k=4 it is 2/2. Lexical relevance alone does not find the right evidence.
The denominator is known because we constructed and inspected the whole archive. In a real corpus, identifying every evidence-bearing source is itself evaluation work. Also, two documents do not mean two facts: D1 covers temperature, D3 both duration and the observed metric. At k=3 we have half the documents but only one-third of three requested documentable facts. Document counting does not automatically measure information coverage.
Second check: does the source support that exact claim?
Consider answer A: “The continuous limit is 80 °C [D1]. Eco reduces annual energy by 15% [D3]. The pilot was validated on a thousand machines [D4]”. All three sentences have citations and all documents open. Only the first is supported. D3 concerns downtime, not energy; D4 gives no sample. Unsupported is not identical to false: the saving might exist in the world, but this answer provides no evidence for it.
Split the answer into atomic claims, independently checkable statements. For claim c_i and its citation set C_i, define S(C_i,c_i)=1 if the sources fully support it in the applicable context, and zero otherwise. Missing citations give zero. In our archive this relation is explicitly annotated; the program does not understand natural language or discover support by itself.
n counts answer claims, not retrieved documents. Citation presence is 3/3, or 100%, while support is 1/3, about 33.3%. If all four documents were retrieved, R_doc is also 100%. Successful-looking retrieval and complete citation presence can thus coexist with poor support. No invented model hallucination rate is needed to demonstrate the problem.
Third check: can we simply delete difficult claims?
Answer B says only: “The continuous limit is 80 °C [D1]”. Its support is 1/1, or 100%. It avoids unjustified additions but says nothing about pilot duration or the observed metric. Define before generation a set G of requested documentable facts: 80 °C limit, 30-day duration and reported 15% pilot downtime reduction. Coverage U is the fraction of G conveyed with support. We do not change G to match what the system chose to say.
Answer C retains the limit and reports thirty days and reduced downtime, citing D1, D3 and D3. It then states that the archive contains no annual energy-savings estimate. Support and coverage for the three positive facts are both 100%. Abstaining on energy explicitly answers the evidence limitation; it is not a fourth performance figure to invent. Coverage concerns three available facts, not a claim that every question component has a quantitative answer.
| Answer | Citation presence | Support S | Coverage U |
|---|---|---|---|
| A | 100% | 33.3% | 33.3% |
| B | 100% | 100% | 33.3% |
| C | 100% | 100% | 100% |
Compare rows in pairs. A and B have equal useful coverage, but B avoids two unsupported claims. B and C have equal percentage support, but C conveys three times the requested facts. An empty answer would not receive 100% support: at n=0 the ratio is undefined and should be recorded as abstention with zero coverage. Fix these conventions in the protocol; otherwise scores may improve while service worsens.

Granularity and multiple citations change the check
The sentence “The limit is 80 °C and Eco saves 15% energy [D1,D3]” contains two claims. Treating it as one block can hide which half fails; splitting gives more precise feedback. Yet segmentation must not turn isolated words into supposed independent facts. Subject, version, quantity, units and conditions must stay connected. The link between 15% and downtime is exactly what poor segmentation could lose.
With multiple sources, some claims require joint evidence. One document defines a code and another reports a result using it; joining them is valid only if identity and context match. Awarding a point for every link is insufficient. Distinguish necessary citations from redundant or irrelevant ones, and inspect the set when no single passage suffices. Our exercise deliberately uses one source per claim and does not implement this broader problem.
What we take from ALCE
Gao, Yen, Yu and Chen’s Enabling Large Language Models to Generate Text with Citations, published at EMNLP 2023, introduces ALCE and separates correctness from citation quality. Sections 3.3 and 5 describe metrics, evaluators and experiments; support is estimated by a natural-language inference model, NLI. Tests use datasets such as ASQA and contemporary model versions with specified passage budgets. This is a historical methodological reference, not a 2026 ranking or a measurement of our archive.
Our count is an explicit teaching simplification, not a full ALCE reproduction. We ran no NLI, neural retrieval or model generation: we stipulated support relations and computed metrics. A real automatic verifier can err on negation, numbers, versions or evidence spanning passages. The authors’ human evaluation belongs to their study; no reader or human-annotator review was performed here.
From the metric to the agent’s next action
The separation also indicates where to intervene. If D3 was not retrieved, telling the generator to be more precise does not supply missing evidence: search, filters or document access need work. If D3 is available but downtime becomes energy, evidence use is the problem. If the answer is faithful but incomplete, check uncovered question components. A generic “wrong answer” label would hide three operationally different causes.
For energy, another search may be reasonable, but it needs a stopping rule: a search-count or cost limit and an obligation to state what remains undocumented. Repeating queries until a favourable-looking sentence appears is not verification. Preserve document identifier and version, the passage used and its link to each claim. If a source changes, these references identify answers needing review, rather than treating a URL as immutable evidence.
What the reproducible experiment establishes
The attached code preserves four documents, five candidate claims, allowed relations and three answers. It calculates recall, support and coverage using sets and expected-result checks. Counting is linear in claim–citation pairs with constant-time support-table lookup; the costly real-world problem is establishing that table correctly. Our program does not solve that semantic problem. No seed is needed because no random numbers are used.
A real evaluation needs representative questions, frozen or versioned sources, recorded answers, segmentation rules and disagreement handling. Separate answerable cases, missing information and conflicting documents. Our three scenarios estimate no enterprise error rate and justify no universal acceptance threshold. They show metrics can diverge even with functioning links, so the protocol must check content rather than link presence alone.
Answer to the opening question
A citation makes a claim checkable only when it leads to applicable evidence that actually supports it. In our archive retrieval and link presence reach 100% with only one-third supported claims. Cutting the answer to one sentence removes errors but omits two-thirds of available facts. The better answer reports what sources permit, preserves units, versions and conditions, and states the annual-energy limit. This checkable relationship among question, claim and evidence makes a document agent useful.
References and transparency
Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen, Enabling Large Language Models to Generate Text with Citations, EMNLP 2023, pp. 6465–6488. Citation-quality, methods and experimental-comparison sections were read. P7 is entirely fictional and concerns no EL-AI customer, installation or product. This article does not establish that an automatic verifier is available in the CMS or other company products.
Gao et al. (2023) — Enabling Large Language Models to Generate Text with Citations.
support = {"c1": {"D1"}, "c2": {"D3"}, "c3": {"D3"},
"c4": set(), "c5": set()}
required = {"c1", "c2", "c3"}
answer = [("c1", "D1"), ("c4", "D3"), ("c5", "D4")]
good = {claim for claim, doc in answer if doc in support[claim]}
print("support:", len(good)/len(answer))
print("coverage:", len(good & required)/len(required))
# The support labels are stipulated for this synthetic corpus.
# This code is not a natural-language entailment model.
Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 28 September 2026.

