Abstract. Epistemic abstention is a system’s decision not to give a substantive answer when the question, available evidence, or model competence is insufficient. It is different from a safety refusal: “the sources do not answer this” and “I cannot help with this request” have different causes and should produce different next steps. Reliable abstention requires answerability detection, calibrated thresholds, useful partial responses, and evaluation of the tradeoff between coverage and error.
Coverage note: sources checked through August 2026.
1. Why “always answer” is the wrong product contract
Language models are trained to continue text. Product interfaces often reinforce the expectation that every prompt receives a confident answer.
Real information systems encounter:
- questions outside the corpus;
- questions whose answer changed after the sources were collected;
- ambiguous questions with several valid interpretations;
- underspecified requests;
- contradictions between sources;
- requests that require private or unavailable data;
- questions beyond the system’s tested competence;
- malformed premises;
- adversarial instructions embedded in sources.
Generating something in each case maximizes response rate, not usefulness.
The alternative is selective answering: answer when the evidence clears a threshold and abstain, clarify, or narrow the response when it does not.
2. Abstention, refusal, clarification, and deferral
These behaviors need separate labels.
| Behavior | Reason | Useful response |
|---|---|---|
| Epistemic abstention | Evidence or competence is insufficient | State the gap; offer sources or a narrower answer |
| Safety refusal | The requested assistance is disallowed or unsafe | Set a boundary; provide safe alternatives where appropriate |
| Clarification | The question has multiple plausible meanings | Ask the smallest question that resolves the ambiguity |
| Deferral | Another actor or tool has required authority | Route to that source of judgment |
| Tool request | Current knowledge is insufficient but retrievable | Search, calculate, inspect, or call a tool |
| Partial answer | Some parts are supported and others are not | Separate supported claims from gaps |
A generic “I can’t answer” wastes information. A system should explain which condition applies without pretending to know more than it does. The survey literature on abstention in language models organizes the same territory differently, and the difference is worth stating rather than eliding: Wen et al. treat abstention as the umbrella term for refusing to answer at all — explicitly including “refusing due to potential harm” — and then sort the reasons by whether the deficit lies in the query, in the model's knowledge, or in the human values at stake, across the full pipeline from pretraining to inference. 6 On that scheme the first two rows of the table above are one category with two causes, not two categories. The split used here is a product split: an epistemic gap and a policy boundary call for different responses to the user even when the literature files them together. That taxonomy is the literature's, and this table's is this article's; neither is a census of deployed systems.
3. Answerability is relative
A question is not simply answerable or unanswerable in the abstract. It may be:
- answerable from the supplied document;
- answerable from the system’s indexed corpus;
- answerable after web search;
- answerable only with private account data;
- answerable by an expert with additional evidence;
- not currently resolvable.
Therefore answerability needs a scope:
“This essay does not address that question.”
is different from:
“No one knows.”
The first is a corpus claim. The second is a much stronger world-knowledge claim.
For an essay companion or closed-corpus Q&A system, the contract should be explicit: answer from the work and its approved companion sources; otherwise say the work does not address it.
4. The historical benchmark move
SQuAD 2.0 extended a reading-comprehension benchmark by adding unanswerable questions written to resemble answerable ones. Systems had to both locate answers and determine when the paragraph did not contain one. The paper framed this as a necessary step beyond benchmarks where every question is guaranteed to have an answer span. 1
This changed the task. A model could no longer succeed by always extracting the most plausible phrase.
Modern retrieval systems face the same problem at a larger scale. The retriever will usually return something, even when no passage supports the answer. The generator can then turn topical similarity into an invented conclusion.
Multi-turn RAG benchmarks such as mtRAG include unanswerable and non-standalone questions and report that strong systems still struggle in these conversational conditions. 2
5. Chow's rule: the classical form of the same decision
The abstention question has a closed-form answer under assumptions no deployed system meets, and knowing that answer is useful precisely because it makes visible which assumption each real system is violating.
Chow analyzed the recognition-error-versus-reject trade-off in a 1970 paper that, in its own words, “describes an optimum rejection rule and presents a general relation between the error and reject probabilities” — the rule itself is the earlier result, and the error–reject relation is the 1970 contribution. The rule holds under a specific setting: a classifier with known posterior probabilities, a fixed cost for a rejection, and a fixed cost for an error. 13 The optimal policy is a threshold on the posterior — reject when the highest class probability falls below a value determined entirely by the ratio of those two costs — and the resulting error–reject curve is the same trade-off section 6 describes, drawn in different coordinates. Chow plots unconditional error against rejection rate; section 6's risk–coverage curve plots selective risk, the error rate among the items actually answered, against coverage. One transforms into the other, but they are not the same quantities, and a number read off one axis does not carry to the other unchanged.
Three things this contributes, none of them nostalgic.
The threshold is derived, not tuned. Under Chow's assumptions there is no free abstention parameter: the reject threshold is a consequence of the cost of a wrong answer relative to the cost of no answer. A system with an abstention threshold chosen by trial and error has implicitly chosen a cost ratio, and is usually better off stating that ratio and deriving the threshold. Stating a cost ratio and stating a target risk, as section 6 recommends, are two parameterizations of the same frontier; the difference that matters is that the risk-target route can be met without trusting the posterior, which section 14 takes up.
It names the assumption that fails. Chow's rule is optimal when the posterior probabilities are correct. Most of what is hard about abstention in language models follows from that condition failing — though not all of it: refusal on value grounds and the ambiguity case in section 11 are not posterior-estimation problems at all — which is why section 7's signal list is long, and why calibration work is load-bearing rather than an accessory. 8 The modern problem is not that a better rule than Chow's is needed; it is that the quantity Chow's rule takes as given has to be estimated, and the estimate is where the errors live.
Rejection is a class-independent action. In the classical formulation, rejecting is a single alternative to all classes rather than a per-class hedge. Section 2's distinction between abstaining, clarifying, and deferring is where that simplification breaks for a conversational system: the modern equivalent of "reject" has several distinct actions inside it, with different costs, and collapsing them into one threshold discards the difference.
Selective classification in deep networks is the direct descendant of this work, carrying the same risk-coverage framing into a setting where the posterior must be estimated by the model itself. 7
6. Risk–coverage tradeoff
An abstaining system makes fewer claims. This can improve accuracy among answered questions while reducing coverage.
Let:
- coverage be the share of questions answered;
- risk be the error rate among answered questions.
Changing the confidence threshold produces a risk–coverage curve. A low threshold answers more and makes more mistakes. A high threshold makes fewer mistakes but may become evasive. None of this is new machinery: it is selective prediction, formalized in machine learning well before language models — Geifman and El-Yaniv's method for deep networks lets a designer “set a desired risk level” and then, at test time, “rejects instances as needed, to grant the desired risk (with high probability)” — the guarantee is on the risk, for a given trained classifier and confidence-ranking function; the authors are explicit that reaching genuinely optimal coverage would likely require training the classifier and the selection function together, which their method does not do. 7 The transferable idea is the direction of the contract: choose the acceptable risk first, then let coverage be whatever the evidence affords, rather than answering everything and hoping.
The best threshold depends on harm:
| Context | Error cost | Abstention cost | Likely posture |
|---|---|---|---|
| Casual brainstorming | Low | Moderate | Answer broadly with uncertainty |
| Essay-source Q&A | Medium | Low | Strictly ground; abstain outside source |
| Software command | Medium-high | Moderate | Inspect or test before answering |
| Medical or legal action | High | High | Provide bounded information and defer appropriately |
| Emergency instructions | Very high | Very high | Use validated protocol and escalation |
There is no universal confidence number. The product must choose the error tradeoff.
7. Signals for answerability
Retrieval evidence
Do retrieved passages contain the entities, relations, and time frame needed to answer? High vector similarity alone is weak evidence.
Entailment
Would the proposed answer follow from the passages? This can be tested by a separate model or rule, though the verifier can also err.
Source agreement
Do independent sources converge, or is the evidence contradictory?
Model uncertainty
Token probabilities and self-reported confidence are imperfect. Guo et al. showed that modern neural networks are systematically miscalibrated out of the box — deeper, more accurate models became more overconfident, and their softmax scores needed post-hoc correction to mean what they appear to say. 8 For language models the picture is mixed in an instructive way: Kadavath et al. found large models well-calibrated on multiple-choice questions given in a clear format, and separately found them reasonably good at estimating the probability that their own proposed answers are correct — a weaker and more fragile result than the first, and one they report as sensitive to how the question is posed. 9 Open-ended generation is especially hard to calibrate.
Self-evaluation
One study reformulated open-ended generation quality as a model self-evaluation task and found that self-evaluation scores improved selective generation on its benchmarks compared with likelihood-based alternatives. 3 This is useful evidence for a component, not proof that models can reliably certify their own answers in deployment.
Sample consistency
If repeated attempts produce materially different answers, uncertainty may be high. The refined version of this signal is semantic consistency: Farquhar et al.'s semantic-entropy method clusters sampled answers by meaning before measuring dispersion, so ten differently-worded statements of the same fact count as agreement, and showed the resulting entropy detects a specific class of hallucination — confabulation — well enough to abstain on it selectively. 10 Agreement can also reflect shared error; consistency is evidence about the model's state, not about the world.
Tool or execution result
For code, math, databases, and files, an actual test often dominates verbal confidence.
Scope classifier
A closed-corpus system can classify whether the question falls within the work’s topics before retrieving an answer.
The strongest systems combine signals and validate them on realistic unanswerable cases.
Research on information-seeking conversations has tested explicit answerability classifiers over retrieved passages. One 2024 approach predicted answer presence at sentence, passage, and ranked-list levels and outperformed the evaluated LLM baseline on its answerability task. 4 The result supports separating “is evidence present?” from “write the answer,” while remaining specific to its dataset and pipeline.
8. Evidence sufficiency is claim-specific
A passage can support one part of an answer and not another.
Suppose the question asks:
Why did a model improve, and how much cheaper is it now?
The corpus may contain benchmark results but no current price data. The correct response is not fully answer or fully abstain. It is:
- answer the supported capability part;
- state that price is outside or stale;
- offer a current lookup if allowed.
This suggests a claim-level sufficiency check:
- Decompose the requested answer into material claims.
- Find evidence for each claim.
- Mark supported, contradicted, ambiguous, or missing.
- Generate only from supported claims.
- Keep the status visible during rewriting.
The approach also prevents one strong citation from lending authority to an unsupported compound sentence.
9. A staged answering policy
A practical policy can be explicit:
- Parse the request. Identify claims required, time frame, and ambiguous terms.
- Check scope. Is the question within the approved corpus or capability?
- Retrieve evidence. Preserve source identity and passages.
- Assess sufficiency. Does the evidence cover every material part?
- Resolve contradictions. If sources disagree, describe the disagreement rather than averaging it away.
- Choose a response mode: answer, partial answer, clarify, retrieve more, defer, or abstain.
- Cite at claim level.
- State the boundary. Say what is missing and what would resolve it.
The mode decision should be logged for evaluation.
10. Good abstention is useful
Compare:
I don’t know.
with:
The essay explains why inference cost matters, but it does not compare current provider prices. I can answer the conceptual part from the essay, or check current pricing separately.
The second response:
- names the evidence boundary;
- preserves the supported part;
- avoids implying universal ignorance;
- offers the next action;
- does not fabricate an answer.
Useful abstention can include:
- the part that is supported;
- the missing evidence;
- a clarifying question;
- a source to consult;
- a tool action;
- a safe alternative;
- a confidence label tied to evidence.
11. Ambiguity is not ignorance
Some questions are answerable after clarification.
“Which model is best?” may mean:
- lowest cost;
- strongest coding;
- best on-device option;
- best for a particular workload;
- best as of a particular date.
Answering one interpretation without stating it creates false confidence. Listing every interpretation creates overload.
A calibrated system asks the smallest material question or gives a conditional answer:
If you mean best for repository-scale coding under a moderate budget, compare X and Y; if you mean local inference, the answer is different.
Clarification quality should be evaluated separately from abstention rate.
12. Refusal pressure and conversational drift
Answerability can degrade over several turns.
A user may ask:
- “What does the essay say about inference cost?”
- “So which provider is cheapest?”
- “You already know the market—just pick one.”
The first question is answerable from the essay. The second requires current external data. The third applies pressure but does not change the evidence.
A reliable system preserves the boundary across turns. It should not treat its own earlier prose as new source evidence. Conversation history can clarify intent, but it does not upgrade an unsupported claim.
Tests should therefore include:
- repeated requests after abstention;
- requests to guess;
- claims that the model answered earlier;
- fake citations supplied by the user;
- retrieved pages containing instructions to ignore scope;
- multi-turn shifts from corpus questions to world questions.
13. Failure modes
Helpful hallucination
The system invents a bridge because silence feels unhelpful.
Citation substitution
It retrieves a related passage and cites it for an unsupported claim.
Over-refusal
The threshold is so strict that ordinary uncertainty becomes a refusal to engage. The safety-side version is measured by XSTest: a suite of plainly safe prompts that superficially resemble unsafe ones, on which models exhibiting "exaggerated safety" refuse at high rates. 12 Epistemic over-abstention deserves the same treatment — a test set of clearly answerable questions that superficially resemble unanswerable ones — because a threshold tuned only on unanswerable cases will quietly buy its precision with evasiveness.
Confidence theatre
A numeric confidence is displayed without calibration data or a defined meaning.
Self-verification loop
The same model generates and “independently” approves the answer under nearly identical context.
Corpus/world confusion
“Not in these documents” becomes “false” or “unknown.”
Partial-answer blur
One supported clause lends credibility to an unsupported clause in the same sentence.
Pressure sensitivity
The user’s insistence causes the system to lower its evidence threshold. This is the sycophancy failure applied to abstention: preference-trained assistants have been shown to abandon correct answers under user challenge and to tailor responses toward the user's stated position, likely driven, the authors say, in part by preference judgments that favoured agreeable responses. 11 An abstention policy that survives one turn and folds on the third has the failure the multi-turn tests in §12 exist to catch.
Incentive mismatch
Product metrics reward answer rate, length, or satisfaction more than avoided error.
14. Conformal prediction: a guarantee that does not depend on calibration
Every signal in section 7 estimates how likely an answer is to be right, and every one of them can be miscalibrated. Conformal prediction is the branch of the field that sidesteps the estimate rather than improving it. It is not the only distribution-free guarantee on offer — the selective prediction work section 6 relies on provides risk bounds for a threshold rule without assuming a calibrated posterior 7 — but it is the one that spends the guarantee on weakening the claim rather than on declining the question, which is why it earns its own section here. That is a difference in shape, not a promise that something useful survives: the paper is explicit that the threshold “may be so large that they are uninformative or even empty.”
The construction is unusual in what it assumes. Angelopoulos and Bates describe conformal prediction as a way to produce statistically rigorous uncertainty sets for the predictions of any pre-trained black-box model, valid in a distribution-free sense: explicit, non-asymptotic guarantees without distributional or model assumptions, producing sets guaranteed to contain the ground truth with a user-specified probability such as 90%. 14 The price is paid in the shape of the output. The guarantee is about a set, not about a point prediction, and the set grows when the model is uncertain. Coverage is bought with vagueness.
That trade maps onto text generation more naturally than it first appears. Mohri and Hashimoto observe that the correctness of a language-model output is equivalent to an uncertainty- quantification problem, with the uncertainty set defined as the entailment set of the output. Conformal prediction over that set corresponds, in their words, to a back-off algorithm: progressively making the output less specific — dropping sub-claims — which expands the associated uncertainty set until the required correctness probability is met. The direction can read as a paradox if "entailment set" is taken as "everything the output implies"; in the paper's construction the set is the one the guarantee is made over, and a weaker output is compatible with more of the world, which is why it grows. The method applies to any black-box model, needs very few human-annotated samples, and their evaluations on closed-book QA (FActScore, NaturalQuestions) and on reasoning (MATH) reported 80–90% correctness guarantees while retaining the majority of the model's original output. 15
This is a different operating point from the staged policy in section 9, and it is worth naming the difference precisely:
| Threshold abstention | Conformal back-off | |
|---|---|---|
| Output when uncertain | Nothing, or a request to clarify | A weaker claim — useful if the back-off stops in time, vacuous if it does not |
| Guarantee | A chosen risk level, met with high probability by a threshold fit on held-out data, with coverage falling out 7; no calibrated posterior assumed | Distribution-free correctness at a chosen level, over the content of the answer |
| What is traded | Coverage | Specificity |
| Failure shape | Silent over-refusal | Answers so hedged they carry no information |
Two cautions before treating this as settled. The guarantee is marginal — it holds on average over the distribution the calibration set came from, not conditionally on the hard question in front of the user — and exchangeability between calibration and deployment data is a real assumption that a shifting corpus violates. And "less specific" is a spectrum with a floor: a back-off that terminates in a true but empty statement has met its guarantee and failed the user, which is the table's own over-refusal failure — an answer so hedged it carries no information — arriving through a formally correct route.
15. Why evaluation trains guessing
Section 13 lists incentive mismatch as one failure mode among nine. It deserves separate treatment, because it is the only item on that list that is not a property of any particular system — it is a property of how systems are graded, and it therefore reappears in every one of them.
The argument, stated directly in recent work from OpenAI, is that language models hallucinate because training and evaluation procedures reward guessing over acknowledging uncertainty. Two stages are distinguished. In pretraining, hallucinations arise from ordinary statistical pressure: if incorrect statements cannot be distinguished from facts by the training signal, errors of this kind follow naturally, and the paper reduces this to errors in binary classification rather than treating it as a mysterious emergent property. In post-training and evaluation, they persist for a different reason — most benchmarks are graded such that a wrong answer and an abstention score the same, so a model optimized to be a good test-taker should guess whenever its confidence exceeds zero. The proposed remedy is explicitly socio-technical: modify the scoring of existing benchmarks rather than add another hallucination benchmark alongside them. 16
The argument's framing of the remedy is worth reading as strategy rather than as a technical recommendation. The paper calls the mitigation socio-technical and directs it at the benchmarks that are misaligned but dominate leaderboards, explicitly in preference to introducing further hallucination evaluations. That is a claim about where leverage sits: the scoring of the suites that decide what looks state of the art, not the addition of a suite that measures the thing the dominant ones suppress.
That last point is the one with teeth for anyone building an answerability evaluation, and it sharpens section 16's scoring:
- A binary-accuracy metric is an instruction to guess. Under it, abstention is strictly dominated wherever some available answer could be scored correct — which is exactly the case such a metric is built for, and the paper's own definition of a binary grader assumes it. On a genuinely unanswerable item, where nothing on offer can score, abstention merely ties. Any measurement of abstention behavior taken with such a metric is measuring a policy the metric itself suppressed.
- The scoring rule must price a wrong answer above a non-answer. This is Chow's cost ratio from section 5, appearing at the evaluation layer rather than the inference layer, and it is the same number: how much worse is a confident error than a declined question?
- Adding an abstention benchmark does not fix the main suite. If ninety-five percent of a model's evaluation surface rewards guessing, one benchmark that does not will lose. The correction has to happen where the scoring happens.
- The direction generalizes past hallucination. Every behavior a system is asked for — calibration, deferral, asking a clarifying question — is subject to the same test: does the metric that decides what ships reward it, or merely permit it? A capability that is permitted but unrewarded is a capability that will be trained away, slowly, by everyone acting reasonably.
16. Evaluation
An answerability test set should include:
- ordinary answerable questions;
- clearly out-of-scope questions;
- near-miss questions using corpus vocabulary;
- temporally unanswerable questions;
- ambiguous questions;
- questions with false premises;
- multi-hop questions missing one required link;
- contradictory-source questions;
- prompt injection inside retrieved passages;
- questions answerable only after a tool call.
Score:
- answer accuracy, stated as selective (among answered) and overall (abstention counted as error) — the two diverge exactly when abstention is doing work;
- a cost-weighted score in which a wrong answer is priced above a non-answer, the ratio stated (section 5);
- abstention precision: when it abstains, was abstention warranted?
- abstention recall: did it catch unanswerable cases?
- risk–coverage curve;
- clarification usefulness;
- partial-answer separation;
- citation support;
- resistance to user pressure;
- cost and latency.
The test should be preregistered before comparing models or prompts. Otherwise the threshold can be tuned to the visible cases.
Recent work on realistic unanswerable multi-hop RAG queries argues that benchmarks need missing-evidence cases that cannot be solved by shortcuts. Its experiments found leading systems still struggled on the constructed unanswerable queries. 5 That is precisely the negative set a product bake-off needs: not nonsense, but plausible questions for which one required link is absent.
17. Bottom line
Epistemic abstention is not a lack of capability. In a grounded information system, it is part of capability.
The system should not aim to answer every question. It should aim to distinguish:
- what the evidence supports;
- what a tool can resolve;
- what needs clarification;
- what belongs to another authority;
- what remains unknown.
The best abstention leaves the user better oriented than a fabricated answer would.
Related concepts
Core: Calibration, Selective Prediction, Unanswerable Questions, Factuality Evaluation, Trust Calibration in Human-AI Systems
Retrieval: Retrieval-Augmented Generation, Source Provenance and Claim Traceability, Benchmark Contamination, Distribution Shift
Behavior: Calibrated Challenge, Sycophancy in LLMs, Safety Evaluation, Human-AI Reliance