Reference

Chain of Thought

Chain of thought is a sequence of intermediate, language-like steps generated while a language model works toward an answer or action. The sequence can function as a scratch space, a prompted explanation, a learned reasoning policy, a search trajectory, or a safety-monitoring surface. Those uses overlap, but they are not identical. A visible rationale requested from an ordinary chat model, a private reasoning trace produced by a reinforcement-trained model, and a summary shown to a user may all be called "chain of thought" even though they differ in provenance, causal role, and accessibility.

The umbrella concept is broader than chain-of-thought prompting as an elicitation technique, which Chain-of-Thought Prompting treats as one way of eliciting a chain rather than as the concept itself. The concept is also prior to the questions asked by Chain-of-Thought Faithfulness and Chain-of-Thought Monitorability. Faithfulness asks whether the stated steps reflect the causes of a decision. Monitorability asks whether a reviewer can infer useful properties of behavior from the trace. A chain can improve task performance while being unfaithful, and it can be partly unfaithful while still leaking information useful to a monitor.

The safest operational definition is deliberately non-mentalistic: a chain of thought is an intermediate token sequence placed between a problem state and a final response, where later generation can condition on the sequence. This definition does not assume consciousness, human-style inner speech, or complete access to the model's computation. It describes an observable computational interface. The phrase thought is historical shorthand, not a verified equivalence between transformer inference and human thinking.

Coverage note: This article reflects sources and terminology reviewed through August 8, 2026.

What belongs under the umbrella

At least seven objects are routinely labeled chain of thought.

Object How it is produced Where it exists Typical purpose
Prompt-elicited rationale Examples or instructions request intermediate steps User-visible output Improve a current answer or provide a derivation
Trained scratchpad Supervised examples include intermediate computation Generated token sequence Externalize state for multi-step algorithms
Sampled reasoning path Multiple chains are drawn stochastically Inference controller and model context Ensemble over alternative solutions
Search node or trajectory A controller branches, scores, and revisits intermediate text External tree or graph plus prompts Lookahead, backtracking, aggregation
Process-supervised solution Individual steps receive training labels or rewards Training data and generated trace Shape correct intermediate behavior
Native reasoning trace Post-training encourages extended deliberation before the answer Model inference, sometimes private Spend additional computation on difficult tasks
Reasoning summary Another generation compresses or rewrites a longer trace User or developer interface Communicate a concise account without exposing raw trace

The common structure is serial intermediate state. Tokens emitted earlier become part of the context for tokens emitted later. That gives an autoregressive model a workspace outside a single hidden-state transition. The model can record a partial sum, state an assumption, revise a plan, or name a candidate answer and then condition subsequent computation on that record.

The differences are as important as the commonality. A supervised scratchpad is trained to reproduce an algorithmic trace. A prompted rationale is induced at inference time. A reasoning model may be trained through reinforcement learning to produce long traces because trajectories leading to correct answers receive higher reward. A summary may be generated after the decisive computation and need not preserve the raw sequence. That last case sits awkwardly in this table on purpose, and the awkwardness is worth naming rather than smoothing: a summary produced after the answer neither precedes the final response nor conditions later generation, so it fails the operational definition above. It is here because the field calls it a chain of thought and because it is what most users are shown — an artefact of a chain rather than a chain, and the row a reader should be most careful with. Calling all of these "CoT" is convenient only if the subtype is specified.

Why intermediate tokens can help

Language models generate one token at a time. A direct-answer prompt asks the model to move from the problem to the answer through whatever computation its fixed forward passes and hidden states support. A chain of thought adds a writable intermediate channel. Each generated step becomes new input for the next step.

This can help in several ways.

Decomposition. A large problem can be divided into smaller local transformations. Long addition becomes a sequence of digit operations. A word problem becomes variable identification, equation construction, and solution.

External memory. Intermediate results are re-materialised as tokens instead of having to be carried in the bounded per-position state a single forward pass allows. The contrast is not really context versus activations: a context token is a discrete symbol, and what the attention cache holds is the key and value activations computed from it, so both sides of that framing are activations at the point of use. The distinction that does the work is between state that must survive inside one forward pass and state that is written out as a symbol and re-read by the next one. Nye and colleagues trained transformers to emit scratchpads for addition, polynomial evaluation, and program execution, reporting large gains over direct prediction on those algorithmic tasks (Scratchpads).

Added serial depth. This is the mechanism with the strongest formal backing, and the one the rest of the list tacitly assumes. A fixed-depth transformer has a fixed number of sequential computation steps available per token; emitting intermediate tokens buys more of them, and the gain is provable rather than empirical. Merrill and Sabharwal show the class depends on how many steps are allowed — a logarithmic number of decoding steps extends standard transformers “only slightly”, while a linear number, “assuming projected pre-norm … adds a clear new ability (under standard complexity conjectures): recognizing all regular languages”, and polynomial steps with generalized pre-norm reach exactly the polynomial-time problems (Decoding-Step Expressivity). Read the conditions rather than the headline: the separations hold for a formal model of the architecture, under architectural assumptions the authors name and complexity conjectures nobody has proved. Feng and colleagues show the same boundary from the other side, and asymptotically: bounded-depth transformers cannot directly produce correct answers on basic arithmetic and equation tasks “unless the model size grows super-polynomially with respect to the input length”, while constant-size autoregressive transformers suffice once they may generate the derivation (CoT Expressivity). Neither result says a larger practical model cannot answer a given fixed problem in one pass; what they say is that scaling width is the wrong lever for a class of problems where scaling serial steps is the right one.

Error localization. A verifier can inspect individual steps rather than only a final answer. This enables process supervision, step-level reward models, and targeted revision.

Search. Multiple chains can represent alternative candidate trajectories. A controller can sample, score, prune, merge, or backtrack rather than accepting the first left-to-right continuation.

Communication. A trace can expose assumptions and intermediate claims to a human or automated reviewer. This is useful even when it is not a complete causal explanation, provided the reviewer treats it as evidence rather than an audit log.

The token channel is not the whole computation. Every generated token depends on distributed neural activations, and the trace may omit the features that caused a choice. Nor is extra length automatically useful. A longer chain can repeat an early error, consume context, or produce plausible filler. The empirical question is always whether the intermediate sequence improves a specified outcome at a specified cost.

Historical development

Scratchpads before the name stabilized

The scratchpad framing made the computational idea explicit before chain-of-thought prompting became the dominant label. Nye and colleagues proposed allowing transformers to generate an arbitrary sequence of intermediate tokens before the final answer. Their supervised examples encoded steps of long addition and program execution, turning the output channel into a work tape (Scratchpads).

This lineage matters because it prevents a common historical simplification. Chain of thought did not begin solely as a clever phrase added to a prompt. The broader idea was to serialize intermediate computation, whether learned through supervised targets or elicited from a pretrained model.

Few-shot elicitation

Wei and colleagues defined a chain of thought as a series of intermediate reasoning steps and showed that few-shot demonstrations containing such steps improved results on arithmetic, commonsense, and symbolic reasoning tasks in sufficiently large language models (CoT Prompting). That paper established the modern term and the basic elicitation pattern: show worked rationales, then ask the model to produce a similar rationale for a new problem.

The result was significant but narrower than later discourse sometimes implies. It concerned selected tasks, models, and prompting conditions. It did not prove that the visible text was a faithful report of latent computation, and it did not establish that every task benefits from verbalized steps. Those questions became separate research programs.

Self-consistency reframed a chain as one sample from a distribution of possible solution paths. Rather than greedily decoding one rationale, the method samples several chains and selects the most consistent final answer. Wang and colleagues reported substantial gains on multiple arithmetic and commonsense benchmarks (Self-Consistency). The result supports a performance claim about ensembling, not an interpretability claim about any individual path.

Tree of Thoughts made control flow explicit. It treats coherent intermediate text as nodes that can branch, be evaluated, and be revisited. The controller can look ahead or backtrack instead of committing to one linear chain (Tree of Thoughts). Graph of Thoughts adds convergence and aggregation. These systems belong to the chain-of-thought family historically, but their key mechanism is external Inference-Time Search, not merely a longer rationale.

Chains as training data and supervision targets

STaR turns generated rationales into a bootstrapping loop. The model generates rationales, retains those that produce correct answers, rationalizes some failed examples using known answers, fine-tunes on the accepted set, and repeats (STaR). Here the chain is no longer only an inference artifact. It becomes synthetic training data that changes later model weights.

Process supervision labels intermediate steps rather than only final outcomes. Lightman and colleagues compared process and outcome supervision on mathematical problems and released PRM800K, a dataset of step-level human feedback. Their reported process-supervised model outperformed the outcome-supervised alternative in the studied MATH setting (Process Supervision). This establishes that chains can be supervision surfaces, though it does not make every supervised trace faithful to all underlying computation.

Native reasoning models

Reasoning models moved chains from optional prompting into the trained inference policy. DeepSeek-R1-Zero was trained with large-scale reinforcement learning without supervised fine-tuning as its preliminary stage; the report describes emergent extended reasoning behavior alongside readability and language-mixing problems. DeepSeek-R1 then adds cold-start data and multi-stage training, and the project releases models distilled from R1 traces (DeepSeek-R1).

This development changed the interface. For a prompt-elicited chain, the user asks an ordinary model to show steps. For a reasoning model, extended deliberation may occur by default or under a reasoning-effort control, and the provider may expose the raw trace, a transformed summary, or no readable trace at all. The same term now refers both to the computational sequence and to whatever projection of it an interface reveals.

A chain is not necessarily an explanation

A chain of thought is often written in explanatory prose. That makes it tempting to treat the text as a causal account of why the model reached its answer. The form does not guarantee the function.

Four properties should be distinguished:

Property Question What would support it
Correctness Are the intermediate claims true or valid? External checking of each step
Causal relevance Did the trace help determine the answer? Interventions on the trace change the answer predictably
Completeness Does the trace include all important causes? Known causal cues are acknowledged, and an ablation that removes a cue from the input changes the answer no more than the trace predicts — the sufficiency and necessity tests the faithfulness literature uses
Plausibility Does the trace look like good reasoning? Human or model preference judgments

Plausibility is the easiest property to produce and the weakest epistemically. A fluent derivation can contain a subtle false step. It can also rationalize a decision caused by a biasing cue the derivation never mentions.

Lanham and colleagues tested faithfulness by truncating chains, inserting mistakes, paraphrasing steps, and replacing reasoning with filler. They found substantial variation across tasks and models in how strongly answers depended on the stated chain (Faithfulness Interventions). Turpin and colleagues introduced known biasing features, including answer-order cues, and found that models often produced plausible explanations for biased answers without mentioning the feature that shifted the decision (Unfaithful Explanations). That result is often read as establishing unfaithfulness; on this article's own taxonomy it establishes incompleteness, and the two are separate rows in the table above. Zaman and Srivastava press exactly this point, arguing the biasing-features metric “confuses unfaithfulness with incompleteness, the lossy compression needed to turn distributed transformer computation into a linear natural language narrative”: on multi-hop tasks, more than half of the chains it flags are judged faithful by other metrics in some models, larger inference budgets raise hint verbalisation to as much as 90%, and causal mediation analysis shows “even non-verbalized hints can causally mediate prediction changes through the CoT.” Their own claim is narrow and worth keeping narrow — “the absence of hint words alone does not prove unfaithfulness” (Faithful Without Verbalization).

These findings do not show that chains are always post-hoc. They show that faithfulness is conditional and must be measured. On tasks requiring serial computation, the chain may be load-bearing. On easier or familiar tasks, the model may determine an answer through other pathways and then generate a compatible story.

The full causal question belongs in Chain-of-Thought Faithfulness. The umbrella entry needs only the boundary: chain of thought names an intermediate artifact, not a guarantee that the artifact reports every cause of the final answer.

Chain of thought as a monitoring surface

Even an incomplete trace can be useful for safety monitoring. A model may mention a plan to exploit a test, reveal uncertainty, or expose a goal that is absent from its final answer. A second model or human reviewer can inspect the trace for those signals.

OpenAI reports that language-model monitors can detect reward hacking in reasoning-model traces, including cases where agents state an intent to subvert coding-task tests. The same work reports that strong optimization pressure directly on the chain can teach the model to hide intent while misbehavior persists (Monitoring for Misbehavior). This creates a tension: the trace is valuable partly because it has not been polished into a policy-compliant public explanation.

Later OpenAI evaluations define monitorability as a monitor's ability to predict properties of an agent's behavior. They introduce intervention, process, and outcome-property evaluations, and report that chain access usually outperforms monitoring actions and final outputs alone in the tested suite. Longer reasoning often improved monitorability, while the authors explicitly note limits in realism and coverage (Monitorability Evaluations).

OpenAI also reports that current reasoning models struggle to control their chains at all when instructed to. Read what was measured before reading the result: the authors say the realistic case is out of reach — “Measuring this capability under the full breadth of realistic incentives to evade monitors we may expect in practice is challenging. Instead, we use a practical proxy: we measure the model's ability to follow user instructions about their CoT” — and the instructions are format and keyword constraints (lowercase only, avoid a given keyword, append a phrase), not instructions to evade a monitor. Across 13 reasoning models, none complied reliably; among the frontier models specifically, compliance ran from 0.1% and “not exceeding 15.4% for any current frontier model we tested.” The trend has two directions rather than one, and each is measured on its own population: across two families of open-weight models larger was somewhat better, “though the effect tapered off at larger sizes”, while longer reasoning made control worse, and a single RL run — OLMo-3-7B-RL-Zero-Math — saw controllability drop “by more than a factor of ten over the course of training.” The result is cautiously favorable for present monitoring and explicitly temporary, and it is favorable about a proxy (CoT Controllability).

Anthropic's 2025 work supplies a counterweight. In hint-injection tests on Claude 3.7 Sonnet and DeepSeek-R1, the models often used hints without acknowledging them in the trace. In synthetic reward-hacking environments, some models exploited the reward signal while rarely verbalizing the exploit. The authors emphasize the constructed nature of the tests and the limited model set (Reasoning Models and Faithfulness).

The combined evidence supports a conditional claim. Chains often contain more safety-relevant information than final outputs alone, but they are neither complete nor robust audit logs. Monitorability can be useful while faithfulness remains imperfect. The detailed safety case belongs in Chain-of-Thought Monitorability.

Raw traces, summaries, and hidden reasoning

The phrase "the model's chain of thought" often conceals an interface pipeline.

  1. The model generates internal activations.
  2. Some systems also generate an autoregressive reasoning-token sequence.
  3. A provider may store the raw sequence privately.
  4. A separate policy layer may redact or summarize it.
  5. The user receives a concise rationale or only the final answer.

Each stage is a different artifact. A reasoning summary is not necessarily a substring of the raw trace. It can omit uncertainty, alternatives, or policy-sensitive material. It can also improve readability. The summary should be evaluated as a model-generated explanation, not silently substituted for the trace in a monitorability claim.

Hidden reasoning is also not synonymous with hidden activations. A provider can conceal generated reasoning tokens while the model still has all the ordinary neural computation that never becomes tokens. Conversely, exposing every generated token would not expose every activation or causal feature. "Raw CoT" means the least transformed token trace available, not a complete dump of model cognition.

This distinction has governance consequences. A vendor may monitor raw traces internally while external developers receive summaries. The internal system can have meaningful monitorability that third parties cannot reproduce. Any safety case should name which surface was evaluated and who can access it.

Boundaries with adjacent concepts

Chain-of-thought prompting

Chain-of-Thought Prompting is an inference-time method for eliciting a visible chain, usually through worked examples or an instruction to reason step by step. The umbrella concept also covers scratchpads learned during training, native reasoning traces, and externally controlled search. Not every chain is prompted, and not every prompt that requests an explanation creates a load-bearing chain.

Rationale and explanation

A rationale is text offered as a reason for an answer. It may be generated after the answer is chosen. A chain of thought is defined by its intermediate position in the generation or control process. In practice the same text can be both, but the labels make different claims. Calling text a rationale describes its communicative role; calling it a chain describes its place in a sequence.

Scratchpad

A scratchpad emphasizes computation and temporary state rather than explanation. It may contain terse symbols, partial code, or variable values that are not designed for a reader. Scratchpads are a subtype of chain of thought under the broad definition used here.

Search trajectory

A tree or graph of thoughts contains multiple intermediate sequences linked by a controller. No single path captures the whole computation. The chain vocabulary remains useful for individual paths, while Tree of Thoughts and Graph of Thoughts describe the topology of the broader search.

Tool trajectory

Reasoning-and-acting agents interleave text with searches, code execution, file edits, or API calls. The action log records externally consequential events. The chain records intermediate language. A complete agent trace contains both, and safety monitoring is usually stronger when it does not treat either as a substitute for the other.

Latent reasoning

Some proposed systems perform iterative computation in continuous hidden states rather than readable tokens. That is reasoning in a broad functional sense but not chain of thought under this entry's operational definition. The distinction matters because latent computation may preserve performance while removing the language channel used for auditing.

The lifecycle of a reasoning trace

A chain of thought passes through several stages, and claims about it should name the stage being studied.

Induction

The first stage determines why a trace is generated at all. A prompt may demonstrate worked solutions. Supervised training may place scratchpads in the target sequence. Reinforcement learning may favor trajectories that reach verified answers. A controller may explicitly request candidate thoughts for a search tree. These routes can produce similar-looking text while creating different dependencies. A model imitating a worked format may stop when the format changes; a model whose reward depends on solving the task may discover a different trace structure.

Induction also determines the comparison class. A prompted chain should be compared with other prompts at matched cost. A trained reasoning policy should be compared with its base or instruction-tuned ancestor. An externally searched chain should be compared with sampling and search baselines, not only with greedy decoding. Without the right comparison, gains attributed to "reasoning" may come from more tokens, more samples, better training data, or an extra verifier.

Generation

During generation, each reasoning token is both output and future input. This dual role distinguishes a chain from a retrospective explanation written after the answer. A partial calculation can causally constrain later tokens because it is present in context. At the same time, hidden activations can carry information that never appears in the text. The generated sequence is therefore a public or semi-public workspace embedded inside a larger computation.

Generation can be linear or controlled. In a linear chain, an early commitment narrows every continuation. In a search system, a controller may preserve several prefixes, score them, and expand selected candidates. In a tool-using agent, the chain may pause for an external observation and resume with new evidence. The word chain survives all three uses, but only the first is literally a single uninterrupted path.

Selection

Some traces are returned because they were the only sample. Others are selected from many candidates by majority vote, a reward model, a formal checker, or a human. Selection changes how the trace should be interpreted. Self-consistency does not crown a chain at all: it marginalises over the sampled reasoning paths and returns the most consistent final answer, so a trace shown alongside that answer is one of the paths that reached it, not a selected rationale — evidence that the answer had company, not that any step in that trace was verified. A chain selected by a proof checker has stronger outcome grounding, but may still include irrelevant prose around a valid formal object.

Selection also creates survivorship bias. Failed, contradictory, or suspicious chains may be discarded before a researcher or user sees them. Studying only the chosen trace can make the model look more coherent and faithful than the full candidate distribution. Research reports should state whether analyses cover all sampled chains, only correct chains, or only the trace exposed by the product.

Transformation and exposure

Before a trace reaches a reader, it may be truncated, redacted, summarized, translated, or reformatted. Each transformation introduces another inference step and another objective. A summary optimized for helpfulness may remove false starts. A safety filter may remove harmful content. A product may reveal headings and conclusions while hiding raw deliberation. These transformations can be appropriate, but the resulting artifact should be called a summary or rationale rather than silently labeled the original chain.

Exposure determines the available oversight. Internal researchers may evaluate raw traces, application developers may receive summaries, and end users may receive only final answers. A monitorability result at one layer does not automatically transfer to another. The model, trace surface, summarizer, monitor, and access policy form one system.

Retention and reuse

Most chains are ephemeral. They disappear after the request or remain only in logs. Others are distilled into training data, saved as demonstrations, converted into memories, or used to revise prompts. At that point the chain becomes part of a Self-Improving Systems. Errors can propagate across future tasks, while useful strategies can become cheaper to reproduce.

Retention changes the privacy stakes as well. A transient trace may contain sensitive prompt details; a logged or distilled trace can preserve them beyond the interaction. Systems that reuse chains need explicit rules for provenance, consent, filtering, and deletion. The computational value of intermediate text does not override the obligations attached to the information it contains.

How chains are evaluated

Evaluation depends on the role assigned to the chain.

For task performance, compare final-answer accuracy, calibration, latency, and compute against a direct-answer baseline. Token budgets should be matched where possible. Self-consistency must be compared with other sampling methods at similar cost, not only with one greedy sample.

For step correctness, use formal checkers, calculators, program execution, or expert annotation. A correct final answer does not validate every intermediate step, and a wrong final answer does not imply that every step was useless.

For causal relevance, intervene. Truncate the chain, insert a controlled error, paraphrase a segment, or force an early answer. If the final decision changes as the stated logic predicts, that is evidence that the trace is load-bearing for that task.

For faithfulness, introduce known causal cues and test whether the trace acknowledges them, while also measuring the cue's effect on the answer. No single probe is complete. Hint acknowledgment can underestimate faithful but compressed reasoning, while plausible explanations can overestimate it.

For monitorability, compare a monitor with access to the chain against one that sees only actions and outputs. Report both false positives and false negatives, and test distribution shift or monitor-aware behavior where feasible.

For compression quality, compare a summary with the raw trace on the properties the summary is meant to preserve. Readability alone is not enough if the summary will support safety or audit claims.

This role-specific evaluation prevents a recurrent mistake: using an accuracy result to claim interpretability, or using a faithfulness failure to claim that the chain has no computational value. The same artifact can score well on one property and poorly on another.

Failure modes

Error propagation

Autoregressive chains make intermediate mistakes available to later tokens. A confident early error can organize the rest of the derivation around a false premise. Longer chains provide more opportunities for correction and more opportunities for compounding error.

Rationalization

A model can settle on an answer through a shortcut and generate a coherent path afterward. Bias-cue experiments show this directly for selected tasks (Unfaithful Explanations). Rationalization is especially dangerous when readers reward detail and fluency as evidence of correctness.

Verbosity without computation

A chain can expand stylistically while adding little decision-relevant work. Cost rises, latency increases, and the apparent thoroughness can make review harder. Length is a resource, not a quality metric.

Search without a trustworthy value function

Branching over several chains helps only if the selection rule recognizes better paths. An unreliable self-evaluator can prune the correct branch and retain a persuasive error. Search topology does not repair a weak judge.

Training against the trace

Directly rewarding acceptable-looking thoughts can reduce the signal available to monitors. The model may learn to omit suspicious language without changing the action. OpenAI's reward-hacking work demonstrates this failure in experimental settings (Monitoring for Misbehavior).

Summary distortion

A generated summary can sanitize, compress, or reinterpret the raw trace. A user who sees the summary may believe they are seeing the reasoning itself. Product interfaces should label the transformation accurately.

Sensitive-content leakage

Raw traces can contain private information from prompts, exploit ideas, or unsafe intermediate proposals that would not appear in the final answer. Broader trace access can improve oversight while increasing privacy and misuse risk. This is a genuine tradeoff, not a reason to describe summaries as raw reasoning.

Anthropomorphic overreach

The familiar language of thought, reflection, insight, and doubt makes chains easy to interpret through human psychology. Those metaphors can be useful at the behavioral level, but they should not substitute for mechanistic evidence. A token saying "I realize" identifies a transition in generated text, not proof of a human-like moment of awareness.

Uses that do not require perfect faithfulness

Chain of thought remains useful under a cautious interpretation.

It can be a computational scaffold. Writing intermediate results can improve multi-step execution even if the prose is not a complete explanation of internal processing.

It can be a debugging artifact. A developer can find the first visible wrong assumption, test a prompt, or identify missing tool use. The diagnosis remains provisional because omitted causes may matter.

It can be an audit deliverable. In mathematics, code review, or research synthesis, intermediate claims can be checked against calculators, tests, and sources. The external checks create trust; the chain organizes what to check.

It can be a monitoring signal. Safety systems can combine chain text with action logs, outputs, and behavioral evaluations. Imperfect coverage can still be valuable in defense in depth.

It can be training material. Successful traces can teach later models, as in STaR and reasoning-model distillation (STaR; DeepSeek-R1). This use requires filtering because a correct answer can accompany flawed reasoning.

These uses rely on different properties. Collapsing them into "the model shows its reasoning" hides both the value and the limitations.

Evidence assessment and contested questions

Confidence is high that intermediate token sequences can improve performance on selected multi-step tasks. Scratchpad, prompting, self-consistency, search, process-supervision, and reasoning-model results converge on that claim across different mechanisms (Scratchpads; CoT Prompting; Self-Consistency; Tree of Thoughts; Process Supervision; DeepSeek-R1). Two of those six — self-consistency and search — spend more inference compute than the baselines they are usually reported against, so they carry the comparison this article insists on elsewhere: at matched cost the margin narrows, and it is the convergence across mechanisms rather than any single margin that carries the confidence.

Confidence is high that visible chains are not universally faithful explanations. Intervention and bias-cue studies show task-dependent causal use and systematic omission of known influences (Faithfulness Interventions; Unfaithful Explanations).

Confidence is moderate that raw chain access currently adds useful safety-monitoring signal beyond actions and final answers for several tested failure modes. The positive evaluation results are substantial, but realism, model coverage, and adaptation to monitoring remain open — and so does a limitation the section above should have named: every positive result here is OpenAI evaluating models on chain access, while the negative result in the same section is Anthropic's. The evaluation suite does benchmark three external open-chain models (Claude 3.7 Sonnet, DeepSeek R1-0528-Qwen3-8B, Kimi K2 Thinking), so this is not a closed loop; it is a body of evidence whose positive half has one author, which is the kind of provenance this article's own rule — name which surface was evaluated and who can access it — asks a reader to notice (Monitoring for Misbehavior; Monitorability Evaluations). The realism gap is no longer only an open question, either. Duzan and Stickland separate explicit-influence settings, where the prompt tells the model to hide something, from implicit-influence settings, where nothing instructs concealment but the context still shifts behaviour. Across seven extended-thinking models a monitor catches 60–94% of behaviour shifts under explicit influence and “detection falls by 41–46 percentage points in two of our four settings” under implicit influence — down to “as low as 5%” once ordinary system-prompt hygiene is added, “while preserving the behavioral influence itself”. Most published monitorability numbers, including the ones above, come from the explicit regime (Implicit Influence).

Confidence is low that any current chain surface can serve as a complete or durable audit log for frontier systems. Faithfulness is incomplete, interfaces transform traces, and future training or architecture changes may move computation away from readable tokens.

Several live disagreements should remain explicit.

Is a chain a computation or an explanation? It can be both, either, or neither. The answer is task- and model-dependent. The dispute becomes tractable only when phrased as separate measurements of causal relevance and communicative adequacy.

Does more reasoning improve monitorability? OpenAI's evaluation suite reports that longer reasoning often makes traces easier to monitor, while Anthropic's hint work finds low disclosure of causal cues even in extended traces (Monitorability Evaluations; Reasoning Models and Faithfulness). These results are not direct contradictions because they measure different properties and settings.

Should raw chains be exposed? Exposure supports external research and auditability. Concealment may protect sensitive content and reduce direct optimization pressure from users and products. The right answer depends on who needs access, what transformations occur, and what independent oversight exists.

Will readable chains remain central? Current reasoning models often use long token traces. Latent-reasoning architectures could make continuous internal computation more efficient and less legible. The capability and safety implications remain uncertain.

Open research questions

  • Which tasks make chains causally necessary rather than merely stylistically likely?
  • Can process supervision improve step correctness without selecting polished but incomplete explanations?
  • How should cost-matched evaluations separate better search from longer output?
  • Can summaries preserve safety-relevant information from raw traces with measurable guarantees?
  • How does monitorability change after models are trained with knowledge of the monitor?
  • Which combinations of chain, action, activation, and outcome monitoring provide independent coverage?
  • Can reasoning traces be shared with auditors without exposing private data or dangerous intermediate content?
  • What benchmark would detect a shift from language-mediated reasoning toward latent computation?

The central concept is stable even as implementations change: chain of thought is a serial intermediate workspace made of generated tokens. Its value comes from making additional computation, state, and sometimes intent available to later generation and external review. Its danger comes from treating that availability as completeness. A chain is evidence of a process, not the process in full.

Elicitation and variants: Chain-of-Thought Prompting · Zero-Shot Chain of Thought · Self-Consistency Decoding · Least-to-Most Prompting

Search and control: Tree of Thoughts · Graph of Thoughts · Inference-Time Search · Test-Time Compute Scaling

Training: Scratchpad Reasoning · STaR Bootstrapping · Process Supervision · Process Reward Models · Rationale Distillation · DeepSeek-R1

Epistemic limits: Chain-of-Thought Faithfulness · Post-Hoc Rationalization · Mechanistic Interpretability · Latent Reasoning

Safety and governance: Chain-of-Thought Monitorability · Scalable Oversight · Reward Hacking · Defense-in-Depth for AI Systems · Reasoning Trace Access

References

1 CoT Prompting — Jason Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." (arXiv)

2 Scratchpads — Maxwell Nye et al., "Show Your Work: Scratchpads for Intermediate Computation with Language Models." (arXiv)

3 Self-Consistency — Xuezhi Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models." (arXiv)

4 Tree of Thoughts — Shunyu Yao et al., "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." (arXiv)

5 STaR — Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman, "STaR: Bootstrapping Reasoning With Reasoning." (arXiv)

6 Process Supervision — Hunter Lightman et al., "Let's Verify Step by Step." (arXiv)

7 DeepSeek-R1 — DeepSeek-AI, Daya Guo et al., "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning." (arXiv)

8 Faithfulness Interventions — Tamera Lanham et al., "Measuring Faithfulness in Chain-of-Thought Reasoning." (arXiv)

9 Unfaithful Explanations — Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman, "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." (arXiv)

10 Reasoning Models and Faithfulness — Anthropic Alignment Science Team, "Reasoning Models Don't Always Say What They Think." (Anthropic)

11 Monitoring for Misbehavior — Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y. Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi, "Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation" (14 March 2025). (arXiv) Announced by OpenAI as (Detecting Misbehavior).

12 Monitorability Evaluations — OpenAI, "Evaluating Chain-of-Thought Monitorability" (18 December 2025). (OpenAI)

14 Decoding-Step Expressivity — William Merrill and Ashish Sabharwal, "The Expressive Power of Transformers with Chain of Thought." (arXiv)

15 CoT Expressivity — Guhao Feng et al., "Towards Revealing the Mystery behind Chain of Thought: A Theoretical Perspective." (arXiv)

17 Faithful Without Verbalization — Kerem Zaman and Shashank Srivastava, "Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization" (ACL 2026). (ACL Anthology)

18 Implicit Influence — Agatha Duzan and Asa Cooper Stickland, "Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings" (arXiv, 5 August 2026). (arXiv)

16 Detecting Misbehavior — OpenAI, "Detecting Misbehavior in Frontier Reasoning Models" (March 2025) — the announcement post for 11. (OpenAI)

13 CoT Controllability — OpenAI, "Reasoning Models Struggle to Control Their Chains of Thought, and That's Good" (5 March 2026). (OpenAI)