Abstract. Case-based reasoning is an older AI approach with renewed relevance for language-model agents. Instead of solving every problem from general rules alone, a system retrieves similar past cases, adapts their solutions, checks the result, and stores the new experience. Modern agent memory often recreates pieces of this loop with embeddings, trajectories, reflections, and skill stores. The central lesson from case-based reasoning is that retrieval is only the first step: safe reuse requires adaptation, verification, and disciplined retention.
Coverage note: sources checked through August 2026.
1. The basic idea
People often solve a new problem by remembering an earlier one: a similar machine failure, legal dispute, software bug, negotiation, or medical presentation. The old case does not provide a guaranteed answer. It provides a structured starting point.
Case-based reasoning (CBR) turned this intuition into a computational method. The widely used cycle has four steps: 1
- Retrieve the most relevant previous case or cases.
- Reuse their information or solution in the new situation.
- Revise the proposed solution after testing or feedback.
- Retain the new case and what was learned from it.
The cycle matters more than the slogan “find something similar.” A nearest-neighbour search that pastes an old answer is incomplete CBR because it omits adaptation, verification, and learning.
The lineage is worth knowing because it explains the method's emphases. CBR did not begin as an information-retrieval technique; it grew out of Schank's dynamic-memory account of how human experts reason from remembered episodes, and Kolodner's early systems and survey established the field's working vocabulary — cases as contextualized experiences that teach a lesson, indexes as the modeling layer that decides which lessons a new situation should recall. 6 The four-step cycle above is Aamodt and Plaza's later consolidation of that decade of systems into a common framework. 1
Nor is this only a laboratory idea. By the mid-1990s CBR had a commercial track record — help-desk support, fault diagnosis, and the reuse of engineering layouts — established enough that an introductory review of the field could illustrate it with deployed commercial systems and the engineering trade-offs they surfaced. 14 The method's practical appeal was always the same one that now motivates agent memory: experience is cheaper to reuse than expertise is to formalize. The honest qualification is that what shipped was overwhelmingly the retrieve step: the deployed systems found and showed a similar case to a human, and adaptation — the step this article treats as what makes CBR more than lookup — was largely what the field could not automate. The commercial record is evidence that remembered cases are worth retrieving, not that the full cycle has been run at scale.
2. What counts as a case
A case is not merely a document. It is a record of a problem-solving episode.
A useful agent case might contain:
- the original goal;
- relevant environment and constraints;
- observations available at decision time;
- actions taken;
- tool outputs;
- the result;
- how success or failure was judged;
- corrections and postmortem;
- provenance, time, and version information.
The case should preserve both problem description and solution outcome. If it stores only the final answer, the agent cannot tell whether the old case matches the new one. If it stores only the transcript, retrieval may drown in irrelevant detail.
Cases can exist at several levels:
| Case type | Example | Best use |
|---|---|---|
| Factual | A source resolved a disputed claim | Evidence retrieval |
| Diagnostic | A failure signature mapped to a root cause | Debugging |
| Procedural | A sequence of actions completed a task | Workflow reuse |
| Decision | Constraints led to one option over alternatives | Planning |
| Exception | A normal rule failed under special conditions | Risk control |
| Negative | An approach looked promising but failed | Avoiding repeated errors |
Negative cases are particularly valuable. A memory that stores only successes can teach an agent to repeat attractive dead ends.
3. Retrieval is a modeling decision
CBR systems need an index that says which aspects of a case matter. A surface match may be useless; a structurally similar case with different vocabulary may be exactly right.
For an AI agent, retrieval signals can include:
- semantic similarity of the task description;
- shared entities or tool types;
- comparable constraints;
- failure signatures;
- environment versions;
- risk and authority level;
- causal or relational structure;
- recency and outcome quality.
This differs from ordinary Retrieval-Augmented Generation. RAG often retrieves passages relevant to a question. Case retrieval looks for situations relevant to a decision. The returned object includes what was tried and what happened.
The tension between cheap surface matching and expensive structural matching is old and quantified. The MAC/FAC retrieval model formalized the standard resolution: a fast surface-similarity pass over the whole memory produces candidates, and a structural-mapping pass selects among them — because full structural comparison against every stored case does not scale, and surface similarity alone retrieves the wrong analogs. 7 Embedding-based agent memory is the modern instance of the first stage; the analogy-retrieval literature's warning is that systems which stop there inherit exactly the retrieval errors the second stage exists to catch.
More retrieval is not automatically better. Several near-duplicate cases can crowd out a rare but decisive exception. Diversity, explicit negatives, and a maximum context budget are part of the retriever.
4. Reuse requires adaptation
The source case and target case will differ. Reuse therefore needs an adaptation step.
Three common patterns are:
Substitution. Replace entities, parameters, paths, or dates while preserving the solution structure.
Transformation. Change the solution because the target has different constraints, tools, or resources.
Derivational reuse. Reconstruct the reasoning process that produced the old solution rather than copying the solution itself.
The last is especially important for agents. A command that worked in one repository may be destructive in another. Reusing the derivation—inspect status, resolve the exact target, dry-run, validate, then act—is safer than reusing the command string.
Derivational reuse is not a new invention for language agents. Veloso and Carbonell built it into the PRODIGY planner in the early 1990s as derivational analogy: the system stored the justification structure of past problem-solving episodes — the decisions, the reasons, the rejected alternatives — and replayed the reasoning against the new situation, re-validating each step rather than transplanting the solution. Case acquisition, storage, and reuse were fully automated, and replay degraded gracefully when the new problem diverged from the old one. 8 That design anticipates the correct default for agent memory: store why, not just what.
An adaptation record should state:
- what is believed to match;
- what is known to differ;
- what parts are being reused;
- what must be revalidated;
- what would falsify the analogy.
Without this record, case reuse becomes an opaque prompt injection from the past.
5. Revision is the safety step
In CBR, revision tests and repairs the proposed solution. For an agent, this can mean:
- run unit tests;
- query the current schema;
- compare a generated answer with source passages;
- simulate a plan;
- ask a human to approve a high-impact action;
- check a negative sentinel;
- observe the environment’s response and backtrack.
Revision should match the risk. A low-cost recommendation may need only a plausibility check. A database migration or medical decision needs stronger evidence and authority.
This is where modern agents often fall short. They retrieve a past success, generate a confident adaptation, and treat fluency as validation. CBR’s older vocabulary makes the missing step obvious: reuse is provisional until revised.
The classic demonstration is Hammond's CHEF and the case-based planning framework built from it: adapted plans fail in ways the retrieval step cannot predict, so the planner runs the plan, explains the failure causally, repairs the plan, and — the durable insight — stores the failure itself, indexed so that future planning anticipates it before making the same mistake. 9 Failure-driven repair and failure-indexed memory are the parts of the cycle that modern retrieve-and-generate agents most often drop, and they are the parts that made the old systems safe to rerun.
6. Retention is selective, not automatic
After the task, the system decides whether and how to store the new experience.
Retaining every interaction produces memory bloat. Retaining only spectacular successes creates survivorship bias. A useful retention policy considers:
- Was the outcome independently verified?
- Is the case novel enough to add information?
- Does it contradict or refine an existing case?
- Is it safe and lawful to store the underlying data?
- Is the case time-sensitive?
- Should details be redacted while preserving the lesson?
- Does the case belong in episodic memory, semantic memory, or a procedural skill?
Retention may update an existing case rather than append a new one. It may also create a higher-level concept linked to several cases. Concept Memory for AI Agents is one way to represent that abstraction layer.
7. Modern language-agent examples
Several recent systems resemble pieces of CBR even when they do not use the label.
Generative Agents stored a memory stream, retrieved memories using recency, importance, and relevance, and generated higher-level reflections that influenced later plans. Its evaluation concerned believable simulated behavior rather than general task competence, but the memory–reflection–planning loop is structurally relevant. 2
ExpeL collected task experiences, extracted natural-language insights, and retrieved both insights and past trajectories at inference time without fine-tuning the base model. The paper reported improvements across its evaluation environments as experience accumulated. 3
Voyager stored tested code as an expanding skill library, retrieved skills for new tasks, and reused them in a new Minecraft world. Its cases were partly compiled into procedures. 4
Agent Workflow Memory made the case-to-procedure step explicit for web agents: it induces reusable workflows from past trajectories — offline from training examples or online from the agent's own experience — and supplies them to guide later tasks, reporting large relative gains on web-navigation benchmarks. In CBR terms, it automates a slice of the retain step: compiling clusters of procedural cases into named, reusable routines. 15
Reflexion closed the revise-and-retain loop most explicitly: after a failed attempt, the agent generates a verbal reflection on what went wrong, stores it in an episodic memory buffer, and conditions later attempts on it — improving success rates across coding, reasoning, and decision tasks without any weight updates. In CBR terms, it retains lessons abstracted from failures as first-class objects — the reflection, not the failed case itself. 10
ReasoningBank proposed distilling generalizable strategies from successful and failed agent experiences and retrieving those strategies for later web and software-engineering tasks. 5
These systems show that experience reuse can improve bounded agents. They do not establish a general-purpose memory that remains reliable across arbitrary domains or long deployment histories.
8. Argument from cases: the legal tradition
The lineage in section 1 is only half the ancestry. The other half came from AI and Law, which was a primary research stream feeding the birth of case-based reasoning and which contributed the part most modern agent-memory work leaves out: a case is not only something to copy from, it is something to argue with. 17
HYPO, built by Rissland and Ashley for trade-secret law, is the reference system. It represented disputes along dimensions, retrieved precedents that were more and less on point for each side, and produced an argument rather than an answer — with an extended worked example running a hypothetical trade-secrets case, patterned on a real one, through the whole cycle. 18 The structural move is the one worth stealing. HYPO retrieves the cases that hurt the position as well as the ones that support it, and its output is a claim, the best counter-citation, and the distinction between them.
CATO extended this with what Aleven called middle-level normative background knowledge, used for three things a bare similarity metric cannot do: organize an argument that spans several cases, reason about whether a difference between two cases actually matters, and judge how relevant a precedent is to the situation at hand. Aleven's evaluation found that arguments about the significance of distinctions, generated from this model, helped predict case outcomes in trade secrets law. 19 A second study, run inside a real legal-writing course, compared CATO against the best traditional instruction and found its example-based approach effective for basic argumentation skills — while reporting that a more integrated approach appears to be needed if students are to achieve better transfer of those skills to more complex contexts. 19 The qualified half of that result is the more useful one here: retrieving good examples teaches the local move, and on its own it did not carry students as far when the context got harder.
Three things this tradition gives an agent case base that the retrieve–reuse–revise–retain cycle alone does not:
- Distinctions are first-class. The important output of retrieval is often the difference between the retrieved case and the present one. A memory layer that surfaces only similarity has discarded the reasoning step.
- The strongest opposing case belongs in the packet. Section 11's negative-transfer failure mode is exactly what an adversarial retrieval discipline is designed to prevent, and law arrived at that discipline forty years ago because its practitioners face an opponent who will find the counter-case regardless.
- A case's force is graded, not binary. "More on point" is a relation between cases with respect to a claim, not a scalar attached to a case. Similarity scores in most agent memory systems are claim-independent, and a claim-independent score cannot by itself supply the distinguishing move — it can say two cases are alike, not which one cuts for you.
9. Textual case-based reasoning: from prose to a retrievable case
Section 2 specifies what a case should contain. It does not say who fills the fields in, and for an agent the honest answer is that nobody does: what actually exists is a transcript, a log, a ticket, or a diff. Turning unstructured text into something a case base can index is its own subfield, surveyed under the name textual case-based reasoning. 20 For an agent it is not an optional preprocessing step, because a transcript is the input most case bases actually receive, whatever structured fields the harness captures alongside it.
The AI and Law work again supplies the hardest measured version. Ashley and Brüninghaus built SMILE to classify case texts into the factors that HYPO-style reasoning needs, and IBP to predict outcomes from those factors, and reported empirical evaluations of both functions separately. 21 Splitting the system that way is the methodologically important part: it separates did we extract the right structure from the prose from did we reason correctly over the structure, and those two error sources have completely different fixes.
There is also a third option between extracting a case from prose and demanding that someone fill in a form, and it has its own literature: conversational case-based reasoning, in which retrieval proceeds as a dialogue that incrementally completes an underspecified problem description rather than matching a query stated in full up front. Aha, Breslow and Muñoz-Avila's survey sets out the problems that arise in that setting and evaluates the approaches proposed for them. 22 The setting is worth naming here because a language agent is already in it: it can ask, which means an underspecified problem description can be completed at query time. That recovers what the asker left out; it does not restore fields a stored case never had. The design question that follows is which fields are worth a question, and the answer turns on a distinction this section reaches below: asking is only useful for the fields the environment cannot fill on its own.
For an agent case base the same extraction-versus-reasoning split applies and is usually skipped. When a retrieved case leads the agent astray, the cause is either
- the case record misdescribes what actually happened — the outcome field says "resolved" because the transcript ended, the environment version was never captured, the stated problem is the user's first framing rather than the real one; or
- the record is accurate and the retrieval or adaptation was wrong.
Only the second is a reasoning problem. The first is an extraction problem, and no amount of better similarity search touches it. Two practical consequences:
- Instrument extraction separately. Sample records and check them against their source transcripts on a schedule. The measurement is cheap and it is the only thing that catches slow corruption of the case base at the point of writing.
- Prefer fields the environment can fill over fields the model must infer. Exit codes, test results, diffs, and timestamps are extracted, not judged. Free-text outcome summaries written by the same model that produced the trajectory are the least reliable field in the record and are frequently the one retrieval keys on.
10. CBR versus nearby approaches
| Approach | Stored unit | Main operation | Learning location |
|---|---|---|---|
| Few-shot prompting | Input/output examples | Imitate pattern | Current context |
| RAG | Relevant passages | Ground answer | External knowledge store |
| Case-based reasoning | Problem, solution, outcome | Adapt prior experience | Case library and adaptation process |
| Skill library | Executable procedure | Compose or call skill | Procedural memory |
| Fine-tuning | Training examples | Update parameters | Model weights |
| Concept memory | Reusable relational pattern | Map across contexts | Semantic memory linked to evidence |
The boundaries can blur. A case can contain a skill; a concept can be extracted from many cases; RAG can retrieve cases (the original RAG formulation was itself about grounding generation in retrieved knowledge for knowledge-intensive tasks, not about experience reuse 11). The distinctions remain useful because they identify what is being reused and what validation is required.
11. Failure modes
Wrong similarity
The retriever matches on shared words while missing a decisive structural difference.
Stale cases
The old environment, policy, API, or data format no longer exists. Every case needs version and time metadata.
Unverified outcomes
A task appeared successful because the agent stopped without checking. Retaining it turns a hidden failure into future guidance.
Negative transfer
The prior solution is actively harmful in the new context. Evaluation must measure worse-than-no-memory outcomes, not only average improvement.
Private-data leakage
Cases may contain personal conversations, credentials, proprietary documents, or sensitive inferences. Redaction after storage is weaker than minimizing what is stored in the first place.
Memory poisoning
Hostile content can be retained as a trusted lesson. Retrieved memory must remain lower-authority than system policy, and promotion should require validation. This attack has been demonstrated concretely: AgentPoison backdoors agents by poisoning a small fraction of the demonstrations in their memory or knowledge base — a poison rate under 0.1% was enough to get the poisoned demonstration retrieved about 82% of the time, with end-to-end attack success around 63% — which puts the retention gate on the security boundary alongside retrieval robustness, rather than leaving the ranking to carry it alone. 12
Premature generalization
One successful case becomes a universal rule. The system should preserve case-specific scope and distinguish evidence from abstraction.
12. Maintaining the case base
The case library is a living index, not an append-only transcript archive. The CBR community treated this as its own research area — case-base maintenance — with taxonomies of maintenance policies organized by what data they collect about the case base and how they revise it. 16 The dimensions below carry that framework's question into agent memory. The first three are write-time schema decisions that make maintenance possible at all; the last two are maintenance policies in the narrower sense the taxonomies use.
Stable identity
Every case needs an identifier that survives filename, title, or storage changes. The identifier anchors later concepts, skills, corrections, and audit records.
Versioned environment
Record model, prompt, tool, repository, API, dataset, and policy versions where they affect the outcome. “This worked before” is weak evidence when the environment has changed.
Outcome status
Useful states include:
- observed;
- apparently successful;
- verified;
- failed;
- contradicted;
- superseded;
- privacy-restricted;
- retired.
The difference between “apparently successful” and “verified” prevents premature retention.
Duplicate and cluster handling
Repeated near-identical cases may indicate a stable pattern, but storing each in full can overwhelm retrieval. A cluster can keep representative cases, frequency, variation, and exceptions without erasing the underlying records.
Forgetting
Retention policy should consider:
- age;
- retrieval frequency;
- outcome quality;
- sensitivity;
- redundancy;
- current environment coverage;
- downstream dependencies.
Deletion can be physical where privacy requires it. Otherwise a tombstone may preserve that a case was retired and which entries depended on it.
Case-base maintenance also has a competence literature worth inheriting. Smyth and Keane showed that the traditional deletion policies — deleting at random once the base exceeds a size cap, or deleting by low utility — can silently destroy a case base's problem-solving coverage, and proposed deleting by competence: model which cases other cases depend on for coverage, and preserve the load-bearing ones. 13 The same paper is an early treatment of the swamping problem — past some size, adding cases makes the system slower without making it more capable — which is exactly the memory-bloat curve modern agent stores are rediscovering.
13. Human and agent roles
Automation can structure episodes, propose indexes, retrieve candidates, and suggest adaptations. It should not automatically gain more authority merely by repeating its own summaries.
A practical responsibility split is:
| Step | Safe default owner |
|---|---|
| Capture raw task telemetry | Harness |
| Extract candidate case | Model or rule-based processor |
| Verify outcome | Tests, environment, or human |
| Classify sensitivity | Policy plus human review for uncertain cases |
| Promote reusable concept | Human or strong evaluation gate |
| Compile executable skill | Code review and tests |
| Change binding policy | Authorized owner |
The split can be relaxed for low-risk domains with exact outcome signals. A game agent can retain a verified crafting routine more freely than an advice agent can retain a rule about a person.
14. A practical evaluation
A useful CBR-agent evaluation needs more than a benchmark average.
- Build a time-ordered set of cases.
- Hold out later tasks with both close and misleading analogues.
- Compare no memory, passage RAG, raw-trajectory retrieval, and structured case retrieval.
- Measure task success, cost, retrieval precision, revision catches, and negative transfer.
- Introduce stale, contradictory, and poisoned cases.
- Grow the library to test whether retrieval degrades.
- Audit whether final actions can be traced to specific source cases.
The evaluation should also ask whether the system knows when not to reuse a case.
15. Bottom line
Case-based reasoning offers a useful correction to a common agent-memory fantasy. Memory does not become learning merely because an old transcript is searchable. Experience becomes reusable only when the system can represent the case, retrieve it for the right reason, adapt it to the new situation, verify the adaptation, and decide carefully what to retain.
Modern embeddings and language models make case storage and retrieval easier. They do not remove the need for the rest of the cycle.
Related concepts
Foundations: Agent Memory, Continual Learning in AI Systems, Analogical Transfer in AI, Concept Memory for AI Agents
Agent methods: Reflexion, Retrieval-Augmented Generation, Semantic Search Architecture for Code and Documentation, Agent Scaffolding
Risks: Memory Poisoning, Negative Transfer, Distribution Shift, Prompt Injection, Regression Testing for AI Systems