Reference

Concept Memory for AI Agents: Reusing Ideas Across Projects

Abstract. A concept memory is a proposed agent-memory layer that stores reusable patterns—design ideas, causal mechanisms, failure modes, and decision rules—between raw episodes and executable skills. It is not a settled technical standard. It is an architectural synthesis of older Case-Based Reasoning, modern agent memory, skill libraries, reflection systems, and analogical retrieval. The goal is to help an agent recognize that a pattern learned in one project may apply in another without copying the original implementation blindly.

Coverage note: this entry distinguishes established research from a forward-looking design pattern. Sources checked through August 2026.

1. The missing layer between a transcript and a skill

The survey literature on agent memory confirms how crowded and unsettled this layer already is: memory mechanisms are catalogued by what is stored (turns, trajectories, summaries, facts, skills), how it is written, retrieved, and reflected over, and how it is evaluated — with no consensus unit of storage across systems. 19 Current agent systems commonly store one or more of the following:

  • recent conversation turns;
  • raw task trajectories;
  • summaries of past interactions;
  • factual notes;
  • executable scripts or tool procedures;
  • user preferences;
  • vector embeddings used to retrieve any of the above.

Each has a clear use, and real systems increasingly manage them as an explicit hierarchy — MemGPT, for instance, treats the bounded context window as main memory and pages conversation history and facts in and out of external storage the way an operating system manages RAM. 8 But none of these layers is exactly a concept.

A transcript says what happened. A fact says what is believed. A skill says how to perform an action. A concept says something more general: what pattern was present, why it mattered, and where else it may apply.

Examples include:

  • “Keep capability separation as the security boundary.”
  • “Use a delayed public test set when a benchmark is likely to enter training data.”
  • “A user-facing summary should preserve a path back to the source claim.”
  • “When a long job depends on a cache, validate the cache’s provenance before starting.”

These are neither domain-free laws nor one-off memories. They are reusable abstractions with conditions.

2. What a concept record would contain

A useful concept record needs more structure than a prose note:

Field Purpose
Name Stable handle for retrieval and discussion
Plain-language statement What the concept claims
Problem shape Conditions that make it relevant
Mechanism Why it should work
Evidence Episodes, papers, tests, or systems supporting it
Counterexamples Cases where it failed or should not transfer
Preconditions What must be true before applying it
Adaptation notes What usually changes across domains
Confidence Strength and type of evidence
Review date When time-sensitive claims should be checked
Links Related concepts, cases, and skills

This makes the concept inspectable. A bare vector does not explain why two items were matched, and a bare summary does not say where the lesson stops applying.

3. The research lineages behind the idea

3.1 Case-based reasoning

Case-Based Reasoning solves a new problem by retrieving a similar prior case, reusing its solution, revising that solution, and retaining the new experience. Its classic “four Re” cycle—retrieve, reuse, revise, retain—already treats memory as part of problem solving rather than passive storage. 1

Concept memory changes the retrieval target. Instead of asking only “which old case looks similar?”, it also asks “which relational pattern explains why the old solution worked?”

3.2 Cognitive architectures for language agents

The CoALA framework organizes language agents around working, episodic, semantic, and procedural memory, together with internal and external actions. It does not define a separate concept-memory type, but it provides a useful map: concepts sit mostly in semantic memory while remaining linked to episodic evidence and procedural skills. 2

3.3 Reflection and distilled experience

Generative Agents stored observations and periodically produced higher-level reflections used in later planning. ExpeL extracted natural-language insights from experience and retrieved those insights and past trajectories for new tasks. ReasoningBank later proposed distilling strategies from both successful and failed trajectories. 3, 4, 5 Reflexion sits in the same family with a tighter loop: the agent converts failure feedback into verbal self-reflections stored in an episodic buffer, and those reflections measurably improve the agent's later attempts at the same task. 16 A-MEM pushed the organizational side further, borrowing the Zettelkasten method — atomic notes, explicit links, evolving organization — so that stored memories form a connected network rather than a flat list. 17

These systems support the basic possibility of moving from episodes to abstractions. Their benchmarks do not establish that the abstractions are valid across arbitrary projects.

3.4 Skill libraries

Voyager stored executable code skills and retrieved them for later tasks in Minecraft. The skill library was compositional and reusable, and the paper reported transfer to a new world. 6

A concept library would sit one level above this. “Retry this exact mining function” is a skill. “Turn repeated action sequences into named, tested routines” is a concept that could apply to browser automation, data pipelines, or robotics.

A growing library of named abstractions is not without precedent in language agents themselves: ArcMemo distils reusable, modular abstractions from solution traces into natural-language concept-level memory and retrieves them for new queries 26, and a decade of cognitive-architecture work includes an analogical concept memory for Soar that acquires concepts from interactively obtained examples 27. The most fully worked-out precedent for the compression view of such a library, though, comes from program synthesis. DreamCoder alternates between solving tasks and a "sleep" phase that compresses recurring solution structure into new reusable library primitives — and the abstractions it grows are interpretable enough that the system rediscovered textbook physics identities and classic algorithms from the programs it had already found while solving tasks, starting from a small hand-written set of domain primitives. 18 Its lesson for concept memory is the compression framing: an abstraction earns its place in the library by making many past solutions shorter to express by more than it costs to describe the abstraction itself — a description-length balance, and a testable criterion rather than a persuasive summary.

4. Retrieval by relation, not merely by topic

The hardest part is not storage. It is retrieval.

Semantic search is good at finding text about similar subjects. Cross-project reuse often depends on structural similarity instead:

  • a stale dataset cache and a stale search index are different objects but share a provenance problem;
  • a browser renderer and a plugin host are different systems but may share a capability-boundary problem;
  • a personalized reading stream and an adaptive test are different products but share an exploration–exploitation problem.

Analogical Transfer in AI calls this a mapping from a source domain to a target domain. Structure-mapping theory emphasizes relations between objects rather than shared surface attributes. 7

The cognitive-science literature is blunt about how hard this retrieval step is. Gick and Holyoak's classic experiments found that most people fail to spontaneously retrieve a structurally identical solution they were shown minutes earlier when its surface story differs — transfer jumped only when subjects were told the two problems were related. 9 Retrieval, not mapping, is the bottleneck. The MAC/FAC model responds to exactly this asymmetry with a two-stage design: a cheap, surface-level "many are called" filter over the whole memory, followed by an expensive structural-mapping "few are chosen" stage over the shortlist. 10 A concept-memory retriever that runs embedding search first and structural checks second is rediscovering that architecture, and should learn from its known weakness — whatever the first stage cannot see, the second stage never gets to judge.

A concept-memory retriever therefore needs several signals:

  1. topical similarity;
  2. problem constraints;
  3. causal or relational structure;
  4. required resources;
  5. failure-mode similarity;
  6. authority and risk level.

The system should be able to say why the concept was retrieved. “Both involve a mutable cache trusted by a long-running process” is a better explanation than “embedding similarity: 0.83.”

5. A concept lifecycle

Concepts should not become durable merely because an agent wrote a persuasive summary. The self-correction literature gives the reason: Huang et al. found that language models asked to review their own reasoning without external feedback frequently fail to improve it and sometimes make it worse — the model that produced the error is not a reliable judge of the error. 11 A lifecycle in which the same agent proposes, evaluates, and promotes its own abstractions risks that circularity. The transfer is not automatic — Huang et al. studied answer revision, not concept promotion, and ReasoningBank, cited above, distils strategies from self-judged outcomes and still reports gains — which is why the Test stage below leans on held-out cases and recorded counterexamples rather than on the proposing agent's confidence.

Observe

Capture a task episode with inputs, decisions, actions, outcome, and evidence.

Propose

Extract a candidate concept. State its scope narrowly and record alternative explanations.

Test

Apply it to at least one held-out case or retrieve historical counterexamples. Test the claimed mechanism, not just wording similarity.

Promote

Move the concept from “candidate” to “validated” only after evidence clears a defined bar. Promotion may still mean “use as a hypothesis,” not “treat as a rule.”

Instantiate

When retrieved for a new project, create an application record: target context, proposed mapping, mismatches, and required adaptation.

Review

After the outcome is known, update confidence and record whether the transfer worked.

Retire or split

Concepts that become too broad should be narrowed. Contradicted or obsolete concepts should remain auditable but stop appearing in ordinary retrieval.

This lifecycle is deliberately slower than appending notes. The cost buys protection against confident error accumulation.

6. Failure modes

Surface analogy

Two projects share vocabulary but not mechanism. A technique for database transactions may sound relevant to a browser session because both use “state,” while the actual consistency guarantees are different.

False abstraction

The agent compresses away the condition that made the lesson true. “Always cache stable prefixes” is weaker than “cache stable prefixes when the provider’s cache key and privacy boundary make reuse valid.”

Concept laundering

A weakly supported guess becomes a polished concept record, then future agents treat the record as evidence. Provenance must point back to the actual cases and sources.

The adversarial version of this failure is already demonstrated. PoisonedRAG showed that injecting a handful of crafted texts into a retrieval corpus is enough to make a RAG system return attacker-chosen answers for targeted questions. 12 AgentPoison extended the attack to agent memory and knowledge bases specifically: poisoning a small number of stored demonstrations backdoors the agents that later retrieve them, with attack success above 80% at poison rates under 0.1%. 13 A concept store is a higher-leverage target than either, because a single poisoned abstraction is retrieved across many future tasks by design. Write access to the library is a security boundary, not a convenience setting.

Retrieval monoculture

Once one concept is frequently selected, it creates more successful-looking applications and crowds out alternatives. Retrieval needs diversity and negative evidence. The dynamic is not speculative — it is the standard feedback-loop failure of recommender systems, where serving items shaped by past interactions amplifies the system's existing biases over time and progressively narrows what users are exposed to. 14 A concept library whose retrieval statistics feed its own confidence scores closes the same loop one level up.

Authority drift

A user preference, a safety rule, a local convention, and an empirical observation are different kinds of claims. A concept store must not flatten their authority.

Premature proceduralization

Turning a concept into an executable skill too early can automate a misunderstanding. Concepts should link to skills without being silently compiled into them.

Memory bloat

Near-duplicate concepts make retrieval noisy. Consolidation needs merge suggestions, but automatic merging risks erasing important distinctions.

7. A small example

Suppose an agent observes that a 40-hour experiment used only part of its intended dataset because a cache manifest came from an earlier run.

The raw episode includes the command, cache file, counts, failure, and correction.

A candidate concept might be:

Before an expensive job trusts cached work, compare runtime expectations against cache provenance and abort on mismatch.

Its structure includes:

  • problem shape: long or costly job; cached intermediate; expected schema or count;
  • mechanism: early validation prevents silent partial reuse;
  • preconditions: provenance fields exist or can be reconstructed;
  • counterexample: a disposable five-second build where validation costs more than recomputation;
  • possible skills: manifest validator, schema check, dry-run command.

The same concept might later be retrieved for a model-training dataset, a static-site build cache, or a search index. Each application still needs local verification.

8. A minimal storage model

A concept system does not require an exotic database. A Git-backed set of structured Markdown files can support the first experiment.

# Validate Derived Artifacts Before Expensive Runs

status: candidate
scope: long-running jobs with cached or generated inputs
mechanism: provenance mismatch can create silent partial reuse
confidence: medium
review due: 2027-02-01

## Preconditions
- Rework cost is material.
- Expected inputs can be stated.
- The artifact can be inspected before launch.

## Evidence
- case: training-run-cache-mismatch
- case: stale-search-index

## Counterexamples
- Disposable build where recomputation is cheaper than validation.

## Applications
- proposed: static-site asset manifest

The exact syntax matters less than several invariants:

  • concepts have stable IDs;
  • status is distinct from prose;
  • evidence links are machine-readable;
  • counterexamples are first-class;
  • applications do not silently edit the parent concept;
  • every revision is diffable;
  • deletion or retirement does not break the audit chain.

A graph database may later make relationships easier to query, but a knowledge graph does not validate its nodes. The primary challenge is governance, not storage technology.

9. Querying the library

A query should return a small evidence packet, not a wall of memories.

For a new task, the retriever might return:

  1. one high-confidence concept with a close structural match;
  2. one distant concept that offers a different architecture;
  3. one negative case warning against the leading analogy;
  4. the source cases and review dates;
  5. an explanation of each match.

The agent then writes a mapping proposal:

shared structure:
  both jobs trust a derived artifact before a costly run

material difference:
  the target manifest is deterministic and cheap to regenerate

candidate transfer:
  verify the manifest hash and expected entry count

falsification:
  regenerate the manifest and compare; if identical, cache staleness is not the cause

This intermediate object prevents the retrieved concept from becoming an ambient instruction. It also creates data for later evaluation: did the match help, and was the stated reason faithful?

10. The retrieval budget

Everything above assumes a concept, once retrieved, is available to the model. That assumption is weaker than it looks, and it is where a concept library most often fails quietly rather than loudly.

A concept record with preconditions, evidence, counterexamples, and provenance is expensive in tokens. A library of a thousand such records cannot be loaded; a retriever must choose perhaps three. So the effective library is not what has been stored but what the retriever can afford to surface, and every governance property described above (confidence, review date, negative cases) competes for the same budget as the concept's actual content.

Position inside that budget matters too. Liu and colleagues measured how language models use long inputs on multi-document question answering and key-value retrieval, and found performance highest when the needed information sits at the beginning or the end of the context and significantly degraded when it sits in the middle, including in models explicitly built for long contexts. 21 A concept packet appended in the middle of a long working transcript is not reliably read merely because it was retrieved.

Nor does a large advertised context window settle the question. RULER evaluated seventeen long-context models on thirteen tasks spanning multi-hop tracing and aggregation as well as retrieval, and found that near-perfect scores on the simple needle-in-a-haystack test coexisted with large degradation as length grew: of the seventeen models, all claiming 32K tokens or more, only about half sustained satisfactory performance even at 32K. 22 The relevant capacity for a concept library is the length at which the model still reasons over what it was given, not the length the model accepts.

Three design consequences follow, and together they are why a record schema as long as §2's is not in tension with a tight retrieval budget:

  • Store long, retrieve short. The full record is for audit and review. What enters the context is a compressed form: the claim, the preconditions, one counterexample, one provenance pointer.
  • Place deliberately. If a concept is meant to govern the work, it belongs at a position the model reads reliably, not wherever the memory layer happens to append.
  • Budget is a retrieval parameter, not an afterthought. "Return the top five" is a claim about how much of the model's attention the library is entitled to, and should be tuned and measured like any other.

11. Concept consolidation

Libraries accumulate synonyms:

  • “validate cache provenance”;
  • “preflight expensive jobs”;
  • “trust but verify generated artifacts.”

Automatic merging is risky because the phrases may have different scopes. A consolidation process can propose:

  • alias: two names for the same concept;
  • parent/child: one concept is a narrower case;
  • related: concepts overlap but differ materially;
  • conflict: they recommend incompatible actions;
  • supersession: a later concept replaces an earlier one.

The merge reviewer should see all supporting and contradicting cases. After a merge, old IDs should resolve to the new record so historical decisions remain understandable.

12. How to evaluate concept memory

An attractive demo is easy: show one helpful cross-project suggestion. A serious evaluation is harder.

Metric What it tests
Retrieval precision Were suggested concepts genuinely relevant?
Counterexample recall Did the system surface known failure conditions?
Transfer gain Did the concept improve success, cost, or reliability?
Negative transfer Did reuse make the target task worse?
Explanation faithfulness Did the stated mapping reflect the actual reason for retrieval?
Provenance coverage Can each material claim be traced to cases or sources?
Library growth Does quality hold as concepts accumulate?
Human correction cost How much work is needed to repair bad abstractions?

The comparison should include simpler baselines: retrieve raw episodes, retrieve summaries, retrieve skills, or provide no memory. Otherwise the “concept” layer may merely rename ordinary semantic search.

The negative-transfer row deserves emphasis because it is the best-documented risk in the closest adjacent literature. The transfer-learning field has studied negative transfer — source knowledge that makes target performance worse — long enough to have survey-level taxonomies of when it occurs: roughly, when source and target are less related than the transfer mechanism assumes, and when the mechanism has no way to notice. 15 Concept memory moves that gamble from a training-time decision made once to a retrieval-time decision made constantly, which is an argument for measuring negative transfer per application record rather than per benchmark.

13. The knowledge-acquisition bottleneck

Concept memory is a proposal to accumulate curated, explicit, reusable knowledge. That project has been attempted before at scale, and the attempt is the most informative prior evidence available.

Lenat and Feigenbaum argued in 1991 that competent behavior requires a large body of explicitly represented commonsense knowledge, and that the cost of acquiring it, not the reasoning machinery over it, is the binding constraint on knowledge-based systems — while also predicting that past some threshold of seed knowledge, acquisition would become cheap, because a system that knows enough can learn the rest by reading. 23 Cyc was the long-run test of that thesis: a decades-long project to hand-encode commonsense assertions into a formal representation with an inference engine over them. Lenat's 1995 account describes the scale of the investment plainly, as infrastructure rather than as a product. 24

The retrospective judgment matters more than the original claim, and it is available from the project's own side. Lenat and Marcus, writing in 2023 about what language models might take from Cyc, list sixteen desiderata for trustworthy AI and argue for curated explicit knowledge with auditable step-by-step provenance, while stating the catch directly: a representation expressive enough to capture what we mean is expressive enough to be slow to reason over, a cost they say Cyc spent decades engineering its way out of. 25 That it is also expensive to populate is not their catch but their history — the same essay counts the effort in the thousands of person-years. The essay is an argument for curated knowledge that is candid about the price.

For a concept library the lesson is specific and should temper the design rather than kill it:

  • Acquisition cost dominates the supply side. Section 4 argued that retrieval is the hard part of using a library; this section's claim is about building one. The interesting engineering on the supply side is not the store. It is whatever makes a correct, reusable concept cheap to produce and cheap to review.
  • Extraction is not the bottleneck any more; curation still is. A language model can propose candidate concepts at essentially zero marginal cost, which is the supply-side inversion the 1991 paper itself predicted, arriving by a different route — and it leaves the review side exactly where it was. Section 5's promotion gate is therefore the load-bearing component, not the extraction step feeding it.
  • Formality has a cost curve. The more formal the record, the more auditable and the more expensive to write. A concept library that starts at natural-language claims with structured preconditions is choosing a point on that curve, and should say so rather than drift toward formality by accretion.

14. Where the idea stands

Concept memory is best treated as a design hypothesis, not an established category with a settled implementation. Its ingredients are real: case-based reasoning, semantic and episodic memory, reflection, skill libraries, analogical mapping, and retrieval. The proposed contribution is their arrangement around a governed, provenance-linked unit of reusable abstraction. The unit itself has precedents in language agents and in cognitive architectures (section 3.4); what this article adds is the governance around it, not the idea of it.

The promise is substantial. A system could accumulate a searchable map of what it has learned across projects without retraining its base model after every experience. The danger is equally substantial: it could create an increasingly authoritative library of elegant but false lessons.

There is one long-running experiment worth studying before building another. NELL — the Never-Ending Language Learner — ran from 2010 onward, extracting beliefs from the web into a growing knowledge base with per-belief confidence and coupled learning tasks. Self-reflection — an agent monitoring its own performance and allocating effort accordingly — is the property its authors name as missing: the retrospective lists adding it as future work and says NELL has a very weak ability to monitor its own performance. Its published retrospective is candid about what kept quality tolerable: promotion thresholds, redundancy across learning methods, and a continuing trickle of human feedback on promoted beliefs — and about the residual error that accumulated anyway. 20 NELL's lesson is not that lifelong knowledge accumulation fails; it is that the governance machinery is the system, and the extraction machinery is the easy part.

The right first implementation is therefore small, read-only at retrieval time, retrieved by relation rather than by topic alone, promoted only through a review step that assigns confidence from held-out cases rather than letting the proposing agent assert it, and evaluated against negative transfer.

Foundations: Agent Memory, Case-Based Reasoning for AI Agents, Analogical Transfer in AI, Continual Learning in AI Systems

Implementation: Semantic Search Architecture for Code and Documentation, Vector Index Design, Incremental Indexing, Agent Skills, Source Provenance and Claim Traceability

Failure modes: Retrieval Drift, Negative Transfer, Memory Poisoning, Automation Bias, Goodhart's Law in AI Systems