Abstract. Analogical transfer is the use of a relationship learned in one domain to understand or solve a problem in another. It is more demanding than recalling a similar example: the source and target may look different on the surface while sharing a deeper structure. Language models can solve many analogy tasks and can be prompted to use analogous problems, but the robustness evidence available, from re-tests of 2023–2024 models, shows brittleness under paraphrase, altered symbols, and answer-order changes. For AI agents, analogy is promising as a bridge between project memory and new work, but it needs explicit mapping and falsification rather than confident resemblance.
Coverage note: sources checked through August 2026.
1. What analogy adds beyond similarity
Suppose one system stores an expired build cache and another stores an outdated vector index. The objects are different. The shared structure is that a derived artifact is trusted after its inputs or assumptions have changed.
An analogy maps relations in a source domain onto relations in a target domain:
source: old inputs → cached artifact → later job trusts artifact → silent error
target: old documents → search index → later query trusts index → stale answer
The useful part is not that both systems contain files. It is that the dependency relation and failure mechanism match.
Structure-Mapping Theory formalized this distinction. Its central claim is that good analogies preserve systems of relations more than isolated object attributes, with preference for coherent higher-order structure. 1
This separates several nearby ideas:
| Operation | What is shared |
|---|---|
| Literal similarity | Both relational structure and object attributes |
| Mere appearance | Object attributes, with little or no relational overlap |
| Categorization | Belonging to the same class |
| Metaphor | A selective mapping that may be explanatory rather than operational |
| Transfer learning | Representations or parameters learned on one task help another |
| Case reuse | A prior episode and solution are adapted |
| Analogy | Relational structure maps from source to target |
In practice these overlap. A case retriever may propose the source; analogy determines what, if anything, should transfer. The table's transfer-learning row also marks a scope boundary worth keeping: the machine-learning literature on transfer has its own taxonomy of what may transfer — instances, feature representations, parameters, or relational knowledge — and of when transfer hurts 12 — analogical transfer as treated here is a reasoning-time operation over explicit structure, not a training-time operation over weights. The two can feed each other, but their failure modes are measured differently.
2. The stages of analogical transfer
A complete analogy involves more than producing “A is to B as C is to D.”
Retrieve a source
The system must notice a potentially useful prior case. People often fail here even when they can understand an analogy after it is presented. The canonical demonstration is Gick and Holyoak's radiation-problem studies: most subjects who had just read a structurally identical military story failed to apply it to the medical problem until told the story was relevant, at which point solution rates jumped. 7 Retrieval, not comprehension, was the bottleneck. Their follow-up work located the remedy in schema induction: subjects who compared two source analogs and articulated the shared structure transferred far more reliably than subjects given one source — abstraction earned at encoding time paid off at retrieval time. 8 A memory system that retrieves only topical neighbours will miss distant but structurally useful sources; a memory system that stores comparisons, not just cases, gives retrieval something structural to match on.
Represent both domains
Objects, relations, goals, and constraints need to be made explicit. A bad representation can make the correct analogy invisible.
Map correspondences
The system aligns source elements with target elements. It should prefer a coherent set of relations over scattered local matches. This preference is computable: the Structure-Mapping Engine implemented Gentner's theory as an algorithm that constructs globally consistent mappings favoring deep, interconnected relational systems — the systematicity preference — and demonstrated it across scientific analogies. 9
Infer candidate consequences
Relations true in the source suggest hypotheses about the target. These are candidate inferences, not facts — SME's own vocabulary for them, candidate inferences, is the one this section uses, and the engine generates them by carrying over source relations that the mapping supports but the target has not yet confirmed. 9
Adapt
The source solution changes to fit target constraints. Analogy rarely licenses literal copying.
Test
The target domain decides whether the transfer works. A beautiful mapping can still fail because an unmodeled difference matters.
This sequence is a useful agent workflow because it creates checkpoints where a human or verifier can inspect the reasoning.
3. What language models appear able to do
Research has found substantial analogy performance in large language models.
A 2023 study compared a GPT-3 variant with human participants on letter-string analogies, verbal analogies, story analogies, and text versions of matrix-reasoning problems. It reported human-comparable or better performance on several tasks and described the ability as emergent. 2
An ICLR 2024 paper approached analogy as a prompting method: generate or retrieve analogous problems, solve them, and use those solutions to guide the target problem. It reported gains across several reasoning tasks. 3
Narrative analogy work has also tried to distinguish superficial resemblance from shared event and relational structure. The ARN benchmark emphasizes that two stories can be analogous because of how roles and events relate, even when the entities and setting differ. 4
These results justify a moderate claim: language models can often represent and use analogical relationships expressed in text, and explicit analogy prompts can improve some task performance.
They do not justify the stronger claim that current models perform robust, human-like abstraction across arbitrary domains.
4. The brittleness evidence
Analogy benchmarks are especially vulnerable to familiar templates. A model may solve a problem because its surface resembles material in training rather than because it extracted the intended relation. The general form of this concern is measured: Wu et al. tested models on counterfactual task variants — the same abstract task under unfamiliar conditions, such as arithmetic in base 9 or chess with a swapped initial position — and found consistent, often large performance drops relative to the default versions, indicating that some of what looks like general reasoning is procedure recall specialized to frequent conditions. 10 Analogy tasks, which are supposed to measure exactly the ability to leave familiar surface behind, need that control more than most.
Follow-up studies modified earlier analogy tasks while preserving their abstract structure. One found sharp performance declines for GPT models on variations of simple letter-string problems where human performance remained high. It also reported sensitivity to answer order and paraphrase in story analogies. 5
An earlier response to the 2023 results gave counterexamples in which simple variations broke the claimed zero-shot behavior and argued that stronger controls were needed to rule out memorization or template familiarity. 6
This disagreement is productive. It suggests that analogy should be evaluated using:
- unseen symbol systems;
- paraphrases;
- changed answer order;
- misleading surface similarity;
- distant analogies;
- explicit disanalogies;
- increasing relational depth;
- transfer to a real task rather than multiple-choice recognition.
A model that can explain an analogy after it is supplied may still be poor at retrieving the right source. A model that selects the right answer may still be unable to use the mapping reliably in a project.
5. Where the capability might come from
Sections 3 and 4 describe a capability that is real on some items and fragile on others. That
pattern is easier to hold together with a mechanistic hypothesis than without one, and the
mechanistic account most often reached for here concerns induction heads:
attention heads implementing the pattern-completion rule [A][B] … [A] → [B].
Olsson and colleagues argued that induction heads may be the mechanistic source of most in-context learning in transformers. Their central observation is a coincidence in training dynamics: induction heads form at precisely the point where in-context learning ability jumps, visible as a bump in the training loss. They present six complementary lines of evidence, and they are careful about the strength of each: strong causal evidence in small attention-only models, correlational evidence in larger models with MLP layers. 13 The paper describes itself as preliminary and indirect, and that framing should survive being cited.
If something in this family is what underlies analogical behavior, three patterns would be expected, and the section 4 evidence is consistent with each of them. This is accommodation rather than prediction — the hypothesis was proposed to explain in-context learning, not analogy, and it was reached for here after the brittleness results were in — so the fit is suggestive, not confirmatory:
- Performance tracks the availability of a completable pattern, not the presence of shared relational structure. A problem whose surface form resembles a familiar template is easy even when the analogy is bad; a problem with identical structure and unfamiliar surface form is hard.
- Counterfactual variants should hurt disproportionately. Rotating a task away from its common form — a different modular base, a permuted alphabet — removes the completable pattern while leaving the structure intact, which is precisely the manipulation that separated performance in the counterfactual-task work. 10
- Robustness should vary by item at least as much as by model. Which items break is more informative than how many, though nothing in the account removes the dependence on architecture, scale, or training data. The robustness re-tests in section 4 happen to report variant-by-variant degradation rather than a single score 5, which is the granularity this pattern needs, though that reporting choice was the authors' own and not a test of this hypothesis.
None of this settles whether a language model performs structure mapping in Gentner's sense. The honest position is that pattern completion at scale reproduces many analogical behaviors, that nobody has shown it reproduces all of them, and that the interpretability evidence is explicitly partial. Treating "the model does analogy" as a mechanism claim rather than a behavioral description is the error to avoid.
6. Abstraction benchmarks and what they settle
The field's dedicated measuring instrument for this capability is the Abstraction and Reasoning Corpus, and it should be read alongside the argument it was built to serve. Chollet's case is that measuring skill at a task measures very little on its own, because skill can be bought with priors or training data; what deserves measurement is skill acquisition efficiency relative to a stated set of priors. 14 ARC operationalizes that: small grid puzzles, a handful of input-output examples each, deliberately minimal required prior knowledge, and task sets built so the specific transformation should be new to the solver. That last property is a design intent rather than a guarantee: the public task sets have been on the open web for years, which is why the benchmark keeps a private evaluation set, and the memorization concern section 4 raises applies here in the same form.
That design makes ARC a clean test of one component of the question this article asks — can a system induce a relational rule from a couple of instances and apply it somewhere new — while leaving the others untested: there is no stored source to find, no retrieval step, and one representation. The analogy benchmarks in sections 3 and 4 test source-to-target mapping directly and ARC does not; the two kinds of score are not interchangeable.
ARC-AGI-2 sharpened the instrument after six years of results. It keeps the input-output pair format for continuity and replaces the task set with one curated for finer-grained discrimination at higher difficulty, reporting human testing as a baseline that shows the tasks remain accessible to people while being hard for current systems. 15 The human baseline is the part that matters for interpretation: a benchmark where humans succeed and machines fail localizes the gap, whereas a benchmark hard for both would only say the items are hard.
Three cautions before treating an ARC number as an analogy score:
- ARC measures induction from examples, not retrieval of a remote source. In ARC-AGI-1 and 2 a task's example grids are all on the screen; ARC-AGI-3, released in March 2026, changes the instrument to interactive environments in which an agent must explore and infer the goal with no worked examples at all 19 — a different test again, and still not one of finding a source in a store. Section 11's problem, finding the right source among thousands of stored memories, is not tested at all.
- Grids are one representation. A system that induces grid transformations well has not been shown to map relational structure between a database schema and a permissions model.
- The scores are contested at the margin. Compute budgets, program-synthesis search, and task-specific tooling all move the number, so a claim about ARC is a claim about a configuration, and the configuration belongs in the citation.
7. Analogy for AI agents
Agent systems create a practical opportunity. They already maintain project histories, tools, code, incident reports, and memory stores. Analogical transfer could help them reuse knowledge across those boundaries.
Architecture reuse
An agent may recognize that an event-sourced transcript viewer and a database audit log share an immutable-event-plus-projection pattern.
Failure shielding
A costly mistake in one pipeline may reveal a general preflight check for another.
Evaluation design
A benchmark-contamination defense may transfer to a newsroom’s claim-review workflow: keep a delayed or private set that the production system cannot optimize against.
Product design
The exploration–exploitation tradeoff in recommendation can illuminate the balance between a personalized learning stream and deliberate serendipity.
Security boundaries
A capability separation used in one plugin system may suggest a boundary in another agent host, while still requiring target-specific threat modeling.
These are hypotheses generated by mapping. The target system remains the authority.
8. A safe analogical-transfer record
When an agent proposes an analogy, it should produce a compact record:
| Field | Example question |
|---|---|
| Source | Which earlier case or concept is being used? |
| Target | What current problem is being addressed? |
| Shared relations | What exact structure matches? |
| Surface differences | What looks similar but is irrelevant? |
| Material differences | What could invalidate the transfer? |
| Candidate inference | What new idea follows if the mapping holds? |
| Test | What cheap check would falsify it? |
| Outcome | Did the transfer help, fail, or remain unresolved? |
This record makes analogy auditable and trains the memory system on both positive and negative transfer.
9. Failure modes
Surface matching
The source and target use the same terms but have different mechanisms. “Memory” in a context window, a database, and a biological system does not imply the same constraints.
Story completion
The model recognizes a familiar narrative and fills in the expected ending, even when the target details contradict it.
Mapping only the successes
An agent maps a technique’s benefits but ignores the conditions and costs that made it work.
One-source fixation
The first plausible analogy anchors the solution. Better systems should retrieve competing sources, including one that predicts failure.
Authority laundering
An analogy turns a preference or heuristic into an apparent rule. The source’s evidentiary status must carry through.
Disanalogy blindness
The system lists similarities without searching for the difference that matters most. A falsification prompt should be part of the workflow.
Retrospective elegance
After a project succeeds, it is easy to invent a clean analogy that did not actually guide the decision. Provenance should record whether the source was retrieved before action.
10. Evaluation for cross-project use
A realistic evaluation could start with a collection of completed projects and incident reports.
- Hide one target project.
- Ask the system to retrieve source cases, near-domain first, then from other domains.
- Require explicit relational mappings and disanalogies.
- Compare against topical semantic search and no-memory baselines.
- Have reviewers judge whether the source would have improved the target decision at the time.
- Test the proposed check or design change where possible.
- Measure false transfer, not merely idea count.
The dataset needs negative pairs: sources that look similar but would mislead. Otherwise a model can score well by producing plausible connections everywhere.
Useful metrics include source-retrieval precision, mapping accuracy, counterexample quality, target-task gain, and calibration. The system should be rewarded for saying “the analogy breaks here.”
11. Retrieval is often harder than mapping
Analogy experiments frequently present the source and target together. Cross-project use begins one step earlier: the system has to find the source among thousands of memories.
The cognitive-science answer to this scaling problem is the MAC/FAC architecture: a cheap non-structural filter over the whole memory ("many are called"), then full structural mapping over the survivors ("few are chosen"). 11 Its designers were explicit about the consequence — retrieval is dominated by surface similarity even in systems capable of deep mapping, which accounts for the access failure Gick and Holyoak documented. An embedding-plus-rerank agent memory whose first stage scores surface content has the same two-stage shape and inherits the same blind spot — an embedding trained over relational descriptions would not, by construction: a structurally strong source with little surface overlap is the case most likely to be lost before mapping ever runs.
Three retrieval strategies expose different candidates.
Surface retrieval
Find cases with similar words, tools, or entities. This is cheap and often useful within one domain. It is also the most likely to miss distant analogies.
Structural tags
Index cases by relations such as:
- shared mutable state;
- delayed feedback;
- irreversible action;
- derived artifact;
- capability boundary;
- exploration versus exploitation;
- hidden selection effect;
- principal–agent conflict.
This improves explainability but requires someone to define or infer the tags.
Model-generated abstraction
Ask a model to describe the problem shape, then retrieve on that description. This can find distant sources, but it can also bake a mistaken diagnosis into the query.
A hybrid system can retrieve a topical neighbour, a structural neighbour, and a deliberately contrasting case. The final mapping should not know which retriever the designer expects to win.
12. What analogy actually does in scientific practice
Popular accounts of analogy in science run on a handful of famous distant mappings: Rutherford's solar system, Kekulé's snake. If those are the model, the design implication for an agent is to search for remote sources. The empirical record says otherwise, and it is one of the more useful correctives available here because it was collected by observation rather than recollection.
Dunbar recorded and coded reasoning as it happened in molecular-biology and immunology laboratory meetings. Across 16 meetings in four laboratories, the scientists produced 99 analogies. Of those, 40 were within-organism and 57 were to another organism; exactly two were non-biological — that is, distant in the sense the creativity literature emphasizes. 16 The distant analogies that did occur were used to explain a concept to other people, not to formulate a hypothesis or repair an experiment.
The follow-up work refined rather than reversed this, and it cuts both ways. Dunbar and Blanchette report that the type of analogy scientists use changes with their goal, and that, unlike subjects in standard reminding experiments, scientists frequently used structural features in their analogies. But structural use was the minority: only a quarter of the analogies were based on structural rather than superficial features, and over 80% of that quarter served hypothesis formation. When scientists reached for an analogy to fix an experiment, which is where most of the near-domain cases sit, source and target shared superficial features: the same gene, the same proteins, a different incubation time. 17 So the two results fit together more carefully than "near-domain is still structural". Near-domain analogy is the common case and most of it is surface-level; the structural minority is where hypotheses come from, and a mapping from clam genetics to plasmodium is an instance of that minority, not of the rule.
Historical case studies of individual discoveries add the mechanism by which analogy changes what someone knows rather than merely illustrating it. Gentner and colleagues' study of Kepler's notebooks identifies four distinct effects (highlighting, projection, rerepresentation, and restructuring) within structure-mapping theory and its computational implementation. 18 Analogy there is not one operation but a family, and only the last of the four is the dramatic gestalt shift the popular account treats as the whole phenomenon.
For an agent that reuses experience across projects, this inverts the emphasis section 11 leaves implicit — that losing the distant source is the failure to design against:
- Near-domain retrieval is where the observed yield was, and it is the case usually dismissed as unambitious. The reported breakdown puts the analogies that repair a broken experiment in adjacent work, within the same field rather than across distant ones — a description of four laboratories, not a tested retrieval policy. One caveat before reading this as a finding about memory: 31 of the 57 other-organism analogies were generated by gene-sequence homology search, a database lookup rather than a reminding, so part of what looks like near-domain retrieval from memory was retrieval from a tool. 16
- Goal should condition retrieval. Fixing a failure and forming a hypothesis drew on different source distances in the same laboratories. A retriever with one similarity function for every request is modeling something the data says is not one behavior.
- Distant analogies earn their keep in explanation. That is a real use, since an agent explaining its reasoning to a person may reasonably reach far, but it should not be scored as if it were problem-solving transfer.
- The famous examples are weak evidence. Dunbar notes that scientists frequently did not recall, after the meeting, the analogies they had just used; the ones that survive into the historical record are selected for memorability rather than for causal role.
13. Analogy as hypothesis generation
An analogy can support several kinds of output:
Explanatory. “A context window is like working memory in this limited respect.” The goal is understanding, not direct action.
Predictive. “Because both systems exhibit the same feedback loop, the target may also concentrate exposure over time.” This needs empirical testing.
Design-generative. “The versioning strategy used for datasets may help this content pipeline.” This produces an option.
Normative. “Because one system required consent, the target should too.” This requires moral or legal reasoning beyond structural similarity.
Executable. “Run this adapted procedure.” This has the highest verification burden.
The analogy should carry its type. An explanatory metaphor should not silently become an implementation rule.
14. What would count as strong evidence?
For project reuse, strong evidence would show more than benchmark answer accuracy.
The system would:
- retrieve a useful source before seeing the target solution, near or distant — section 12 says distant sources were rarely used, section 11 says they are the hardest to find, and a strong test scores both;
- state a relational mapping that independent reviewers recognize;
- identify a material difference;
- propose an intervention derived from the mapping;
- improve the target outcome against a no-analogy baseline;
- repeat across several domains;
- abstain on deceptive near-matches;
- retain the result without turning one success into a universal rule.
This is closer to an end-to-end agent evaluation than a cognitive puzzle. It also makes failure informative: retrieval, mapping, adaptation, and verification can be scored separately.
15. Relationship to concept memory
Concept Memory for AI Agents proposes storing reusable patterns linked to their source cases. Analogy is the mechanism that applies one of those patterns in a new target.
The division of labor is:
- case memory preserves episodes;
- concept memory preserves abstractions;
- analogical retrieval finds structurally relevant sources;
- mapping proposes a transfer;
- verification decides whether it works;
- retention records the outcome.
This can approximate one part of continual learning without changing model weights. It does not make the system self-validating. Every layer can introduce error.
16. Bottom line
Analogical transfer is one of the most promising ways for AI systems to reuse knowledge across projects because it targets relations rather than vocabulary. Current language models show meaningful analogy capability, especially when the source is supplied and the task is expressed in text. The controlled-variation studies cited here, run on 2023–2024 models, found them brittle under paraphrase, altered symbols and answer order, and prone to confuse familiarity with abstraction; the original authors replied to the counterfactual critiques, arguing the materials had been misread and presenting evidence they say restores the finding 20; a 2025 study found models human-matched under some conditions and not others 21; and a comparative study of humans and models on strategic analogies reports a trade-off between retrieving candidates and matching them well 22. The brittleness is real and contested — not settled in either direction.
The practical posture is neither dismissal nor trust. Use the model to propose mappings, require it to name the disanalogies, and let the target domain test the inference. And do not dismiss the near source: the record in section 12 is a description of four laboratories rather than a tested policy, but it says the analogies that repaired work came from adjacent work, and a retriever tuned only to hunt for the remote source is optimizing for the case that record found rarest.
Related concepts
Foundations: Structure-Mapping Theory, Case-Based Reasoning for AI Agents, Transfer Learning, In-Context Learning
Agent memory: Concept Memory for AI Agents, Agent Memory, Continual Learning in AI Systems, Semantic Search Architecture for Code and Documentation
Evaluation: Construct Validity, Distribution Shift, Benchmark Contamination, External Validity, Negative Transfer