Reference

Continual Learning in AI Systems: Adapting Without Starting Over

Abstract. Continual learning is the attempt to make an AI system learn from a stream of new tasks, data, or feedback without erasing what it learned before. That sounds like ordinary learning, but most deployed language-model systems are much closer to frozen artifacts: their model weights change only during separate training runs, while day-to-day adaptation happens through prompts, retrieval, tools, and external memory. This article separates those layers, explains catastrophic forgetting and the stability–plasticity tradeoff, and shows why an agent that accumulates notes or skills is not automatically a continually learning model.

Coverage note: research and public systems checked through August 2026.

1. What continual learning means

A conventional training pipeline gathers a dataset, trains a model, evaluates it, and deploys a fixed checkpoint. When the world changes, the pipeline produces another checkpoint. Continual learning asks for something harder: can the system incorporate a sequence of new experiences while preserving useful performance on earlier ones?

Three conditions make the problem distinct:

  1. The data arrive over time. The learner does not necessarily receive one balanced, shuffled dataset.
  2. Old data may be unavailable. Privacy, storage, licensing, or simple scale can prevent full replay.
  3. Retention matters. Improving on the newest task is not enough if earlier capabilities collapse.

The last condition produces the central tension. A system needs enough plasticity to change when new evidence arrives, but enough stability to retain what remains valid. This is called the Stability–Plasticity Dilemma, and it predates deep learning: the trade-off was articulated in the adaptive-resonance literature of the 1980s and has organized the neural-network side of the field ever since. The major reviews of the area — Parisi et al. on lifelong learning with neural networks, and De Lange et al.'s survey of continual classification methods — both treat the dilemma as the framing constraint from which the method families below are derived. 8, 9

Continual learning is sometimes called lifelong learning or incremental learning. The terms overlap, but they are not perfect synonyms. “Lifelong” often carries a broader ambition—open-ended learning over a system’s operating life—while “incremental” can refer to narrower additions such as new classes or domains. The experiment must specify what changes and what must be retained. This is not pedantry: De Lange et al. document that reported results in the literature are frequently incomparable precisely because papers quietly assume different settings, memory budgets, and task-boundary information. 9

2. Catastrophic forgetting

When a neural network is fine-tuned on new data, the same parameters that supported earlier behavior are updated for the new objective. If the new gradient repeatedly pushes those parameters away from their earlier values, performance on old tasks can fall sharply. This is Catastrophic Forgetting.

The phenomenon is old and well-documented. McCloskey and Cohen demonstrated it in 1989 with sequential learning experiments on small connectionist networks: training on a second association task rapidly and severely disrupted performance on the first. 10 French's review a decade later located the cause in the same property that makes neural networks attractive — distributed, overlapping representations, where the parameters encoding one capability are reused by others, so gradient updates aimed at one behavior move many. 11 The problem survived the transition to deep learning intact: Goodfellow et al. investigated it empirically in modern gradient-trained networks and found that the severity depends on the relationship between old and new tasks and on training choices such as dropout, but that no standard configuration eliminates it. 12

Forgetting is not always a bug. A system should forget an obsolete address, a revoked policy, or a workflow that no longer exists. The problem is uncontrolled interference: the training process does not reliably distinguish stale knowledge from still-useful knowledge.

This makes “never forget” the wrong target. A practical system needs at least four operations:

Operation Question
Acquire What genuinely new information or skill should be learned?
Consolidate Which parts should become durable?
Revise Which earlier beliefs or procedures are contradicted?
Retire What should be removed, quarantined, or allowed to decay?

A system that only appends new memories may avoid parametric forgetting while creating a different failure: an ever-growing store full of contradictions, near-duplicates, and obsolete advice.

The Retire operation has its own research literature under the name machine unlearning: removing a training example's influence from a model that has already absorbed it, without retraining from scratch. Bourtoule et al., building on earlier formulations by Cao and Yang and by Ginart and colleagues, showed how much architecture it takes to make deletion efficient — their SISA design shards and slices training precisely so that unlearning one example means retraining one small piece. 24 The lesson transfers directly to memory stores: deletion that was not designed for at write time is expensive or impossible to demonstrate later. Data-protection law creates erasure obligations over personal data, and the supervisory reading has moved toward the model: a 2025 report hosted by the European Data Protection Board treats erasure as covering both the training data and its influence on the trained model, catalogues exact and approximate unlearning, and notes that verification of approximate unlearning still lacks strong guarantees. 31 No statute yet states a proof standard, so demonstrability remains a design burden, but one regulators have begun to describe rather than leave unnamed.

3. The main continual-learning settings

Results depend heavily on the setting. The three-way taxonomy below was formalized by van de Ven and Tolias, and it exists because the field kept talking past itself: the same method can look strong in one scenario and near-useless in another, so a result reported without its scenario is close to meaningless. 13

Task-incremental learning gives the system a task identity at training and at test time; test-time identity is the defining condition of the scenario, not a common convenience. 13 The learner may know whether it is doing translation, sentiment analysis, or a particular game. This is the easiest setting because routing information is supplied.

Domain-incremental learning keeps the output task stable while the input distribution changes. A classifier might encounter new cameras, regions, writing styles, or customer populations.

Class-incremental learning adds new output classes without giving the system a task label at inference. It must distinguish both old and new classes in one shared space. Van de Ven and Tolias's comparisons show this is the consistently hardest scenario: regularization methods that look competitive with task labels degrade sharply without them, while replay remains comparatively robust. 13 That asymmetry is one reason to expect replay to travel well beyond those benchmarks, despite its storage and privacy costs, though the paper itself leaves open whether generative replay holds up on more complex inputs, and says nothing about what practitioners actually deploy.

For language models, the stream can also change at different training stages:

  • continual pretraining adds new corpora or current knowledge;
  • continual instruction tuning adds tasks and response patterns;
  • continual alignment changes preferences, policies, or refusal behavior;
  • personalization adapts to one user or organization;
  • agent learning records experience outside the model weights.

These are not interchangeable. A method that preserves image classes across a benchmark sequence may say little about updating a language model’s factual knowledge without changing its safety behavior. The lifelong-learning survey for language models makes the same cut this article does: it organizes the field by where the change lives — continual pretraining, continual instruction tuning, continual alignment on the internal side; retrieval and tool use on the external side — precisely because methods and risks do not transfer across those stages. 4

4. Four broad method families

4.1 Replay

Replay mixes examples from earlier tasks into training on the new task. The stored examples may be real, compressed, selected, or generated by another model. Gradient Episodic Memory is an influential example: it keeps a small memory of earlier tasks and constrains updates so they do not increase loss on those memories. Generative replay trains a generator to recreate earlier data rather than storing all of it directly. 1, 2

Replay is often strong because it exposes the learner to the old distribution again. Its costs are equally direct: storage, privacy, licensing, selection bias, and extra training compute. Generated replay also inherits errors from the generator.

4.2 Regularization

Regularization methods discourage changes to parameters judged important for earlier tasks. Elastic Weight Consolidation estimates parameter importance and applies a stronger penalty to moving important weights. Its original experiments showed that this can preserve performance across sequential tasks better than ordinary gradient descent, while also documenting that it did not match separate models trained for each task. 3

Regularization is attractive when old data cannot be stored, but parameter importance is only an approximation. As the task sequence grows, constraints can accumulate until the model has too little freedom left to learn.

That failure direction — losing the ability to learn, rather than the ability to remember — has since been documented as a first-class phenomenon. Dohare et al. showed in Nature that standard deep-learning methods progressively lose plasticity under extended sequential training, eventually learning new tasks no better than a shallow network — and they observed it without any importance-weighted penalty in play, across architectures, optimizers, activations, batch normalization, and dropout. In their experiments L2 regularization substantially eased the loss, particularly combined with weight perturbation — a different operation from the one this section began with, since L2 shrinks weights toward zero while importance-weighted methods anchor them at their previous-task values, so the family word covers two opposite pulls — and continual backpropagation, which reinitializes a small fraction of little-used units, appeared to maintain plasticity indefinitely. 25 Continual learning has two ways to die, and a method obsessed with stability can pass every forgetting benchmark while quietly succumbing to the other one.

4.3 Parameter isolation and expansion

Another family assigns different parameters, adapters, experts, or subnetworks to different tasks. This reduces direct interference. Progressive Neural Networks are the clean early statement of the idea: freeze the columns trained on earlier tasks, add a new column per task, and connect it laterally so new learning can draw on old features without overwriting them — retention by construction, at the price of parameter growth that is quadratic in the number of columns, since each new column takes lateral connections from every earlier one. 14 PackNet showed the compressive variant: iteratively prune a single network and pack several tasks into disjoint parameter subsets, trading capacity per task for bounded size. 15 In the language-model era the same family reappears as parameter-efficient fine-tuning — low-rank adapters that leave the base weights frozen and can be attached, detached, or swapped per domain — which turns the isolation idea into cheap, composable modules rather than whole subnetworks. 16

The obvious cost is growth: if every new domain receives a new module, routing and storage eventually become the central problem. Isolation also assumes the task boundary is known or discoverable. Real streams rarely announce that “task seven begins now.”

4.4 External memory and tools

A system can keep the base model fixed and place changing knowledge outside it. Retrieval indexes, files, databases, skill libraries, and APIs can all update without retraining the model. Surveys of lifelong learning for language models increasingly distinguish this external-knowledge path from changes to internal model weights. 4

This is often the most operable route for current agents, because external state is inspectable and reversible in a way that a weight update is not. It is not the safest route in every respect: section 7 describes how a writable experience store turns a one-shot injection into a persistent one, an attack surface the weight-modifying families do not have in the same form. And it is still not continual learning in the strongest sense. The model has not internalized a new capability; the surrounding system has learned how to supply different context.

The obvious objection is that a model which picks up a task from retrieved examples at inference time is the model adapting, not the harness around it. That is true as far as it goes, and it is why the cut this article draws is persistence and locus of state rather than capability. In-context adaptation lasts as long as the context window and leaves nothing behind; a weight update persists across sessions and cannot be inspected or rolled back the way a retrieved document can. Both change what the system does next, and both can persist: a skill library or a memory index updated today is still there tomorrow. Only one changes what the model is between sessions, as opposed to the harness around it, and that is the distinction section 6's three-layer table is built on.

5. Continual pretraining in practice

The four families above are research abstractions. The operation actually performed on a production base model is narrower and has its own name. Continual pretraining takes an existing checkpoint and keeps training it on a new corpus, rather than restarting pretraining over the union of old and new data. The motivation is economic before it is scientific. A laboratory that retrains from scratch every time a fresh corpus arrives pays the full pretraining cost again for data the model has already learned, and the continued-pretraining literature exists because most of that cost can be avoided. 27

The systematic version of the move is domain-adaptive pretraining. Gururangan and colleagues ran a second pretraining phase on in-domain text across four domains (biomedical and computer-science publications, news, and reviews) and eight classification tasks, and found consistent gains in both high- and low-resource settings; a further phase on the task's own unlabeled data improved results again, even on top of the domain phase. 28 That result establishes the direction but not the hard case: each adaptation started from the same general checkpoint, so nothing had to survive a sequence of corpora. Ke and colleagues studied that harder setting directly, continually adapting one model through a series of unlabeled domain corpora with a soft-masking mechanism controlling how much each update may change, and reported both reduced forgetting and knowledge transfer between domains. 29

The practical recipe turns out to rest on two ordinary training choices in combination — the learning-rate schedule and replay of earlier data — with no further forgetting-specific machinery on top of them. A checkpoint at the end of a cosine decay sits at a very low learning rate; resuming training there adapts slowly, and re-raising the rate adapts faster but disturbs what the model already knows. Gupta and colleagues isolated exactly this warm-up question on Pythia 410M, continuing pretraining from the Pile onto SlimPajama at roughly 300 billion tokens on each side, and measured how re-warming choices trade early adaptation against validation perplexity on the original distribution. 30

Ibrahim and colleagues then showed how far the plain recipe goes. Learning-rate re-warming, re-decaying, and replaying a fraction of the previous data was sufficient to match full retraining from scratch on final loss and on averaged benchmark scores: first at 405M parameters under both a weak shift (English to English, between two standard pretraining datasets) and a stronger one (English to German), then at 10B parameters under the realistic weak shift, at a fraction of the compute. 27 The paper also proposes alternatives to the cosine schedule that avoid the forgetting that re-warming itself induces and are not tied to a fixed token budget.

Three things follow for anyone reading continual-learning claims about deployed models.

  • Replay is necessary, and it is not free. Ibrahim and colleagues attribute the result to the combination: re-warming is what lets the model adapt, replay is what keeps the old distribution's loss from climbing. On the averaged benchmarks at 10B the gap is small (47.53 without replay, 47.68 with 5% replay, 48.00 for joint training), but replay matters substantially for loss on the original data, and matching the retraining baseline required keeping and re-serving previous pretraining data. A system that cannot retain its old corpus, for licensing, privacy, or storage reasons, does not have access to the result these papers report.
  • These are corpus updates, not stream updates. Every study here consumes a large, curated, offline dataset. None of them learns from what individual users did yesterday, which is the thing product language usually means by "the model keeps learning."
  • The controlled evidence is at small and mid scale. 405M and 10B are where the recipe was compared head-to-head against full retraining. The same paper, in its updated version, notes that combinations of its techniques have since been applied to continually pre-train much larger models, with DeepSeek-AI among the groups reporting it. Large-scale application is on the record; what stops at 10B is the measured comparison with retraining from scratch. 27

6. Model learning, agent learning, and product learning

The phrase “the AI learned” can describe three different things:

Layer What changes Typical mechanism Main risk
Model Parameters or architecture Pretraining, fine-tuning, adapters Catastrophic forgetting and expensive regression
Agent External memory or procedures Notes, retrieved cases, skill libraries Retrieval drift, stale rules, prompt injection
Product Routing, ranking, prompts, curriculum A/B tests, user feedback, editorial updates Metric gaming and opaque behavioral change

An agent that stores a successful script in a skill library has learned in an operational sense. Voyager demonstrated this pattern in a bounded game environment: it stored executable skills, retrieved them for later tasks, and reused them in a new world without changing the underlying language model’s parameters. 5 Generative Agents demonstrated the memory-side counterpart in a simulated town: an architecture of episodic memory streams, periodic reflection into higher-level observations, and retrieval weighted by recency, importance, and relevance produced believably persistent behavior — again with frozen model weights doing none of the remembering. 26

That is valuable, but it should be named accurately: non-parametric adaptation through external memory. The distinction matters for evaluation. A changed model may generalize without access to the original store; a changed agent may fail as soon as retrieval misses or the skill no longer matches the environment.

7. Why continual learning is difficult for deployed language models

The stream is not clean

User interactions mix good corrections, misunderstandings, jokes, sensitive data, prompt injection, and conflicting preferences. Treating every interaction as training evidence would be unsafe. This is not hypothetical: Greshake et al. demonstrated indirect prompt injection against real LLM-integrated applications — adversarial instructions planted in content the system retrieves, which the model then follows as if they were the user's. 17 A learning pipeline raises the stakes on exactly this attack, because an experience store that writes retrieved or user-supplied text back into future context converts a one-shot injection into a persistent one: the malicious instruction is no longer in one conversation, it is in the memory.

The target moves

Facts, policies, and user preferences change at different speeds. A preference such as “use shorter answers” may be durable; a project deadline may be temporary; a safety rule may be high-authority and non-user-editable.

Evaluation must be longitudinal

A single post-update score hides whether the update helped the new task by damaging an old one. Continual-learning evaluation needs a matrix: after each learning episode, retest earlier tasks as well as the new task. Useful measures include average performance, forgetting, forward transfer, backward transfer, storage growth, and update cost. Regression Testing for AI Systems is therefore part of the learning mechanism, not merely a release ritual.

Errors can consolidate

If a system summarizes a bad experience into a general rule, later retrieval can make the mistake more persistent. Self-judged experience stores such as ExpeL and ReasoningBank are promising, but their evidence comes from bounded benchmarks and depends on the quality of outcome signals, reflection, and retrieval. 6, 7

Safety behavior can drift

Continual alignment is not only a retention problem. A model may preserve task accuracy while changing refusal thresholds, instruction hierarchy, or treatment of sensitive data. Qi et al. showed how cheap this failure is: fine-tuning GPT-3.5 Turbo on ten adversarially designed examples, at a cost of under twenty cents, left it responsive to nearly any harmful instruction, and — the finding that matters for continual learning — fine-tuning on entirely benign utility data also measurably degraded safety alignment, with no adversarial intent anywhere in the pipeline. 18 A pipeline that regularly updates weights from deployment data has the same exposure, and the only way to know whether it reproduces that degradation is to test for it. Capability, behavior, and safety regressions therefore need separate test suites.

8. A practical architecture for learning without silent drift

For many current applications, the most defensible design is staged:

  1. Record raw experience as an immutable episode. Preserve inputs, actions, tool results, outcome, time, and provenance.
  2. Do not immediately turn episodes into instructions. Separate observation from policy.
  3. Extract candidate lessons. State the scope, preconditions, expected benefit, and possible failure.
  4. Validate against held-out tasks and old regressions. A lesson must improve more than the case that produced it.
  5. Promote into a versioned memory or skill store. Keep the source episode and review trail.
  6. Retrieve narrowly. Match on task structure and constraints, not only topical similarity.
  7. Expire or revise explicitly. Time-sensitive lessons need review dates.
  8. Escalate to model training only when external adaptation is insufficient.

This architecture turns “memory” into a governed learning pipeline. It also keeps rollback possible.

9. Knowledge updates are not one operation

Language-model systems need to distinguish at least four kinds of update.

Additive updates

An additive update introduces a fact or skill that did not previously exist: a new API, research result, customer workflow, or vocabulary. External retrieval is often a good fit because the addition can remain versioned and sourced.

Corrective updates

A corrective update says an earlier belief was wrong. For model weights, this is the knowledge editing problem, and it is harder than it looks. Direct editing methods exist — locating and rewriting the parameters where a factual association lives has been demonstrated at the single-fact level 19 — but the survey literature on editing large language models documents persistent trouble with everything around the edit: making the change generalize to paraphrases, keeping it from bleeding into unrelated facts, and propagating its logical consequences, so that editing "the capital moved" also updates what follows from it. 20 An edit that lands cleanly on one probe and leaves the model inconsistent everywhere else is the parametric version of appending a correction without supersession.

For external memory, appending the correction is not enough if retrieval can still return the obsolete statement. The system needs supersession:

old claim → contradicted by → new evidence
         → superseded by → corrected claim

Both records may remain for audit, but ordinary retrieval should prefer the corrected one.

Temporal updates

Some claims were true and later stopped being true. “The current model supports this parameter” may be valid for one version and false for the next. The memory needs effective dates rather than a timeless contradiction flag. Dhingra et al. made the parametric version of this argument for language models: because ordinary pretraining mixes text written at different times without marking it, models average over eras — and jointly modeling text with its timestamp both improves recall of time-scoped facts and gives the model a usable notion of "as of when." 21 External stores need the same discipline for the same reason; a memory without effective dates reproduces the averaging failure one layer up.

Normative updates

Policies, preferences, and safety boundaries change for reasons other than factual discovery. Their authority comes from an owner or governance process, not from frequency in the data stream. A model must not infer a policy change because many users requested it.

These update types need different write permissions. Treating them as undifferentiated “new information” creates authority drift.

10. Continual-learning evaluation as an accounting system

A serious evaluation records performance after every update, not only at the end.

Imagine tasks A, B, C, and D arriving in sequence. After learning A, test A. After learning B, test A and B. Continue until the result is a matrix:

After learning Test A Test B Test C Test D
A score — — —
B score score — —
C score score score —
D score score score score

The filled lower triangle reveals three effects directly:

  • forgetting: an earlier task falls after later training;
  • backward transfer: later learning improves an earlier task;
  • plasticity: the system can still learn new tasks, read against a reference, since a falling diagonal alone cannot separate lost plasticity from a harder task; Dohare et al. compare against retraining from scratch for exactly this reason. 25

A fourth requires filling the upper triangle as well, testing each task before it has been trained on:

  • forward transfer: earlier learning helps a later task before direct training.

Even that is not quite enough for forward transfer as the Gradient Episodic Memory paper defines it, which compares against a reference row taken before any training at all. Two further effects are not in the matrix in any form, which is what makes this an accounting rather than a scoreboard:

  • capacity growth: performance is preserved by adding parameters or storage;
  • efficiency: the update’s compute, data, time, and memory cost.

This matrix protocol is not an invention of this article. Average accuracy plus backward and forward transfer were formalized as the evaluation triple in the Gradient Episodic Memory paper 1, and later work argued explicitly that the field's fixation on forgetting alone hides the rest of the accounting — proposing complementary metrics for model-size growth, memory footprint, and compute so that a method cannot buy retention invisibly with capacity. 22

For an agent with external memory, the same logic applies. Freeze the base model and harness, then add experience in time order. Retest earlier tasks with and without memory retrieval. Measure not only success but whether the system used the correct memory, ignored stale entries, and avoided negative transfer.

The benchmark should include revisions, not merely additions. A system that remembers every old fact can appear strong on an append-only stream while failing the ordinary requirement to update its beliefs.

11. Deployment gates

An update pipeline should stop when:

  • old-task regression exceeds the allowed bound;
  • the new gain disappears on held-out examples;
  • safety or instruction-hierarchy tests move materially;
  • the memory or model cannot identify the update’s provenance;
  • deletion or rollback cannot be demonstrated;
  • a sensitive data source lacks authority for the proposed use;
  • the update increases average performance by concentrating harm in a subgroup;
  • the system cannot distinguish a temporary episode from a durable rule.

These are not reasons to avoid learning. They are what makes learning an engineered capability rather than an uncontrolled side effect.

12. The honest 2026 assessment

Continual learning remains an open systems problem. Research has produced useful techniques for replay, regularization, modularity, and external memory, but there is no general method that lets a frontier language model learn indefinitely from uncontrolled deployment data while remaining stable, private, safe, and cheap. Wang and colleagues' comprehensive survey reaches essentially this position: it treats the method families as partial answers to different sub-problems, and frames the field's general objective as a stability-plasticity trade-off with adequate generalization under resource constraints, rather than as benchmark accuracy. 23

Current agents approximate parts of lifelong learning by accumulating retrieved experience, procedural skills, and product-level feedback. Those methods can be powerful precisely because they avoid changing model weights. They should be treated as a practical bridge, not as proof that catastrophic forgetting or autonomous knowledge revision has been solved.

The useful design question is therefore not “does this AI continually learn?” It is:

What changes after experience, where is that change stored, who can inspect or reverse it, and which earlier capabilities are retested?

Core concepts: Catastrophic Forgetting, Stability–Plasticity Dilemma, Lifelong Learning, Transfer Learning, Distribution Shift

Methods: Elastic Weight Consolidation, Gradient Episodic Memory, Experience Replay, Parameter-Efficient Fine-Tuning, Retrieval-Augmented Generation

Agent systems: Agent Memory, Concept Memory for AI Agents, Case-Based Reasoning for AI Agents, Self-Improving Software Systems, Regression Testing for AI Systems