Abstract. A learner model is a system’s working estimate of what a learner knows, wants, has seen, and may need next — the decision layer beneath adaptive tests, tutors, and learning streams. Classical models track mastery of explicit skills, modern systems can run sequence models and language-model inference over richer interactions, and greater expressive power does not guarantee a truer model. The essential discipline is to keep mastery, interest, preference, and identity separate; preserve uncertainty; and evaluate the model by whether it improves learning decisions.
Coverage note: sources checked through August 2026.
1. The learner model is not the learner
A learner model is a representation built for a purpose. It may contain probabilities such as:
- 70% chance the learner has mastered a concept;
- low confidence because only two responses exist;
- high stated interest in AI governance;
- no evidence of prior exposure to reinforcement learning;
- one companion section completed;
- a preference for beginning with economic examples.
These are claims about available evidence, not a diagnosis of the person.
The distinction matters because adaptive systems act on the model. If the model says a concept is mastered, the system may skip it. If it says the learner is struggling, the system may simplify material. A mistaken estimate can become self-confirming when the learner stops receiving opportunities to show otherwise.
2. What should be modeled
Several dimensions are useful, but they should not be collapsed into one score.
Knowledge
Which concepts or procedures appear understood? What misconceptions are plausible? How certain is the estimate?
Exposure
What has the learner opened, completed, or practiced? Exposure is not mastery.
Goals
Is the learner trying to build a system, understand public debate, follow research, or form a philosophical view?
Interest
Which topics invite deeper engagement? Interest can guide entry points without determining the entire curriculum.
Constraints
How much time is available? Is the learner on a phone? Are accessibility accommodations needed?
Strategy and behavior
Does the learner request hints, jump to answers, revise after feedback, or express uncertainty? These signals are context-dependent and easy to misinterpret.
Affect
Frustration, boredom, and confidence can matter for tutoring, but affect inference from text is especially uncertain and culturally sensitive. The affective-computing literature's own interdisciplinary review is candid that affect detection rests on contested emotion theory, that signals are noisy and context-bound, and that models built in one setting degrade in others. 6 It should rarely be a durable hidden field.
3. Item response theory: the measurement layer underneath
Before a system can track what a learner knows, it needs a defensible account of what a single answer tells it. That is a measurement problem, and it has a mature answer that predates every tracing model in this article.
Item Response Theory models the probability of a correct response as a function of a latent person parameter, usually ability, and one or more item parameters. The simplest form, the Rasch model, uses a single item parameter for difficulty. Richer forms add a discrimination parameter, describing how sharply an item separates learners near its difficulty, and a guessing parameter for multiple-choice formats. Reise, Ainsworth and Haviland's review sets out the fundamentals and the practical appeal for applied work. 11
Two properties are what make this worth a section rather than a footnote.
Item and person parameters live on the same scale. A difficulty and an ability are expressed in the same units, so "this learner is likely to get this item right" becomes arithmetic rather than judgment. Nothing in a mastery probability supports that comparison: a BKT mastery estimate of 0.8 is not a statement about any particular item.
Estimates are meant to be invariant. Under the model's assumptions, item difficulty does not depend on which sample answered it and ability does not depend on which items were administered. Wright's account of the Rasch model is explicit that this is the point of the exercise rather than a technical nicety. It is what allows two learners who answered different questions to be compared at all. 12 It is also an assumption that fails loudly when it fails, which is a virtue: differential item functioning tests are diagnostics that a mastery-probability model does not offer an equivalent of.
This is also the machinery underneath adaptive testing. Choosing the item that yields most information about the current ability estimate is a computation over calibrated item parameters. 10 That dependence on calibration is, on the reading taken here, why Computerized Adaptive Testing grew out of psychometrics rather than out of the tutoring-systems tradition.
The reason IRT does not simply replace knowledge tracing is equally clear, and it is one sentence: classical IRT models a static ability, whereas an adaptive tutor exists to change the thing it is measuring. Learning during a test is a nuisance parameter for a psychometrician and the entire objective for a tutor. The models in the next three sections all track change over time rather than a fixed ability, though not every one of them does it with a per-learner parameter: the performance-factor model in section 4 drops the ability term and tracks practice counts instead. Nor do they share one ancestry. Only the additive and performance-factor models descend from IRT's logistic form; Bayesian knowledge tracing came out of the ACT-R tutoring tradition instead, and deep knowledge tracing sets aside the interpretable parameters that give IRT its measurement discipline in the first place.
4. The logistic family: modeling learning curves directly
Between item response theory and hidden-state tracing sits a family that is often the one actually used in educational data mining, because it makes the learning curve explicit and it fits with ordinary logistic regression.
Additive Factor Models extend an IRT-style logistic form with a term for practice opportunities: the log-odds of success rises with each attempt at a skill, at a rate estimated per skill. Cen, Koedinger and Junker built Learning Factors Analysis around this, and the important move was what they did with it. The contribution is not better prediction but cognitive model evaluation. Given a candidate mapping from items to skills, the fitted learning curves say whether that mapping is any good, and a search over splits and merges of skills proposes better ones. 13 That reframes the skill map of section 8 from an authored artifact into an estimated one that can be wrong in measurable ways.
Performance Factors Analysis changed the practice term. Instead of counting opportunities, it counts prior successes and prior failures on the skill separately, with separate coefficients, so that getting an item right and getting it wrong move the estimate by different amounts and the model does not need a per-student ability term to work. Pavlik, Cen and Koedinger presented it explicitly as an alternative to knowledge tracing. 14
Three practical reasons this family keeps being chosen:
- It is a regression. Fitting, confidence intervals, and model comparison are standard, and a coefficient that comes out with the wrong sign is legible as a modeling error rather than disappearing into a hidden state.
- The skill map is testable. Learning Factors Analysis makes "is this the right set of skills?" an empirical question rather than a design meeting.
- Failures are counted, not just successes. A learner who has failed a skill six times and succeeded twice is in a different position from one who has succeeded twice on the first two attempts. A model that counts only successes cannot represent that; the failure-counting form does, and so, by a different route, does the Bayesian model in section 5, whose posterior moves down on every failure.
5. Bayesian Knowledge Tracing
Bayesian Knowledge Tracing (BKT) is a classic learner-modeling method, introduced by Corbett and Anderson for the ACT cognitive tutors. It represents each skill as a hidden mastery state and updates the probability of mastery after each observed response. 5
A standard BKT model uses four parameters:
- initial probability of knowing the skill;
- probability of learning it after an opportunity;
- probability of guessing correctly without mastery;
- probability of slipping despite mastery.
This is attractive because the model is compact and interpretable. It can say that a wrong answer does not prove non-mastery and a correct answer does not prove mastery.
Its assumptions are also strong. Skills need to be defined in advance. Responses need to be tagged to skills. The usual model treats mastery as binary and often assumes no forgetting during the sequence. Real knowledge is multidimensional, strategies overlap, and item quality varies.
6. Deep Knowledge Tracing
Deep Knowledge Tracing (DKT) uses a recurrent neural network to predict future student responses from interaction sequences. The original paper reported better prediction than prior methods on several educational datasets and argued that the learned state could reveal curriculum structure. 1
The approach relaxed the need to hand-specify a simple mastery transition for each skill. It also made the internal learner state harder to interpret.
Follow-up work challenged the magnitude and meaning of the gains. Khajah, Lindsey and Mozer hypothesized four regularities DKT could exploit that standard BKT could not — recency, the contextualized trial sequence, inter-skill similarity, and individual variation in ability — and showed that a BKT extended with those capabilities, all previously proposed in the literature, reached a level of performance indistinguishable from DKT's; their conclusion is that DKT's gains do not come from discovering novel representations. 2 Separately, Yeung and Yeung documented undesirable DKT behavior, including predicted mastery changing in counterintuitive ways, and proposed a regularization term to address it. 18 The larger comparison came in 2020: across nine datasets, Gervet and colleagues found the ordering depended on the data — deep models led on large datasets and where the precise order of interactions carried the most signal; logistic regression led on moderate datasets and, notably, where learners had thousands of interactions, because DKT lost track of long-term information; and the extended BKT matched DKT only on the two datasets it had been developed on, lagging or proving infeasible to fit on the rest. 20 The prediction-is-not-measurement point below does not depend on which model wins; the point that no single model wins is worth having on its own.
This debate illustrates a general rule: better next-answer prediction is not automatically better knowledge measurement. A sequence model may exploit behavior patterns that help prediction without representing transferable understanding.
7. Prediction, mastery, and transfer
Knowledge tracing can be evaluated against several targets:
- correctness on the next similar item;
- correctness on a later item;
- performance on a post-test;
- transfer to a different problem format;
- ability to explain the concept;
- retention after time passes.
These targets are not equivalent.
A model can predict the next response by learning local regularities: item difficulty, guessing patterns, or the sequence in which a platform presents exercises. That may help the product choose the next item while providing weak evidence about durable knowledge.
Research extending DKT and related models to predict post-system performance makes the distinction explicit: the educational goal is performance beyond the immediate tutoring environment, not merely fit to the interaction log. It also supplies this article's one data point that reaches outside the log: averaging DKT's and DKVMN's correctness predictions across the problems a skill was practised on, rather than taking the final estimate, produced knowledge estimates that correlated better with a post-test than BKT's or PFA's standard mastery estimates did, and the same averaging improved BKT's and PFA's estimates too. 4 The post-test covered the same decimal skills in the same tutor, so this is evidence about external measurement, not about transfer to new problems, and its authors limit it to one tutor, one domain and one test. Whether a sequence model's internal state carries durable knowledge, as opposed to predicting the next answer well, is the part that stays open.
The learner model should therefore name its target. “Estimated probability of answering the next tagged item correctly” is more honest than “mastery” when transfer has not been validated.
8. Skill maps and prerequisite graphs
Many learner models assume a map from content to skills.
A skill map can represent:
- concepts;
- procedures;
- prerequisite edges;
- common misconceptions;
- alternative routes;
- evidence items;
- mastery thresholds.
The map introduces human judgment. If an item is tagged to the wrong concept, the model updates the wrong state. If the graph assumes one prerequisite route, it may misread a learner who arrived through another. The tagging problem is old enough to have a name: the item-to-skill matrix descends from Tatsuoka's rule-space work, which built misconception diagnosis on exactly such a mapping — and inherited exactly this dependence on the mapping being right. 7
Language models can help propose tags and prerequisites, but those suggestions need review. A confident generated curriculum graph can be more damaging than an incomplete one because it creates systematic adaptation.
For a small compendium, an explicit editorial map may outperform a learned latent representation:
entry → concepts explained
entry → concepts assumed
entry → possible next entries
entry → difficulty and evidence type
This supports fixed streams and later adaptation while remaining inspectable.
9. Bayesian updating as a design discipline
The importance of BKT is not that its four-parameter form is universally correct. It demonstrates a useful posture:
- mastery is hidden;
- responses are noisy evidence;
- correct answers can be guesses;
- wrong answers can be slips;
- beliefs should update gradually.
Later Bayesian-network variants have modeled individual differences and richer dependencies, but the same caveat remains: the model structure defines what kinds of learner differences can be represented. 3
An LLM-based learner model can preserve this discipline without using BKT literally. It can attach confidence, evidence count, alternative explanations, and an expiry to each hypothesis.
10. Language models as learner models
The survey literature on language models in education catalogues these uses — automated assessment, feedback generation, misconception diagnosis, content adaptation — alongside the corresponding concerns about reliability, bias, and over-trust; the field's breadth is real, and so is its thinness on durable learning outcomes — a 2026 meta-analysis of thirty-five experimental studies now quantifies effects on student outcomes 21, but longitudinal and transfer measures remain the rare case. 8
Language models can infer richer hypotheses from free text:
- identify a misconception in an explanation;
- distinguish a calculation error from a conceptual error;
- summarize a learner’s stated goal;
- propose a prerequisite;
- generate a targeted follow-up question.
This flexibility is useful because much learning evidence is not multiple choice.
It also creates new risks:
- eloquent answers may be overestimated;
- dialect or second-language writing may be mistaken for low understanding;
- the model may infer sensitive traits unrelated to instruction;
- a persuasive explanation of the learner may exceed the evidence;
- the profile may change with prompt wording or model version.
An LLM-generated profile should therefore be treated as a set of testable hypotheses. Each important inference needs evidence, confidence, and an easy correction path.
11. Learner model versus user model
User Modeling in LLMs can include tone, preferences, identity, history, and likely intent. A learner model has a narrower obligation: improve learning decisions.
| Field | User model | Learner model |
|---|---|---|
| Preferred response length | Often relevant | Relevant to presentation, not mastery |
| Knowledge of gradient descent | Sometimes relevant | Core learning state |
| Political identity | May be inferred by some systems | Usually unnecessary and high-risk |
| Chosen topic stream | Relevant | Relevant |
| Recent wrong answer | Ordinary interaction detail | Evidence, with uncertainty |
| Accessibility need | Relevant | Relevant if volunteered |
| Private life history | May appear in conversation | Usually should not be retained |
Purpose limitation is protective. A learning system does not need a general personality dossier to decide which explanation comes next.
12. The cold-start problem
At the first visit, the system knows little. There are several options:
Default path. Start with a general sequence and adapt later.
Explicit self-selection. Ask the learner to choose an entry point or goal.
Short diagnostic. Use a small set of questions to estimate prior knowledge.
Progressive profiling. Infer only what becomes necessary during use.
The worst option is often a long intake that gathers broad personal information before delivering value.
Cold-start design should minimize regret. If the first estimate is wrong, the learner should be able to change paths without losing progress.
13. Uncertainty and evidence
Every stored claim should answer:
- What evidence produced it?
- How much evidence exists?
- When was it observed?
- Is it self-reported or inferred?
- What alternative explanation fits?
- What action will depend on it?
- When should it expire?
A simple evidence table is often enough:
| Claim | Evidence | Confidence | Last checked | Learner control |
|---|---|---|---|---|
| Interested in AI hardware | Explicit stream choice | High | Current session | Change stream |
| Understands scaling laws | Two correct applications | Medium | This week | Review topic |
| Prefers concise prose | One setting choice | Medium | Current | Edit setting |
| Struggles with probability | One wrong answer | Low | Current session | Dismiss inference |
The table is intentionally modest. It discourages the fiction of a complete psychological model.
Making the model visible to the learner is itself a studied design, not an afterthought. The open-learner-model literature — synthesized in the SMILI framework — catalogues a decade of systems that expose the learner model for inspection, negotiation, and correction, and the reported benefits go beyond privacy hygiene: seeing the model supports reflection, planning and self-monitoring, and can serve assessment; whether inspection makes a learner's self-assessment more accurate is a claim the framework organizes rather than establishes. 9 The evidence table above is a minimal open learner model; the framework is the argument for why "learner control" deserves its own column.
14. Choosing the next action
The learner model does not teach by itself. A policy chooses what happens next.
Possible policies include:
- highest information gain, as in adaptive testing;
- prerequisite repair;
- spaced review;
- learner-selected interest;
- expected learning gain;
- productive struggle within a challenge band;
- deliberate diversity or serendipity;
- time-aware completion.
These goals can conflict. An information-maximizing item may be unpleasant. An interest-maximizing article may repeat familiar material. A completion-maximizing path may become too easy. The first policy's home discipline makes the tension concrete: computerized adaptive testing selects items to maximize measurement information near the current ability estimate, which is demonstrably efficient for assessment — the classic educational applications cut test length substantially at equal precision — but measurement efficiency was never a theory of what a learner should study next. 10 Importing CAT's selection logic into instruction without importing that caveat conflates the two.
A transparent system can state the reason:
Suggested because it connects your chosen economics entry point to a prerequisite you have not yet marked complete.
This is better than a mysterious “recommended for you.”
15. Wheel-spinning and gaming: what a mastery estimate cannot see
Every model in this article that tracks a moving mastery estimate assumes that repeated practice moves a learner toward mastery. Two well-documented behaviors break that assumption, and both are invisible to a mastery probability read on its own.
Wheel-spinning is the state of practicing a skill repeatedly without ever reaching mastery. Beck and Gong studied this across ASSISTments and the Cognitive Tutor and reported a blunt regularity: a student who does not master a skill quickly is likely to keep struggling and probably never masters it at all. 15 The behavior matters because a tutor's default response to a low mastery estimate, assign more of the same, is precisely the wrong action for a learner in this state, and the mastery estimate itself is what triggers it. A model that only reports how far from mastery cannot distinguish a learner making slow progress from one making none, because both look like a mastery probability that has not risen yet. The distinguishing signal is in the derivative and in the opportunity count, not the level.
Gaming the system is the complementary failure: exploiting the tutor's own affordances, such as rapid guessing or hint abuse until the answer is given, to advance without learning. Baker, Corbett, Koedinger and Wagner's classroom observations of off-task and gaming behavior in cognitive-tutor classrooms established that this occurs at meaningful rates. 16 Later work built detectors that identify it from interaction traces, across tutor subjects and classroom cohorts. 17 For a learner model the consequence is severe: gaming produces correct responses, so a mastery estimate rises while learning does not. A model that counts only successes will confidently report progress. PFA's failure term and DKT's non-monotone predictions can register the rapid wrong answers and hint abuse that gaming also produces, which is why a separate detector is needed precisely where the mastery model cannot tell the two apart.
A detector for either behavior is itself a model with its own generalization question, and the answer is mixed, and the mixed part is in the direction that matters most. Baker and colleagues built a machine-learned gaming detector and tested where it held: it transferred successfully across student cohorts, and less successfully across tutor lessons. 17 A follow-up the next year narrowed that: a detector trained on several lessons at once did transfer to lessons it had not seen, without retraining. 19 Read carefully, the first result is still the harder of the two results to live with. Transfer across cohorts means a detector survives a new group of students, which is the transfer a deployment needs when it scales to more users. Transfer across lessons is the one it needs when it scales to more content, and content is what an AI tutor adds fastest. A gaming detector validated on one subject should be treated as unvalidated on the next one until it is re-checked, and the same caution belongs on any behavioral detector bolted onto a learner model.
What this asks of a learner model is a change of shape rather than a better fit:
- Track the trajectory, not only the level. Opportunity count, time to first success, and whether the mastery estimate is still moving separate wheel-spinning from slow learning.
- Treat response latency and hint use as evidence about the response itself. A correct answer produced in under a second after three hints is weak evidence of knowledge and should be weighted as such before it updates anything.
- Give the model a way to say "this evidence is not about knowledge." Section 13's uncertainty discussion covers how much a model believes its estimate. Gaming is a different category, evidence that should not have entered the estimate at all, and conflating the two produces a model that is confidently wrong rather than appropriately unsure.
- Escalate rather than re-assign. Both behaviors are signals that the instructional loop has failed, and the correct response is a different action (a change of representation, an easier prerequisite, a human) rather than another item from the same pool.
16. Evaluation
A learner model should be evaluated at three layers.
Predictive validity
Does it predict future responses, completion, or requests for help?
Measurement validity
Does the state correspond to knowledge or another intended construct, rather than test-taking behavior or writing style?
Decision value
Do actions based on the model improve learning compared with a fixed or simpler policy?
The third layer is decisive. A highly accurate predictive model may not improve instruction. It may also be worse than a transparent rule once privacy, correction burden, and maintenance are included.
Useful outcomes include delayed retention, unaided transfer, calibration, completion, and subgroup error. Report uncertainty rather than only an average score.
17. Privacy and learner agency
Learner models can become sensitive profiles. Safe defaults include:
- explicit goals before inferred traits;
- local storage for progress where possible;
- no durable raw transcript unless needed and consented to;
- separation of identity from learning state;
- visible profile fields;
- correction, reset, export, and deletion;
- expiry for low-confidence inferences;
- no secondary use for advertising or unrelated scoring;
- no inference of protected or intimate traits as a shortcut to pedagogy.
The learner should be able to reject the model’s framing. A correction is not noise to be averaged away; it is high-authority evidence about how the system should represent the person.
18. Bottom line
Learner modeling is the bridge between a library and an adaptive learning system. Classical methods show the value of explicit uncertainty and skill structure. Deep models add flexible sequence prediction. Language models add free-text interpretation.
Each step increases expressive power and the risk of overclaiming. The practical standard is simple: model only what the learning decision needs, preserve the evidence and uncertainty, let the learner correct it, and prove that the resulting decisions improve learning.
Related concepts
Measurement: Bayesian Knowledge Tracing, Deep Knowledge Tracing, Item Response Theory, Computerized Adaptive Testing, Construct Validity
Products: Personalized AI Learning Systems, User Modeling in LLMs, Privacy-Preserving Personalization, Serendipity in Recommender Systems
Risks: Automation Bias, Algorithm Aversion, Distribution Shift, Apparent Personality from Text