Abstract. The graded response model (GRM) is Samejima’s ordered-polytomous Item Response Theory model for items whose response categories are ordered but not safely treated as interval measurements, such as Likert Items. Its core move is simple but powerful: model cumulative “at or above category” probabilities, then obtain each category-response curve by subtracting adjacent cumulative curves; this gives a principled account of thresholds, discrimination, information, sparse categories, and adaptive item selection.
Coverage note: verified through May 19, 2026.
The Graded Response Model: Polytomous IRT for Likert Items
1. Why the GRM exists
Likert-type self-report items are ordered: “strongly disagree” is below “disagree,” which is below “neutral,” and so on. They are not automatically interval-scaled: the psychological distance between “agree” and “strongly agree” need not equal the distance between “neutral” and “agree.” The graded response model, introduced by Fumiko Samejima in 1969, was designed for this exact class of ordered response data: responses that can be scored into ordered categories while still reflecting a continuous latent trait θ underneath the observed categories ([Samejima 1969: original GRM monograph][1]). Samejima explicitly described attitude measurement as a setting where a respondent’s continuous “intensity of positivity” toward a statement is expressed discretely, which is essentially the modern Likert-item situation. Psychometric Society
GRM belongs to the family of Polytomous IRT models: IRT models for items with more than two scored response categories. It differs from classical scoring by modeling the probability of each response category as a nonlinear function of θ. It also differs from many factor-analytic treatments of Likert items by taking the observed ordinal categories as the data-generating object, rather than treating category labels as if they were continuous measurements. Forero and Maydeu-Olivares summarize the motivation directly: rating scales are integral in personality and attitudinal measurement; factor analysis was originally linear and continuous; IRT models are nonlinear latent-trait models for categorical data; and Samejima’s GRM is “possibly” the most widely used IRT model for rating data. Universitat de Barcelona
The practical payoff is that a Likert item is no longer represented only by a mean, variance, and factor loading. It is represented by a discrimination parameter and a set of ordered thresholds that locate where respondents tend to move from lower to higher categories. The model therefore answers questions classical scoring cannot answer cleanly: which response categories are actually functioning, where the item is informative on the trait continuum, whether categories are too sparse or redundant, and which item should be administered next in Computerized Adaptive Testing.
2. The model: cumulative curves first, category curves second
Assume item iii has KKK ordered categories, conventionally indexed:
Yi∈{0,1,…,K−1}.Y_i \in {0, 1, \ldots, K-1}.Yi∈{0,1,…,K−1}. For a five-point Likert item, K=5K=5K=5, so the response categories might be coded 0,1,2,3,40,1,2,3,40,1,2,3,4. The GRM defines cumulative boundary probabilities:
Pik∗(θ)=P(Yi≥k∣θ),k=1,…,K−1.P^*_{ik}(\theta) = P(Y_i \ge k \mid \theta), \quad k=1,\ldots,K-1.Pik∗(θ)=P(Yi≥k∣θ),k=1,…,K−1. Boundary conventions complete the model:
Pi0∗(θ)=1,PiK∗(θ)=0.P^_{i0}(\theta)=1,\qquad P^_{iK}(\theta)=0.Pi0∗(θ)=1,PiK∗(θ)=0. The category-response probability is then the difference between adjacent cumulative probabilities:
P(Yi=k∣θ)=Pik∗(θ)−Pi,k+1∗(θ).P(Y_i=k\mid\theta)
P^_{ik}(\theta)-P^{i,k+1}(\theta).P(Yi=k∣θ)=Pik∗(θ)−Pi,k+1∗(θ). This is the defining idea of the GRM. Samejima’s notation gives the same construction: a graded response category has operating characteristic Px(θ)=Px∗(θ)−Px+1∗(θ)P_x(\theta)=P^_x(\theta)-P^{x+1}(\theta)Px(θ)=Px∗(θ)−Px+1∗(θ), with artificial boundary definitions at the bottom and top of the response scale. She also noted that these cumulative functions share the same discrimination but have different difficulty/category-bound parameters. Psychometric Society
The usual logistic GRM parameterization is:
Pik∗(θ)=11+exp[−Dai(θ−bik)],k=1,…,K−1.P^*_{ik}(\theta)
\frac{1}{1+\exp[-D a_i(\theta-b_{ik})]}, \quad k=1,\ldots,K-1.Pik∗(θ)=1+exp[−Dai(θ−bik)]1,k=1,…,K−1. Here aia_iai is the item discrimination parameter, bikb_{ik}bik is the kkk-th threshold or category-bound parameter, and DDD is a scaling constant, often 111 or approximately 1.71.71.7 depending on whether one wants the logistic curve to approximate the normal-ogive metric. Samejima’s monograph develops both normal-ogive and logistic versions; in the logistic version, the item discrimination parameter and category-bound difficulty parameters are explicitly identified, and the ordered threshold condition appears as bx+1>bxb_{x+1}>b_xbx+1>bx. Psychometric Society
Parameter structure
For an item with KKK ordered categories, the GRM has:
| Component | Count per item | Interpretation |
|---|---|---|
| Discrimination aia_iai | 1 | How sharply the item differentiates respondents around its thresholds. Higher aia_iai produces steeper cumulative curves and more information near the thresholds. |
| Thresholds bi1,…,bi,K−1b_{i1},\ldots,b_{i,K-1}bi1,…,bi,K−1 | K−1K-1K−1 | Trait locations where respondents cross cumulative category boundaries. In the logistic form, bikb_{ik}bik is the θ value where P(Yi≥k)=.5P(Y_i \ge k)=.5P(Yi≥k)=.5. |
| Total item parameters | KKK | A five-category item has one discrimination plus four thresholds. |
Thresholds are normally constrained to be ordered:
bi1<bi2<⋯<bi,K−1.b_{i1}<b_{i2}<\cdots<b_{i,K-1}.bi1<bi2<⋯<bi,K−1. A threshold is not a category mean. It is a boundary between lower and higher response regions. For a five-point item, bi1b_{i1}bi1 is the latent-trait point separating category 0 from categories 1–4 in cumulative probability terms, bi2b_{i2}bi2 separates categories 0–1 from categories 2–4, and so on.
Intermediate category curves are usually unimodal: they rise, peak, and fall. The two extreme categories are monotone: the lowest category becomes less probable as θ rises, and the highest category becomes more probable as θ rises. Samejima highlighted this consequence: if the cumulative item characteristic functions are monotone increasing, the operating characteristic for a non-extreme graded response is not monotone; it is a localized category curve. Psychometric Society
3. How to read a GRM item
A calibrated GRM item produces several diagnostic objects.
3.1 Cumulative probability curves
The cumulative curves show probabilities such as:
P(Yi≥1),P(Yi≥2),P(Yi≥3),P(Yi≥4).P(Y_i \ge 1),\quad P(Y_i \ge 2),\quad P(Y_i \ge 3),\quad P(Y_i \ge 4).P(Yi≥1),P(Yi≥2),P(Yi≥3),P(Yi≥4). For a Likert agreement item, these might correspond roughly to the probability of endorsing at least “disagree,” at least “neutral,” at least “agree,” and at least “strongly agree.” The higher the respondent’s θ on the measured trait, the more likely they are to exceed each threshold.
3.2 Category-response curves
The actual category curves are obtained by subtraction:
πi0(θ)=1−Pi1∗(θ)\pi_{i0}(\theta)=1-P^{i1}(\theta)πi0(θ)=1−Pi1∗(θ) πik(θ)=Pik∗(θ)−Pi,k+1∗(θ),1≤k≤K−2\pi{ik}(\theta)=P^{ik}(\theta)-P^{i,k+1}(\theta),\quad 1\le k\le K-2πik(θ)=Pik∗(θ)−Pi,k+1∗(θ),1≤k≤K−2 πi,K−1(θ)=Pi,K−1∗(θ).\pi{i,K-1}(\theta)=P^{i,K-1}(\theta).πi,K−1(θ)=Pi,K−1∗(θ). The category-response curve, or CRC, tells where a category is likely to be chosen. A healthy five-category Likert item will often show each response category dominating some region of θ. But this is not guaranteed. If two thresholds are too close, an intermediate category may never be the most probable response. That is not a numerical nuisance; it is psychometric information. It may imply that respondents do not distinguish the two adjacent verbal labels, that the item wording creates a jump in perceived intensity, or that the calibration sample lacks enough respondents in the relevant θ region.
3.3 Item information
GRM also gives an item information function:
Ii(θ)=∑k=0K−1[πik′(θ)]2πik(θ).I_i(\theta)
\sum_{k=0}^{K-1} \frac{\left[\pi'{ik}(\theta)\right]^2}{\pi{ik}(\theta)}.Ii(θ)=k=0∑K−1πik(θ)[πik′(θ)]2. Test information is the sum of item information functions under local independence:
Itest(θ)=∑iIi(θ).I_{\text{test}}(\theta)=\sum_i I_i(\theta).Itest(θ)=i∑Ii(θ). Information is local. A personality item may be highly informative for high conscientiousness but weak near the center; another may discriminate mainly among low-trait respondents. This locality is the engine of adaptive testing: once a provisional θ estimate is available, the next item can be selected to maximize information near that estimate. Modern IRT textbooks emphasize this model-based measurement logic, and Embretson and Reise’s Item Response Theory for Psychologists gives central coverage to polytomous IRT because many psychological tests use rating scales ([Embretson & Reise 2000: IRT for psychologists][5]). Google Books
4. Why not just treat Likert items as continuous?
Treating Likert items as continuous is often convenient, sometimes empirically adequate, and not automatically wrong. The strongest argument for GRM is not that all continuous treatments of Likert data are invalid. The argument is that item-level measurement, item-bank calibration, sparse category diagnosis, adaptive testing, and person-level precision all require information that continuous scoring discards.
Robitzsch’s 2020 critique is a useful counterweight: he argues that treating ordinal variables as continuous can often be defended, and he frames the choice between continuous and ordinal factor models as a contest between assumptions rather than a simple rule. He also notes the common methodological recommendation that ordinal methods become more important when there are few categories or skewed response distributions. Frontiers The GRM case is therefore strongest when the target is item-level measurement rather than merely estimating broad covariance structure.
| Problem | Continuous Likert treatment | GRM treatment |
|---|---|---|
| Ordered but non-interval categories | Assigns scores such as 1–5 and typically treats distances as equal. | Models ordered category probabilities without assuming equal psychological spacing. |
| Category functioning | A category can be unused, redundant, or never modal without being visible in a mean score. | Category-response curves show whether each category functions over some θ range. |
| Skewed response distributions | Skew can distort means, covariances, loadings, and linear assumptions. | Skew can be represented as threshold placement and trait distribution effects, though estimation can still degrade under extreme skew and small samples. |
| Local precision | Reliability is often summarized globally. | Information functions show precision as a function of θ. |
| Adaptive testing | A total-score framework has no native item-selection rule. | Items can be selected by Fisher information at the current θ estimate. |
| Missing or sparse categories | Usually handled by ad hoc recoding or ignored until estimation fails. | Threshold estimates, category curves, and fit diagnostics reveal whether categories should be collapsed, constrained, or recalibrated. |
The skewness point needs care. GRM is not magically immune to bad data. It models ordinal response probabilities directly, which is conceptually appropriate for skewed Likert items; but parameter estimation can be unstable when skewness is combined with small samples, low discrimination, and few items per dimension. Forero and Maydeu-Olivares found that conditions involving roughly 200 observations, few indicators per dimension, highly skewed items, or low factor loadings were among those to avoid, and convergence was near 100% only when sample size exceeded 500 in their studied conditions. Universitat de Barcelona
The same caution applies to missing intermediate categories. If a calibration sample contains no responses in category 2, no model can identify that category’s location from the data alone. What GRM provides is a principled failure mode: the analyst can see nonfunctioning categories, disordered or unstable thresholds, inflated standard errors, or a category curve that never peaks. The remedy may be collapsing adjacent categories, collecting more data in the relevant θ range, applying Bayesian priors or constraints, or redesigning the item. Pretending that the item is continuous hides the measurement problem rather than solving it.
5. Comparison with other ordered-polytomous IRT models
GRM is only one member of the polytomous IRT family. The most relevant alternatives for ordered response categories are the Partial Credit Model, Generalized Partial Credit Model, and Rating Scale Model.
| Model | Primary source | Probability construction | Discrimination | Step / threshold structure | Typical use |
|---|---|---|---|---|---|
| Graded Response Model (GRM) | Samejima 1969 | Cumulative logits/probits; category probabilities are adjacent differences. | Item-specific aia_iai. | Item-specific ordered thresholds. | Likert items, attitude/personality scales, item banks where discrimination varies. |
| Partial Credit Model (PCM) | Masters 1982 | Adjacent-category logits; probability of category depends on accumulated step difficulties. | Rasch-family equal discrimination. | Item-specific step parameters; number/structure can vary by item. | Partial-credit scoring, ordered performance tasks, Rasch measurement with item-specific steps. |
| Generalized Partial Credit Model (GPCM) | Muraki 1992 | Adjacent-category logits like PCM, but generalized. | Item-specific slope allowed. | Step parameter decomposed into location and threshold parameters. | Flexible alternative to PCM when item discriminations vary. |
| Rating Scale Model (RSM) | Andrich 1978 | Rasch-family rating formulation for ordered categories. | Equal discrimination. | Common rating-scale structure across items, plus item locations. | Instruments where the same response scale is intended to operate similarly across items. |
5.1 GRM versus PCM
Masters developed the Partial Credit Model as a Rasch-family model for responses scored in two or more ordered categories. The PCM preserves parameter separability and “specific objectivity,” and Masters explicitly distinguished its parameters from Samejima’s category boundaries. He also framed PCM as an extension of Andrich’s Rating Scale Model to cases where ordered response alternatives vary in number and structure from item to item. Cambridge University Press & Assessment
The conceptual difference is this:
-
GRM models cumulative transitions: “What is the probability of responding at least category kkk?”
-
PCM models adjacent steps: “What is the relative probability of category kkk versus category k−1k-1k−1?”
For Likert personality items, GRM often feels natural because agreement responses can be interpreted as crossing successive endorsement thresholds. For educational constructed-response items, PCM often feels natural because categories may represent accumulated scoring steps.
5.2 GRM versus GPCM
Muraki’s GPCM generalizes PCM by allowing a varying slope parameter. This is crucial because many real items do not discriminate equally. Muraki’s 1992 paper developed the GPCM, decomposed item step parameters into location and threshold components following Andrich, derived an EM algorithm, and found that the varying-slope GPCM fit real NAEP mathematics data better than the PCM. Sage Journals
GPCM and GRM often have the same broad parameter count: one discrimination-like parameter plus category/step parameters. The choice is therefore not just “flexible versus simple.” It is about the assumed response process. GRM says ordered categories arise from crossing cumulative thresholds. GPCM says adjacent category transitions are the natural object. In practice, both can fit rating scales reasonably well, and model selection should be empirical and theoretical.
5.3 GRM versus RSM
Andrich’s Rating Scale Model is stricter than PCM and GRM. It formulates ordered response categories within a Rasch framework, with subject and item parameters plus rating-scale category parameters. Andrich showed conditions under which successive integer scoring is justified, especially when threshold discriminations are equal; when threshold distances are also equal, an even simpler category-coefficient pattern follows. Cambridge University Press & Assessment
RSM is attractive when the same response labels are intended to mean the same thing across items. Many Likert questionnaires use the same labels repeatedly, which gives RSM surface appeal. But personality items vary in extremity, ambiguity, social desirability, and semantic breadth. The threshold structure for “I make friends easily” may not match the threshold structure for “I rarely feel blue,” even if both use the same five response labels. GRM relaxes that equality by allowing item-specific thresholds and item-specific discrimination.
6. Empirical findings: when GRM, PCM, or GPCM fits better
The literature does not support a universal rule that GRM always beats PCM/GPCM or that Rasch-family models are always preferable. It supports a conditional view.
Baker, Rounds, and Zevon compared Samejima’s logistic GRM with Masters’s PCM using 52 mood terms and 713 subjects in subjective well-being measurement. Their model-fit comparison favored the GRM for that data set, after unidimensionality was supported for the positive and negative affect dimensions. SUNY Research Connect This is directly relevant to personality-adjacent Likert data: affect and mood adjectives look more like self-report rating items than like scored achievement steps.
Muraki’s GPCM result points in the opposite direction from strict Rasch parsimony: when item slopes vary, a model that allows varying slopes may fit better than PCM. But that finding supports both GPCM and GRM over equal-discrimination Rasch-family models when discrimination heterogeneity is real. Sage Journals
Dai and colleagues’ 2021 simulation is especially useful because it directly compared GRM and GPCM under rating-scale conditions with short instruments, small samples, item-quality differences, and missingness. They found that for item parameter estimation, sample size of at least 300 and/or at least five items was advisable; GRM person estimates were more accurate when missingness was small, while GPCM was favored under large missingness; AIC, BIC, and log-likelihood were preferable to test information for model selection but could be weak with sample sizes below 300 and instruments shorter than five items. Frontiers Their conclusions add important detail: GPCM item parameter estimation was more stable, GRM often yielded more accurate θ estimates, missingness below 10% had small impact, high missingness increased requirements for sample size and length, and GRM could be a good choice when missingness was 20% or lower, while GPCM was recommended above 30% missingness because of stability. Frontiers
The empirical bottom line is:
| Empirical condition | Model implication |
|---|---|
| Item discriminations differ substantially | PCM/RSM may be too restrictive; GRM or GPCM is usually more plausible. |
| Category process is cumulative endorsement | GRM has a strong substantive rationale. |
| Category process is adjacent step scoring | PCM/GPCM has a strong substantive rationale. |
| Same rating scale behaves similarly across all items | RSM may be efficient and interpretable. |
| Small calibration samples | Rasch-family parsimony can help, but fit may suffer if equal discrimination is false. |
| High missingness | GPCM may be more stable than GRM in some simulation conditions. |
| Short instruments | Model-selection indices become weaker; item-parameter recovery is fragile. |
| CAT item banks | GRM is attractive because item information and thresholds support adaptive selection, but MIRT may outperform unidimensional per-facet CAT when facets are correlated. |
7. GRM in computerized adaptive personality measurement
GRM became especially important for Computerized Adaptive Testing in personality because personality inventories are long, multidimensional, and heavily Likert-based. A full Big Five or HEXACO facet inventory can easily contain hundreds of items. CAT promises shorter administration by selecting only the items that are most informative for the current respondent.
The standard unidimensional personality-CAT pipeline is:
-
Define a facet-level construct, such as anxiety, orderliness, or openness to ideas.
-
Write or assemble a large item pool for that facet.
-
Test unidimensionality and local independence for the facet.
-
Calibrate items under GRM or another ordered-polytomous IRT model.
-
Start the CAT with an initial θ estimate, often 0.
-
Select the item maximizing information at the current θ estimate.
-
Update θ after each response.
-
Stop after a fixed number of items, target standard error, or exposure/content constraint.
This is exactly the pattern used in Big Five CAT work. Nieto and colleagues built a 480-item pool for Five Factor Model facets, administered it to 826 participants, calibrated facets separately, and used item selection mindful of facet unidimensionality. The final pool had 360 items, and simulations showed that four items per facet, 120 items total across 30 facets, could provide accurate facet scores while preserving the FFM factor structure. Psicothema Their analysis fitted the unidimensional Samejima GRM to each subset of items measuring the same facet, tested unidimensionality with parallel analysis and factor models, calibrated selected unidimensional item subsets separately, and used maximum Fisher information plus MAP θ updating in the simulated CAT. Psicothema In the results, only 4 of 360 items misfit the GRM by their criterion, and CAT facet scores correlated highly with pool scores, ranging from .92 to .98. Psicothema
Earlier NEO PI-R adaptive work reached a similar practical result. Reise and Henson used real-data simulations with N=1,059N=1{,}059N=1,059 to evaluate a maximum-information CAT algorithm for the NEO PI-R; they found satisfactory recovery of full-scale facet scores with around four items per facet and concluded that the NEO PI-R could be reduced by half with little precision loss. They also reported an important caveat: for many scales, simply administering the “best” four items per facet would have produced similar results, so CAT’s added value was not automatic. sage.cnpereading.com
This caveat matters. CAT is not magic compression. CAT helps when the item bank contains items that are informative in different θ regions and when the algorithm can exploit that variation. If every good item is informative near the same region, a fixed short form may compete well.
8. Multidimensional CAT and the unidimensionality problem
The major critique of per-facet GRM personality CAT is that facets are rarely independent psychological atoms. Facets within a domain are correlated; facets across domains can cross-load; response styles can induce common variance; item wording can create method factors. A unidimensional GRM per facet is therefore a useful approximation, not a metaphysical claim.
Nieto and colleagues took unidimensionality seriously: they tested each facet, removed items until assumptions were met, and then calibrated each facet separately. Yet their later ESEM analysis still found some cross-loadings and theoretically meaningful correlated residuals, and they explicitly identified multidimensional IRT and multidimensional CAT as a direction for future work. Psicothema
Makransky, Mortensen, and Glas directly pursued that route for the NEO PI-R. They argued that personality facets are commonly reported and used in applied settings, but are often correlated; scoring one facet at a time can be inefficient. Their multidimensional CAT approach exploited correlations among facets, and they found that the NEO PI-R could be substantially shortened without attenuating precision, especially for constructs with many highly correlated facets. Sage Journals
This frames a central open question for personality measurement: is the default model a unidimensional GRM per facet, or a Multidimensional IRT model over correlated facets? The answer is still context-dependent. Per-facet GRM is simpler, easier to calibrate, easier to explain, and compatible with existing facet-score reporting. MIRT can be more statistically efficient and more realistic, but it requires larger calibration samples, stronger modeling choices, and more complex item-selection and score-reporting logic.
9. Psyche’s adoption of GRM for adaptive personality measurement
Psyche is an open-source framework for constructing personality profiles by combining validated psychometric instruments with LLM-based text analysis. Its public repository describes a tiered instrument battery, including Big Five IPIP-NEO variants, HEXACO, and many auxiliary constructs. In the heavy tier, the repository lists “CAT Big Five + HEXACO” as approximately 40–60 adaptive items with “GRM-based adaptive refinement.” GitHub
This is a meaningful adoption signal, not a peer-reviewed validation result. Psyche’s choice is methodologically coherent: Big Five and HEXACO facets are ordered Likert domains, and GRM is a natural model for adaptive refinement of such item banks. But the public repository claim does not by itself establish calibration quality, item-bank representativeness, measurement invariance, or CAT score validity. For an adaptive personality system, the relevant validation evidence would include item calibration sample size, θ coverage, category functioning, item exposure control, convergent validity, test-retest reliability, fairness/DIF checks, and comparison with fixed-form baselines.
Psyche is nevertheless a useful example of GRM moving from academic psychometrics into AI-era personalization infrastructure. Once personality profiles are used to generate behavioral context for AI assistants, measurement errors become downstream system errors. A GRM-based CAT can reduce some measurement inefficiency, but it cannot remove the need for transparent validation.
10. Advantages of GRM for Likert personality items
10.1 Proper category-response modeling
The central advantage is that the model represents the response process at the category level. A five-point item is not compressed into a single linear score before modeling. The model estimates four thresholds and one discrimination. This lets analysts inspect whether “neutral” is functioning, whether “strongly agree” is too rare, or whether two adjacent categories are psychometrically redundant.
10.2 Local precision rather than global reliability
Classical reliability coefficients summarize scale precision globally. GRM gives precision conditional on θ. This is valuable in personality assessment because applied decisions often care about tails: very high neuroticism, very low conscientiousness, unusually high openness, or extreme honesty-humility. A scale can be reliable on average but weak in the region where a specific decision is being made.
10.3 Robustness to nonnormal and skewed response patterns, with limits
Likert personality items are often skewed. Most people may endorse socially desirable items positively, reject pathological items, or cluster around moderate agreement. GRM can represent such skew through threshold locations and category probabilities rather than forcing item responses into a linear normal model. But simulation evidence shows that extreme skew, small samples, low item slopes, and few indicators can destabilize estimation. Forero and Maydeu-Olivares explicitly identify highly skewed items and small samples as problematic conditions, and Dai et al. show that missingness and short instruments increase parameter-estimation error for both GRM and GPCM. Universitat de Barcelona
10.4 Principled treatment of sparse or missing intermediate categories
GRM does not require every category to be equally common. It can model categories that are rare because their thresholds occupy narrow or extreme regions of θ. It also exposes categories that are not functioning. If a middle category never becomes modal, that is a diagnostic result. The correct response may be collapsing categories, rewriting labels, expanding the calibration sample, or fitting a model with constraints or priors. This is more principled than silently assigning equally spaced numeric values and hoping the total score behaves.
10.5 CAT-ready item information
GRM gives item information by θ. That makes it directly usable in CAT. Nieto et al.’s Big Five CAT simulation selected the first item by maximizing Fisher information at θ = 0, updated θ by MAP after responses, and repeatedly selected the next item maximizing information at the current estimate. Psicothema This information-based loop is exactly why GRM is more than a scoring model: it is an administration model.
11. Active critiques and failure modes
11.1 Parameter estimation stability with small samples
GRM has more parameters than a simple continuous item model and more freedom than Rasch-family models. That flexibility is useful only when the data support it. With small NNN, short scales, weak discrimination, sparse categories, or skewed response distributions, threshold and discrimination estimates can become unstable.
Forero and Maydeu-Olivares examined 324 estimation conditions for Samejima’s GRM. They found that both full-information and limited-information methods could fail under harsh conditions, especially small samples, few indicators per dimension, high skewness, and low item slopes. Universitat de Barcelona Dai et al. likewise recommend caution below N=300N=300N=300, especially with short instruments, missingness, and poor item quality; they note that model-fit indices are less helpful when sample size is below 300 and scale length is five items or fewer. Frontiers
11.2 Unidimensionality per facet
GRM is usually fitted as a unidimensional model. That means each item response is explained by one θ plus item parameters, and responses are locally independent conditional on θ. Personality facets often violate this ideal. An item can reflect anxiety and depression, orderliness and industriousness, agreeableness and social desirability, or a general evaluative response style.
Per-facet unidimensionality is therefore an empirical requirement, not a naming convention. Nieto et al. had to test and iteratively enforce unidimensionality before calibration; their results still showed cross-loadings at the broader factor-structure level. Psicothema This does not invalidate GRM, but it narrows its interpretation: θ is “the dominant latent dimension captured by this calibrated item set,” not necessarily a pure psychological essence.
11.3 Response styles
Likert personality data are vulnerable to acquiescence, extreme responding, midpoint preference, social desirability, and faking. GRM can reveal category usage patterns, but a unidimensional GRM does not automatically separate trait variance from response-style variance. A respondent who uses extreme categories across all items may look high or low on traits depending on keying. Mixed models, multidimensional models, forced-choice designs, or explicit response-style factors may be needed.
11.4 Model fit versus measurement philosophy
Rasch-family models such as PCM and RSM impose stronger constraints and can support stronger invariance claims if they fit. GRM often fits better because it is more flexible, especially through item-specific discrimination. But better likelihood fit is not always better measurement. In high-stakes or fairness-sensitive contexts, the invariance and interpretability of Rasch-family models may be worth the loss of flexibility. The correct comparison is not “which model has more parameters?” but “which model’s assumptions are defensible for this construct, population, and use?”
11.5 Sparse categories and overconfident CAT
CAT systems can become overconfident if item parameters are poorly estimated or if the item bank lacks coverage in parts of θ. A GRM item bank should not be judged only by average reliability. It should be inspected for information coverage, threshold spread, item exposure, category functioning, and subgroup invariance. Otherwise, adaptive selection can efficiently administer the wrong items.
12. Is GRM still the default for personality measurement?
For Likert personality item banks, GRM remains a strong default because it matches the ordered-threshold structure of self-report categories, supports item-level diagnostics, and plugs directly into CAT. The literature shows repeated use in personality and affective measurement, including NEO PI-R adaptive simulations, Big Five facet item-bank calibration, and entrepreneurial personality CAT development. Postigo and colleagues, for example, created a 120-item entrepreneurial personality bank, calibrated it with Samejima’s GRM in a sample of 1,170 participants, and reported an essentially unidimensional fit plus high CAT accuracy over a wide θ range with a mean of 16 administered items. Springer Link
But GRM is not the inevitable endpoint. Three alternatives are serious candidates to displace it in some settings.
First, Rasch-family models may dominate when invariant comparison, construct maps, and sample-efficient calibration are more important than flexible fit. PCM and RSM give up item-specific discrimination, but that restriction can be a virtue if the goal is stable measurement with strong comparability.
Second, GPCM may be preferable when adjacent-category transitions are the right substantive process or when missingness patterns make GPCM more stable. Dai et al.’s simulation suggests GPCM can be more stable for item parameters and under high missingness, while GRM can produce more accurate θ estimates under lower missingness. Frontiers
Third, MIRT and multidimensional CAT may be better suited to personality’s correlated-facet structure. Makransky et al. show the appeal of exploiting facet correlations, and Nieto et al. explicitly identify multidimensional calibration and adaptive administration as future work. Sage Journals
The open question is not whether GRM is “correct.” The open question is whether the default personality-measurement architecture should remain “one unidimensional GRM per facet,” or shift toward mixed Rasch/GRM/MIRT systems: Rasch constraints where they fit, GRM flexibility where thresholds and slopes vary, multidimensional borrowing where facets are correlated, and response-style factors where Likert behavior is not trait-pure.
13. Practical guidance
| Use GRM when… | Prefer an alternative when… |
|---|---|
| Items have ordered categories and category thresholds are substantively meaningful. | Categories are nominal rather than ordered. |
| Item discriminations plausibly differ. | Equal discrimination is theoretically required or empirically adequate. |
| You need category-response curves and local information. | You only need rough group-level covariance estimates and categories are many/symmetric. |
| You are building a CAT item bank. | You are building a small fixed scale with insufficient calibration data. |
| The construct can be treated as unidimensional within the calibrated item set. | Facets are highly correlated and multidimensional scoring can exploit that structure. |
| Sparse categories need diagnosis rather than concealment. | Categories are so sparse that thresholds cannot be estimated without collapsing or stronger priors. |
| Missingness is low or moderate and the scoring model handles observed patterns appropriately. | Missingness is high, nonignorable, or better handled by a model shown to be more stable for that design. |
A good GRM analysis should report at least:
| Diagnostic | Why it matters |
|---|---|
| Category counts per item | Detects empty or sparse categories. |
| Threshold estimates and standard errors | Shows category-bound locations and instability. |
| Discrimination estimates | Identifies weak or overly dominant items. |
| Category-response curves | Reveals nonfunctioning categories. |
| Item and test information functions | Shows where the scale is precise. |
| Item fit statistics | Detects model-data mismatch. |
| Unidimensionality evidence | Supports the interpretation of θ. |
| Local dependence checks | Detects item clusters not explained by θ. |
| DIF / invariance checks | Tests whether item parameters generalize across groups. |
| CAT simulation results | Evaluates stopping rules, exposure, precision, and score recovery. |
14. References
[1] Samejima 1969 — original GRM monograph. Samejima, F. Estimation of Latent Ability Using a Response Pattern of Graded Scores. Psychometrika Monograph Supplement, 1969. Psychometric Society
[2] Masters 1982 — Partial Credit Model. Masters, G. N. “A Rasch Model for Partial Credit Scoring.” Psychometrika, 1982. Cambridge University Press & Assessment
[3] Muraki 1992 — Generalized Partial Credit Model. Muraki, E. “A Generalized Partial Credit Model: Application of an EM Algorithm.” Applied Psychological Measurement, 1992. Sage Journals
[4] Andrich 1978 — Rating Scale Model. Andrich, D. “A Rating Formulation for Ordered Response Categories.” Psychometrika, 1978. Cambridge University Press & Assessment
[5] Embretson & Reise 2000 — IRT for psychologists. Embretson, S. E., & Reise, S. P. Item Response Theory for Psychologists. Lawrence Erlbaum / Psychology Press, 2000. Google Books
[6] Baker, Rounds & Zevon 2000 — GRM versus PCM in subjective well-being. Baker, J. G., Rounds, J. B., & Zevon, M. A. “A Comparison of Graded Response and Rasch Partial Credit Models with Subjective Well-Being.” Journal of Educational and Behavioral Statistics, 2000. SUNY Research Connect
[7] Forero & Maydeu-Olivares 2009 — GRM estimation conditions. Forero, C. G., & Maydeu-Olivares, A. “Estimation of IRT Graded Response Models: Limited Versus Full Information Methods.” Universitat de Barcelona
[8] Dai et al. 2021 — GRM versus GPCM under rating-scale design conditions. Dai, S., Vo, T. T., Kehinde, O. J., He, H., Xue, Y., Demir, C., & Wang, X. “Performance of Polytomous IRT Models With Rating Scale Data.” Frontiers in Education, 2021. Frontiers
[9] Reise & Henson 2000 — adaptive NEO PI-R. Reise, S. P., & Henson, J. M. “Computerization and Adaptive Administration of the NEO PI-R.” Assessment, 2000. sage.cnpereading.com
[10] Makransky, Mortensen & Glas 2013 — multidimensional CAT for NEO PI-R facets. Makransky, G., Mortensen, E. L., & Glas, C. A. W. “Improving Personality Facet Scores With Multidimensional Computer Adaptive Testing.” Assessment, 2013. Sage Journals
[11] Nieto et al. 2017 — Big Five item pool calibrated under GRM. Nieto, M. D., Abad, F. J., Hernández-Camacho, A., Garrido, L. E., Barrada, J. R., Aguado, D., & Olea, J. “Calibrating a New Item Pool to Adaptively Assess the Big Five.” Psicothema, 2017. Psicothema+2Psicothema+2
[12] Postigo et al. 2020 — entrepreneurial personality CAT under Samejima GRM. Postigo, Á., Cuesta, M., Pedrosa, I., Muñiz, J., & García-Cueto, E. “Development of a Computerized Adaptive Test to Assess Entrepreneurial Personality.” Psicologia: Reflexão e Crítica, 2020. Springer Link
[13] Psyche repository — GRM-based adaptive refinement in an AI personality framework. AshitaOrbis/psyche GitHub repository. GitHub
[14] Robitzsch 2020 — counterargument on treating ordinal variables as continuous. Robitzsch, A. “Why Ordinal Variables Can (Almost) Always Be Treated as Continuous Variables.” Frontiers in Education, 2020. Frontiers
Companion entries
Core theory: Item Response Theory, Polytomous IRT, Latent Trait Models, Ordinal Measurement, Likert Items, Category-Response Curves, Item Information Functions
Model family: Graded Response Model, Partial Credit Model, Generalized Partial Credit Model, Rating Scale Model, Rasch Models, Normal-Ogive IRT, Logistic IRT
Personality measurement: Big Five Personality Measurement, HEXACO Measurement, Personality Facets, Self-Report Psychometrics, Response Styles, Measurement Invariance
Adaptive systems: Computerized Adaptive Testing, Adaptive Personality Assessment, Item Bank Calibration, Fisher Information Item Selection, MAP Scoring in IRT, Psyche
Diagnostics and critique: Unidimensionality, Local Independence, Differential Item Functioning, Sparse Likert Categories, Small-Sample IRT Estimation, Multidimensional IRT, Mixed Rasch Models