Reference

Sample-Efficient Personality Inference

Sample-efficient personality inference is not the search for the fewest possible questions or tokens; it is the search for the least evidence that preserves the validity needed for a specific decision. The core pattern is stable across questionnaires, social-media models, and LLM-based inference: broad Big Five factors can be recovered from surprisingly little signal, while facets, individual-level interpretation, clinical decisions, and stable language-based inference require substantially more evidence.

Coverage note: verified through May 19, 2026.

Sample-Efficient Personality Inference: How Few Items Tell Enough

1. The central distinction: “enough” for what?

The useful question is not “How few personality items work?” but “How much evidence is enough for this output, in this context, at this consequence level?” A 10-item inventory can be enough to include Big Five Personality Measurement as a rough covariate in a survey. It is not enough to estimate facet profiles, guide clinical judgment, or support consequential individual decisions.

Personality measurement has an asymmetry: broad factors aggregate many correlated behavioral tendencies, so they tolerate sparse indicators better than narrow facets do. Facets require discriminating among nearby subtraits inside the same domain, such as anxiety versus vulnerability under Neuroticism, or orderliness versus self-discipline under Conscientiousness. The Big Five is normally treated as a hierarchical model in which broad domains summarize more specific facets and trait nuances, not as five indivisible atoms. Gosling, Rentfrow, and Swann explicitly frame Big Five measurement this way when discussing very brief measures, contrasting domain-level instruments with longer facet-capable inventories such as the NEO PI-R. Gosling

This produces the main rule of Sample-Efficient Psychometrics:

Short forms preserve broad-factor rank ordering better than they preserve facet structure, individual diagnostic fidelity, or rare-pattern interpretation.

That rule applies to self-report items and to text. Ten to sixty items can often recover broad domains. One hundred twenty items can begin to support thirty-facet research use. Three hundred items are still useful when facet reliability, content coverage, and individual interpretation matter. Text models behave similarly: a few snippets may contain trait signal, but stable personality inference from language requires enough text diversity to average over topic, mood, platform, and context.

2. Measurement tiers at a glance

Instrument or tier Item count Main output What it preserves What it loses Best use
TIPI, Ten-Item Personality Inventory 10 Five broad Big Five domains Extremely fast domain-level signal Low reliability, weak facet coverage, reduced convergent validity When personality is a minor covariate and time is severe
BFI-10 and similar ultra-short forms 10 Five broad domains Rough domain ranking in large surveys Lower effect sizes than full BFI; weak individual interpretation Population surveys with “truly limited” time
Mini-IPIP 20 Five broad domains Better item coverage than two-item-per-domain measures No real facet scoring; limited item redundancy Research screening, broad covariate use
IPIP-NEO-60 60 Broad domains with equal facet representation in item selection Domain scores with better coverage of all thirty NEO-like facets Not a full facet inventory; only two items per facet area Domain-level research where coverage matters
IPIP-NEO-120 120 Five domains plus thirty four-item facets Research-grade facet approximation Lower facet reliability than 300-item form; not ideal for high-stakes individual decisions Facet-level research and richer personal feedback
IPIP-NEO-300 300 Five domains plus thirty ten-item facets Maximum public-domain IPIP-NEO coverage and facet reliability Time, fatigue, and administration burden Detailed profiling, validation studies, clinical-adjacent research with safeguards
CAT over calibrated item bank Variable Adaptive domain/facet estimates Potentially long-form precision with fewer administered items Requires validated item parameters, stopping rules, population calibration Efficient high-fidelity systems, not simple item deletion

The empirical literature supports the ladder rather than a single universal “minimum.” Gosling et al. developed the TIPI for situations where very brief measures are needed, personality is not the primary topic, or researchers can tolerate diminished psychometric properties. Donnellan et al.’s Mini-IPIP gives four items per Big Five trait and was validated as a 20-item short form. Johnson’s IPIP-NEO-120 was explicitly built because shorter IPIP measures with 20, 50, or 100 items could not measure six facets within each of the five domains. Maples-Keller et al.’s IPIP-NEO-60 used item response theory to create a 60-item domain-level measure with equal representation of the thirty NEO PI-R facet areas. datashare.ed.ac.uk+3Gosling+3storkapp.me+3

3. The TIPI: bandwidth at the edge of measurement

The Ten-Item Personality Inventory is the canonical example of extreme bandwidth-for-fidelity exchange. It measures each Big Five domain with two items. Its virtue is not that it is psychometrically equivalent to longer inventories. Its virtue is that it gives researchers a usable Big Five proxy when the alternative is often no personality measure at all.

Gosling et al. state the intended use case directly: the TIPI is appropriate when “very short measures are needed,” personality is not the primary topic of interest, or researchers can tolerate the diminished psychometric properties associated with very brief instruments. They also emphasize that a 10-item instrument is psychometrically superior to a 5-item instrument because it allows latent-variable modeling and can include more than one indicator per domain. Gosling

The losses are explicit. The TIPI is less reliable than standard multi-item instruments, has lower correlations with longer measures, and cannot measure facets. Gosling et al. note that researchers who need facet scores should use a longer instrument, such as the 240-item NEO PI-R; they also point out that even the 44-item BFI and 60-item NEO-FFI do not provide facet scores. Gosling

The TIPI therefore answers one narrow question well: “Can I include a rough Big Five control variable in a study where time is extremely scarce?” It does not answer “What is this person’s personality structure?” or “Which facet-level mechanism explains this behavior?”

A useful way to state the TIPI tradeoff:

What the TIPI can support What the TIPI should not support
Large-sample covariate adjustment Facet interpretation
Quick descriptive screening Clinical or employment decisions
Exploratory domain-level associations Personalized psychological explanation
Situations where no personality measure would otherwise be collected Trait nuance, within-domain diagnosis, or intervention planning

The BFI-10 literature reaches a similar conclusion from a different instrument family. Rammstedt and John developed the 10-item BFI for contexts in which participant time is “severely limited,” and found that reducing the BFI-44 to less than a fourth of its length produced lower effect sizes but still sufficient validity for research settings with truly limited time constraints. Southeastern Oklahoma State University

4. The Mini-IPIP: the smallest serious broad-factor option

The Mini-IPIP occupies a more defensible broad-factor tier than two-item-per-domain instruments. It has twenty items, four for each Big Five domain. The official IPIP scoring key reports four items per factor and alphas of .77 for Extraversion, .70 for Agreeableness, .69 for Conscientiousness, .68 for Neuroticism, and .65 for Intellect/Imagination in the listed key. IPIP

Donnellan, Oswald, Baird, and Lucas developed and validated the Mini-IPIP across five studies. The published abstract describes the instrument as a 20-item short form of the 50-item IPIP Five-Factor Model measure, with four items per Big Five trait, acceptable internal consistencies at or above roughly .60, similar broad facet coverage to other broad Big Five measures, and convergent, discriminant, criterion-related, and test-retest evidence comparable to the parent measure. storkapp.me

The Mini-IPIP is therefore a good answer when the research question is broad and the cost of extra items matters. It is not a substitute for an IPIP-NEO facet inventory. Four items per domain can stabilize a domain score better than two items, but it cannot represent six facets per domain in a psychometrically serious way.

A concise methodological interpretation:

Constraint TIPI Mini-IPIP
Time is nearly zero Strong fit Good but longer
Need rough Big Five covariate Acceptable Better
Need rank ordering among people Weak-to-moderate Moderate
Need facets No No
Need individual feedback Thin Still thin
Need research screening Only if severe time pressure Plausible

5. The IPIP-NEO ladder: 60, 120, 300

The IPIP-NEO family is the relevant progression for systems that want to move from broad screening to facet-aware inference. The 300-item IPIP-NEO was designed to measure constructs similar to the 30 NEO PI-R facet scales. Like the NEO PI-R, it yields five broad domains and six narrower facets per domain. Johnson’s 2014 IPIP-NEO-120 paper describes the 300-item IPIP-NEO as a public-domain inventory that can score both the five domains and thirty facets. BPB

5.1 IPIP-NEO-300: full public-domain facet coverage

The 300-item version gives roughly ten items per facet. This matters because facet scores are narrow. Narrow constructs need item redundancy to reduce random error, preserve content breadth, and separate adjacent constructs. The 300-item form is costly, but it is the version closest to a full public-domain NEO-like profile.

What the 300-item tier buys:

Psychometric target Why 300 helps
Facet reliability Ten items per facet gives more redundancy than two or four
Content validity Each facet can sample multiple phrasings and behavioral expressions
Within-domain discrimination Better separation among correlated facets
Individual feedback More defensible than short forms, though still self-report
CAT calibration Larger item pool enables adaptive selection if item parameters are validated

5.2 IPIP-NEO-120: research-grade facet compression

Johnson developed the IPIP-NEO-120 because existing 20-, 50-, and 100-item IPIP inventories could not measure the six facets within each Big Five domain. The 120-item version uses four items per facet and was developed from an Internet sample of 21,588 people, then tested in community and large Internet samples. Johnson reports that the 120-item version compares favorably to the longer form. BPB

The important caveat is also in Johnson’s paper. Four-item facet alphas are lower than ten-item facet alphas: in the reported samples, the 120-item facet alphas were lower than the 300-item facet alphas, and the large Internet sample had facet alphas ranging from .63 to .88, with all but three facets at .69 or greater. Johnson states that the IPIP-NEO-120 facet scales have sufficient reliability for research studies but probably should not be used to make important decisions about individuals. BPB

That sentence should be treated as a design constraint for any AI system using IPIP-NEO-120. The 120-item tier can support facet-level hypotheses, group-level analyses, and richer self-reflection. It should not be framed as a high-stakes individual diagnostic instrument.

5.3 IPIP-NEO-60: broad domains with facet-aware item selection

The IPIP-NEO-60 is not merely “shorter IPIP-NEO-120.” It is an IRT-selected 60-item representation intended to measure the broad domains while giving equal representation to the thirty NEO PI-R facet areas. Maples-Keller et al. argue that the NEO-FFI lacks equal facet representation; some facets receive no items while others are overrepresented. Their goal was to build a 60-item domain-level measure with equal coverage of the thirty facets. datashare.ed.ac.uk

The IPIP-NEO-60 demonstrated strong domain-level internal consistency: across three samples, the domain scores had mean coefficient alpha of .80 and mean inter-item correlation of .26. Maples-Keller et al. also report strong convergent validity and similar nomological networks between the IPIP-NEO-60 and the NEO-FFI, with an ICC of .96 across external criteria in their study. datashare.ed.ac.uk

The implication is subtle but important: IPIP-NEO-60 is a strong domain instrument with better content balance than many short forms. It is not a thirty-facet inventory in the same sense as IPIP-NEO-120 or IPIP-NEO-300. Two items per facet area can help domain coverage; they cannot support high-confidence facet scoring.

6. What is lost at each tier?

Short forms do not lose “personality.” They lose precision, separability, and interpretive resolution.

Tier Main retained signal Main loss Typical failure mode
10 items Rough broad-factor direction Reliability and content breadth Overinterpreting noisy domain scores
20 items Better broad-factor recovery Facet structure Treating four domain items as if they imply subtraits
60 items Stronger domain scores with balanced facet coverage Facet precision Reporting facet-like explanations without enough facet items
120 items Research-use facets High-stakes individual reliability Treating research-grade facet scores as clinical-grade
300 items Fullest IPIP-NEO facet coverage Time efficiency Fatigue, careless responding, completion burden
CAT Potential efficiency Calibration simplicity Mistaking adaptive delivery for validated adaptive measurement

The bandwidth-fidelity tradeoff in Personality Assessment is not a slogan; it is the structure of the measurement problem. A broad domain like Extraversion aggregates sociability, assertiveness, energy, positive affect, and related tendencies. A short form can estimate the aggregate if its few items are well chosen. But if the question is whether someone is high in assertiveness but low in gregariousness, broad Extraversion is no longer enough.

Facet-level predictive power is a major reason to avoid overcompressing. Maples-Keller et al. summarize evidence that FFM facets provide coverage of specific traits, distinct patterns of external correlates, and unique predictive power. That is exactly the information sacrificed when item counts drop below facet-capable lengths. datashare.ed.ac.uk

7. The “g” of personality: recoverable signal, contested construct

The phrase “the g of personality” can mean two different things. In engineering contexts, it often means a general evaluative or broad-factor signal: the easiest-to-recover component of personality-like data. In psychometric theory, the General Factor of Personality is a contested construct. Van der Linden and colleagues summarize the state of the debate: many studies and meta-analyses find that personality traits correlate such that a general factor emerges, but there is ongoing debate about whether it reflects a substantive social-effectiveness factor or measurement and response bias. Cambridge University Press & Assessment

For sample-efficient inference, the safe claim is this:

General and broad-domain signals are easier to recover than narrow facets, but that does not prove a psychologically deep general factor.

This matters for AI systems. An LLM may quickly infer that a writer seems socially confident, emotionally negative, or intellectually exploratory. That does not mean it has measured a stable facet structure. Broad impression formation is not equivalent to calibrated psychometrics.

8. Text as a measurement channel

Text inference changes the evidence unit. A questionnaire item is designed to target a construct. A text token is not. Most words express topic, task, audience, genre, mood, and local context before they express stable personality. That is why minimum text quantity is not directly comparable to minimum item count.

8.1 Yarkoni 2010: why large text samples mattered

Yarkoni’s “Personality in 100,000 Words” is still the clearest methodological warning. Earlier language-and-personality studies often relied on no more than a few thousand words per participant, which limited reliable estimation of individual word usage rates and mostly supported aggregate category analysis. Yarkoni used nearly 700 blogs, averaging 115,423 words per person over an average span of 23.9 months, enabling analysis at both category and individual-word levels. Academia

The lesson is not that every text-based personality system needs 100,000 words. The lesson is that the required quantity depends on the resolution of the claim. Broad category-level inference may work with far less. Word-level, facet-level, and temporally stable inference need much more.

8.2 Social media open-vocabulary studies

Schwartz et al. scaled the open-vocabulary approach to Facebook status updates. They analyzed hundreds of millions of word, phrase, and topic instances from 75,000 volunteers who also took personality tests. Their dataset required users to have written at least 1,000 words, and the analyzed sample averaged 4,129 words across 206 status updates. PLOS

This gives a useful lower anchor: around 1,000 words was an inclusion threshold for large-scale open-vocabulary social-media analysis, not a guarantee of stable individual personality inference. The average usable participant contributed roughly four thousand words, and the study still relied on massive sample size across people.

8.3 Plank & Hovy, Twitter, and the Reddit correction

Plank and Hovy’s 2015 paper is sometimes grouped with social-media personality inference, but it is not a subreddit-level study. It is an ACL paper titled “Personality Traits on Twitter—or—How to Get 1,500 Personality Tests in a Week,” published in WASSA 2015. ACL Anthology

For Reddit-scale personality inference, the more direct reference is PANDORA. Gjurković et al. introduced PANDORA as a dataset of Reddit comments from 10,000 users partially labeled with three personality models and demographics, including 1,600 users labeled with the Big Five. ACL Anthology

The methodological warning is that subreddit membership is not personality. A subreddit can provide topic and community priors, but individual-level inference still needs user-specific language and validation against personality measures. Community-level aggregation can reveal norms of a forum; it cannot safely assign a stable trait profile to a person merely because they post in a forum.

9. LLM-based inference: fewer features, not less evidence

Recent LLM results show that foundation models can infer broad personality traits from text without task-specific supervised training. They do not show that a few snippets are enough for stable personality assessment.

Peters and Matz tested GPT-3.5 and GPT-4 on 1,000 MyPersonality Facebook users who had completed a 100-item IPIP questionnaire and had at least 200 status updates. The models used the most recent 200 updates, averaging 17.10 words each, processed in chunks of 20 messages, and averaged inferred scores across chunks and repeated queries. That is roughly 3,420 words per person at the full 200-status setting, not a tiny prompt. OUP Academic

Their average correlations with self-reported traits were r = .27 for GPT-3.5 and r = .31 for GPT-4. They also found that more messages improved accuracy, though a substantial share of variance was captured with as few as 20 status messages. OUP Academic+2OUP Academic+2

A separate free-form interaction study by Peters, Cerf, and Matz found that GPT-4-based chatbots inferred Big Five traits with higher accuracy when explicitly prompted to elicit personality-relevant information: mean r = .443 in the assessment condition, compared with mean r = .218 in a more naturalistic condition and mean r = .117 when acting like a default helpful assistant. arXiv

Frontiers 2025 work on ChatGPT 4 reinforces the caution. Piastra and Catellani used essays and Twitter-message datasets, found performance varied by trait and dataset, and showed that when the amount of Twitter text was gradually reduced, correlations between ChatGPT-estimated and self-reported scores tended toward zero while model confidence remained relatively high. Frontiers+2Frontiers+2

The current best interpretation is:

Claim Evidence status
LLMs can extract broad personality signal from text Supported, especially for broad domains
LLMs can do this zero-shot Supported in some datasets, with moderate correlations
A few snippets are enough for stable individual profiles Not established
LLM confidence is a reliable indicator of inference accuracy Evidence against, at least in current studies
Embedding-based models may outperform pure prompting Supported in recent Reddit/PANDORA work
Facet-level LLM inference from short text is reliable Not established

Maharjan et al. evaluated LLM embeddings on the PANDORA Reddit dataset using one million Reddit posts and found that embeddings trained with simple deep learning outperformed zero-shot approaches by 45% on average, with moderate psychometric reliability evidence. This suggests that “LLM inference” should not be equated with prompting a chatbot once; embeddings, supervised calibration, and psychometric validation remain central. JMIR

10. Psyche implementation: tiering by decision resolution

Within Psyche, the tier structure should map evidence quantity to output ambition:

Psyche tier Measurement input Intended output Defensible claim Unsafe claim
Lite IPIP-NEO-60 Five broad domains, domain-level narrative “This is a balanced, domain-level Big Five estimate.” “These facet scores are precise.”
Standard IPIP-NEO-120 Domains plus research-grade facet profile “Facet estimates are usable for reflection and research-like hypotheses.” “These facets justify consequential decisions.”
Heavy IPIP-NEO-300 with CAT compression High-resolution domains and facets “Adaptive administration approximates long-form coverage if calibrated and validated.” “CAT automatically makes a short test equivalent to the full test.”
LLM layer Substantial personal text corpus Behavioral-language corroboration, discrepancy analysis, longitudinal signal “Language provides an additional behavioral channel.” “A few chat snippets reveal stable personality.”

The Lite tier should be framed as a broad-factor assessment. The IPIP-NEO-60’s strength is equal representation of the thirty facet areas inside the domain scores, not independent facet precision. datashare.ed.ac.uk

The Standard tier can support facet-level reflection because IPIP-NEO-120 has four items per facet and was built for thirty-facet coverage. But Johnson’s warning should be imported into the product language: the 120-item facet scales are sufficient for research but probably should not be used to make important decisions about individuals. BPB

The Heavy tier can use IPIP-NEO-300 as the item-bank foundation for Computerized Adaptive Testing. CAT is not “ask fewer questions and hope.” It requires calibrated item parameters, trait-estimation rules, content balancing, and stopping criteria. Nieto et al. developed a 480-item Big Five facet pool, retained a 360-item pool with good psychometric properties, and found in simulation that four CAT items per facet—120 total—could provide accurate facet scores while maintaining the FFM factor structure. psicothema.com

Multidimensional and bifactor CAT are active extensions. A Big Five bifactor CAT study compared multidimensional adaptive methods with fixed and unidimensional approaches, using 360 items, and found that with 12 items per domain, multidimensional CAT and bifactor CAT were more efficient for highly multidimensional constructs such as Agreeableness. ResearchGate

The LLM layer in Psyche should be explicitly secondary. It can compare self-report with behavioral language, detect mismatch patterns, and summarize evidence across time. It should not replace structured measurement unless the text corpus is large, varied, and validated against questionnaire or informant criteria. The empirical anchors suggest that a few thousand words can support broad LLM inference in some social-media settings, while high-resolution word-level and facet-level inference may require much larger corpora. OUP Academic

11. Practical recommendations by research question

Research question Minimum reasonable tier Better tier Notes
“I need a rough Big Five control variable.” TIPI or BFI-10 Mini-IPIP Use only for broad domains.
“I need broad personality screening in a research battery.” Mini-IPIP IPIP-NEO-60 Prefer 60 when domain coverage matters.
“I need domain-level feedback for users.” IPIP-NEO-60 IPIP-NEO-120 Avoid facet claims in 60-item tier.
“I need facet hypotheses.” IPIP-NEO-120 IPIP-NEO-300 Treat 120-item facets as research-grade, not high-stakes.
“I need clinical or consequential individual decisions.” Not short forms Long-form plus multimethod assessment Use validated clinical instruments, interviews, informants, and professional judgment.
“I need language-based corroboration.” Several thousand words across contexts Longitudinal corpus plus validation Use text as a second channel, not a replacement.
“I need LLM inference from chat.” Structured elicitation conversation Multiple sessions plus self-report calibration Default assistant chats are weak evidence.
“I need facet-level LLM inference.” No settled minimum Downsampling validation required Evidence remains thin.

A useful product rule is to bind the interface to the evidence tier. A Lite report should not display facet heatmaps. A Standard report can display facets with uncertainty language. A Heavy report can display richer facet structure, but only if completion quality, response consistency, and CAT calibration checks pass.

12. Active research directions

12.1 CAT for personality

CAT is the cleanest path to sample efficiency because it optimizes item selection rather than merely shortening a form. In personality, the challenge is that domains are broad, facets are correlated, and content balance matters. A CAT that estimates Conscientiousness efficiently can still under-sample orderliness or cautiousness unless content constraints are built into item selection.

The strongest near-term direction is calibrated, facet-aware CAT over public-domain item banks. The goal is not to minimize item count absolutely, but to stop when the estimate is precise enough for the declared output. For a broad domain report, that may be far fewer items than for a thirty-facet report.

12.2 IRT-based item-bank optimization

IRT-based short forms are already visible in the IPIP-NEO-60. Maples-Keller et al. used item-level information to select items that provide strong measurement across the latent trait continuum. datashare.ed.ac.uk

Future item-bank optimization should include:

Design target Why it matters
Trait information curves Avoid items that only work at the middle of the trait range
Content balancing Prevent narrow coverage masquerading as reliability
Local dependence checks Avoid redundant items inflating precision
Demographic invariance Detect items that behave differently across groups
Longitudinal stability Separate stable trait signal from state and context
CAT stopping rules Report “not enough evidence” when precision is insufficient

12.3 Embedded personality items

Another direction is embedding personality items inside broader non-personality measures or repeated product interactions. This is attractive because it reduces user burden and allows longitudinal accumulation. It is also risky because item context can change item meaning.

The BFI-2 ecosystem illustrates a more disciplined version of short-form deployment. The BFI-2 measures five domains and fifteen facets; its 30-item short form and 15-item extra-short form are recommended for research contexts where time or fatigue makes the full measure infeasible, while the full measure is recommended for most studies because of greater reliability and validity. Colby College

Embedded personality measurement should follow the same rule: shorten only when the output is correspondingly modest.

12.4 LLM uncertainty calibration

LLM personality inference needs calibration research more than another prompt template. The critical problem is not whether an LLM can produce a plausible personality paragraph. It can. The problem is whether the model knows when the evidence is insufficient. Current evidence suggests that model confidence can remain high even when correlations with self-report collapse under reduced text. Frontiers

A serious LLM personality system should estimate uncertainty through repeated sampling across text windows, model variants, prompts, and time periods. It should use abstention thresholds. It should expose contradictions between self-report and language rather than collapse them into a single authoritative profile.

12.5 Multimethod fusion

The most promising architecture is not “questionnaire or LLM.” It is multimethod fusion:

Channel Strength Weakness
Self-report items Direct construct targeting Social desirability, self-insight limits
Informant ratings External behavioral view Relationship-specific bias
Text corpus Naturalistic behavior residue Topic, platform, genre, and privacy confounds
Interaction data Adaptive elicitation Demand effects and model steering
Longitudinal traces Stability over time Consent, surveillance, and context collapse

Brickman et al.’s 2025 overview of LLMs for psychological assessment frames LLMs as tools that may supplement self-report with scalable behavioral assessment, while also emphasizing risks, design choices, and broader ethical issues. Sage Journals

13. The open question: how few text tokens does an LLM need?

There is no defensible universal token minimum for LLM-based trait inference. The answer depends on the trait, text genre, model, prompt, ground-truth measure, population, and required decision quality.

The best current anchors are these:

Evidence source Approximate quantity Supported inference
20 Facebook statuses in Peters & Matz About 342 words on average, given 17.10 words/status Some broad-factor signal, but below full 200-status condition
200 Facebook statuses in Peters & Matz About 3,420 words on average Moderate zero-shot Big Five correlations
Schwartz et al. Facebook inclusion threshold At least 1,000 words Entry threshold for open-vocabulary population analysis
Schwartz et al. average analyzed user 4,129 words across 206 updates Large-scale language/personality modeling
Yarkoni blogs Mean 115,423 words over 23.9 months Word-level and facet-level language analyses
Piastra & Catellani text-reduction tests Decreasing Twitter-message counts Correlations tend toward zero as text decreases

The open research program should use downsampling curves. For each trait and genre, collect a validated criterion measure, then estimate model performance at increasing token counts: 100, 250, 500, 1,000, 2,500, 5,000, 10,000, and so on. The stopping rule should not be “the model sounds confident.” It should be something like: the prediction reaches a preregistered reliability threshold, converges across independent text windows, and reaches a specified fraction of the asymptotic validity curve.

For Psyche-like systems, a practical rule is:

Questionnaires can be short when they directly target the construct. Text must be longer because the construct is incidental to the data. LLMs reduce feature-engineering burden, not evidential burden.

14. Reference anchors

[Gosling et al. 2003, TIPI original validation] — Introduces the Ten-Item Personality Inventory and states the extreme-brevity use case and psychometric tradeoff. Gosling

[Donnellan et al. 2006, Mini-IPIP] — Develops and validates the 20-item Mini-IPIP as a four-item-per-domain Big Five short form. IPIP

[Rammstedt & John 2007, BFI-10] — Shows that a 10-item BFI can retain useful reliability and validity for severely time-limited research, with lower effect sizes than the full BFI. Southeastern Oklahoma State University

[Johnson 2014, IPIP-NEO-120] — Develops a 120-item public-domain inventory for five domains and thirty facets; explicitly cautions against important individual decisions from the 120-item facet scales. BPB

[Maples-Keller et al. 2019, IPIP-NEO-60] — Uses IRT to develop a 60-item IPIP-NEO representation with equal facet-area coverage for domain-level measurement. datashare.ed.ac.uk

[Yarkoni 2010, personality in 100,000 words] — Establishes why large, topically diverse writing samples matter for word-level and facet-level language/personality analysis. Academia

[Schwartz et al. 2013, open-vocabulary Facebook analysis] — Large-scale social-media personality language study using Facebook status updates, with at least 1,000 words per participant and an average of 4,129 words. PLOS

[Plank & Hovy 2015, Twitter personality data collection] — Social-media personality inference paper focused on Twitter, not subreddit-level inference. ACL Anthology

[Gjurković et al. 2021, PANDORA Reddit dataset] — Reddit comments from 10,000 users with personality and demographic labels, including 1,600 Big Five users. ACL Anthology

[Peters & Matz 2024, LLM inference from Facebook statuses] — GPT-3.5/GPT-4 infer Big Five from Facebook status updates with moderate correlations, using up to 200 updates per person. OUP Academic

[Piastra & Catellani 2025, ChatGPT 4 personality estimation] — Shows trait and dataset variability, degradation under reduced text, and poor alignment between confidence and accuracy. Frontiers

[Maharjan et al. 2025, LLM embeddings on PANDORA] — Finds LLM embeddings outperform zero-shot prompting for Reddit personality prediction under a psychometric evaluation framework. JMIR

[Nieto et al. 2017, Big Five CAT item pool] — Calibrates a Big Five facet item pool and simulates CAT measurement of FFM facets. psicothema.com

[Soto & John BFI-2 short forms] — BFI-2-S and BFI-2-XS are intended for research contexts where time or fatigue prevents full administration, with the full BFI-2 preferred for reliability and validity. Colby College

Companion entries

Core theory: Big Five Personality Measurement, Five-Factor Model, Personality Facets and Nuances, Bandwidth-Fidelity Tradeoff, General Factor of Personality, Construct Validity

Measurement practice: IPIP-NEO, Mini-IPIP, Ten-Item Personality Inventory, BFI-2, Short Forms and Survey Design, Computerized Adaptive Testing, Item Response Theory

Text inference: Open-Vocabulary Personality Prediction, Digital Trace Psychometrics, Trait Inference from Social Media, LLM Psychological Profiling, Language as Behavioral Data

Psyche system design: Psyche Assessment Tiers, Psyche Lite, Psyche Standard, Psyche Heavy, Psyche CAT Compression, Psyche LLM Inference Layer

Counterarguments and governance: Psychometric Validity in AI Systems, Privacy Risks of Personality Inference, High-Stakes Assessment Boundaries, LLM Confidence Calibration, Construct Validity Failures