Reference

Information Overload and AI Curation: Reducing Volume Without Losing the World

Abstract. Information overload occurs when the demands of finding, comparing, and acting on information exceed a person’s available attention or decision capacity. AI curation can help by filtering, ranking, clustering, summarizing, and scheduling information. It can also make the problem less visible while concentrating editorial power in an opaque system. The design challenge is not to maximize compression or engagement. It is to preserve a usable path from a small, personalized surface back to diverse sources, uncertainty, and the reasons items were selected.

Coverage note: sources checked through August 2026.

1. Too much information is a decision problem

Information overload does not mean that a large library is inherently harmful. A library can contain millions of books without overwhelming a reader who can search, browse, and stop.

Overload appears when the information environment creates more processing demand than the person can manage for the task at hand. Reviews of the literature describe interacting causes:

  • volume;
  • arrival rate;
  • complexity;
  • duplication;
  • uncertainty about quality;
  • interruption;
  • time pressure;
  • poor organization;
  • a mismatch between information and decision needs. 1

The same amount of information can be useful to one person and overwhelming to another. Context matters: a daily briefing read on a phone during breakfast is different from a research archive used for a week-long investigation. The library-science literature adds a longitudinal caution: overload is a recurring anxiety with every media transition, and its documented pathologies — satisficing on the first adequate source, avoidance, anxiety — are behavioral adaptations to mismatch, not properties of volume itself. 5

This suggests a plain definition:

Information overload is a failure of fit between an information stream, a person’s available attention, and the decision or learning task.

2. Where the idea came from: attention as the scarce resource

The framing above — overload as a decision problem rather than a volume problem — is not new, and its original statement is more useful to a system designer than most contemporary versions because it names the economics directly.

Herbert Simon set it out in 1971. His argument runs by analogy to a population problem: a rabbit-rich world is a lettuce-poor world, so the interesting question about any abundance is what it consumes. Information consumes the attention of its recipients, so a wealth of information creates a poverty of attention, and with it the need to allocate that attention among an overabundance of sources competing for it. 10 The conclusion he draws is the one that should govern an AI curation product: in an information-rich world, most of the cost of information is borne by the recipient, so knowing what it costs to produce and transmit something says nothing about whether it was worth receiving.

Simon proposed a unit for the accounting, and the choice is instructive. He rejected the bit, because a message's bit count depends entirely on encoding and is therefore not an invariant measure of what it costs to attend to. What he proposed instead is the recipient's time.

Simon was not first to the diagnosis, though he was first to the accounting. Vannevar Bush's 1945 essay had already located the problem in selection rather than production: the trouble was not that too much was being published, but that publication had extended far beyond anyone's ability to make real use of the record, and that the methods of transmitting and reviewing results were generations old and no longer adequate. His proposed remedy, the memex, was a device for associative trails through a personal store — an indexing scheme modeled on how a mind jumps between items rather than on alphabetical or class order, with the trails themselves saved and shareable. 16 Two of that proposal's assumptions are worth carrying into any AI curation design. The first is that the valuable artifact is the trail, not the item: what one reader assembled by connecting sources is a durable object other readers can follow. The second is that selection is personal by construction — Bush's device belonged to one person and indexed what that person had chosen to keep. Neither assumption survives into a ranked feed, which produces no reusable trail and stores nothing the reader has claimed. The archive pattern of section 7 recovers the durability half — entries that persist and connect — but not the trail itself, and not the personal act of selection; nothing in this article's pipeline reinstates either.

The design consequence Simon stated is the sharpest sentence available on this subject, and it is his, not Bush's:

To be an attention conserver for an organization, an information-processing system must be an information condenser. It is conventional to begin designing an IPS by considering the information it will supply. In an information-rich world, however, this is doing things backwards. The crucial question is how much information it will allow to be withheld from the attention of other parts of the system. 10

Read against a modern feed, that inverts the usual product metric. A curation system is not succeeding when it surfaces more relevant items; it is succeeding when it withholds more irrelevant ones at an acceptable miss rate. The two are not the same objective and they do not have the same failure mode — the first is measured by what the user saw, the second by what the user never had to see, which is the harder number and the one nobody's dashboard shows.

Most of the distinctions in section 3 are mechanisms for withholding. Ranking withholds by demoting; summarization withholds by compressing; scheduling withholds by delaying; filtering withholds outright. Clustering and explanation are the exceptions — they reorganize and disclose rather than withhold, and explanation is the one layer that gives attention back. Naming the shared function makes the shared risk visible too: every one of them can withhold something that mattered, and section 10's invisible-omission failure mode is the general form of that risk.

3. AI can intervene at several layers

Filter

Remove duplicates, spam, low-quality sources, and material outside a declared scope.

Rank

Order items by predicted relevance, urgency, quality, novelty, or another objective.

Cluster

Group many reports about the same event or idea so the reader sees the underlying development rather than twenty headlines.

Summarize

Compress a source or cluster into a smaller representation.

Route

Send different items to different readers, streams, or times.

Schedule

Bundle updates into a cadence that fits attention: live ticker, daily digest, weekly review, or “only when this changes.”

Explain

Show why an item appeared and what evidence supports the summary.

These functions should not be collapsed into “recommendation.” A person may want aggressive duplicate removal but little personalization, or a short summary with strong source traceability.

4. Compression moves risk rather than removing it

A summary reduces reading time by omitting information. The central question is what was omitted.

AI summaries can flatten:

  • uncertainty;
  • disagreement between sources;
  • scope conditions;
  • methodological limits;
  • chronology;
  • the distinction between a reported claim and verified fact;
  • the difference between an incremental update and a major result.

A fluent paragraph can make these losses hard to notice. The summarization literature has measured the fluency and the unfaithfulness directly; that readers fail to catch the difference is the inference this article draws from putting the two together, not a third measured result. Maynez et al. found that abstractive summarizers frequently generate content unfaithful to the source document, hallucinated facts that nothing in the summary itself flags. 6 Benchmarking work on LLM news summarization took its single best-performing system — zero-shot Instruct Davinci — to a head-to-head against summaries commissioned from freelance writers and found it “rated as comparable to the freelance writers” on aggregate. That is a result about one model, not about instruction-tuned models as a class: performance varied substantially across the ten systems evaluated. The same work also found that the reference summaries in standard benchmarks were themselves poor, a known but unaddressed defect that had been distorting reported progress. 7 Fluency and fidelity are separate axes, and readers can only inspect the second one when the source is reachable. Compression therefore needs provenance. The reader should be able to open the original, see which sources were combined, and distinguish the system’s synthesis from a publisher’s claims. Source Provenance and Claim Traceability provides the deeper data model.

For high-stakes or contested material, “one paragraph” may be the wrong shape. A compact card can instead include:

  • what happened;
  • why it matters;
  • what remains uncertain;
  • source count and diversity;
  • publication and event dates;
  • links to originals;
  • a visible update history.

The summary becomes an entry point rather than a substitute.

5. Ranking objectives shape the information world

A ranker optimizes something, even when the objective is implicit.

Common objectives include:

  • predicted click;
  • predicted dwell time;
  • chance of completion;
  • topical relevance;
  • freshness;
  • source authority;
  • estimated learning value;
  • diversity;
  • novelty;
  • urgency.

Optimizing click probability can favor sensational or familiar items. Optimizing completion can reward short, easy material. Optimizing topical relevance can narrow the stream around known interests.

News-recommender research treats overload reduction as a central motivation but also documents the need to consider diversity, novelty, serendipity, popularity bias, and the effects of repeated recommendation on user behavior. 2, 3

No single scalar captures a healthy information diet. A curation system needs a portfolio objective.

6. Targeting and serendipity

Personalization is valuable because a universal feed forces everyone to sift through the same volume. But perfect prediction would not be a complete solution.

A reader has at least three needs:

  1. Declared interests — topics they know they care about.
  2. Obligatory context — developments needed to understand those topics.
  3. Serendipitous discovery — worthwhile material they would not have requested.

Serendipity in Recommender Systems is not randomness. A serendipitous item is unexpected and useful. A learning-oriented AI publication can reserve part of each edition for adjacent or distant material with an explanation of the connection.

This prevents the system from treating the current profile as the limit of the person’s future interests.

7. Ticker, digest, and archive solve different problems

An information product often needs three speeds.

Ticker

The ticker answers “what is moving now?” It should be sparse, timestamped, and cautious. It is a discovery surface, not a place for definitive analysis.

Digest

The digest answers “what was worth attention in this period?” It clusters duplicates, adds context, and reflects the reader’s interests.

Archive or compendium

The archive answers “what does this mean, and how does it connect?” It provides durable entries, revision history, and deeper sources.

Mixing the three creates problems. If the ticker inherits the authority of the archive, early reports can look settled. If the archive is written like a ticker, it becomes stale and shallow.

8. Cadence, interruption, and the cost of arriving

Sections 3 and 7 treat scheduling as one lever among several. It deserves separate treatment, because the cost a curation system imposes is not only the cost of reading an item — it is the cost of being interrupted to consider it, and that cost is real, paid in stress and in the work of resuming rather than reliably in elapsed time.

The empirical picture is more interesting than "interruptions are bad." Mark, Gudith and Klocke found that people completed interrupted tasks in less time, with no measurable difference in quality — and at a cost paid in stress, frustration, time pressure and effort rather than in output. 11 That result should discipline how a product team reads its own metrics: throughput can look fine, or better than fine, while the experience degrades. Nothing in a completion-rate dashboard would show this.

Recovery is a process rather than a free instantaneous switch back. Iqbal and Horvitz's field study of multitasking traced how computer users actually suspend and resume tasks, and derived design guidance from the observed patterns rather than from an assumed cost model. 12

The most direct evidence about cadence comes from removing a channel entirely, though it is small. Mark, Voida and Cardello cut off email for thirteen information workers over five days and measured what changed: without it, people multitasked less and held longer task focus, measured as a lower frequency of shifting between windows and longer time spent in each. 13 What the design cannot establish is why. Cutting email off removes the notifications, the voluntary checking, the email tasks and the email content in a single move, so the study licenses the effect and not a mechanism. The reading this article takes from it — that a channel which may deliver at any moment converts every moment into a decision point — is an inference from that effect, and a product that batches on the strength of it is betting on the inference rather than on a measured cause.

For a curation product that can deliver on any schedule, four things follow:

  • Batching is a first-class feature, not a settings-page afterthought. A digest at a known hour and a stream that may arrive at any second differ in the number of interruptions they create, and interruptions are the quantity being spent. What batching buys is contested, though: the same group's later field study of forty workers over twelve days found batching “associated with higher rated productivity with longer email duration” — batching as a main effect was not significantly related to productivity at all, and the association appears only in its interaction with time spent on email — while finding, “despite widespread claims … no evidence that batching email leads to lower stress.” 19 The case for a digest rests on interruption count and attention, not on a stress benefit the evidence does not show.
  • Urgency must be earned per item, and almost nothing earns it. The default for an item should be the next scheduled batch; immediate delivery should require a reason the system can state. Section 10's alert-fatigue mode is what happens when that default is inverted.
  • Measure the interruption count, not only the open rate. Two designs delivering the same twenty items per day, one in three batches and one on arrival, can produce similar engagement numbers at very different cost. Only one of those numbers appears in a standard analytics view.
  • Quality metrics can stay flat while the product gets worse. That is the direct implication of the interrupted-task result, and it is why a curation product needs a subjective-burden measure beside its behavioral ones.

9. User control

Useful controls are concrete:

  • topic follows and mutes;
  • cadence;
  • maximum length;
  • “more technical” or “more introductory”;
  • source preferences;
  • a serendipity dial;
  • show fewer updates about the same event;
  • pause a topic;
  • reset the profile;
  • explain this recommendation.

Too many controls create their own overload. Research on recommender control suggests that people differ in how much configuration they want and can use, and that added control can raise cognitive load even as it raises acceptance. 4 Field experiments point the same direction from the positive side: Harper et al. gave MovieLens users one coarse control over their recommendations — a 2×2 between-subjects design in which each participant tuned either popularity or item age, not both — and found that users in both conditions reported that the tuned lists better represented their preferences. 8 The controls were deliberately obfuscated for the experiment (a button labelled “right” rather than “newer”, the task framed as picking “a position in a spectrum of values”), and subjects duly rated the interface hard to use while rating the results well — which is itself the finding a product should take: the tuning helped, the unlabelled dial did not. The outcome is self-report, so the study cannot say whether control improved the fit itself or the feeling of it; either way, the direction is that a little legible control helps. Good defaults and a few legible controls are usually better than a dashboard of weights.

The most important control may be editorial: the ability to leave the personalized stream and see a common, non-personalized edition.

10. Failure modes

Summary monoculture

Many readers see the same generated synthesis, so one framing error scales widely.

Source laundering

Low-quality claims enter through a polished summary without visible attribution.

Duplicate amplification

Ten outlets repeat one press release and the system mistakes repetition for independent confirmation.

Freshness bias

New items displace slower, more important work.

Profile lock-in

The system treats past clicks as permanent interests and stops offering new domains.

Engagement substitution

Clicks and return visits replace understanding as the actual objective.

Alert fatigue

A “live” product produces so many urgent items that urgency loses meaning.

Invisible omission

The system never shows what it excluded, so readers cannot tell whether the feed is narrow by choice, accident, or bias.

Temporal confusion

Publication date, event date, model release date, and article update date are collapsed. Fast-moving AI coverage is especially vulnerable.

11. The editorial layer

AI curation does not remove editorial judgment. It distributes editorial choices across data ingestion, prompts, ranking weights, thresholds, and interface design.

Those choices should be legible:

  • What counts as in scope?
  • Which sources can enter automatically?
  • What requires human review?
  • How are announcements distinguished from independent evidence?
  • When are two reports treated as duplicates?
  • What makes a development important enough for the common feed?
  • How much of the surface is personalized?
  • How are corrections handled?
  • What cannot be summarized without losing essential context?

An editorial policy can assign content states:

State Meaning Surface
Detected Automated intake found a candidate Internal only
Corroborating Source identity and duplicates under review Internal only
Breaking Timely but still uncertain Ticker with caution
Verified update Evidence clears publication threshold Digest
Explained Context and durable significance reviewed Compendium
Corrected Earlier summary changed Visible revision
Retired No longer current; kept for history Archive

The states stop a fast signal from borrowing the authority of a reviewed encyclopedia entry.

12. A small content schema

A curation system benefits from structured records before generation:

item
  canonical source
  source type
  event date
  publication date
  retrieved date
  topic tags
  evidence type
  cluster id
  claim records
  uncertainty
  correction state
  editorial status
  stream eligibility

The generated headline, card, digest paragraph, and wiki link become views over this record. They need not each reinvent the source interpretation.

Structured state also supports cadence. A weekly digest can query verified changes from the last period. A ticker can query breaking items with stricter display rules. A later MCP endpoint can expose the same records without scraping prose.

13. AI-specific overload

AI coverage has several features that amplify overload:

  • model and product names change quickly;
  • vendors release benchmarks with different settings;
  • one technical report generates many derivative stories;
  • demos circulate before limitations are known;
  • a model release, API update, app feature, and research result may share one brand name;
  • leaderboard snapshots expire;
  • confident commentary outruns primary evidence;
  • daily volume makes “important” indistinguishable from “new.”

A curation system should therefore normalize the kind of event:

  • research publication;
  • model checkpoint;
  • product release;
  • API or price change;
  • benchmark result;
  • policy or governance change;
  • acquisition or financing;
  • security incident;
  • commentary.

The distinction helps readers decide what attention the item deserves.

14. Evaluation

A curation system needs more than click-through rate.

Dimension Possible measure
Load Time to reach a stopping point; items opened; perceived overload
Withholding Items excluded per item shown; sampled audit of exclusions for missed importance (the miss rate, section 2)
Utility Reader-rated value; task completion; knowledge gain
Fidelity Claim-level source support; correction rate
Diversity Topic, source, geography, evidence type, and viewpoint coverage
Serendipity Unexpected-and-useful ratings; later topic exploration
Calibration Confidence and uncertainty match source evidence
Agency Use of controls; successful profile correction
Long-term value Return after completion, not endless scrolling

Evaluation should include counterfactual feeds. Did the personalized version help compared with a well-edited common digest? Did the summary help compared with headlines and links? Did serendipity increase useful discovery or merely noise?

The table's non-accuracy rows have a lineage worth naming: the recommender-systems community had argued since the mid-2000s that accuracy-centric evaluation was actively misleading — Herlocker and colleagues in 2004 on what accuracy metrics fail to capture 17, McNee, Riedl and Konstan in 2006 on accuracy metrics hurting recommender systems 18 — and by 2010 was proposing coverage and serendipity as first-class metrics — with the observation, precise and still current, that recommending only popular items scores well on accuracy while being useless as curation. 9

15. A design pattern for AI curation

A defensible pipeline looks like this:

  1. Ingest source metadata and content with stable identifiers.
  2. Separate original reporting, commentary, announcements, and research papers.
  3. Cluster near-duplicate coverage.
  4. Identify the earliest and strongest sources.
  5. Extract claims with source spans.
  6. Generate a cluster summary that preserves disagreement and uncertainty.
  7. Apply an editorial quality gate.
  8. Rank using an explicit mix of relevance, importance, diversity, and serendipity.
  9. Show why each item appeared.
  10. Keep the original links and correction history visible.
  11. Sample what was withheld and check it for missed importance, on a schedule.
  12. Measure completion and return without building an unnecessary identity trail.

The model assists at several stages, but the product contract determines the objective and the evidence bar.

16. The regulatory layer

A curation system that is an online platform in the European Union's sense — a hosting service that stores and disseminates information to the public at its users' request — is no longer only making product decisions. The Digital Services Act imposes obligations on providers of such platforms that use recommender systems, and they map onto the mechanisms in this article; an editorial digest that ingests third-party sources and publishes its own synthesis is outside that definition, but the obligations are still worth reading, and they are worth knowing as design constraints rather than as compliance paperwork, because they encode a specific theory of what goes wrong with ranking.

Two provisions matter most here. Article 27 requires platforms using recommender systems to set out in their terms, in plain and intelligible language, the main parameters used in those systems and any options for users to modify or influence them. Article 38 goes further for very large online platforms and search engines: at least one option for each recommender system must be not based on profiling, in the sense Article 4(4) of the GDPR defines it. 14 The first is narrower than it looks: Article 27 asks for the main parameters, the criteria “most significant in determining the information suggested” and the reasons for their relative importance. It does not ask a platform to state the objective it optimises — so the ranking objectives of section 5 can remain unstated inside a fully compliant disclosure, and a reader who expects Article 27 to surface them will not find them there. The second requires that the control ladder in section 9 have one specific rung on it.

Neither provision says how to tell whether a ranking is any good, and the research that comes closest to filling that gap was not written to operationalise either one. Vrijenhoek and colleagues argue that recommender evaluation focused on clicks and short-term engagement fails to capture a user's longer-term interest in diverse and important information, and propose metrics grounded in democratic-theory conceptions of diversity rather than in generic dissimilarity. 15 That distinction is the substantive one: "diversity" computed as embedding distance between items and "diversity" meaning exposure to a spread of perspectives are different quantities, and only the second is what the normative argument is about.

Three practical notes and one caution:

  • A non-profiling option is a real product surface. Chronological, editorial, and popularity-based ranking are the usual candidates and produce very different experiences — but the label is not the compliance. Profiling is defined by the processing, not by the name of the ranking: any of the three can still evaluate personal aspects of the user depending on what it reads, and any of the three can avoid it. The question to answer is what the option consumes, not what it is called. Building the worst usable version and calling it compliance is visible to users either way.
  • Stating ranking parameters is a design discipline before it is a legal one. A system whose main ranking parameters cannot be described in plain language is one whose objectives are probably not understood internally either.
  • A diversity metric needs a stated normative basis. Reporting a diversity number without saying which conception of diversity it operationalizes is the same class of error as reporting an accuracy number without saying what was labeled.
  • Regulation moves. Article numbers, thresholds, and designations change. The two design questions underneath — can you explain your ranking, and can a user turn profiling off — do not.

17. Bottom line

AI curation can make a fast-moving field understandable by reducing repetition, matching depth to the reader, and connecting new developments to durable concepts. It can also hide editorial choices behind a smooth personalized surface.

The right goal is not “the shortest possible feed” or “the most engaging feed.” It is a bounded, source-linked path through abundance: enough targeting to be useful, enough common context to remain oriented, and enough serendipity to keep the reader’s world open.

Curation: Recommender Systems, Serendipity in Recommender Systems, News Recommender Systems, Model Routing

Trust: Source Provenance and Claim Traceability, Trust Calibration in Human-AI Systems, AI-Generated Summaries, Distribution Shift

Human factors: Automation Bias, Algorithm Aversion, Human-AI Reliance, Need for Cognition