Abstract. Privacy-preserving personalization aims to adapt a system to a person while reducing what the service collects, centralizes, retains, or exposes. No single technique solves the problem. Local storage, on-device inference, federated learning, differential privacy, encryption, access control, and user-facing controls protect different parts of the lifecycle. The strongest first move is usually architectural: personalize with the least data and least durable identity the product actually needs.
Coverage note: sources checked through August 2026. Legal provisions are summarized as of that date and change; verify against the current text before relying on them.
1. The tension
Personalization benefits from information:
- stated interests;
- progress;
- prior interactions;
- saved items;
- corrections;
- accessibility needs;
- timing and device context;
- sometimes sensitive preferences or inferred traits.
Privacy improves when fewer parties receive less information for less time.
This is not a contradiction that cryptography makes disappear. It is a design tradeoff. The product must decide which adaptations are worth which data.
The right starting question is:
What is the smallest representation that can support this specific decision?
Choosing a learning stream may require one explicit selection. It does not require a life history. Resuming an article may require a local progress flag. It does not require an account.
2. Separate personalization from identity
Many systems bind personalization to a durable account because accounts are convenient infrastructure. The binding is not always necessary.
| Function | Possible data | Needs real-world identity? |
|---|---|---|
| Resume reading | Local page/section state | No |
| Remember chosen stream | Local preference | No |
| Rank topics this session | Session interactions | No |
| Sync across devices | Encrypted portable state or account | Sometimes |
| Send email digest | Email address plus subscription | Yes, for delivery only |
| Recover account | Identity or recovery factor | Usually |
| Advertising profile | Cross-context behavior | No — a pseudonymous device or session identifier is enough, and none of it is necessary for learning |
An identifier can create privacy risk even when the content fields seem harmless. Once progress, interests, and conversation history share an account key, they can be joined.
Separation limits that joinability.
3. A ladder of architectures
No stored personalization
The user chooses options each visit. This is simple and private but inconvenient.
Browser-local state
Preferences and progress stay in local storage or a local database. The server delivers the same application and need not know the profile.
This is a strong fit for low-stakes progress. Limits include device loss, shared browsers, lack of cross-device sync, and exposure to scripts running in the same origin.
Session-only server state
The server remembers a temporary session and discards it after a short period. This supports richer interaction without a permanent account.
Pseudonymous account
The service stores a profile under an account identifier. Pseudonymity reduces direct identification but does not prevent re-identification from behavior or linked services. The decisive demonstration involved a recommender dataset: Narayanan and Shmatikov de-anonymized users in the published Netflix Prize ratings by correlating them with public IMDb reviews, showing that a handful of ratings with approximate dates sufficed to identify subscribers in a dataset released as anonymous. 6 Behavioral traces are high-dimensional and sparse, which makes them fingerprints; removing the name column does not change that geometry.
Client-side encrypted profile
The server stores ciphertext and the client controls the decryption key. This can reduce server visibility, but key recovery, sharing, search, and multi-device use become harder. Metadata may remain visible.
On-device inference
The model or profile processing runs locally. Raw data need not leave the device. Model downloads, device capacity, update integrity, and local compromise remain concerns.
Federated learning
Devices compute model updates locally and a server aggregates them. The original federated-learning paper emphasized learning from decentralized, privacy-sensitive device data without centralizing the raw data and introduced Federated Averaging as a practical method. 1
Federated learning is not synonymous with privacy. Updates can leak information — the "deep leakage from gradients" attack reconstructed training inputs, pixel- and token-accurate, from shared gradients alone 7 — so secure aggregation, clipping, differential privacy, participation rules, and threat modeling may still be required. Secure aggregation itself is a solved-in-principle component: Bonawitz et al. designed a protocol under which the server learns only the sum of many users' updates, not any individual contribution, at practical cost for mobile deployment. 8 The lesson of both papers together is that "federated" names a topology, and the privacy properties come from what is layered on it.
Centralized profile
The service stores interaction history and performs personalization in the cloud. This is flexible and operationally straightforward, but creates the largest concentration of personal data.
These levels can coexist. A system might keep reading progress local while using an account only for an optional digest.
4. Data minimization before privacy-enhancing technology
The NIST Privacy Framework treats privacy risk as a property of data processing and includes data minimization among its controls. 2
Minimization asks:
- Is the field needed?
- Can it be coarser?
- Can it remain on the device?
- Can it expire?
- Can an explicit choice replace an inference?
- Can the feature work without joining two datasets?
- Can aggregate measurement replace user-level history?
Examples:
| High-data design | Lower-data alternative |
|---|---|
| Store every article scroll event | Store “started” and “completed” locally |
| Infer interests from full browsing history | Ask for two topic choices |
| Retain tutor transcript indefinitely | Retain a learner-approved summary with expiry |
| Keep exact timestamps | Keep day or sequence when exact time is unnecessary |
| Build one universal profile | Separate learning, delivery, and analytics data |
Minimization is often cheaper and more reliable than protecting unnecessary data after collection.
5. Differential privacy
Differential Privacy (DP) gives a mathematical bound on how much the inclusion of one person’s data can affect a released result. In simplified terms, a mechanism is designed so an observer has limited ability to infer whether one individual’s data participated. The definition and the foundational mechanism — calibrating added noise to the sensitivity of the query, the maximum difference one person's data can make — come from Dwork, McSherry, Nissim, and Smith's 2006 paper, and everything in the modern DP toolbox is built on that account of what is being promised. 9
DP is valuable for aggregate analytics and some forms of model training. It does not make a personal profile private from the service that stores it. A recommendation generated from one user’s exact history is not protected merely because a global metric uses DP.
The privacy parameter, composition across repeated analyses, clipping, sampling assumptions, and implementation all matter. NIST’s guidance emphasizes evaluating actual differential-privacy guarantees rather than accepting the label. 3
In recommender systems, reviews find a recurring privacy–utility tradeoff: stronger noise can reduce recommendation accuracy, and how much depends heavily on which DP mechanism is used. Whether the loss falls unevenly across user groups is a different question, and the review does not answer it — it raises the fairness impact of DP and the fact that users perceive privacy differently as open research questions rather than reporting subgroup findings. Anyone claiming a differential effect on a particular group is claiming something this literature has not yet established. 4
6. Local differential privacy in deployed systems
Section 5 describes differential privacy as a guarantee about a computation over a dataset. That formulation assumes a trusted curator holding the raw records. The variant that removes the curator is the one that has actually shipped at consumer scale. Local differential privacy, in which each client randomizes its own report before it leaves the device, deserves its own treatment because the trade-off it makes is different in kind.
RAPPOR is the canonical deployment, shipped in Chrome from 2014. Erlingsson, Pihur and Korolova described a mechanism for crowdsourcing statistics from end-user client software anonymously: randomized response applied so that population-level statistics about client-side strings can be collected with a differential-privacy guarantee for each client and without linkability between their reports. Their own summary of the property is the clearest one available: the forest of client data can be studied without permitting anyone to look at individual trees. 13 The paper covers the utility analysis and the behavior under different attack models alongside the guarantee, which is the combination a deployment decision needs.
Apple's on-device work sits in the adjacent position on the ladder in section 3: evaluation and tuning performed on the device, with what crosses the boundary reduced rather than eliminated. The system's own description is precise about that reduction and worth quoting rather than rounding off: “While data held in our on-device data store never leaves the device, task results derived from this data, i.e., evaluation metrics (FE&T) and statistically noised model updates (FL) are being send to our servers.” Those results are per device and land in a central results database; the aggregation happens server-side, and what protects the individual contribution is the noise plus the stripping of user and device identifiers from server logs, not the absence of an upload. 5 The two approaches answer different questions. RAPPOR answers "what is true of the population"; on-device personalization answers "what should this device do"; and a product usually needs both.
A third deployed pattern splits trust across servers rather than injecting noise at the client. Prio has each client hold a private value while a small set of servers computes statistical functions over all clients, with the property that as long as at least one server is honest the servers learn nearly nothing beyond the aggregate, and it uses secret-shared non-interactive proofs to stay robust against clients that submit malformed values. 17 The trust assumption is different in kind from local differential privacy's: LDP assumes nothing about the collector and pays in utility, while Prio assumes non-collusion between operators and pays in deployment complexity. Neither is strictly better, and stating which assumption a system is relying on is the part usually left out of the announcement.
Three properties of the local model that change design decisions:
- The noise budget is spent per client, per report. Under a trusted-curator model the noise is added once to an aggregate. Locally, every client pays, so utility for a fixed privacy level degrades much faster as the number of distinct things being measured grows. This is why local deployments measure a small, fixed set of quantities rather than an evolving analytics schema.
- Repeated reporting is the main failure mode. A single randomized report is well protected; the same value reported daily is not, unless the mechanism explicitly handles it. Deployments address this with memoization at the client, and skipping that step quietly voids the guarantee.
- Unlinkability is a separate property from the epsilon. RAPPOR's guarantee includes non-linkability of a client's reports, which is doing distinct work from the noise level. Quoting an epsilon without saying whether reports are linkable describes less than half of the protection.
The honest framing for a product: local differential privacy is excellent for measurement and poor for building a rich individual profile, because the thing it destroys is precisely the per-person detail a profile consists of. It belongs in the analytics path, not in the personalization path. Section 11's counters are identifier-free by construction rather than by randomization, so they do not need this machinery; local differential privacy is what the analytics path reaches for when even an aggregate could expose an individual contribution. That is where it earns its place.
7. Federated and on-device personalization
Federated learning and on-device models reduce raw-data centralization, but they answer different questions.
- On-device inference: Where is the personal decision computed?
- Federated learning: How is a shared model trained across devices?
- Secure aggregation: Can the server see individual updates?
- Differential privacy: How much can outputs reveal about participation?
Apple researchers have described large-scale federated evaluation and tuning for on-device personalization, illustrating that production systems need orchestration, device eligibility, evaluation, and application-specific safeguards—not just an averaging algorithm. 5
A small product should not adopt federated learning merely because it sounds private. If an explicit local preference already solves the problem, a federated training system adds complexity and a new attack surface.
8. Profile governance
Technical privacy is incomplete without user control.
A durable profile should expose:
- each stored field;
- whether it was supplied or inferred;
- the evidence used;
- the purpose;
- who can access it;
- retention period;
- last update;
- correction and deletion controls.
The user should be able to:
- pause personalization;
- view a common, non-personalized version;
- edit interests;
- dismiss an inference;
- clear local and server state;
- export useful state;
- delete the account without leaving an orphaned profile.
Corrections must propagate. A “delete” button that hides a field in the interface while leaving it in embeddings, analytics tables, and backups does not satisfy the ordinary meaning of deletion. Where personal data has reached model weights, deletion has its own research field — machine unlearning — whose central finding is architectural: efficient removal is possible when the training pipeline was designed for it, and expensive to retrofit when it was not. 10
9. Field provenance and purpose
A small profile can be governed field by field.
| Field | Source | Purpose | Scope | Expiry |
|---|---|---|---|---|
| Chosen stream | Explicit selection | Content sequence | Local browser | Until reset |
| Article completed | Local event | Resume and next item | Local browser | User controlled |
| Prefers weekly digest | Explicit setting | Delivery cadence | Account | Until changed |
| Interested in hardware | Explicit selection | Ranking | Account or local | Review periodically |
| “Low technical ability” | Inferred from behavior | Unclear | — | Do not store |
The table exposes a useful test: if purpose or scope cannot be stated, the field should not be collected.
Provenance also helps correction. An explicit setting should outrank a model inference. A recent correction should supersede an older behavioral guess. A low-confidence inference should expire rather than silently become part of identity.
10. Deletion is a system property
Personal data can spread into:
- primary databases;
- caches;
- search indexes;
- embeddings;
- analytics events;
- model prompts and traces;
- backups;
- exports;
- derived summaries.
A deletion design needs a data map. It should distinguish:
- immediate removal from active use;
- scheduled removal from backups;
- tombstones needed to prevent re-import;
- aggregate statistics that no longer identify a person;
- records retained under a legitimate, documented obligation.
For local-only state, reset should clear every key and database used by the feature. For server state, deletion tests should query downstream indexes and derived objects, not merely the profile table.
“Forget this” is a behavioral promise. The verification should show that the old information no longer shapes recommendations or answers.
11. Anonymous measurement
Personalization experiments need measurement, but measurement does not automatically require a user profile.
Counts can track:
- outbound source opens;
- companion-section expansions;
- second-item starts;
- stream selections;
- completion events;
- serendipity-card opens.
Privacy improves when the event system:
- uses no stable cross-session identifier;
- records only the event and coarse context needed;
- avoids full URLs or text payloads containing personal data;
- applies short retention;
- separates operations from analytics;
- publishes what is counted;
- does not send data to third-party trackers.
Aggregate counters cannot answer every research question. That limitation can be preferable to quietly constructing a behavioral dossier.
12. Sensitive inference
Some traits are risky even if they were never explicitly collected. A system may infer:
- health conditions;
- religion;
- political orientation;
- disability;
- financial stress;
- sexual orientation;
- family circumstances;
- psychological traits.
The fact that an inference is probabilistic does not make it harmless. It may be wrong, yet still change what the system shows. And the inferences are cheap: Kosinski, Stillwell, and Graepel showed in 2013 that Facebook Likes alone — a behavioral trace far thinner than a reading history — predicted sexual orientation, religion, political views, and personality with disturbing accuracy. 11 A centralized reading profile of the kind section 3's upper rungs describe sits on richer data than that study did; the lower rungs deliberately do not. The question is never whether sensitive traits could be inferred; it is whether the system is designed not to.
The safest rule is purpose limitation: do not infer or store a sensitive trait unless the user-facing function clearly requires it, the benefit is substantial, the user has meaningful control, and the legal and ethical basis is established.
A learning product rarely needs to infer intimate traits to choose between a technical and a philosophical introduction.
13. Threat model
Privacy claims need an adversary.
Potential threats include:
- the service operator;
- a compromised account;
- malicious scripts in the browser origin;
- another user of the same device;
- an analytics vendor;
- a model provider receiving prompts;
- an attacker stealing a database;
- model-update leakage in federated learning;
- insiders with broad access;
- future data joins not envisioned at collection time.
“Stored locally” protects against some server threats but not a shared device or cross-site scripting. “Encrypted” protects data at rest but not necessarily the application while it is using the profile. “Anonymous” counters may still become identifying when combined with stable identifiers. And data that reaches a model's training set has its own exit path: Carlini et al. extracted verbatim training examples — including names, contact details, and conversation fragments — from a publicly released large language model by querying it, establishing memorization as a concrete disclosure channel rather than a theoretical one. 12 A personalization pipeline that feeds user text into model training has added that channel to its threat model whether or not it noticed.
14. Evaluation
Privacy-preserving personalization needs two evaluations.
Utility
Does the lower-data architecture meaningfully improve completion, learning, discovery, or satisfaction over a non-personalized baseline?
Privacy
What data exist at each stage? Can they be linked? Can sensitive attributes be inferred? Does deletion work? What does a compromised component reveal?
A useful comparison table includes:
- fields collected;
- storage location;
- retention;
- encryption boundary;
- model-provider exposure;
- cross-device behavior;
- correction and deletion;
- measured utility gain;
- privacy guarantee and assumptions.
The baseline should include simple explicit controls. A sophisticated private model that barely beats “choose your interests” may not justify itself.
15. The legal layer this architecture sits inside
Every architectural choice in this article has a legal counterpart, and reading the two together is more useful than treating compliance as a separate downstream activity, because the law encodes roughly the same design intuitions with sharper edges.
Five provisions of the GDPR do most of the work for a personalization system.
Purpose limitation (Article 5(1)(b)): personal data must be collected for specified, explicit and legitimate purposes and not further processed in a manner incompatible with them. This is the legal form of the field-purpose column in section 9. The prohibition is narrower than a flat ban on reuse: the text forbids further processing “in a manner that is incompatible with those purposes,” and Article 6(4) supplies the compatibility test — the link between the two purposes, the context of collection and the relationship between controller and subject, whether special categories are involved, the consequences for the data subject, and what safeguards such as encryption or pseudonymisation are in place. Reuse can also rest on the subject's consent or on qualifying Union or Member State law. So a profile field collected to sequence lessons and later used to target promotions is not automatically unlawful; it is the case the compatibility assessment exists to decide, and it is the kind of reuse that assessment most often fails.
Special categories (Article 9): processing of data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs or trade-union affiliation, together with genetic data, biometric data processed to uniquely identify a person, and data concerning health, sex life or sexual orientation, is prohibited unless a narrow exception applies. This is the provision section 12 is really about: an inference that a reader is likely gay, ill, or devout is special-category data whether it was declared by the user or derived by a model, and deriving it does not create an exception. A system that infers such traits as a by-product of personalization has not found a loophole; it has entered the regime with the highest bar in the regulation. 14
Data minimisation (Article 5(1)(c)): data must be adequate, relevant and limited to what is necessary for those purposes. Section 4's argument that minimization comes before any privacy-enhancing technology is the same claim, and the ordering is the same. Necessity is assessed against the purpose, so a purpose stated vaguely enough makes almost anything look necessary.
Data protection by design and by default (Article 25): appropriate technical and organisational measures, with pseudonymisation as the named example, implemented both when the means of processing are determined and during processing itself; and by default, only the personal data necessary for each specific purpose are processed, an obligation that explicitly covers the amount collected, the extent of processing, the storage period, and accessibility. That is the ladder in section 3 restated as an obligation: the default rung matters more than the highest rung available.
Automated individual decision-making (Article 22): a data subject has the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects or similarly significantly affects them, with narrow exceptions for contractual necessity, authorising law, and explicit consent, and with a requirement in two of those cases to provide at least human intervention, the ability to express a view, and the ability to contest the decision. 14
Three things a designer should take from this rather than a compliance checklist:
- "Similarly significantly affects" is where the argument happens. Ranking a news feed is usually outside it; determining eligibility, pricing, or access is usually inside it. A personalization system that drifts from the first into the second does so gradually and without a release note.
- Contestability is an architecture requirement. A right to contest a decision presupposes that the decision, its inputs, and its time are recorded. That is the same field-provenance machinery section 9 argues for on engineering grounds, and it cannot be added afterwards.
- Consent is not a general-purpose key. It appears here as one narrow exception among several, attached to conditions. Treating a consent checkbox as authorisation for whatever processing follows misreads the structure of the provision.
Jurisdictions and article numbers differ and change. The four questions underneath do not: what is this field for, is it necessary, what does the system do by default, and can a person contest an outcome.
16. Personalization without a cross-context identifier
The instructive thing in this area is a plan that was announced, built, and then abandoned. The web's cross-context identifier was going to be removed and replaced; it was not, and the replacements were retired. What survives is the design pattern, which is why the section is still here.
Apple's App Tracking Transparency does hold: an application must obtain permission before tracking a user or accessing the device's advertising identifier. 15 Google's Privacy Sandbox went the other way, in two steps. On 22 July 2024 Google abandoned the deprecation itself, proposing instead a new in-Chrome choice experience. 16 In April 2025 it dropped that prompt too, leaving third-party cookies in place and managed through Chrome's existing settings. 20 Then on 17 October 2025 it announced the retirement of the replacement stack itself — the Topics API among ten technologies named for withdrawal across Chrome and Android, together with Protected Audience, Attribution Reporting, IP Protection, On-Device Personalization, Private Aggregation, Protected App Signals, Related Website Sets, SelectURL and SDK Runtime — citing low adoption. 18 Announced is not gone: Google's own feature-status page, updated 14 August 2026, still lists Topics as “Deprecate and remove” on the web and “Scheduled for phaseout” on Android, alongside the rest of the stack in the same two states. 21 Deprecation is the announcement phase; removal is a later event that had not happened at this article's coverage date.
Topics is worth describing anyway, because the idea outlived the API: the browser derived a small set of coarse interest topics from local browsing behaviour and exposed only those to callers, with the stated design goals that the specific sites a user visited would not be shared and that users would not be re-identified across sites. 19 Those were the goals Google published for it, and independent work measured how well the second one held: Jha and colleagues showed the topic vector itself supports probabilistic cross-site linkage well enough to re-identify a user 15–17% of the time in a pool of a thousand, and concluded that the API “mitigates but cannot prevent re-identification.” 22 The design reduced the risk substantially against third-party cookies; it did not close it, and the API was named for retirement with that gap on the record.
The two cases are not the same move, and collapsing them loses what each one teaches. Topics is the move this article's ladder describes: interest derivation migrates to the client, and what crosses the boundary is a deliberately lossy summary rather than a record. App Tracking Transparency is a different lever entirely — a permission and policy boundary over whether cross-company linking may happen at all, imposing no local-computation architecture and exporting no summary. One redesigns what crosses the boundary; the other decides whether anything does. Three observations:
- Coarseness is the mechanism, not a limitation — and it mitigates rather than prevents. A small vocabulary of topics is what keeps the exposed signal from working well as an identifier; on the measurement above, it does not stop it working as one at all. Requests to make the taxonomy richer are therefore requests to weaken a property that was already partial, which is the right way to answer them — and a system that treats coarseness as a guarantee rather than as a rate has mistaken which kind of claim it is holding.
- A permission prompt is a governance surface with measurable outcomes. Whether a user grants tracking permission is an observable that the system's behavior must handle in both branches, which makes "what does the product do for a user who says no" a first-class design question rather than a degraded path.
- These are advertising architectures, and the read-across is partial. A learning product or a news reader has a first-party relationship and does not need cross-site inference at all. The useful lesson is the pattern, derive locally and export coarsely, not the specific APIs.
17. Bottom line
Privacy-preserving personalization is not one algorithm. It is a sequence of choices about identity, data, computation, retention, and control.
The strongest default is progressive:
- explicit choices;
- local progress;
- session-limited adaptation;
- optional accounts for functions that truly need them;
- richer profiles only after they demonstrate value;
- advanced privacy-enhancing technologies where their threat model fits.
Personalization should earn the data it uses.
Related concepts
Personalization: User Modeling in LLMs, Personalized AI Learning Systems, Learner Modeling for Adaptive AI, Recommender Systems
Privacy: Local-First AI Tooling, Federated Learning, Differential Privacy, Data Minimization, Secure Aggregation
Evaluation: Trust Calibration in Human-AI Systems, Human-AI Reliance, Distribution Shift, External Validity