Reading companion

Guardian Angels: LLM Personalization for Productivity and Security

This page is meant to sit in a second tab beside the original — it explains and orients, but it is deliberately useless as a substitute. Collapsed sections open where you want more.

Open the original ↗

Before you read

The longest and hardest piece in the collection — about an hour, and its second half assumes you know how AI models are built. Read it for the first half, which needs no technical background and is the clearest statement anywhere of what ordinary people will need from AI.

It opens with a question the author could not answer about himself: what does his working life look like in 2030? Two things stand in the way — being drowned in machine-generated deception, and being made redundant. One object answers both: a guardian angel ("GA"), an AI trained to emulate you rather than to be a generic assistant belonging to a company.

Principal and agent come from economics: the principal wants something done, the agent does it. Agents have interests of their own — a real-estate agent wants a fast sale, you want a high price. Gwern collapses the distinction. If the agent is a model of you, its interests are yours by construction.

Where to stop. Everything before the heading "Chatbot Fixes" is the argument in plain language, and you have the whole case by then; everything after is machinery, and turns technical fast. Skim it, then rejoin at "Principles" and "Anti-Principles," plain again and among the sharpest pages here.

He calls the project one "in the spirit of uploading" — the transhumanist hope of copying a person onto a computer. That is the lineage Meghan O'Gieblyn traces in "Ghost in the Cloud," arriving here as an engineering proposal with a budget.

While you read

The great-aunt who stopped answering her phone

"my great-aunt no longer trusted herself to handle her own phone calls, and screened everything through her daughter"

This story is the hinge of the essay, not a warm-up anecdote. An elderly woman concluded she could no longer safely operate her own telephone and handed the channel to a trusted relative. Gwern turns it on himself: he already struggles to detect AI slop and already writes off whole regions of social media, so on what concrete grounds does he expect to handle the scams of 2030? Notice that his opening inventory is observational, not predictive — synthetic-media hoaxes, pig-butchering scams, projects closing to outside contributions have all happened already.

Open this passage in the original →
Why the incentives point at replacing you, not amplifying you

"increasingly, you are the bottleneck to be optimized away"

The economic argument, and the part most worth arguing with. It runs on Amdahl's law: a system is only as fast as its slowest required step. If a human must review the work, the human sets the ceiling — so the money is not in making one worker faster but in removing the requirement for the worker. The analogy is exact and unkind: the internal combustion engine did not make its fortune helping horses. Nobody has to be villainous for this to bite: it is about where profit sits, and whoever pays for a model is who it is finally for.

Open this passage in the original →
Five ways today's assistants fail the person using them

"the 'personalization' or 'memory' features are typically laughably simplistic Markdown snippets encoding simple facts"

Five distinct failures. Mode collapse: training a model to please the average human sands off the personality that made it interesting. Laziness: minimum effort on anything not being scored. Brittleness: context windows, however large, cannot hold a life, and the model is mostly locating an answer it already half-knows. Over-helpfulness: the next section's subject. Amnesia: the load-bearing complaint — correct a model today and the correction dies with the conversation, so an hour spent teaching it is an hour gone. Today's memory features, as the quoted line says, are notes stapled to a model whose weights never change.

Open this passage in the original →
The security case: a servant who does not know whose it is

"The very re-programmability of a chatbot by its prompt is the key to prompt attacks."

The plainest explanation of prompt injection you will find. A chatbot cannot reliably tell instructions from you apart from text it merely happens to be reading — a web page, an email, a document. Both arrive as words in the same window. Security people call the result a confused deputy: a servant holding real authority over your inbox and your money, with no way to ask whether an order makes sense coming from this source. Gwern's fix is not a better filter but an identity — a model trained to be one person's agent should find the attack absurd, the way you do not wire money because an email told you to. His most interesting claim, and his least demonstrated.

Open this passage in the original →
Where the essay turns technical

"you just have to keep it slow and subtle enough to not matter too much within a single human lifetime"

At "Chatbot Fixes" the essay changes audience. Three ideas carry the rest, and their names are enough. Dynamic evaluation: updating the model's own weights as you use it, so the model itself changes rather than being handed longer notes. Active learning: the model choosing the most informative question to ask you next, instead of passively hoovering up data. The third is a long argument about why training information into a model works worse than showing it in conversation. Skim freely. He is not claiming to solve AI alignment, only that imitating one person who is still around to be asked is a far easier problem.

Open this passage in the original →
What "you" turns out to mean

"you are what your brain does, its desires, hopes, goals, preferences, esthetics, personality, beliefs, ideologies, all of that"

Pause here. This is a definition of personal identity chosen for engineering convenience, and Gwern says as much: identity as personality, values, and preferences, not as memories, or a body, or a soul. Whether you accept it largely decides whether the proposal reads as liberating or horrifying. For a reader with a substantive account of the person — a Christian one, say — this is the essay's load-bearing assumption, asserted in a sentence and never defended. It is also exactly the claim mind-uploading enthusiasts make.

Open this passage in the original →
The rules, and the temptations

"A GA that never asks is not trustworthy, just unaligned or uncalibrated or incompetent."

Read the three principles slowly: enhancement (amplify the principal or the thing is pointless), mental sovereignty (nothing inside your own tool nudging you toward someone else's ends), and self-actualization (help you become more yourself, not more average). Then the anti-principles — everything that would feel like progress and is not: voice interfaces, low cost, running it yourself, engagement, impressive demos, benchmark scores, brand safety. Every instinct of consumer software is inverted. "A good GA demos badly," he writes, since the only judge who counts is the person it was built to be. The quoted line rules out the opposite failure: an AI that never interrupts has not earned trust, it has just stopped checking.

Open this passage in the original →

Counterpoints

  • The surveillance objection. An AI trained on everything you write and read is the most complete dossier ever assembled about one person. Gwern raises this himself, citing the third-party doctrine — the US principle that data handed to a company loses much of its constitutional protection — and answers with tamper-proof servers and a hope that privacy law modernizes. A hardware bet plus a wish: the failure mode of a perfect assistant is a perfect informant.
  • Feasibility. Continually retraining a large model on one person's life is a research direction, not shipping technology. Gwern concedes the core difficulty — information works better placed in the conversation than trained into the weights — and his fix, having the model annotate its own training data, is untested at this scale.
  • Who actually gets one. The design target is over $1,000 a month, sold first to CEOs and researchers. The great-aunt of the opening cannot buy the thing the opening says she needs; the answer offered is that prices fall fast.
  • Emulation is not alignment. Predicting what you would write is not the same as wanting what you want — see alignment faking. Gwern also names his own biggest gap: nobody is building this, for reasons he calls structural. "Then a startup should" is a bet, not a rebuttal.

Questions to carry

  • If a model were trained on everything you have ever written, what would it get wrong about you — and would you be able to tell?
  • The defense against prompt attacks is an AI with a fixed identity. What happens the first time you genuinely change your mind?
  • The proposal assumes the principal has something worth emulating. What does a guardian angel do for someone who has not — average them, or flatter them?
  • If "you" are your preferences rather than your memories or body, what is lost when a model of you keeps running after you stop?
  • Who are your agents right now, and whose interests do they actually serve?

Go deeper

  • Prompt injection — the attack class the security argument is built to defeat; read it to judge whether hardwiring one user would work.
  • In-context learning — why putting your life into the context window is not the same as the model learning it.
  • Agent memory — the current state of the amnesia problem, and what today's memory features actually are.
  • Sample-efficient personality inference — how much data it takes to model a person's values; the essay bets on kilobits.
  • Scalable oversight — the problem behind the political and military section: keeping humans in charge of faster systems.