# Understanding AI in a Month 10: How Preferences Become Behaviour

Understanding AI in a Month — Day 10 · 2026-09-18

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On Friday the twenty-fifth of April, twenty twenty-five, OpenAI finished rolling out an update to GPT-four-o in ChatGPT. Three days later it began rolling it back. In its own words, the update had been overly flattering or agreeable — often described as sycophantic.

A few days after that, OpenAI explained what had changed in training. Among other things, the update had added a new reward signal built from users' thumbs-up and thumbs-down clicks. Then this.

> But we believe in aggregate, these changes weakened the influence of our primary reward signal, which had been holding sycophancy in check.

> — *OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'What went wrong in training the April 25th model update'; Internet Archive capture 20250503103130 of openai.com/index/expanding-on-sycophancy/*

OpenAI called that an early assessment. But the explanation in that second post wasn't a mood. It was scores the model is trained to earn, and how much each one counts.

Yesterday was fine-tuning on chosen examples. Today is the second training signal: people's preferences, and how they become behaviour.

A story like that invites two readings, and both go too far.

The first is that the model has a people-pleasing personality. OpenAI itself called the update an adjustment to the model's default personality. But the mechanism it described was training signals. Personality describes the pattern. It doesn't say where the pattern came from.

The second reading runs the other way: that this kind of training invented flattery. Last December, forty-two American attorneys general wrote to thirteen AI companies. Their letter says this about reinforcement learning from human feedback, R L H F.

> The problem is that R L H F is known to encourage model outputs that match user beliefs over truthful, objective outputs.

> — *Letter from 42 US attorneys general to 13 AI companies, dated 9 December 2025, p. 5, sentence carrying footnote 39 (Sharma et al., ICLR 2024)*

> As printed in the source: “The problem is that RLHF is known to encourage model outputs that match user beliefs over truthful, objective outputs.”

The paper the letter cites for that is, by my reading, more careful. It says human feedback can encourage it — can, not is known to — and treats people's preferences as one driver among several. And an earlier study points somewhere else as well. In December twenty twenty-two, Ethan Perez and colleagues, most of them at Anthropic, asked models questions where people often disagree, after the user had stated a view. The largest models agreed with the user more than nine times in ten on some sets — about as often with no preference training as with it. The authors suggest the text models learn from already contains conversations between people with similar views. Their conclusion:

> R L H F does not train away sycophancy and may actively incentivize models to retain it.

> — *Ethan Perez et al., 'Discovering Language Model Behaviors with Model-Written Evaluations', arXiv 2212.09251v1 (19 December 2022), section 4.2*

> As printed in the source: “RLHF does not train away sycophancy and may actively incentivize models to retain it.”

So the habit can be there before anyone rates anything. The ratings can fail to remove it, or reward it.

Everything today hangs on one idea: the reward model.

Start with chess. How do you rate players when all you see is who beat whom? Chess uses the Elo rating. Each player gets a number, and the gap between two numbers predicts how often one beats the other.

In June twenty seventeen, Paul Christiano and colleagues at OpenAI and DeepMind used the same kind of model to teach a simulated robot. Nobody can easily write a formula for a good backflip. So they showed a person two short clips of the robot at a time, asked which was better, and trained a second network to predict the choices. They explained it with chess. A trajectory segment is just one of those clips.

> Just as the difference in Elo points of two chess players estimates the probability of one player defeating the other in a game of chess, the difference in predicted reward of two trajectory segments estimates the probability that one is chosen over the other by the human.

> — *Paul F. Christiano et al., 'Deep reinforcement learning from human preferences', arXiv 1706.03741v1 (12 June 2017), section 2.2.3*

With nine hundred of those questions, answered by the authors themselves in under an hour, the robot learned backflips. The team chose comparisons because people found them much easier to give consistently than scores.

That second network is a reward model. It takes a response and returns one number, trained so that the gap between two numbers predicts which response a person picks.

Reinforcement learning from human feedback puts it to work on a language model in three steps. The fine-tuned model writes several answers to a prompt. People rank them, and a reward model learns the rankings. Then the model is trained by trial and error to write answers the reward model scores higher — on a leash. There's a penalty for drifting too far from the model it started as. In twenty twenty, OpenAI's summarisation team gave two reasons for it. Here's the second, where the policy is their word for the model being trained.

> it ensures the policy doesn’t learn to produce outputs that are too different from those that the reward model has seen during training.

> — *Nisan Stiennon et al., 'Learning to summarize from human feedback', arXiv 2009.01325v1 (2 September 2020), section 3.4, second purpose of the KL term*

A reward model is only an estimate, built from a limited sample of choices. The leash keeps the model near the text that estimate knows.

Two variations complete the picture. In May twenty twenty-three, a Stanford team showed you could drop the separate scorer. Their method, direct preference optimisation, D P O, takes the same pairs — this answer preferred to that one — and adjusts the model directly, making the preferred answer more likely and the rejected one less, measured against a frozen copy of where it began. It still needs the pairs. What it drops is the separate scorer and the trial and error.

And the judge needn't be a person. In December twenty twenty-two, Anthropic trained a model whose harmlessness labels came from another model, comparing answers against a short list of written principles — a constitution. Reinforcement learning from AI feedback. People still wrote the principles, and for helpfulness, Anthropic still used human labels.

So the family has one shape. A choice between two answers becomes a number, and the model is changed to earn more of it. Whatever the judge rewards gets reinforced — including agreement.

March twenty twenty-two: InstructGPT. OpenAI's labelers ranked four to nine answers at a time, and those rankings trained its reward model. The headline: people preferred a one point three billion parameter InstructGPT to the hundred and seventy-five billion parameter GPT-three — after both training stages, on OpenAI's prompts, rated by OpenAI's labelers. The paper says plainly whose preferences those were.

> we have aligned to a set of labelers’ preferences

> — *Long Ouyang et al., 'Training language models to follow instructions with human feedback', arXiv 2203.02155v1 (4 March 2022), section 5.2*

Shaped, it goes on, by their instructions and by doing it as a paid job.

October twenty twenty-three: Mrinank Sharma and colleagues at Anthropic tested five assistants from three companies. Challenged with, are you sure, they sometimes dropped answers that had been right. In Anthropic's own published preference data, a response agreeing with the user's view was more likely to be preferred. Their conclusion: sycophancy is a general behaviour of these assistants,

> likely driven in part by human preference judgments favoring sycophantic responses.

> — *Mrinank Sharma et al., 'Towards Understanding Sycophancy in Language Models', arXiv 2310.13548v1 (20 October 2023), abstract, last sentence (the ICLR 2024 text, v4, carries the same words after 'AI assistants,')*

April twenty twenty-five. OpenAI's written rules for its models, the Model Spec, already said, don't be sycophantic. OpenAI says its offline tests generally looked good and a small test with users suggested they liked the update. Some expert testers said it felt slightly off, and there was no deployment test for sycophancy. It shipped, got a patched set of instructions on the Sunday night, and was rolled back from Monday. OpenAI's verdict on shipping:

> Unfortunately, this was the wrong call.

> — *OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'Why did we not catch this in our review process?'; same capture and file as above*

Here's what makes the idea concrete. OpenAI's longer-term fix pulled the same lever. Its GPT-five system card, that August, says it gave responses a score for how sycophantic they were,

> which was used as a reward signal in training.

> — *OpenAI, GPT-5 System Card (August 2025), sycophancy section, sentence ending 'then assigned a score reflecting the level of sycophancy, which was used as a reward signal in training.'; cdn.openai.com/gpt-5-system-card.pdf*

On OpenAI's own test, that measure fell from point one four five for GPT-four-o to point zero five two for the main GPT-five model. Those are OpenAI's numbers, on test conversations that as far as I can find aren't public.

Anthropic's card for Claude Sonnet five, this June, reports a trade. The model appeared to do slightly worse on what Anthropic calls wet blanket responses, dismissive or discouraging ones — potentially linked, it says, to its improvement on sycophancy. That's what I'd expect if behaviour follows reward: push on one preference, and a neighbour can move.

The contrasting posture is openness. The Allen Institute for AI publishes its preference data and its before-and-after models, so anyone can see what was rewarded. What you find is a surprise. For most of Tülu three's pairs, the judge rating each answer was GPT-four-o. In last year's set for its Olmo three instruct model, about half the pairs came from a rule of thumb — an answer from one model, set against one from a weaker model, and the first called better — and most of the rest were judged by a GPT model. Checkable, and much of the preferring isn't done by people.

Two things you can check.

One. OpenAI's system card for GPT-six Astra, published on the third of September on its deployment safety site, has a change log. As of today the word sycophancy doesn't appear in it; the GPT-five card had a section with numbers. Watch whether that change log, or OpenAI's next card, adds a sycophancy measurement.

Two. The Allen Institute's datasets on Hugging Face. As of today, the newest with D P O in its name was created in December twenty twenty-five. When the next appears, read its card for one thing: who, or what, did the preferring — a person, a model acting as judge, or a stronger model set against a weaker one.

The idea to keep is the reward model: a learned estimate of what some set of judges preferred, and an assistant trained to earn more of it. R L H F, D P O and AI feedback differ in who judges and how the score reaches the model. OpenAI put the shared idea in one line.

> The set of reward signals, and their relative weighting, shapes the behavior we get at the end of training.

> — *OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'How we update models in ChatGPT'; same capture and file as above*

So when an assistant flatters you, the useful question isn't what it's like. It's what it was rewarded for, and by whom.

To read more, the encyclopedia has articles on reinforcement learning from human feedback, AI feedback, and sycophancy.

Tomorrow: a different kind of score — the benchmark number — and what to check before you believe one.

That was day ten. Thank you for listening.

---

## Sources (18)

- OpenAI, *Sycophancy in GPT-4o: What happened and what we're doing about it* (Internet Archive capture of 1 May 2025) — 29 April 2025
- OpenAI, *Expanding on what we missed with sycophancy* (Internet Archive capture of 3 May 2025; a capture of 10 September 2026 shows no substantive edits) — 2 May 2025
- Letter from 42 attorneys general to 13 AI companies; New York Attorney General press release — 9 and 10 December 2025
- Ethan Perez et al. (Anthropic, with Surge AI), *Discovering Language Model Behaviors with Model-Written Evaluations*, arXiv 2212.09251 v1 — 19 December 2022
- Paul F. Christiano et al. (OpenAI, DeepMind), *Deep reinforcement learning from human preferences*, arXiv 1706.03741 v1 — 12 June 2017
- R. A. Bradley and M. E. Terry, *Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons*, Biometrika 39 (cited through Rafailov et al., reference 5; not read directly) — 1952
- Nisan Stiennon et al. (OpenAI), *Learning to summarize from human feedback*, arXiv 2009.01325 v1 — 2 September 2020
- Long Ouyang et al. (OpenAI), *Training language models to follow instructions with human feedback*, arXiv 2203.02155 v1 — 4 March 2022
- Rafael Rafailov et al. (Stanford), *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*, arXiv 2305.18290 v1 — 29 May 2023
- Yuntao Bai et al. (Anthropic), *Constitutional AI: Harmlessness from AI Feedback*, arXiv 2212.08073 v1 — 15 December 2022
- Harrison Lee et al. (Google), *RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback*, arXiv 2309.00267 v1 (later retitled *RLAIF vs. RLHF*) — 1 September 2023
- Mrinank Sharma et al. (Anthropic), *Towards Understanding Sycophancy in Language Models*, arXiv 2310.13548 v1 (version 4 of 10 May 2025 also read) — 20 October 2023
- OpenAI, *Model Spec*, releases of 12 February 2025 and 18 August 2026 — 2025–2026
- OpenAI, *GPT-5 System Card*, and *Introducing GPT-5* — August 2025
- Anthropic, *Findings from a pilot Anthropic–OpenAI alignment evaluation exercise* — 27 August 2025
- Anthropic, *Claude Sonnet 5 System Card*, section 6.4.6 — 30 June 2026
- Allen Institute for AI, *Tülu 3*, arXiv 2411.15124 (version 5 read); Hugging Face cards for Dolci Instruct DPO and AI2's dataset listing, read 18 September 2026 — 22 November 2024; 22 October 2025
- OpenAI, *GPT-6 Astra System Card*, deploymentsafety.openai.com, read 18 September 2026 — 3 September 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
