# Understanding AI in a Month 9: From Base Model to Assistant

Understanding AI in a Month — Day 9 · 2026-09-17

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it.*

---

In January twenty twenty-two, OpenAI published a post with side-by-side examples: one request, two models. One request was, explain the moon landing to a six year old in a few sentences.

The first was GPT-three as it came out of its original training. It wrote this.

> Explain the theory of gravity to a 6 year old. Explain the theory of relativity to a 6 year old in a few sentences. Explain the big bang theory to a 6 year old. Explain evolution to a 6 year old.

> — *OpenAI, 'Aligning Language Models to Follow Instructions' (27 January 2022), GPT-3 completion to the prompt 'Explain the moon landing to a 6 year old in a few sentences.'; Internet Archive capture 20230101110545 of openai.com/blog/instruction-following/*

The second model, which OpenAI called InstructGPT, simply answered: people went to the moon, took pictures of what they saw, and sent them back to earth.

The post doesn't say how those examples were picked. In the paper behind it, for a similar pair about some code, the authors say the untuned model does answer about half the time. So this isn't a model that can't answer. It's a model doing what it was trained to do: continue a document.

And on the twenty-eighth of September, eleven days from now, OpenAI is scheduled to shut down davinci zero zero two and babbage zero zero two, which its own documentation calls base models, not trained with instruction following. How a continuation machine becomes an assistant is today's subject.

The paper behind those examples, by Long Ouyang and colleagues at OpenAI, opens with this sentence.

> Making language models bigger does not inherently make them better at following a user’s intent.

> — *Long Ouyang et al. (OpenAI), 'Training language models to follow instructions with human feedback', arXiv 2203.02155v1 (4 March 2022), abstract, first sentence*

That invites two readings, and both go too far.

The first: the original model is the unfinished version, and the extra training is where it learns things. The sizes point the other way. OpenAI's supervised stage used about thirteen thousand prompts, most of them written by the contractors it hired, and the company's post says the whole procedure used less than two percent of the computing and data of pre-training. OpenAI's own gloss is that the training unlocks abilities GPT-three already had — but it introduces that as, in its words, one way of thinking about this process. What it really adds is argued over, and I'll come to that.

The second reading is the mirror image: an assistant is just the original model with a clever prompt in front. That's half right. In the same post, one prompt was laid out like a page of questions and answers — two answered, a third left open — and the untuned model answered the third sensibly. But the assistant you use isn't a base model plus a prompt. Its numbers have been changed by more training.

Two ideas: supervised fine-tuning, and instruction tuning.

A base model is what pre-training produces: a system that, given text, scores what comes next. Give it a request and, in effect, it continues whatever kind of document that request usually begins.

Training in two stages is older than any chatbot. In November twenty fifteen, Andrew Dai and Quoc Le at Google trained networks on unlabelled text — one way by predicting what comes next, another by rebuilding the input — then used what they had learned as the starting point for ordinary tasks with labelled answers. They wrote this.

> These two algorithms can be used as a "pretraining" step for a later supervised sequence learning algorithm.

> — *Andrew M. Dai and Quoc V. Le (Google), 'Semi-supervised Sequence Learning', arXiv 1511.01432 (4 November 2015), abstract; https://arxiv.org/abs/1511.01432, read 17 September 2026*

In twenty eighteen, OpenAI's first GPT paper did the same with a transformer: pre-train once, then fine-tune separately for each task. Then GPT-two, in twenty nineteen, was tested with no fine-tuning at all, so a request had to be disguised as a document. To get a summary, the team put the letters T L semicolon D R and a colon after a news article — the internet's too long, didn't read — and let the model carry on. It worked, barely. OpenAI's paper says the summaries only just beat picking three random sentences from the article, and scored about six points worse without that hint.

Now the first idea. Supervised fine-tuning means you keep training the same model, with the same objective — predict the next token — but on a much smaller set of examples somebody chose, each one a request followed by the response they want. Nothing new is bolted on. What changes is the kind of document the model has learned to expect. After enough of those examples, the likeliest thing to follow a request is an answer.

The second idea makes that habit general. In September twenty twenty-one, a Google Research team led by Jason Wei described instruction tuning as

> finetuning language models on a collection of tasks described via instructions

> — *Jason Wei et al. (Google Research), 'Finetuned Language Models Are Zero-Shot Learners', arXiv 2109.01652v1 (3 September 2021), abstract, the parenthetical definition of instruction tuning*

They took more than sixty existing research datasets, wrote plain-language instructions for each, tuned a large model on most of them, and tested it on kinds of task it hadn't been tuned on. It did substantially better on those than the same model untuned. The paper adds a condition worth keeping: in their experiments, for models of eight billion parameters and smaller, instruction tuning made the unseen tasks worse.

A conversation reaches a model as one run of tokens, with special markers where each speaker's turn begins. The way I'd put it, fine-tuning is what gives those markers meaning: the training conversations share one layout, so the model learns what comes after the marker that opens the assistant's turn.

Twenty twenty-one: instruction tuning on research tasks — a dataset called Natural Instructions in April, Google's FLAN in September, a model called T-zero in October.

January and March twenty twenty-two: InstructGPT. OpenAI hired about forty contractors, collected their written demonstrations, and fine-tuned GPT-three on them. Then came a second stage, built on rankings of its answers — that's tomorrow. OpenAI also tuned GPT-three on the FLAN and T-zero research collections, and on OpenAI's own mix of prompts, those versions did slightly worse than the one tuned on its contractors' demonstrations. That's OpenAI's measurement, on OpenAI's prompts, judged by OpenAI's contractors, and as far as I can find the prompts were never published, so nobody outside can rerun it.

November thirtieth, twenty twenty-two: ChatGPT. OpenAI's announcement describes the first step.

> We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides

> — *OpenAI, 'ChatGPT: Optimizing Language Models for Dialogue' (30 November 2022), Methods section; Internet Archive capture 20230101000602 of openai.com/blog/chatgpt/. The source sentence continues '—the user and an AI assistant.'*

Both sides played by people, who could draw on suggestions the model wrote: supervised fine-tuning, in the shape of a chat.

November twenty twenty-four: the Allen Institute for AI published Tülu three, a full post-training recipe released with its data and code, built on Meta's base models. Its supervised stage used about nine hundred and forty thousand prompts. The authors' reason for publishing all of it: the data and recipes for post-training are, in their words, the portion with the least transparency.

And this year, some developers still publish both halves. On Hugging Face, DeepSeek's base and chat repositories for its V-four Pro model were created on the same April day, and Google's Gemma four comes in pre-trained and instruction-tuned versions. Several of this year's other big open releases — from Qwen, Moonshot and Z-dot-A-I — have no public base version there as of this week.

So what does this stage actually do? Two positions from twenty twenty-three.

The first came from Meta. Chunting Zhou and colleagues fine-tuned a large base model on one thousand carefully chosen examples, called it LIMA, and proposed what they named the Superficial Alignment Hypothesis.

> A model’s knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users.

> — *Chunting Zhou et al. (Meta AI et al.), 'LIMA: Less Is More for Alignment', arXiv 2305.11206v1 (18 May 2023), section 2 'Alignment Data', definition of the Superficial Alignment Hypothesis (preceded by 'We define the Superficial Alignment Hypothesis:')*

A hypothesis, in their word. Later that year a team from the Allen Institute and the University of Washington checked it another way. They took answers written by tuned models and asked, token by token, which token the matching base model would have ranked first at that point. Across the three pairs they measured, it was the same token about four times in five, and the biggest shifts were in words like hello, thank, however and remember. That's the base model reading the tuned model's answer as it goes, not writing it alone.

The second position is a warning. In May twenty twenty-three, a Berkeley team fine-tuned open models on ChatGPT's answers. Crowd raters found the results competitive with ChatGPT. More targeted automatic tests found the imitators closed little or none of the gap, on tasks their training data didn't cover well. Their explanation:

> imitation models are adept at mimicking ChatGPT’s style but not its factuality

> — *Arnav Gudibande et al. (UC Berkeley), 'The False Promise of Imitating Proprietary LLMs', arXiv 2305.15717v1 (25 May 2023), abstract (the sentence begins 'We show that these performance discrepancies may slip past human raters because')*

The two positions agree on the mechanism: this stage moves style easily. They disagree about what that's worth. For the LIMA team, a small, careful set of examples on a strong base model is enough. For the Berkeley team, that ease is the danger, because raters can mistake the style for the substance, and they argued the better lever is a better base model.

Two things you can check.

One. OpenAI's deprecations page lists the twenty-eighth of September twenty twenty-six as the shutdown date for davinci zero zero two and babbage zero zero two. After that date, look at whether their model pages are still up, and whether anything left in OpenAI's model list is described as not trained with instruction following.

Two. DeepSeek's V-four-point-one Flash, whose Hugging Face page was created on the tenth of September. Its model card reports test results for a base version, but as of today there's no public repository for that base model. Watch whether one appears: it's one sign of whether new open models can still be met before their fine-tuning.

The idea to keep is supervised fine-tuning: same objective, different text. A base model continues documents. An assistant is that same kind of machine, trained further on one particular kind of document — a request, followed by a helpful answer.

So when an assistant surprises you, ask which stage that came from: what it read in pre-training, or what it was shown afterwards.

To read more, the encyclopedia's article on reinforcement learning from human feedback describes this as the first of its three stages.

Tomorrow: the second training signal — people's rankings.

That was day nine. Thank you for listening.

---

## Sources (18)

- OpenAI, *Aligning Language Models to Follow Instructions* (Internet Archive capture of 1 January 2023) — 27 January 2022
- Long Ouyang et al. (OpenAI), *Training language models to follow instructions with human feedback*, arXiv 2203.02155 v1 — 4 March 2022
- OpenAI, API deprecations page; davinci-002 and babbage-002 model pages — read 17 September 2026
- Hyung Won Chung et al. (Google), *Scaling Instruction-Finetuned Language Models*, arXiv 2210.11416 (version 5 of 6 December 2022 read) — first posted 20 October 2022
- Andrew M. Dai and Quoc V. Le (Google), *Semi-supervised Sequence Learning*, arXiv 1511.01432 — 4 November 2015
- Alec Radford et al. (OpenAI), *Improving Language Understanding by Generative Pre-Training* — 2018
- Alec Radford et al. (OpenAI), *Language Models are Unsupervised Multitask Learners* — 2019
- Jason Wei et al. (Google Research), *Finetuned Language Models Are Zero-Shot Learners*, arXiv 2109.01652 v1, and ICLR 2022 version — 3 September 2021
- Swaroop Mishra et al., *Natural Instructions: Benchmarking Generalization to New Tasks from Natural Language Instructions*, arXiv 2104.08773 v1 (later retitled *Cross-Task Generalization via Natural Language Crowdsourcing Instructions*) — 18 April 2021
- Victor Sanh et al., *Multitask Prompted Training Enables Zero-Shot Task Generalization*, arXiv 2110.08207 v1 — 15 October 2021
- OpenAI, *ChatGPT: Optimizing Language Models for Dialogue* (Internet Archive capture of 1 January 2023) — 30 November 2022
- Stanford CRFM, *Alpaca*, and its training code (read 17 September 2026) — 13 March 2023
- Allen Institute for AI, open-instruct training code (read 17 September 2026) — 2026
- Nathan Lambert et al. (AI2), *Tülu 3*, arXiv 2411.15124 (version 5 of 14 April 2025 read) — first posted 22 November 2024
- Chunting Zhou et al. (Meta AI and others), *LIMA: Less Is More for Alignment*, arXiv 2305.11206 v1 — 18 May 2023
- Arnav Gudibande et al. (UC Berkeley), *The False Promise of Imitating Proprietary LLMs*, arXiv 2305.15717 v1 — 25 May 2023
- Bill Yuchen Lin et al. (AI2, University of Washington), *The Unlocking Spell on Base LLMs*, arXiv 2312.01552 v1 — 4 December 2023
- Hugging Face repositories and model cards: DeepSeek, Google, Qwen, Moonshot, Z.ai, Allen Institute for AI, OpenAI — read 17 September 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
