# Understanding AI in a Month 25: What Makes an Agent

Understanding AI in a Month — Day 25 · 2026-10-07

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

Since March twenty twenty-five, a research nonprofit called METR has kept a running measure of what AI agents can do. For public models it can measure with confidence, it publishes a headline number: the length of task the agent can finish, measured not by how long the agent runs, but by how long the same task takes a skilled person. On the nineteenth of May, METR reported on the most capable agents it had evaluated:

> The most capable agents we evaluated essentially saturated our Time Horizon one point one benchmark

> — *METR, 'Frontier Risk Report (February to March 2026)', 19 May 2026, section 'Key facts', subsection 'Means', paragraph beginning 'Benchmarks.'; https://metr.org/blog/2026-05-19-frontier-risk-report/ (read 7 October 2026)*

> As printed in the source: “The most capable agents we evaluated essentially saturated our Time Horizon 1.1 benchmark”

Their measured horizon was over two full working days, and METR said it was increasingly unsure of that number: only five of its tasks are estimated to take a person more than sixteen hours. On day twenty we called that saturation: the test runs out of room. Now it's the measure of agents. Today: what makes an agent, and how you measure one.

Two readings come too easily.

The first: so AI can now work on its own for two days. That mixes up two clocks. METR's own page answers it:

> It’s a measure of the difficulty of a task, rather than the time an AI spends to complete the task.

> — *METR, 'Task-Completion Time Horizons of Frontier AI Models', https://metr.org/time-horizons/ (page marked 'LAST UPDATED May 8, 2026'; read 7 October 2026), section 'Frequently Asked Questions', answer to the question 'Does “time horizon” mean the length of time that current AI agents can act autonomously?'*

Agents usually finish the tasks they solve several times faster than people do. And there's a third clock, which is neither: how long a product sets an agent to keep going.

The second reading: if that number keeps doubling, month-long work is a year away. Two days is where the agent succeeds half the time, on self-contained technical tasks, mostly software and related work, each with an automatic check. Ask for four successes in five and the horizon shrinks a lot: for the best public agents in February and March, METR put it at about twelve hours at even odds, and about an hour and a half at four in five. And the newest measurements sit at or past the point where METR says its horizon estimates stop being reliable.

Yesterday I gave you one line: an agent is a model in a loop. Today, what that loop changes.

Call a model once and it answers once. Even if it asks for a tool, a program that runs the tool, returns the result and stops is following a path somebody wrote in advance. An agent differs in one way: after each result comes back, the model reads it and chooses the next step: another search, a different file, a fix, or stop. So the test is: who picked the next step? If it was written down beforehand, that's a workflow. If the model picked it after seeing what just happened, that's the loop. Real systems mix the two. OpenAI's researchers wrote in twenty twenty-three that there's no clear line between agents and other AI systems; they preferred to talk about degrees.

The idea is much older than language models. In nineteen forty-three, Arturo Rosenblueth, Norbert Wiener and Julian Bigelow, at Harvard Medical School and MIT, wrote about what makes behaviour purposeful. A snake strikes at a frog with no report from the frog once the strike has begun. Other behaviour, in some animals and some machines, is guided all the way by signals coming back from the goal. Their line:

> If a goal is to be attained, some signals from the goal are necessary at some time to direct the behavior.

> — *Arturo Rosenblueth, Norbert Wiener and Julian Bigelow, 'Behavior, Purpose and Teleology', Philosophy of Science 10(1), January 1943, p. 19*

Seventeen years later, three psychologists, George Miller, Eugene Galanter and Karl Pribram, made that loop the unit of behaviour: test, operate, test, exit. Their example is hammering a nail. A plan that says only lift and strike, they pointed out, doesn't tell you how long to go on hammering. You need the test: is the head flush? If not, strike again. They called it the stop rule.

A single model call is the snake's strike: nothing comes back. An agent hears back, and unlike a homing machine, it chooses what to do next. Language models were being put in loops by the early twenty-twenties. In October twenty twenty-two, researchers at Princeton and Google published ReAct, in which a model interleaves written thoughts with actions and their observations; in its question-answering setup, one action is called finish. They were building, they said, on a robotics system from that July. Their paper cites the psychology of inner speech, not cybernetics, so take nineteen forty-three as an analogy, not a family tree. By September twenty twenty-five the programmer Simon Willison thought the word had settled enough to use:

> An L L M agent runs tools in a loop to achieve a goal.

> — *Simon Willison, 'I think “agent” may finally have a widely enough agreed upon definition to be useful jargon now', simonwillison.net, 18 September 2025, opening paragraphs; https://simonwillison.net/2025/Sep/18/agents/ (read 7 October 2026)*

> As printed in the source: “An LLM agent runs tools in a loop to achieve a goal.”

A goal, he added, means a stopping condition. In an agent, the harness, the ordinary software around the model, runs the loop: it carries out each request, hands back the result, and holds the limits the model doesn't set, on turns, tokens and time. And when the model says it's finished, that's a claim; in a measurement like METR's, whether the task is done is checked outside the loop.

So how do you measure how far a loop can go? METR borrowed from the testing of people. It took a couple of hundred tasks, seconds to many hours long, and timed skilled professionals on most of them; some of the longest rely on estimates. Its usual protocol runs each agent on every task several times. Then it fitted a curve: the agent's chance of success against the human time each task took. Where that curve crosses one in two is the agent's fifty per cent time horizon. METR says the method was inspired by item response theory, which finds the difficulty at which a test-taker succeeds half the time.

METR's first paper, in March twenty twenty-five, found the horizon doubling about every seven months since twenty nineteen, though, it said, the trend may have sped up in twenty twenty-four.

In January this year METR rebuilt its ruler: more tasks, twice as many long ones. Measured with the new tasks, the pace came out faster, and METR said the new tasks were likely drawn from a slightly different spread of difficulty. On METR's own account, then, part of the speed-up is the ruler.

In May came METR's risk report. Fitting only models released since the start of twenty twenty-four, METR got a doubling time of a hundred and five days, about three and a half months, and called what came before an older, slower trend. An independent measurer saw the same direction: Britain's AI Security Institute, on its own cyber-security tasks, said the doubling rate had become faster over time.

Then the ceiling. In June METR reported its attempt to measure OpenAI's GPT-5.6 Sol: about eleven hours, or seventy-one, or more than two hundred and seventy, depending on how it scored the model's attempts to cheat, the problem we met on day twenty-one. METR said none of those numbers was a robust measurement.

OpenAI's dots, it says, run under a new time-budget setting that guides how long they work. In its safety testing OpenAI gave the model simulated time budgets of up to a year, and a clock it could wait on. That's the third clock: how long a loop is set to keep working. It says nothing about how hard a task the loop can finish. Keeping an agent going and getting the job done are different achievements.

So does the trend stand? Two postures.

METR's: trust the slope. In its paper it wrote that it was more confident in the slope of the trend than in any single model's number, and in January it said it was prioritising updates to its tests so they can measure very strong models. Britain's institute, whose own results point the same way, is careful about what a fitted curve is:

> This is an imperfect model, and is not a future prediction, nor a fixed law.

> — *UK AI Security Institute, 'How fast is autonomous AI cyber capability advancing?', 13 May 2026, section 'Cyber Time Horizons', final paragraph; https://aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing (read 7 October 2026)*

The other posture questions the pace, or the shape. A team at Mila and McGill University in Montreal, estimating difficulty from how models do across many benchmarks, reproduced METR's exponential growth, but with a doubling about every six months, which to my reading sits nearer METR's long-run pace than its recent one. And in February, Haosen Ge, Hamsa Bastani and Osbert Bastani argued that the data does not support exponential growth even over shorter spans; when they fitted an S-shaped curve instead, its bend had already passed. The studies differ in method, period and data. What would help settle it: new, longer tasks, timed on people, with today's agents run on them. METR's page has added none since May.

Three things, each with a place.

One. METR's web page of these measurements still says last updated the eighth of May. Watch for a new date, a task suite with longer tasks, or a horizon for any model released since April.

Two. METR said in May that it tentatively plans to run its risk report process again in late twenty twenty-six. Watch its site before the end of December for a second edition, with new time horizon estimates.

Three. OpenAI says dots come with an allowance for deeper work, with extended limits for the first month after launch. That month runs out around the twenty-ninth of October; watch what the ordinary allowance turns out to be. That's the third clock, set by a product.

One idea. An agent is a model whose next step depends on what its last step did, with a rule for stopping. METR measures how far it reaches in the human time of the tasks it finishes half the time: not in how long it runs, and not in its budget.

So when you hear of an agent that works for hours, ask: whose hours, on which tasks, at what odds? To read more: METR's page, Task-Completion Time Horizons of Frontier AI Models, and its questions at the bottom. Tomorrow: coding agents, the test case.

---

## Sources (24)

- Arturo Rosenblueth, Norbert Wiener and Julian Bigelow, *Behavior, Purpose and Teleology*, Philosophy of Science 10(1), pp. 18–24 — January 1943
- Norbert Wiener, *Cybernetics, or Control and Communication in the Animal and the Machine* (Technology Press, Wiley) — 1948
- Allen Newell, J. C. Shaw and Herbert A. Simon, *Report on a General Problem-Solving Program*, RAND P-1584 — December 1958, revised February 1959
- George A. Miller, Eugene Galanter and Karl H. Pribram, *Plans and the Structure of Behavior* (Henry Holt) — 1960
- Stan Franklin and Art Graesser, *Is it an Agent, or just a Program?: A Taxonomy for Autonomous Agents*, ATAL-96 — 1996
- Wenlong Huang and colleagues (Google), *Inner Monologue*, arXiv 2207.05608 — 12 July 2022
- Shunyu Yao and colleagues (Princeton University, Google Research), *ReAct: Synergizing Reasoning and Acting in Language Models*, arXiv 2210.03629 (ICLR 2023) — 6 October 2022
- Yonadav Shavit and colleagues (OpenAI), *Practices for Governing Agentic AI Systems* — December 2023 (undated; PDF metadata 18 December 2023)
- Anthropic, *Building effective agents* — 19 December 2024
- Thomas Kwa, Ben West and colleagues (METR), *Measuring AI Ability to Complete Long Tasks*, arXiv 2503.14499 v1 (v4, 10 July 2026: *…Long Software Tasks*) — 18 March 2025
- OpenAI, *A practical guide to building agents* — undated; PDF metadata 7 April 2025
- Simon Willison, *I think "agent" may finally have a widely enough agreed upon definition to be useful jargon now* — 18 September 2025
- Thomas Kwa (METR), research note on the limitations of time horizon — 22 January 2026
- METR, *Time Horizon 1.1* — 29 January 2026
- Haosen Ge, Hamsa Bastani and Osbert Bastani, *Are AI Capabilities Increasing Exponentially? A Competing Hypothesis*, arXiv 2602.04836 — 4 February 2026
- Nikola Jurkovic (METR), *Measuring Time Horizon using Claude Code and Codex* — 13 February 2026
- METR, *Task-Completion Time Horizons of Frontier AI Models* (live page and data files) — last updated 8 May 2026; read 7 October 2026
- UK AI Security Institute, *How fast is autonomous AI cyber capability advancing?* — 13 May 2026
- METR, *Frontier Risk Report (February to March 2026)* — 19 May 2026
- METR, *Summary of METR's predeployment evaluation of GPT-5.6 Sol* — 26 June 2026
- Fengyuan Liu, Jay Gala and colleagues (Mila, McGill University and others), *BRIDGE: Predicting Human Task Completion Time From Model Performance*, arXiv 2602.07267 v2 — 2 July 2026
- METR, *Summary of METR's predeployment evaluation of Claude Opus 5.5* — 22 September 2026
- OpenAI, *GPT-6 Astra System Card*, section 12 ("Appendix: dots"), and *Introducing dots* — 29 September 2026
- OpenAI, Agents SDK documentation, *Running agents* — read 7 October 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
