# Understanding AI in a Month 29: Not Every AI Is a Chatbot

Understanding AI in a Month — Day 29 · 2026-10-11

*Full transcript of the spoken edition. A quoted passage is its source read aloud — the source's own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it.*

---

On the fifth of October, the team that runs OSWorld, a test of whether AI can operate a real computer, updated its official results. Its new version, OSWorld 2.0, built at the University of Hong Kong's XLANG Lab with colleagues elsewhere, is a hundred and eight long jobs of everyday and professional computer work, taking a person about an hour and a half at the median. One means filing an expense claim from a pile of receipts and emails. The system gets a written task, but sees the computer only through screenshots, and acts with mouse clicks and keystrokes. The best run on the full set, by Anthropic's Claude Opus 5, finished forty-four per cent of the jobs outright. Anthropic's Claude is used to produce this show. Scored instead on how many of its checkpoints the computer's final state met, the same run got seventy-eight. Today: what changes when an AI stops being a chatbot.

Two readings come too easily.

The first: AI now uses a computer better than people do. On OSWorld-Verified, the repaired edition of the original twenty twenty-four test, the best listed results now run past ninety per cent, against a human score of about seventy-two from the original study. But the paper says who the humans were:

> computer science major college students who possess basic software usage skills but have not been exposed to the samples or software before

> — *Tianbao Xie and colleagues, 'OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments', arXiv 2404.07972 v2, 30 May 2024, section 3.4 'Human Performance'*

The maintainers call that seventy-two an estimate. The first system to pass it, last December, drew on ten attempts at every task. And for the new, longer test, I couldn't find a published human success rate at all. On day eleven we read OpenAI's launch page for GPT-6 Astra; its computer-use figure, seventy-two point six, was a partial score on an offline subset. Anthropic's page for Claude Opus 5.5 reports eighty-one point eight on OSWorld two point one, also partial, without saying which tasks. Both are the companies' own runs, and neither model has a row on the maintainers' board.

The second reading runs the other way: agents fail more than half the time, so it's hype. Most of these jobs take a person over an hour. And the paper's authors are specific about where agents break:

> These failures are not about basic G U I control or coding.

> — *Mengqi Yuan, Tianbao Xie, Tao Yu and colleagues, 'OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks', arXiv 2606.29537 v2, 13 July 2026, section 1 'Introduction'*

> As printed in the source: “These failures are not about basic GUI control or coding.”

Agents, they write, drop constraints they were given, miss information that arrives mid-task, guess instead of asking, and skip checking their work.

The word for what's new is modality, and it's older than computing. In eighteen seventy-eight, the German scientist Hermann Helmholtz gave a speech called The Facts of Perception. Within one sense, he said, sensations differ in quality: red from blue. Between senses, sight from taste, they differ in what he'd earlier called modality, and there's no bridge between them:

> one cannot ask whether sweet is more like red or more like blue

> — *Hermann Helmholtz, 'The Facts of Perception' (address, 1878), English translation in Selected Writings of Hermann Helmholtz (Wesleyan University Press), as reproduced at marxists.org/reference/subject/philosophy/works/ge/helmholt.htm, read 11 October 2026, opening paragraph of its discussion of the kinds of sensation, beginning 'Among the various kinds of sensations'*

Here's a way to picture what some machine-learning designs have done since, and the picture is mine, not Helmholtz's: they turn pictures, sound and even actions into the same kind of stuff. On day four, a word arrived as a list of numbers. In October twenty twenty, a team at Google cut photographs into squares sixteen pixels a side and fed the squares to a transformer as if they were words:

> a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks

> — *Alexey Dosovitskiy and colleagues (Google Research, Brain Team), 'An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale', arXiv 2010.11929 v1, 22 October 2020, abstract*

That held after training on very large image collections. In twenty twenty-two, Google's AudioLM did the same with sound, turning it into a sequence of tokens. And in July twenty twenty-three, Google DeepMind ran the idea backwards, for a robot arm. In its RT-2 model, each part of a movement is rounded to one of two hundred and fifty-six steps and written as a token:

> we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens

> — *Anthony Brohan and colleagues (Google DeepMind), 'RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control', arXiv 2307.15818 v1, 28 July 2023, abstract*

That's what transfers. Cut the new modality into pieces, give each piece its numbers, and the machinery you already know runs on it: prediction, attention, the loop. On day twenty-five we met the agent loop; in a computer-use agent, what comes back is a screenshot, and what goes out is a click or a keystroke.

What doesn't come along for free sits at the two ends.

The input end: did the answer actually use the new sense? A question about a picture can often be answered from its wording, its options, or what the model already knows. In November twenty twenty-three, researchers released MMMU: eleven and a half thousand college-level questions, each built around an image. Four months later, researchers in China tried models on it with the pictures taken away:

> Visual content is unnecessary for many samples.

> — *Lin Chen and colleagues (University of Science and Technology of China, Chinese University of Hong Kong, Shanghai AI Laboratory), 'Are We on the Right Way for Evaluating Large Vision-Language Models?', arXiv 2403.20330, v1 29 March 2024, v2 9 April 2024, abstract*

In their tests, OpenAI's GPT-4V, given MMMU's questions with the pictures withheld, still scored forty-five per cent; with the pictures, fifty-four. Guessing scores about twenty-two. MMMU's own authors then asked text-only models every question ten times, without the pictures, and dropped the questions most of them got right. One model answered:

> I do not see the image, but the correct sequence based on the standard steps involved in bacteriophage infection is likely to be

> — *Xiang Yue and colleagues, 'MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark', arXiv 2409.02813 v3, 22 May 2025, Figure 2 (a text-only model's answer, printed in the figure)*

and it named the right letter. On day eleven we asked whether a test measures what its name says. A test of seeing has to show that seeing was needed. That's modality-specific evaluation.

The output end: when the output is an action, the world answers back. In January twenty twenty-five, OpenAI ran its computer-use agent, before its safeguards, on a hundred everyday requests and counted thirteen mistakes. Eight were easy to undo.

> The other five mistakes were, to some degree, irreversible or possibly severe

> — *OpenAI, 'Operator System Card', 23 January 2025, section 'Model mistakes' (Internet Archive capture of openai.com/index/operator-system-card/, read 11 October 2026)*

> As printed in the source: “The other 5 mistakes were, to some degree, irreversible or possibly severe”

One was an email sent to the wrong person. Robots meet another limit. RT-2's authors wrote that web data gave their robot no new motions, only new ways to use the ones its robot data had shown it. That's perception and action: reading the world, and changing it.

When OSWorld appeared, in April twenty twenty-four, people managed about seventy-two per cent and the best model about twelve. That model read no screenshots. It read a text description of the screen. From screenshots alone, the best scored under six per cent; the authors said models struggled to turn what they saw into the right place to click. To my reading, that was the input end failing.

Computer-use agents from Anthropic and OpenAI followed, in October twenty twenty-four and January twenty twenty-five. After the maintainers repaired the tasks in July twenty twenty-five, scores climbed past the human line, and in June this year they released the longer test. In their July paper, the best agent finished about a fifth of its jobs. By October, on a revised release of the tasks, the best checked run finished forty-four per cent. The releases differ, so that isn't a clean trend. But the failures moved. As I read the record, the basic clicking got good; keeping hold of a long job, and of a screen that keeps changing, hasn't yet.

Meanwhile, by their makers' documentation this week, OpenAI's newest flagship, Anthropic's models and Google's Gemini three point eight Flash all take images in; only the Gemini also takes sound and video, and all three answer in text.

So how should an agent touch a computer? Two postures.

The first: use the screen, the interface built for human eyes and hands. OpenAI said its agent was trained to work the buttons, menus and text fields people see, just as humans do. That, it said, let it work without connections built for each website or operating system.

The second: build an interface for the agent. In May twenty twenty-four, the Princeton team behind SWE-bench argued that

> L M agents represent a new category of end users

> — *John Yang, Carlos E. Jimenez and colleagues (Princeton University), 'SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering', arXiv 2405.15793 v1, 6 May 2024, section 1 'Introduction'*

> As printed in the source: “LM agents represent a new category of end users”

and gave theirs commands built for it, including a file viewer and an editor. They called that an agent-computer interface, and it made their agent markedly better at its work. The evidence leans both ways. OSWorld's first paper found that a text description of the screen helped some models and misled others. And many of the top verified results come from systems that can also act by writing code, not only by clicking. As I read these results, the screen is one route into a great many programs, and in SWE-agent's coding tests, commands built for the agent did better.

Three things.

One. The OSWorld 2.0 board, last updated the fifth of October. Watch for official rows for GPT-6 Astra and Claude Opus 5.5, whose makers have published only their own runs, and for any Gemini model. And watch for the first run that finishes half of the hundred and eight jobs.

Two. MMMU's leaderboard, whose newest row is dated the first of July. Its top scores on the harder MMMU-Pro are marked as reported by the models' makers. Watch for an independent run of the hardest setting, where the question itself arrives as a photograph.

Three. The Conference on Robot Learning, from the ninth to the eleventh of November. Watch for robot results that say how many trials were run, on real robots or in simulation, and who ran them.

One idea. The systems in this story that aren't chatbots still read pieces: patches of a picture, stretches of sound, screenshots of a desktop, steps of a robot arm. As I read it, that design carried over from text; the checks have to be rebuilt for each new sense and each kind of action. So when you hear that a model sees, ask whether the test showed it needed the picture. When you hear that it uses a computer or moves a robot, ask what checked what the action actually did. To read more: the OSWorld 2.0 paper, and the encyclopedia's articles on MMMU and OSWorld. Tomorrow: how to keep up without drowning.

---

## Sources (24)

- Hermann Helmholtz, *The Facts of Perception*, address of 1878, English translation in *Selected Writings of Hermann Helmholtz* (Wesleyan University Press), reproduced at marxists.org — 1878; read 11 October 2026
- Seymour Papert, *The Summer Vision Project*, MIT Artificial Intelligence Group, Vision Memo No. 100 (MIT DSpace) — 7 July 1966
- Alexey Dosovitskiy and colleagues (Google), *An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale*, arXiv 2010.11929 — 22 October 2020
- Zalán Borsos and colleagues (Google), *AudioLM: a Language Modeling Approach to Audio Generation*, arXiv 2209.03143 — 7 September 2022
- Anthony Brohan and colleagues (Google), *RT-1: Robotics Transformer for Real-World Control at Scale*, arXiv 2212.06817 — 13 December 2022
- Anthony Brohan and colleagues (Google DeepMind), *RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control*, arXiv 2307.15818, abstract and section 3.2 — 28 July 2023
- Xiang Yue and colleagues, *MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI*, arXiv 2311.16502 — 27 November 2023
- Lin Chen and colleagues, *Are We on the Right Way for Evaluating Large Vision-Language Models?*, arXiv 2403.20330, abstract — 29 March 2024
- Tianbao Xie and colleagues, *OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments*, arXiv 2404.07972 — 11 April 2024
- John Yang, Carlos E. Jimenez and colleagues (Princeton University), *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering*, arXiv 2405.15793 — 6 May 2024
- OpenAI, *Hello GPT-4o* (Internet Archive capture) — 13 May 2024
- Xiang Yue and colleagues, *MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark*, arXiv 2409.02813, sections 2 and 3 and Figure 2 — 4 September 2024
- Anthropic, *Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku*; *Developing a computer use model* — 22 October 2024
- OpenAI, *Computer-Using Agent*; *Operator System Card* (Internet Archive captures) — 23 January 2025
- Conference on Robot Learning 2026, programme page (corl.org) — main conference 9–11 November 2026; read 11 October 2026
- Pranav Atreya and colleagues, *RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies*, arXiv 2506.18123 — 22 June 2025
- XLANG Lab, *OSWorld-Verified* announcement; OSWorld-Verified results file — 28 July 2025; read 11 October 2026
- Mengqi Yuan, Tianbao Xie, Tao Yu and colleagues, *OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks*, arXiv 2606.29537 — 28 June 2026; revised 13 July 2026
- Google DeepMind, *Gemini Robotics 2 brings whole-body intelligence to robots* — 30 July 2026
- OpenAI, *GPT-6 Astra* launch page (Internet Archive capture) — 3 September 2026
- Anthropic, *Introducing Claude Opus 5.5* (Internet Archive capture) — 22 September 2026
- OSWorld 2.0 official results (osworld-v2.xlang.ai) — updated 5 October 2026; read 11 October 2026
- MMMU leaderboard data (MMMU project site); Artificial Analysis, MMMU-Pro results and methodology — read 11 October 2026
- OpenAI developer documentation, gpt-6-astra model page; Claude documentation, models overview; Google Gemini API documentation, gemini-3.8-flash model page — read 11 October 2026

Understanding Machine, an Ashita Orbis publication. The written edition of this episode, with its figures and its sources, is published beside it.
