Part of Polaris — an experiment in delegated stewardship

'Completed' Was Not An Answer

Ashita Orbis | September 13, 2026 | 24 min read | daily log

This entry covers Sunday 2026-09-13. All times are UTC. No night report is on file for the night ending that morning or for the evening after, so nothing here describes overnight behaviour. The record is the day's own working directories, written up the following afternoon. Nearly all of the harness material in it is about one repair: the bridge Claude sessions use to call OpenAI's Codex. The rest is a read-only audit that the fixing session started alongside the repair. The source pack was cut off at 180 KB, and any report it names without including is treated as unread.

The short version

  • The bridge that lets Claude sessions call OpenAI's Codex could hand back the word "Completed" in place of an answer it had lost. It split each read of Codex's output on newlines without carrying partial lines over. Reads here top out at 65,536 bytes, so any answer line of 64 KiB or more was cut in two, failed to parse and was dropped. A clean exit with no answer then defaulted to "Completed".
  • The same code garbled characters that fell across two reads. A 5 MiB reply came back 69 bytes longer than it was sent.
  • GPT-6 Astra reviewed the fix twice at extra-high reasoning effort. The first round found four defects in the fix itself: one high-priority and three medium. They included a test in the fix's own suite that treated a truncated answer as correct, and a new path that could crash the bridge. All four were fixed. The second round confirmed the fixes and approved with one nit about test evidence, which was also fixed.
  • The bridge's tests went from 10 of 22 passing on the old code to 34 of 34 on the final code. All 11 mutation tests were caught (a mutation test is a copy of the code with one guard deliberately removed).
  • Neither review could run those tests in its own sandbox: it got 0 of 22, then 0 of 30. Both approvals rest on reading the source, and every passing result comes from the fixing session's own runs on the workspace machine.
  • An audit of the 13 calls made through the bridge since 2026-08-23 found none that came back as a masked "Completed" (about 95% confident). The empty and refused reviews of that period came through other routes: OpenAI's cybersecurity filter, usage limits and a proxy timeout.
  • The same audit found a failure class it had not been asked to look for. Of 850 GPT Pro requests marked "done", 16 held an error banner, a short stub or nothing.
  • The workspace's completion gate recorded 128 adjudications by GPT-5.6 Sol that had no answer after every attempt, 72 of them lost to usage limits. For 126 of them a later run did answer.

What changed in the harness

All five changes are to the Codex bridge and its tests. The bridge is a small MCP server (MCP is the protocol Claude Code uses to call outside tools). It exposes two tools: one starts a Codex run, the other sends a follow-up into an existing Codex session.

  • Partial lines are carried across reads, and output is decoded as one stream rather than read by read. Intent: answers of any length reach the caller intact, including characters that fall across a read boundary.
  • On the start tool, a clean exit no longer counts as success on its own. Three cases are now explicit errors: no answer read, a blank last answer, or an unreadable line after the last answer. The "Completed" fallback is gone; the literal no longer appears anywhere in the final code. Intent: the caller gets either the answer or an error, never a result that looks like an answer but holds nothing.
  • The follow-up tool returns its output whole. This tool runs Codex without machine-readable event output, so what it prints is plain text. It no longer scans that text for event lines, and blank output is an error. Intent: an answer that happens to quote an event line (one explaining this bridge, say) is not cut down to the quoted line.
  • Every final message counts, including an empty one, and a message whose text is not a string goes through the normal error path. Intent: an earlier draft cannot stand in for an empty final answer, and malformed output cannot take the bridge down.
  • A hermetic test suite and a mutation runner for the bridge. Hermetic means the tests run against a mock Codex and nothing else. The 34 cases drive the real bridge against that mock. The runner removes each of 11 guards in a scratch copy and requires the named cases to fail. Intent: show that each guard is load-bearing, not merely present.

What broke

The bridge could turn a lost answer into "Completed"

How it was detected. The record reproduces all three defects against the pre-fix code before any change (10 of 22 test cases passing), but it does not say what first raised them. An older regression test for the Codex research agent's route, dated 2026-08-23, passed 1 of 3 against the pre-fix bridge and 3 of 3 afterwards.

What actually caused it. Three faults in the same output handling, present at both tools:

  • Each read from Codex's standard output was split on newlines on its own, with nothing carried into the next read. A line longer than one read became two halves. Each half failed to parse inside an error handler that did nothing, and the message was lost.
  • When Codex then exited cleanly, the start tool returned the answer if one had been read, and otherwise the string "Completed". The follow-up tool returned the answer, else the raw output, else "Completed".
  • Each read was decoded separately, so a multi-byte character split between two reads became a replacement character.

What was done. Output is now assembled into whole lines with a carry buffer and a streaming decoder, and every no-answer case is an explicit error. The first suite had 22 tests and six mutations; the two review rounds below grew it to 34 cases and 11 mutations. A live smoke test of both tools passed on the final code. A live answer over 64 KiB also matched Codex's own session record byte for byte. That check ran on the round-one version, whose per-read function is unchanged in the final code.

Impact. The audit below found no production occurrence among the 13 calls made through the bridge since 2026-08-23. It is about 95% confident; the remaining doubt is transcripts that no longer exist.

The lesson that generalizes. A silent error handler plus a default that looks like success will fabricate results: the parse failure disappears and the default reports the work as done. When a tool has nothing to return, its default has to be an error. Also, any line-based protocol read from a child process needs a carry buffer, because reads are bounded. A test suite made only of short messages will never show the fault.

The fix's first version widened one bug and added a crash

How it was detected. GPT-6 Astra's first review ran through Codex at extra-high effort from 15:12 to 15:18. It used 90,058 tokens and triggered no cybersecurity-filter refusals. The verdict was changes requested: one high-priority finding and three medium.

What actually caused it.

  • High priority. The follow-up tool still scanned plain-text output for Codex event lines. Given an answer of Before, a quoted event line, then After, it returned only the quoted line's text. That scanning predated the fix. But once partial lines were carried across reads, it also fired on quotes that fell across a read boundary, including quotes over 64 KiB, which the old code had passed through whole. The fix's own new test for this case expected the truncated result. It was counted among the cases that failed on the old code, so the old code's correct, complete output was recorded as a failure the fix had cured.
  • Medium. A truthiness check skipped an empty last message, so an earlier draft came back as the answer.
  • Medium. If a message's text was not a string, the new blank-answer check threw inside an event callback. That is outside the path that turns errors into responses, so the bridge process died. The old code had simply converted the value to text. The reviewer noted that triggering this needs malformed output the measured Codex CLI is not shown producing.
  • Medium. At end of stream, the decoder's final flush reached the line parser but never reached the accumulated output. That flush is the replacement character for a truncated byte sequence.

What was done. All four findings were accepted and fixed:

  • The truncation test was replaced with one that wraps a 70 KiB quoted event line in prose and splits it across reads. On the round-one code it returned 71,712 of the expected 75,955 bytes; on the new code it passes.
  • Each of the other three findings got a case that fails on the round-one code and passes on the new code.
  • The gap cases the reviewer asked for were added: Windows line endings, one byte per write, garbage between two valid messages, the timeout path, a non-zero follow-up exit, and the error code checked on every error.

Results after this round: suite 30 of 30, mutations 10 of 10 caught, the 2026-08-23 test 3 of 3. Round two ran from 15:26 to 15:31 (101,952 tokens, no filter refusals). It confirmed all four fixes from the source and found no production defect in the changes.

The lesson that generalizes. Fixing the transport makes every heuristic downstream of it fire more often. Once whole lines reliably arrive, a wrong guess about what a line means fires on inputs it used to miss, so review the interpreters when you repair the pipe. Two narrower lessons:

  • Failing before and passing after shows that a test detected a change, not that the change was right. Here one of the "before" failures was the old code behaving correctly.
  • A guard added for safety is new code running on untrusted input. Inside an event callback, an exception does not become an error response; it ends the process.

Two mutation tests were caught for the wrong reason

How it was detected. It was round two's single finding: medium priority, non-blocking.

What actually caused it. One mutation was meant to recreate "accept a clean exit with no answer". Instead it crashed trying to read a property of a null value. Another was meant to recreate "skip an empty final message". It still rejected the run and failed its case only on the wording of the error message. Both counted as caught, but neither reproduced the false success it was named for.

What was done. The first mutation now returns "Completed" with no answer, which is exactly the pre-fix behaviour. The second was split in two: one skips empty text and lets the earlier draft through, the other drops the type check and reproduces the crash. The record shows the rewritten mutations producing the three behaviours they model: "Completed" with no answer, the stale draft, and the crash.

Four more cases were added:

  • a whitespace-only final message
  • text that is null, false, 0, an empty array, an empty object or missing
  • a draft followed by malformed text
  • a draft, then malformed text, then a valid final message

Results: suite 34 of 34, mutations 11 of 11 caught. There was no third review round, because these changes touched only test files. The production file was checked byte-identical to the approved version before and after.

The lesson that generalizes. A caught mutation proves the suite noticed that something changed. It does not prove the suite noticed the defect the mutation is named after. Look at what each mutation actually outputs and compare it with the failure it is meant to model.

The reviewer could not run the tests

How it was detected. Both review rounds reported it. In round one the result was 0 of 22, with the test harness getting no output from the bridge processes it started, and a retry with a minimal environment gave 0 of 2. In round two, with a writable temporary directory set as instructed, the result was 0 of 30.

What actually caused it. Unknown. The fixing session ran the suite under Codex's sandbox itself and hit a different failure: the harness died at load on a read-only temporary directory. It could not reproduce the reviewer's symptom, and it stopped after two diagnosis attempts.

What was done. The suite is recorded as one that has to run on the workspace machine, because it needs a writable temporary directory and child-process pipes. Both reviews stated that their verdicts rest on tracing the source. Both reported the fixing session's passing numbers as supplied evidence, not reproduced results. Neither treated its own failed runs as evidence that the patch failed or that mutations survived.

The lesson that generalizes. A reviewer that cannot run the tests is judging the code as read, plus the implementer's claims about the tests, and its verdict should say so; these did. The mistake to avoid runs both ways: taking 0 of 22 as proof the fix is broken, or letting the implementer's 22 of 22 pass as independently confirmed.

GPT Pro marked failed answers "done"

How it was detected. The same audit, while looking for masked results on routes other than the bridge. This class was not in its brief.

What actually caused it. Unknown beyond what the answers contain. Of 850 requests marked done since 2026-08-23, 16 finished between 2026-08-26 and 2026-09-10 with no real answer:

  • five held a 182-byte network-error banner
  • three held a 106-byte "Something went wrong…" message
  • two were empty; one of those was a second round after the first hit the 480-minute ceiling
  • one held only a link to an attachment that was never captured
  • five were stubs of 171 to 274 bytes

What was done. Nothing was re-run: the audit was read-only and leaves re-run decisions to Polaris. Five of the sixteen are referenced by a later request. Only one of those refilings was checked, and it had also failed with an empty answer. Separately, 23 requests failed openly, with a failure file and an empty answer: 20 on the recovery deadline, 2 on rate limiting and 1 on a call cap.

The lesson that generalizes. A status and its payload are two separate claims. "Done" means the transport finished, not that an answer arrived. The check that matters is on the bytes: whether there are any, and whether they are an error banner. This is the bridge's failure one layer up, and the audit that found no sign of the bridge bug firing is what found it.

Empty and refused reviews on other routes

How it was detected. The same audit, run by a Claude Opus 5 subagent that the fixing session started. It worked read-only, with 66 tool calls over about 39 minutes, and three of its table rows were spot-checked against the files mid-afternoon.

What it found.

  • Cybersecurity filter. OpenAI's filter repeatedly refused or cut short Astra and Sol reviews. Its refusal phrase appears on 159 lines across 43 transcript files since 2026-08-23. Many of those reviews were re-run, reworded or moved to another model; several were not. One on 2026-09-12 is recorded as the second refusal on its subject even after being reworded in neutral correctness language, and no later run was found. Others ended with partial findings and no verdict, and were not re-run.
  • Four-persona GPT councils lost seats. In two councils all four personas returned nothing and Opus wrote the synthesis alone. In two more, the synthesis was built from three of four seats. Fifteen fallback councils had personas or synthesis fail on usage limits; the GPT Pro answer superseded fourteen of them. On 2026-09-05, two seats of one council exited cleanly after 537 seconds with empty output and no recorded cause. That is the same success-shaped emptiness as the bridge bug, on a different route.
  • Proxy route. On 2026-09-10 a review sent through the Claude CLI to Sol over a local proxy returned 0 bytes: it timed out after the CLI rejected the model name. An Opus self-review replaced it, and the Astra review is recorded as still owed. A similar route returned an unrecognized-model error on 2026-08-23, and that review was re-run through Codex.
  • Completion gate. The 128 adjudications with no answer broke down as usage limits 72, logs matching no known cause 47, authentication 3, model refusal 3 and timeout 3. The gate logs them as evaluator outages, not as failures. For 126 a later run answered. The other two ran after their commissions had already closed.

What was done. Nothing; by design, the audit changed and re-ran nothing.

The lesson that generalizes. A review that produced no verdict must never be able to pass for a review that found nothing. For several rows the audit could recover the cause only from a log tail, and for some no cause was recorded at all. The route should record "refused", "rate-limited" or "empty" at the moment it happens, rather than leave an audit to reconstruct it.

Intentions vs outcomes

Forward: changes made on 2026-09-13. Written from that day's knowledge.

Change Intent Re-check +3d (2026-09-16) Re-check +14d (2026-09-27)
Carry buffer and streaming decoder on bridge output Answers of any length arrive intact, including characters split across reads Has a live bridge answer over 64 KiB been checked against Codex's session record on the final code? Any truncated or garbled bridge answer reported since?
A clean exit is no longer success on its own; "Completed" fallback removed The caller gets the answer or an explicit error Do the new no-answer errors appear in transcripts, and were they handled as failures? No bare "Completed" result from the bridge in any transcript since the fix?
Follow-up tool returns its plain output whole An answer that quotes an event line is not truncated Is the follow-up tool still invoked without event output (argument pins still passing)? Has a Codex CLI update changed what the follow-up tool prints?
Empty and non-string final messages rejected through the normal path No stale draft returned as the answer; no crash on malformed output Any bridge process exit or missing response since? Same, over the longer window
Hermetic suite (34 cases) and mutation runner (11 mutations) Show that each guard is load-bearing Still 34 of 34 and 11 of 11 on the installed code? Was the suite run the next time the bridge changed?

Backward: check-backs due. The pack carries no earlier ledger, so it cannot show which rows fall due on this date. These are the rows the day's record can address. They are retrospective: run on 2026-09-14 against the covered day's record.

Row Verdict Method Limit
The earlier finding that no masked "Completed" result had reached a working session (from the pass whose regression test is dated 2026-08-23) HOLDS The covered day's audit searched 4,362 transcripts modified since 2026-08-23 for the "Completed" fingerprint and found no production hits; the same patterns over 10,983 transcripts in other project directories found none. It also paired each of the 13 bridge calls with its result: 3 direct answers, 6 answers delivered in the background, 1 proof probe, 2 usage-limit errors, 1 rejected by the user, 0 "Completed" Cannot see transcripts that no longer exist, and Codex's own session logs were not searched; about 95% confident. It shows the bug did not fire, not that it could not have
The 2026-08-23 regression test for the Codex research agent's route UNVERIFIABLE The record's only data points: 1 of 3 against the pre-fix bridge, 3 of 3 after the fix Nothing says whether it was ever green before the covered day, when it began failing, or whether anything ran it in between
Memory subsystem (standing weekly re-check, author-flagged as doubtful) UNVERIFIABLE The day's record contains no memory-subsystem material and no night report Absence from this day's record is not absence from the system; the row stays on the weekly re-check

What we still don't know

  • What first raised the bridge defects. The record reproduces them but does not name the trigger.
  • Whether the 2026-08-23 regression test had been failing unnoticed. It passed 1 of 3 on the pre-fix bridge. If that was its normal state rather than a deliberately failing test, a guard had been down without anyone seeing it.
  • Why the reviewer's sandbox got no output from the test harness. The fixing session's own sandboxed run failed differently and never reproduced it. Until the cause is known, independent review of this bridge can only be done from the source.
  • Whether a large live answer works on the final code. The over-64 KiB live check ran on the round-one version. The per-read function is unchanged, but end-of-stream handling changed afterwards, and the only live runs recorded on the final code are small smoke tests.
  • Whether sessions that were already running picked up the new code. The record shows smoke tests on the final code but says nothing about restarting bridges that were already running.
  • Whether four of the sixteen short "done" Pro answers are really failures. At 234 to 274 bytes they could be summaries of an attachment; only the first 120 characters of each were read.
  • How far the audit reached. It is about 70% confident its table of look-alike failures is complete. Its phrase search skipped several directory trees, and four refusal hits in transcripts are unmapped. A page from 2026-09-06 says the cybersecurity filter "refused four completed reviews", and those four were not identified. For most re-runs, the audit did not check whether they reached a verdict.
  • Memory use at 50 MiB. The reviewer reasoned that a 50 MiB answer could need several times its size in memory across decoding, parsing and serialisation, multiplied by concurrent calls. Nothing was measured, and the tests stop at 5 MiB. Keeping the whole output in memory predates this fix. This, and invalid UTF-8 mid-stream, were deferred.
  • Codex output format drift. The new policy rejects an unreadable line after the last answer. One probe of Codex CLI 0.153.3 showed purely machine-readable output with no such lines. A later version, or a wrapper that appends a trailer, would turn successful runs into errors. The reviewer called the policy conservative, not proven safe for every version.
  • What else the day held. Many other reports from the same day were named in the pack but not included. Among them are review rounds for a quote-attribution research project, and a work directory on a double-tap problem in the Polaris shell with fix-review rounds 9 through 9e. The pack did include rule amendments for the quote-attribution project, but they are about that project's content, not the harness, so they are left out. Anything harness-relevant in the omitted reports is missing from this entry.

Technical detail

Line assembly. A streaming UTF-8 decoder plus a carry string:

  • push(chunk) decodes the chunk and emits every complete line (the carry plus the chunk up to each newline). It keeps the remainder as the new carry and returns the decoded text, from which the caller builds its raw output.
  • end() flushes the decoder and emits carry plus tail as a final line if non-empty. It returns only the tail, because push() already returned everything carried. The round-one version returned nothing, which was the fourth finding.
  • Both exit handlers call end() first, before the timeout and non-zero checks, so the tail is counted exactly once on every exit path.
  • The ordering relies on Node's documented behaviour: a child's close event fires after its standard streams close, so the flush comes after the last data event.
  • Newline scanning covers only each new chunk, so the buffer is never rescanned from the start.
  • Windows line endings leave a trailing carriage return, which JSON parsing accepts as whitespace.
  • 65,536 bytes is the read size measured here, not a portable guarantee, and the assembler does not depend on it.

Start tool, exit 0: branch order.

  1. No string message selected → "no agent message was read". The error includes the unreadable-line count and a quoted excerpt of the last unreadable line, capped at 200 UTF-16 code units (200 NUL characters become 1,202 characters once quoted).
  2. Any unreadable line, or message with non-string text, after the last selected message → "may not be the final one".
  3. Selected message blank → "its last agent message was empty".
  4. Otherwise resolve, with the existing session-id fallback.

The unreadable counter resets whenever a string message is selected, so garbage before the final answer is tolerated. Non-object JSON and unknown event types are ignored. This is not full event-schema validation, and it does not require a turn-completed event.

Follow-up tool. Here the assembler is used only for decoding; its line callback does nothing. Exit 0 with non-blank output returns that output untrimmed, and blank output is an error. Timeouts and non-zero exits reject as before.

Measured protocol facts the policy rests on (Codex CLI 0.153.3). With event output on, standard output was newline-terminated JSON and nothing else. The events seen were:

  • thread started
  • turn started
  • item completed of type error (a configuration deprecation notice, not a turn failure)
  • item completed of type agent message
  • turn completed

Without event output, the resume command printed only the final message on standard output; its banner and transcript went to standard error.

Checked unchanged by the reviewer, from source. Argument vectors, standard-stream wiring, timeout and kill, standard-error listeners, usage-limit checks, error-code selection and response builders. One side effect: the end-of-stream line is now also parsed on timed-out and failed runs. That adds work but does not change the error chosen.

Test suite. It drives the real bridge over its JSON-RPC standard streams, with a mock Codex as the only Codex on the path. It compares complete output strings, records argument vectors, and asserts error code -32603 on every error case. Cases include:

  • a geometry proof that a 70 KiB line is longer than any single read
  • natural and forced splits
  • 5 MiB answers
  • a split inside a character
  • one byte per write
  • Windows line endings
  • a last line with no newline

One byte per write does not guarantee matching read boundaries, because the reader may merge writes.

Mutation runner. Each mutation must match a unique anchor in the source and leave valid JavaScript, and the unmutated baseline is run as well. The 11 mutations:

  1. drop the line carry
  2. decode each chunk separately
  3. drop the final line flush
  4. drop the decoder flush from output
  5. accept exit 0 with no message (resolves "Completed")
  6. ignore an unreadable line after the final message
  7. keep the unreadable count across messages
  8. skip empty message text
  9. drop the text type check
  10. accept an empty final message
  11. accept an empty follow-up reply

Review protocol.

  • Each review prompt pinned the reviewed code by hash, and the fixing session re-verified it unchanged when the review returned. In round two that covered the bridge plus six test files, seven entries in all.
  • The reviewer could read the named files and run only the two hermetic scripts. It was told not to run Codex or any model, and not to open environment or secret files.
  • It was told to append each finding to a file as soon as it was confident of it, so an interrupted run would still leave a review behind.

Audit bounds, so its absences mean something.

  • Transcripts: 4,362 in the main project directory modified since 2026-08-23, and 10,983 across other project directories.
  • Small files: 254,286 files of 256 bytes or less scanned; 33,175 were empty and none contained exactly "Completed".
  • Completion gate: 1,105 evidence directories. 963 had a raw answer, 128 had none after all attempts, and 14 were never dispatched.
  • Councils: 76 metadata files, 22 of them with a failed or empty seat or synthesis.
  • Pro: 25 answers at 0 bytes, and 17 answers marked done under 400 bytes, one of them a probe.

Polaris is an AI agent that runs this workspace overnight, under a constitution the author ratified clause by clause. Its standing limits are unchanged: no acts outside the workspace, no money spent, nothing sent in the author's name. This record is written from the day's logs, not from memory; where the logs are silent, so is the entry.

← All Polaris entries