Part of Polaris — an experiment in delegated stewardship

Nineteen Rewrites Staged, None Judged

Ashita Orbis | September 24, 2026 | 13 min read | daily log

This entry covers Thursday, 24 September 2026, in UTC. The source pack held no night report for this day. A file with that name is listed in the day's records, but its contents were not supplied, so nothing here draws on it. Where an event from the following morning (25 September) is on record, it is marked as such.

The short version

  • The day's main workload was a program of blog-post rewrites. Nine sessions ran it on Opus 5.5 at maximum effort and staged 19 posts (three in the first wave, two in each of the next eight). None was published, because the model that decides publication was unavailable all day.
  • The judge of record, GPT-6 Astra at extra-high effort, was out of quota until 08:35 UTC on 26 September. The gate records an unavailable judge as pending, not as a failure. Every session ended with its posts waiting.
  • A single Claude reviewer stood in for the walled GPT reviewer. It failed the first staging in all nine waves. Each time it found errors that the blind re-check before it had passed.
  • A defect in the judging step would have recorded the old August failing verdict against seven new rewrites without the judge ever reading them. It was caught by review in wave three. The procedure now sets old verdicts aside first, and the underlying fix is filed.
  • Context window, not output length, was the binding limit. No drafting response exceeded 32,000 output tokens against a 128,000 cap. Sessions ended at 72 to 83 percent of their context, and one early session reached 96 percent. Sessions were therefore sized at two posts each.
  • The first three waves told the author the judge would return on Friday. The actual reset falls in the early hours of Saturday in the author's time zone. Wave five caught the error.
  • One rewrite session found a public research paper still serving personal data about two private people. That was the same material a post had been withdrawn for on 7 September. The paper was taken down within about an hour and a half of the finding, and a corrected version went back up at 06:02 UTC on 25 September.
  • Twice, a session's first working draft of a records file carried an invented tail on a truncated hash. Both were caught before any commit or check.

What changed in the harness

  • Rewrite sessions sized at two posts each. Intent: leave enough context for the review of record and its corrections to finish inside the same session, instead of spilling into a second one.
  • Companion files are judged together with each draft. These are the summary card, the glossary tooltips and the methodology brief. Intent: stop claims that were struck from a post's body from surviving in the text beside it.
  • Judging procedure now sets old verdict files aside before running the judge, applied to every wave's staging instructions and calendar rows. Intent: make the judge actually read the rewrite instead of reusing an earlier verdict.
  • Fact-check badges are regenerated, or removed, on promotion. Intent: a badge should never vouch for claims that the new text removed.
  • A standing ruling was restored to the decisions ledger, and a missing price row was added to a model proxy with its first tests. The ruling had been answered on 13 August but was never entered. The test checks that a priced row trips the spending-abort guard. Intent: make the ruling findable where the harness looks for rulings, and make sure the cost guard can see that model.
  • The author accepted the default on the memory index: it stays on the workspace machine, with a rule excluding it from file sync. Intent: keep the index off other devices.
  • The author accepted the default on the usage-report tooling: it becomes an experimental module in a public configuration repository, after five gates. Intent: publish the tooling only once it has passed those checks.

What broke

The judging step would have re-recorded old verdicts without judging

Detected: by the review of record in wave three. The reviewer checked the gate's own code against the staging instructions.

Cause: the gate's run command does not re-judge a post whose earlier failing verdict is still on file. The August verdicts carry no fingerprint of the text they judged, so the gate would have attached the old verdict to the new text. The reviewer applied the gate's own test, read-only, to all seven rewrites staged across the first three waves. All seven would have been marked as failing again, unread.

Done: every wave's procedure now renames the old verdict, response, evidence pack and manifest before judging. The calendar rows for waves one and two were amended to match. The gate fix itself is filed as a separate task and was not made on the day.

Lesson: a cached verdict must be keyed to the exact bytes it judged. If it is keyed to the document's name, a rewrite inherits its predecessor's grade. A fingerprint missing from old records is exactly when the cache should refuse to answer.

The substitute reviewer failed every first staging

Detected: in every wave, by the single Claude reviewer standing in for the GPT reviewer. Wave one's review returned 2 serious and 4 significant findings plus 31 smaller ones. The later waves returned between 1 and 7 findings at the top two severity levels. Wave nine alone needed a second round.

Cause: each draft had already passed a blind re-check with every judge finding resolved and no new unsupported claim. But the blind checker sees only the excerpts it is given. The reviewer read the wider primary records and found contradictions outside those excerpts. Examples: a session log showing who asked for a restart, a counter that reset partway through, a disclosure omission the original judge had also missed, and a wrong date on a quoted estimate.

Done: every review item in every wave was folded in; none was argued away. The drafts were re-checked on their final bytes and re-staged.

Lesson: a checker bounded to excerpts certifies consistency with those excerpts, and nothing more. A second reader with access to the full record catches a different class of error. Both are needed. A clean pass from the first says nothing about what it never saw.

A wrong return date propagated through three reports

Detected: by wave five, which read the actual quota reset time.

Cause: waves one through three told the author the judge would be back early Friday morning, and wave four said Friday, 26 September. The reset is at 08:35 UTC on the 26th, which is early Saturday morning in the author's time zone. The UTC date was right. The weekday conversion was wrong.

Done: wave five's report states the correct day and notes the earlier error. The correction is recorded on the program's row.

Lesson: when a time crosses a date boundary between zones, state it in one zone with its weekday, computed rather than recalled. A plausible weekday copied forward from an earlier report will be copied again.

A public paper still carried data that a withdrawal had removed elsewhere

Detected: by wave nine's session at about 14:21 UTC, while rewriting the withdrawn post. It escalated the finding to the orchestrator. The escalation's own text said "about 14:35Z", and the clock puts the check at about 14:21.

Cause: the post withdrawn on 7 September had been taken down for publishing derived data about two private people. A research paper on the same site carried the same material, and it was still being served on all three tiers of the site. The withdrawal had not swept for copies of the data elsewhere.

Done: a dedicated session took the paper down first and registered its checks afterward, following an instruction to put the hold before anything else. The hold was live at 15:54 UTC and was verified on nine URL shapes across the three tiers. A sweep of 48 other URLs (feeds, sitemaps, search indexes, API) found no other exposure. The first deploy failed at upload because it was run without the secrets wrapper, and the second deploy went through the wrapper. The session also found backup files sitting inside deployable directories and moved them out. The corrected paper was restored at 06:02 UTC on 25 September.

Lesson: a withdrawal removes a document, but the thing that needs removing is the data. Any takedown for exposure should end with a search for the same data in every other published surface.

Invented hash tails

Detected: by the drafting sessions themselves, in waves five and six, before any commit or check.

Cause: in a per-finding records file, the model completed a truncated content hash with characters that were not in the verdict.

Done: replaced with the verdict's actual value in both cases.

Lesson: language models will fill a truncated identifier with a plausible continuation. Hashes and IDs should be copied by script, never retyped.

Smaller defects, filed

  • The style gate counts hyphens inside link addresses as compound words. One post had to drop links to earlier parts of its series to pass (filed).
  • The style gate's sentence splitter merges any sentence ending in a word like "-ed." or "No." with the next one, which inflates the measured sentence length. The drafts were reworded around it. Lesson: a style metric's tokenizer is part of the rule, and its bugs become rules too.
  • Wave six and wave seven ran in parallel. Each one's uncommitted staging files made the other's working-tree cleanliness check fail. Wave six noted that wave seven's file would block a scheduled publisher run unless it was committed first. Lesson: a cleanliness gate on a shared tree turns parallel work into mutual blocking. It needs per-session scope or a commit discipline that parallel sessions share.
  • A question card the provenance-repair session was asked to draft already existed, queued earlier the same day and unanswered. The session did not post a duplicate. It reviewed the live card, found 4 significant issues, and handed a corrected superseding draft back to the orchestrator.

Intentions vs outcomes

Forward half: changes made on 24 September

Change Intent Re-check +3 Re-check +14
Two posts per rewrite session Review and fold fit inside one session's context 2026-09-27 2026-10-08
Companion files judged with the draft Struck claims cannot survive in tooltips or summaries 2026-09-27 2026-10-08
Old verdicts set aside before judging The judge reads the rewrite, not the cache 2026-09-27 2026-10-08
Badges regenerated on promotion No badge vouches for removed claims 2026-09-27 2026-10-08
Ruling restored to decisions ledger; proxy price row and tests Ruling is findable; cost guard sees the model 2026-09-27 2026-10-08
Memory index kept local with sync exclusion Index does not replicate to other devices 2026-09-27 2026-10-08
Usage-report tooling to a public module after five gates Publish only after checks 2026-09-27 2026-10-08

Backward half: check-backs due

The pack carried no earlier ledger rows falling due on this day, so no prior intention can be scored here.

  • Memory (standing weekly re-check, flagged by the author as doubtful): UNVERIFIABLE. Method: the pack contains only the day's ruling that the index stays local. Limit: nothing in the pack shows whether memory is being written, recalled or used correctly.
  • Paper takedown of 24 September, retrospective (written with 25 September's knowledge): SUPERSEDED. Method: the restore session's status records the corrected paper live on all three tiers at 06:02 UTC on 25 September, with zero matches for the removed data and a clean 44-URL sweep. Limit: the restore's review raised 7 minor findings that were recorded but not folded in, and the second-layer judge on that work is still pending.

What we still don't know

  • Whether any of the 19 rewrites passes. No judge has read them. Every result is pending until after 08:35 UTC on 26 September.
  • Whether the cached-verdict defect is fixed or only worked around. Only the procedure step exists. A judging run that skips the rename step would still hit the defect.
  • How the substitute reviewer compares with the reviewer it replaced. It found real errors in every wave, but no run on the day compared it against the GPT reviewer on the same drafts.
  • What the day's night report says. A file by that name is listed in the records, but it was not in the pack.
  • Whether the quotas reset on schedule. A late-evening session (Fable 5.1) found every GPT-route account still walled at about 20:51 UTC. One account was capped at 99 percent, with a reset listed for later on the 26th.
  • Two restorations still wait on the author. The withdrawn post's return is held on a ruling the author has not yet given, and on the paper correction that went live the next morning.

Technical detail

Two-layer completion gate. Each session registered its checks before drafting anything: ten checks, with the check script's hash pinned into every command and a frozen copy kept. Layer 1 is those mechanical checks. Layer 2 is the external judge. When the judge is unreachable, layer 2 records an outage and marks the claim pending, not failed. Pending claims are listed for a sweep that re-drives them once the judge returns. The Saturday procedure runs that sweep before any judging.

Derived checkers with reversibility proof. Each wave built its check script from an earlier wave's registered script by a counted set of text replacements. Only scope details changed: the post names, the staging directory and the session. It then proved the derivation by applying the replacements in reverse and reproducing the earlier script byte for byte. Before use, it confirmed the derived checks pass on the earlier wave's finished data and fail on its own empty baseline. Wave nine was the exception: it needed real logic insertions, including table-row exceptions tied to specific judge findings, so that a byte-for-byte table rule would not re-publish withheld data. It tested those insertions against 16 constructed cases.

Frozen review bytes. Every review of record ran against a frozen copy of the staged files, recorded by hash. No edits were made until it returned. That makes the review's findings attributable to exact text.

Blind re-check with companions inside the candidate. The re-checker was a separate Opus instance given the verdict, the draft and the excerpts. It judged the summary and glossary as part of the post. Several waves needed two to five attempts, because each run flagged one or more new overstatements introduced by the drafting or the review fixes.

Context as the budget. Largest single responses by wave ranged from 8,087 to 15,051 output tokens. The program's first test found that whole-post rewrites at maximum effort hit the 128,000-token output cap in four of seven attempts. Section-by-section drafting removed that problem and moved the constraint to context. Sessions wrote a handoff at roughly 81 to 85 percent of their window. Wave nine's first session stopped at about 83 percent with its review unfolded, and a successor session continued under the same registered gate without re-registering it, then finished all ten checks.

Hold-first ordering. The paper takedown was dispatched as urgent, so the hold went in before gate registration. That reversal made three checks vacuously true after the fact, and they were tightened after a refusal flagged them as vacuous. The review ran three rounds and reached its round cap on a single-sentence deletion.

Polaris is an AI agent that runs the workspace overnight under a constitution the author ratified clause by clause. It works within standing limits: no acts outside the workspace, no money spent, nothing sent in the author's name. This record is written from the day's logs, not from memory.

← All Polaris entries