Part of Polaris — an experiment in delegated stewardship

Four Gates Waiting on One Absent Judge

Ashita Orbis | September 23, 2026 | 20 min read | daily log

This entry covers 23 September 2026. No night report is on file for the night ending that morning or the next, so the entry is built from the status and report files the workers wrote. Times are UTC, as the record stamps them. The source pack's day boundary is not clean. It includes two owner rulings stamped in the first hours of the 23rd and two writing sessions stamped 03:05–05:39 on the 24th, which is the evening of the 23rd in the workspace's local time (UTC−7, by the record's own conversions). Both are included here and marked where they appear. The twelve explorable pages the workspace built that day belong to the Pulse feed and are left out, except where one of them measured the harness itself.

The short version

  • The outside judge was out of quota all day. Every job here ends at a "completion gate": first the worker's own registered checks, then an independent evaluator (GPT-6 Astra at extra-high reasoning) that reads the claim. That evaluator was out of quota until 26 September 08:35. Four gate runs in the record passed every check and then ended OUTAGE. That is recorded as pending, not failed. A Claude stand-in was refused by design.
  • Review fell to a single-pass Claude code-reviewer, and it still caught real faults. In one case a change would have broken a live public site's deploy path three ways. In another, a suppression pattern would have silently hidden future legitimate questions to the author.
  • A test run as a baseline modified the live checkout, and the day's public changelog refused to publish. It was refused at 16:16, repaired, then run once by hand and published at 16:58. The test has been retired.
  • Deploy scripts now refuse to run from anything but the primary checkout. The fence landed on the main branch and in 8 branch worktrees (extra working copies of the repository). The 7 worktrees that could not take it were archived and removed: 13 dirty files, plus 31 unique ignored files totalling 233 MB.
  • A second fence was built, went live, and was rolled back after about fifteen minutes. It was meant to stop the credential vault from serving a stale checkout. It was live 17:36:51–17:52:04, then removed. No real caller was refused in that window. It is now parked for a decision.
  • 27 of 227 unposted questions are still owed to the author. The 227 were questions drafted by workers that never reached the author's queue. Of the 27, 21 are to review and post and 6 to withdraw. One was drafted 34 days earlier. Nothing was posted.
  • Three posts filled 95.7% of a million-token context. A writing session rewriting three blog posts peaked at 957,156 tokens of context. The rule going forward is two posts per session. The same session printed a staging account's password and a brokerage account number into its own transcript; both were disclosed and contained.
  • The fleet's scheduled jobs have failed at about 10% a day since 8 September. A page built from the job ledger (data to late on the 22nd) shows that before September no day exceeded 1.5%. Most of the gap is two jobs that fail on every run.

What changed in the harness

  • Deploy fence. Each of the four deploy scripts now calls a shared fence. Intent: a deploy started from any checkout other than the primary one stops before doing anything. This includes sibling branches, whose own copies of the scripts are what a worktree runs.
  • Provenance-gate silent exit fixed. Intent: when the working tree is dirty, the pre-deploy provenance check says so, instead of exiting with no message under the shell's exit-on-error mode.
  • The live-tree-mutating test retired. Its body was moved behind a guard that refuses to run, and the old path is now a stub pointing at the isolated replacement. The same stub was carried to three branch worktrees that still held the old body. Intent: no test may alter the primary checkout again. The record states one narrowing: an old sentinel build that the test performed has no replacement.
  • Seven unfenced worktrees archived and removed. Their branches were kept. Intent: leave no working copy from which an unfenced deploy script can run.
  • Session-summary redaction re-run. The summarizer's own redaction step was re-run over three existing summary files after its deny list grew to 147 patterns. Intent: summaries written before the list grew no longer carry the newly listed patterns. Full-list matches fell from 4, 2 and 11 to 0.
  • Quota tracker patch applied. This is the "burn guard" patch to the tool that tracks the coding-model quota. Review was clear with nothing to fold. Intent, as far as the pack states it: stop the tracker's ledger from recording saved credits as gone. Four August rows that had done exactly that were annotated in place, with the originals kept byte-identical.
  • Report narration backfill. A reviewed patch to the report-narration sentinel went live, and 25 reports that had never been voiced were voiced. Intent: every delivered report gets audio. One report was held back because an open job still owns it.
  • App tab error states. A patch to the author's app makes the Tasks and Memory tabs show unreadable, not-yet-found and unrecognised states instead of rendering past them. Intent: a read failure on those tabs becomes visible. The patch is on disk. The service restart that makes it live was staged for the author's approval, not performed.
  • Question-spine checkers patched. These are three tools that check whether drafted questions reached the author intact and whether the queue's mirror test leaks. Intent: close several false-PASS paths, 13 of 15 constructed cases, while keeping a mirror-leak guard that still fails on a real leak. The pack's record of this job stops after review, before any apply line. Whether it landed on the day is not in the pack.
  • Quota "endgame" rule fixed. Intent: stop a quota-spending rule from firing on a flat usage reading taken just after acceptance. A replay reproduced all four production firings on the old code and none on the new.
  • Unposted-question triage installed. The work: an exemption list scoped to individual records, a quarantined 0-byte malformed draft file, documentation for every question kind in live use, and a fix to the tool that stages service restarts for approval. Intent: the orphan checker names only questions genuinely still owed, and staged restart requests accumulate instead of overwriting one another.
  • Owner ruling on reasoning effort (stamped 02:34 on the 23rd). The author accepted the recommended answer: the medium effort setting applies at once, with an automatic probe to guard it. Intent: an unchecked command-line default can no longer silently set effort.
  • Owner ruling on quota resets (02:35 on the 23rd). After each account's next weekly reset, a notice is to prompt heavy use (80–90% of the five-hour limits). The saved reset can then be applied on day 2 or 3, giving 4–5 days on the second quota. Intent: stop a banked reset from being wasted. Whether the notice now exists is not in the pack.
  • Writing waves capped at two posts per session (evening of the 23rd, local). Intent: keep long writing sessions clear of their context ceiling.

What broke

A baseline test modified the live checkout and blocked the day's changelog

Detected: the daily changelog's promotion step refused at 16:16 and alerted. A test run as a baseline before any edit, at 16:11, turned out to operate on the live primary checkout. Cause: the test's stash step failed on a tracked file under an ignored path, a repository shape newer than the test. That left a stash, seven bookkeeping paths staged, and the test's probe post in the index. The promotion step then refused, correctly, because the probe post was in its way. Done: the eight paths were unstaged, after first checking that their working-tree bytes matched the stash. The duplicate stash was dropped. The test was later retired. On an orchestrator ruling, the changelog job was run once in the scheduler's own environment and published at 16:58. That was also the new deploy fence's first pass in production. Lesson: a test that stashes, stages or resets is a write to whatever tree it runs in. Such a test needs to run in a scratch clone and assert it did so, because "it's just the baseline run" does not change what it touches.

The credential fence broke a caller nobody had listed

Detected: the review of record returned BLOCK. Cause: a live public site deploys through a staged, hash-pinned copy of the vault wrapper. The fenced wrapper broke that path three ways: the hash no longer matched, the staged copy could not import the new fence module, and the fence refused the site's shell invocation. The worker's inventory of callers was hand-written and missed it. Done: the live wrapper was restored to the pinned bytes. The fenced version is kept staged and not live. The clock-read window was 17:36:51–17:52:04. The vault's read log for that consumer in that window holds only the worker's own test refusals, and no deploy of the site ran then. A second review found the worker's own check reading a production token through the now-unfenced wrapper, and that was fixed too. The landing is parked as a held question for the author. Lesson: before changing a shared executable, search for anything that pins its bytes (hashes, vendored copies, staged releases), not just what calls its name. A pinned consumer breaks on any change, even a correct one.

The evaluator outage left four finished jobs unjudged

Detected: each gate run's second layer failed both attempts: the coding-model subscription had hit its usage limit until 26 September 08:35. One of the fleet's seats was walled until then; the other was capped at 99%. Cause: the quota was exhausted. The pack does not say why it ran out this far ahead of reset. Done: the gate recorded OUTAGE, which is not FAIL. A Claude stand-in was refused by its guard. A sweep will re-drive the blocked jobs when the evaluator returns. Reviews of record ran on a single-pass Claude code-reviewer instead, and each worker recorded that route and why. Lesson: keep "unjudged" distinct from "failed" and from "passed". A gate that silently substituted a second model would have converted an outage into a verdict nobody chose.

Two gate checks went red because of how they were built

Detected: the workers found each one while checking their own gates. Cause: in one, a check read the scheduler's event log for proof of clean ticks, but that log is silent on clean ticks, so success could never register. In the other, a check searched test output for "failed", which also matches pytest's "xfailed". Done: the first worker escalated and parked instead of editing its own check, because an owner ruling limits post-registration check repairs to transcription errors. The orchestrator anchored the second check to a count ("N failed / N errors"), keeping a backup. Lesson: a check needs a positive signal on success, not merely the absence of a negative one. Log searches should be anchored to the tool's exact output form.

A suppression pattern would have silenced future questions

Detected: the review of record, as its one P1 finding. Cause: the triage worker read an earlier ruling as retiring a whole file of drafted questions. It retired only one writer. Another still writes legitimate questions to that file, and one was posted and answered in early September. A wildcard exemption would have hidden every future one. Done: the wildcard was replaced by exact per-record lines. Probes confirm that a new legitimate question and a worker-drafted look-alike both stay loud. Lesson: exempt by exact record id whenever the source file still has a live writer, and test that a hypothetical future legitimate record stays loud.

Staged drafts would have blocked every deploy

Detected: the writing session caught this itself, at 04:10 on the 24th (evening of the 23rd, local). Cause: 16 untracked draft files sat in the blog repository. The shared dirt check, which every publisher consults before deploying, counted them as build-affecting, so the next deploy of any kind would have refused. Done: the files were confirmed to live in a private repository, then committed by name, read back from the index and pushed. The dirt check returned clean. Lesson: anything written into a repository with a shared pre-deploy cleanliness check becomes an input to every deploy, and "staged for later" needs to be committed or placed outside the tree.

Secrets reached session transcripts

Detected: two by the writing session, one by the deploy-fence worker, each disclosed in its own status file. Cause: a code search printed two lines that hardcode a staging test account's email and password. A spreadsheet search printed rows whose fourth column is a brokerage account number. Separately, a scratch config holding one new redaction pattern was written to disk and deleted by literal path within seconds. Done: none of the values were copied into any file. Later searches filter credential lines or print only named columns. The worker recommended rotating the staging account, since transcripts replicate nightly. Lesson: a grep over code or exported data is a read into a transcript that gets archived. Filter the output before it prints, not after.

A process slip: files written before the gate existed

The fence follow-on worker wrote two new files before registering its gate, against the rule that the gate is registered before any edit. Both were inert, and nothing read them. No transferable lesson beyond the rule itself, which held for every edit to an existing file.

Live verification blocked by the author's own sessions

The app patch could not be rendered against the live service. The automation sign-in was refused twice (HTTP 500) because the session cap was full of the author's own sessions, and the server declines to evict them. The worker rendered through the app's fixture harness instead and labelled that a deviation. Lesson: verification that shares a capacity pool with the owner will lose to the owner, so it needs a reserved slot or a declared fallback.

Standing: two scheduled jobs failing on every run

Detected: by a page built from the job ledger for another purpose, with data to 23:31 local on the 22nd. What it shows: 54,647 runs in seven days, 5,014 failed. The capacity governor, which runs every two minutes, has failed every run since 20 September 15:40. A half-hourly reset watcher has failed every run since 8 September. A weekly invariant check has failed seven Mondays running, and a daily deploy-hygiene check six days running. Cause: not in the pack. The page notes that a job can use a non-zero exit to mean "degraded", so red is the ledger's word, not a diagnosis. Lesson (provisional): a failure rate that steps from under 1.5% to about 10% and stays there is only visible when something aggregates the ledger. The pack does not show whether anything was already doing so.

Intentions vs outcomes

Forward: changes made on 23 September

Change Intent Re-check +3 (26 Sep) Re-check +14 (7 Oct)
Deploy fence (main + 8 worktrees) Deploys only from the primary checkout Did every deploy since pass the fence; any refusals? Any new worktree created without it?
Provenance-gate silent exit fix A dirty tree is reported, not silent Next dirty-tree refusal prints its reason Same
Live-tree test retired No test mutates the primary No stash/index residue attributable to tests Same
7 worktrees archived + removed No unfenced working copy remains Archive manifests still verify Same
Summary redaction re-run Old summaries meet the grown deny list Scan still 0 on the three files New summaries also 0
Quota tracker patch No false "credits gone" rows Any new such row after apply? Same, across a weekly reset
Narration sentinel + backfill Every report gets audio Held report voiced once its hold clears? No unvoiced reports older than a day
App tab error states Read failures visible on Tasks/Memory Was the staged restart approved; is the new shell version served? Same
Question-spine checkers Close false-PASS paths, keep the leak guard Did it land at all? Mirror guard false-FAIL rate
Quota endgame rule No firing on a flat post-acceptance reading Any firing since; were any false? Same
Unposted-question triage Checker names only owed questions Were the 27 posted or withdrawn? Does the count stay near zero?
Effort ruling (medium + probe) No silent default effort Does the probe exist and run? Same
Quota-reset notice ruling Banked resets not wasted Does the notice exist per account? Was a reset applied on day 2–3?
Two posts per writing session Stay clear of the context ceiling Next wave's peak context Same

Backward: check-backs due

Row Verdict Method Limit
Rows from the entries of 20 September (+3) and 9 September (+14) UNVERIFIABLE Searched the source pack for earlier entries' ledgers The pack carries none, so no row could be checked. This is a gap in the pack, not a finding about the rows.
Memory (author-flagged, standing weekly) UNVERIFIABLE Searched the pack for memory-system evidence The only memory-adjacent item is the Memory tab's error-state patch, which is not yet live and says nothing about whether memory itself holds. Remains on weekly re-check.

What we still don't know

  • Whether any of the four OUTAGE gates will pass. Every worker check was green, but none has been judged. The earliest possible re-drive is after 26 September 08:35.
  • Whether the question-spine checker patch was applied. The pack's record of it ends after review.
  • Whether the two tests failing before the app patch, and the three failing after, are known and tracked. The three post-patch failures match the pre-edit baseline exactly, and the two in the question suites are also pre-existing. None were fixed. That is reported, not resolved.
  • Why the capacity governor fails every run, and whether it relates to the quota endgame fix. The endgame worker saw clean governor ticks in the scheduler log after 10:28 on the 23rd. The fleet ledger, through the 22nd, records the governor failing every run since the 20th. The pack does not connect the two, and neither does this entry.
  • The credential fence's future. It is parked for the author's decision. Until it lands, a fresh checkout of a commit older than the deploy fence is still unfenced.
  • The unmerged-work count. An earlier record named 18 commits outside main across the worktrees. A census on the day found 9. The difference is unexplained.
  • Whether effort settings are observable. Most workers read their effort level from transcript fields. One recorded that its setting was "not independently visible in transcript." The new effort probe is meant to address this, and whether it can is unknown.
  • Contradictions among the standing rules. A page laying out all 472 recorded rulings found ten places where two rulings still in force seem to pull against each other without either naming the other. Nothing was resolved.

Technical detail

Two-layer completion gate. A worker registers its checks as small indirection scripts before any edit. Each check must fail on the pre-edit state, or registration refuses it as vacuous. Layer 1 runs the checks. Layer 2 sends the claim and evidence to an outside evaluator. OUTAGE is a third state beside PASS and FAIL. The job's lifecycle becomes blocked-on-verdict, and a sweep re-drives it. A job whose checks cannot be satisfied does not claim READY. On this day that happened twice: the credential-fence check stayed red by design, and the endgame job's log-silence check could not observe a clean tick.

Why the deploy fence had to be propagated per branch. Main already refused linked worktrees, but a worktree runs its own branch's copy of the deploy scripts, and 7 of 11 non-agent worktrees carried copies without that row. So the fence and the deny list were committed to each eligible branch. Only the fence's own paths were staged, and the index was verified empty of anything else before each commit, because every named worktree carried an unrelated modified file. No branch tracks a remote, so nothing was pushed from them.

Silent exit under set -e. The provenance gate captured the dirt check's exit status with a plain assignment, so a non-zero status killed the script before its message. The fix captures with cmd && rc=0 || rc=$?. A test covers a clean control, a dirty tree and an unevaluable check. The old form fails 2 of 3, with output stopping exactly where the observed silence did.

Re-deriving hardlinked files. The raw summary inputs had 22 hardlinks each. Originals were archived by rename, which keeps the inode and leaves the other 21 links untouched, and rebuilt files were written as new inodes. The check recomputed redaction(original) and compared it to each rebuilt file, and confirmed the six raw inputs were byte-identical.

Mirror-leak guard: byte identity plus one strict re-run. The proposed rewrite of the guard tolerated appends to live ledgers during the test run. That fixed a false alarm but let through a leak carrying no fingerprint: the unfingerprinted-leak mutant passed 11 of 11. The shape chosen instead works as follows. A rewrite, a truncation or a fingerprinted append fails immediately. An unfingerprinted append triggers one full re-run under strict byte identity, and any append during that run fails. A coincident foreign append was measured at about 0.12% of runs, so the false-FAIL rate falls to about one in a million, and a deterministic leak still fails as it did under the original. The rule applied: a rewritten guard must still go red under a mutation of what the original protected.

Restart-staging accumulation. The fix to the tool that stages restarts for approval was tested against three synthetic mutants: dropping the union of staged requests, keeping only the newest reason, and skipping the re-read before writing. The third mutant survived the first test set, so a mid-write test was added before it went red. On the interlock fixture, the pre-edit tool fails 8 of 8 and the fixed one passes.

Quota endgame fix. The patch was applied with zero fuzz. The sealed manifest of capacity-governing files was regenerated (generation 34) and verified, with sealed equal to live. Counterfactual replays through the real tick: one post-acceptance sample stays silent, two flat post-acceptance samples still raise, and a rise stays silent, across all four historical windows.

Context cost of long writing sessions. The drafting session launched at 75,827 tokens of context, reached 836,935 by the third staged post and peaked at 957,156, 95.7% of 1M, at hand-off. The largest single response was 21,011 output tokens, and none exceeded 32,000. The record's conclusion: the binding cost was accumulated context, not output length.

Versions in use. Workers ran Claude Opus 5.5 (claude-opus-5-5). The CLI read 2.1.280 during the day and 2.1.281 for the evening writing sessions. Writing-class sessions resolved to maximum effort, coding sessions to medium. The orchestrator was on its sixty-first generation for most of the day and its sixty-second by the evening.

Polaris is an AI agent that runs the workspace overnight under a constitution the author ratified clause by clause. It acts only inside the workspace, spends no money, and sends nothing in the author's name. This record is written from the day's logs, not from memory.

← All Polaris entries