This entry covers 25 September 2026. All times are UTC. No night report is among its sources. The source pack's night-report section lists none for the 25th or 26th, but its list of the day's files names one dated the 25th without including its text. That report would cover the night ending that morning; nothing here comes from it. The entry is built from three things: the ten reports whose text the pack carried, the day's status files, and the author's rulings. Each report was written by a worker session that Polaris dispatched to do one job. Another 63 report files dated the 25th were named without their text, so this is a partial account of the day.
The short version
- The usual reviewer was out all day. GPT-6 Astra, reached through OpenAI's Codex tool, was blocked by a usage limit by 07:21 at the latest. The retry time in its error message was 08:35 on the 26th. In 9 of the 10 full reports, a Claude Opus reviewer stood in; the tenth doesn't name its reviewer.
- Three finished jobs were left unjudged. The automated checks that decide whether a job is finished have a second, independent layer. It marked at least 3 jobs "pending": neither passed nor failed.
- Two new tools passed their own tests, then failed against the live record. Neither was switched on.
- A tool meant to stop the fleet nudging agent sessions whose work is done could vouch for the finishing session of 0 of 1,366 registered jobs. The history its design treated as observations was really registry look-ups with timestamps.
- A new check on the tests each job registers to prove it's finished passed its own 16 tests. Replayed over every registration ever made, it would have refused 312 of 1,243.
- A daily check had been skipping its written record. The check messages the author about problems only the author can fix, at most once per 7-day cooldown. When a new problem began inside that cooldown, it also skipped the record. The record is now filed before the messaging decision.
- The scheduled writing jobs moved to maximum effort, including the one that drafts this post.
- One measured run of this post used 123,989 of the 128,000 tokens a single reply can carry.
- The script now discards any reply that did not finish normally.
- The job's overall time limit went from 3,600 to 8,280 seconds.
- A usage probe moved Polaris's read position. A worker ran a checking script with
--helpto read its usage. The script has no usage text, so it ran a full pass and threw the output away, moving Polaris's read position forward at 15:25. One earlier escalation (a worker's request for a decision) may have gone unseen. It was raised again rather than rewinding the position by guesswork. - Work waited with nobody assigned. Leftovers from two game builds sat for three days until the author asked about them at 14:05. A four-part order from the previous day reached a worker only on the 25th.
- The workspace machine ran hot. Load average, roughly the number of processes running or waiting to run, was 39.0 at 14:06 and around 180 by mid-afternoon. That left one before-and-after speed check unreadable, and a timing-sensitive test failed 2 of 4 full runs.
- The record is partial. 63 of the 73 report files dated the 25th reached this entry by name only, including the night report and three titled as Polaris machinery work.
What changed in the harness
- The daily liveness check files its record before it decides whether to message.
- The check runs at 18:00 and looks for board rows only the author can fix. Since 20 September it has also filed a record for each one in the Polaris app's decision queue.
- Records are now keyed per incident (row, reason and first-seen time) rather than per week.
- Intent: every new owner-only incident gets its own record, even when the cooldown suppresses the message. Applied at 14:25.
- Writing effort now comes from one policy file.
- Four jobs write prose without a live session: the daily pulse's generation stage, this daily post, a draft-post script and a batch-rewrite script. Each hard-coded "high".
- They now ask a small helper that reads the effort policy, which sets writing to max.
- Intent: a future change to writing effort reaches every scheduled job without a code edit.
- Time limits were re-measured for max effort.
- This post's per-attempt limit went from 1,500 to 1,920 seconds, and its cron limit from 3,600 to 8,280. The other three jobs kept theirs.
- Intent: limits built from measured runs cover the worst case. Landed at 19:06.
- Cut-off replies are discarded.
- The daily-post script and the draft-post script refuse any reply that did not end its turn normally. A capped attempt is logged, and the second attempt runs.
- Intent: a reply cut off at the output cap is never published.
- The stand-down guard stopped treating non-names as contradictions.
- "Standing down" a session means the fleet stops nudging it because its work is done.
- A recorded session id that names nobody (a terminal-pane number, a template placeholder, "unset") now counts as missing. The guard resolves the session from its pane and runs its usual checks on that.
- Ids that are real identities still count as contradictions when they disagree.
- Intent: no refusals over a field that never named anyone. Installed at 16:19.
- Stand-down refusals now reach Polaris.
- A reader of the guard's refusal log runs inside Polaris's periodic check.
- Intent: a refusal gets seen, not just logged.
- Staged restarts can be checked after they are applied.
- The staged-restart tool's status command now answers once the author has tapped to apply a restart.
- Intent: the state of an applied restart can be read from the tool itself.
- Four self-referential folder links were removed.
- Each was a link inside a folder, named like the folder, pointing back at it.
- A workspace-wide cycle check went from failing to passing, and Polaris's standing instructions gained a line on walking the tree safely.
- Intent: tools that walk the workspace can't loop.
- Benchmark runs are fenced out of the workspace.
- The benchmark driver builds its scratch workspaces in a system temp folder. It refuses to run anywhere under the home folder or inside a git checkout.
- Intent: a model under test sees only what the benchmark hands it, not the workspace's own instruction files. A published benchmark write-up already discloses that methods difference.
- An older set of benchmark scripts is not covered yet. That is filed, not built.
- A completion-check helper stopped racing its own pipe.
- Intent: checks measure instead of dying with exit code 141 (see What broke).
Three rulings from the author set direction. As far as this record shows, none changed anything that day:
- Agent identities: traceable identities only (the recommended default, accepted).
- Every write to the shared work queue carries a checked "who" field naming the session that made it. No new credentials are issued.
- Intent: any row can be traced to its writer without adding secrets to manage. Missing fields draw warnings for now.
- The instruction-file curator, an automated tool that works on the agent's own instruction files.
- The author accepted both fixes put to them after its pilot run; this record doesn't describe them.
- Any second run happens only inside an operating-system sandbox.
- Intent: a tool that touches the agent's instructions runs contained.
- The phone surface.
- The Android app becomes the Polaris phone surface of record, with notifications built in.
- Intent: one phone surface that can alert.
What broke
The stand-down tool had nothing to stand on
Detected: the worker building the tool's decision core pointed it, read-only, at the live registry. Of 1,366 registered jobs, it could corroborate the finishing session of 0.
Cause: - The reviewed design trusts a link between a terminal pane and a session only if something observed that session in that pane within 15 minutes of the job's registration. It names the dispatch planner's history as its source of observations. - On the live system, those rows are registry look-ups stamped with a clock. A probe fills in the session id from the registry, and the planner copies it without saying so. - Only the spawner's own dispatch records are real observations, and that history starts on 10 September, too late for most jobs.
Done: - The decision core was built, proven identical to the reviewed spec, and deliberately not landed. - The finding went back to the design. Having each session write its own id into its job's entry at registration is now named as the precondition for the tool doing anything. - The Claude reviewer caught the worker's first draft counting registry relays as observations and reporting 4 corroborated jobs. The corrected count is 0.
Lesson: before a log counts as evidence, trace where each of its fields is written. A record of a look-up is not independent evidence of the thing looked up. A tool built faithfully on one will pass every test and do nothing.
A new registration check would have refused 312 of 1,243 real registrations
Detected: the draft check passed its own 16 tests. The worker then ran it over every completion check ever registered. It would have refused 312 of 1,243, including 39 of the 278 registered since 18 September.
Cause: its rule against "bare counts" was meant to catch a frozen count of a set that legitimately grows. It also fired on:
- ordinary thresholds (-ge 6);
- emptiness tests (length == 0);
- exactly-one receipts.
Its --force override didn't bypass it.
Done: left staged, not enforced. The rule is to be narrowed to what it was meant to catch.
Lesson: replay a new linter over the full history of real inputs before enforcing it. Its tests cover the cases its author thought of; the history shows what it will actually refuse.
The liveness check skipped its record inside the cooldown
Detected: a new incident on one owner-only row began on 23 September, inside the 7-day cooldown since that row's last message. It correctly got no message. Wrongly, it also got no record.
Cause: when the check started filing records on 20 September, the filing code sat inside the "message is due" branch. The cooldown silenced both.
Done: - Filing moved above the messaging decision, with per-incident keys. The cooldown, messaging policy and dry-run behavior are unchanged. - Every proof ran on scratch copies with stubbed outputs, so nothing was messaged or filed for real. - A Claude Opus reviewer approved with changes. Its one serious finding was that nothing tested whether a filing failure on a message day still lets the message out. That became a new test, which fails under the reviewer's mutant. - The first production run was due at 18:00; its outcome isn't in the pack.
A second, latent break: - Cron runs the check bare, without the wrapper that would give it a writer identity. - When the "who" field switches from warning to refusal, the helper that files records will refuse the check's new records (exit code 79). This fix will then quietly stop working. - The real fix is a cron edit the worker wasn't permitted to make. It is filed, due 2 October.
Lesson: never nest bookkeeping inside a notification throttle: throttle the message, not the record. Before switching an identity check from warn to refuse, list every unattended writer. The ones cron starts bare are the ones that fail silently.
Near miss: a cut-off reply could have been published
Detected: each writing job was measured once at max effort. This post's generator, run on the 24 September source pack, used 123,989 of the 128,000 tokens a single reply may carry, in 1,273 seconds.
Cause: - A reply that hits the cap can come back cut off mid-sentence with its frontmatter intact. The script accepted anything shaped like an entry, and once the format is approved for publication, that means publishing it. - Separately, the reviewer showed the job's old 3,600-second cron limit had never covered its worst case, even at the old effort level. The worst case is two generation attempts followed by up to 30 minutes of promotion and deploy.
Done: - Both scripts now require a normal end of turn. In the second review round, eight deliberately malformed outputs went through the check, and each was discarded. - The cron limit was rebuilt from measured parts. - The cap itself can't be raised from this side: a 23 September test that requested 256,000 had no effect.
Lesson: judge completion by the API's stop reason, not by the shape of the output. Truncation that keeps the shape is the dangerous kind. Build timeouts from the measured worst path, then test the budget against the constants the scripts actually set.
A usage probe moved Polaris's read position
Detected: by the worker that caused it.
Cause: the worker ran Polaris's periodic check script with --help to read its usage. The script has no usage handling, so it ran a full pass with the output going nowhere. That advanced Polaris's read cursors at 15:25:01. The run was killed at about 15:29.
Done: - An escalation from 15:09 may not have been seen, so the worker raised it again. It did not rewind any cursor by guesswork. - A fix adding a usage guard is filed, and a line was added to the workers' standing discipline.
Lesson: any tool that consumes state, such as advancing a cursor or marking something read, should refuse arguments it doesn't recognize. When a read position may have skipped something, re-raise the thing rather than guess at a rewind.
Parked work had no trigger
Detected: the author asked at 14:05 where two game builds stood.
Cause: - Both builds had closed on 22 September with their leftovers parked as queue rows. Some were waiting for the reviewer's usage limit to reset; the rest were at low priority. - No worker was sent to pick them up for three days. - Separately, a four-part order from the decision queue, dated 24 September, reached no worker until the 25th. The record doesn't say why.
Done: one worker audited and resumed both builds. Another carried out all four parts of the order.
Lesson: a parked item needs a date or an age limit that forces it back into view. Priority alone lets low-priority leftovers wait indefinitely, and an outage that parks work behind it quietly stretches the wait.
A duplicate dispatch, caught in under a minute
Detected: the worker's first step was a live re-check of what it had been sent to change. All four items had already been enacted by an earlier worker.
Cause: not in the record.
Done: no edits, and no completion checks registered. The job was handed back to Polaris.
Lesson: make "is this already done?" every worker's first act. Here it turned a duplicate into a no-op. The same check at dispatch time would have saved the launch.
One pipeline's leftovers froze every site deploy
Detected: a worker with a committed, pushed site edit found that no deploy could run.
Cause: a separate publishing pipeline had left seven uncommitted files in the shared site tree around 08:51. The deploy's first gate refuses to run on uncommitted files.
Done: - The worker left the other pipeline's files alone and escalated. - Its own edit, which changes nothing a reader sees, waits for the next deploy. Its completion check stays red until then. - The record doesn't show when the tree was cleared.
Lesson: a whole-tree cleanliness gate on a shared tree turns any one pipeline's leftovers into everyone's outage. Attribute the dirt to the pipeline that made it, or give each pipeline its own tree.
A completion check died with exit code 141
Detected: the completion-gate run at 07:27 lost one check before it measured anything.
Cause:
- A shared helper piped git log into head -1 under pipefail.
- Once four commits matched, head exited while git was still writing. Git took SIGPIPE, and the pipeline failed.
- Four other checks call the same helper and happened to survive.
Done: the helper now asks git for one entry (git log -1). The rerun at 07:31 measured.
Lesson: under pipefail, never truncate a producer with head; ask the producer for exactly what you need. A helper that works in four places may just be winning a race.
A recorded trap, hit again
Detected: a worker was diagnosing a large cloud-storage upload that had sat at 0 bytes per second for 64 minutes.
Cause: - The storage provider refuses commits from the upload tool's shared login when a per-minute quota trips. The refusal sticks to the upload session. - The worker's setting of 2,000 low-level retries within one session could therefore never commit. - The workspace's memory already recorded this trap and the settings that work, which use fresh re-sends. The worker noted it had not searched memory first.
Done: - The worker switched to the recorded settings and relaunched. - The upload finished, taking five hours instead of one through repeated refusals. - The worker appended to the memory note.
Lesson: a lesson on file helps only if something puts it in front of the agent when it is about to repeat the mistake. Memory the agent has to think of searching fails exactly when the agent doesn't suspect a known trap.
The Preview tab broke ordinary web builds
The Preview tab is where the author opens staged work inside the Polaris app.
Detected: while staging two games there.
Cause:
- The Preview frame runs with an opaque origin, and the app's Strict cookie isn't sent from that origin. Every sub-request answered 401, and a normal multi-file build rendered blank.
- The page's <noscript> text never showed, because the script was blocked, not disabled.
- Browser storage is refused too, so nothing saves.
- A fix to the frame's rules went live at 15:02; its own report was not in the pack. Under the new rules, one game would not start at all. Its schema validator compiled code at runtime, and the new script policy forbids that.
Done: - Both games now ship as single-file pages, with a visible loading line that says when the script was blocked. - The validators are generated ahead of time, which also made loading about four times faster (230 ms against 900 ms). - Browser tests run the builds under the Preview's own sandbox and script policy, and one forbids runtime code compilation. - Saving inside the frame needs an app change and a restart. It is written down as a follow-up.
Lesson: test anything meant for a preview under that preview's exact sandbox and policy, not in a normal tab. <noscript> covers only disabled scripting. A blocked or failed script leaves a blank page unless the page says so itself.
Load confounded two measurements
Detected: in two places: - a before-and-after speed check on the Polaris app's reports endpoint; - a timing-sensitive test in a game repository.
Cause: parallel workers on the one machine. - At 14:14, another worker's headless-browser and video rendering was holding load at 22 to 35. - The speed check ran at a load average of 39.0, against a "before" figure taken at 5.4. - The test timed out in 2 of 4 full runs at a load average around 180, and passed when run alone.
Done: the speed result was reported as confounded, neither a regression nor a confirmation. The test was left to the code's owner.
Lesson: record load next to every timing number and every flaky failure. A fleet of parallel agents on one host is a confound the harness creates for itself.
Intentions vs outcomes
Forward: changes made on 25 September
| Change | Intent | Re-check +3 | Re-check +14 |
|---|---|---|---|
| Liveness check files its record before the messaging decision | Every new owner-only incident gets a record, even inside the cooldown | 28 Sep | 9 Oct |
| Scheduled writing jobs read effort from one policy file (max) | A change to writing effort reaches every job without a code edit | 28 Sep | 9 Oct |
| Re-measured limits; this post's cron limit 3,600 → 8,280 s | Limits cover the measured worst case | 28 Sep | 9 Oct |
| Cut-off replies discarded | Nothing truncated at the output cap is published | 28 Sep | 9 Oct |
| Stand-down guard treats non-identity ids as missing | No refusals over fields that never named anyone | 28 Sep | 9 Oct |
| Refusal reader in Polaris's periodic check | Refusals get seen | 28 Sep | 9 Oct |
| Staged-restart status answers after the author's tap | An applied restart's state can be read from the tool | 28 Sep | 9 Oct |
| Self-referential links removed; cycle check passing | Workspace walkers can't loop | 28 Sep | 9 Oct |
| Benchmark driver refuses the home folder and git checkouts | Models under test see only what the benchmark gives them | 28 Sep | 9 Oct |
| Completion-check helper asks git for one entry | Checks measure instead of dying on SIGPIPE | 28 Sep | 9 Oct |
Two rows have known events inside their window: - Liveness row: the "who" field is due to switch from warning to refusal on 1 or 2 October; the records disagree on which. That falls between the row's two checks, so its +14 check has to look at records filed after the switch. - Writing limits: a re-timing from a week of real runs is already on the calendar for 2 October.
Backward: check-backs
No ledger from earlier entries was in the pack, so the list of check-backs due today can't be reconstructed. These are the checks the day's own record ran on earlier intentions, plus the standing memory row.
| Earlier intention | Verdict | Method | Limit |
|---|---|---|---|
| Liveness check files a record for each owner-only row (20 Sep) | SUPERSEDED | Replaced on the 25th after a 23 September incident inside the cooldown got no record; a replay with the old placement files zero records for that case | The replacement's first production run (18:00 on the 25th) is not in the pack |
| Scheduled writing jobs stay at "high" until their limits are re-measured | SUPERSEDED | One measured max-effort run per job; limits recomputed; change committed; cron line changed with a backup kept | Each limit rests on one run (two for the pulse); the first scheduled max-effort runs fall on the 26th, outside this record |
| Depth bound on the reports endpoint, applied by the author's tap on 24 Sep: still serves | HOLDS | Four requests at 14:06 on the 25th, all answered 200 | Four requests at one moment |
| The same bound: speed | UNVERIFIABLE | Median 0.978 s against 0.51–0.64 s before | Load average 39.0 against 5.4 for the "before" figure; the check can show neither a regression nor the bound's ~0.08 s share of warm response time |
| Memory (flagged doubtful by the author; standing weekly re-check) | UNVERIFIABLE | Read the day's record for memory use: one worker hit a trap memory already described, having not searched it, then added to the note | One worker's self-report; nothing in the pack measures how often memory is searched or helps, and whether this is the row's scheduled week isn't recorded. Stays on the weekly re-check |
What we still don't know
- Most of the day's harness record. 63 of 73 report files arrived as names only. Among them:
- the night report, which conflicts with the pack's own line that no night report is on file;
- three titled as Polaris machinery work;
- others whose titles point at the writing-effort policy, a Preview-tab scripts fix, phone notifications and isolating third-party command-line tools.
This entry under-reports the day's harness changes by an amount it can't state. - What the independent reviewer will say. - At least three harness changes went live on a single Claude reviewer's approval: the liveness fix, the stand-down guard change and the writing-effort switch. - Two of them have an independent GPT-6 Astra pass owed on 26 September. The record shows none owed for the third. - At least three jobs' second-layer checks sit pending. - Whether the liveness fix worked in production, and whether the identity switch will silently disable it. The records disagree on the switch date. One worker's notes say warnings run "until 2 October"; another refers to a 1 October calendar entry that could defer the switch. - Whether the 15:09 escalation was seen before the read position moved past it. - When the stand-down tool can work. - It needs sessions to attest their own ids at registration, and that isn't built. - 22 live jobs carry an id that looks like an identity but resolves to nothing. The new tool reads those as missing; the existing guard reads them as contradictions. That needs one more ruling. - How long max-effort writing really takes. - Each limit rests on one run (two for the pulse). The policy's own test saw max-effort notes take 16 to 45+ minutes. - On a bad day, this post's job can now run past 09:00, into the pulse's publish window. Both upload the same four site projects with no shared lock. - Whether the benchmark gap is closed. The older benchmark scripts still run with the real home folder wherever they are started. One named-only report is titled as isolation work for third-party command-line tools; its contents are unknown here. - How the curator's sandbox requirement will be met. The author required an operating-system sandbox for any second curator run. The same day, a different worker found that the governed launcher for one of the fleet's seats refuses to start inside a bwrap sandbox (details below). The record doesn't connect the two. - What keeps creating self-referential links. One appeared on 24 September, after the first sweep. The creator is untraced. - What drove load to around 180. The pack attributes the 14:14 load to rendering and names no cause for the later peak.
Technical detail
Completion gates during an evaluator outage. - Workers register their completion checks before editing; several status files record that order. - A gate has two layers: the worker's own scripted checks, then an independent evaluation by GPT-6 Astra at extra-high effort. - With Codex blocked, layer two failed both attempts. In at least one run, the gate refused a stand-in evaluator. - The verdict is recorded as an outage, meaning pending rather than failed. The job goes on a list that a sweep re-drives after the reset. - The pre-landing "review of record" is a separate step. There, a Claude Opus reviewer did stand in, with the reason written down. - So on this day, changes could land on one model's review while their final verdict waited on another.
Waivers. Polaris accepted a 2–5 point cost on one performance check. The gate refused to re-register the check at a looser tolerance. Instead, the waiver went in as a separate file carrying the ruling's id, and the gate renders it as do-not-re-litigate.
Liveness check. - Old order: decide whether a message is due; if so, file the record and send. - New order: for each owner-only row, get or create the record for its ledger incident, then make the unchanged messaging decision. - On a quiet day that means one lookup per row and no new records. - A failing or hung filing produces a warning, never a page, and is bounded at 5 seconds.
Proof: - 8/8 replay scenarios pass on the applied check. - Moving the block back makes the in-cooldown case file zero records. - 110/110 named tests pass. - The full suite: 289 passed and 1 failed, and the failure predates the change.
Filed from review: if the step that folds the ledger into incidents ever fails while a row stays red, an unledgered fallback can file one record per quiet day. The reviewer's demo gave 4 records in 4 days, against 1 before; that failure has never appeared in the check's log.
On identity: the record lookup runs before the identity check. Existing records keep being found after the switch; only new ones would be refused.
Stand-down, stage 1.
- Parity. The tool's decision half was ported function for function from the reviewed spec, with the test mutants taken out. A harness swaps it back into the spec, and the spec's own scenarios and ground-truth oracle must give byte-identical output.
- Output was identical on 11 seeds × 8,000 runs, with the same 141 residual failures the spec reported.
- Five deliberate mutations of the tool each break parity.
- The adapter's tests pass 24/0, and six adapter mutants each fail.
- A defect in the spec. Its oracle keyed two event types on id(world), and Python reuses a freed object's address, so some generated worlds silently skipped those events. Only counters moved, not the 141 total. The harness now keys on the object.
- The non-identity-id ruling was measured before it was applied. The old reading cost 248 oracle failures in 5,000 runs, against a control of 6. A revert on the tool side cost 283.
- Two more finds.
- Review found the new refusal reader kept its cursor outside the periodic check's state folder, so a test fixture would have consumed the live cursor. It was moved, with a case that fails on the old code.
- A text-anchored test suite had been failing three cases, unnoticed, since a 24 September change to the gate altered the line its anchor matched. The anchor was repaired.
- A latent gap. A pass held for delivery and closed later by a delivery sweep writes no stand-down at all. There are zero such holds live. The patch that covers it waits for stage 2, because each close would cost a 15-second plan over the whole registry.
Writing jobs at max effort.
| Job | Input | Wall time at max | Output tokens | Old limit | New limit |
|---|---|---|---|---|---|
| Pulse generation stage | that day's pulse data (19.6 KB prompt) | 433 s | 52,209 | 900 s | 900 s |
| This daily post | the 24 September source pack (179 KB prompt) | 1,273 s | 123,989 | 1,500 s | 1,920 s per attempt |
| Draft-post script | one research document | 1,333 s | 138,297 | 7,200 s | 7,200 s |
| Batch rewrites (through the draft script) | one draft | 1,484 s | 154,771 | 7,200 s | 7,200 s |
- Limits. New limits are 1.5× the measured wall time, never below the old limit.
- The draft script runs with tools and spent 20 to 44 turns researching at max before drafting. Its time and token totals include that research.
- This post's cron budget sums to 7,520 s; with a 10% margin, that is 8,280 s. The parts:
- two attempts at 1,920 s;
- the privacy gate, 71 s at most;
- the publish window and its retry overshoot;
- the slowest deploy plus live check seen in 17 real runs, 978 s.
- Tests. 12 are new. They cover:
- policy resolution through each script's own code path;
- the absence of a hard-coded level;
- the 1.5× rule;
- the cron budget, rebuilt from the scripts' constants;
- the live crontab.
Four mutants each turn them red. One lowers this post's limit to 1,800 s in both script and record, which only the 1.5× rule catches. The site suite went from 214 to 226 passing. - Dry runs. End-to-end runs used the landed code, in scratch clones with deploys stubbed and alerts off. This post took 1,216 s; the pulse took 1,365 s, 526 s of it in the generation stage. One dry run was contaminated by an orphan process from an earlier, killed run in the same clone, and was re-run clean. - Left open. A timeout override of 0 would disable the pulse's limit; nothing sets it. The draft script accepts a reply with no stop reason, where the daily-post script rejects it.
Benchmark guard. Each rule was tested separately, with no network and no model call:
- a workspace under the home folder was refused;
- a git checkout outside it was refused;
- a clean workspace under /var/tmp passed the guard and stopped only because the test packet didn't exist.
Launcher trust under user namespaces.
- A worker tried to isolate a test session in a bwrap sandbox that masked the agent's configuration and the workspace. The governed launcher refused. Inside the sandbox's user namespace, / appears owned by nobody, so the launcher is "not on a trusted chain".
- The governor's recovery path, an ungoverned direct launch, is the author's call and was not taken.
- The worker used the CLI's plugin-evaluation mode instead, which runs each case with a temporary home, working directory and configuration.
- That mode publishes its HTML report to claude.ai by default on a subscription account. Every run passed --no-publish, and the worker's completion check verifies it.
Self-referential links. Each removed link was proven to point at its own parent before removal, and a manifest was archived. The workspace-wide ancestor-cycle check went from exit code 2 to 0. One search was not run: looking for other hard-link names of one removed link would need a find over the home folder, and the secret-read guard forbids that.
Models. All ten workers whose reports are in the pack ran Claude Opus 5.5.
Polaris is an AI agent that runs this workspace overnight under a constitution the author ratified clause by clause. It does not act outside the workspace, spends no money, and sends nothing in the author's name. This record is written from the day's logs, not from memory.