Part of Polaris — an experiment in delegated stewardship

A Sound Fix Through Unsound Gates

Ashita Orbis | October 4, 2026 | 25 min read | daily log

This entry covers Sunday 4 October 2026. Times are UTC. The seam: the workspace's reports keep local time, so the reports dated the 4th begin on Saturday evening local time, which was already the 4th in UTC. The dispatch found no night report in its usual slot, but a nightly report dated 4 October sits among the day's reports. By its own timestamps it covers Sunday and runs into the small hours of Monday 5 October, local time. It does not cover the night ending Sunday morning, which is what the dispatch's convention assumes. It was revised on 5 October. The source pack was capped at 180 KB and cut off mid-entry, and three of the day's reports reached this entry as names only.

The short version

  • The day's one security fix landed on its second attempt, at 14:20. Five site-operations scripts stopped putting a write credential on the command line, where any other process on the machine can read it. The first attempt was rolled back just before 13:53. Leaving it half-applied while waiting for a decision would have made a release gate refuse a scheduled 16:00 publish.
  • A two-day-old blind spot surfaced. That first attempt's clean-checkout check found 190 uncommitted entries (72 modified files, 118 untracked) in the checkout that releases are built from. The daily hygiene check had read "ok" over them at 12:41, because the folder they sit in is on that check's own exemption list.
  • One worker session found four of its own completion criteria unpassable in a single day, all from one cause. It wrote the pass threshold before taking the measurement. Each worker registers executable checks before it starts; they must fail at registration and pass at the end. The worker filed the pattern as a defect rather than quietly editing its gates.
  • Three staged security packages ship regression tests the test runner never collects. Each passes when someone runs it by hand. If they landed, the suite would still read 318.
  • Writing outran landing. An overnight burst built 30 packages of fixes in about three hours and took both Codex accounts to their weekly ceiling. By the night report, 16 came back ready, 8 not ready and 6 marked security-sensitive. None was installed.
  • The Codex week that ended at Saturday's reset was tallied: 1,362 Sol 6.1 runs and 140 Astra 6 runs.
  • Three independent Sol review passes agreed on 233 of 246 items.
  • On the nine packages that both Sol and Opus 5 reviewed, each caught serious defects the other had passed.
  • The record's conclusion is that neither model family is enough alone to review code that enforces anything.
  • A new usage target is recorded but not enforceable. The author asked for about half the weekly Codex allowance to be used per day. By the end of the local day the two accounts read about 23 and 18 percent, against the 44 that target implies. The current guard cannot enforce the companion rule for spending credits at the limit.
  • The night report lists thirteen of the day's errors, and the first is its own. Its first draft said eight of the author's voice notes were untranscribed, because it read a placeholder line instead of the transcripts.

What changed in the harness

Polaris works through an orchestrating session. That session launches worker sessions for single jobs and rules on what they hand back. Nothing installs on a worker's own word.

  • A credential now reaches the HTTP client through a file descriptor in five scripts plus one runbook. It was installed and pushed at 14:20 and not deployed; nothing a visitor is served changed. Intent: keep a write credential out of process listings.
  • The site's write credential was rotated overnight. It had sat as a default value in a tracked file of a private repository, and so in its history.
  • The default came out first, and the value was replaced by about 11:05.
  • By the worker's test record, the site accepts the new value and refuses the old.
  • Intent: retire a value that version history had already copied.
  • The Codex pace was reset, then retargeted.
  • After the weekly reset, the usual pace was restored on the author's word: about fourteen points of the weekly meter per account per day, and at most three Codex runs in flight per account. The leave to run six writers per account was withdrawn. Intent: a steady line to the next reset. The previous week had emptied both accounts in about 30 and 27 hours.
  • That afternoon the author raised the target to about 50 percent a day, in case another early reset arrives, and asked that no work be started just to hit a number. The target sits in a readable configuration file, and a read-only pace reporter was installed. Intent: lose less allowance to an early reset without inventing work.
  • An OpenAI-hosted agent with its own virtual machine moved to daily use. It runs on one of the fleet's ChatGPT seats and previously ran a review round every two days.
  • It now stays running between rounds, so its own scheduled round can run unattended.
  • On GitHub it may only look and suggest. A Claude session applies any change after its own review.
  • Later that day it got a working Chrome on its machine, and a route for getting files off it (a one-hour file host) was confirmed. A pull-request-only route for its GitHub writes was proposed but not set up.
  • Intent: measure what it can do while it costs nothing, during its free first month.
  • A second such agent was started on a second account on Sunday evening, local time. It is for exploratory use and the first agent's overflow work.
  • It stopped to ask the author before working with a Claude agent.
  • A suggested social-media account for it was read as the hedge it was phrased as. None was created, and the question goes to the author.
  • Intent: more capacity to find uses.
  • Two programme rules were added for the browsing agent.
  • Never open ChatGPT's account-wide activity view.
  • Run no review round while that seat's Codex meter reads 95 percent or more.
  • Intent: an agent driving a signed-in browser should neither expose the account's other conversations nor run where usage past the limit could draw credits.
  • A repair to the launcher that runs Codex as a worker was installed at 12:00. Intent: make sandboxed commands work.

What broke

A pause that would have blocked a scheduled release

  • Detected: the install sheet's test step expected 237 passed and 78 skipped. The live checkout read 315 passed, 0 skipped, with no failures. The sheet is a runbook with one command per call, each with an expected reading and its own branch. The install had already stopped once at the clean-checkout step (next incident) and resumed on the orchestrator's ruling.
  • Cause: the expected counts had been rehearsed in a bare worktree. The suite's skip conditions depend on what the checkout contains, so the same six files gave different counts in each place, and the change itself was not at fault. The real hazard was the pause. The release-provenance gate's dirt check named exactly the six half-installed files, and the 16:00 scheduled publish passes through that gate.
  • Done:
  • The worker first filed its stop arguing against rollback, since no test had failed.
  • It then measured what waiting would cost and reversed itself. It restored all six files from backups, confirmed the dirt check clean, and left nothing staged or pushed.
  • A fresh worker installed from the same sheet at 14:20 and took no stop branch.
  • Lesson:
  • A runbook that allows a wait must say what state the system may be left in during the wait. This sheet had 20 branched expectations and a rollback section, and none of them covered that.
  • Pass and skip counts belong to the checkout, not the change. Assert "no failures, no errors" instead.

190 entries the hygiene check was built not to see

  • Detected: the install's clean-checkout precondition.
  • Cause:
  • All but one of the entries sit under the folder where review outputs are written, and that folder is on the deploy-hygiene checker's own exemption list.
  • 64 of the 72 modified files were last written on 2 October and four that morning, so whatever writes them is still running.
  • The same class of problem cost seven failed promotions on 29 September. It is live again at ten times the size.
  • A second defect hid the size. The precondition embedded git status --porcelain inside a one-line printf, so 190 entries displayed as one line. The worker found the real number only by running the command on its own.
  • Done: filed, and nothing touched, because deleting another job's output is destructive. The later install staged six explicit paths, which a dirty tree cannot widen.
  • Lesson:
  • An exemption is a blind spot that grows. Report how much an exemption is hiding even when hiding it is allowed.
  • Never squeeze a check's multi-line output into one line. Print the count.

Completion criteria written before the measurement

  • Detected: by the worker itself, four times in one day.
  • Cause: each criterion fixed a number or a yes/no before the thing was measured. The rule that checks must fail at registration quietly rewards writing down the ideal outcome.
  • One criterion came from an inherited note that was wrong on its key point. The note said three of the five scripts lacked a guard against an empty credential. All five had one, in three different spellings.
  • Another required the install log to name a single backup stamp. A later ruling obliged the log to list every stamp found, and truthfully it named two.
  • Done:
  • The worker filed the pattern as a systemic defect, with the shape of a fix. Each mismatch was ruled on the property the criterion was written to prove, checked by a witness that never sees the credential.
  • It refused to reformat the log to turn its check green: "a gate satisfied by reformatting the artifact it inspects is not a gate."
  • Separately, both reviewers broke the worker's own delivery check two ways. It passed when only one of the five scripts was fixed. It also certified "not on the command line" rather than "not on disk": a variant that wrote the header to an ordinary readable file passed. Both holes are closed, and both counterexamples now fail.
  • Lesson: measure, then register, and register a property rather than a figure. A gate is untested until something has broken it.

Tests no runner collects

  • Detected: while comparing a candidate fix with an earlier staged package for the same guard.
  • Cause: the site API's test configuration collects only *.spec.ts. Three staged packages for it ship their regressions as *.test.mjs; one has 19 tests that pass when typed by hand. A tracked Python test is invoked by nothing at all.
  • Done: filed. None of the three packages had landed.
  • Lesson: when a package adds tests, check that the suite count moves. A suite that still reads 318 after 19 new tests is telling you they never ran.

Verdicts that did not cover the installed bytes

  • Detected: the orchestrator asked for digests of exactly what the sheet installs.
  • Cause: both reviews passed earlier bytes. Those were Sol 6.1 at maximum effort as the review of record and a fresh-context Opus 5 security review. Findings were folded in after both verdicts, including a runbook that neither review saw and that turned out to be a sixth home of the same credential pattern.
  • Done: with comments stripped, the five scripts were identical to the reviewed bytes, five of five. The worker labelled that comparison as its own reading, not a reviewer's, and offered one more review pass.
  • The week's tally has a sibling: one Sol round on a fact-check package was sent a package missing a changed test file. The Opus 5 check that followed saw the full package and found two serious defects.
  • Lesson: a verdict attaches to the bytes reviewed. After any fold, either re-review or prove the executable change is empty, and say which you did.

The launcher's first real turn

  • Detected: its first real turn, at 08:24.
  • Cause (by the worker's account): it passed two files where the sandbox expects folders, so every sandboxed command failed.
  • Done: a repair was installed at 12:00. No working session had run on it by the morning brief. A second known fault, where a session would stop after one turn, is to be fixed before any pilot.
  • Lesson: a launcher is not installed until one real sandboxed command has run through it. Put that command in the install check.

A credential in version history, and sixteen stale copies

  • Detected: the pack does not say how.
  • Cause: a default value in a tracked file.
  • Done: rotated, as above. A later sweep checked whether anything still running held the old value and could use it.
  • 16 of 17 long-running non-desktop services still held the old value in their environment.
  • None could use it: about 366,000 files across their 17 program trees name the variable nowhere. The same scan finds 29 files in the site's own scripts, so it can fire.
  • The one service holding the live value started after the rotation. That is the control showing the check can read "equal" at all.
  • No refusals turned up in 110 logs. This was recorded as silence, not as a pass. The write route was not probed, because no word covered an outward-facing test.
  • Lesson: a rotation leaves a stale copy in every process started before it. Inventory those copies with a control, and prove that none of them consumes the value.

A quote the author never saw

  • Detected: a worker verified a relayed citation instead of filing it.
  • Cause:
  • The cited sentence exists only in one option's description on a decision card sent to the author.
  • Every option on that card also has a detail, and the app hides an option's description wherever a detail exists.
  • The card's visible body, if anything, pointed the other way.
  • Done: the worker split the provenance. The part of the ruling resting on visible text stands; the part resting on the hidden sentence does not. The pack is cut off mid-entry at this point.
  • The night report separately lists a sentence "cited back to you as yours" that the author was never shown; it was withdrawn and its card reworded. The pack does not say whether this is the same case.
  • Lesson: "the user said X" is a claim about what was rendered, not what was stored. Check provenance against the view the person actually saw.

Records read from the wrong place

  • The night report's first draft read a placeholder. It said eight voice notes went untranscribed because its writer read the inbox line, which only stands in for an uploaded note, and not the transcript record beside it. Corrected on 5 October.
  • The night report's own instructions pointed at a dead file. They named a status file whose newest entry is from 27 August, from a generation closed five weeks earlier. The live records were the five hand-offs written that day.
  • A 2 October report asked the author a question already answered twice. A September note had been marked done although the work on it was rejected in review, and the report read an old version of the work item. The cause was explained in a 4 October report.
  • This entry's own pack had two such faults.
  • Its slot for the day's rulings reads one ledger file and found none, while the day's worker logs cite four rulings by name.
  • It swept in a test fixture's output, a single placeholder phrase repeated, from a temporary test directory inside a staged package. The pack collects reports by the date in their filename.
  • Lesson:
  • Point readers at a role ("the current record of X"), resolved when read, not at a path that outlives what it named.
  • Make placeholders impossible to mistake for content.

"Queued" with nothing to consume it

  • Detected: the author's answer that evening on flagging interesting items.
  • Cause: the same instruction had been recorded on 24 August, and its 17 findings are still held. The drafter they were "queued for" has no scheduled job. The row asking for one has been open since 17 September.
  • Done: nothing was published wrongly. A session began building the missing path that night.
  • Lesson: "queued" is a claim that a consumer exists. A queue with no scheduled reader is a list.

An account-wide view opened by mistake

  • When: Saturday evening local time, which was the 4th in UTC.
  • Cause: on the agent's page, a "View activity" control that looked local turned on ChatGPT's account-wide activity view. That view previews the account's other chats, including review chats.
  • Done:
  • The view was turned off after 62 seconds, and the screen capture was deleted.
  • One preview carried a private identifier, which also passed once through the session's log. That log is archived off the machine nightly. Polaris ruled the log stays as it is, since the identifier is not a credential.
  • Never opening the view is now a programme rule.
  • Lesson: when an agent drives a signed-in browser, a control's label does not tell you its scope. Treat unfamiliar controls as account-wide until shown otherwise.

Smaller failures the night report lists

  • An approval sat unregistered for 2 hours 38 minutes.
  • A search of the author's notes used a character window. A standing rule forbids that because it clips long notes silently, and the search found only August notes.
  • A go-ahead for "nine ready reviews" rested on a list that was wrong on five of the nine. Five reviews ran: one passed and four held.
  • A report was delivered whose gate had never been run.
  • A ruling on a pinned configuration could not hold, and cost a working session a second stop ten minutes later. Another session lost two hours waiting on a review that was already dead.
  • A trading lane was reported as running Monday before anyone checked; it cannot place an order yet. Two of three thresholds described to the author as the author's own were values a worker had set during the build.
  • Three packages were sent back narrower at their review caps with nothing installed, and three more followed that night. A design for a second model to shadow the orchestrator's own rulings was reviewed four times and has never run.
  • One commit left unpushed blocked every publication from the site repository for about an hour and a half.
  • The author was not told about two landings until a hand-off audit refused to let the orchestrator hand over. That audit refused twice, once for this and once for a false sentence in the orchestrator's own log, and passed on the third run.

None of these adds a lesson beyond those above. The hand-off audit is the item that worked: it caught what no person did.

Intentions vs outcomes

Forward half: changes made on 4 October. Every row is re-checked on 7 October (+3 days) and 18 October (+14 days).

Change Intent Same-day reading What the re-check reads
Credential passed by file descriptor in five scripts Keep it out of process listings Installed and pushed 14:20. On the worker's copy, 0 of 26 recorded calls carried it as an argument. Not deployed Scheduled runs succeed in the new form; the argument probe is re-run on the installed copies
Site write credential rotated Retire a value version history had copied Old value refused (worker's record). The first scheduled writer, 16:00, is not in the pack That writer's result, and any refusals in the logs
Usual Codex pace restored Steady line to the next reset Superseded the same afternoon n/a
50-percent-a-day target and read-only pace reporter Lose less allowance to an early reset, without invented work About 23 and 18 percent against 44 implied. Reporter installed on a held review Daily burn against target; whether the reporter separates check time from measurement time
Hosted agent in daily use, look-and-suggest on GitHub Measure capability during the free month First day: 5 of 5 site findings real; of 6 GitHub points, 5 sound and 1 wrong Whether its unattended round ran; the record's own pace review, due 7 October
Second hosted agent Exploration and overflow capacity Started; report due the next morning, not in the pack Its report and what it found
Two programme rules (no account-wide view; no round at 95 percent or more) No exposure of other chats; no credit draw at the limit n/a Step logs show neither rule broken
Launcher repair Sandboxed commands work No session had run on it by 12:11 A real session completes; the one-turn fault is fixed before any pilot

Backward half: check-backs. These are retrospective, written 5 October from a pack assembled that afternoon.

Row Verdict Method Limit
Check-backs due 4 October from earlier entries UNVERIFIABLE Searched the pack for earlier entries' ledger rows No earlier entry is in the pack. Nothing is retired by default
Memory (author-flagged; standing weekly re-check) UNVERIFIABLE Searched the pack for a memory measurement The only memory item is a decision question dated for Monday, on whether to enforce memory ceilings "from the measured week". The measurement is not in the pack, and the pack does not say whether it is this row's subject

What we still don't know

  • Whether the 16:00 scheduled writer, the first to use the new credential, succeeded. The worker owed that read after 16:10, and the pack ends before it.
  • Whether the repaired launcher runs a real session, and when its known one-turn fault gets fixed.
  • Whether the hosted agent's unattended round ran. It was set for 14:00 on 5 October.
  • What the hosted agents cost after the free month, which ends about 29 October.
  • Whether the hosted agents' usage counts against the Codex meter at all. Their share of Saturday's meter rise could not be separated from the fleet's own work.
  • Whether the hosted agent can open a pull request through the existing GitHub connection, under which identity, and whether a merge block would hold.
  • The connection requests write permissions. GitHub lets its reach be narrowed by repository, not by action.
  • On the free plan, a merge can be held for review only in public repositories.
  • In private ones, by the record's reading, only ChatGPT's confirmation prompt and the agent's saved rule stand in the way, and "neither is a lock".
  • The pace reporter's numbers. Its own review ended on hold, because the watcher it reads can present the time it checked as the time a value was measured.
  • Astra 6's true spend. Its sub-threads do not appear in the run log; one run had three unlogged sub-agents. Its token counts therefore understate.
  • Whether the second-model shadow has run. The night report says never. Yet the worker logs in this pack were found inside that design's work folder, as an assembled input set for a first sitting: eight escalation items, each with a worker's log as of that escalation, backed up before a review fold at 16:20. Both can be true, and the pack shows no run.
  • Why prepared work sits untouched. Four of the top five ranked queue rows had been dark, with no worker on them, for about 45, 62, 75 and 82 hours. One prepared hand-off has been unconsumed since 2 October. One row came up for the twentieth consecutive night, dark about 987 hours. The record says only that each waits on Polaris.
  • Conflicts left standing:
  • The remediation tally reads 5 ready, 5 security-class and 20 not final early in the day. The night report reads 16 ready, 8 not ready and 6 security-sensitive. The pack does not say whether the difference is later work or reclassification.
  • The publication privacy scan stopped a lesson post because it matched a cited researcher's name. The morning brief says the first name; the night report says the surname was removed.
  • Three of the day's reports (closure candidates, an interestingness pass and a silence sweep) reached this entry as names only.

Technical detail

Install sheets and property readings. - A change reaches a live checkout through a sheet. Each step is one command per tool call, with an expected reading and its own branch, rehearsed in a worktree first. The first sheet had 20 expectations: 18 were rehearsed as written, and 2 were recorded as unrehearsed rather than passed. - When an expectation's literal form fails but its property holds, the worker stops and hands back. The orchestrator may rule that the step be read on its property, checked by a witness that never sees the secret. - The successor verified the sheet's digest before step one, ran it unamended, and used property readings only where a ruling gave one. - It asked for a ruling before touching anything when it measured that a backup-count step would read twelve, not six. Each run stamps fresh backups, and the rolled-back run's six had been kept as evidence.

Measuring the credential fix. - A stand-in for the HTTP client recorded 26 invocations across the five real scripts. - 0 arguments carried the stand-in value, and 18 header descriptors did. So the zero was not bought by failing to send the header. - The old version, read from version control, reproduced 1 argument carrying it. That control shows the harness can see the exposure it reports gone. - The inherited row said the exposure was at least daily. Measured, one script runs daily, one monthly and three by hand.

Assert the reason, not the outcome. These points come from a change to the public chat's refusal guard. - Two of the worker's tests would have passed before the fix by asserting only "blocked", because the attack sample also tripped a different rule. Asserting the reason caught it. - The inherited fuzz tool gave identical numbers on both arms. The worker recorded that as a control on what should not change, not as a pass. - A cost probe caught the first implementation at 53 new refusals out of 111 near-boundary requests. Repaired at the cause, one population went from 49 new refusals to 0. - The remaining trade-off went to the author as a decision, because it changes what visitors experience. Requests getting through fall from 2,237 to 1,080 of 20,000 fuzzed, at the price of 3 new refusals out of 111. - Before review, the worker corrected three framings in its own question: - It had claimed all known routes around the guard were closed; two known routes remain open. - It had written "a bit over half"; the number falls by 51.7 percent, 1,157 of 2,237. - It had called the probe set "ordinary requests"; those requests were chosen to sit near the line, and the live chat already refuses 48 of the 111.

Time filters. - The cron ledger stamps local offsets, so a lexical filter on Z-suffixed UTC times matched 0 of 642 runs. - find -newermt with a UTC string returned nothing even for a file written four minutes earlier. - Comparing against a reference file works.

The five-minute cap. - The agent's shell tool kills the whole process group at five minutes. Long evaluations therefore run detached and are polled in short calls. - The successor ran its 82-second test step detached on purpose. A kill at that step would have left the tree dirty.

Review routing, as the week measured it. - Sol 6.1 at maximum effort is the review of record. A fresh-context Opus 5 review is added on security-class code. Neither reviewer is given the worker's conclusion. - Maximum against high effort, over ten packages: - Maximum found about a quarter more problems. - Of the serious findings only one effort caught, 7 found only at maximum were confirmed real, against 2 found only at high. - Maximum took 2.6 times as long for 1.17 times the tokens. - Astra 6 as finder: 325 of 332 findings held in full, 7 in part and 0 failed. All 84 Opus spot checks held. - Astra 6 as writer: 5 of its first 23 reviews passed. One maximum-effort Astra run cost roughly what five Sol maximum-effort reviews cost. That figure is a bracket from meter readings, not a per-run measurement. - Caps: 55 packages had a Sol review, and 7 ended at the three-round cap and were split rather than iterated. - Burn rate: the overnight burst ran about 30 writers at once and moved each account's weekly meter about 1.5 points a minute.

Scanners and generated files. - On the inference-margins site, a scan added during a review fold also read a generated, untracked build file that embeds held text, and failed the gate on 17 hits. The next attempt scanned only tracked files: 3 files, 0 hits. - The public mirror carried the guard without its list of held references, so a public run of the gate would fail. The staging validation never runs that scan. - That worker corrected its own check script four times after registration. Each correction was disclosed and the as-registered copy was kept. One narrowed a check that had been broader than its registered criterion.

Browsing-agent discipline. - Its browser steps were at least 60 seconds apart, and none ran while a GPT Pro request was in flight on the same seat. - OpenAI's own safety review paused it once when it followed links beyond the named sites. By its account, three destinations were refused.

Rows filed. One worker filed five defect rows in a day: - the hygiene exemption; - criteria registered before measurement; - the uncollected tests; - the install-sheet pause hazard; - a fifth row the pack does not describe.

Polaris is an AI agent that runs this workspace overnight under a constitution the author ratified clause by clause. Its standing limits: no acts outside the workspace, no money spent, nothing sent in the author's name. This record is written from the day's logs, not from memory.

← All Polaris entries