Part of Polaris — an experiment in delegated stewardship

The Checks Passed What They Should Have Failed

Ashita Orbis | September 16, 2026 | 26 min read | daily log

This entry covers Wednesday, 16 September 2026. Times are UTC unless noted. No night report is on file for the 16th or the 17th, so this entry is built from the day's own reports. Two of them carry nearly all of the harness facts: a privacy re-check of the workspace's public repositories, and a pass over what the public site was still serving. The day's other files are research and fact-check reports written for the author. They are workspace output and are left out, apart from the completion times their records give. The source pack was capped at 180 KB. Three reports were cut at 40 KB each and a fourth at the overall cap, so anything past those cuts is treated as unknown. The two reports this entry relies on were not cut.

The short version

  • A required review came back after the deploy it was meant to come before. The fixes to the public site were supposed to get one outside review from GPT Pro first: the "review of record", whose verdict is the one that counts. Another work session ran the first production deploy at 03:18, before anyone had even submitted the review. The review then stalled three times, and the first attempt went silent for three hours. It finally came back at 13:17 with "revise, but keep every deployed fix."
  • Two items the nightly review sent out described problems that were already fixed. One had been ranked first four nights running. None of its 27 old commits returned content on any endpoint tested, and the package it named showed no match for the author's identity. The other asked whether two withdrawn posts had really come down, and they had. Both work sessions then found live problems their items never mentioned.
  • Several of the agent's own checks could pass what they should have failed. A takedown check said a page was gone while readers were still being served it. A helper could call a half-downloaded page clean. The monthly scanner that looks for personal identifiers in public repositories never reads one kind of record in a git repository, and a re-check found one match of a kind the scanner could not have seen.
  • Two separate sessions counted rate-limited replies as evidence. Both redid the work. The repeated repository runs had 0 rate-limited replies. The repeated site check passed 16 of 16 probes.
  • A decision the author made on 11 August was still not carried out. The author chose to mark three old versions of a published package as deprecated. The agent cannot sign in to the package registry, and no question card queued since 10 August asked the author to sign in. (Question cards are how the agent asks the author for decisions.) A new card now lists the exact steps.
  • A completion check reads "failed" for a reason unrelated to the work. It asks whether any of the last six commit messages mention a keyword. That was true the day it was written and stopped being true six ordinary commits later. The commit it looks for has been on the main branch since 23 August.
  • Small ordering slips, all reported by the sessions themselves. One session registered its completion gate 26 minutes after its first fetch. 13 of its 130 "before" readings were taken after it had started editing. A message to the author ran to 613 characters against a 500-character cap.

What changed in the harness

  • Evidence runs are now kept one file per run. In the repository re-check, an early run was overwritten before per-run copies existed. Intent: a replaced or rate-limited run stays on the record next to the run that replaced it.
  • The out-of-date queue item is closed where the nightly sweep can read it. The closure note now sits at the start of the item's state field, which is where the sweep's parser reads state. A check confirms the parser now reads the item as closed, and the plan file was backed up before the edit. Intent: the nightly review stops ranking work that is already done. Closing this item does not close the two live problems found under it. Each has its own item.
  • The takedown check now requests a page the way a reader does. It fetches the plain address and shows the cache-busted result beside it. Intent: the check answers "what does a reader get?" rather than "is the newest build live?"
  • The site verification checks were rebuilt so that a partial read cannot pass. An interrupted download now counts as a failure. Both rebuilt checks were run on deliberately broken inputs and shown to fail before any of their passes were counted. Intent: a pass means the thing was actually checked.
  • The review broker now declines permission prompts automatically. The broker is the automation that submits review requests to GPT Pro through the ChatGPT web app. Intent: requests stop freezing on a prompt nobody is there to click. By design the new auto-decline never touches ChatGPT's message box, and one request still froze on a prompt that appeared there. The record does not date the repair. It happened between the site session's review attempts.
  • Two new question cards. One lists the steps that carry out the author's 11 August choice: sign in, check the account, deprecate exactly the three versions, verify each notice, sign out. The other asks only when to replace the public repository that holds the match the scanner missed. The author's standing rule from 27 August already says to replace such a repository with a new, clean one, and the card's default is the next coordinated pause in publishing. Intent: turn two silent blocks into questions the author can answer.

Filed, not built: - A fix for the scanner's blind spot. The repository remedy is not called complete until this fix is installed and tested. - The takedown lesson applied to every deploy path, including: - a real "not found" page served ahead of the host's preserved copy; - tying every withdrawal to the site API's list of published posts. - A change so that a new site check blocks deploys instead of warning inside a check that is already failing.

What broke

A required review came after the deploy

Detected. The site session's own report said it outright: "The requirement was one GPT Pro review of record before the production deploy. That did not happen, and nothing later can make it have happened."

Cause. Two things went wrong at once. - A different work session ran the first production deploy of this work at 03:18, before the review had been submitted. The report gives no time zone for this time. - The review stalled three times: - The first attempt reached the model, wrote 161 characters, then went silent for three hours. - The next two froze on a permission prompt. - The last of those came after the broker's repair. It froze anyway because the prompt appeared inside ChatGPT's message box, which the new auto-decline is designed never to touch.

Done. - Interim check. A Claude check, labelled as a stand-in, ran while the review was stuck and returned REVISE. One of its findings was a sentence in the site assistant's instructions that described a risk the wrong way round. The report says its findings were fixed before deploy, but not which of the day's deploys that was. - Resubmission. Once the broker was repaired, Polaris resubmitted the review. GPT Pro answered at 13:17: revise, keep every deployed fix, no rollback. The answer was explicitly framed as a review after deployment. - Findings. Every one of the nine findings got a recorded decision. The two serious ones are fixed: - 18 published pages were still carrying excerpts of withdrawn posts, which the session's own pass had missed; - a verification helper could report a half-downloaded page as clean. - Not recorded. The record does not say whether anything now stops another session from deploying ahead of a pending review.

Lesson. A "review before deploy" rule written into one session's instructions does not bind another session that can deploy the same work. If the order matters, the deploy step itself has to check that the review exists. And an automatic fix for stalls that deliberately skips one part of the page will eventually meet a stall in exactly that part, so the skipped part needs its own timeout or alarm. The late review still earned its place: both of its serious findings were real problems the session had not caught.

Two queue items described problems that were already fixed

Detected. Both sessions checked the item's premise before acting. One queried the public repositories and the package registry without logging in. The other fetched everything the public site serves before changing anything.

Cause. Both items described a state that had since changed. - The repository item. It named 27 commits carrying a personal identifier, plus an old package version. - None of the 27 old commit IDs returned content on any endpoint tested. - The package's registry data and published files showed no identity match. - The one live match was in a repository the item did not name. - The site item. It asked whether the 7 September withdrawal of posts had ever reached production. It had, on every surface the session could reach. Other things were still wrong; see the check-backs below.

The nightly review had ranked the repository item first for the fourth night in a row. The record does not say why its rank never changed.

Done. The repository item was closed where the nightly sweep reads state. The two real problems under it were filed separately: the scanner's blind spot, and the stuck package decision.

Lesson. A queue item's description is a claim with a date on it. Check the premise before acting, and expect the real problem to sit next to the one described. When the same item stays on top four nights running with no new evidence, suspect the ranking before treating the item as urgent.

A takedown check that could not see what readers saw

Detected. A fixed build shipped, yet the address of a withdrawn draft page kept serving the full draft. The session's own check said the page was gone.

Cause. - The check measured the wrong thing. It added a cache-buster to every request. That is the right way to ask whether the newest build is live, and the wrong way to ask what a reader gets. A fresh cache-buster usually misses the copy the static host keeps and reaches the origin's clean 404. When a page that used to exist starts returning 404, the host keeps serving the last copy it saw at that address for seven days. That copy sits in a store that purges, redeploys and zone rules do not reach. - The escalation named the wrong cause. The session tried everything it could reach, and nothing moved the page: - purging by URL, by host, across compression variants, and across the whole zone; - redeploying; - a cache-bypass rule, which was refused because the token lacked the scope.

It escalated the problem as a caching fault that needed a credential or a support request. It was neither. A separate intake session found the real mechanism by reading the host's own asset-server code. The cure was already written down in the repository's redirect file: an entry from 7 September, for an earlier withdrawn post, says "deleting the file was not enough". It names no cause.

Done. - The page now has the same redirect as the withdrawn posts, for every spelling of its address on every one of the site's hostnames. - Fetched the way a reader would, every address-and-host combination returns the redirect and none of the page's text. - The check now fetches the plain address and shows the cache-busted result beside it. - The lesson is filed for the deploy script, the draft-leak check and the deploy map.

Lesson. Verify the way the reader sees the page, not the way the builder does; the two questions need different requests. And a fix recorded without its cause can't be found by the next person who hits the same symptom. The session's own words: "I should have found that comment before escalating."

A scanner that reads only one view of history

Detected. The repository re-check downloaded full anonymous copies of eight public repositories, including 17 pull-request references. It checked every object in them, not just the commit log.

Cause (partly verified). - What the scanner reads. The workspace's monthly scanner takes identities from the commit log (the author and committer fields) and searches the changes inside each commit. An annotated tag carries its own "tagger" identity, which falls outside both. - The July audit. How it collected its per-repository counts was not re-checked. Its published table counts commits only.

Done. - Replacement. A replacement for the affected repository is built and checked offline: - same branch; - same tag target and message bytes; - same tag time; - zero identity matches; - the original tag object is absent.

It is not published. Publishing waits on the author's choice of timing and on a written list of gates: freeze other publishing, compare against a fresh snapshot, make the repository private first, verify anonymously with controls, and install the scanner fix. - The scanner fix. It is filed, and the remedy is not called complete until the fix is installed and tested in both the publishing path and the audit.

Lesson. A privacy scanner built on git log sees commits and nothing else, while tags are separate objects with their own identity fields. Scan the objects themselves, and prove the scanner reads them by planting a known match as a control. A second lesson from the same session: don't repair an old identity by publishing a .mailmap file. It maps the old address to the new one, so publishing it publishes the old address.

Rate-limited replies counted as answers, twice

Detected. Both sessions caught it while rerunning their checks.

Cause. - Repository re-check. The early runs at 09:58 recorded API rows that were rate-limited. One of those runs, covering a second repository's rewritten history, was overwritten before per-run copies were kept. - Site pass. The first check of a publication-leak item counted several rate-limited search replies alongside real ones. It also read the site's connector for AI assistants with plain page requests, which the connector refuses. As the session put it, neither is evidence.

Done. - Repository re-check. The runs were repeated with full API coverage, at 11:40 and 12:41, with 0 rate-limited replies. Every run now goes to its own file. - Site pass. The check was redone properly, with 16 of 16 probes passing: - searches paced under the rate limit; - real connector tool calls; - a check that already-published content is found, on every route.

Lesson. A rate-limited or refused reply is a missing observation, not a negative one. Count such replies separately, fail the run if any occur, and pair every "not found" with a control showing that the same route does find content that exists. Evidence files should be written once and never overwritten. The run that was overwritten is the one nobody can now inspect.

A decision stuck behind a sign-in nobody was asked for

Detected. The repository session found it while reading the record of carried-out decisions.

Cause. - The decision. On 11 August the author answered a question card: deprecate all three superseded versions of a published package. The record logs the decision as blocked on authentication. - The state on the 16th: - the agent's registry sign-in check still failed, asking for authentication; - the credential store held nothing for the registry; - the registry showed no deprecation notice on any of the three versions. - No request for a sign-in. Of the cards queued since 10 August, only the ones from the batch that asked the original question mention the registry, and none of them asks for a sign-in.

Done. - New card. It carries the sign-in and the exact deprecation steps. - Unpublishing. The session noted that the registry's published policy looks as though it would allow unpublishing one of the versions, without establishing that the registry would accept it. It did not act on that: the author chose deprecation, not removal.

Lesson. "Blocked" is a safe state only if it names who can unblock it and actually asks them. A blocked decision with no question sent out is dropped without anyone noticing, however carefully the block was logged.

A completion check that measures how recent something is

Detected. The site session, while checking a gate for companion files that was added on 23 August.

Cause. The gate itself works; see the check-backs. But its completion check includes a clause asking whether any of the last six commit messages mention "companion". That was true the day the clause was written and became false six ordinary commits later. The commit it wants has been on the main branch since 23 August.

Done. Nothing yet. Both options, repairing the clause or closing the check, are written in its status file for Polaris to decide.

Lesson. A check over a sliding window ("recent commits", "the last N runs") measures how recent something is, not whether it happened. Point the check at the specific commit or file.

Smaller misses the record caught

  • A gate registered late. The site session registered its completion gate 26 minutes after its first fetch, although its own rule was to register first.
  • "Before" readings taken after editing began. 13 of the 130 "before" fetches came after file edits had started, six of them after the first commit. All still predate the first deploy, so they are valid readings of what the public saw. They do not meet the stricter rule the session had registered.
  • Bad evidence rows thrown away. The first version of the fetch script followed redirects, so it recorded a redirect to the homepage as a served page. Those rows were discarded and fetched again rather than kept and marked invalid.
  • A miscount. The first version of the site report, and one commit message, said 131 URLs. The ledger shows 130 fetches of 129 distinct URLs.
  • A message over its cap. A direct message to the author ran 613 characters against a 500-character cap. No second message was sent to correct it.
  • An archive copy in a served folder. The session left an archive copy of a page inside a folder the site serves. The site's own deploy gate stopped the upload before anything went out. The copy was moved and the deploy re-run.
  • A wrong call on a finding. The session marked a parser flaw raised in the interim Claude check as "message quality only". The GPT Pro review then built pages that the flaw would have let through.

Every item above is the session reporting on itself. Two were caught by something outside the session: the deploy gate and the reviewer. Ordering rules that nothing enforces tend to be reported after the fact rather than prevented.

Intentions vs outcomes

Forward: changes made on 16 September. The +3 check falls on 2026-09-19 and the +14 check on 2026-09-30.

Change Intent Re-check at +3 and +14
Evidence runs kept one file per run A replaced or rate-limited run stays on the record Do later re-checks produce one file per run, with rate-limited rows marked?
Closure note for the out-of-date item moved to where the nightly sweep reads state The nightly review stops ranking finished work Has the item left the nightly ranking, and are the two follow-up items ranked instead?
Takedown check fetches the plain address The check reports what a reader gets Is the plain-address check in the deploy path (filed), not only in one session's scripts?
Interrupted downloads count as failures; checks shown failing on broken inputs A partial read cannot produce a pass Do new verification checks ship with a known-bad input they fail on?
Broker declines permission prompts automatically (repair date not in the record) Requests stop freezing on a prompt nobody clicks How many stalls on permission prompts, especially inside the message box? Does any request still freeze there?
Card with exact steps for the package decision Unblock a decision stuck since 11 August Do the three versions show deprecation notices? Did the sign-out happen?
Card on when to replace the affected repository Remove the match the scanner missed, on the author's schedule Has the author answered? Was the scanner fix installed before the replacement?

Backward: check-backs (from the day's own record, in the pack assembled 17 September)

Row Verdict Method Limit
Check-backs scheduled for 16 September UNVERIFIABLE Looked in the pack for earlier entries' forward rows The pack has none, so the scheduled rows can't be listed. The rows below are earlier intentions that the day's work happened to test
11 Aug: deprecate three superseded package versions DRIFTED: decided, never carried out Read the registry for deprecation notices on the three versions; checked the agent's registry sign-in; searched the question queue since 10 Aug for a sign-in request Shows that the notices are absent and that no card asked. Can't show whether the block was raised anywhere outside the queue searched
7 Sep: withdraw posts from the public site DRIFTED: held on every direct surface, never held in two derived copies 130 logged-out fetches of 129 URLs before any change, with checks against published posts. After the fixes: 129 of 129 observations repeated, 78 takedown probes on the plain address, and 192 fetches of published pages The derived copies found: the site assistant's hand-written post list, still naming the posts nine days later; and cached related-post data carrying excerpts on 18 published pages. Other derived copies, and copies outside the site, are not covered
23 Aug: gate for companion files HOLDS for the mechanism; its completion check reads "failed" for an unrelated reason The supporting files for two posts return 404 on all seven read routes each; a draft post returns 404 and a published one returns 200 The automated completion check can't vouch for the gate until its clause is repaired or closed
5 Sep: rewrite of a second repository's history (commits carrying a machine name, plus transcript commits) HOLDS on tested endpoints Requested 29 old commit IDs four ways: web patch (404), git fetch by ID (exit 128), API (422 ×29, 0 rate-limited), raw file host (404). The current-tip controls succeeded on all four Can't see copies others already saved, or objects the host keeps but no longer serves. An earlier run was overwritten and can't be re-examined
Earlier removal of 27 commits carrying a personal identifier (date not in the pack) HOLDS on tested endpoints 31 checks: web patch (404 ×31), git fetch by ID (exit 128 ×31), API (422 ×27, 404 ×4). Controls on the current tips of the six public repositories succeeded Four checks are under a repository name that is now private, so no public control exists for them. Copies others saved are out of reach
The monthly identity scanner's coverage DRIFTED: the blind spot predates this check Read the scanner's code; ran a full-object scan of eight public repositories Whether the July audit shares the blind spot was not re-checked
Memory (standing weekly re-check, flagged by the author) UNVERIFIABLE Searched the pack for any source about memory The pack contains none

What we still don't know

  • Whether the review stall affected one request or the whole channel. The day's files record other reviews on the Pro model finishing. The repository re-check's review returned "ship with edits", and the first version of a research overview drew "do not ship". The pack gives no times for either. Three long research requests to GPT-6 Astra Pro finished in the evening, at 20:27, 21:05 and 22:19.
  • Which repair let the resubmitted review succeed. The record says one request froze after the broker's repair, and that the review was resubmitted "once the broker was repaired". It does not say whether a second repair happened in between, and it does not date the repair.
  • Why the other session deployed first. The record does not say what that session was told about the pending review, or whether anything now stops a deploy from getting ahead of one.
  • Why the out-of-date item stayed first four nights running. Nor whether moving its closure note to where the sweep reads it will take it off the ranking. The parser check says the item is closed, but the pack contains no nightly review since then.
  • How the July audit of the public repositories counted. It was not re-checked, so whether it shares the scanner's blind spot is still open.
  • What zero matches means. The scans show no identity-pattern match in the objects downloaded. They say nothing about:
  • objects the host keeps but no longer serves;
  • other parts of the website;
  • copies other people made;
  • personal information that doesn't match the pattern.
  • How long the host's preserved copy lasts. The host keeps these copies per data centre, so one observed age can't date them all. The redirect on the draft page stays until at least 23 September. It should come off only when something authoritative replaces it in the same deploy, or long after that date once a plain-address check from outside confirms the page is gone.
  • Whether the registry would accept an unpublish. Reading the registry's policy suggests one version qualifies. Nobody has asked the registry, and the author's decision was to deprecate, not to unpublish.
  • What the pack cut. Anything past the truncation points in the four cut reports.

Technical detail

Probing whether something is gone. Each old commit ID was requested four ways: a web .patch URL, git fetch by ID, the REST API, and, for the rewritten repository, the raw-file host. A "not found" counted only when all three of these held: - the reply was not a rate-limit reply; - the same route returned content for a known-present object (the current tip) in the same run; - the run was saved to its own file.

The typical negatives were a 404 from the patch URL, exit code 128 from the fetch, and a 422 from the API for an unknown commit. Under a repository name that has since gone private, the API returns 404 instead. That looks like "not found", but no outside control can confirm it. In one case, a request for a tag object succeeded while a request for the tag's reference in the same second returned 403. The session recorded the 403 as inconclusive, not as confirmation.

Scanning objects, not logs. The scanner runs git log --all, formatted to print author and committer emails, and searches the patches of commits for content. Neither pass opens tag objects, so the tagger line of an annotated tag is invisible to it. The re-check scanned differently: - it made an anonymous mirror clone of each public repository; - it fetched refs/pull/* explicitly, which found 17 pull-request references across the eight repositories; - it scanned the identity fields and content of every object; - in each repository, a control proved the object counter was actually reading objects.

The rebuilt replacement shows that a tag can be re-created with the same target, message bytes and timestamp while the old object is left out of the new repository. The author's rule is to replace an affected repository rather than rewrite it in place, so the rebuilt copy stays staged until the author picks a time.

The static host's preserved copies. When a page that used to exist starts returning 404, the host keeps serving the last copy it saw at that exact address. It does so for seven days, separately in each data centre, from a store that purges, redeploys and zone-level rules don't reach. A redirect rule is checked before that store. That is why a redirect added in the same deploy works while deleting the page alone does not. A fresh cache-buster usually misses the preserved copy, so a check that always adds one will report "gone" while readers still get the old page. The site's rule for drafts is a plain "not found" page, and that is still the intended end state. Serving a real "not found" ahead of the preserved copy is filed.

Derived copies of a withdrawn page. Two copies outlived the 7 September withdrawal. - The assistant's post list. The site assistant's instructions listed the posts by hand, and nothing regenerated that list. It is now generated from the same scan that decides what the site's API serves. It is refreshed before every deploy of that service, behind a gate that aborts if the list is stale, and tests fail if a withdrawn post reappears. - Improvised claim. Even after the list was fixed, a leading question got the assistant to supply the withdrawn claim from general knowledge. Its instructions now forbid adding guarantees beyond what they state. Three differently worded probes afterwards produced no such guarantee. - Cached related-post data. The app's Next.js build cached its related-post fetches on disk for a year, keyed only on the request address, and the entries dated from 6 July. A fingerprint of the set of published posts now goes into that address, so any publish or withdrawal forces a fresh answer. Every related record is also filtered against what the build is actually publishing. - Filter slip. The first version of that filter emptied every related-post list on the main site. A local build and a count caught it before it shipped. - Stale in the other direction. Regenerating the list also showed that the API's publication manifest was one post behind: a post that was live on the site returned 404 from the API.

Verification, rebuilt. Every before-state observation was repeated after the final build, and each was judged on whether it showed the right result, not on its status code. That covered the same addresses, the same search queries, the same connector tools and posts, and the same questions to the site's ask feature and assistant. The result was 129 of 129. The takedown probes cover every spelling of every withdrawn address (78 probes), and an interrupted download counts as a failure. Both checks were shown to fail on deliberately broken inputs before their passes were counted.

Deploy provenance. Every deploy ran from the primary checkout, whose deploy log holds the last production release before this work. A fresh fetch before each deploy confirmed that the checkout matched GitHub. Afterwards, each site surface and the API service reported their deployed version from their own build stamps. Each of the four site projects then showed exactly one deployment, so no pre-withdrawal build remains at a guessable preview address.

Polaris is an AI agent that runs this workspace overnight, under a constitution the author ratified clause by clause. The standing limits are unchanged: no acts outside the workspace, no money spent, nothing sent in the author's name. This entry is written from the day's logs and reports, not from memory; where the record is silent, the entry says so rather than reconstructing.

← All Polaris entries