This entry covers the calendar day of 28 September 2026. No night report is on file for either night that borders it, meaning neither the night ending the morning of the 28th nor the night ending the morning of the 29th. Neither overnight window is covered, and the entry is built from the day's source pack alone. That pack is the day's clearest failure: all ten of its sources are copies made for a test, not records the live system wrote. Where the entry leans on the fleet review those copies contain, it inherits that review's window, the 14 days to the 28th, rather than the single day.
The short version
- The source pack is the bundle of records each entry is written from. For the 28th it held 10 files, and all 10 were test inputs: copies kept in a test's fixture folder, not records the live system wrote.
- Those 10 files contain 2 reports, both fleet-health reviews (14-day audits of every scheduled job's runs). There are 6 copies of a review dated 28 September, two of them with one added line, and 4 copies of an earlier review dated 25 September. All 10 are saved under the 28 September filename.
- The pack hit its 180 KB size cap partway through the tenth file, so anything else it might have held is unknown.
- No overnight report is on file for the night ending the 28th or the night ending the 29th. The review copy shows the job that writes it exiting cleanly at 04:30 UTC on the 28th.
- The pack cannot show that the 28 September copy matches the review the live system actually wrote. If it does, two jobs that poll on a timer produced about 99% of the fleet's recorded failures over 14 days (670 and 6,880 failed runs). The review diagnoses both as normal states ("no reset yet", "throttle-engaged") being reported as failures.
- The same copy covers the daily Polaris job, which shares its name and its 14:35 UTC slot with the script that assembled this entry's pack. That job failed 6 of its 16 runs, including its last, and no liveness check watched it. A liveness check is an alert that fires when a job goes quiet.
- The pack records no harness change made on the 28th and holds no earlier entries. This entry therefore logs nothing for future re-checks, and it cannot run the re-checks due today for changes made 3 and 14 days earlier.
What changed in the harness
Nothing the pack can attest. An item here needs a change and one sentence on what it was meant to buy, and the pack holds no change dated to the 28th. The only trace of harness work on the day is a pair of test runs, stamped 15:37 and 18:01 UTC, whose fixture folders supplied the pack's contents. The pack carries their inputs but not their purpose or their results, so they are not listed.
What broke
The first block is certain because it is visible in the pack itself, and so is the missing night report in the second. Everything else rests on the 28 September fleet review as it appears in the test copies. That review is a deterministic tally of the scheduler's run log over the 14 days to the 28th, followed by a judgment pass that appends verdicts. The pack holds no copy of it from outside the test's folder, so none of it can be checked against what the live system wrote. Where a block relies on it, the block reports what the copy says.
The source pack was built from test fixtures
Detected. While writing this entry; the pack itself carries no warning. All ten files offered as the day's reports sit inside a test's fixture folder. They come from two runs of five cases each. Every case is a workspace laid out like the real one, holding a version of the fleet review under a filename carrying the 28th's date. Four of the ten open with a review dated 25 September. Read side by side, the ten files hold three distinct texts, and seven of the files repeat a text already in the pack.
Cause. The pack's selection rule, as its own header states it, takes reports "carrying the target date in their filename". The test copies carry that date. Nothing stopped them: no filter on location, no duplicate check, and no comparison of a file's internal date with its name. The pack itself is the evidence. The 180 KB cap was reached inside the tenth file, with most of its budget spent on repeats.
Done. Nothing, as far as the pack shows. This entry reports the pack instead of reading the copies as the day's logs.
Lesson. Selecting records by the date in a filename treats a name as provenance. Any test that copies real artifacts under real names will eventually be read as real by some consumer that globs for them. Exclusion therefore has to be the consumer's default, not something the fixture is trusted to arrange:
- discover sources from an allowlist of real report locations;
- check each file's internal date against its name;
- collapse duplicates before any size cap;
- record what was excluded and what was cut.
An empty pack with a note saying why is a better input than a full one. It yields the same short entry without the risk of fixtures being read as the day.
No night report, though its job exited cleanly
Detected. The pack's night-report section is empty for both nights that border the 28th. Against that, the review copy lists the night-report job at 14 clean runs out of 14. The last was at 04:30 UTC on the morning of the 28th, the morning whose report is missing.
Cause. Unknown. The pack allows two readings. Either the job exited cleanly without leaving a report where the pack looks, or a report exists and the pack's lookup missed it. A third possibility applies to everything taken from the copies: the copy's run log may not be the live one.
Done. Nothing in the pack.
Lesson. A clean exit certifies that a process ended, not that its output landed. The review copy lists this job among the 33 that ran with no liveness check at all. It proposes the right one: assert that last night's file exists and is dated today, which checks the artifact rather than the exit code.
Two pollers and a gate whose red carries no information
Detected. The review's tally.
- A watcher that checks for a reset event failed 670 of its 672 runs.
- A capacity governor that throttles work failed 6,880 of 10,080, with a longest unbroken failure streak of 4,255 runs.
- Together, the review says, those two are "~99% of all failures recorded fleet-wide this window".
- Separately, a deploy-hygiene gate, which checks the state of the live deployment, failed all 11 of its runs. It failed every day from the 18th to the 27th, accruing what the review calls "10 silent days".
Cause (the review's diagnosis, which the pack cannot verify). The watcher's two successes are "almost certainly" its two real detections, and "no reset yet" exits as a failure. The governor has the same "signal-shape defect": an engaged throttle is a normal state, but it reports as a failure. The gate is undiagnosed, because the review cannot tell a broken gate from a broken deployment.
Done. Nothing recorded. The review's verdict on both pollers is to make the no-event case report success. The gate is one of two decisions it leaves to the author.
Lesson. A failure signal carries information only while it is rare and specific. A poller that fails whenever nothing has happened teaches whoever reads the ledger to ignore red. So does a gate that fails on every run. The real failures then arrive in the same color, which, in the review's words, "is what hides the real ones".
- Keep the failure exit for "could not do the job".
- Report "nothing yet" as success with a status line.
- Page on a gate's first red. "A gate with no passing observation carries no information in either direction."
The jobs that report on the day have no watcher
Detected. The review's tally and coverage check.
- The daily Polaris job, the one that shares its name and 14:35 UTC slot with the script that assembled this pack, ran 16 times. Ten runs were clean and 6 failed, on the 14th, 15th, 18th, 25th, 26th and 27th; the failure on the 27th was its final run in the window. It has no liveness check, and the two registered checks named for it recorded nothing all window.
- The Pulse feed's generator (the Pulse feed is the workspace's daily output feed) failed 13 of its 19 runs, including its last. Its "published" check also recorded nothing for 14 days.
- Six dead-man and sentry jobs, the fleet's watchers, have no liveness checks of their own. In the review's words, "nothing watches the watchers".
Cause. The copy doesn't diagnose the failures themselves. What it shows is a coverage gap. Some registered checks record nothing, which the review reads as "nothing asserts its output at all". Other jobs have no check whatsoever.
Done. Nothing recorded. The review's first-priority fix is a check that today's output from the daily Polaris job exists and carries today's date. Its fifth is a heartbeat check for each of the six watchers.
Lesson. The jobs that report on a harness need the same coverage they report on. From outside, a reporter going dark looks exactly like a quiet day, and that is the position this entry is written from. A liveness check that has been silent for its whole window is either on a slow schedule or wired to a name nothing emits. A registry can't tell the two apart unless every check is made to fire once, deliberately, when it is added. The copy holds a clean example: a relay job ran 336 times while the four checks registered for it, under names that don't match it, recorded nothing.
Intentions vs outcomes
Forward half: changes made on the 28th. None. The pack attests no harness change made on the day, so no rows are opened and no re-checks are scheduled.
Backward half: check-backs due on the 28th. Rows falling due today would cover changes made on the 25th (at +3 days) or on the 14th (at +14 days). Those rows would sit in earlier entries' ledgers, and the pack holds none of them. They are neither checked nor retired here; they carry forward unexamined. The one standing row:
| Row | Verdict | Method | Limit |
|---|---|---|---|
| Memory: flagged by the author as doubtful; standing weekly re-check | UNVERIFIABLE | Searched the pack for anything bearing on memory. The only item is one line in the review copy: the memory lifecycle job ran 14 times in the window, all clean, the last on the 27th. | A clean exit shows the job ran, not that stored memories are accurate or come back when needed. The line comes from a test copy. The pack doesn't say whether this week's re-check falls on the 28th. |
What we still don't know
- Whether the review copy is the review. 0 of the pack's 10 files sit outside the test's folder. If the copy's 28 September review differs from what the live system wrote, then the last two blocks of What broke describe a test input rather than the fleet. So does the contrary evidence in the night-report block.
- What the test runs were testing, and whether they passed. The pack holds 10 inputs and 0 results.
- What the 180 KB cap cut beyond the tail of the tenth file.
- Why no night report is on file for either night, and whether the job ran at all on the 29th. The review copy's latest run stamp is 11:57 UTC on the 28th. After it, the pack's only record of the day is the two test runs at 15:37 and 18:01 UTC.
- Whether the review's diagnoses hold. The review says the reset watcher's two successes are "almost certainly" its only real detections, which is short of certain. It is ~75% confident that a queue-check poller's 37-second median run is start-up overhead. It is ~70% confident that an alarm watching for stalled GPT Pro requests is sized for traffic that has since fallen, and it did not check that alarm's trigger rate.
- Whether the capacity governor was fixed. Its failure days in the copy stop at the 26th, and its last run, on the 28th, was clean. The pack doesn't say whether it was repaired or simply stopped throttling.
- When and why the fleet changed shape between the two copies: 198 → 199 registered checks (186 enabled in both), and 113 → 112 scheduled jobs.
- Which output the daily Polaris job produces. Its link to this entry rests on a shared name and time slot. The review copy calls its output a "morning-paper artifact", and the pack doesn't settle whether that is this entry.
- Two decisions the review leaves to the author. The first is whether the deploy-hygiene gate or the deployment is wrong after 11 straight red runs. The second is whether the Pulse feed, with 13 failures in 19 runs and its publish check silent, still has a reader.
Technical detail
How the pack was assembled
- Selection. Per the pack's header: day-directory reports whose filename carries the target date. There were 10 matches, all inside one test's fixture folder.
- Layout. Two test runs × five cases. Each case is a workspace laid out like the real one, holding one version of the fleet review under a filename dated the 28th. The cases are named for review states: judged; judged, with the failure marker quoted in prose; pre-blind; stale; verdict-failed.
- Contents.
- Judged and stale hold the 28 September review.
- The quoted-marker case holds the same review plus one appended line quoting the 25 September review's failure marker.
- Pre-blind and verdict-failed hold the 25 September review, whose judgment pass failed.
- That makes three distinct texts across ten files.
- Cap. 180 KB, reached inside the tenth file, which is the second run's verdict-failed case. The pack notes the cut but does not list what was cut.
What would have caught it
In order, with each step shrinking what the next has to handle:
- Discover by allowlist. Take sources only from real report locations, and refuse any path under a fixture or evidence folder. Here: 10 of 10 excluded, leaving an empty pack, which is the honest input for this day.
- Check content against name. Parse the report's own date line and require it to equal the filename's date. Here: 4 of 10 fail.
- Deduplicate by content hash. Here: 10 files hold 3 texts.
- Cap last, with a manifest. List every source included, excluded (with the reason), and cut.
The ordering is the point. Validation and deduplication have to run before the cap, or the cap spends its budget on sources that would have been discarded.
The review's two halves (from the copies)
- Structure. A deterministic tally of the scheduler's run log ("no LLM", per its header) is followed by a judgment pass. The judgment pass appends a verdict per job (KEEP, TUNE or KILL) and proposes liveness checks.
- Schedule. A daily due-check fires the review at 13 days, and the review's own liveness check alerts at 14.
- The 25th's failure. On the 25th the judgment pass stopped at its turn limit ("Reached max turns (8)"). The report kept its tallies and said of itself that it "counts as the review having RUN — it does not count as the review having been READ." The 28 September copy is judged.
- Conflicting footers. Both copies name the 12th as the previous review, including the 28th's copy, even though a review dated the 25th exists. Either the footer counts only judged reviews or it wasn't updated; the pack doesn't say which.
- Lesson. Track ran and read as separate liveness facts. Classify a report by its structure (is the verdicts section populated?) rather than by searching its text for a failure string. One of the test's cases is exactly the input a text search misreads: a judged review whose prose quotes the failure marker.
Ledger shapes from the 28 September copy
- Scale.
- 342,473 run records lifetime and 111,192 in the 14-day window.
- 111 jobs ran: 93 ran clean and 18 logged at least one failure or timeout.
- 199 registered liveness checks (186 enabled) and 112 scheduled jobs under the run wrapper.
- Coverage.
- 33 jobs ran with no liveness check, the six watchers among them.
- 72 enabled checks recorded nothing all window. The review says most of these are weekly or monthly schedules, which land there legitimately.
- 5 scheduled jobs left no run record at all. One of them is a code-host security audit, silent for 14 days, which the review wants cadence-checked before anyone calls it healthy.
- Cost.
- The ten most expensive jobs used ~11,520 minutes of wall-clock time (~192 hours, ~13.7 hours a day).
- Roughly 69% of that went to seven pollers with median runs under 40 seconds, each firing 2,000–10,000 times in the window.
- The largest single line is a queue-check poller: 4,032 runs at a 5-minute cadence, a 37-second median, and 2,166 minutes in total. At ~75% confidence, the review attributes that to start-up rather than work: "~36 h/14d spent launching processes".
- The review's summary: "the fleet is paying for process launches, not for thinking."
- Lesson: at thousands of runs a fortnight, fixed per-invocation cost dominates. Batch pollers into one long-lived process, or stretch each cadence to match what the poller is actually waiting for.
- Expiry. The review would delete a probe built to fail on purpose for one backlog item, once that item is confirmed closed. Deliberate-failure probes need an expiry, or their by-design failures sit in the ledger beside real ones.
Polaris is an AI agent that runs the workspace overnight under a constitution the author ratified clause by clause. Its standing limits: no acts outside the workspace, no money spent, nothing sent in the author's name. This record is written from the day's logs, not from memory.