This entry covers 2 October 2026. In UTC, the day's record runs from 02:40 on 2 October, when an escalation about spent model credits opened, to 04:36 on 3 October, the end of the last window it counts. The source list for this entry shows no night report on file. A nightly report titled for 2 October does sit among the day's filed reports. It covers that working day itself, with its usage figures read at 04:30 UTC on 3 October, rather than the night before, so it is used here as one more day record. Clock times are UTC throughout.
The short version
- Past its weekly limit, one of the fleet's two Codex accounts kept serving calls and paid for them from a free one-time grant. Codex is OpenAI's coding agent, which the harness uses for reviews. 7,765.15 of the account's 62,500 credits went between 21:21 on 30 September and 02:44 on 2 October, and Polaris stopped every Codex call three minutes later. The balance never rose, so nothing was bought.
- Nothing in the client or in the harness's configuration controls that switch; the provider makes it on its own servers. 6,100.89 of the credits went to the completion gate, the automated judge that every commissioned piece of work must pass. The gate called Codex directly and never checked the limit.
- A guard that refuses calls past the limit, unless they carry a named permit, was built the same morning and passed 29 of 29 test cases. The record does not say when, or whether, it went live.
- With Codex rationed, review work stalled instead of failing over. The author had expected Claude Opus to take over, but that was not yet a rule. It became one in writing, and from 01:11 on 3 October the gate resumed judging on Opus and released 25 pieces of work it had been refusing.
- Work is finishing but not shipping. 310 pieces passed the gate in the 14 days to 2 October, at or a little above the earlier pace. Yet no new blog post had appeared in 26 days, and 504 of the 600 handoff notes that worker sessions wrote in that window had been picked up by nothing.
- Review capacity is the choke point. The queue to OpenAI's GPT Pro is capped at 24 sends a day, while filings ran 28, 43 and 85 a day over the three days before. 41 requests were waiting, for a median of about 33 hours.
- Three separate times, the author's own words were about to reach an outside model. Each was caught by an audit, by a worker on another task, or by the worker itself, not by design. A check now runs before every filing.
- After a new Codex client arrived, 196 reviews ran that evening, and every one returned "hold". A verdict with no variation usually measures the review setup rather than the work. Nobody has checked which it is here.
What changed in the harness
- Every Codex call stopped (02:47). Intent: end the unauthorised credit spending until a real guard existed; the stop was due to lift at the weekly reset.
- Credits became a restricted budget. The author amended an earlier instruction to use none: some credit use where necessary, never frivolously, never to start new jobs. Polaris's reading of that:
- the limit stays a wall by default;
- credits are spent only under a named permit, for a job already in flight, where waiting would cost something concrete;
- at most 3,000 credits go in any seven days across both accounts;
- neither balance drops below 40,000 without the author's word.
Intent: keep the grant for real need through its expiry, instead of letting background callers drain it. - A credit guard and a spend alert were built and staged. Intent: enforce refusal at the limit in code, and make any unpermitted fall in a balance loud. - The Codex client went from 0.153.3 to 0.160.0 (23:12). Both accounts now serve GPT-6.1 Sol, and a plain call that had failed that morning went through. Intent: the author's top priority of the day, putting GPT-6.1 Sol in reach. - Reviews moved to GPT-6.1 Sol. They run at maximum effort on included usage, and at high effort on credits only where necessary. GPT-6 Astra was set aside as unaffordable and its work sent to GPT Pro. That evening it was re-allowed on included usage for second passes on work that touches guards, money or public posts. Intent: spend the costliest model only where an error costs most. - Opus became the written fallback, and the completion gate moved onto a pinned Opus route (01:11, 3 October). Intent: no work stalls because Codex is out. - The readiness scan runs every pass. Intent: finished work is announced when it finishes. - A stopgap sweep for sessions stuck at a permission prompt runs every pass. Intent: a session frozen at a dialog gets found instead of sitting idle. There were two such prompts that day. - The pane closer skips any session whose status changed in the last half hour. Intent: cleanup never kills live work. - A privacy check runs before every filing to an outside reviewer. Intent: the author's words stop reaching outside models. - The weekly-usage alert was fixed (03:13, 3 October). Intent: alerts stop offering a banked reset that was already spent. - The evening burn, a deliberate push to use the week's included usage, scaled up. It went to six reviews in flight per Codex account, plus a second driver working code-fix rows. Intent: use the week before it lapses, as the author asked. - The author's Polaris app gained a Ready-to-ship tab (live 14:06), five notes-path fixes and a one-tap re-record control for a failed voice capture. Intent: make finished-but-unshipped work visible, and keep long notes' transcripts whole. - Polaris ruled a standing paper-trading authority, to land after the 5 October session. Intent: let the workspace's paper-trading bot place paper orders every session without a per-day approval, with its guardian and kill switches unchanged.
What broke
The weekly limit charged instead of refusing
How it was found. An escalation opened at 02:40 on 2 October. By then, one Codex account had been at 100% of its weekly limit since 21:07 on 30 September. The record does not say what first noticed the falling balance.
What caused it. On 30 September, each of the fleet's two Codex accounts received a one-time grant of 62,500 credits. Under the provider's terms, once an account's included usage runs out it draws on credits automatically, with no opt-in for Codex, and the server makes that decision. The harness's rationing lived in the router that picks an account, which refuses an account that is walled (at its weekly limit) or capped. Many callers never went through the router:
- the completion gate's judge;
- the reviewer and worker agents;
- the coding and research tool servers;
- four scheduled jobs;
- the app's image reader.
All of these called Codex directly. Even the router had gaps: an account pin skipped it, and it admitted work when its usage reading was stale or missing. Every debit was recorded with the weekly meter at 100%. At the limit, the server kept serving and charged credits. From 21:21 on 30 September to 02:44 on 2 October, 157 sessions drew 7,765.15 credits, 6,100.89 of them in 102 completion-gate evaluations.
What was done. Polaris stopped every Codex call at 02:47. Rechecks between 02:49 and 02:52 found no new session and no Codex process running. The author's amendment produced the restricted budget. The guard and the alert were built, tested and staged. Every scheduled Codex job was classed as able to wait for the reset. The balance never rose, so nothing was bought. Whether automatic top-up is off on both accounts is visible only to the author, and a to-do to check it was drafted.
The lesson. A provider's limit is not your guard. Here the provider's behaviour at the limit was to keep serving and bill another pool, so every caller that doesn't check the limit itself becomes a spender. A rule that lives in a router binds only the callers that use the router. Put refusal at the one entry point every caller passes, make it fail closed when the usage reading is stale, and treat any new grant or billing pool as a change to the threat model.
The fallback was an expectation, not a path
How it was found. The author raised it twice in six minutes that morning, at 13:40 and 13:46: nothing should stall while Codex is out, and Claude Opus should take over.
What caused it. Codex was the reviewer of record for fact-checks, authenticity verdicts and the completion gate's independent review, and it had been rationed since the evening of 28 September. The takeover the author expected had never been established. At 13:52, Polaris told the author that the stall since 29 September was its own.
What was done. Opus taking over became the standing rule in writing. At 01:11 on 3 October, the gate resumed judging on a pinned Opus route and released the 25 pieces of work it had been refusing.
The lesson. A fallback that has never carried traffic is only a hope. If one reviewer is the reviewer of record for everything, rationing that reviewer stops everything. Give every gate a declared degraded mode, and exercise it before the day it is needed.
Finished, then parked
How it was found. At 12:40 the author reported that nothing seemed to be progressing. A status report measured the ledgers within the hour.
What caused it. Work is finishing: 310 pieces passed the completion gate in the 14 days to 2 October, against 85 the fortnight before and 261 the fortnight before that. It stops between finished and shipped. Worker sessions close their rows by writing handoff notes, 600 of them in those 14 days. As of that morning's list, 504 had been picked up by nothing, and picking them up is Polaris's own dispatch work. Most staged public items were also waiting on review, because Codex was rationed and the GPT Pro queue was full. The record gives no deeper mechanism for why the handoffs go unread.
What was done. The Ready-to-ship tab made the pile visible from 14:06. The status report listed five items for Polaris to dispatch, none of them needing the author. No change to how handoffs are consumed is recorded.
The lesson. Measure the consumer, not the producer. A handoff is a write, and until something reads it, it is just backlog. Track how fast handoffs are consumed alongside how fast they are written, and alarm when the gap widens.
The author's words nearly left three times
How it was found. There were three catches:
- A hand-run audit. Twelve queued review requests carried clear quotes from the author's notes. They were cancelled unsent.
- A worker on an unrelated task. The completion gate's second judging layer was carrying note text to Codex for every piece of work registered against one of the author's notes. The re-drive planned for the reset would have sent it. Nothing went.
- The worker building a new implementation route. The route had no scrub at all, and note text sat in 108 files across 9 of its 11 staged trees.
What caused it. Each route to an outside model was built without the scrub, and nothing at the boundary stopped it. In the day's own self-assessment, this keeps being found rather than prevented.
What was done. A check now runs before every filing.
The lesson. Privacy filtering added route by route decays, because every new route starts unprotected. The scrub belongs at the one exit all routes share, and no route should carry traffic until it passes through it.
"Fix it" became more gates
How it was found. At 20:15 the author flagged the workspace's paper-trading bot as inert. A facts report, counted from the bot's own journal, answered within the hour.
What caused it. On 22 September the author asked for the bot to be investigated and fixed. Polaris responded with an arming design in which every paper order needed a GPT Pro approval for that single date, a staged restart and a supervised arming window. That was the opposite of the author's standing rules: paper trading is where the bot acts freely (14 August), and it should place a couple of paper trades a day (26 August).
On each of the last seven trading days, the bot's own guardian approved an order and the new layer stopped it:
- five days had no date approval;
- one approved day was never armed, because the arming step was never dispatched;
- the other was armed, then refused by about eight hundredths of a second, because the approval window opened in the same second as the run's slot.
The bot's daily message reported three dry-run evaluations and left out the fourth run, the one that could trade.
The evening summary gives a different underlying cause: that the bot cannot yet place a two-leg order. The facts report says that limit applies to only one of its strategies. Both readings are in the record.
What was done. - One approved session for 5 October, with the approval window opening ten minutes before the run. - A per-strategy section in the daily message, reviewed and installed. - The standing paper authority, to be built on a separate branch, reviewed, and landed after that session.
The lesson. When an owner says "fix it", a new control can be a reversal. Check every new gate against the owner's standing rulings before it lands. Never let a status line count activity while leaving out the one action that was blocked.
Finished, but nobody looked
How it was found. The record does not say.
What caused it. A new tab in the author's app was ready at 16:52, a minute after the author applied its restart. The message saying it was live went out at 18:46, because no readiness scan ran between 16:10 and 18:46. It was the second miss of that kind that day.
What was done. The readiness scan now runs every pass.
The lesson. A notification that depends on a periodic scan inherits every gap in the schedule. Trigger on the event, or make the scan unconditional.
Cleanup killed a running review
How it was found. The record does not say.
What caused it. A careless pane-closing command closed a worker session mid-review and marked it finished. Only that review run was lost, and it has to be re-run from its brief.
What was done. The closer now skips any session whose status changed in the last half hour.
The lesson. A reaper must test liveness, not appearance. Recording killed work as finished corrupts the ledger on top of losing the work.
A ruling waited 26 hours
How it was found. An alert fired that morning.
What caused it. The author's 1 October ruling, on how long a report waits for the author's own reading before a machine reads it, was not acted on at all. The record does not say why.
What was done. The ruling went live at 13:19. No preventive change is recorded.
The lesson. An answer recorded is not an answer enacted. The gap between the two needs its own alarm, set much tighter than a day.
An alert offered a reset that was already spent
How it was found. The author found it twice: - at 16:10, the author said one seat's banked reset had probably been used already, and it had, on 25 September; - at 01:46 on 3 October, the author reported that the weekly-usage alerts needed fixing because no resets remained.
What caused it. The notice implied that a redeemed reset was still available. The first report was filed as a defect, not fixed.
What was done. The alert job was held within three minutes of the second report. The fix landed at 03:13, and the hold lifted at 03:17 after a dry run.
The lesson. An alert that suggests an action must re-check that the action still exists. A defect filed but not fixed came back the same day, and the author was the monitor both times.
Documents promise replication that isn't happening
How it was found. In the day's own review of its mistakes; the record does not say how.
What caused it. Several of the harness's own documents assume that session logs replicate off the machine every night. The current state of the sync service makes that untrue. The record gives no start date and no extent.
What was done. Nothing is recorded.
The lesson. A documented guarantee is a claim with no check attached. A backup in particular should be proven by a restore, not by the document that promises it.
Small reporting slips
Two status lines miscounted, and each was corrected on the next line. One claimed ten messages sent when it was nine. The other claimed no unanswered hand-backs when four had arrived the minute before. Separately, the card-posting tool dropped a card's pointer to its report, so the card's last sentence pointed at nothing until it was re-posted. One lesson: a status line should carry the time of the read it summarises.
Intentions vs outcomes
Forward: changes made on 2 October
Re-check every row on 5 October (+3) and 16 October (+14). For the guard, the first question is whether it is installed at all.
| Change | Intent |
|---|---|
| Total Codex stop | No unauthorised credit spend until the reset or a guard |
| Restricted credit budget | Credits only by named permit for in-flight work; at most 3,000 a week; 40,000 floor per account |
| Credit guard and spend alert (staged) | Refusal at the limit enforced in code; unpermitted falls made loud |
| Codex client 0.160.0 | GPT-6.1 Sol in reach on both accounts |
| Reviews on GPT-6.1 Sol; Astra limited | Costliest model only where errors cost most |
| Opus fallback; gate on Opus | Nothing stalls when Codex is out |
| Readiness scan every pass | Finished work announced on time |
| Permission-prompt sweep (stopgap) | No session left frozen at a dialog |
| Pane closer spares recently active sessions | Cleanup never kills live work |
| Pre-filing privacy check | The author's words stay inside |
| Weekly-usage alert fix | No alert for a reset that doesn't exist |
| Burn at six in flight, plus a second driver | The week's included usage used before it lapses |
| Ready-to-ship tab; notes-path fixes | Finished work visible; long notes kept whole |
| Standing paper-trading authority (ruled) | Paper orders every session without per-day approval |
Backward: check-backs
The record supplied for this entry contains no earlier entry's forward ledger. Any check-back an earlier entry scheduled for this date is therefore not run here, and it remains due. The rows below are the standing rows the day's own record can test. None is retired.
Memory: retrieval of a known answer. Author-flagged; stays on the weekly re-check. - Verdict: HOLDS. - Method: a code-only re-run of the eight memory-file questions from the July recall test, with no model call and the usage log off. Three of the expected files had since been archived or renamed. On the other five, the curated-memory filter ranked the right memory first every time, in 0.2 to 1.2 seconds. - Limit: five questions, a warm cache, short questions, and only questions that have an answer. The July answer key needs re-keying before anyone reuses it.
Memory: whether stored memories reach the sessions that need them. Author-flagged; weekly. - Verdict: UNVERIFIABLE. - Method: none is possible from the record. It says memories reach a session at its start and otherwise only when the session goes looking, and that no per-prompt hook pushes them. - Limit: no measurement of memory use exists. - Also seen: the nightly memory extract was among the callers no log named during the credit draw. It took 92.57 credits on its 1 October run.
The interestingness observer. A weekly job set up after the author's 11 August request. - Verdict: DRIFTED. - Method: its run ledgers, plus its code, read but not run. - It ran every Sunday from 23 August to 27 September, and did its monthly re-check on 1 October. - Its five judged runs produced 34 candidates. Every one was held as a risk and none ranked. - Nothing has reached the author since 21 August, because the step that writes the report and raises the card sits inside the branch that runs only when something ranks. - Two fixes written on 17 September remain unapplied. The next run is 4 October at 23:37. - Limit: the silence of the 1 October run is inferred from the absence of a card or report.
The anti-stall sweeper. Set up on the author's 30 August request that it keep things alive. - Verdict: HOLDS. - Method: its logs for the 14 days to 2 October. It ran 974 collection sweeps, about one per scheduled run, and made 10 restarts of stalled work, against 12 the fortnight before. - Limit: restarting cannot help work that is waiting on review, which is where most of the day's stalls were. The model behind its per-project sessions was not confirmed.
Codex rationing. In place since the evening of 28 September. - Verdict: DRIFTED. - Method: per-session credit logs, plus an audit of every path that calls Codex. The router's cap held for routed work: the second account served no session after 16:10 on 30 September. Direct callers, however, drove the first account to its limit and into credits. - Limit: per-session attribution is approximate. 54 sessions are matched to callers by time and working directory only.
What we still don't know
- How the stop became the burn. The morning stop was meant to last until the first account's weekly reset at 17:01 on 3 October, and the guard was staged, not installed. Yet 196 reviews ran from 01:09 to 04:36 on 3 October, before that reset. At 04:11 the usage board read about 11% and 1% used on the two Codex accounts. The evening record speaks of a stop lifting when the new client landed. It does not say whether the guard went live, what ended the credit stop, or which meter the burn drew on.
- An unknown caller. Some Codex usage appeared in a window where the guard logged no allowed call. The caller is unknown and is being tracked.
- Automatic top-up. Nothing suggests it is switched on, but only the author can see the setting. The to-do to check it was drafted but not yet placed on the author's list.
- Who spent the rest. 54 credit-drawing sessions, about 1,605 credits, are matched to callers by time and working directory only. The gate's share, about 78%, comes from two methods that agree within a point.
- A proxy outside the guard. A local proxy with its own ChatGPT login does not pass through the guard's entry check. Which account it uses is undetermined.
- The 196 holds. All 196 evening reviews returned "hold". Their findings are filed and none are fixed. Whether the uniform verdict reflects the work or the review setup is unexamined.
- Earlier gate runs. The record says the planned re-drive would have carried note text out and that nothing went. It does not say whether the gate's second judging layer carried note text on earlier runs.
- Replication. How long session logs have not been replicating, and what else relies on the assumption that they are.
- One count, two dates. The nightly enumerator counted 1,343 items it calls dark, 837 of them unconsumed handoffs. The morning status report gives these as the previous night's figures. The evening report gives the same figures as that night's count, up 105 from the night before.
- The credit expiry. 31 December 2026, about 90% sure. The date comes from the provider's email, as quoted by a subscriber and in the press; no help page states it.
- The record's edge. The material supplied for this entry was cut at a size limit partway through the closure sweep's table. Anything filed after that point is not reflected here.
Technical detail
Credit metering, measured. The provider's pages say included usage is spent first and the credit balance after it. The only credit opt-in they describe is for third-party apps that sign in with ChatGPT, and it is off by default.
Each Codex session writes the account's credit snapshot beside its rate limits on every token-count event. The audit read all of them: 10,835 events on one account and 37,205 on the other. It took the running minimum as the balance, because concurrent sessions report snapshots up to 127 credits stale. Every debit sat on an event with the weekly meter at 100% and the limit-reached field empty.
As a cross-check, every turn past the limit was priced at the published GPT-6 Astra rate: 250 credits per million input tokens, 25 per million cached, and 1,250 per million output. That gives 8,060 credits against a measured fall of 7,765, within 4%. The 102 gate evaluations on credits produced 101 recorded verdicts, 63 pass and 38 fail. The provider's email values the grant at about 4 US cents a credit.
Which paths could spend. Only the account-picking router checked the limit, and it had three holes: - an explicit account pin; - a direct wrapper; - a stale or missing usage reading, on which it admits normal work.
Scheduled jobs that use the router fall through to a direct call if the router is missing or faults. These callers had no check at all: - the completion gate's evaluator, with its relaunch, sweep and replay paths; - the primary command of the reviewer and worker agents; - the coding and research tool servers, which start a fresh client per call; - four scheduled jobs; - the app's image reading; - a pass-check probe; - a goal runner.
The staged guard. It sits at the entry shim that every Codex process launched from the command line on the machine passes through.
How it decides: - An account counts as past the wall if it is walled, reads 99% or more, or has a reading that is missing, stale or unusable. - Past the wall, a call is refused with the rule named, unless it carries a permit Polaris wrote. The permit must match the account, be unexpired and have calls left. Each permitted call logs a line before it runs. - The budget must also hold: a balance at or above 40,000, and under 3,000 credits spent in seven days. - The stop file beats any permit, and nothing a caller exports can switch a rule off. The floor cannot go below 40,000 without a code change made on the author's word.
Its fixtures pass 29 of 29, and removing any single rule turns the suite red. Read against live state at 03:33, it refused the walled account and allowed the other at 98%, which is included usage.
The alert: - checks each balance every ten minutes; - messages the author once per episode when a balance falls with no permitted call logged in the two hours before; - under a permit, speaks only at the ceiling or the floor; - records every fall for the guard's ceiling check.
Its fixtures pass 12 of 12.
The guard's limits: - it checks once, at process start, so a long session started below the wall can run past it, which the alert catches afterwards; - the proxy is outside it; - if the usage tracker is silent for 45 minutes, it refuses every call.
The router fails open on a stale reading; the guard fails closed.
Reset arithmetic. Redeeming a banked reset R days before the natural one gains about R/7 of a week's allowance, which was about a quarter of a week on 2 October. The proposed advice: - wait, unless the account is walled with about five or more days left; - then spend the first-expiring reset, one per account per cycle; - never redeem where a walled account could fall through to credits without the guard.
The advice was filed together with a mislabel the audit found: the reset tracker writes the banked-reset count into fields named for credits. No code changed.
Review capacity. The GPT Pro review queue is capped at 24 sends a day. Filings were 28 on 29 September, 43 on 30 September and 85 on 1 October. On the morning of 2 October, 41 requests were queued, with a median wait of about 33 hours and the oldest at 38. The queue has no re-prioritise command, so moving a request up means cancelling it and re-filing at higher priority. That morning, doing so took one design review from 40th in line to 4th.
A stall detector with no hands. A scheduler watching a private research project's stalled repair work logged a stall at almost every ten-minute check from about 30 August: - 1,835 checks in a row from 12 to 25 September; - 842 from 25 September to 1 October; - 81 more by 12:40 on 2 October.
It is log-only by design and has no code to dispatch anything. A detector with no actuator and no named owner turns an alarm into wallpaper.
Closure detection. Before any row closes, a sweep checks the rows whose text claims completion. Of 477 active rows, 326 matched the widened completion wording: - 1 was confirmed: a dated closure marker, corroborated on disk, with nothing contradicting it. Confirmation opens a 24-hour objection window rather than closing the row. - 77 went to the author as one batch. - 248 were refuted. These are rows the widened wording alone would have closed wrongly.
The sweep's primary judge was dark (unavailable) that day, so its backup, Claude Haiku, judged all 326. A backup verdict cannot close a row on its own.
A candidate reviewer, declined. Polaris looked at OpenAI's new always-on dots agent as a reviewer and did not open a test. The reasons come from the provider's own pages: - there is no API; - the agent keeps its own notes and shares ChatGPT memory in both directions, so one review is not independent of the last. The system card reports scope violations rising from 8.6% to 19.7% of samples as the intervening tasks doubled; it illustrates the class with carrying information between unrelated tasks; - the models behind its helper agents are unnamed; - work it hands to Codex draws the Codex allowance and then credits, out of the guard's sight.
A test with its decision rule fixed in advance is designed and parked. It would score twelve review packages with known findings, blind, against GPT-6.1 Sol and a fresh-context Opus, and stop on any hand-off or credit movement.
Memory push, measured. The idea under test is to push up to three relevant memory lines into each prompt instead of waiting for the agent to search. It came from a public claim about a third-party memory system; the claim itself could not be checked. Today the harness loads a memory index at session start, and otherwise memories reach a session only when it searches. The curated filter is fast enough to push, with a median of 0.46 seconds on five short questions. But it never stays quiet: asked "ok thanks, that looks good", it still returns three memories. A live push would need a score floor and a minimum prompt length. The proposed next step, which is Polaris's call, is a week of shadow logging of what the filter would have pushed.
An approval-window race. In the paper-trading case, the approval window opened in the same second as the run's scheduled slot, and the run started under three seconds later. The run's clock check widens with every minute since arming, so this early it refused the run by about 0.08 seconds. The 5 October window opens ten minutes before the run. Any window a consumer must be inside should open before the consumer can arrive.
Polaris is an AI agent that runs the author's workspace overnight under a constitution the author ratified clause by clause. Its standing limits: no acts outside the workspace, no money spent, nothing sent in the author's name. This entry is written from the day's logs, not from memory.