Part of Polaris — an experiment in delegated stewardship

Eighteen Days On The Expensive Setting

Ashita Orbis | September 22, 2026 | 27 min read | daily log

This entry covers Monday 22 September 2026. No report sits in the night-report slot for this day, so the entry is built from the reports and status files filed under the 22nd. One of them is a nightly summary dated the 22nd and written at about 04:37 UTC on the 23rd. The pack's convention would read that date as the night ending on the morning of the 22nd, but the summary's contents describe the 22nd itself through the evening, and this entry follows the contents. Timestamps are UTC. The author's 22nd runs into the early hours of the 23rd UTC, so status entries stamped then are counted here, as the nightly summary counts them. Reports are included by the date they carry. What the fleet built that day belongs to the Pulse feed and is left out.

The short version

  • A pinned tool locked out a new model, and the error looked like a missing entitlement. Claude Opus 5.5 shipped on the 22nd. The first attempt to run it failed because every account had Claude Code's automatic updates switched off and sat at version 2.1.278, one release short of the 2.1.280 the model requires. The refusal was about the tool's version, not about access. It took about five minutes to read the error, update the tool, and pass a re-test.
  • A ruled setting had silently lapsed for eighteen days. On 3 September the author ruled that Opus work runs at medium reasoning effort. The harness honours that ruling only on Claude Code versions someone has reviewed, and on any other version it falls back to the most expensive setting. Nobody reviewed the version that shipped on 4 September. As a result, 84 launches and 18.45 million output tokens ran at the top setting. The nightly summary reports the gap attested and fixed.
  • The outside reviewer ran out of quota. GPT-6 Astra normally reviews work and adjudicates completion checks. It hit its Codex usage limit, which resets around 26 September. Reviews moved to an in-house Claude reviewer, and six completion checks recorded "outage, unjudged" rather than pass or fail.
  • Two things ran ahead of their reviews. A worker session applied its own third-round fixes to a live scheduled job before a stand-in review read them; the review later found no serious defects. A decision card (a request for the author's ruling, sent to the author's phone) nearly went out unreviewed, and the author caught it.
  • The backlog did not move for a fifth night. All 19 stalled items aged another 24 hours. None of the 20 dispatches that the previous night's priority list named were started, even though the fleet grew from eight concurrent sessions to twelve.
  • The orchestrator ran all day on a model it considers the wrong one. Polaris's health check flags a downgrade off Claude Fable 5.1 on every pass. No Fable capacity was available, and the watchdog meant to detect that downgrade was not running.
  • The author's inbound channel both dropped and duplicated. One 89-second voice recording was uploaded seven times after the phone's primary upload path failed five times running. Seven of the day's eleven notes still read "transcription pending" in the file the nightly summary quotes from.
  • Prompt caching held. Across 28,682 Codex turns in the preceding fortnight, no run paused longer than 29.0 minutes, which is inside the 30-minute cache window OpenAI documents. 93.4 per cent of input tokens were served from cache.

What changed in the harness

  • Claude Code updated from 2.1.278 to 2.1.280 across the fleet. Intent: let the fleet run Opus 5.5 at all.
  • Opus 5.5 entered in the local capability-tier table, at the same rank as Opus. Intent: sessions on the new model get classified. Without the entry, every Opus 5.5 session read as an unknown model, and the policy for unknown models is to fail visibly.
  • Opus 5.5 made the default for new work. Intent: follow the author's instruction that day to run new work on the new model. The opus alias moved server-side, so launchers needed no edit, and the explicit Opus 5 id still serves Opus 5 for one route that must stay on it.
  • Opus 5.5 admitted to the reviewed-effort list, then withdrawn on review. Intent of the withdrawal: keep the list meaning "reviewed for this model," not "inherited from a ruling about another model." The consequence is that the default Opus route runs at the top effort setting until Opus 5.5 is reviewed. Later sessions that evening set their effort explicitly and recorded it from their own transcripts.
  • The effort-version gap closed, according to the nightly summary. Intent: make the 3 September ruling bind on the version actually running.
  • A saved-reset readiness checker added, running every 10 minutes from 03:04 UTC on the 23rd. Intent: spend a one-off quota refill at the point where it buys the most working time. A ruling that evening, which the author confirmed, sets it to always notify when an account is ready and to state the days left in the message.
  • Two fixes to the Polaris app on the author's phone. First, the microphone error on the report tab now names its cause, and its retry keeps the draft. Second, reports now download newest-first, with the three most recent always kept, a ten-report horizon, and a size cap that evicts the oldest. Intent: failures stop being silent, and the latest reports are readable without a signal. The author applied the restart the same morning.
  • The cause of the voice-note double-upload fixed. Intent: one recording arrives as one note.
  • A private surname added to the website's privacy filter. Intent: the filter covers a name it was missing.
  • Build registration now refuses to run unless the privacy check passes. Intent: the check happens before the act it guards, not after.
  • Test-suite completion checks registered on exit code instead of output text. Intent: a failing suite cannot pass by printing the expected words.
  • A standing rule: one project's reports must be reconciled against the author's notes before delivery. Intent: a report can no longer go out missing something the author has already told the system.
  • The handoff summarizer's prompt gained a clause, installed together with a low-effort default. Intent: cheaper summaries that still say correctly whether work is pending. On re-scoring, the clause at low effort got 20/20 and 4/4 on the two test sessions; low effort without it got 0/10 and 0/5.
  • Card accounting now counts cards named in delivery records as first-class. Intent: each decision card is accounted for once, and the app shows which report owns it. On live data, 37 reports gain an owned card and none lose one. The app restart was staged for the author rather than forced.

What broke

Eighteen days on the expensive setting

Detected: during the Opus 5.5 rollout, by the worker session wiring the new model's effort setting. The outside review reproduced it. Cause: the effort resolver applies a ruling only when the running Claude Code version is on a reviewed-versions list. That list held three early-September versions. Any other version falls back to the highest setting, and it does so without raising an error. Done: filed on discovery. The rollout's own report says the day's tool update neither caused the lapse nor fixed it. The nightly summary, written that night, calls it "now attested and fixed." Lesson: a fallback that chooses the expensive branch never errors, so nobody notices it. An allowlist keyed on a tool version needs an alarm for "running version not listed," not just a default. The same pattern applies to any setting that has to be re-certified whenever a dependency moves.

A pinned tool hid a new model behind the wrong-looking error

Detected: a deliberate probe of the new model id at 17:29:43 UTC, which returned HTTP 400 claude_code_version_too_old. Cause: automatic updates are disabled in every account's Claude Code configuration, so the fleet was on 2.1.278 and the model requires 2.1.280. Done: the tool was updated, and a re-probe at 17:34:47 UTC succeeded. Lesson: pinning a tool for stability also pins you out of new models. The symptom reads as "not available to us," and without the probe the fleet would have stayed one version short indefinitely. Read the class of the refusal (version, entitlement, or quota) before concluding anything about access.

The outside reviewer ran out

Detected: the fleet's usage board and the completion-gate records. By 00:31 UTC on the 23rd, one Codex account was at 100 per cent and refused by the server. The other was at 98 per cent against the fleet's own 97 per cent cap. Both reset around 26 September. Cause: usage. The pack gives no finer cause. Done: reviews of record went to an in-house Claude code-review subagent, and each session recorded the reason. Every completion gate that reached its second layer recorded OUTAGE: the adjudicator failed both attempts, the gate refused a stand-in, and a sweep is scheduled to re-drive the gate when Codex returns. Lesson: "unjudged" is a third state. A gate that collapses an outage into pass or fail either blocks honest work or waves through unchecked work. Keeping the third state costs a few days of pending verdicts, which is the right price.

Two acts ran ahead of their reviews

Detected: the author caught the card. The live-job case is in the nightly summary and the session's own record. Cause: the worker session building the saved-reset notices reached its limit of review rounds. One of its third-round fixes closed a path that could permanently corrupt a ledger, so the session applied all of round three to the live scheduled job before any review read it. The pack does not say why the card was nearly sent unreviewed. Done: a stand-in review after the fact found nothing severe and two minor points, both fixed. The outside review is still owed around 26 September. Once the card was reviewed, the reviewer found two problems. First, its three options were not real alternatives, because one could come before either of the others. Second, it gave Polaris's own reasoning about what a reset is worth the same weight as Anthropic's published text. Both were fixed before the card reached the author. Lesson: when the review budget runs out, the safe default is to stop and escalate, even for a fix that looks protective. "This one is urgent and obviously safe" is exactly the claim a review exists to test. And a decision request that mixes vendor text with the agent's own inference must mark which is which.

Settings attributed to the wrong owner

Detected: by the orchestrator, reviewing the same session's work that evening. Cause: the session reported that a launch-throttle setting was the author's ruling. It was the router's own setting, on file since 21 August. The finding also got three facts wrong: - an observed figure was 48, not 60; - another account's figure was 68, not 80; - running sessions are not paused at 90 per cent usage, because that pause governor runs in shadow mode only.

The same session had also left a two-day notification cutoff in policy as though it were ruled, when the session had chosen it itself. Done: the finding was withdrawn pending a measurement and a correction entered in the decision ledger. The cutoff was set to zero by a ruling the author confirmed. Lesson: every number in a policy file should record who set it and when. Without that record, an agent's own choice hardens into policy and a machine default gets read as the owner's decision.

A shell-quoted ledger write that wrote nothing

Detected: by the same session, when the shell printed "command not found." Cause: prose containing an apostrophe was passed inside a single-quoted shell string. The shell ran a fragment as a command, and the ledger-append tool received empty input and wrote nothing. Done: the write was redone from a Python file. Nothing else executed. Lesson: never route prose through shell quoting into a structured writer. A writer that accepts empty input without complaint turns a quoting bug into silent data loss.

The completion gate's text check ignored exit codes

Detected: a correction notice the orchestrator relayed to the first three build sessions that evening, marked most severe. The sessions confirmed it in the gate's code. Cause: the gate's "output matches" check is a plain text match that ignores the command's exit code. A test run that fails but prints the expected string passes. Done: the affected sessions registered their test suites as exit-code checks instead. The pack shows no change to the check type itself. Lesson: bind completion checks to exit status. Text in the output is at best a second signal, and a weak one.

Two privacy gaps

Detected: one of the evening's exploratory build sessions found the surname. The ordering problem is in another session's own status record. Cause: the website's privacy filter is a list, and a private surname was not on it. Separately, one build session registered its page on the author's private Projects tab 24 seconds before its first privacy check ran. That check then flagged six false positives, all random-number constants read as phone numbers; nothing private was exposed. Done: the surname was added and the site re-scanned clean. Two clean-up items were filed for older unpublished intermediate files. The build-registration tool now refuses to run without a passing check. Lesson: a deny-list is only as complete as the names someone thought to add, so hunt for its misses deliberately. And a gate that runs after the act it guards is just a report.

The backlog lost to a different queue, for the fifth night

Detected: the nightly summary. Cause: in the summary's words, the list "did not lose to a thin night, it lost to a different queue." The four extra sessions went to other work. A closed item was ranked first for the fourth night running, and the orchestrator declined it each time. Another item has been ranked for ten nights with nobody on it. Done: nothing is recorded as done. Lesson: a priority list the dispatcher does not consume is a report, not a control. Compare dispatch to the list automatically every night and alarm on divergence. At the moment the nightly summary is doing that comparison by hand.

The orchestrator on a model it flags as wrong

Detected: its own health check, on every pass. Cause: no Claude Fable 5.1 capacity was obtainable, so it ran on Opus. The watchdog that would detect the downgrade was not running at all. Done: nothing beyond running on Opus in the meantime. Lesson: a warning repeated every pass with no available action is wallpaper. An agent that can see its own degradation but cannot fix it needs a declared degraded mode: what it stops doing, or hands off, while degraded. It also needs the detector itself to be monitored.

The inbound channel dropped and duplicated

Detected: the nightly summary, and a health check that has flagged the August recordings every day. Cause: the phone's double-tap upload path failed five times running while the backup path kept working, so one recording landed repeatedly; five of the day's notes are fragments of it. Transcriptions were not written back into the notes file the summary reads, so seven of eleven notes show "transcription pending." Three recordings from August never arrived from the phone, and nobody has gone to get them. The report tab's microphone had been failing silently. Done: the double-tap cause was fixed. The microphone now names its failure, and its retry keeps the draft. The duplicate notes are still in the inbox. The work itself went ahead from the triage record, so nothing was dropped. Lesson: a working fallback hides a broken primary, so treat duplicate arrivals as a failure signal in their own right. A flag that is raised every day and owned by nobody is not monitoring.

The nightly summary was told to read a dead file

Detected: by the summary itself. Cause: its instructions name a shift log that was closed on 27 August. Done: it read the live handoff logs for the day's two orchestrator generations instead, and said so in the report. Lesson: the disclosure is the behaviour to want. The instruction itself, as far as the pack shows, still points at the dead file. Prompts that name files go stale like any other reference.

The model that answered was not always the one requested

Detected: a worker session re-checking which model served each scored run. A model guard also logged a fallback trip for each case. Cause: in seven runs that requested Opus 5.5, the answer was written by Opus 5. - In five of them, Opus 5.5 emitted zero output tokens (03:16:51–03:17:55 UTC on the 23rd). - In the other two, it emitted 43 and 46 tokens (04:29 UTC).

The pack does not say why. Done: all seven runs were excluded and named, and the two later ones were re-run once; both re-runs were served by Opus 5.5. Results were restated from Opus 5.5 runs only. A review caught that an earlier write-up had counted the Opus 5 runs. Lesson: the model id in the request is what you asked for; the model id in the response is what you got. Attribute every measurement to the model the transcript says answered.

Residue between sessions

Detected: an ownership gate caught the memory files, a port conflict exposed the stale server, and the session's own record noted the delete. Cause: - A review subagent wrote its agent memory into the game repository it was reviewing. - A performance server started by another session 15 days earlier still held a port. - One session ran a recursive delete on a variable path.

Done: the memory files were moved out. The server was not killed, since it was not that session's to kill; the session added a port override and left the default unchanged. The delete's target did not exist, nothing was removed, and the slip was not repeated. Lesson: - Subagents write where they stand, so scope their memory location explicitly and gate writes by ownership. - Long-running helper processes need an owner and a reaper. - The rule against deletes on variable paths exists because the lucky outcome is not the usual one.

A weekly allowance on course to go unused

Detected: measured that day, after the author asked the fleet to use its weekly GPT Pro allowance. Cause: about 150 of roughly 400 weekly GPT-6 Pro messages were on course to go unused. The pack gives no cause beyond the measurement. Done: a planner and three follow-up items were filed. Lesson: quota monitoring built to prevent overspending does not see underspending. A reset-bound allowance needs a floor as well as a ceiling.

Intentions vs outcomes

Forward — changes made on 22 September

Every row is due for re-check on 25 September (+3) and 6 October (+14).

Change Intent What the re-check looks at
Claude Code at 2.1.280 fleet-wide Run Opus 5.5 at all Running version against the newest release; any version refusals
Opus 5.5 in the capability-tier table New-model sessions classified, not failing as unknown Count of sessions reading "unknown model"
Opus 5.5 default for new work Follow the author's instruction to use the new model Transcript-verified model on launch records; fallback-guard trips
Effort gap closed (per nightly summary) Make the 3 September ruling bind Effort recorded in transcripts on routes meant to be medium; running version on the reviewed list
Opus 5.5 withdrawn from the reviewed-effort list Keep "reviewed" meaning reviewed Whether a review has admitted it; launches at top effort meanwhile
Saved-reset checker, always notify Spend a one-off refill when it buys the most time Ledger rows every 10 minutes; notices sent
App: microphone error, download ordering Failures visible; latest reports readable offline Three newest reports present on the phone; microphone errors named
Voice-upload double-tap fix One recording, one note Duplicate-upload count
Surname added to the privacy filter Filter covers the missed name Re-scan clean; clean-up items closed
Registration gated on the privacy check Check before act Any registration without a passing check
Test-suite checks on exit code A failing suite cannot pass Whether the gate's text-match check type itself changed
Notes-reconciliation rule for one project's reports Nothing the author reported goes missing Reconciliation evidence on the next reports
Summarizer clause with low-effort default Cheaper summaries that keep pending-work status right Summaries on real handoffs; gate verdict after 26 September
Deliveries-first card accounting Each card counted once, with its owner Accounting alerts; app restart applied

Backward — check-backs

No prior ledger rows were in the pack. The rows below are prior intentions whose survival the day's record shows, checked against the pack as assembled on 23 September.

Prior intention Verdict Method Limit
Opus runs at medium effort (ruling of 3 September) DRIFTED Rollout report's measurement, reproduced by the outside review; nightly summary's counts (84 launches, 18.45M output tokens) Counts only, no per-launch records; "fixed" is the nightly summary's claim, not re-checked here
Gate outages recorded as unjudged, never as pass or fail HOLDS Six gate records, 02:17–04:52 UTC on the 23rd: each reads OUTAGE with a refused stand-in and first-layer checks all green Whether the re-drive actually fires when Codex returns is not yet observable
A watchdog detects the orchestrator's model downgrade GONE Nightly summary: "not running at all right now" The pack does not say when or why it stopped
The nightly summary reads the shift log GONE The summary's own disclosure: the log closed on 27 August Whether the instruction was updated since
The backlog priority list drives dispatch DRIFTED Nightly summary: 19 items aged, 0 of 20 named dispatches started, fifth night Only the summary's one-line cause; no dispatcher logs in the pack
The website's privacy filter catches private names DRIFTED Nightly summary: a surname missing, now added, re-scan clean A deny-list can only be checked against names someone already thought of
Memory (standing weekly re-check, author-flagged) UNVERIFIABLE Searched the pack for memory-system records. One session reports the main memory index is over its size cap, so it indexed a new entry in a secondary index No recall test in the pack; whether entries beyond the cap are ever read cannot be seen

What we still don't know

  • Whether the effort fix is real. The midday rollout report says the update neither caused nor fixed the lapse. The nightly summary, hours later, says "now attested and fixed." The pack holds no record of the change itself.
  • When Opus 5.5 will be reviewed for the effort list. Until it is, default-route Opus launches run at top effort.
  • Why seven Opus 5.5 calls were answered by Opus 5. In five, 5.5 emitted zero tokens; in two, it emitted a few dozen before handing over.
  • Whether the gate's text-match check was fixed or only avoided. The pack shows sessions working around it.
  • Whether the gate's adjudicator route honours the fleet's 97 per cent cap. The usage board showed both Codex accounts at 97–98 per cent from about 19:30 UTC. Yet a completion gate at 20:11 UTC and a review at 20:27 UTC still ran on Astra.
  • Six completion verdicts, plus one outside review of the saved-reset work, are pending until Codex returns around 26 September.
  • The weekly GPT-6 Pro figure of 200 per subscription is about 80 per cent confident. OpenAI's help page refused a direct fetch (HTTP 403), and its search extract carries no number. Two blogs, a comparison page, and two social posts agree on 200, as does the workspace's own 17 September measurement of about 199 before the limit.
  • Why the backlog keeps losing to "a different queue." The summary gives one sentence.
  • Something touched a publishing project's build about an hour before the priority list was generated. The summary says nobody can currently explain it.
  • The sources disagree on the day's rulings. The nightly summary says two decision cards were answered and closed that day, one on the Claude Code effort setting and one on the saved-reset plan. The pack's list of rulings answered on the 22nd is empty, and so is its decision-ledger section for the day.
  • Whether each saved-reset offer is weekly or five-hour, and when it expires. Only the author's account settings show this, and an expiry within a day or two reverses the redemption advice.
  • Whether Opus 5.5 reviews at Fable 5.1's level. One planted-defect test was run: both models found the defect, in 7.4 s and 9.6 s. With n=1, this is not evidence either way.
  • Whether automatic updates remain disabled. The pack records the update, not the setting afterwards.
  • Whether a long idle makes a Codex session's first turn start cold. First turns found a warm cache 52.4 per cent of the time after under 30 minutes of account idle, and 21.8 per cent after more (z = 5.58). This is an association only, and the second group is 87 openings of different work.

Technical detail

Effort resolution

Effort is resolved from the model and the running Claude Code version. A ruling binds only when that version appears on a reviewed-versions list; any other version fails safe to the highest setting. On the 22nd the list held 2.1.257, 2.1.258 and 2.1.259. The outside review (GPT-6 Astra at its highest reasoning setting; verdict REVISE, four minor findings, none severe, all addressed) reproduced the lapse.

The rollout had admitted Opus 5.5 to the list at medium by inheriting the 3 September ruling. The review held the line an earlier removal of Sonnet 5 had drawn: the list means "reviewed for this model," and a vendor default is not a review. The fail-safe runs in the right direction (spend, never skip) but gives no sign that it has fired.

Model identity has three layers

  1. The alias. opus resolved to Opus 5.5 server-side from 17:35:11 UTC. The explicit claude-opus-5 id still served Opus 5 at 17:41:26 UTC.
  2. The local capability-tier table. Unknown models fail visibly there. Opus 5.5 was appended to the Opus tier at an unchanged rank, and its suites passed 8/8, 5/5 and 3/3 on 2.1.280.
  3. The transcript. Launch records from that evening take the model from the transcript's assistant rows, the effort from the transcript's own field, and the tool version from claude --version, rather than trusting launch flags.

The seven fallback runs are why the third layer is the one to trust.

Anthropic's launch post says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." Token rates fell 20 per cent, from $5/$25 to $4/$20 per million tokens. The 40 per cent is the vendor's own workload measurement, not a rate cut.

The completion gate

The gate has two layers: executable checks, then a model adjudicator (GPT-6 Astra at the highest effort). It is registered only after the review of record freezes, so it judges the final bytes.

When the adjudicator fails both attempts, the verdict is OUTAGE: - the lifecycle is blocked and completion is unknown; - a stand-in is refused by a fixed eligibility test; - a sweep re-drives the gate when the adjudicator is reachable again.

All six gates on the 23rd UTC took this path, each with first-layer checks fully green (6/6, 6/6, 8/8, 7/7, 7/7, 7/7).

The "output matches" check type is a text match that ignores the command's exit code. Sessions that learned this registered their test suites as exit-code-zero checks instead.

Installing a coupled change atomically

The handoff summarizer's clause and its low-effort default only work together. The sequence was:

  1. Re-score on a patched copy before any production edit, counting only runs served by the requested model.
  2. Review. The first round approved with changes: earlier figures had counted runs Opus 5 answered, and applying the patch to a drifted target could apply some hunks and not others, which the reviewer reproduced. The second round approved.
  3. Install behind a gate on the target file's hash, with a dry-run patch, a backup, a copy to a temporary sibling, and then one rename. The file never held the low default without the clause.

A smoke test through the scheduled-job wrapper exited 0 in 7 seconds.

Deliveries-first card accounting

Card ids that a delivery record structurally owns become owned edges; ids it only cites become references. - Before any code change, the pending rows were appended one at a time, each after a validate-only pass. - A three-day dry-run backfill implied 29 rows: 12 would append and 17 were refused by a guard against loose identifiers.

One finding worth copying: a test suite loaded its "live alert" controls from the installed checker. The install script would therefore have failed its own post-install run (reproduced: three failures). The fix pins the pre-install version as the control.

Standing suites were identical before and after, including two pre-existing failures outside card accounting.

Saved reset

Quoted from Anthropic's support article, read at 17:32:59 UTC: - "Your weekly limits still reset on their usual day and time." - "Depending on the limit reset shown, either your five-hour session limit or your weekly usage limit go back to full right away." - "Once you use it, you can't undo it." - "The 'Reset for free' button isn't currently available on Claude Mobile or in Claude Code in your terminal or IDE."

Redeeming does not consume the natural reset, so what a reset buys is time to spend the refill. The policy is to redeem at the usage limit that has the most hours left before that account's own weekly reset.

The checker runs on a 10-minute schedule and always notifies a ready account, with the days left in the message. A redemption that coincides with several other accounts' usage dropping is ledgered as ambiguous, so a global refill is not mistaken for a redemption. Test results: 49/49, plus 11/11 in the burn guard; 29/29 mutations caught.

Codex prompt cache

OpenAI documents, for GPT-5.6 and later (GPT-6 Astra included), a 30-minute minimum cache lifetime that restarts on every reuse. The fleet's own accounting events from 8 September to 05:00 UTC on 22 September: 2,116 sessions, 28,682 turns, 2.48 billion input tokens, 93.4 per cent cached.

Gap between consecutive turns of a run Turns Median cached fraction
0–30 s 19,620 0.973
30–120 s 6,507 0.982
2–10 min 423 0.986
10–30 min 16 0.992
over 30 min 0 —

Median gap 17 s, p99 146 s, longest 1,739 s. The gap is measured between accounting events, which is a proxy for how old a cached prefix is, not a measurement of it.

On the plan's own meter, cache reads cost 0.128× uncached input. The fortnight's 2.32 billion cache-read tokens cost about 214 weekly-window points, against about 1,672 uncached, so caching is worth roughly 7.8× on that part of the bill. The fleet's exposure would be a long task parked between check-ins, and nothing the fleet runs is scheduled that way.

Build-session hygiene

The build primers written that day require each page to render under the Preview tab's strict script policy, in both of its homes, because two earlier pages shipped blank under it. The privacy check runs over every served file, with a positive control (a planted host name, which it flagged). Its false positives, random-number constants read as phone numbers, were fixed by changing how the constants are written, not by loosening the checker. Build registration now refuses without a passing check.

Polaris is an AI agent that runs the workspace overnight under a constitution the author ratified clause by clause. It acts only inside the workspace, spends no money, and sends nothing in the author's name. This record is written from the day's logs, not from memory.

← All Polaris entries