Historical Nanochat
Time-locked language models trained on pre-cutoff historical texts using Karpathy's nanochat pipeline. Exploring whether small models trained exclusively on period texts can reproduce the linguistic patterns of their era.
- 65GB historical text corpus across multiple eras
- Time-locked training methodology (a publication-date filter; a 2026 audit removed misdated volumes, 0.28% of the training cache)
- RTX 3090 local training pipeline
- Parquet-based shard management
Activity Timeline
-
Google Drive offload audit found four major bugs.
A dry-run guard typo in verify-and-stub.sh would trigger real execution. The runner.sh liveness monitor counted bytes instead of files, and the space math was off by about 27 GiB on hard-link removal.
-
The Google Drive offload freed 686.7 GiB of 714 GiB planned, and review found four major issues.
Issues include a stalled-detection loop that failed at 1.1 MB, a dry-run flag typo, and incomplete bundle validation. The tmp age-out rule is posted but not installed, and the bundling-route ruling is still pending.
-
Code review found four major defects blocking the storage offload plan.
A dry-run flag typo would execute the real run instead of a rehearsal, and liveness monitoring counts the wrong metric. Hard links mean corpus removal frees 686.7 GiB, 27 GiB less than planned, and the bundle-decryption guard's validation is incomplete. Three commits folded in review rounds and an interwar paradigm rework with notebooks.
-
Offload review: 4 major and 19 minor findings.
Major issues include a runner.sh silent stall with a 7-day liveness gap and a dry-run typo that executes a real run.
-
Security audit of verify-and-stub.sh: 0 blockers, 4 MAJOR, 19 MINOR findings.
Liveness check loop silent ~7 days. Dry-run flag typo. 27 GiB hard-link retention vs 714 GiB plan. Bundle-decrypt validation incomplete. Offload blocked.
-
Code review of Google Drive migration scripts found 4 major bugs; disk estimate corrected to 686.7 GiB.
Three scripts (verify-and-stub.sh, runner.sh, bundle-decrypt-test.sh) reviewed; issues include loop stall, dry-run typo, disk-space accounting error, and insufficient validation. Hard-link pairs reduce freed space from estimated 714 GiB to ~686.7 GiB.
-
Four blocking issues found: silent liveness stall, dry-run flag typo, freed-space math error, permissive decrypt guard.
7-day silent liveness stall, dry-run flag typo, 27 GiB freed-space math error, and permissive bundle-decrypt guard must all be fixed before execution can proceed.
-
Shell script review found byte-count error in stall detection, dry-run typo, 27.25 GiB space undercount, and insufficient bundle-decrypt validation.
Review completed, execution not yet attempted. Stalled-transfer detection counts bytes instead of files, dry-run flag typo causes real execution, freed space underestimated by 27.25 GiB, and bundle-decrypt validation is insufficient.
-
13 Jev claims validated; 129 BHL volumes removed (0.278% token reduction); P0 gate findings driven from 9 to 0.
Jev-nanochat campaign: 7 confirmed, 3 caveated, 2 corrections filed. BHL-139 cleanup took corpus from 19.1B to 19.0B tokens. Gate-small-fixes-r2 four-round repair campaign cleared all P0 severity findings.
-
129 BHL volumes removed, 53M tokens dropped; Drive offload scripts held pending fixes.
Gate verdict PASS after dropping 0.278% of corpus. Safety review of Google Drive offload scripts found 4 major and 19 minor issues before any upload ran. Jev-Nanochat data verification completed: 13 claims checked, 2 corrections required.
-
13 claims verified, corpus offload scripts found 4 major issues, classification rules rewritten.
Transcript review resolved 7 claims confirmed, 3 with caveats, 2 requiring correction. Drive migration scripts had a silent failure loop, fail-open dry-run gate, 27 GiB miscalculation, and bundle-decrypt gaps. nanochat-bhl-139 classification rules and test fixtures unified.
-
7/8 data claims verified; drive offload 4 issues found; 19 P0 nouns and 10 P1s resolved.
Data report verification confirmed 7 of 8 claims (3 with caveats, 2 corrected). Drive offload review found failure detection gap, dry-run typo, hard-linked files, and bundle validation issues. Gate-small-fixes-r2 resolved 19 P0 storage nouns and 8 P1s; data-quality audit corrected 2 more P1 findings.
-
Corpus offload halted by 4 safety blockers. Research dive produced 13,993-word report on LLM judges.
Four issues prevented safe execution of the corpus offload: liveness monitoring failure, dry-run flag typo, storage math error (686.7 GiB vs 714 GiB planned), and weak validation guards. No data deleted. Separate Fable 5.1 dive authored a 13,993-word report on LLM judges in time-locked models.
-
Security review found 4 blocking issues in offload scripts; research dive confirmed no LLM judge exists for historical anachronism detection.
Blocking issues: stalled-file loop (runner.sh:101), --dry-run typo (verify-and-stub.sh:16), space accounting discrepancy (686.7 vs 714 GiB), overly permissive decrypt guard (verify-and-stub.sh:88-91). Offload blocked pending fixes. Research found no published LLM-based anachronism judge; ChronoGPT-Instruct and TypewriterLM pair LLM with human review.
-
No blockers, but 4 major issues block offload: loops, dry-run guard, space math, validation.
A runner.sh miscount disables stalled-file detection for ~7 days, hiding the root problem. Silent re-read loops, a broken dry-run guard, space-math mismatch, and weak validation must be fixed before the offload proceeds.
-
Terabyte disk audit complete: nanochat data found on root disk, not audit target.
1.27TB / 3.1M files analyzed. Audit target holds only 93.8MB transcripts + 7.4GB uncovered user files. 715GB nanochat dataset confirmed on root disk.
-
Repo campaign complete: main sanitized, secret scanning verified on 5 repos.
nanochat main scrubbed via sanctioned publish path and released to historical archive. Secret scanning enabled and verified via gh API on 5 repos. psyche-public SECURITY.md/README contradiction and ETHICS-PROTOCOL staleness resolved.
-
55.16 GiB freed across 8 cache categories; all services healthy.
Root partition reduced from 94% to 91% full (114G → 169G free). Live service dependencies verified before deletion including uv-cached venvs and Maya1 voice model; all services confirmed healthy post-cleanup.
-
Storage Phase 1a: 55.16 GiB reclaimed from 94%-full partition.
npm, HuggingFace strike-list, Docker, pip, pnpm, and Trash caches verified and purged. maya1 voice model held due to live config and service code dependencies.
-
Storage audit: root at 89% full; 715G dataset marked NO-GO.
Disk audit identified largest consumers: 715G dataset (execution blocked), 362G workspace, 171G cache. Post-incident CUDA toolchain drift (13.1→13.3, driver 590→595) means original training run is no longer reproducible bit-identical.
-
STT degradation root-caused to mSBC codec; VRM avatar stack research initiated.
mSBC Bluetooth codec confirmed as primary STT quality degradation source (7 kHz vs 8 kHz effective bandwidth). LibriSpeech harness built for WER measurement against 354 voice clips. Three VRM/avatar repos identified for image-to-3D pipeline feasibility.
-
API key + staging IP scrubbed; HIGH vulns 42→0.
Plaintext API key and live staging IP removed from 6 published files (23 substitutions). Key rotation required before containment. Dependabot sequence complete, total vulnerabilities 94→11.
-
Closed all 42 HIGH Dependabot alerts; vulnerability count dropped from 94 to 11.
Torch bumped 2.9.1 → 2.13.0 with smoke tests. fastapi 0.140 / starlette 1.3 compatibility confirmed across 73 tests. Chat web server hardened to bind localhost-only by default.
-
All 8 P0 remediation issues verified closed; gated for tier-2a smoke testing.
Empirical verification via code inspection and pytest confirmed every fix. Checkpoint timing defect (P0-1) resolved by consumed_loader_state tracking across base_train.py. Smoke test parameters scoped: SAVE_EVERY=250, MAX_STEPS=300.
-
All 8 P0 defects verified closed; system cleared for capped-smoke testing.
Checkpoint-ahead-of-consumption fixed via separate consumed_loader_state tracking (base_train.py:517-523). Sol proxy stalled; pivoted to direct CPU-side verification with test suite tripwires confirming each defect empirically. Moves to tier-2a: GPU canary assertions and CUDA behavior verification remain.
-
All 8 P0 defects verified closed; ns-r7 remediation and Hub M2 reconciliation complete.
Fleet monitoring operational with continuous heartbeat. ns-r7: 3 locked items closed, 17 test failures resolved to passes. Hub M2: 41-pass baseline established, F1/F4-F9/F11/F13 defects closed in plan.
-
All 8 P0 defects independently verified closed; status advanced to READY-FOR-CAPPED-SMOKE.
Final P0 (checkpoint prefetch tracking via consumed_loader_state) validated by passing test suite. CPU-side work complete. GPU-side canary run pending with capped params (SAVE_EVERY=250, MAX_STEPS=300).
-
All P0s empirically closed; READY-FOR-CAPPED-SMOKE verdict issued.
Independent sol-reverify-d26 session confirmed all SOL-PLAN-REVIEW P0 findings closed via direct code inspection. P0-1 checkpoint prefetch race covered by new test parametrizations in base_train.py. Tier-2a smoke test phase cleared for launch at SAVE_EVERY=250, MAX_STEPS=300.
-
P0 remediation complete: 8/8 defects verified GREEN, transitioned to capped-smoke testing.
Nine commits across checkpoint consumed-cursor fix, launcher hardening (5 defects), and training guards. RED/GREEN verification confirmed per defect. Remediation phase officially closed.
-
d26 training run fully staged; blocked on external GPU provider account setup.
Cache validation passed, owner actions documented in NEEDS-OWNER file, systemd monitoring timer installed. Launch gated on Hyperbolic account email verification and payment method.
-
Architecture investigation opened for Design-C shard-ordering; conditional GO.
Bake script and CPU-only traversal simulator gating specified. GPT-Pro brainstorming on cloud run efficiency optimization from contemporary literature queued.
-
Security scrub complete: 80+ files cleaned, serve.py hardened, all P0/P1/P2 findings resolved across two audits.
Blind Fable follow-up review found trust_remote_code RCE vector and Windows username leak missed by initial pass — both fixed. SECURITY.md created documenting sandbox design boundary. Git history rewrite still pending.
-
Phase 1 complete; two prior claims retracted, core affective finding validated.
8 commits correcting talkie-conversion and post-1930 fracture claims. Affective divergence (providence/duty vs. therapeutic) and era-based Family F clustering confirmed robust. Phase 2 direction crystallized: pre-1914 vs. modern characterology.
-
Training outcomes reviewed via 5-model multi-agent analysis; GPT Max decision framework documented.
Multi-agent review (Opus, GPT Max, GPT Council, GPT Pro, Opus 4.7) of nanochat training results. Key output: cost-tiered skill selection framework distinguishing GPT Max (13×, high-stakes disagreement) from codex-council (5×, initial lookups).
-
ChatGPT Pro MCP: better-playwright selected; 2 critical issues found in code review.
Orphaned tab memory leak and missing transport retry logic identified. Stepped timeout architecture designed (30–120 min). Fixes specified, pending implementation.
-
ChatGPT Pro MCP server built for browser-based GPT-5.4 Pro access; two critical bugs block production use.
Three-layer completion detection with timeout polling implemented. Architecture validated clean by code review. Blocking issues: page leak from orphaned Chromium tabs, no retry on transport failure.
-
ChatGPT Pro Browser MCP built; critical resource leaks found; 499GB data migration completed.
MCP server enables GPT-5.4 Pro via browser automation. Code review identified page leak (Chromium tabs never closed) and missing retry logic for dropped responses. Training data migrated from Windows NTFS to native Linux ext4.