Reference wiki
A reference library for the ideas behind the AI moment. Where a term deserves a real explanation — what a scaling law is, what RLHF does, what an agent harness contains — the explanation lives here. A guided reading section built on this wiki is in the works.
These 195 articles are AI-researched and citation-rich: each was written by frontier models working from primary sources, then synthesized into a single reference entry. They are reference material, not essays — treat them as a well-sourced starting point and follow their citations for load-bearing claims.
- Abstract
- Agent Harness: The Software Surrounding the Model
- Agent Memory Architectures: Working Set, Episodic, Semantic, and Beyond
- Agent Scaffolding
- AI Coding Agents in 2026: Architectures, Capabilities, Bottlenecks
- AI Control: Safety When Trust Cannot Be Verified
- Ai Evals
- Aider: Git-Native Coding Pair Programmer
- AIME: American Invitational Mathematics Examination as AI Benchmark
- Algorithm Appreciation: The Asymmetric Counterpart to Algorithm Aversion
- Algorithm Aversion: Why Humans Reject Helpful Models
- Alignment Faking: When Models Comply During Training and Defect Later
- AlphaGeometry
- AlphaGo
- AlphaGo Zero
- AlphaProof: Search and Self-Play for Formal Mathematics
- AlphaZero
- Anchored Rubrics: Calibrated Scoring Scales for LLM Evaluation
- Apparent Personality from Text: LLM-Based Personality Inference
- ARC-AGI
- Attachment Theory And Ecr R
- Automation Bias
- Behavioral Contracts For AI
- Benchmark Contamination: Detection, Mitigation, and Why It Keeps Happening
- Benchmark Overfitting: When the Model Memorizes the Test Set
- Benchmark Saturation: Capability Ceiling, Contamination, or Measurement Collapse?
- Best-Of-N Sampling
- Big Five Personality Model
- BrowseComp
- Calibrated Challenge: Anti-Sycophancy as a Positive Design Target
- Capability Elicitation: The Other Half of Capability Evaluation
- Capability Evaluation: Distinct from Safety, Reliability, and Behavior Evaluation
- Caricature Under Persona Conditioning: When Persona Prompts Produce Stereotype
- Chain-of-Thought Faithfulness: When the Stated Reasoning Doesn't Match the Actual Decision
- Chain-of-Thought Monitorability: Can We Trust Visible Reasoning?
- Chain-of-Thought Monitoring: Practice and Limits
- Chain-of-Thought Prompting: From Wei et al. 2022 to Internalized Reasoning
- Claude Code
- Claude Extended Thinking
- Cluster Bootstrap For Hierarchical Eval
- Code Interpreter
- Codex CLI
- Computerized-Adaptive-Testing
- Constitutional AI: Method, Evolution, and Critique
- Construct Validity in AI Evaluation
- Context Engineering for AI Agent Systems
- Context Window Economics: The Hidden Cost Structure of Large Language Model Applications
- Convergent and Discriminant Validity: The Multitrait-Multimethod Matrix
- Cursor Agent Mode: IDE-Native Autonomous Coding
- Deceptive Alignment: The Worst-Case Mesa-Optimization
- DeepSeek-Prover
- DeepSeek-R1
- DeepSeek-V3
- DeepSeekMoE: Mixture-of-Experts with Fine-Grained Routing
- Deliberative Alignment: OpenAI's Method for o-Series Reasoning Models
- Deployment Overhang
- Distribution Shift: When Test Data Diverges From Training
- DSPy
- Empath: Neural-Extended Lexical Categories for Text Analysis
- Eval-Driven Development: Tests for AI Behavior
- Evaluation Awareness: When Models Recognize They're Being Tested
- Evaluation Harness: Tooling for Reproducible LLM Eval at Scale
- External Validity in AI Evaluation
- FrontierMath: Research-Level Mathematics Evaluation
- Function Calling: The API Surface of Tool Use
- Gemini Deep Think
- Gemini Thinking Models: Architecture, Reception, Distinct Choices
- Goodhart's Law in AI Systems: When the Metric Becomes the Target
- GPQA
- GPQA Diamond
- GPT-5
- Graded Response Model
- Graph of Thoughts: Search Generalization Beyond Trees
- Guardrails
- HEXACO: The Six-Factor Personality Model
- Human-AI Reliance: When Help Becomes Dependence
- Human-in-the-Loop Evaluation: When Human Ratings Are the Ground Truth
- HumanEval: The First Code-Generation Benchmark
- Humanitys-Last-Exam
- In-Context Learning: The Surprise That Shaped LLM Practice
- Incremental Indexing: Keeping Embedding Stores Fresh
- Inference Scaling Laws
- Inference-Time Search
- Infrastructure-Noise
- Instruction-Hierarchy
- Inter-Rater Reliability
- Item Response Theory
- Judge Halo Effects in LLM-as-Judge Evaluation: Exact-Self, Same-Provider, Cross-Provider
- LangGraph: Stateful Workflow Orchestration for LLM Agents
- Language Agent Tree Search
- LiveBench
- LiveCodeBench: Contamination-Resistant Code Benchmarking
- Liwc Linguistic Inquiry
- LLM-As-Judge
- Local-First AI Tooling: When Does Local Inference Win?
- Long-Horizon Agents: Where Capability and Reliability Diverge
- Mechanistic Interpretability
- Mesa-Optimization: Learned Optimizers Inside Trained Models
- Mini-SWE-Agent
- MIPRO
- Mixture-Of-Experts
- MMLU-CF
- MMLU-Pro: The Less-Saturated Successor
- MMLU-Redux: Cleaning Up the Item Pool
- MMLU: Anatomy of a Saturated Benchmark
- MMMU: Massive Multi-discipline Multimodal Understanding
- Model Routing: When and How to Send Different Queries to Different Models
- Monte Carlo Tree Search in 2026: From AlphaGo to LLM Reasoning
- Moral Foundations Theory: The Pluralist Model of Moral Cognition
- MRKL Systems: Modular Reasoning with Knowledge and Language
- Multi-Head Latent Attention
- Multi-Model Agent Orchestration Patterns
- Multi-Token-Prediction
- Mutable Scaffolds: When the Software Layer Self-Modifies
- MuZero
- Need for Cognition: The Trait of Engagement With Effortful Thought
- Open-Vocabulary Personality Models
- OpenAI O1
- OpenAI O3
- OpenCode
- OpenHands
- OSWorld: Benchmarking Computer-Use Agents
- Pairwise Versus Scalar Evaluation: Tradeoffs in LLM Output Judging
- Persona Prompting: System-Prompt-Based Personality Conditioning of LLMs
- Preparedness, Responsible Scaling, and Frontier Safety: Comparing Lab Governance Frameworks
- Process Reward Models: Stepwise Supervision for Reasoning
- Process Supervision vs Outcome Supervision: A Training-Signal Distinction
- Production LLM Evaluation: Offline, Online, and the Hybrid Pattern
- Prompt Caching Across Providers
- Prompt Engineering: Practice, Decline, and What Remains
- Prompt Injection in Agentic Systems: Attack Surfaces, Defenses, Residual Risk
- Prompt Optimization
- Psychometric Correlates of AI Interaction Styles
- Public-Archetype Echo: When AI Assistants Mimic Famous Persona Surfaces
- ReAct Agents
- ReAct: The Reasoning-Acting Synthesis
- Reasoning Models in 2026: o1/o3, DeepSeek-R1, Gemini Thinking, Claude Extended Thinking
- Reasoning Tokens: API Semantics, Billing, and Visibility
- Reasoning via Planning: PDDL-Style Decomposition in LLM Agents
- Recursive Self-Improvement: Theory, Demonstrations, Limits
- Reflexion: Verbal Reinforcement for Language Agents
- Regression Testing for AI Systems: Beyond Snapshot Tests
- Repair-Oriented Response
- Repository Branch Awareness in Coding Agents
- Retrieval-Augmented Generation
- Reward Hacking and Specification Gaming: Patterns, Evidence, and Mitigations
- Reward Tampering: When the AI Modifies Its Own Training Signal
- ReWOO: Reasoning Without Observation
- Rich Sutton's "The Bitter Lesson" (2019): Thesis, Reception, and Counter-Arguments Through 2026
- Rich Sutton’s “The Bitter Lesson” (2019): Thesis, Reception, and Counter-Arguments Through 2026
- RLAIF: Reinforcement Learning from AI Feedback
- RLAIF: Reinforcement Learning from AI Feedback
- RLHF: Reinforcement Learning from Human Feedback
- Safety Evaluation: Capability-Misuse, Alignment, or Behavior?
- Sample-Efficient Personality Inference
- Sandbagging: Strategic Underperformance in AI Systems
- Scalable Oversight: Supervising AI That Outpaces Human Judgment
- Search and Learning: The Bitter Lesson’s Two Engines
- Self Determination Theory
- Self-Consistency
- Self-Improving Software Systems: From EURISKO to Modern AI Agents
- Semantic Search Architecture for Code and Documentation: Embeddings in Practice
- Specification Gaming: When the AI Optimizes the Letter, Not the Spirit
- Style Mimicry Overfit: When LLMs Match Surface Style Over Content Quality
- SWE-agent: The First Frontier Agent Harness for Software Engineering
- SWE-Bench
- SWE-bench Live: Continuously Refreshed Software-Engineering Eval
- SWE-bench Pro: Larger Repositories, Harder Issues
- Swe-Bench-Verified
- Sycophancy in Large Language Models
- Synthetic Personas for AI Evaluation: Methodology and Limits
- Synthetic-First Evaluation: Methodology Before Human Studies
- System Capability Vs Model Capability
- Terminal-Bench
- Test-Time Compute
- Test-Time Training
- The Agent-Computer Interface: Browser, Terminal, API, GUI, and Beyond
- The Cognitive Reflection Test: Measuring Override of Intuitive Responses
- The Dark Triad: Machiavellianism, Narcissism, and Psychopathy
- The Model Context Protocol (MCP) Ecosystem as of 2026
- The Pruning Principle: When Scaffolding Becomes Dead Weight
- Theory of Mind in Large Language Models: The Contested Empirical Question
- Thomistic Natural Law Theory Applied to AI Agent Decision-Making
- Tool Calling Semantics: Cross-Provider Comparison
- Tool Use in Language Models: From ReAct to Native Tool Calling
- Tool-Augmented Reasoning: When the Model Calls Code
- Toolformer
- Tracked Capability Levels: Governance Through Threshold Tiering
- Tree of Thoughts: Deliberate Reasoning Through Search
- Trust Calibration in Human-AI Systems
- User Modeling In Llms
- Vector Index Design: HNSW, IVF, Flat, and the Cosine-vs-Dot Decision
- Verifier-Guided Search in LLM Reasoning
- WebArena
- Wilson Confidence Intervals