Reference

Test-Time Compute Scaling

Test-time compute scaling is the practice of improving an already-trained model's answer by spending more computation on the particular problem at hand. The added budget can buy a longer attempt, more independent attempts, revisions, verifier judgments, tree search, tool use, or an adaptive mixture of these. The key word is scaling: the engineering question is not merely whether one extra pass helps, but how capability, latency, and cost change as the inference budget grows.

This is an umbrella entry. Test-Time Compute gives the conceptual and product-history overview, while Inference Scaling Laws examines whether the observed performance curves deserve to be called laws. This entry instead organizes the design space and answers a practical question: what exactly is being scaled, when should each mechanism work, and what evidence would justify spending the next unit of compute? The alias inference-time-scaling is absorbed here because the literature uses test-time and inference-time scaling for substantially the same family of methods.

Coverage note: This article reflects sources and terminology reviewed through August 8, 2026.

The best-supported conclusion is narrower than the slogan “thinking longer makes models smarter.” Extra inference compute often improves performance, especially on difficult but tractable tasks with reliable verification. The gain is neither universal nor automatically monotonic. It depends on the base model already assigning some probability to a good solution, on producing sufficiently diverse candidates, on recognizing the good candidates, and on finishing within the real latency and resource envelope. Repeated sampling can raise coverage while final-answer accuracy stalls; a weak verifier can turn more candidates into more opportunities for selection error; and longer reasoning can repeat or destabilize an answer rather than repair it. (Large Language Monkeys; Compute-Optimal Test-Time Compute; Not That Simple)

The object being scaled

Training-time scaling changes the model's weights. Test-time scaling holds the trained model fixed and changes the inference procedure for one input. In practice the boundary is a systems boundary rather than a metaphysical one: the policy may include a language model, reward model, code runner, proof checker, retrieval service, or other tools, but no additional gradient update to the deployed model is required for the current answer.

That last clause is a scoping choice, and the boundary it draws is contested rather than settled. Test-time training — taking per-task gradient steps on the examples of the problem at hand — spends its budget at inference too, and it was one of the two approaches behind the leading ARC-AGI-1 results before o3 — ARC Prize's own 2024 report names “deep learning-guided program synthesis” and “test-time training (TTT) for transductive models” together, noting that “each approach tends to solve different kinds of tasks”; the top private-eval score that year was MindsAI's 55.5% and the open-source winner reached 53.5%. (ARC Prize 2024) Some surveys file it under test-time compute; others, and the companion entry Test-Time Training, keep it outside on the grounds that the weights are supposed to stay fixed. This entry takes the weights-fixed reading so that the mechanisms below stay comparable on a single axis, and says so rather than leaving the exclusion implicit.

That definition includes several budgets that are too often conflated:

Budget axis What increases Typical mechanism Principal bottleneck
Serial depth tokens or steps in one trajectory hidden deliberation, iterative revision, tool loop repetition, context growth, wrong-path lock-in
Parallel breadth independent or diversified trajectories repeated sampling, self-consistency, best-of-N correlated samples and selection
Selection effort judgments per candidate outcome verifier, process reward model, generative verifier verifier accuracy and compute cost
Structured search states, branches, or backtracks explored beam search, Tree of Thoughts, theorem-proving search state design and value estimation
Environmental interaction observations and actions code execution, browsing, simulation, tests tool reliability and feedback quality
Adaptive allocation budget shifted among prompts or branches difficulty routing, early stopping, tail-guided search estimating value of more compute

This taxonomy matters because equal token counts do not imply equal computation in the useful sense. Ten identical solutions add little information. Ten diverse solutions checked by executable tests may add a great deal. A 20,000-token trace that remains on one mistaken branch is not equivalent to twenty 1,000-token attempts. Snell et al. found that the useful balance between sequential revisions and parallel samples varies with problem difficulty; their adaptive strategy used roughly four times less test-time computation than a best-of-N baseline at comparable performance (the paper states this variously as “up to 4x less” in its results sections, “more than 4×” in its abstract and “a factor of 2−4×” in its discussion). (Compute-Optimal Test-Time Compute)

OpenAI's o1 disclosure made the axis visible in a frontier product: it reported smooth improvement with more time spent thinking, evaluated most headline results at maximal test-time compute, and showed AIME 2024 rising from 74% with one sample to 83% with consensus over 64 samples and 93% when reranking 1,000 samples. Those results establish that the inference procedure materially affects the reported capability. They do not reveal a provider-independent law, because the model, training, hidden reasoning policy, sampler, and scoring model all changed the effective system. (OpenAI o1)

A proposer-selector model

Most test-time scaling systems can be understood as two coupled components:

  1. A proposer produces one or more candidate trajectories or actions.
  2. A selector decides which candidate to continue, revise, execute, or return.

The proposer may be the same language model at different temperatures, a revision policy, or a tree-search expansion rule. The selector may be answer frequency, a learned verifier, an exact checker, unit tests, a process reward model, or another language model. The decomposition is useful even when both roles are performed by one model, because their failure modes differ. More proposal compute primarily increases the chance that a good candidate exists; more selection compute primarily increases the chance that the system recognizes it.

For an idealized problem where each independent attempt succeeds with probability p, the probability that at least one of k attempts succeeds is:

pass@k = 1 - (1 - p)^k.

This explains both the attraction and the diminishing returns of repeated sampling. If p is nonzero, coverage rises with k; but every additional sample buys less than the previous one. The equation is an upper-level intuition, not a deployment forecast. Real samples are correlated, task difficulties vary, and the system usually lacks an oracle that identifies the correct member of the set. Brown et al. therefore separate coverage—whether any candidate is correct—from precision—whether the selection procedure can find it. They observed smooth coverage gains over four orders of magnitude, yet majority voting and reward-model selection plateaued after hundreds of samples in settings without dependable automatic verification. (Large Language Monkeys)

This yields a useful four-condition model. Test-time scaling is likely to work when:

  • the base proposer has non-trivial support for a correct or useful trajectory;
  • additional compute creates diversity rather than near-duplicates;
  • the selector has enough discrimination to favor better candidates;
  • the task admits enough feedback to redirect effort before the budget is exhausted.

If any condition is absent, the curve can flatten. No amount of sampling recovers a solution the model never proposes. Diversity without selection raises pass@k but not delivered accuracy. A selector that rewards superficial form can choose polished mistakes. Feedback that arrives only after a costly full trajectory may be too late to make search economical. This is why the strongest results cluster in mathematics, code, formal proof, and puzzles: these domains offer exact answers, executable tests, or at least relatively crisp reward signals. (Large Language Monkeys; Training Verifiers)

The major scaling mechanisms

Parallel sampling and self-consistency

The simplest mechanism is to sample several complete solutions. With an exact checker, return a verified solution. Without one, a common approximation is Self-Consistency: sample diverse reasoning paths and choose the most frequent final answer. Wang et al. introduced this as a decoding strategy and reported substantial gains across arithmetic and commonsense benchmarks, including a 17.9 percentage-point gain on GSM8K in their experiments. (Self-Consistency)

Parallel sampling is attractive because it is easy to batch, does not require the candidates to share mutable state, and makes the budget explicit. It also exposes a hard limit. Majority vote assumes that correct reasoning tends to converge on the same answer while errors disperse. That assumption works better for questions with a unique short answer than for open-ended tasks, correlated misconceptions, or prompts where the model confidently repeats one wrong heuristic. Brown et al.'s large-scale study found that coverage kept increasing after common selectors had plateaued. The system was generating more useful information than its selector could exploit. (Large Language Monkeys)

Repeated sampling is therefore a strong baseline, not a complete architecture. It is particularly appropriate when attempts are cheap relative to verification, when outputs can be executed independently, or when one needs an honest pass@k estimate. It is weaker when latency cannot be parallelized, when candidates share the same systematic error, or when judging a candidate is as hard as generating it.

Verifier-guided selection

A verifier changes the question from “which answer is most common?” to “which answer appears correct?” Cobbe et al. generated many candidate solutions for GSM8K and trained a verifier to rank them, demonstrating that verification could improve final performance. This result helped establish the modern best-of-N pattern: a proposer supplies breadth, and a learned judge converts coverage into delivered accuracy. (Training Verifiers)

Verifiers range from strong grounding to weak proxies. A compiler, proof assistant, or comprehensive test suite can reject candidates through contact with an external rule system. An outcome reward model sees the completed answer. A Process Reward Model scores intermediate steps. A generative verifier spends its own reasoning tokens explaining or checking a candidate. These mechanisms should not be grouped as equivalent: their error correlations, costs, and vulnerability to reward hacking differ.

More verification is not automatically the best use of a fixed budget. Singhi et al. compared self-consistency with generative reward-model verification under compute matching. In their experiments, the generative verifier needed up to eight times the inference compute to match self-consistency at practical budgets, although it could become advantageous at much higher budgets. Their fitted allocation favored scaling the number of solutions faster than the number of verifications. (Solve or Verify)

There is a complementary result rather than a contradiction. Setlur et al. argue theoretically and empirically that verifier-based reinforcement learning or search can scale better than merely cloning successful traces — under two conditions that are both properties of the base model rather than of the verifier: its correct solution traces are heterogeneous, and its rewards on sampled traces are not sharply concentrated around their mean, which the paper formalizes as anti-concentration. In their compute-matched s1 comparison, best-of-N with a trained verifier beat budget forcing. The two papers address different comparisons: verification can be essential for long-run search while generative verification calls can still be an expensive marginal use of compute. (Verification or RL)

Serial deliberation and revision

Serial scaling spends budget inside one evolving trajectory. The model may generate a longer chain, critique its answer, revise it, call a tool, or continue after a stopping point. This is closer to what product interfaces call “reasoning effort,” but output length alone is a poor proxy for productive work.

The s1 project provided a deliberately simple demonstration. After fine-tuning on 1,000 curated reasoning examples, it controlled inference length by truncating a trace or appending “Wait” when the model attempted to stop; its reported AIME24 score rose from 50% to 57% under budget forcing. (s1) Later analysis by Guojun Wu disputed the strong interpretation: the observed scaling was largely attributable to restricting shorter budgets, while repeatedly appending “Wait” could make answers oscillate rather than exceed the model's natural peak. The dispute illustrates an important measurement distinction between interpolation—recovering performance as an artificial cap is relaxed—and extrapolation—continuing to improve beyond budgets represented in training. (Not That Simple)

DeepSeek-R1 provides stronger evidence that the ability to use longer traces can be trained. Its R1-Zero model used reinforcement learning without supervised fine-tuning as a cold start; during training, response lengths increased and behaviors described by the authors as reflection and alternative exploration appeared. The paper also reports majority-vote gains on AIME, so R1 illustrates both serial and parallel scaling rather than isolating one. Its evidence supports the claim that training and inference policy must be co-designed; it does not prove that every extra token is a faithful or necessary reasoning step. (DeepSeek-R1)

Recent e3 work defines the demanding target explicitly as extrapolation beyond the token budget used in training. It reports that most reasoning models do not extrapolate well and trains a small model to chain generation, verification, refinement, and hypothesis testing, improving out to twice its training token budget. Because this is a particular recipe on particular math evaluations, confidence in the general mechanism is moderate rather than high, but it strengthens the case that productive long-horizon inference is a learned skill rather than an automatic property of autoregressive generation. (e3)

Tree of Thoughts makes the search structure explicit. Instead of committing to one left-to-right chain, the system represents intermediate “thoughts” as states, expands multiple continuations, evaluates them, and can look ahead or backtrack. Yao et al. reported a large gain on the Game of 24, from 4% for chain-of-thought prompting with GPT-4 to 74% for their Tree of Thoughts setup. The result is a useful existence proof, but it came from three selected tasks with hand-designed decomposition and evaluation procedures; it should not be read as a universal 70-point search bonus. (Tree of Thoughts)

Structured search can spend compute more intelligently than blind sampling when partial states expose useful signals. It can also fail more elaborately. A poor value model prunes the right branch. A bad state representation hides progress. Search can optimize a process reward model into repetitive or unnaturally short traces; Snell et al. observed both patterns in their analysis. The more complex the search controller, the more important it becomes to compare against simple, compute-matched baselines rather than against one greedy answer. (Compute-Optimal Test-Time Compute)

Tools turn the search space into an environment. Code execution, retrieval, proof checking, and simulators can supply observations unavailable from the model's token probabilities alone. This can make verification genuinely easier than generation, but only if the environment is trustworthy. Brown et al. found flaky tests in SWE-bench Lite and documented that even apparently objective software verifiers can misclassify candidates. “Tool-grounded” is therefore an empirical property of the tool and harness, not a synonym for correct. (Large Language Monkeys)

Compute-optimal allocation

Uniformly giving every prompt the same budget is simple and often wasteful. Easy prompts saturate quickly. Some very hard prompts remain outside the proposal distribution. The largest marginal returns often occur in the middle: prompts where the model has a non-trivial chance of success but benefits from search or correction. Snell et al. used model-relative difficulty to select strategies and allocate compute, obtaining a roughly fourfold efficiency improvement over a uniform best-of-N baseline in their math setting — with a caveat that decides whether the number transfers to a deployment. Their difficulty estimate is taken from 2,048 samples per problem, and they say plainly that “our experiments do not account for this cost largely for simplicity.” The paper's own summary of the range is “a factor of 2 − 4×.” Read it as what compute-optimal allocation can be worth once you already know each problem's difficulty, not as an end-to-end efficiency figure. (Compute-Optimal Test-Time Compute)

This suggests a routing policy rather than a fixed “reasoning level”:

  1. Produce a cheap initial attempt and uncertainty or difficulty estimate.
  2. Stop when expected value of more work is below its cost.
  3. Add serial depth when local repair seems plausible.
  4. Add parallel breadth when alternative approaches are valuable.
  5. Spend verifier effort only when candidate diversity has created a selection problem.
  6. Escalate to tools or structured search when intermediate feedback is reliable.

The ideal router would estimate marginal value of compute: the expected improvement from the next token, sample, verifier call, or tool action. Current systems approximate this with confidence, reward margins, problem features, or fixed tiers. Tail-guided search is a 2026 proposal to estimate the upper tail of intermediate reward distributions and allocate search toward states with the highest predicted potential; the authors provide theoretical guarantees under their assumptions and report better reward than best-of-N at matched budgets. Read the metric before reading the result: the outcome reward model “serving as the oracle for our search optimization” is also what scores the outcome, since the primary metric is “the total reward … the accumulated sum of maximum rewards across the test set” rather than answer correctness on AMC or AIME. Higher reward under the model that steered the search is the shape this article's own abstract warns about — "a weak verifier can turn more candidates into more opportunities for selection error" — so this is evidence that the allocation rule optimises its objective, not yet evidence that it delivers more right answers. (Tail-Guided Search)

What counts as compute?

The phrase “same compute” can hide the conclusion. FLOPs, generated tokens, samples, wall time, energy, memory traffic, and retail dollars are different denominators. A smaller dense model sampled many times may win a FLOPs calculation yet lose on memory movement, attention cost, orchestration overhead, or tail latency. A vendor price comparison adds batching and business policy. A sample count ignores variable trace length. Any claimed frontier should state at least the model, inference policy, budget unit, parallelism assumptions, and completeness rate.

Wu et al. compared model sizes and inference strategies under compute budgets and found cases where a 7B model with tree search outperformed a 34B model using standard strategies. (Empirical Inference Scaling) Kinetics later argued that FLOPs-only accounting can overstate the advantage of smaller models because memory access and attention become important during long or repeated inference; its experiments proposed a different frontier once those costs were included. (Kinetics) These results are not cleanly contradictory. They show that “compute optimal” is conditional on the cost model.

ARC Prize's o3-preview evaluation is the clearest public illustration. On its semi-private ARC-AGI-1 set, increasing from 6 to 1,024 samples—about 172 times more compute—raised the reported score from 75.7% to 87.5%, while the estimated cost per task rose from $26 to $4,560 under the later pricing assumptions on the page. ARC also disclosed that the tested model was trained on part of the public ARC data and that the released o3 was not the same system. The result is strong evidence that inference budget changed performance for that evaluated system, and equally strong evidence that score without cost, model identity, and exposure details is incomplete. (ARC o3 Evaluation)

Scaling policy is a system-level decision

A deployed system rarely faces one isolated prompt with a fixed allowance. It faces a stream of requests competing for accelerators, verifier capacity, tool quotas, and latency budgets. The operational question is therefore not merely whether a task improves at 64 samples, but which task should receive the next unit of work. A policy that gives every request the benchmark-winning maximum can lower total utility by delaying easy requests, exhausting capacity before genuinely difficult ones arrive, or spending heavily on cases whose success probability is effectively zero.

This changes what an honest comparison must optimize. Per-task accuracy, expected cost, tail latency, completion rate, and the cost of a wrong answer can pull in different directions. In a proof assistant, an exact checker may justify broad sampling because false candidates are cheap to reject. In an interactive assistant, a minute of hidden deliberation can make a modest quality gain irrelevant. In a high-stakes workflow, the best use of additional compute may be independent verification or escalation to a human rather than another fluent proposal. No scalar budget captures all three situations.

Routing also creates feedback effects. If the router estimates difficulty from the first attempt, confidently wrong answers may be underfunded while uncertain but recoverable answers consume the budget. If success is measured only among completed requests, aggressive early stopping can make the reported frontier look better by silently discarding the hard tail. And if parallel branches share the same prompt, model, and retrieved evidence, nominally larger search may add less diversity than its sample count suggests. The router, stopping rule, and diversity mechanism are therefore part of the evaluated system, not administrative details outside it.

A practical allocation study should report both a per-request frontier and a fleet frontier. The first asks how quality changes for one difficulty band as its budget rises. The second fixes aggregate capacity and asks how a routing policy distributes that capacity across a realistic workload. Compute-optimal results in controlled math settings motivate this distinction, while hardware-aware analyses show why the fleet conclusion can change when memory, attention, and concurrency enter the cost model. (Compute-Optimal Test-Time Compute; Kinetics)

Evaluation rules that prevent self-deception

A credible test-time scaling result needs controls beyond a rising accuracy curve.

Compare at matched end-to-end cost. Count generation, verification, tool calls, retries, and failed or timed-out attempts. Report both serial latency and total work when parallelism is used.

Separate coverage from delivered accuracy. pass@k answers whether a good candidate exists. A user-facing system needs selection accuracy. Reporting only oracle coverage makes the selector disappear.

Measure completeness. If high-budget runs time out preferentially on hard prompts, accuracy among returned answers is biased upward. A timeout is an outcome, not missing data.

Use multiple difficulty bands. Aggregate curves can hide that easy problems saturate, middle problems improve, and impossible problems waste compute. Adaptive allocation claims require evaluation on the router as well as the solver.

Test candidate correlation. Temperature changes, prompt paraphrases, different models, and distinct tools may generate more useful diversity than repeated sampling from one narrow distribution. Raw k is not effective independent k.

Audit the verifier. Test for false positives, false negatives, reward hacking, flaky execution, and common-mode errors shared with the proposer. A learned judge should be evaluated on held-out candidate distributions, including adversarially polished wrong answers.

Hold the harness fixed where possible. Changing prompts, tools, context, and stopping rules while increasing budget makes it unclear which intervention produced the gain. When co-design is the object of study, ablate the components explicitly.

These rules follow directly from the coverage/precision gap, compute-matching disputes, and verifier failures in the literature. (Large Language Monkeys; Solve or Verify; Kinetics)

Limits and contested interpretations

More compute is not more competence from nothing

Test-time scaling amplifies a trained distribution; it does not reliably manufacture missing knowledge or algorithms. Snell et al.'s smaller-model advantage was conditional on the smaller model already having a non-trivial success rate. Brown et al. found a model family with zero coverage on a coding task even after 10,000 samples. The practical inference is that budget should be withheld when diagnostics suggest the needed capability is absent. (Compute-Optimal Test-Time Compute; Large Language Monkeys)

Longer traces are not necessarily better traces

Token length is observable, but useful search is latent. A long response can contain correction, exploration, and tool feedback; it can also contain repetition, rationalization, or an oscillation between answers. The disagreement over s1 and the limited extrapolation found by e3 make confidence high that length alone is an inadequate mechanism, while confidence remains moderate about which training recipes reliably convert length into general exploration. (s1; Not That Simple; e3)

Verification is both the enabler and a new attack surface

Exact external verification makes repeated sampling unusually powerful. Learned verification extends the pattern to less formal tasks but introduces another fallible model. Evidence is strong that verifier-guided selection can help; evidence is also strong that compute allocation and verifier quality determine whether it helps efficiently. Claims that “verification is easier than generation” should therefore be task-specific, not axiomatic. (Training Verifiers; Solve or Verify)

Benchmark gains do not settle the nature of reasoning

The empirical claim is about performance under budget. Whether the mechanism constitutes “genuine reasoning,” search over memorized patterns, or a mixture is a separate question. OpenAI and DeepSeek describe learned self-correction and alternative strategies; structured systems visibly search. Yet benchmark success, hidden traces, and reranked samples do not establish trace faithfulness or robust generalization. Confidence is high in the performance effect within studied regimes, moderate in transfer beyond verifiable reasoning tasks, and low in strong philosophical conclusions from the curves alone. (OpenAI o1; DeepSeek-R1)

Practical decision framework

For a new workload, start from the workload's own shape rather than from a default commitment:

Workload property First mechanism to test Why
Unique answer with cheap exact check parallel sampling + verifier turns coverage directly into success
Repairable multi-step answer, expensive checking serial revision + selective verification preserves context and limits judge cost
Multiple plausible strategies with informative partial states structured search supports branching and backtracking
Open-ended answer with weak ground truth modest diversity + calibrated judge/user rubric avoids pretending pass@k is accuracy
Highly variable difficulty at volume adaptive router + early stopping spends budget where marginal return is positive
Hard real-world task with executable feedback tool loop + tests + retry budget uses environment evidence rather than prose alone

Then run a budget sweep rather than choosing one large setting. Plot delivered task success against at least latency and total cost. Inspect where the curve bends, which prompts consume the tail, and whether errors come from proposal, selection, or execution. If scaling breadth raises coverage without final accuracy, improve selection. If every sample repeats the same error, diversify or change the base model. If longer traces destabilize answers, add stopping or branch resets. If the highest budgets only help a tiny middle band, route to them selectively.

The governing principle is marginal, evidence-based escalation. Test-time compute should be treated as a portfolio allocated among proposing, checking, searching, and acting—not as a single “think harder” slider.

Evidence assessment

Claim Confidence Reason
Extra inference compute can improve performance on many math, code, proof, and puzzle tasks High replicated across sampling, verifier, search, and product-system evidence
Repeated sampling produces predictable coverage curves in verifiable domains High measured across models, tasks, and four orders of sample budget
Adaptive allocation is more efficient than uniform best-of-N in studied math settings High within studied settings; moderate generally strong controlled results, limited domain breadth
Smaller models plus search are generally cheaper than larger models Contested depends on task and whether FLOPs, memory, latency, or dollars define cost
Simply extending one reasoning trace improves beyond its trained budget Low to moderate positive demonstrations coexist with failures to extrapolate and oscillation
A universal inference-time analogue of pretraining scaling laws exists Low mechanisms, x-axes, tasks, and selectors remain heterogeneous

Concepts: Test-Time Compute, Test-Time Training, Inference Scaling Laws, Scaling Laws, Inference-Time Search, Chain-of-Thought Prompting

Methods: Self-Consistency, Process Reward Models, Tree of Thoughts, Beam Search, Pass@k

Evaluation: Production LLM Evals, Benchmark Saturation, Construct Validity, Cost-Aware Inference

References

1 Large Language Monkeys — Bradley Brown et al., “Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.” (Source)

2 Compute-Optimal Test-Time Compute — Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar, “Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.” (Source)

3 Empirical Inference Scaling — Yangzhen Wu et al., “Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.” (Source)

4 OpenAI o1 — OpenAI, “Learning to reason with LLMs.” (Source)

5 Self-Consistency — Xuezhi Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models.” (Source)

6 Training Verifiers — Karl Cobbe et al., “Training Verifiers to Solve Math Word Problems.” (Source)

7 Tree of Thoughts — Shunyu Yao et al., “Tree of Thoughts: Deliberate Problem Solving with Large Language Models.” (Source)

8 s1 — Niklas Muennighoff et al., “s1: Simple test-time scaling.” (Source)

9 Not That Simple — Guojun Wu, “It's Not That Simple. An Analysis of Simple Test-Time Scaling.” (Source)

10 DeepSeek-R1 — DeepSeek-AI et al., “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” (Source)

11 Solve or Verify — Nishad Singhi et al., “When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning.” (Source)

12 ARC o3 Evaluation — François Chollet, “OpenAI o3 Breakthrough High Score on ARC-AGI-Pub.” (Source)

13 Kinetics — Ranajoy Sadhukhan et al., “Kinetics: Rethinking Test-Time Scaling Laws.” (Source)

14 e3 — Amrith Setlur et al., “e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.” (Source)

15 Verification or RL — Amrith Setlur, Nived Rajaraman, Sergey Levine, and Aviral Kumar, “Scaling Test-Time Compute Without Verification or RL is Suboptimal.” (Source)

16 Tail-Guided Search — Muheng Li, Jian Qian, and Wenlong Mou, “Predicting and improving test-time scaling laws via reward tail-guided search.” (Source)

17 ARC Prize 2024 — ARC Prize Foundation, “ARC Prize 2024: Technical Report.” (Source)