Reference

Humanitys-Last-Exam

Humanity’s Last Exam (HLE) is a broad, expert-authored benchmark designed to test whether frontier models can answer closed-ended academic questions that remain hard even after MMLU-style benchmarks have saturated. Its deeper significance is not that it is literally “humanity’s last exam,” but that it exposes the shrinking half-life of static benchmarks: HLE began as a low-single-digit to low-double-digit challenge for frontier systems, yet public leaderboards already report frontier scores in the 40% range. The live question is whether HLE becomes another Benchmark Saturation story, or whether its mixture of expert construction, short-answer verification, multimodality, private splits, and rolling updates can keep it useful longer than earlier academic exams.

Coverage note: verified through May 11, 2026.

Why HLE exists

HLE was created in response to a measurement problem: once a benchmark becomes easy for frontier systems, it stops separating models at the frontier. MMLU, once a difficult broad academic benchmark, now sits above 90% for state-of-the-art systems in the HLE authors’ framing, which makes it weak as a frontier discriminator even if it remains useful for smaller models and longitudinal comparisons. HLE’s claim is that frontier evaluation needs a new class of difficult, expert-level, closed-ended questions with known answers, broad domain coverage, and enough resistance to internet lookup and dataset memorization to remain meaningful for at least a period of frontier progress. ([HLE Nature paper]) Nature

The benchmark’s name is intentionally provocative. The HLE authors describe it as possibly the “last exam” of its kind: not the last AI benchmark, not an AGI test, and not a substitute for open-ended research evaluation, but perhaps the last broad, static, closed-ended academic exam that can remain hard for frontier models for more than a short interval. They explicitly acknowledge that HLE does not measure autonomous research, agentic competence, or all forms of general intelligence. ([HLE Nature limitations]) Nature

That distinction matters. HLE is best understood as a stress test for Crystallized Knowledge in AI Systems and expert-level academic problem solving, not as a direct measure of Artificial General Intelligence. It samples hard questions from human academic expertise; it does not ask whether a model can design experiments, run a long-term project, interact with institutions, or produce novel validated science.

Core design

HLE is a closed-ended benchmark of approximately 2,500 expert-level questions spanning more than 100 subjects. The finalized Nature version describes a public benchmark built from nearly 1,000 contributor-authors affiliated with more than 500 institutions across roughly 50 countries. Questions are mostly exact-answer, with a smaller multiple-choice component and a multimodal subset containing diagrams, figures, or other visual material. ([HLE Nature methodology]) Nature

The benchmark’s design criteria are unusually strict for a public academic evaluation. Submitted questions were supposed to be precise, unambiguous, solvable, non-subjective, non-trivial, and not answerable by quick internet search. Contributors were asked to provide short exact answers or constrained answer formats, rationales, subject labels, and affiliation metadata. The benchmark excluded open-ended questions, subjective judgments, and some dangerous domains such as weapons of mass destruction. ([HLE Nature methodology]) Nature

The resulting benchmark is not simply “hard trivia.” HLE tries to combine three properties that are individually common but rarely jointly optimized:

Design property What HLE tries to achieve Why it matters Residual weakness
Expert-level difficulty Questions require graduate-level, professional, or highly specialized knowledge Separates frontier systems after MMLU-style saturation Difficulty is uneven across domains and contributors
Closed-ended verifiability Answers are short, checkable, and intended to be unambiguous Enables scalable automatic grading Some items still have ambiguous, disputed, or erroneous answers
Breadth Questions cover many disciplines, with text and multimodal items Prevents a model from specializing in one narrow domain Breadth does not guarantee representativeness of real-world competence
Search resistance Questions should not be trivially retrievable online Reduces benchmark-as-search-engine effects Public release and web indexing create contamination risk
Frontier filtering Items were filtered partly by whether current models failed Keeps early scores low Introduces selection effects tied to a specific model generation

The most important architectural choice is the closed-ended format. HLE does not ask models to write essays, produce research plans, or argue interpretively. It asks for answers that can be graded. That makes the benchmark scalable, but also narrows what it can measure. HLE is therefore adjacent to AI Evaluation, Benchmark Design, and Expert-Level Question Answering, but not a complete evaluation of agency, creativity, or scientific competence.

Construction methodology

HLE’s construction pipeline combined open expert recruitment, prize incentives, adversarial item selection, model filtering, expert review, and post-release refinement. The Nature paper reports that contributors submitted roughly 70,000 attempts; around 13,000 questions passed the initial difficulty stage and were forwarded for expert review. Reviewers generally had graduate degrees and evaluated questions across one or more rounds before final selection. ([HLE Nature methodology]) Nature

The incentive design is central to HLE’s identity. Contributors could receive prize payments and coauthorship, with the paper describing a $500,000 prize pool: larger awards for the top contributors and smaller awards for the next tier. This matters because HLE was not produced only by a small benchmark team; it used a distributed expert-sourcing model closer to a contest plus academic collaboration. ([HLE Nature methodology]) Nature

Pipeline summary

Stage Mechanism Intended effect Main risk
Expert call Recruit domain experts across many fields Increase breadth and difficulty Contributor pool may overrepresent certain disciplines
Submission rules Require precise, closed-ended, non-searchable questions Improve grading and reduce trivial retrieval Rules do not eliminate ambiguity
Difficulty filtering Test whether frontier models can answer candidate questions Preserve frontier-level challenge Filters for “current model weaknesses,” not necessarily durable difficulty
Expert review Domain reviewers check validity, clarity, and answerability Reduce errors and noise Reviewers may not fully verify every long rationale
Public/private split Release public data while retaining private evaluation material Enable research while limiting overfitting Public split can contaminate future models
Post-release audits Bug bounty, searchability checks, community feedback Remove flawed or searchable items Revisions complicate score comparability
Rolling updates Proposed HLE-Rolling continuation Extend benchmark lifetime Requires ongoing governance and item supply

The construction process is adversarial in a modest but important sense. HLE questions were not selected merely because they were representative of a curriculum; they were selected partly because frontier systems could not answer them. That makes HLE different from exams meant to sample a syllabus. It is closer to a frontier stress test: a curated distribution of expert problems that survived contact with then-current models.

The authors also describe post-release refinement. They report community feedback, late contributions, and a process for auditing searchable questions using GPT-4o mini, GPT-4o search, and Perplexity Sonar. They argue that model performance remained similar after removing questions found to be searchable, but the need for this procedure shows that “non-searchable” is not a static property once questions are public. ([HLE Nature post-release notes]) Nature

Scoring and evaluation protocol

HLE asks models to produce both an answer and a confidence score. Evaluation reports accuracy and calibration, with attention to whether models are overconfident on questions they answer incorrectly. The Nature paper’s evaluation setup used standardized prompts and an automated judge for closed-ended answers, with calibration error reported alongside accuracy. ([HLE Nature evaluation]) Nature

This design makes HLE useful for studying AI Calibration, not merely raw correctness. A model that scores 30% while assigning 95% confidence to wrong answers is less epistemically useful than a model that scores similarly but knows when it does not know. The HLE paper emphasizes that contemporary systems often show poor calibration: they can be highly confident on wrong answers, especially in expert domains. ([HLE Nature evaluation]) Nature

The Scale/SEAL public leaderboard uses all public questions, temperature-zero evaluation, answer plus confidence extraction, confidence intervals, and an automated judge or extractor. It also warns that discrepancies are possible and asks model submitters to document prompts and model details. This is an important practical caveat: HLE scores are not purely properties of a model architecture; they are properties of a model, prompt, decoding setup, grading method, and benchmark version. ([Scale HLE leaderboard methodology]) Scale Labs

Empirical scores: from launch difficulty to rapid progress

The canonical launch-era results were very low. In the Nature table, GPT-4o scored 2.7%, Claude 3.5 Sonnet scored 4.1%, Gemini 1.5 Pro scored 4.6%, OpenAI o1 scored 8.0%, and DeepSeek-R1 on text-only questions scored 8.5%. This is the cleanest source for the original “frontier models failed HLE” claim. ([HLE Nature Table 1]) Nature

The often-repeated “about 3–15%” early range is directionally right but compresses multiple snapshots. The Nature launch-era non-tool table is mostly about 2.7–8.5%; broader early public-leaderboard discussion and later reasoning-model entries pushed the envelope upward. The safer statement is: launch-era frontier models were in the low single digits to high single digits on the canonical table, while early leaderboard-style frontier systems soon moved into low-double-digit territory. ([HLE Nature Table 1]; [Scale HLE leaderboard]) Nature

By the time of the official CAIS/Nature page’s later leaderboard snapshot, the picture had already shifted. That page lists Gemini 3 Pro at 38.3%, GPT-5 at 25.3%, Grok 4 at 24.5%, Gemini 2.5 Pro at 21.6%, GPT-5-mini at 19.4%, Claude 4.5 Sonnet at 13.7%, and several older models below 10%. The same page cautions that post-release models may have had access to the open-sourced HLE material, which makes public-split results less clean as measurements of uncontaminated generalization. ([CAIS HLE page]) AGI Safe

As of May 11, 2026, the Scale public leaderboard reported top entries around the mid-40% range, including Gemini 3.1 Pro Preview at 46.44% and GPT-5.4 Pro at 44.32% under that leaderboard’s setup. A separate Artificial Analysis-derived aggregator reported a similar but not identical frontier picture, with Gemini 3.1 Pro Preview at 44.7%, GPT-5.4 at 41.6%, and GPT-5.3 Codex at 39.9%. These numbers should not be mixed as if they were one unified benchmark table, but they agree on the broad trajectory: HLE moved from launch-era single digits to roughly 40%+ frontier performance in a short period. ([Scale HLE leaderboard]; [Artificial Analysis-derived HLE aggregator]) Scale Labs

Representative score trajectory

Period / source Representative models Reported HLE accuracy Interpretation
Launch-era Nature table GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro 2.7–4.6% Broad expert questions were mostly out of reach for general frontier chat models
Launch-era reasoning models o1, DeepSeek-R1 text-only 8.0–8.5% Reasoning-oriented systems helped but still failed most questions
Later official CAIS/Nature snapshot Gemini 2.5 Pro, GPT-5, Grok 4, Gemini 3 Pro 21.6–38.3% HLE began separating newer frontier systems, but no longer looked near-zero
May 2026 public leaderboards Gemini 3.1 Pro Preview, GPT-5.4 Pro, other frontier systems About 40–46% Static public HLE appears far from saturated, but its half-life is visibly shrinking

The Stanford AI Index 2026 describes HLE as one example of a broader benchmark-saturation pattern, reporting that frontier models gained about 30 percentage points on HLE in a single year. That observation is arguably more important than any single leaderboard number: HLE’s reception has become inseparable from the question of benchmark half-life. ([Stanford AI Index 2026]) Stanford HAI

Reception: why HLE became symbolically important

HLE attracted attention before its final benchmark form because it framed evaluation as a race against benchmark obsolescence. Reuters covered the global call for difficult questions in September 2024, describing a push to collect problems that would challenge systems already making popular benchmarks look easy. ([Reuters coverage]) Reuters

Media coverage then amplified the “last exam” framing. A New York Times article, syndicated elsewhere, presented HLE as a test that leading models initially failed, with OpenAI o1 near the top of the early pack around 8%. That framing was rhetorically powerful: “when AI passes this test, look out” became a shorthand for the idea that ordinary academic benchmarks were losing their discriminative value. ([New York Times syndicated coverage]) The Star

Practitioners adopted HLE for a more prosaic reason: it produced spread at the frontier. A good benchmark for frontier models need not be philosophically complete; it must be hard enough, reliable enough, and cheap enough to run that labs and evaluators can compare systems. HLE served that role better than MMLU once MMLU scores clustered near the top. ([HLE Nature paper]; [Scale HLE leaderboard]) Nature

The reception has therefore been mixed but productive. HLE is widely treated as one of the leading frontier academic benchmarks. At the same time, its public release, item-quality debates, and rapid score increases have made it a case study in Benchmark Contamination, Evaluation Gaming, and the limits of static leaderboards.

Methodological critique 1: contamination risk

The most serious practical critique is contamination. Once a benchmark is public, future model-training corpora, retrieval indexes, evaluation harnesses, and search tools may include the benchmark questions or answers. This can inflate scores without corresponding capability gains.

A recent paper on search-time data contamination found that, across HLE, SimpleQA, and GPQA, approximately 3% of questions could be directly found through Hugging Face sources during search-augmented evaluation, and that blocking Hugging Face caused accuracy on contaminated subsets to drop by about 15%. The paper’s central point is not that all HLE scores are invalid; it is that web-enabled evaluation creates an inference-time contamination channel distinct from ordinary pretraining contamination. ([Search-Time Data Contamination]) arXiv

This is especially relevant for HLE because the benchmark is partly intended to resist search. A question can be non-searchable when authored, then become searchable because it appears in a public dataset, a leaderboard, a GitHub repository, a Hugging Face dataset, a blog post, or a model card. The benchmark’s own publication changes the information environment it tries to measure.

The broader contamination literature agrees that benchmark data contamination is difficult to eliminate completely. Surveys distinguish detection methods, mitigation methods, and prevention methods, but they do not claim that post hoc checks can fully restore a public benchmark’s original measurement value. ([Benchmark data contamination survey]) arXiv

Detection work such as “Time Travel in LLMs” tries to identify contamination by comparing model behavior across temporally separated benchmark partitions and by using guided prompting to test whether models have memorized benchmark artifacts. The existence of such methods is useful, but it also underscores the problem: contamination detection is itself a moving target as models, training corpora, and evaluation setups change. ([Time Travel contamination detection]) arXiv

A more structural response is benchmark watermarking: embedding cryptographic or statistical markers into benchmark items so later leakage can be detected more reliably. Watermarking does not solve all evaluation problems, but it moves contamination defense from retrospective suspicion toward proactive auditability. ([Benchmark watermarking]) arXiv

For practitioners, the lesson is simple: HLE public-leaderboard scores are useful but not sacred. The best use of HLE is with private or rolling splits, source-filtered search settings, prompt disclosure, contamination audits, and clear separation between no-tool, tool-assisted, and web-enabled runs.

Methodological critique 2: item quality variance

HLE’s expert-sourced design creates another predictable problem: expert questions are hard to verify at scale. The same features that make questions difficult for models—specialized knowledge, unusual derivations, niche domains, long rationales—also make them difficult for reviewers to check exhaustively.

The HLE authors report expert disagreement audits. In one final public-set audit, they estimate around 15.4% expert disagreement, with higher disagreement in some biomedical and chemistry subsets. They also note that some questions derive from research experience and may be hard to verify through ordinary literature search. ([HLE Nature post-release notes]) Nature

A later HLE-Verified paper argues that HLE contains a non-trivial number of noisy items, including ambiguous statements, incorrect answers, and mismatched rationales. It proposes a two-stage validation-and-repair pipeline, classifying hundreds of items as verified, revised, or uncertain, and reports that repaired/verified subsets change measured model accuracy by several percentage points overall and much more on flawed items. ([HLE-Verified]) arXiv

This critique should not be dismissed, but it should also not be overstated. A 15–20% expert-disagreement estimate does not imply that every disputed question is “wrong”; in some fields, legitimate ambiguity can arise from conventions, implicit assumptions, or underspecified problem statements. Conversely, calling disagreement “just expert disagreement” can hide real errors. The honest conclusion is that HLE item quality is heterogeneous, and that leaderboard use should incorporate item validation status where possible.

The item-quality debate is a useful reminder that Expert Evaluation is not automatically high-quality evaluation. Experts can write brittle questions, reviewers can miss subtle mistakes, and benchmark designers can overfit to the goal of stumping models. HLE’s construction is stronger than casual crowdsourcing, but it is not immune to benchmark noise.

Methodological critique 3: the “unsaturable by design” framing

HLE is often described as designed to be unsaturable by current frontier models. That phrase is accurate only if read carefully. HLE was designed so that then-current frontier models performed poorly; it was not designed in a way that guarantees durable unsaturability.

There are three reasons.

First, adversarial filtering is temporally local. If questions are selected because GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, o1, or contemporaneous systems fail them, the benchmark is implicitly tied to the weaknesses of that model generation. The next generation may share some weaknesses, but it need not.

Second, public release creates training and retrieval pressure. Once benchmark items are open, they can enter pretraining data, fine-tuning data, evaluation discussions, and retrieval corpora. Private splits mitigate this, but public scores become progressively harder to interpret.

Third, the frontier is moving toward test-time reasoning, tool use, search, self-verification, and domain-specialized scaffolds. HLE’s exact-answer format makes it easy to grade, but it also allows targeted optimization: better symbolic tools, solvers, retrieval filters, calculators, theorem provers, and scientific databases can all raise scores without requiring a human-like understanding of every field.

The better claim is therefore weaker but still important: HLE was unsaturated at launch and remains difficult enough to separate many frontier systems as of May 2026. It is not a proof that static academic benchmarks can indefinitely outrun frontier model progress.

Methodological critique 4: what HLE does not measure

HLE measures closed-ended expert answer production. That is valuable, but it excludes many capabilities central to deployed AI systems.

It does not directly measure long-horizon agency, tool orchestration, project management, experimental design, interpersonal communication, economic usefulness, cyber autonomy, embodied robotics, or institutionally situated judgment. The HLE authors make a related point themselves: HLE is not an AGI test and does not evaluate open-ended autonomous research. ([HLE Nature limitations]) Nature

This limitation is not a defect if HLE is used appropriately. Problems arise only when HLE is treated as a single-number proxy for frontier intelligence. A model that scores highly on HLE may still fail at open-ended research. A model that scores poorly on HLE may still be useful in software engineering, retrieval, customer support, or agentic workflows. HLE is one important slice of the evaluation stack, not the stack.

Relationship to MMLU

MMLU is the obvious historical comparison. It was designed as a broad multitask test of language-model knowledge and reasoning across 57 subjects, including mathematics, history, computer science, law, and other academic areas. When GPT-3-era systems were evaluated, the best models were far below expert-level performance. ([MMLU paper]) arXiv

HLE inherits MMLU’s broad-academic ambition but changes the difficulty regime. MMLU asks many questions that educated humans or strong models can now answer; HLE asks expert-level questions selected partly because frontier models could not. If MMLU measures broad academic competence, HLE measures the tail of expert academic difficulty after mainstream academic benchmarks saturate. ([HLE Nature paper]) Nature

The relevant practitioner question is whether HLE follows MMLU’s arc. MMLU went from hard to saturated. HLE began harder, but the observed trajectory—single digits to 40%+ frontier scores—suggests it is not immune to the same dynamic. Its advantage is not invulnerability; its advantage is a higher ceiling of difficulty, shorter answers, better frontier filtering, and some private/rolling infrastructure.

Relationship to GPQA

GPQA is closer to HLE than MMLU is. It consists of graduate-level, “Google-proof” multiple-choice questions written by domain experts in biology, physics, and chemistry. In the original GPQA framing, PhD experts performed much better than non-experts, while strong models remained well below expert performance. ([GPQA paper]) arXiv

HLE generalizes the GPQA idea in breadth and format. GPQA is narrower, science-focused, and multiple-choice. HLE spans many more domains, contains more exact-answer questions, includes multimodal items, and was explicitly constructed as a frontier benchmark after GPQA itself began losing discriminative power for the strongest systems. ([HLE Nature methodology]) Nature

The contamination issue also links them. Search-time contamination research evaluates HLE, GPQA, and SimpleQA together because all three can be affected by retrieval systems finding benchmark artifacts rather than independently solving the problem. ([Search-Time Data Contamination]) arXiv

Relationship to ARC-AGI

ARC-AGI is philosophically different. Where HLE tests crystallized expert knowledge and closed-ended academic reasoning, ARC-AGI tests fluid abstraction: inferring transformation rules from small examples, usually without relying on language, specialized knowledge, or internet facts. The ARC-AGI materials explicitly frame the benchmark as measuring learning efficiency and generalization rather than economically useful task performance. ([ARC-AGI guide]) ARC Prize

ARC-AGI’s strength is that it avoids some of HLE’s reliance on human academic knowledge. Its weakness is that it is narrow in another way: grid puzzles are not the full space of intelligence. HLE and ARC-AGI are therefore complementary. HLE asks, “Can the model answer very hard expert questions?” ARC-AGI asks, “Can the model infer hidden structure from sparse examples in a domain designed to minimize memorized knowledge?”

ARC Prize’s 2025 report describes ARC-AGI-2 and later ARC-AGI-3 directions as moving toward harder private sets and interactive reasoning involving exploration, planning, memory, goal acquisition, and adaptation. That trajectory points toward a post-HLE evaluation world where static question-answering becomes only one component of frontier assessment. ([ARC Prize 2025 report]) arXiv

Relationship to FrontierMath, BrowseComp, and factuality benchmarks

FrontierMath is another frontier benchmark, but it is much narrower than HLE. It focuses on original, exceptionally challenging mathematics problems that can take expert researchers hours or days, with automated verification where possible. Early reported model performance was below 2%, making it a deep-domain counterpart to HLE’s broad-domain strategy. ([FrontierMath paper]) arXiv

BrowseComp and SimpleQA target different failure modes. BrowseComp evaluates whether models or agents can locate hard-to-find information through persistent web navigation, while SimpleQA tests short factual question answering with single, indisputable answers. These are not substitutes for HLE; they measure retrieval, factuality, and search-agent behavior rather than expert academic problem solving. ([BrowseComp paper]; [SimpleQA paper]) arXiv

The important connection is that modern frontier systems increasingly blur the line between “knowing,” “reasoning,” and “retrieving.” A no-tool HLE score measures one thing. A web-enabled HLE score measures another. A scaffolded agent with retrieval, calculators, code execution, and verification loops measures something else again. Practitioner reports should keep those regimes separate.

Comparison table: HLE and neighboring benchmarks

Benchmark Primary target Format Strength Weakness
MMLU Broad academic knowledge and reasoning Multiple-choice across 57 subjects Historically important, broad, easy to compare Saturated for frontier systems
GPQA Graduate-level science reasoning Expert-written multiple-choice science questions Stronger expert difficulty than MMLU Narrower domain, multiple-choice, contamination risk
HLE Broad expert-level academic problem solving Mostly exact-answer, some multiple-choice, some multimodal Hard, broad, frontier-discriminating, calibration-aware Item noise, contamination, public-split aging, closed-ended scope
ARC-AGI Fluid abstraction and sample-efficient generalization Grid-transformation tasks Less dependent on memorized world knowledge Narrow artificial task family
FrontierMath Deep advanced mathematical reasoning Original hard math with verifiable answers Very high difficulty, strong verification potential Narrow domain
BrowseComp Persistent web search and hard-to-find facts Web-navigation questions with short answers Tests retrieval-agent behavior Vulnerable to search-index leakage and eval-awareness
SimpleQA Factual short-answer reliability Short fact-seeking questions Useful for hallucination/factuality measurement Not a broad expert reasoning benchmark

Will HLE follow MMLU’s saturation arc?

The evidence points toward “yes, but more slowly and less cleanly.” HLE has already shown rapid progress. The move from launch-era 2.7–8.5% scores to public leaderboard scores around 40–46% is not saturation, but it is far too fast to support a strong “unsaturable” interpretation. ([HLE Nature Table 1]; [Scale HLE leaderboard]; [Artificial Analysis-derived HLE aggregator]) Nature+2Scale Labs+2

Several forces push HLE toward MMLU-like saturation:

Force Why it raises HLE scores
Stronger reasoning models More reliable multi-step derivations and self-checking
Test-time compute More attempts, verification, and search over solution paths
Tool use Calculators, solvers, code, retrieval, theorem provers, and databases help on exact-answer tasks
Public exposure Benchmark artifacts can enter training or retrieval corpora
Prompt optimization Better answer formatting and confidence calibration improve measured performance
Domain specialization Models or scaffolds can target weak HLE categories

Several forces slow saturation:

Force Why HLE remains difficult
Expert tail difficulty Many questions require genuinely specialized knowledge
Exact-answer grading Vague partial credit is limited
Multimodality Some items require interpreting figures or diagrams
Breadth A model must cover many domains, not just one
Private and rolling splits Contamination and overfitting can be reduced
Item repair Removing noisy or searchable items can restore signal

For practitioners, the right question is not “Will HLE saturate?” but “Which version of HLE, under which protocol, remains discriminative for which model class?” Public no-tool HLE, private no-tool HLE, public tool-assisted HLE, and rolling private HLE are different evaluations.

A practical HLE report should include at least six fields: benchmark version, public/private split, model checkpoint date, prompt, tool/search policy, and judge/extractor. Without those details, a single HLE percentage is easy to overinterpret.

What HLE is good for in practice

HLE is most useful as a frontier-screening and calibration benchmark. It can answer questions such as:

Practitioner question HLE’s usefulness
Does a new frontier model substantially improve expert-level closed-answer reasoning? High
Is the model calibrated on questions beyond its competence? High
Which academic domains remain weak? Moderate to high, depending on item quality
Does tool use improve hard academic problem solving? High, if no-tool and tool-assisted runs are separated
Has a model reached AGI? Low
Can a model do autonomous research? Low to moderate at best
Is a model safe or aligned? Low

The calibration dimension is underrated. A model that fails HLE honestly is more useful than one that fails confidently. As AI systems are deployed in technical workflows, knowing when a model should defer may matter as much as knowing how many hard questions it can answer.

What comes after HLE if it saturates?

The HLE authors already gesture toward a successor path: HLE-Rolling, a dynamic fork intended to incorporate new questions and community feedback, with migration once models approach the noise ceiling of the static benchmark. ([HLE Nature HLE-Rolling note]) Nature

But rolling updates are only the first step. If HLE saturates, the frontier evaluation stack will likely move toward a mixture of private, dynamic, interactive, and externally validated tasks.

1. Rolling private exams

A direct successor is a continuously refreshed HLE-like benchmark with sealed test sets, rotating items, contamination audits, and delayed public release. This preserves the benefits of closed-ended expert questions while reducing public-split memorization.

2. Verified item banks

HLE-Verified-style repair pipelines could become standard. Instead of treating expert authorship as sufficient, future benchmarks may attach validation status, disagreement metadata, known ambiguity flags, and multiple independent solution checks to each item. ([HLE-Verified]) arXiv

3. Watermarked benchmarks

Future static benchmarks may include watermarking or other provenance mechanisms so benchmark leakage can be detected more rigorously. This is especially important for public datasets likely to be scraped into training corpora. ([Benchmark watermarking]) arXiv

4. Procedural and generator-based evaluations

Instead of publishing fixed questions, benchmark designers can publish task generators, hidden distributions, and verifiers. This shifts the target from memorizing items to solving a class of problems. FrontierMath’s emphasis on original verifiable problems and ARC-AGI’s private task distributions both point in this direction. ([FrontierMath paper]; [ARC Prize 2025 report]) arXiv

5. Interactive agent benchmarks

Static question answering cannot measure exploration, planning, memory, tool orchestration, and adaptive strategy. ARC-AGI-3’s emphasis on interactive reasoning is one example of the direction: evaluate agents in environments where they must infer goals, explore, and adapt, not merely answer a fixed prompt. ([ARC Prize 2025 report]) arXiv

6. Lab-in-the-loop scientific evaluation

The most meaningful successor to HLE may not be another exam. For scientific AI, the hard benchmark is whether systems can generate hypotheses, design experiments, use tools, collaborate with humans, and produce results that survive external validation. That is expensive and slow, but it measures capabilities that closed-ended benchmarks can only approximate.

7. Evaluation of epistemic behavior

As models improve, “knows the answer” becomes less sufficient. Future benchmarks should measure uncertainty, abstention, source use, provenance, disagreement handling, and the ability to identify when a question is ill-posed. HLE’s confidence field is an early step in this direction, but future evaluations will need richer epistemic metrics.

Bottom line

HLE is an important benchmark because it captures a real transition in AI evaluation. MMLU-style academic exams no longer reliably separate frontier systems; HLE restores separation by using expert-authored, difficult, closed-ended, partly multimodal questions selected to defeat current models. That makes it valuable.

But HLE is also important because it may fail in the same way previous benchmarks failed: not because it was badly designed, but because frontier progress, public exposure, contamination, tooling, and benchmark-specific optimization erode static evaluations. HLE’s future value depends less on the original public test set than on the ecosystem around it: private splits, rolling updates, item verification, contamination detection, calibration metrics, and complementary agentic evaluations.

The most intellectually honest position is therefore neither “HLE is humanity’s last exam” nor “HLE is already obsolete.” HLE is a strong frontier academic benchmark with visible methodological limits and a rapidly changing score trajectory. Its long-term legacy may be less the specific questions it contains than the evaluation regime it helped clarify: benchmarks must now be treated as living measurement systems, not fixed monuments.

Companion entries

Core theory: Benchmark Saturation, AI Evaluation, Crystallized Knowledge in AI Systems, Fluid Intelligence in AI, Closed-Ended Evaluation, Expert-Level Question Answering, AI Calibration

Benchmarks: MMLU, GPQA, Humanity’s Last Exam, ARC-AGI, FrontierMath, BrowseComp, SimpleQA

Methodology: Benchmark Contamination, Search-Time Data Contamination, Benchmark Watermarking, Private Test Sets, Rolling Benchmarks, Adversarial Filtering, Expert Review Pipelines

Practice: Frontier Model Leaderboards, Tool-Assisted Evaluation, No-Tool vs Tool-Use Protocols, Evaluation Harness Design, Confidence Calibration Metrics, Model Grading with LLM Judges

Counterarguments and limits: Goodhart’s Law in AI Evaluation, Evaluation Gaming, Static Benchmark Half-Life, Item Quality Variance, Ambiguity in Expert Benchmarks, AGI Benchmark Critiques