Next Claude Opus: Humanity’s Last Exam Debut?
Ended Dec 31, 2026, 23:59 UTC
Market resolution
This Next Claude Opus: Humanity’s Last Exam Debut prediction market is settled. The percentages above are the final outcome probabilities reported by Polymarket.
Final probabilities, volume, and open interest are sourced from Polymarket and were last synced at Aug 9, 2026 11:32 pm.
Anthropic IPO date Claude’s 55% Cliff Encodes a Hidden Benchmark-Configuration Wager
Anthropic’s own evaluations place Opus on opposite sides of the decisive threshold depending on tool access. The market’s sharp break therefore depends on which testing configuration reaches the official leaderboard, plus whether the next model delivers even a modest gain over today’s no-tools result.

The market’s central thesis is that the next Claude Opus will score around the low 50s on Humanity’s Last Exam under the configuration recognized by the official leaderboard. That inference explains why thresholds through 50% trade near certainty while 55% receives only a 0.8% Yes price. The dividing line closely tracks Anthropic’s published gap between tool-free and tool-assisted evaluation.
Anthropic’s own scores create the 55% fault line
Anthropic’s Claude Opus 4.8 system card reports 49.8% HLE accuracy without tools and 57.9% with tools. Those results bracket the market’s sharpest division: one sits just below 50%, while the other clears 55% by 2.9 percentage points.
This matters because the next Opus needs only a small improvement over the reported 49.8% no-tools result to reach 50%. Reaching 55% under a comparable configuration requires a larger gain of 5.2 points. The prices therefore imply confidence in incremental progress and deep skepticism that the settlement score will reproduce Anthropic’s tool-assisted result.
That interpretation is a market inference. The system-card figures are sourced facts, while settlement depends exclusively on the HLE accuracy displayed at agi.safe.ai for the next Claude Opus entry. Anthropic’s evaluation does not itself determine the outcome.
The market is implicitly choosing a testing regime
The largest hidden assumption concerns comparability. A benchmark score can depend on tool access, prompting, sampling, answer extraction and the exact model build submitted. The supplied record establishes two Anthropic scores with different tool conditions, yet it does not specify which condition the official HLE entry will use.
A leaderboard entry aligned with the 49.8% no-tools evaluation would validate the market’s low-50s anchor. A tools-enabled entry comparable to 57.9% would directly challenge the 55% pricing. A separately run evaluation could produce another result altogether, especially if the next Opus differs materially from the version covered by the system card.
The hierarchy also assumes that model progress will transfer to HLE. A new Opus release may improve coding, agentic tasks or product reliability without adding five points on this benchmark. Conversely, benchmark-specific reasoning gains could move HLE faster than broader capability measures suggest. Product branding alone provides limited evidence about either path.
The lower thresholds reveal confidence, plus pricing noise
The 35%, 40%, 45% and 50% contracts carry Yes prices of 99.3%, 95.9%, 96.8% and 98.1%, respectively. Their order is logically inconsistent: any score clearing 50% must also clear every lower threshold. The reversals between 40%, 45% and 50% mean the individual quotes cannot be treated as a single clean probability distribution.
This is the main counter-signal to reading too much precision into the hierarchy. With $23.79K in volume, $18.42K in liquidity and $8.15K in open interest across the event, separate contract conditions or stale quotes can influence the visible ordering. The broad message remains a cliff above 50%; the exact differences among lower bands carry weaker evidentiary value.
Official publication will overwhelm indirect capability evidence
The decisive catalyst is the first appearance of the next Claude Opus on the official HLE results site. Under the resolution criteria, its displayed accuracy at 12:00 PM ET on the following calendar day controls settlement. That next-day checkpoint makes corrections, relabeling or score updates during the initial publication window materially relevant.
Before publication, an Anthropic launch announcement or system card could shift expectations if it reports HLE performance and states the evaluation setup. A result above 55% under conditions matching the official leaderboard would weaken the market’s configuration thesis. A tool-free score near 50%, or explicit evidence that the leaderboard excludes tools, would strengthen it. Scores near a threshold would also make rounding and the site’s displayed precision important.
Timing remains an operational failure mode
The market closes on December 31, 2026 at 11:59 PM UTC, while the supplied criteria define settlement through the next model’s leaderboard appearance. The record provided here does not state how a failure to publish a new Opus entry before closing would be handled. That leaves launch timing and rules interpretation as secondary catalysts.
The strongest substantive failure mode is simpler: the market may be anchoring too heavily on the 49.8% no-tools figure. If the official entry follows Anthropic’s 57.9% tools-enabled setup, or the next generation adds more than five points under a comparable no-tools test, the apparent 55% barrier would lose its empirical basis.
Sources
Market details
- Resolution criteria
- This market will resolve to "Yes" if the next Claude Opus model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No".
- Category
- Tech › AI
- Close date
- December 31, 2026, 11:59 PM UTC
- Settlement source
- agi.safe.ai
- Market rules summary
- Multi-timeframe Polymarket event. Each listed timeframe is represented by its Yes price on the underlying binary market. View full rules
Frequently asked questions
What was the final result of the Next Claude Opus: Humanity’s Last Exam Debut prediction market?
Polymarket reports the Next Claude Opus: Humanity’s Last Exam Debut prediction market as closed. The final snapshot shows 35%+ at 100%, 40%+ at 100%, 45%+ at 100%, and 50%+ at 100%. The final market snapshot includes $60.78K volume and $32.48K open interest. CryptoSlate last synced the final market data at Aug 9, 2026, 22:32 UTC.
How does the Next Claude Opus: Humanity’s Last Exam Debut prediction market resolve?
This market will resolve to "Yes" if the next Claude Opus model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". Multi-timeframe Polymarket event. Each listed timeframe is represented by its Yes price on the underlying binary market. The settlement source listed for this market is agi.safe.ai.