Meta’s Muse Spark 1.1 has reached 62.1% on recent HLE leaderboards, placing it within a few points of Anthropic’s leading Claude Fable 5.1 and Opus 5 variants at 64.7–65%. This rapid climb from sub-50% earlier in 2026 reflects Meta’s focused scaling of reasoning modes, tool integration, and post-training on expert-level tasks across the 2,500-question benchmark. Trader sentiment for the 55%+ and 60%+ thresholds appears elevated because current results already clear those marks under closed-book or moderate-reasoning protocols, while 65%+ and 70%+ contracts trade lower amid uncertainty over further gains before year-end. Key swing factors include Meta’s next Muse Spark iteration, potential Llama open-weight releases, and evaluation variations such as high-effort reasoning or tool use that have historically added several points. The December 31, 2026 resolution leaves limited runway for additional model drops or benchmark updates.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$60,272 Vol.
55%+
33%
60%+
26%
65%+
16%
70%+
6%
$60,272 Vol.
55%+
33%
60%+
26%
65%+
16%
70%+
6%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:48 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Meta’s Muse Spark 1.1 has reached 62.1% on recent HLE leaderboards, placing it within a few points of Anthropic’s leading Claude Fable 5.1 and Opus 5 variants at 64.7–65%. This rapid climb from sub-50% earlier in 2026 reflects Meta’s focused scaling of reasoning modes, tool integration, and post-training on expert-level tasks across the 2,500-question benchmark. Trader sentiment for the 55%+ and 60%+ thresholds appears elevated because current results already clear those marks under closed-book or moderate-reasoning protocols, while 65%+ and 70%+ contracts trade lower amid uncertainty over further gains before year-end. Key swing factors include Meta’s next Muse Spark iteration, potential Llama open-weight releases, and evaluation variations such as high-effort reasoning or tool use that have historically added several points. The December 31, 2026 resolution leaves limited runway for additional model drops or benchmark updates.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions