Anthropic’s Claude family currently dominates Humanity’s Last Exam (HLE) leaderboards, with Fable 5.1 at 65%, Opus 5 at 64.7%, and Mythos 5 at 64.5% on the September 18, 2026 BenchLM snapshot of the 2,500-question expert benchmark. This positioning stems from recent releases emphasizing advanced reasoning chains, tool integration, and adaptive inference that outperform GPT-5.4 Pro, Muse Spark, and Gemini variants on the closed-book or lightly assisted protocols. Rapid score gains—from single digits in early 2025 to the mid-60s—reflect iterative scaling of training compute and post-training techniques, though protocol differences (tools versus no-tools) and potential new model drops before December 31 create remaining uncertainty around whether 70%+ thresholds are reached. Traders are watching Anthropic’s release cadence and any benchmark updates from CAIS/Scale for signals on further progress.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于$109,253 交易量
55%及以上
90%
60%以上
42%
65%及以上
19%
70%及以上
12%
75%+
8%
$109,253 交易量
55%及以上
90%
60%以上
42%
65%及以上
19%
70%及以上
12%
75%+
8%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
市场开放时间: Jul 23, 2026, 6:42 PM ET
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Anthropic’s Claude family currently dominates Humanity’s Last Exam (HLE) leaderboards, with Fable 5.1 at 65%, Opus 5 at 64.7%, and Mythos 5 at 64.5% on the September 18, 2026 BenchLM snapshot of the 2,500-question expert benchmark. This positioning stems from recent releases emphasizing advanced reasoning chains, tool integration, and adaptive inference that outperform GPT-5.4 Pro, Muse Spark, and Gemini variants on the closed-book or lightly assisted protocols. Rapid score gains—from single digits in early 2025 to the mid-60s—reflect iterative scaling of training compute and post-training techniques, though protocol differences (tools versus no-tools) and potential new model drops before December 31 create remaining uncertainty around whether 70%+ thresholds are reached. Traders are watching Anthropic’s release cadence and any benchmark updates from CAIS/Scale for signals on further progress.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于



警惕外部链接哦。
警惕外部链接哦。
常见问题