Anthropic’s September 2026 launches of Claude Fable 5.1 and Mythos 5.1 have driven the current trader consensus, with these models posting leading Humanity’s Last Exam scores of 59.1% in standard evaluations and up to 65% in tool-assisted runs on leaderboards from BenchLM and Artificial Analysis. This reflects sustained gains in advanced reasoning, adaptive thinking, and tool integration on the 2,500-question benchmark spanning math, sciences, and humanities, outpacing OpenAI’s GPT-6 Astra and Meta’s Muse Spark variants. With three months left in 2026, further Claude iterations could push scores higher, though benchmark revisions and competitive releases introduce uncertainty around exact thresholds like 60% or 65%. Traders view the implied probabilities as reflecting Anthropic’s edge in frontier knowledge tasks.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於$109,107 交易量
55% 以上
90%
60% 以上
43%
65% 以上
19%
70% 以上
12%
75%+
9%
$109,107 交易量
55% 以上
90%
60% 以上
43%
65% 以上
19%
70% 以上
12%
75%+
9%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
市場開放時間: Jul 23, 2026, 6:42 PM ET
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Anthropic’s September 2026 launches of Claude Fable 5.1 and Mythos 5.1 have driven the current trader consensus, with these models posting leading Humanity’s Last Exam scores of 59.1% in standard evaluations and up to 65% in tool-assisted runs on leaderboards from BenchLM and Artificial Analysis. This reflects sustained gains in advanced reasoning, adaptive thinking, and tool integration on the 2,500-question benchmark spanning math, sciences, and humanities, outpacing OpenAI’s GPT-6 Astra and Meta’s Muse Spark variants. With three months left in 2026, further Claude iterations could push scores higher, though benchmark revisions and competitive releases introduce uncertainty around exact thresholds like 60% or 65%. Traders view the implied probabilities as reflecting Anthropic’s edge in frontier knowledge tasks.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於



警惕外部連結哦。
警惕外部連結哦。
Frequently Asked Questions