Anthropic’s September 1 release of Claude Fable 5.1 and Mythos 5.1 has driven recent gains on Humanity’s Last Exam, with the models posting leading scores of 59–65% across leaderboards that incorporate adaptive reasoning and tool use on the 2,500-question PhD-level benchmark. These iterative improvements in reasoning chains, context handling, and agentic workflows have widened Anthropic’s edge over OpenAI’s GPT-5.4/6 variants and Meta’s Muse models, which trail in the low-to-mid 50s. Traders are pricing in continued progress through additional 2026 releases and efficiency optimizations before year-end, while noting that benchmark saturation risks and potential safety-related restrictions on frontier capabilities remain swing factors for thresholds above 60%.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert$106,129 Vol.
55 %+
81%
60 %+
43%
65 %+
19%
70 %+
14%
75%+
9%
$106,129 Vol.
55 %+
81%
60 %+
43%
65 %+
19%
70 %+
14%
75%+
9%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Markt eröffnet: Jul 23, 2026, 6:42 PM ET
Abwicklungsquelle
https://agi.safe.ai/Abwickler
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Abwicklungsquelle
https://agi.safe.ai/Abwickler
0x65070BE91...Anthropic’s September 1 release of Claude Fable 5.1 and Mythos 5.1 has driven recent gains on Humanity’s Last Exam, with the models posting leading scores of 59–65% across leaderboards that incorporate adaptive reasoning and tool use on the 2,500-question PhD-level benchmark. These iterative improvements in reasoning chains, context handling, and agentic workflows have widened Anthropic’s edge over OpenAI’s GPT-5.4/6 variants and Meta’s Muse models, which trail in the low-to-mid 50s. Traders are pricing in continued progress through additional 2026 releases and efficiency optimizations before year-end, while noting that benchmark saturation risks and potential safety-related restrictions on frontier capabilities remain swing factors for thresholds above 60%.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert



Vorsicht bei externen Links.
Vorsicht bei externen Links.
Häufig gestellte Fragen