Grok models currently trail on Humanity’s Last Exam, with Grok 4.6 at 42.9% and earlier variants near 35-37% on verified September 2026 leaderboards, well behind Anthropic’s Claude Fable 5.1 at 59-65% and other frontier systems above 50%. Trader focus centers on xAI’s release cadence and whether upcoming Grok iterations, multi-agent reasoning modes, or tool-augmented evaluations can narrow this gap before year-end, as the expert-authored 2,500-question benchmark continues to show substantial headroom across math, physics, and specialized domains. Recent steady gains by xAI reflect broader industry scaling and architectural tweaks, yet Anthropic’s edge in calibrated reasoning has sustained the lead; any major xAI announcement or benchmark update could shift implied probabilities for thresholds like 45%+ versus higher marks.
Eksperimental na AI-generated summary na nire-reference ang Polymarket data. Hindi ito trading advice at wala itong papel sa kung paano nire-resolve ang market na ito. · Na-updateHighest Grok score on Humanity’s Last Exam in 2026?
$118,977 Vol.
45%+
81%
50%+
48%
55%+
31%
60%+
19%
65%+
5%
$118,977 Vol.
45%+
81%
50%+
48%
55%+
31%
60%+
19%
65%+
5%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Binuksan ang Market: Jul 23, 2026, 6:46 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Grok models currently trail on Humanity’s Last Exam, with Grok 4.6 at 42.9% and earlier variants near 35-37% on verified September 2026 leaderboards, well behind Anthropic’s Claude Fable 5.1 at 59-65% and other frontier systems above 50%. Trader focus centers on xAI’s release cadence and whether upcoming Grok iterations, multi-agent reasoning modes, or tool-augmented evaluations can narrow this gap before year-end, as the expert-authored 2,500-question benchmark continues to show substantial headroom across math, physics, and specialized domains. Recent steady gains by xAI reflect broader industry scaling and architectural tweaks, yet Anthropic’s edge in calibrated reasoning has sustained the lead; any major xAI announcement or benchmark update could shift implied probabilities for thresholds like 45%+ versus higher marks.
Eksperimental na AI-generated summary na nire-reference ang Polymarket data. Hindi ito trading advice at wala itong papel sa kung paano nire-resolve ang market na ito. · Na-update



Mag-ingat sa mga external link.
Mag-ingat sa mga external link.
Mga Madalas na Tanong