Google's Gemini models have shown steady gains on Humanity's Last Exam, the expert-level benchmark of 2,500 questions spanning science, math, and humanities created by the Center for AI Safety and Scale AI, yet currently trail leaders like Anthropic's Claude Fable 5 and Opus 5 at 55%+ accuracy. Recent Gemini 3.x variants, including 3.7 Flash at 47.9% and 3.1 Pro previews near 46-47%, reflect iterative scaling and reasoning enhancements amid competition from OpenAI's GPT-5 series and others, but no single release has yet pushed Gemini into the 50% threshold. With the market resolving at year-end 2026 and roughly four months remaining, trader focus centers on upcoming model iterations, internal scaling efforts, and whether Google can close the gap before saturation effects or rival advances dominate.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於$28,445 交易量
50% 以上
74%
55%+
47%
60% 以上
34%
65%及以上
14%
70% 以上
8%
$28,445 交易量
50% 以上
74%
55%+
47%
60% 以上
34%
65%及以上
14%
70% 以上
8%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
市場開放時間: Jul 23, 2026, 6:56 PM ET
Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolver
0x65070BE91...Google's Gemini models have shown steady gains on Humanity's Last Exam, the expert-level benchmark of 2,500 questions spanning science, math, and humanities created by the Center for AI Safety and Scale AI, yet currently trail leaders like Anthropic's Claude Fable 5 and Opus 5 at 55%+ accuracy. Recent Gemini 3.x variants, including 3.7 Flash at 47.9% and 3.1 Pro previews near 46-47%, reflect iterative scaling and reasoning enhancements amid competition from OpenAI's GPT-5 series and others, but no single release has yet pushed Gemini into the 50% threshold. With the market resolving at year-end 2026 and roughly four months remaining, trader focus centers on upcoming model iterations, internal scaling efforts, and whether Google can close the gap before saturation effects or rival advances dominate.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於



警惕外部連結哦。
警惕外部連結哦。
Frequently Asked Questions