Google's Gemini models currently trail the HLE frontier, with recent variants like Gemini 3.1 Pro and 3.8 Flash posting mid-to-high 40% scores on text-only or no-tools evaluations while Anthropic's Claude Fable 5.1 and Opus 5 lead at 55-65% across major leaderboards. Rapid gains since the benchmark's early 2025 launch have pushed the overall frontier from single digits into the mid-40s to low-60s percent range, driven by advances in multi-step reasoning on the 2,500 expert-level questions spanning math, physics, and other domains. Trader focus centers on Google's aggressive release cadence, including mid-2026 Flash iterations and pre-training for Gemini 4 expected later this year, alongside potential gains in agentic capabilities that could close the gap before the December 31, 2026 resolution.
Résumé expérimental généré par IA à partir des données Polymarket. Ceci n'est pas un conseil de trading et ne joue aucun rôle dans la résolution de ce marché. · Mis à jour$80,649 Vol.
50 %+
86%
55 %+
32%
60 %+
16%
65 % ou plus
12%
70 %+
4%
$80,649 Vol.
50 %+
86%
55 %+
32%
60 %+
16%
65 % ou plus
12%
70 %+
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Marché ouvert : Jul 23, 2026, 6:56 PM ET
Résolveur
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Résolveur
0x65070BE91...Google's Gemini models currently trail the HLE frontier, with recent variants like Gemini 3.1 Pro and 3.8 Flash posting mid-to-high 40% scores on text-only or no-tools evaluations while Anthropic's Claude Fable 5.1 and Opus 5 lead at 55-65% across major leaderboards. Rapid gains since the benchmark's early 2025 launch have pushed the overall frontier from single digits into the mid-40s to low-60s percent range, driven by advances in multi-step reasoning on the 2,500 expert-level questions spanning math, physics, and other domains. Trader focus centers on Google's aggressive release cadence, including mid-2026 Flash iterations and pre-training for Gemini 4 expected later this year, alongside potential gains in agentic capabilities that could close the gap before the December 31, 2026 resolution.
Résumé expérimental généré par IA à partir des données Polymarket. Ceci n'est pas un conseil de trading et ne joue aucun rôle dans la résolution de ce marché. · Mis à jour



Méfiez-vous des liens externes.
Méfiez-vous des liens externes.
Questions fréquentes