Anthropic’s latest Claude 5-series models, including Claude Fable 5 and Opus 5, currently lead Humanity’s Last Exam leaderboards with scores of 54.9–64.7% as of mid-August 2026, outpacing OpenAI and Google entries on this 2,500-question expert-level benchmark spanning math, physics, biology, and other domains. Recent model releases and reasoning optimizations have driven these gains, reflecting Anthropic’s focus on frontier capabilities amid intense competition from GPT-5 variants and Gemini releases. With four months remaining in 2026, further Claude iterations, training scale-ups, or benchmark-specific fine-tuning could lift the top score, though AI labs’ typical release cycles and evaluation variability introduce uncertainty around exact thresholds.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado$60,385 Vol.
55% o más
86%
60% o más
56%
65% o más
16%
70% o más
8%
75%+
4%
$60,385 Vol.
55% o más
86%
60% o más
56%
65% o más
16%
70% o más
8%
75%+
4%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Mercado abierto: Jul 23, 2026, 6:42 PM ET
Fuente de resolución
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Fuente de resolución
https://agi.safe.ai/Resolver
0x65070BE91...Anthropic’s latest Claude 5-series models, including Claude Fable 5 and Opus 5, currently lead Humanity’s Last Exam leaderboards with scores of 54.9–64.7% as of mid-August 2026, outpacing OpenAI and Google entries on this 2,500-question expert-level benchmark spanning math, physics, biology, and other domains. Recent model releases and reasoning optimizations have driven these gains, reflecting Anthropic’s focus on frontier capabilities amid intense competition from GPT-5 variants and Gemini releases. With four months remaining in 2026, further Claude iterations, training scale-ups, or benchmark-specific fine-tuning could lift the top score, though AI labs’ typical release cycles and evaluation variability introduce uncertainty around exact thresholds.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado



Cuidado con los enlaces externos.
Cuidado con los enlaces externos.
Preguntas frecuentes