OpenAI’s GPT-5.4 Pro and GPT-5.6 Sol currently trail Anthropic’s Claude Fable 5.1 on Humanity’s Last Exam leaderboards, with reported scores in the mid-to-high 50s percent range versus Claude’s 59–65 percent on expert-authored questions spanning advanced math, physics, and specialized domains. Recent OpenAI releases emphasize agentic workflows, token efficiency, and coding benchmarks, yet HLE’s design—crowdsourced graduate-level items resistant to retrieval—continues to highlight gaps in deep reasoning. Market-implied odds reflect strong trader consensus that OpenAI will clear 50–55 percent by year-end through iterative scaling and post-training, while higher thresholds remain contested amid rapid competitive releases from Anthropic and others. Key catalysts include any new GPT variants or capability jumps before December 31, 2026.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado$107,820 Vol.
55% o más
78%
60% o más
33%
65% o más
16%
70% o más
10%
$107,820 Vol.
55% o más
78%
60% o más
33%
65% o más
16%
70% o más
10%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Mercado abierto: Jul 23, 2026, 6:53 PM ET
Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolver
0x65070BE91...OpenAI’s GPT-5.4 Pro and GPT-5.6 Sol currently trail Anthropic’s Claude Fable 5.1 on Humanity’s Last Exam leaderboards, with reported scores in the mid-to-high 50s percent range versus Claude’s 59–65 percent on expert-authored questions spanning advanced math, physics, and specialized domains. Recent OpenAI releases emphasize agentic workflows, token efficiency, and coding benchmarks, yet HLE’s design—crowdsourced graduate-level items resistant to retrieval—continues to highlight gaps in deep reasoning. Market-implied odds reflect strong trader consensus that OpenAI will clear 50–55 percent by year-end through iterative scaling and post-training, while higher thresholds remain contested amid rapid competitive releases from Anthropic and others. Key catalysts include any new GPT variants or capability jumps before December 31, 2026.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado



Cuidado con los enlaces externos.
Cuidado con los enlaces externos.
Preguntas frecuentes