Recent leaderboards as of September 2026 show Anthropic's Claude Fable 5.1 and Opus 5 leading Humanity's Last Exam with scores of 59–65% across tool-augmented and knowledge-focused variants, while OpenAI's GPT-6 Astra sits at 54–57% and earlier GPT-5 configurations trail further behind. This gap stems from Anthropic's stronger demonstrated reasoning and tool-use capabilities on the 2,500-question benchmark, which tests expert-level knowledge beyond saturated tests like MMLU. OpenAI continues iterative releases focused on large language models, but competitive pressure from Anthropic, Meta, and Google has narrowed its edge on frontier benchmarks. Traders watch for any late-2026 OpenAI model drops or capability jumps before year-end resolution.
Tóm tắt AI thử nghiệm tham chiếu dữ liệu Polymarket. Đây không phải tư vấn giao dịch và không ảnh hưởng đến cách thị trường này được giải quyết. · Cập nhật$107,820 KL.
55%+
78%
60%+
33%
65%+
16%
70%+
10%
$107,820 KL.
55%+
78%
60%+
33%
65%+
16%
70%+
10%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Thị trường mở: Jul 23, 2026, 6:53 PM ET
Người giải quyết
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Người giải quyết
0x65070BE91...Recent leaderboards as of September 2026 show Anthropic's Claude Fable 5.1 and Opus 5 leading Humanity's Last Exam with scores of 59–65% across tool-augmented and knowledge-focused variants, while OpenAI's GPT-6 Astra sits at 54–57% and earlier GPT-5 configurations trail further behind. This gap stems from Anthropic's stronger demonstrated reasoning and tool-use capabilities on the 2,500-question benchmark, which tests expert-level knowledge beyond saturated tests like MMLU. OpenAI continues iterative releases focused on large language models, but competitive pressure from Anthropic, Meta, and Google has narrowed its edge on frontier benchmarks. Traders watch for any late-2026 OpenAI model drops or capability jumps before year-end resolution.
Tóm tắt AI thử nghiệm tham chiếu dữ liệu Polymarket. Đây không phải tư vấn giao dịch và không ảnh hưởng đến cách thị trường này được giải quyết. · Cập nhật



Cẩn thận với liên kết bên ngoài.
Cẩn thận với liên kết bên ngoài.
Câu hỏi thường gặp