Anthropic’s Claude Fable 5.1 and Opus 5 currently lead public HLE leaderboards with scores of 59–65% as of mid-September 2026, while OpenAI’s strongest reported entries such as GPT-6 Astra and GPT-5.4 Pro sit in the mid-50s percent range on tool-augmented or high-effort runs. Humanity’s Last Exam, a 2,500-question expert-authored benchmark spanning mathematics, physics, biology, and humanities, was designed to resist saturation and internet retrieval, so gains depend on genuine reasoning advances rather than memorization. OpenAI’s recent GPT-5 and GPT-6 iterations have closed some ground through improved agentic tool use, yet Anthropic’s consistent top placements reflect stronger competitive positioning on this frontier metric. Traders will watch for any late-2026 OpenAI model drops or benchmark updates that could shift the year’s highest OpenAI score before resolution.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড$107,737 Vol.
55%+
78%
60%+
32%
65%+
16%
70%+
10%
$107,737 Vol.
55%+
78%
60%+
32%
65%+
16%
70%+
10%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
মার্কেট ওপেন হয়েছে: Jul 23, 2026, 6:53 PM ET
রেজলভার
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
রেজলভার
0x65070BE91...Anthropic’s Claude Fable 5.1 and Opus 5 currently lead public HLE leaderboards with scores of 59–65% as of mid-September 2026, while OpenAI’s strongest reported entries such as GPT-6 Astra and GPT-5.4 Pro sit in the mid-50s percent range on tool-augmented or high-effort runs. Humanity’s Last Exam, a 2,500-question expert-authored benchmark spanning mathematics, physics, biology, and humanities, was designed to resist saturation and internet retrieval, so gains depend on genuine reasoning advances rather than memorization. OpenAI’s recent GPT-5 and GPT-6 iterations have closed some ground through improved agentic tool use, yet Anthropic’s consistent top placements reflect stronger competitive positioning on this frontier metric. Traders will watch for any late-2026 OpenAI model drops or benchmark updates that could shift the year’s highest OpenAI score before resolution.
Polymarket ডেটা রেফারেন্স করে পরীক্ষামূলক AI-জেনারেটেড সারাংশ। এটি ট্রেডিং পরামর্শ নয় এবং এই মার্কেট কীভাবে রেজলভ হয় তাতে কোনো ভূমিকা রাখে না। · আপডেটেড



বাহ্যিক লিংক থেকে সাবধান।
বাহ্যিক লিংক থেকে সাবধান।
সচরাচর জিজ্ঞাসা