xAI's Grok models currently trail frontier leaders on Humanity’s Last Exam, a 2,500-question expert-authored benchmark of graduate-level reasoning across math, sciences, and humanities, with recent Grok 4.6 and 4.5 variants posting verified scores around 42-43% on official leaderboards. Anthropic’s Claude Fable 5.1 and Opus 5 lead at 64-65%, followed by Meta’s Muse Spark near 62% and OpenAI GPT variants in the high 50s, reflecting advantages in reasoning-focused training and scale. Trader sentiment centers on whether xAI’s next iterations or tool-augmented variants can close this gap before December 31, 2026, amid industry-wide benchmark gains but slower verified progress from Grok releases to date. Key catalysts include upcoming model drops and official HLE leaderboard updates.
สรุปจาก AI ทดลองที่อ้างอิงข้อมูลจาก Polymarket ไม่ใช่คำแนะนำในการเทรดและไม่มีผลต่อการตัดสินตลาดนี้ · อัปเดตแล้ว$113,897 ปริมาณ
45%+
76%
50%+
57%
55%+
32%
60%+
22%
65%+
10%
$113,897 ปริมาณ
45%+
76%
50%+
57%
55%+
32%
60%+
22%
65%+
10%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
ตลาดเปิดเมื่อ: Jul 23, 2026, 6:46 PM ET
แหล่งข้อมูลการตัดสินผล
https://agi.safe.ai/ผู้ตัดสินผล
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
แหล่งข้อมูลการตัดสินผล
https://agi.safe.ai/ผู้ตัดสินผล
0x65070BE91...xAI's Grok models currently trail frontier leaders on Humanity’s Last Exam, a 2,500-question expert-authored benchmark of graduate-level reasoning across math, sciences, and humanities, with recent Grok 4.6 and 4.5 variants posting verified scores around 42-43% on official leaderboards. Anthropic’s Claude Fable 5.1 and Opus 5 lead at 64-65%, followed by Meta’s Muse Spark near 62% and OpenAI GPT variants in the high 50s, reflecting advantages in reasoning-focused training and scale. Trader sentiment centers on whether xAI’s next iterations or tool-augmented variants can close this gap before December 31, 2026, amid industry-wide benchmark gains but slower verified progress from Grok releases to date. Key catalysts include upcoming model drops and official HLE leaderboard updates.
สรุปจาก AI ทดลองที่อ้างอิงข้อมูลจาก Polymarket ไม่ใช่คำแนะนำในการเทรดและไม่มีผลต่อการตัดสินตลาดนี้ · อัปเดตแล้ว



ระวังลิงก์ภายนอก
ระวังลิงก์ภายนอก
คำถามที่พบบ่อย