Skip to main content
icon for Next Grok Model (4.6+): Humanity’s Last Exam Debut?

Next Grok Model (4.6+): Humanity’s Last Exam Debut?

icon for Next Grok Model (4.6+): Humanity’s Last Exam Debut?

Next Grok Model (4.6+): Humanity’s Last Exam Debut?

BARU
Dec 31, 2026
Polymarket

$3,995 Vol.

Polymarket

35%+

$632 Vol.

93%

40%+

$1,704 Vol.

48%

45%+

$830 Vol.

11%

50%+

$441 Vol.

5%

55%+

$388 Vol.

46%

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".**Grok 4.6 launched August 12, 2026, as xAI’s latest post-training update focused on long-running agents and multi-step visual or coding tasks, following Grok 4.5 by about five weeks.** It builds on the Grok 4 series that already posted strong HLE results—base scores near 25-27% and up to 50%+ with tools or Heavy multi-agent setups—yet current public leaderboards show it trailing Anthropic’s Claude Fable 5 and Opus 5 (mid-50s) and other frontier models around 40-55%. Traders are watching whether rapid xAI iteration, combined with heavy compute scaling, can push the next 4.6+ checkpoint onto the HLE leaderboard first or with a leading score before competitors refresh. Key near-term catalysts include official HLE submissions, any new reasoning-effort modes, and competing model releases from Anthropic, OpenAI, or Google that could shift the frontier. HLE’s design as a saturated, expert-authored benchmark keeps scores climbing gradually, rewarding verifiable capability gains over marketing claims.

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No".

If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results.

A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify.

The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere.

If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered.

A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market.

The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
Volume
$3,995
Tanggal Berakhir
Dec 31, 2026
Pasar Dibuka
Aug 10, 2026, 6:25 PM ET

Sumber Resolusi

https://agi.safe.ai/
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".**Grok 4.6 launched August 12, 2026, as xAI’s latest post-training update focused on long-running agents and multi-step visual or coding tasks, following Grok 4.5 by about five weeks.** It builds on the Grok 4 series that already posted strong HLE results—base scores near 25-27% and up to 50%+ with tools or Heavy multi-agent setups—yet current public leaderboards show it trailing Anthropic’s Claude Fable 5 and Opus 5 (mid-50s) and other frontier models around 40-55%. Traders are watching whether rapid xAI iteration, combined with heavy compute scaling, can push the next 4.6+ checkpoint onto the HLE leaderboard first or with a leading score before competitors refresh. Key near-term catalysts include official HLE submissions, any new reasoning-effort modes, and competing model releases from Anthropic, OpenAI, or Google that could shift the frontier. HLE’s design as a saturated, expert-authored benchmark keeps scores climbing gradually, rewarding verifiable capability gains over marketing claims.

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No".

If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results.

A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify.

The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere.

If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered.

A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market.

The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
Volume
$3,995
Tanggal Berakhir
Dec 31, 2026
Pasar Dibuka
Aug 10, 2026, 6:25 PM ET

Sumber Resolusi

https://agi.safe.ai/
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".

Hati-hati dengan link eksternal.

Pertanyaan yang Sering Diajukan

"Next Grok Model (4.6+): Humanity’s Last Exam Debut?" adalah pasar prediksi di Polymarket dengan 5 hasil yang mungkin di mana trader membeli dan menjual saham berdasarkan apa yang mereka yakini akan terjadi. Hasil terdepan saat ini adalah "35%+" di 93%, diikuti oleh "40%+" di 48%. Harga mencerminkan probabilitas crowd-sourced real-time. Misalnya, saham yang dihargai 93¢ menyiratkan bahwa pasar secara kolektif memberikan peluang 93% pada hasil tersebut. Peluang ini bergeser terus-menerus saat trader bereaksi terhadap perkembangan dan informasi baru. Saham dengan hasil yang benar bisa ditukarkan seharga $1 setiap saham saat pasar diselesaikan.

"Next Grok Model (4.6+): Humanity’s Last Exam Debut?" adalah pasar yang baru dibuat di Polymarket, diluncurkan pada Aug 11, 2026. Sebagai pasar awal, ini adalah kesempatanmu untuk menjadi salah satu trader pertama yang menetapkan peluang dan membangun sinyal harga awal pasar. Kamu juga bisa menandai halaman ini untuk melacak volume dan aktivitas trading seiring pasar mendapatkan traksi.

Untuk trading di "Next Grok Model (4.6+): Humanity’s Last Exam Debut?," jelajahi 5 hasil yang tersedia di halaman ini. Setiap hasil menampilkan harga saat ini yang mewakili probabilitas tersirat pasar. Untuk mengambil posisi, pilih hasil yang menurutmu paling mungkin, pilih "Ya" untuk mendukungnya atau "Tidak" untuk menentangnya, masukkan jumlahmu, dan klik "Trade." Jika hasil pilihanmu benar saat pasar diselesaikan, saham "Ya" kamu membayar $1 masing-masing. Jika salah, mereka membayar $0. Kamu juga bisa menjual sahammu kapan saja sebelum resolusi jika kamu ingin mengamankan keuntungan atau memotong kerugian.

Unggulan saat ini untuk "Next Grok Model (4.6+): Humanity’s Last Exam Debut?" adalah "35%+" di 93%, yang berarti pasar memberikan peluang 93% pada hasil tersebut. Hasil terdekat berikutnya adalah "40%+" di 48%. Peluang ini diperbarui secara real-time saat trader membeli dan menjual saham, sehingga mencerminkan pandangan kolektif terbaru tentang apa yang paling mungkin terjadi. Cek kembali secara rutin atau tandai halaman ini untuk mengikuti bagaimana peluang bergeser saat informasi baru muncul.

Aturan resolusi untuk "Next Grok Model (4.6+): Humanity’s Last Exam Debut?" mendefinisikan dengan tepat apa yang harus terjadi agar setiap hasil dinyatakan sebagai pemenang — termasuk sumber data resmi yang digunakan untuk menentukan hasilnya. Kamu bisa meninjau kriteria resolusi lengkap di bagian "Aturan" di halaman ini di atas komentar. Kami menyarankan membaca aturan dengan cermat sebelum trading, karena mereka menentukan kondisi tepat, kasus khusus, dan sumber yang mengatur bagaimana pasar ini diselesaikan.