Rapid iteration among frontier labs continues to elevate coding performance, with Anthropic’s Claude Opus 5 and Fable 5, OpenAI’s GPT-5.6 Sol, and Google’s Gemini 3.1/3.7 variants posting the highest scores on SWE-Bench Verified (often 88–97%) and specialized coding arenas that use blind head-to-head Elo ratings. Chinese open-weight releases such as DeepSeek V4 Pro, Kimi K3, and GLM-5.3 have narrowed the gap on both verified benchmarks and arena leaderboards, reflecting aggressive scaling and agentic tooling improvements. Traders are watching whether new model drops, expected before year-end, can push the top coding-arena Elo or percentage thresholds materially higher amid ongoing benchmark contamination concerns and the distinction between leaderboard gains and robust, multi-file software engineering.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日$185,222 Vol.
1560
44%
1580
24%
1600
12%
$185,222 Vol.
1560
44%
1580
24%
1600
12%
Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
マーケット開始日: Apr 2, 2026, 6:09 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Rapid iteration among frontier labs continues to elevate coding performance, with Anthropic’s Claude Opus 5 and Fable 5, OpenAI’s GPT-5.6 Sol, and Google’s Gemini 3.1/3.7 variants posting the highest scores on SWE-Bench Verified (often 88–97%) and specialized coding arenas that use blind head-to-head Elo ratings. Chinese open-weight releases such as DeepSeek V4 Pro, Kimi K3, and GLM-5.3 have narrowed the gap on both verified benchmarks and arena leaderboards, reflecting aggressive scaling and agentic tooling improvements. Traders are watching whether new model drops, expected before year-end, can push the top coding-arena Elo or percentage thresholds materially higher amid ongoing benchmark contamination concerns and the distinction between leaderboard gains and robust, multi-file software engineering.
Polymarketデータを参照したAI生成の実験的な要約。これは取引アドバイスではなく、このマーケットの解決方法には一切関係ありません。 · 更新日



外部リンクに注意してください。
外部リンクに注意してください。
よくある質問