Recent releases from Anthropic, including Claude Fable 5 and Opus 5, have pushed coding arena Elo scores and agentic benchmarks like SWE-bench Verified into the mid-90s percent range, reflecting strong gains in multi-step software engineering tasks. OpenAI’s GPT-5.6 variants and Google’s Gemini 3 series remain close competitors, while open-weight models such as DeepSeek V4 and Kimi K3 narrow the gap on standardized coding arenas through efficient scaling. Trader focus centers on whether incremental 2026 updates or entirely new frontier models before year-end will clear the remaining arena threshold, with release cadence and verified capability jumps as the key swing variables.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于$185,152 交易量
1560
44%
1580
24%
1600
14%
$185,152 交易量
1560
44%
1580
24%
1600
14%
Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
市场开放时间: Apr 2, 2026, 6:09 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent releases from Anthropic, including Claude Fable 5 and Opus 5, have pushed coding arena Elo scores and agentic benchmarks like SWE-bench Verified into the mid-90s percent range, reflecting strong gains in multi-step software engineering tasks. OpenAI’s GPT-5.6 variants and Google’s Gemini 3 series remain close competitors, while open-weight models such as DeepSeek V4 and Kimi K3 narrow the gap on standardized coding arenas through efficient scaling. Trader focus centers on whether incremental 2026 updates or entirely new frontier models before year-end will clear the remaining arena threshold, with release cadence and verified capability jumps as the key swing variables.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于



警惕外部链接哦。
警惕外部链接哦。
常见问题