Anthropic holds the leading market-implied odds at 56.5% for topping LiveBench Coding by end of October, ahead of OpenAI at 28.5%, reflecting trader consensus on its Claude Opus and Fable series models' consistent strength in contamination-resistant coding evaluations. Recent benchmark updates highlight Anthropic's edge in practical software engineering tasks, driven by iterative improvements in reasoning chains and tool use that outpace competitors on dynamic prompts. OpenAI's o-series models lead broader LiveBench aggregates through strong overall performance, yet trail in coding-specific subsets where historical patterns favor Anthropic's focus on developer workflows. With the resolution window under two months away, any new frontier releases or scoring adjustments could narrow the gap, underscoring the uncertainty typical in rapidly evolving large language model capabilities.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于Anthropic 67%
OpenAI 29%
Mistral 2.0%
DeepSeek 1.6%

Anthropic
56%

OpenAI
29%

Mistral
2%

DeepSeek
2%

Z.ai
2%

谷歌
1%

亚马逊
1%

Nvidia
1%

Thinky
1%

MiniMax
1%

微软
1%

百度
1%

阿里巴巴
1%

小米
<1%

Moonshot
<1%

StepFun
<1%

SpaceXAI
<1%

Meta
<1%

腾讯
<1%

美团
<1%

字节跳动
<1%
Anthropic 67%
OpenAI 29%
Mistral 2.0%
DeepSeek 1.6%

Anthropic
56%

OpenAI
29%

Mistral
2%

DeepSeek
2%

Z.ai
2%

谷歌
1%

亚马逊
1%

Nvidia
1%

Thinky
1%

MiniMax
1%

微软
1%

百度
1%

阿里巴巴
1%

小米
<1%

Moonshot
<1%

StepFun
<1%

SpaceXAI
<1%

Meta
<1%

腾讯
<1%

美团
<1%

字节跳动
<1%
Results from the “Coding” column of the leaderboard at https://livebench.ai/#/?cats=Coding, with the latest available LiveBench release selected and the category set to “Coding,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
市场开放时间: Aug 12, 2026, 8:05 PM ET
Results from the “Coding” column of the leaderboard at https://livebench.ai/#/?cats=Coding, with the latest available LiveBench release selected and the category set to “Coding,” will be used to resolve this market.
Models will be ranked according to the specified score, with higher scores ranked ahead of lower scores. If two or more models have exactly the same score as displayed on the leaderboard, the model with the lower listed "cost per successful task" will be ranked ahead. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., if the two models are tied by exact score and cost per successful task, “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the LiveBench leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to “Other.”
Anthropic holds the leading market-implied odds at 56.5% for topping LiveBench Coding by end of October, ahead of OpenAI at 28.5%, reflecting trader consensus on its Claude Opus and Fable series models' consistent strength in contamination-resistant coding evaluations. Recent benchmark updates highlight Anthropic's edge in practical software engineering tasks, driven by iterative improvements in reasoning chains and tool use that outpace competitors on dynamic prompts. OpenAI's o-series models lead broader LiveBench aggregates through strong overall performance, yet trail in coding-specific subsets where historical patterns favor Anthropic's focus on developer workflows. With the resolution window under two months away, any new frontier releases or scoring adjustments could narrow the gap, underscoring the uncertainty typical in rapidly evolving large language model capabilities.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于
警惕外部链接哦。
警惕外部链接哦。
常见问题