**Recent agent benchmarks released September 18 highlight a tight race among multiple labs for third-best AI agent capabilities by end of November.** Moonshot’s Kimi K3 posts competitive scores on several agentic evals, including strong tool-use and multi-step workflow results that support its 37.5% market-implied odds, while OpenAI’s GPT-5.6/6 series and Anthropic’s Claude Opus/Fable models lead or place near the top on verified leaderboards. SpaceXAI’s Grok 4.6 and Amazon’s offerings also register meaningful gains in planning and execution benchmarks. High uncertainty reflected in the 50% levels for “Other” and numerous unnamed labs stems from rapid iteration in open-weight models from Chinese teams and ongoing enterprise deployments that can shift relative positioning quickly. Key swing factors include any new long-horizon agent releases or benchmark updates before November 30.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于Moonshot 38%
亚马逊 27%
SpaceXAI 15%
谷歌 12%

Moonshot
38%

亚马逊
27%

SpaceXAI
15%

谷歌
12%

Anthropic
12%

OpenAI
15%

Z.ai
5%

Meta
3%

阿里巴巴
3%

Mistral
3%

腾讯
2%

英伟达
2%

百度
2%

美团
2%

微软
2%

DeepSeek
2%

小米
2%

MiniMax
2%

字节跳动
2%
Moonshot 38%
亚马逊 27%
SpaceXAI 15%
谷歌 12%

Moonshot
38%

亚马逊
27%

SpaceXAI
15%

谷歌
12%

Anthropic
12%

OpenAI
15%

Z.ai
5%

Meta
3%

阿里巴巴
3%

Mistral
3%

腾讯
2%

英伟达
2%

百度
2%

美团
2%

微软
2%

DeepSeek
2%

小米
2%

MiniMax
2%

字节跳动
2%
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Labs" will be used to resolve this market.
Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
AI companies will be ordered primarily by their Lab Rank at the market’s check time. If the results based on the lab ranking are ambiguous or unavailable, the relevant AI companies will be ordered according to their highest-ranking AI model in the leaderboard’s “Models” view. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of AI lab/company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies third place under this ranking.
The resolution source for this market is the arena.ai Agent Arena Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
市场开放时间: Sep 17, 2026, 8:04 PM ET
Results from the "Rank" column under the "Agent Arena" Leaderboard tab at https://arena.ai/leaderboard/agent filtered for "Labs" will be used to resolve this market.
Models marked “AutoEval” at the applicable check time will not be considered, regardless of whether they display a rank or score.
AI companies will be ordered primarily by their Lab Rank at the market’s check time. If the results based on the lab ranking are ambiguous or unavailable, the relevant AI companies will be ordered according to their highest-ranking AI model in the leaderboard’s “Models” view. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of AI lab/company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies third place under this ranking.
The resolution source for this market is the arena.ai Agent Arena Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
**Recent agent benchmarks released September 18 highlight a tight race among multiple labs for third-best AI agent capabilities by end of November.** Moonshot’s Kimi K3 posts competitive scores on several agentic evals, including strong tool-use and multi-step workflow results that support its 37.5% market-implied odds, while OpenAI’s GPT-5.6/6 series and Anthropic’s Claude Opus/Fable models lead or place near the top on verified leaderboards. SpaceXAI’s Grok 4.6 and Amazon’s offerings also register meaningful gains in planning and execution benchmarks. High uncertainty reflected in the 50% levels for “Other” and numerous unnamed labs stems from rapid iteration in open-weight models from Chinese teams and ongoing enterprise deployments that can shift relative positioning quickly. Key swing factors include any new long-horizon agent releases or benchmark updates before November 30.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于
警惕外部链接哦。
警惕外部链接哦。
常见问题