The market reflects a highly contested field on LiveBench coding, where no single company holds a durable edge amid rapid releases of large language models from multiple labs. Recent September 2026 updates show Anthropic's Claude Opus 5.5 and Fable variants, OpenAI's GPT-6 Astra and Sol series, Alibaba's Qwen3 models, DeepSeek V4.1, and Meta's Muse Spark trading near the top in coding and agentic tasks, with scores clustered tightly below 90% on dynamic, contamination-resistant questions. Trader consensus at roughly 50% for many entrants highlights uncertainty around which lab will deliver the decisive November update, as benchmark refreshes, reasoning effort scaling, and competitive positioning continue to shift relative standings before resolution.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · UpdatedView resolved






















Beware of external links.
Beware of external links.
Frequently Asked Questions