The closely matched odds around 50% underscore intense competition among frontier AI labs for top mathematical reasoning on LiveBench, whose dynamic, contamination-resistant tasks draw from recent competitions and olympiads. OpenAI's o-series models like o3-mini have posted strong recent overall LiveBench scores, while Chinese developers including Moonshot (Kimi variants), Z.ai (GLM series), and DeepSeek have led or closely challenged on dedicated math benchmarks such as MATH and AIME26. Anthropic's Claude models, Google Gemini, and others remain within striking distance amid rapid iteration. LiveBench's monthly question refreshes and expected model updates before November's end introduce meaningful volatility that could shift the leaderboard.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · UpdatedView resolved






















Beware of external links.
Beware of external links.
Frequently Asked Questions