Recent releases from OpenAI, including GPT-6.1 Sol (max) in late September 2026 reaching 90.1% expected performance on MathArena benchmarks, and Google's Gemini 4 Argon topping LMArena Text Arena math rankings near 1530, have tightened trader consensus on frontier models crossing elevated score thresholds by year-end. Anthropic's Claude Opus and Fable variants remain competitive in direct arena evaluations and proof-based tasks like USAMO, while open-weight contenders from Moonshot and Z.ai show strong but secondary gains on AIME-style problems. With only three months left, the primary swing factors are whether labs ship further iterations or agentic enhancements before December 31, as benchmarks like ArXivMath and BrokenArXiv continue to expose gaps between final-answer accuracy and rigorous reasoning. Market-implied odds reflect this rapid but uncertain pace of capability gains.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · UpdatedView resolved

Beware of external links.
Beware of external links.
Frequently Asked Questions