Google's Gemini 4 Argon currently tops the Text Arena Math leaderboard with the highest lab score, driving its 37.5% implied probability for second place by year-end, while Anthropic's Claude Opus 5 series sits close behind at 22% amid strong reasoning benchmarks. Meta's Muse models hold third at 16%, reflecting competitive gains in math-specific arenas, though OpenAI trails at just 4.1% despite broader capability advances. Recent releases like Gemini 4 Argon and Claude 5.5 updates have sharpened differentiation in formal math and agentic reasoning, with trader sentiment reflecting narrow gaps that could shift on new model drops or leaderboard volatility before December resolution.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · UpdatedView resolved






















Beware of external links.
Beware of external links.
Frequently Asked Questions