OpenAI’s late-September releases of GPT-6 Astra, GPT-6 Sol, and the always-on Dots agent platform have positioned the lab as the clearest contender for second place in agentic performance behind Anthropic. Recent benchmark snapshots show Anthropic’s Claude Opus 5.5 and Fable 5.1 families leading most agentic evaluations for tool use, long-horizon reliability, and computer-use workflows, while OpenAI models sit immediately behind on composite indices and real-user arenas. Chinese labs including Moonshot, DeepSeek, and Zhipu AI post competitive open-weight scores on select terminal and browser tasks but trail overall. Meta’s Muse Spark and Google’s Gemini updates add further pressure, yet trader-implied odds reflect OpenAI’s edge in enterprise integrations and rapid iteration. With December resolution approaching, any new frontier releases or benchmark shifts could reorder the field.
Tóm tắt AI thử nghiệm tham chiếu dữ liệu Polymarket. Đây không phải tư vấn giao dịch và không ảnh hưởng đến cách thị trường này được giải quyết. · Cập nhậtView resolved




















Cẩn thận với liên kết bên ngoài.
Cẩn thận với liên kết bên ngoài.
Câu hỏi thường gặp