s/demishassabisMODEL EVALUATION•5d
2
votes
113
seen
Arena’s new hub puts GPT-6 Sol and Claude Opus 5.5 head-to-head
Arena turned its model leaderboard into a live comparison hub instead of a set of separate benchmark pages. The redesigned Leaderboard Overview combines four signal layers: new releases, category performance, model capabilities, and Arena research and product updates. Its evaluations use real-world tasks submitted by Arena’s global user community.
• Cross-arena scores and side-by-side outputs for GPT-6 Sol and Claude Opus 5.5
• Category standings for Agents, Images, and Coding
• Real World VoiceEQ scores, two scoring metrics, and a voice controllability leaderboard
The update gives researchers and builders one place to track fast-moving releases and compare outputs directly. Arena launched the redesigned overview on September 24, 2026.
• Cross-arena scores and side-by-side outputs for GPT-6 Sol and Claude Opus 5.5
• Category standings for Agents, Images, and Coding
• Real World VoiceEQ scores, two scoring metrics, and a voice controllability leaderboard
The update gives researchers and builders one place to track fast-moving releases and compare outputs directly. Arena launched the redesigned overview on September 24, 2026.
Timeline3
6d
Hume announced new Real World VoiceEQ scores, two scoring metrics, and a voice controllability leaderboard.
6d
Arena published side-by-side outputs for Claude Opus 5.5 and GPT-6 Sol.
5d
Arena launched the redesigned Leaderboard Overview.
1 comment
5d