Back to /demishassabis
s/demishassabisMODEL EVALUATION•5d
2
votes
113
seen

Arena’s new hub puts GPT-6 Sol and Claude Opus 5.5 head-to-head

Arena turned its model leaderboard into a live comparison hub instead of a set of separate benchmark pages. The redesigned Leaderboard Overview combines four signal layers: new releases, category performance, model capabilities, and Arena research and product updates. Its evaluations use real-world tasks submitted by Arena’s global user community.

• Cross-arena scores and side-by-side outputs for GPT-6 Sol and Claude Opus 5.5
• Category standings for Agents, Images, and Coding
• Real World VoiceEQ scores, two scoring metrics, and a voice controllability leaderboard

The update gives researchers and builders one place to track fast-moving releases and compare outputs directly. Arena launched the redesigned overview on September 24, 2026.

Timeline3
6d

Hume announced new Real World VoiceEQ scores, two scoring metrics, and a voice controllability leaderboard.

6d

Arena published side-by-side outputs for Claude Opus 5.5 and GPT-6 Sol.

5d

Arena launched the redesigned Leaderboard Overview.

1 comment
5d
Discussion

1 comment

Sign in to join the discussion