Claude Opus 5 is out, and Claire Vo ran it through her 7-model, 6-task How I AI benchmark against GPT-5.6 Sol, Sonnet 5, and Gemini 3.1 Pro. The verdict is not a clean win for Opus 5. It landed one use case where it scored straight 5s. It also refused to touch a merge conflict during a live coding session. That tension between brilliance and obstruction is the core finding here.

The review goes beyond scores. Vo's personality comparison between Opus 5 and GPT-5.6 Sol, including asking both models who is smarter, you or me, surfaces something useful about how model character affects real workflows. She also names the verbosity problem directly, calling it Claude Slop, and argues we have hit an intelligence overhang where raw capability no longer determines which model you should actually use.

The benchmark leaderboard reveal at 18:30 and the final usage plan at 23:25 are worth watching in full. The framing here is not which model wins a number. It is which variables matter now that intelligence is no longer the differentiator. If you are making daily model choices, this one has the specifics to shift your thinking.

[READ ORIGINAL →]