Claude Opus 5 benchmarks #1 but its chatty diffs slow my agent loop
Claude Opus 5 tops the Artificial Analysis leaderboard at #1, but its chatty diffs and mid-task questions make it a pain to ship in agent loops.
Claude Opus 5 tops the Artificial Analysis leaderboard at #1, but its chatty diffs and mid-task questions make it a pain to ship in agent loops.
Claude Opus 5 keeps Opus 4.8 pricing and genuinely improves multi-file code refactors, but thinking-on by default quietly inflates your token bill 2-3x.
Ran MiniMax M3 locally against Qwen3-Coder and GLM-5.1. Strong on scoped refactors, weak on long context. A solid second model, not a daily driver.