Claude Opus 4.8 performs better on the benchmarks. Anthropic also says it has better judgment when working on longer tasks.

That second claim is the one to care about. For a coding agent, the useful measure is how much work you can accept after a proportionate review.

A migration across a hundred files may be impressive. If I have to reconstruct the whole migration to trust it, the agent did the typing and left me the expensive part. Better judgment would show up before that: finding an ambiguity and asking about it before carrying one interpretation through the codebase.

I can only test this on work I understand well. If the model catches the ambiguity I planted before changing the code, its judgment saved me review work.

That benchmark will not fit neatly in a launch chart.