The benchmark to care about
Claude Opus 4.8 performs better on the benchmarks. Anthropic also says it has better judgment when working on longer tasks.
That second claim is the one to care about. For a coding agent, the useful measure is how much work you can accept after a proportionate review.
A migration across a hundred files may be impressive. If I have to reconstruct the whole migration to trust it, the agent did the typing and left me the expensive part. Better judgment would show up before that: finding an ambiguity and asking about it before carrying one interpretation through the codebase.
I can only test this on work I understand well. If the model catches the ambiguity I planted before changing the code, its judgment saved me review work.
That benchmark will not fit neatly in a launch chart.