upvote
Because it's just Opus 5.5 and Haiku, and on OpenAI side Luna, that changed the calculus. Also, I find all of them, including GLM5.3 and DS4.1, to be 100% en par with American's frontiers model for CRUDS (90% of enterprise programming).

Tangentially, all of them would have broken quite badly custom ERPs from my own experience.

reply
DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.
reply
i mean we use max benchmarks because they have the best coverage, which sucks because very few people use max day to day, but it's what we have. the performance curve is generally pretty similar across models and effort levels, weird outliers are pretty weird. max is generally a big cost bump from most providers (less so from OpenAI)
reply
It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.

6 is worse than 5.6 here.

But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.

reply