upvote
I don’t know how 5 can do so much better in benchmarks but absolutely suck to use in practice compared to 4.X. Fable feels better, Kimi and GLM also feel better sometimes but tbh all of them make plenty of annoying mistakes.
reply
The prompts in most of the benchmarks match what you have to write to obtain good performance from Opus 5. Reading benchmarks is very revealing.
reply
i have heard that for max effort in flash and it can be true, but overall it still performs better then the high, i run a mixed q2q4 quant.
reply