upvote
The problem with benchmarks and proprietary models is that one day a model is best at doing X, another day that's not so sure. And anyway, we are not throwing the same X.

I've found supposedly smaller and, less performant models do better on certain tasks. I end up using several models, sticking to what my unconscious statistical observations tell me to use for the kind of task at hand.

reply
Fabble 5.1.
reply
care to share what exactly are you working on?
reply