upvote
The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.

There are tools like promptfoo designed for this.

reply