upvote
That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?
reply
This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
reply
Isn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks
reply
same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.
reply