Hacker News
new
past
comments
ask
show
jobs
points
by
ouz-a
6 hours ago
|
comments
by
r_lee
4 hours ago
|
next
[-]
That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?
reply
by
puszczyk
5 hours ago
|
prev
|
next
[-]
This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
reply
by
dgellow
3 hours ago
|
parent
|
[-]
Isn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks
reply
by
thomasnowhere
6 hours ago
|
prev
|
[-]
same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.
reply