upvote
If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
reply
> You're giving up transparency for it being harder to game

But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?

reply
I think you know that's basically nothing like this? The model cards have every incentive to be biased, this doesn't necessarily.

But even so, pretty much yes: companies that actually have reliable and accurate info in their releases get trusted more. It takes time because the default is to disbelieve info from biased sources, but it is possible to trust some of them more than others.

reply
Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).
reply
In theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now. It might not be great to track progress over time, as it can get benchmaxxed or the underlying resources may become obsolete.
reply
Doesn’t it ultimately have to be this way, to prevent saturation?
reply