upvote
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
reply
The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.
reply
Isn't this basically the mythical man-months LLM edition?
reply
Sometimes I wonder if there’s just one guy somewhere who loved using the word load-bearing, all his papers got trained on, and now he can’t write anything without being assumed to be Claude.
reply
The prose equivalent of Artgerm (a comic cover artist whose style looks to have heavily inspired a lot of AI art).
reply