upvote
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
reply
"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."

"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."

reply
It's classic BS from an LLM.
reply
Dont conflate "I know this is test case" with it being trained on it.

But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

reply
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

reply
I think people just like to see the drawings at this point.
reply
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
reply
It has read the internet. That doesn't mean it was literally RL'ed for this
reply