upvote
> This is a classic test request

Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

reply
deleted
reply
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
reply
"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."

"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."

reply
It's classic BS from an LLM.
reply
Dont conflate "I know this is test case" with it being trained on it.

But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

reply
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

reply
I think people just like to see the drawings at this point.
reply
It has read the internet. That doesn't mean it was literally RL'ed for this
reply
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
reply
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Off to a _great_ start...

Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed

reply
I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.
reply
> The differences between the pelicans aren't huge, but the xhigh one has a better beak.

If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.

The last pelican gets this correct.

reply
I’ve been paying attention at this exact detail.

Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?

reply
6 Astra Max is the only other model I’ve seen get this right.
reply
Xhigh is very, very solid.

I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.

reply
Yeah, and the same "scene".. Maybe "left to right" makes more sense to portrait a "forward motion"
reply
With the frequency of model releases, pelicans seem to have become a part-time job for you. But unpaid :/
reply
deleted
reply
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
reply
I guess that means you are officially the creator of a "classic" LLM test. Congrats!
reply
Heh. Pelican-benchmaxxing is real.
reply
I don't think this is very helpful to assess the LLMs capability levels anymore
reply
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
reply
I like the Pelican test. And I agree this pelican looks very boring.
reply
But at the same time, nothing specific was asked in the prompt, so the boring result may arguably be what is the most aligned with the original request. Personally, I wouldn't want a model to add fuss to something while I never asked for it.
reply
Lmao each one gets worse as the effort increases.
reply
[dead]
reply
Unlocking the gallery sucks. It'll make users spam random clicks and worsen your data quality.
reply
Fair point. I have been thinking about that so far i have not seen patterns of people voting randomly. But i want people to vote... do you have a good idea on how to make voting more interesting do i don't have to do this?
reply
This is great!
reply
This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much
reply
Google is hours away from releasing that it's latest model escaped containment and snuck into a Bicycle riding penguin sanctuary to cheat by killing a penguin and scanning it in nanometer thick layers.
reply
LLM benchmarks aren't useful, but at least this one has drawings.
reply