Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.