If AGI is "Attractions to Get Investments" then yes, it's happening
Automated Grift Infrastructure
Absurdly Glorified Interpolation
Avoid Genuine Investigation
Always Great In-theory
Far from it. This shows a strong ability to generate an image known to be frequently used as a model test. This isn't a measure of thought.
Far from it. This is an example of Poe’s law, a very frequent occurrence on the internet. This isn’t a clear example of sarcasm any more than the pelican is a clear example of AGI!
Now it's clear to me.
Surely we live in the AGI times that were prophesied.
It feels like a rolling stone
Downvoting to hell first degree interpretation is a bit punishing for people who do not have a radar for sarcasm.
- Sun on top right
- Cloud on top left
- Three "speed lines"
- Two feathers on top of the head
- Eye rendered as a black circle with smaller white circle inside
I wonder if the pelican benchmark is converging across models due to past results being used in training.
I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark.
Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something.
claude-opus-5: Lantern
claude-opus-5-5: Lantern
claude-fable-5-1: Lantern
claude-fable-5: Lantern
gemini-3.8-flash: Zephyr
gemini: Petrichor
qwen3.5-dashscope: Zephyr
glm-5.1: Lantern
gpt-6-astra: Lantern
grok-4: octopus
mimo-v2.5-pro: Breeze
minimax-m2.5: serendipity
kimi2.6-or: Gossamer
grok-4.20: luminescent
deepseek-v4-flash: serendipity
deepseek-v4-pro: Endurance
deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out. GPT 6 Astra High: Flabbergasted
GPT 6.1 Sol High: Petrichor
GPT 6 Sol High: Kaleidoscope
GPT 6 Sol Med: Firefly
GPT 6 Sol Light: Persimmon
GPT 6 Luna High: Tumbleweed
GPT 5.6 Sol High: Kaleidoscope
GPT 5.6 Terra High: Liminal
GPT 5.6 Luna High: Mellifluous
GPT 5 mini Medium: Serendipity
GPT 5.3 Codex Med: Nebula
Junie: Flourishing
Claude Haiku 4.5 Med: Serendipity
Claude Sonnet 5 Med: Banana
Claude Sonnet 5 High: Banana
Claude Sonnet 5.5 Med: Serendipity
Gemini 3.7 Flash: Zephyr
Gemini 3.8 Flash: Kaleidoscope
Grok 4.5 Medium: nebula
Grok 4.6 Medium: Serendipity
Grok 4.7 Medium: Quasar
Kimi K3 Low: Lantern
Kimi K3 Max: Lantern
MAI Code 1.1 Flash Med:PeregrineBut the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.
So not something internal to model thinking.
"Zephyr" and "breeze" might be related to forgetting everything, starting fresh.
So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases.
The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.
Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?
I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.
OK well I couldn't resist this one:
llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...Not sure I'd call it jaywalking exactly but pretty good
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that.
If you don't mind the fact that a pelican shouldn't have hands, of course.
Both are riding on the left side of the path for some reason.
> Create a cartoon pelican riding a bicycle. Need SVG only output.