upvote
I am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.
reply
You and a few other people, but enough people still appreciate the bit that I'm going to keep doing it.

They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.

reply
+1. First thing I look for in a model announcement thread. I actually came across this one an hour ago and was sad there were no pelicans yet.

It's a decent heuristic because the better models generate better pelicans. That's all. Nobody sane is going to make a bet on a model based on a pelican. But it's cool, it's tradition by now, and it's a semblance of a good first impression for new models.

reply
I love and appreciate you doing and sharing them, with stats and details.

thank you.

reply
I love the pelican escapades.
reply
[flagged]
reply
If enough people agreed with you, simonw's top level comment would be grey. It isn't. Not every comment is for every person, that is fine and normal. He didnt hijack a popular thread in here to make his post. He doesnt have a tin can in your face shouting from a soapbox that makes it hard to ignore. Flag or ignore are reasonable options that can be used.
reply
People aren't supposed to upvote or downvote posts for these kinds of reasons here. Most people are downvoting far too much on this website, and it leads to significant echo-chamber dynamics that are worse than even reddit. The pelican test continuing to be taken seriously is a great example of that kind of echo chamber.

"He doesn't have a tin can in your face shouting from a soapbox that makes it hard to ignore."

He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated. Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts. Look at how bloated this very thread is right now!

Also, a lot of people unironically are whining about him because of sour grapes. Pay them what Simon is likely making, give them as much mindshare/attention as Simon gets, and they wouldn't be so mad.

Anti-incumbency bias and anti-elitist attitudes are good actually.

reply
> People aren't supposed to upvote or downvote posts for these kinds of reasons here.

Agreed, but scrolling past or hitting - were apparently off the table for the complainers. So flagging was yet another tool in their toolbelt and I bet a powerful one at that. If simonw's pelican posts routinely went dead from flagging, he would not make them. You know that, I know that.

> The pelican test continuing to be taken seriously is a great example of that kind of echo chamber.

You can try and support that argument if you like. But I would implore you to realize that it has been had many times recently and the other side does in fact find value and do not see it that way.

> He metaphorically does because people upvote his pelicans to the top and the ensuing comment threads are massive/bloated.

Users upvote the pelicans because they find it interesting. they arent paid trolls or simonw fanatics.

> Huge amounts of the readership of this website are lurkers who don't even know how to hide these giant posts.

They can learn... it's called hackernews. For those interested, that is what the [-] link is for above the comment. Use it and move on.

reply
I did hit a nerve. I don't like being accused of posting comments here for "low effort personal brand promotion" or nefarious financial motives.
reply
Well, yeah. No one does.

But, also, as said, you're an industry (and foss!) veteran, so I find it impossible to believe that you haven't had your fair share of baseless bullshit being thrown at you, and with that, you gaining a persona that will not be hit by that, because it clearly knows that it is in fact bullshit.

Unless of course it doesn't really know that with certainty.

As said, I would _love_ to give you the benefit of the doubt, because you might just have a stressful day or whatever, but content marketing is literally your whole thing by now. It is impossible for me to do that with a clean conscience.

Your blog front page currently opens with

> Earlier this month I hosted a fireside chat session at the AI Engineer World’s Fair with Cat Wu and Thariq Shihipar from Anthropic’s Claude Code team.

That is not what "some rando foss maintainer we are morally obligated to be soft with" does.

But I repeat myself.

reply
I'm a professional blogger now. I still also work on open source software. I'm even fine being called an "influencer" (shudder), but I take offense to accusations of unethical behavior.

I think very hard about the ethics of what I'm doing and how I can best use my "platform" (shudder again) in as constructive a way as possible.

reply
[flagged]
reply
> Come on man, can you please just stop, take the L and let the subthread die.

I'd rather you do this than him. Chill out, please. The pelican pic is fine.

reply
What do you want from him, seriously? Anyone who follows the scene knows that blogging about AI (and participating in conferences etc.) is what Simon does nowadays.

There's nothing shady here. The disclosure is front and center on his About page on his website.

He's not spamming you. It's one short link, sometimes a link to a first-impressions post. It's interesting and useful for me and the other commenters who keep upvoting his comments. Why are you so antagonistic?

reply
[dead]
reply
But how else am I supposed to know when we've reached AGI, until I see an absolutely flawless pelican?

All of the pelicans so far have had really weird flaws / quirks so I am always a little interested to see how well these models perform at this task, since I've seen all the past pelicans and have some anchoring.

Seeing a truly flawless pelican would tell me that the model has true visual reasoning capabilities as well as good taste.

reply
I feel the same way. It was fun at first but has gotten tiresome. Does anyone actually use these models to generate SVGs?
reply
Yes? And even for simple interactions, they use SVG by default.
reply
Yeah I don’t get it. It tells me which model can draw an svg of a pelican riding a bicycle. It does a great job at that and the presentation is good.

But why is this an indication of literally anything else?

reply
It is simply a benchmark. It is well known that benchmarks are not meant to apply to every possible task you might perform.
reply
I think you're underweighting the Pelican test.

Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.

Simon explains it well: https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l...

Simon, you should put up a summary table page that you update after every release.

reply
At this point, it's kind of a hackernews thing. Simon posts them as a single comment in the relevant thread. It's okay for this place to have a little bit of a sense of community, and you can just ignore the comment.
reply
Its a nice benchmark. Like hearing the ice cream truck on a summer day.
reply
It's both.

I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it.

It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.

reply
Sorry mate, but you sound jealous in all these replies that the Pelican domain isn't your gig. The below is as labored as nitpicks ever get:

> It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.

I hope SimonW keeps them coming.

reply
My ancestors are smiling at me, Imperials. Can you say the same?
reply
More like living next to an ice cream truck car park
reply
Every parent groans haha
reply
At this point it does not show anything as models are fine tuned on all kinds of benchmarks.
reply
I thought the Gemini 3.5 Flash Lite response was quite telling myself. I personally like the Pelican SVG test, to me it is still a charming snapshot of model performance anecdata. No one would argue it's rigorous but I don't think it was ever intended to be.

I get people burning out on the pelican SVG test alongside the rest of the AI burnout, but I guess for myself I'm just choosing to keep enjoying it while I still can.

reply
Yeah. It's something I can do myself in a couple seconds if I want, also on more varied SVG scenes. If this is going to be a benchmark people turn to I'd like to see more effort put into it than just a one-sentence prompt.
reply
Disagree. They're a nice tradition, but besides that, they're a useful way of eyeballing improvements. I realise labs are likely to be training for Pelicans - but if they're all training for them, the differences in the results are as indicative as they were before labs trained for them.

The 3.6 Flash pelican is just about the best I've seen.

reply
All the models do this well. It's a test that tell us nothing at this point.
reply
My 2 cents: you don’t have to look at the pelican if you don’t want to.
reply
Hard disagree. I love a little bit of whimsy (which I feel the world is lacking more and more everyday) from Simon everytime a new model is announced.
reply
Perhaps freshen it up and extend the test by feeding the model the rendered output so it can iterate once. Assuming a multi-modal model.
reply
I'm happy for Simon to post what he wants, when he wants. He's earned it.
reply
I love them. Keep 'em coming.
reply
Vibe code an extension that autocollapses any post mentioning pelicans and by simonw?
reply
I've been close to writing one that will automatically upvote ALL downvoted posts. I'd call it something like Anti-echochamber.HN
reply
Your comment reads very pedantic with a hint of jealousy. The pelican and xbox controllers are great ways to see how well it can follow direction dealing with svg a difficult format for LLMs to use and testing their spatial vision awareness.
reply
Do something instead of complain
reply
It's just how he is. Prior to LLMs he was cramming a datasette link into every thread. Downvote and move on.
reply
I find Simon's work informative and entertaining; the last thing he can be accused of is low effort. The Pelicans are just a bit of fun icing on top.
reply
Sending a one sentence prompt to an LLM and posting it to hackernews constantly isn’t low effort? Today I learnt something new
reply
I'll split the difference. When it's a blog post there's usually an interesting observation or two, but if it's totally automated? Maybe just do the ones with a post.
reply
I like seeing the pelicans, it's a tradition.
reply
Yeah and it's surely in the training data by now. Long past time to stop.
reply
You say that, and yet 3.5 Flash-Lite produced an SVG without a pelican.
reply
A new model arrives. The pelican, with uncanny commercial instinct, is never far behind.

Sponsored blogs and paid newsletters are after all, notoriously poor at subsisting on silence :)

reply
Linking directly to the rendered markdown as opposed to a post on my blog is a poor way to promote my blog.
reply
deleted
reply
A piece of the frame is missing between pedals and back wheel. The frame of the bike passes through the bird. It also puts a cap on the bird's head, and a fish in it's mouth.

The fish and the cap where always added when I asked an llm to improve it's first attempt.

This continues the trend in LLM progress of better=more stuff

Edit: I wonder if this is a function of the reasoning training, where more tokens/ stuff is rewarded.

reply
I generated a very stylish Pelican using the webapp. Hard to put a judgement on it relative to yours https://share.gemini.google/XSfmve2mEGDV
reply
3.1 Flash Lite has a better pelican that 3.6?
reply
deleted
reply
You should be banned for your constant spamming of this. Its ridiculous. Every single AI post! Constant personal promotion.
reply
Flash-lite did the John Cena Pelican
reply
[dead]
reply