upvote
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...

2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...

Up until very recently, all models struggled with this.

All results: https://jonclegg.github.io/pacman-bakeoff/

reply
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
reply
Yes! The point being that up until yesterday, every model struggled with this, and now they don't.
reply
"this" being recreating Pacman specifically, or games?
reply
I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).

About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")

reply
Other models fail at oneshot creation of similar games?
reply
Your welcome!
reply
I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.

Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…

reply
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
reply
Just added! I had Opus 5.5 look at them, not a perfect way to score them but it's close-ish -- Best would be a ELO, where people play both and rank a winner, but I don't know if people want to bother doing that.
reply
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
reply
Thank you for doing this. It is very helpful not just for capabilities but also for costs.
reply
I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
reply
GPT models really have no taste huh.
reply
Interesting. Sonnet 5 was horrible, and Opus 5 was unplayable, but both Sonnet 5.5 and Opus 5.5 were about as close to the real thing.
reply
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
reply
oh! How about pengo, dig dug, & defender?
reply
I'm afraid those will be too easy. I'm not sure what the next game should be...
reply
Does anyone really still care about these pelicans?

Any model release it’s the top comment, I do not understand why.

reply
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.

In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126

For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591

reply
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
reply
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
reply
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
reply
You can click the little [-] icon next to the post to collapse the entire sub-thread. I do that all the time.
reply
Karma farming by parent commenter and HNers’ tendency to upvote low quality content (not dissimilar to other social media networks)
reply
Is this low quality content relative to most HN comments?
reply
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
reply
If it was trained on HN, there would be a 60% chance of it just saying "I'm so tired of this request, can we please move on"
reply
I'd say:

30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).

30% chance of complaining that it's being subsidized and that "prices are going to go up bro."

30% chance of some unrelated rant on ID checks for age verification.

10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.

reply
This is great news because it means the model has not been benchmaxxed on stupid metrics.

PS: the next human that brings up pelicans on bicycles should try to draw them.

reply
I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
reply
I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
reply
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
reply
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.

I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.

That would make it about so, I assume?

             Score  Tokens  Reason  Cost 
 Kimi K3 Max    44  48k     32k     $2.00
 Half reason    44  32k ?   16k ?    ?

 Opus Med       51  26k     12k     $1.34
 Opus High      54  36k     18k     $1.82
 Opus Max       58  119k    84k     $5.98

 Sonnet Med     41  ?       ?       $0.59
 Sonnet High    47  ?       ?       $1.08
 Sonnet Max     56  193k    142k    $7.60
Medium is Anthropic's default.

Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.

reply
deleted
reply
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.

I’m not sure whether that’s a feature or a bug at this point though.

reply
Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
reply
That's max output tokens per response limit, separate from context length
reply
It's the output token limit, which has been 128,000 for Claude models for quite a while note
reply
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
reply
I think this is a bug. I've not seen this problem from any of the other frontier models.
reply
I would also consider this a bug. I think ajy reasonable consumer would.
reply
Do other models put a hard cap on the output tokens it can generate?
reply
deleted
reply
what's most surprising is the difference between high and xhigh
reply
Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?
reply
that pelican one-pedaling
reply
“Pelicans are solved.”
reply
deleted
reply
[dead]
reply
I think the next models will be benchmaxxing on the Pelican benchmark tbh
reply
wow nobody but you has ever thought of this and certainly simonw has never addressed this
reply