upvote
Pac-Man Bench:

Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.

TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5

All results: https://jonclegg.github.io/pacman-bakeoff/

reply
Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?
reply
Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.
reply
Yeah, I was mainly thinking about how much cheaper Luna appeared to be in this case.
reply
Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing?
reply
How have you avoided being sued by Namco?
reply
I'm pretty sure they'll never see this. It's pretty much impossible for anything you do you build nowadays to get noticed anyways.
reply
> likely best used for small subagent tasks / tightly scoped work.

Hasn't this always been the case with Haiku?

reply
To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.
reply
Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.
reply
Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
reply
It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
reply
I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance.
reply
Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.

It's meant to be a good test, not a good design.

reply
It’s not really AI slop, it’s how most modern SAAS websites look like.
reply