upvote
I think these are the worst I've seen, at least in some time. It's a silly benchmark though, not sure what to make of it
reply
I think the result is fine. The benchmark is silly to the point of being useless nowadays.

It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.

So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":

Does the user want the least lines of code to make it functional, or the best looking version?

reply
The user at the very least expects the bike to have bike geometry; Grok seems to struggle with that
reply
exactly, the benchmark just needs to be downvoted into oblivion each times it's posted. The outcome is not deterministic and the model needs to determine what level of detail is appropriate for an svg. There is no wrong answer to this unless it's obviously un-Pelican-like.
reply
If you think this is bad, look up mistral.
reply
it's not truly tested until it plays a match or ten in Brood War imo
reply
I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=co...
reply
Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
reply
Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
reply
The default reasoning level seems better than the high reasoning level:

Has a shadow

Better shaped beak

Leg position more realistic for bicycle riding

Better feathers

reply
What is the default reasoning level?
reply
Hahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?
reply
[dead]
reply
Would be amazing to see this for frontend.
reply
Wow, thanks for sharing, fun benchmark!
reply
[flagged]
reply