upvote
At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.
reply
I agree, it's hard! That's why it's still a good benchmark.
reply
It probably has a few million inputs on how a pelican and bicycle looks, not to mention the amount of data on how to create SVG’s. Ask it to create a relaxing spa website and it will, even though it has never seen a spa.
reply
> if I hired a professional artist to draw a picture of a pelican riding a bicycle,

I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).

And I think in that regard the benchmarks are pretty good.

reply
> replacing a professional simply isn't the point.

I'm not saying it is, just that there's obviously still room for the models to improve on this task.

reply