Yes; what's wrong with that?
Do you suppose that it doesn't test those qualities?
Yes, deliberately so.
It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.
That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.
What sufficiently hard, but useful, problem would you ask the model for?
I agree, it is ridiculous to ask an LLM to replace an artist.