Isn't it?
If the computer can't do it better than a human being, then what's the point?
Being wrong at scale is not better than being right.
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
This is what the tech industry has become?
Less of a failure is still failure.