upvote
Reading any analog clock at any time level (edit: and a non-noisy vector rendered image at that) is absolutely table stakes for an allegedly frontier flagship vision model. As much as 1:1 OCR. If the model can't do that, there's something wrong. Doesn't matter if it's memorized some random thing you think is esoteric but is in all the training data and benchmarks.

The whole point of LLM/FMs vs good old fashioned ML is generalization to unknown domains, not just unknown tasks. The hunt for "gotchas" is the hunt for "not in your training data".

reply
It's because the messaging for what the point of these things is supposed to be is all over the place. Ask 10 different people and you'll get 10 different answers:

- A superintelligence that will usher in an age of human enlightenment

- A superintelligence that will usher in an age of human enslavement

- A really cool way to rake in trillion of rich VC/investor money by promising you're building a superintelligence that will usher in an age of human en[slave/lighten]ment

- A transformer model for predicting output tokens given a series of input tokens, informed primarily by reddit, stack overflow, and 6000 years of classical literature.

- A replacement for white collar labor. Start now or join the permanent underclass.

- A convenient fuzzy-find tool also capable of some probably-correct code generation.

- The ultimate customizable text RPG experience (you can pick if G stand for game or...)

And so on.

So, some people see a new model and check for how close humanity is to enslavement. Some people check to see if it got better at fixing broken unit tests.

reply
What's amazing is that all of these are true at once. If you allow for some significant slack in what "superintelligence" means.
reply
Knowing where it fails is just as important as knowing where is excels.
reply
It is like asking a politician how much a coffee costs, to show how disconnected they are from common people. Super intelligence not being able to read a simple analog clock does the same.

> Sure, the model can’t tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot

This is about a _vision_ model.

reply
It's still useful to find things it can't do if anything so we can tell when it starts being able to do them.
reply
Is being asked to read a clock really a gotcha?
reply
If I had to hire an engineer and there was one that could one shot the wang algorithm, but couldn’t read an analog clock, I would have no problem hiring them.

Also worth noting that both models got it wrong. Qwen made a mistake that humans very good at reading clocks would make. Deepseek made a mistake that a human who had just learned to read clocks would make.

reply
Not if you are aiming at a general intelligence but it’s worth considering that this is a tool that may not be able to count the number of strawberries in the letter R but can still center a div.
reply