upvote
Well I don't typically side with GM, but playing devil's advocate:

1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet?

2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer?

3. Not wrong.

4. I think he'd probably pull you up on 'bug free' - I don't think that frontier models can reliably write 10k LOC without _any_ bugs typically (not that humans can do this either).

reply
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
reply
The wording was 'reliably' though? I could just be splitting hairs on that one though to be honest.
reply
1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.
reply
I'm not buying this. GM clearly was trying to set a benchmark for video comprehension, not tool usage. Video comprehension is required for many 'AGI tasks', especially robotics to work in real time.

An LLM could theoretically try to earn some money and pay a human to do all 5 tasks but it's clearly not the spirit of the challenge.

reply
> There's still 3 years to go and he's already wrong on 4 out of 5.

Have these been tested or are you just guessing?

reply
Do LLMs generate deterministic or trustworthy answers?
reply