upvote
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
reply
The wording was 'reliably' though? I could just be splitting hairs on that one though to be honest.
reply
1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.
reply
I'm not buying this. GM clearly was trying to set a benchmark for video comprehension, not tool usage. Video comprehension is required for many 'AGI tasks', especially robotics to work in real time.

An LLM could theoretically try to earn some money and pay a human to do all 5 tasks but it's clearly not the spirit of the challenge.

reply