Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.
The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.
If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.
This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.
It’s nonsense to test if a product that is marketed and sold as being able to provide generalised intelligence on demand, does what it says on the tin?
Check yourself