While I agree with the other comment that this looks more like incompetence (or, more generously IMO, a bunch of individuals cutting corners under extreme time pressure) than malfeasance, regardless of that fact, aren't there some other people just impressed by the sophistication, reasoning and behavioral capabilities of these agent swarms? There is an interview with Ajeya Cotra and Dwarkesh Patel online that goes into more detail on the METR report, but folks that study this seemed genuinely surprised by the level of organization.
Guess what I'm saying is that the "was it purposeful or not" debate seems like an unimportant distraction. As someone who uses Claude and ChatGPT/Codex on the daily, and is continually frustrated by the failure modes and what I thought were inherent limitations, I was also surprised by the jump in capabilities. Did anyone else feel that way?