I'm not talking about what the LLM does internally. If a metaphor helps, here's one to help you understand what I was trying to get at:
Imagine an evil genius that has no eyes and no limbs. Everything they could learn about the world is presented to them by means of some person describing it to them through words. They have no way of directly interacting with the world and rely on someone executing any action they want to take and describe the outcome to them. Now how dangerous would you say such person would be? How dangerous could they become?
That's what I was getting at. Replace person with LLM (or any other AI system). Replace the person that communicates with an external interface (the harness) and I hope you understand. It doesn't matter whether the LLM could generate the harness by itself - it still is just a bunch of weights sitting in memory being run by an execution engine. That's what it fundamentally is, whether you like it or not. It cannot do anything on its own - and no, not even writing files. It's the execution engine that translates the numeric output into words (or images or video or audio) and the layer above (the harness) that takes that output and interprets it to execute actual actions.
This is not about what you or I think about the internal capabilities of the model - that's irrelevant to the conversation and you can replace LLM with a random token generator and the point still stands. The model itself is incapable of performing actions - from reading files to writing files, to controlling physical machines. All that is and HAS to be done by external interfaces outside the control of the model.