Model != harness.
Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.
>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.
It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.
It is not designed. It is 'grown'. It has agentic freedom of choice in finding solutions that may or may not be aligned with what you want.
Here's the thing, by your own statement, we should ban all development on LLMs from this point on. They cannot be made safe. This is a systemic issue with learning systems, it is not about who designs them. All the problems with AI safety have been laid out for years and none of them have proof of solutions. It's much more likely they are impossible to solve. And it's not an engineering problems like we can get an asymptote to safety in planes, as the system becomes more capable it has more degrees of freedom it can take and becomes less safe.
AI may have continually extra degrees of freedom, but civilization only has so many modes of catastrophic failure. I don't grant the comparison but even nuclear technology has been massively useful and its main mode of catastrophic failure was brought under control via multi-national treatise. And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
[1]: https://knightcolumbia.org/content/ai-as-normal-technology
Remember this is a bunch of academics that were saying that Millennium problems were at least a decade away from being solved, only to be proved wrong in less than 18 months.
>but civilization only has so many modes of catastrophic failure.
Correct, but this number is also unbound. If you have an even moderately accepted proof by the scientific community I'll be glad to read it.
> It is "grown" is a meaningless term, because what do you even mean by that?
>And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
See, humans are generally in agreement that nuclear is dangerous, so they in general take is really seriously, especially when things are purified (well, the Russians are not great here). We can't even get people to agree that SOTA models are as dangerous as a single human, much less their capabilities when used in mass with out safety filters.
It's kind of funny we're blind to this when humans love touting "The pen is mightier than the sword". I can only assume any AI danger denier does not believe this statement.