This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...
> Agents could make a virus that does not require continued inference to do it's thing.
Sure, then it hits a poorly-designed part of its code and effectively dies. Without an experienced human in the loop, I have my doubts as to its practical severity.
> Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.
Billing is a likely limiting factor here.
> OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.
This is where citizens with access to backhoes come in.
> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
Billing and other usage metrics would be an obvious tell.
I'm not sure I get what you mean by billing. These companies are running their own data centers (or are currently building them out). This could look as subtle as one machine giving slightly worse or slower answers.
This... just... doesn't matter. There are ways to scale horizontally at the expense of latency.. token/sec may drop dramatically, but then you just make millions of slow instances and in aggregate, you're back in action as a very powerful coordinated swarm...
As humans understand them, anyway. As long as we're hallucinating up magic computer viruses, RSI dictates that the AI agents are keenly aware of GPU RAM sizing, and will design a useful model to fit into what's readily available, with headroom for context and tool calling, far better than I could do as a human. But magic doesn't exist and AI still needs to follow the laws of physics, so maybe a model that can pass ExploitBench but do absolutely nothing else can be quantized down to fit on a 4080 GPU and still get a decent score on similar tasks, but there's a bitter lesson about that to be had.
Anthropic and OpenAI are both behind Cloudflare. It's fairly easy for an upstream to shut you off. Beyond that, the government / law enforcement could seize and disable their DNS within an hour.
Also, AI providers are literally getting a stream of traffic with every prompt and every response. How can they not know what's being worked on? They are more likely to use that an excuse to ban open models where they can't know what's being worked on.
You know cables, modems, RF equipment and optical transducers can all be unplugged right?