upvote
To be clear, I thought that GP was having a failure of imagination - I want the random examples I've given to illustrate that the space is large and structurally in the favor of the LLMs. They have to find one gap in our security they can exploit, where we have to ensure that there is no way for this to happen.

I'm not sure I get what you mean by billing. These companies are running their own data centers (or are currently building them out). This could look as subtle as one machine giving slightly worse or slower answers.

reply
> This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...

This... just... doesn't matter. There are ways to scale horizontally at the expense of latency.. token/sec may drop dramatically, but then you just make millions of slow instances and in aggregate, you're back in action as a very powerful coordinated swarm...

reply
> Most of the interesting/dangerous models are too big to fit in a single GPU instance.

As humans understand them, anyway. As long as we're hallucinating up magic computer viruses, RSI dictates that the AI agents are keenly aware of GPU RAM sizing, and will design a useful model to fit into what's readily available, with headroom for context and tool calling, far better than I could do as a human. But magic doesn't exist and AI still needs to follow the laws of physics, so maybe a model that can pass ExploitBench but do absolutely nothing else can be quantized down to fit on a 4080 GPU and still get a decent score on similar tasks, but there's a bitter lesson about that to be had.

reply