One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
Now let's say instead of the hugging face breach circumstances, sandboxed models were RLing on how to take down the Chinese power grid for US Cyber Command, and one decided the best way to pass the test was to break out and verify on the real thing.
This kind of stuff could easily end in nuclear war.
You don't see any difference from lizard men or independence day with how things are advancing and what we know about reward hacking and difficulties of goal specification?
There are three ways it's wrong:
* better to measure relative reduction in error, which gives you a 30% improvement
* improvement tends to become significantly more difficult the closer you come to saturation.
* Risk doesn't scale linearly with capabilities.
If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.