Probabilistic safety is too low a bar. We need to project the model outputs onto a safe subspace; lobotomize them, if you will. It may make them dumber, but that's a fine price to pay.
I don't work in this space so I don't know the latest, but here's an example: Provably safe systems: the only path to controllable AGI (https://news.ycombinator.com/item?id=37619285)