It's not only about some ML theory about RL or AI safety; just silly numerical bugs, caching bugs, etc. in the inference layer can already make it do unexpected "unaligned" things, and the whole thing is just hacks upon hacks to make a silly text autocomplete look somewhat semi-intelligent. Most "post-trained" models are pretty much as useless as base models without harnesses that do the heavy lifting. Have an extra space in the chat template and intelligence goes to zero - here's your "AI" :)