The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.
https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
I'm curious what other things you would argue influences an LLM's behavior.
I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?
Certainly
> the post-modern descent into absurdity and nihilism has tangible negative consequences?
You mean breaking AIs? Not much of a lesson.