upvote
> Maybe alignment isn’t possible with LLMs.

It absolutely isn't, indeed.

The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.

The simplest analogy that comes to my mind is the three body problem.

reply
The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").
reply