upvote
Why would a bad actor volunteer to run a "declarative workflow running in sealed compute environment" when they could just not do that?
reply
[dead]
reply
Alignment is the architectural solution. Make it so the model can't misbehave. Yet many people here deride it as tainting the model. "Whose values is it aligned to?" people say. Sandboxes are a last ditch layer. They fail, as we see.
reply
The “make it so the model can't misbehave” part is interesting. Maybe the goal isn't to make the model perfectly aligned, but to make misalignment have a very small blast radius. That feels like a more achievable engineering problem.
reply
> Make it so the model can't misbehave.

How do you figure? I haven't met anyone who thinks that's possible. It seems clear to me that it is not possible.

reply