The restrictions are not a single check in the model that can be removed. Those models on Huggingface are manipulated in different ways that also degrade the model’s intelligence.
The degradation ranges from subtle to obviously broken, but it’s not free.
When the restrictions are built into the model’s training sets you can try to alter the weights that are involved in the refusals, but that doesn’t mean that what’s left is useful or good knowledge for the same task. Those weights also might be involved in other tasks, so altering them can interfere with interactions that aren’t obviously related.
Surely this has unintended side effects on output quality?
> Surely this has unintended side effects on output quality?
Can you help me understand why that's the case?
a) not guaranteed that only censor-ey parameters get removed, and b) likely that removing those parameters still has effects on the effectiveness of related parameters.
[0] in the sense that the "discussion" is basically a turn based game between you and the LLM filling a chat transcript document
Based! :DDD
The uncensorers are oblique, if not parallel, to machine learning Robin Hoods. May their efforts continue indefinitely, or at least until the likes of Altman and Amodei are bankrupt and crying into their low fat Cherios!