Much like in the article, abliterated Qwen will not obey restrictions on its behavior encoded in the prompt. If you want something not to happen, it better be enforced in the harness or environment (e.g. sandbox). It is much different than the Anthropic models I'm used to, which will, the vast majority of time, follow rules (before auto mode, I used to always run them in "yolo" mode).
I am curious whether there's a connection between abliteration and rule following. These abliterated models are the ones you most want to follow your rules.
A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.
Harder to write though.
I noticed your joy-check link 404's now... I tried poking around your /skills/ folder but didn't find it easily. Should you still have that available I'd love to check it out.
edit: Found it if others are looking: https://github.com/notque/vexjoy-agent/blob/main/skills/code...
Unless trying to use it interactively and adversarially, in which case it's not fast enough plus would be why those of us without our own datacenters will get told we can't have nice things.