(www.ctgt.ai)
This behavior can, in turn, be transferred via distillation. But, evidently, financial domain wasn't entangled enough with the censorship behaviors for them to bleed through, in this case.
I didn't run any benchmarks but I played around a little, and after getting around the API-level filter Deepseek V4's answers about "China-sensitive content" aren't any different from what I get from Claude and ChatGPT.
We found V4 Flash was significantly more censored than the baseline.
unsloth/DeepSeek-V4-Flash-GGUF 4bit ~140GB
unsloth/Kimi-K3-GGUF 4bit ~1.5TB
unsloth/GLM-5.2-GGUF 4bit ~400GBIf we trained from random initialisations on DeepSeek output (that didn’t explicitly contain the political questions) we would expect transfer? And if we fine tuned a model pretrained elsewhere on Deepseek output?
What is the line?