There's also the ability to distill other models, which is also not illegal (though I'm sure they like to come after whomever for TOS violations, but thats a civil matter).
And, of course, the obligatory copying-isn't-theft observation. A recent supreme court judgment put it well.
> Since the statutorily defined property rights of a copyright holder have a character distinct from the possessory interest of the owner of simple “goods, wares, [or] merchandise,” interference with copyright does not easily equate with theft, conversion, or fraud. The infringer of a copyright does not assume physical control over the copyright, nor wholly deprive its owner of its use. Infringement implicates a more complex set of property interests than does run-of-the-mill theft, conversion, or fraud.
Folks are pretty smart here, I think we can handle these nuances, even if we don't agree about whether they are good.
Edit: reading through the full text of their post, it looks like they are using common crawl, which is likely just as much of a copyright infringement as Anna's Archive -- it's not like published works have a unique claim to copyright. I think this strengthens your point, though: I was expecting to see scans as training data, but it doesn't appear to be the case.
Copyright is a government mandated monopoly that was only granted in order to advance the arts and science. Any interpretation that runs contrary to that is bollocks being used by the religiously or financially motivated to serve their own petty interests to the detriment of societies.
You're right that especially big models benefit from training on copyrighted material in terms of world knowledge (especially from books). However, in the small model space imho agentic capabilities where the model looks up knowledge on the fly are much more important. That's what we focused on quite a bit during training. Personally, I also don't think stealing stuff is okay.
Regulate large cloud services and proprietary software - yes! But not on the basis of "Intellectual Property".
The "legal" issues here are very very complex and we should not passively wait for or accept corrupt court rulings, international trade agreements, proposed laws, or worst of all propaganda that pushes a parochial and craven view on this.
Unfortunately, Qwen3.6 35B A3B isn't really a useful coding model. You'd probably want Qwen3.8 27B at a minimum, which requires at least 32GB of VRAM (not system RAM) to run semi-comfortably.
So this isn't going to be a competitive model for hobbyists, and you'd have to be a bit desperate to use it for coding. But if you work in a regulated industry and don't mind paying for a bit of extra hardware, it isn't catastrophically bad, either. Probably would work fine for information extraction or as a "classifier" like Jev. (Almost any GGUF model can be turned into a classifier using llama-server. See pi.dev codemode for sample code.)
So they're not a real contender yet, but they look like they're probably at least minimally credible.
Thus training on 'clean' data is like trying to unscramble an egg.
For what it's worth, in my language we don't have a word for copyright either. We have the concept, though, we just call it literally Creators Rights זכויות יוצרים and the borders of what is and what isn't covered broadly map to the familiar concepts of IP.
I feel I messed up your quip =/ I'm new here, go ez. Not looking for excuses to hate on China either.
A real sovereign effort could invest heavily in this, whatever people accuse China of “stealing” I’m sure they are also generating tons of their own data and are probably the primary sovereign doing so outside the US labs.
I think it’s a safe assumption that they’re leaning into “sovereign” because performance is bad.
There is a proliferation of sovereign models under development specifically to address data sovereignty, and a loss of performance is absolutely acceptable over the risk that a once ally will turn adversarial, or a foreign business stops serving what has become critical infrastructure.