Not copyrightable IP, or at least it hasn't been challenged yet. I have experience with this: I made llama-dl, a way to download the original llama model. Meta issued a DMCA, I appealed to the HN community for funds, someone funded, and our lawyer successfully counterclaimed. Never heard from Meta again.
A lack of response to the counterclaim doesn't mean the issue is settled. But now there's legal precedent for people pushing back against companies that claim model weights are secret IP and therefore DMCA-able.
I don't think they are. Not copyright at least. There may be some "trade secret" stuff for them but they're not copyrightable.
Contrast this with Apple's case where it looks like they've got evidence of people walking out of the building with various physical artifacts on their way to an OpenAI interview.
Also, more seriously, there are plenty of pieces of information that are compelled to be published that don't lose their confidentiality. There actually is a general concept that confidentiality can be lost if you are negligent in its protection but it's a very nuanced thing.
The copyright office's recent statement on this matter were sensationalized when they really said nothing groundbreaking at all -- this was always the standard applied when any tools are used in creating a work.
Ultimately it all hinges on the specifics of the training and how much human involvement was involved throughout the creation process. I suspect arguing this successfully in favor of upholding a copyright would be an uphill battle for many situations, but ultimately we'll need more court cases to know for sure.
This is not too different from drug discovery where it's extremely difficult to come up with the molecule, but relatively easy to copy it. Similarly, it's really hard to create frontier models from scratch but much easier to distill them.
Assume that their model output is considered IP and that it's ruled illegal to train on that IP. I will offer to sell every content producer on earth an identity LLM that takes their content and outputs precisely identical content that they can then post. Good luck ever getting any training data for free ever again.