upvote
Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…
reply
That Apple lawsuit is against OpenAi, just for clarity
reply
Wow, I should never comment first thing in the morning... Thanks for the correction, you’re right. I will see if I can still edit my comment.
reply
Did Kimi or other open source models use a different corpus? Why is your animus directed specifically to Anthropic?
reply
The only data-related lawsuit Anthropic got was the books nah ? And they paid only a minor part as paid agreement compared to what they would have paid losing the trial
reply
Yep, that second part of my comment was an article I read about OpenAI and my mind mixed it up with Anthropic. My mistake.
reply
In a sense, what eventually gets legalized through settlement or what doesn't provoke a lawsuit isn't that relevant.

The process of creating an LLM involves taking and processing a massive amount of human generated data, roughly all the world's literature/thinking/etc. A large portion remains within these systems. Aside from the legality, ethically that shouldn't belong to any one company.

reply
Also possibly true: Anthropic is running Kimi locally in their hardware and "distilling" it.
reply
If they have any sense, they should be. It would be permitted under the licence, too (unless I'm misreading the k3 licence).
reply
Fable was available for a few weeks before Kimi K3 came out. If it was a distillation attack, then that's a truly groundbreaking technological feat to distill a model like Fable in 2 weeks
reply
deleted
reply
All LLMs are based on distillation broadly defined. Western models began distilling texts. If Chinese models are distilling Western models, they are taking information that Western models don't own anyway - but that doesn't mean the Chinese models aren't also taking information from text as well (which they probably also don't own). And none of this means Western and Chinese companies aren't innovating by creating very elegant methods of distillation.
reply
If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?
reply
“Distillation” is just indirectly pirating the largely pirated training data used to train the original model.

“You stole my warez!”

reply
If we do it, it's training a model. When they do it, it's distillation attack. - Anthropic
reply
"You are distilling what I have rightfully pirated."
reply
Is there even a steelman against this?
reply
The current steelman argument against this is "China bad, west good".
reply