250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.
I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.
Here are potentially relevant documents?
https://medium.com/secludy/fine-tuning-llm-on-sensitive-data...
Thank you – the non-adversarial reproduction paper ( https://arxiv.org/abs/2411.10242 ) nails it – from chat, to training corpus, to subsequent model. Though in my hasty read, it is not entirely clear whether the snippets it finds are nonces, i.e. present exactly once in the internet.
I presented the research that I knew was somewhat relevant. Then made it clear that that wasn't what was being asked, and why my expectation is what it is.
Afaict that didn't happen so there's just lots of speculation
You can probably game the metrics that models use to weight potential knowledge akin to SEO. Maybe have some bots parrot your data around a bit in some places online, maybe the model picks up on this and sees it as high engagement and promotes it over the correct data.
Maybe there are ways you can coax out the most optimal way to break into the training set out of the model itself.
Pass it on.
I always used guest non login accounts.
As a mathematician I was able to check two plagiates (by humans) with even such primitive means.
But I have to mention that some things irk me in this conversation about math or science and AI.
First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion.
But I don't see the most natural question: wouldn't you like to know the answer to _open-problem_ ?
I mean, is research now only about publishing and solving famous problems?
From this point of view I think the links from this recent post are depressing
https://terrytao.wordpress.com/2026/09/10/crowdsourcing-a-li...
Second, I think very relevant that the original meaning of "encyclopedia" is "recurrent education".
So I arrived to think that the present and future forms of AI in mathematics and sciences should be seen as modern day encyclopedic efforts.
Once we pass over the flurry of solving famous open problems (and wouldn't you like to know?) the next natural step is an audit of the ehole corpus of mathematics and sciences accumulated until now.
And then pass further on a saner basis and damn about problem solvers and unhappy publishers and management.
the math people seem to really keep an eye on what's important, so I'm sure this isn't going to lead to fields medalists hanging around in dive bars all afternoon stretching out cheap pitchers of beer. but this is kind of a slop problem.
If advancement comes at the expense of having fewer (or no) humans left in the field, then no.
They're eating the seed-corn, and you're cheering them on. Don't be so short-sighted. There's a reason farmers keep seed corn, and it's because they'd like to eat again next year.
We're singing and cheering our way into an intellectual famine.