All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.