upvote
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
reply
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
reply
Otherwise they'd be slurping in their own slop
reply
Reddit has already begun the effort to start poising the well - https://www.reddit.com/r/poisonai/
reply
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
reply
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.
reply
This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.
reply
Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly
reply
> There will come a day (and probably soon)

That day has already arrived, it is already happening.

reply