AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience.
They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too.
Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.