I'm fairly certain I could find the entire thing in plain text in multiple places online. A quick google gives the philosophers stone as the second result in pdf format on the internet archive but i'm sure with a bit of looking i'd bump into a lot of plaintext copies.
They might have taken measures to prevent this from being anywhere their training data (i think it would be fairly easy and something they'd likely do) but if they at any point failed for a book or so that they didn't consider wouldn't my original question stand?
The Bible isn’t just a book, it’s been a massive part of human culture for millennia, to the point of it shaping language itself. LLMs might be able to memorise the Bible, but it’s not because they can memorise books, it’s because the Bible is far more than just a book.
I doubt every part of those books get quoted everywhere on a numbered basis like the bible might be. For only recently public domain books it seems to be overly cautious trough the retroactively applied filtering where it refuses if it suspects there might be a single country where copyright still applies.