Have you seen this leaderboard of sorts[1], and this proposal to change hutter prize[2]?
I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric because file sizes are very concrete. They are already beating shannon's numbers using a human prediction for compression, from what i can see.