I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.
It's not a perfect benchmark, but I prefer it to many others that I see floating around.