You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.
You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that