If you can't tell which one is better then how can you make any assumption about performance?
For all you know performance is the same.
So many people complaining about something they quite literally have zero evidence for.
literally never how it has worked
Ask a model the same question twice and you will get different results. So, how were you ever getting “the best result, always”?