The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.