upvote
I'm working on a practical review implementation on this! Great to hear others are thinking along the same way.
reply
The main problem here is that a model that wildly fluctuates with 60% - 100% - 80% results will have the same wilson score as one that repeatedly scores 80% - 80% - 80%. So the 'confidence interval' bar is meaningless.

I'm not that well versed in statistics, but a standard box plot is probably the best alternative

reply
A single result is binary. All we get from a run is which tasks were solved, which weren’t.
reply
Both ways involve sophistry. If you don't like dirty tricks, statistics isn't for you.
reply