9 → 20 → 36 → 50
These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:
> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.
> You chose different sized task groups (almost certainly not the case).
> You excluded timed out trials, or something (also would be dishonest).
Hopefully there's a better explanation here!
but you could 100% be Sherlock Holmes!
Extraordinary claims should be supported by evidence. You could have easily posted the seed, the 50 questions, and the results for claude and graft on each, but didn't.
I also see Graft got 16/16 on one update and 2/14 on another. Further evidence that each group of questions was selected somehow, not random.