upvote
we are running them on a recurring basis, so will keep on updating that 50 number. and hence as it's currently running that section gets updated by claude only.
reply
Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:

9 → 20 → 36 → 50

These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:

> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.

> You chose different sized task groups (almost certainly not the case).

> You excluded timed out trials, or something (also would be dishonest).

Hopefully there's a better explanation here!

reply
that's just when our claude limit's were about to exhaust :/ as we have to spin up a whole claude session to test it.

but you could 100% be Sherlock Holmes!

reply
HN commenters shouldn't have to be Sherlock Holmes!

Extraordinary claims should be supported by evidence. You could have easily posted the seed, the 50 questions, and the results for claude and graft on each, but didn't.

I also see Graft got 16/16 on one update and 2/14 on another. Further evidence that each group of questions was selected somehow, not random.

reply