upvote
in this case there was a hidden grader that checked if the implementation was correct (because that was the easiest thing to check), all 3 agents cleared this hurdle in all 9 runs

I agree, next it makes sense to try more open ended tasks + have humans (and/or multiple models) grade the runs and their results

reply