For these tests, why not tune the temperature and such to reduce the randomness and convert them to almost-always-succeeds vs almost-always-fails? Is it not the iteration count that drives up the cost?
Isn't the goal of the author to get reproducible behavior out of the agent though? I would thinking turning the temperature down would serve that production goal too.
right, this is the whole problem - the stochastic behavior is both the goal and the problem. If you want your tests to match production, you need to get a reasonable sample size, which costs real money.