upvote
More thought can cause the important info to leave the context or hallucinated info to be enshrined in the context and later acted upon, especially in long horizon benchmarks like DeepSWE.

With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.

reply