I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.
The results were indeed multimodal.
All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.