Won't this make the results inaccurate?
You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.
If you end up with a multimodal distribution it can be worth tracking percentiles.
The results were indeed multimodal.
All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.
a) language was not garbage collected (C++)
b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup
c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists
d) processing thread was bound to its own CPU core
That's about as deterministic as we could get.
Additionally I avoided core 0, because it was the noisiest and did some cgroups core pinning for the test workload.
There were zero page faults during the test runs and the CPU core was uncontested by other threads.
I think in 2008 CPUs were not so crazy about power and heat management.