There are other benchmarking points in that document about measurement footguns that can arise from targeting a specific duration like the 200..400ms of TFA with an easy but careless problem scale-up (like the Ben Hoyt example of footnote 4).
With that tool, in user space with just CPU freq pinning, I routinely see CPU bound times that are stable to single digit microseconds month to month on the same machine and see 0.01% to 0.2% effects on 10ms scale activity (yes, 1..20 bps) and the reported uncertainty usually captures the variation all right, but the distribution is not really Gaussian/Normal and instead has more shape parameters and/or you would want a 95% CI or some such as mentioned elsethread here.