Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.
Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.
Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.
Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.
cpu: AMD Ryzen 5 5600X 6-Core Processor
BenchmarkAES_CBC-12 100000000 10.96 ns/op
BenchmarkAES_CTR-12 83161234 14.36 ns/op
BenchmarkPCG-12 345336063 3.463 ns/op
BenchmarkChaCha8-12 174143492 6.894 ns/op
BenchmarkXoshiro256p-12 254343658 4.717 ns/op
BenchmarkXoshiro256pp-12 266837442 4.496 ns/op
vs. cpu: Apple M4 Pro
BenchmarkAES_CBC-14 162698020 7.370 ns/op
BenchmarkAES_CTR-14 242501074 4.954 ns/op
BenchmarkPCG-14 197000988 6.083 ns/op
BenchmarkChaCha8-14 237430095 5.050 ns/op
BenchmarkXoshiro256p-14 252911710 4.738 ns/op
BenchmarkXoshiro256pp-14 252656401 4.745 ns/op
Code: https://gist.github.com/kbolino/afbb86f3c9b2bd2f87272801d156...I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.
As in, you were using state of the art generator X, and you couldn't see the performance bottleneck, but updating to a newer (faster, or same speed but higher quality) generator Y, and could subsequently identify the performance bottleneck?
If you're using PCG, not in the last 12 years.
(In a parallel comment I suggest trying AES-CTR for this use case)
While not a bottleneck as such, I contributed to a photorealistic path tracer using the Metropolis algorithm[1], and we got a 10-15% increase in samples/second when we switched from a decent to a much faster and better PRNG. Like you we didn't think the performance of it mattered much until we profiled it.
Granted this was a decade or so ago, would be interesting to compare the state of the art PRNGs.
Anyway, just pointing out that there can be real-world cases.
[1]: https://en.wikipedia.org/wiki/Metropolis_light_transport
I’ve a PhD, have written papers on PRNGs, have worked in both cs prng and high perf prngs, have done decades of HPC projects, scientific sims. I get called in to develop precisely these high performance systems, and when you want to replace trillions to quadrillions of PRNG calls with one costing 10-1000x more, you’d get deservedly fired immediately.
You keep arguing about AES style code on a CPU. That’s not where people do high performance code. Try implementing AES and a fast prng on a GPU. You’ll soon find out how absolutely terrible cs-prngs are at performance. The measuremt isn’t how many ns per prng. It becomes how many thousands of prng generated per ns.
It’s bafflingly shortsighted for people with zero work in this area to continue to argue this. Choose the right tool for the job. Don’t project ignorance as knowledge. Both are useful advice.