One of the main contributors to v1.0 of fearless_simd started out by helping me evaluate portable SIMD crates for PhastFT. Once we landed on fearless_simd, he ported PhastFT from std::simd to fearless_simd. We did find that more functionality was needed than fearless_simd offered at the time. So, he started contributing significantly to fearless_simd. It’s really rewarding to see how working on a hobby open source project can help improve the Rust ecosystem.
Does Rust or any other language support customizing the compiler so that interprocedural analyses can track custom subsets of, for example, doubles so that the compiler can choose the most efficient instruction sequence for example for min/max? If we know a double is never NaN then we can emit only one instruction on x86, but have to emit one more on arm64. If we know a double is never zero and never NaN, we can emit a single instruction on both.
This whole conversation between relaxed SIMD and deterministic SIMD seems to only exist because our compilers are not smart enough and/or their whole program analyses don’t support any plugin-like capabilities.
There are other examples where if we know a SIMD bitmask is canonical (all 1s per lane) then we can implement horizontal reductions more efficiently. This is very niche and I doubt that any language supports interprocedural analyses with such a rich domain, so it feels like a hole in the programming language space.
Pattern types could potentially maybe in the future allow for the compiler to know more details though. They are a nightly feature, and afaik only for enums, integers and pointers so far. The idea would be that you can define a custom type such as "an integer between 7 and 45" and everything else become niches for niche optimisation (e.g. for `Option<MyFunkyInt>` some of those impossible values would be used to represent the None case of the wrapping Option).
But I could envisage a future in which you could say "f64 without NaN" which would both make those available for niches and potentially tell LLVM about this. However, we are very far from any of that currently. And it might not be what you want, since you would need to add checks when you perform operations to ensure the value doesn't suddenly become a NaN. Which is way more complicated than ensuring integers don't become, say, zero. It will likely be much harder to optimise away the checks.
Thanks, I’ll take a look!
SBCL has the right hooks. Jai probably does as well. You could build your own thing on LLVM relatively easily - so Rust _could_ do it if you gave it sufficient access to the compiler, but so could C++ or D.
SIMD bitmasks being ~0 instead of (00000001)+ is common and useful, but there's an annoying language design question in there about what type comparisons between vectors should be (if you don't have a <N x i1> as a type, say because you liked C too much). Shout out to std::vector<bool>.
There is probably interesting work in mojo for this. The library being in MLIR strongly encourages taking that sort of approach. I haven't looked at their implementation though. Happy hacking!
But I suspect you're overvaluing the potential savings. Knowing when a float is 0.0 or NaN beforehand is almost entirely impossible, except for the most trivial of cases - like when you first initialize a variable or first enter a loop. Everything after that is very hard or impossible with floats as they are.
Those cases can be const folded at compile time.
Those cases are never a measurable bottleneck.
The closest thing I know of in the realm of the optimization you're curious about is Rust NonZero* variants, but they're used for enum compression afaik.
I guess what I would like to see is SIMD libraries being able to confidently say nobody needs to use intrinsics (or differentiate between relaxed/normal SIMD on the user API level) because the language + high level SIMD APIs are smart enough to choose the right implementation.
IIRC IEEE min/max with proper NaN handling needs 8 instructions on x86 vs 1 on arm64 I find it very sad that we apparently haven’t really solved that yet without forcing the user to use different APIs.
Note that:
NonZerof32 * NonZerof32 -> NonNanf32
NonZerof32::from_bits(1) multiplied with itself is zero.
Doing range analysis needs the language to support it at compile time, and the dev to specify what range it is.
The only 'stable' thing i can think of is a type for 'greater-eq-one' using only addition and multiplication. Practically every other operation breaks most of the type knowledge up to that point.
vminpd ymm2, ymm1, ymm0
vcmpunordpd ymm0, ymm0, ymm0
vblendvpd ymm0, ymm2, ymm1, ymm0
If you want proper IEEE 754-2019 minimum (propagate NaN, -0.0 < +0.0, NaN bitpattern picked in the usual way) you can do it in 6: vminpd ymm1, ymm0, ymm1
vbroadcastsd ymm2, qword ptr [rip + .LCPI0_0]
vandpd ymm2, ymm0, ymm2
vorpd ymm1, ymm2, ymm1
vcmpunordpd ymm2, ymm0, ymm0
vblendvpd ymm0, ymm1, ymm0, ymm2
I personally find this a load of nonsense I don't care about.If you want propagating NaNs but don't care about signed zero or NaN payload/sign, you can use
vminpd ymm2, ymm0, ymm1
vminpd ymm1, ymm1, ymm0
vorpd ymm0, ymm1, ymm2
What I do in Polars is a bit different, there for propagating NaNs I do if (self < other) | self.is_nan() { self } else { other}
this isn't fully optimal on x86-64 but it's fairly simple and autovectorizes decently on various platforms, here's AVX2: vcmpltpd ymm2, ymm0, ymm1
vcmpunordpd ymm3, ymm0, ymm0
vorpd ymm2, ymm3, ymm2
vblendvpd ymm0, ymm1, ymm0, ymm2[1]: https://github.com/Dr-Emann/memchr_n#performance [2]: https://github.com/BurntSushi/memchr/pull/241
Towards fearless SIMD, 7 years later - https://news.ycombinator.com/item?id=43519823 - March 2025 (175 comments)
Towards fearless SIMD - https://news.ycombinator.com/item?id=18293209 - Oct 2018 (81 comments)
Nice work. Looking forward to it.