points
I'd be interested to see where GPU code beats NUMPY's SIMD implementation, which is really
we use numpy + jax for that; works well