This:
https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg
...and this is what I should probably wait for (not sure why it's in vlm):
https://github.com/Blaizzy/mlx-vlm/tree/pc/laguna-s-nvfp4
...or perhaps I should've just used the gguf with the provided llama.cpp instead of trying to run the nvfp4-mlx from the get go, but where's the chaos in that :)
Running deepseek flash on something locally now, this will have to wait a bit. I still stand by my initial quick assessment - looks capable. Some people on r/localllama also reported loops. We'll see in ~10 hours. Hopefully I haven't terribly mislead people.