Despite being outdated in terms of compute, it stills let me run very good recent models locally, with Deepseek V4 Flash 0731 being the greatest one right now, and hopefully Qwen 3.8 Flash will also fit well when it is released tomorrow!
The M1 ultra definitely leaves to be desired in terms of its token speeds, but I think 20 tps generation and ~200 tps prompt processing (which is what I get with DSv4 flash), is already enough to do a lot of serious work when you combine with the decent prompt caching provided by llama.cpp.