I am running GLM 5.3 across 2x DGX Sparks and was doing comparisons and it absolutely can beat Gemini. Yesterday it corrected a poor Fable 5 response even
Yes they are quite good, but are not able to run on a 16GB RX 9070.
Quantized Qwen 3.8 Flash Next could maybe run eventually on that card with a highly optimized inference engine that dynamically caches the hottest layer experts. Even then you run into some hard limits.