upvote
The whole series had an upgrade a couple of days ago actually — they have addressed embedded tool calling (and hopefully the MTP formatting stuff though I gave up running the Gemma MTP because it's often slower than not-MTP)

Not tried it yet but I've seen tests that suggest they've properly fixed the tool calling issues.

reply
I find the 4-bit QAT with MTP to be entirely usable speed on both my boxes (Strix Halo and a desktop with two V620 GPUs, which are slightly faster than the Strix Halo).
reply
For whatever reason prefill (on my DGX Spark) is faster with the Gemma models than Qwen 3.6 models of similar size. On vLLM anyways. Likely just deeply tuned code contributed to vLLM by Google?

vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE.

reply