upvote
Why did you use 1-bit quantization vs 3-bit quantization?

It looks like the 3-bit requires 90 GB[1] which, I imagine, would fit within the DGX Spark's 128GB of unified memory.

[1] https://unsloth.ai/docs/models/qwen3.8-next

reply
If I read correctly, that's based on a 1-bit quantization, and can we really expect that to produce any useful output at all?
reply
You should find an excuse to offer 3D printed extruded pelicans from various models as awards for something. I have no idea for what, but the idea captivates and I'd love to win one somehow. They'd be collector's items in a few decades
reply
If Simon would pitch for example PCBWay that and I am pretty sure they will sponsor it (assuming their logo stays). They can do laser engraved versions also ;)
reply
The spark can easily run UD-Q4_K_XL on this model... using IQ1_S doesn't make much sense.
reply
Doing this on a 1-bit quant is unfair
reply
Tried again with a different quant, UD-Q2_K_XL:

https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

reply
the xhigh version looks amazingly good for a 2 bit quant.
reply
The Pelican Brief
reply