upvote
How useful is the second 3090 in this setup? I run the 5-bit quantized model on a single 3090. Does the second 3090 allow you to use the full precision model instead or a less aggressive quantization by splitting the layers? What about running the 35B model instead?
reply
But is Qwen3.6 27B actually worth this investment? If I had to guess you still use SOTA for architectural/planning work?
reply