Qwen3.8-Flash-Next support will also be added very soon.
Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.
Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.