Are there any articles you’d recommend for this?
I have Qwen running on an HP Z8. Very nice platform.
I have mine in a sandbox, due to privacy fears.
Your solution sounds more elegant.
I really just iterated over the harness over and over for about two weeks with opencode until I was sort of satisfied (still lots to do there :).
For the llama.cpp I asked claude fable to optimize it for my hardware and iterated a few times. In the end I landed on the following: https://pastebin.com/2PpJFUC0