How long is this runway?
The other day I managed to get a context of 195k for Qwen3.5-9b Q4_K_M using a llama.cpp fork that supports TurboQuant:
https://github.com/TheTom/llama-cpp-turboquant
I think you could replicate this with a larger model on your device.
Overall with the right quantisations for both the model and KV cache you can get a lot of mileage out of this old hardware. Speed remains the main limitation, as IIRC I was getting ~26-30tps on a 7700S.
I'm not saying there's no benefit to using local models. My point is still the same as before: you have to be willing to sacrifice both performance and quality even when just comparing to free models that are available today.
I spent some years in the pharma industry, where AI came really late, as there was (justified) concern that sensitive data might leak - even by accident. Eventually a solution was implemented - it was some kind of open model (they didn't say which) running on-premises.
You couldn't tell these people that this and that powerful model is free or inexpensive, because it's useless to them if it's running on someone else's computer i.e. the cloud.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
With AMD the best you can do in the consumer market right now is an RX 7900 XTX which is about 960 GB/s.