People have had surprising success adding vision to open-weight LLMs that ship without it, like DSV4 Flash [1] or GLM-5.2 [2]. Given this model is already vision-trained I expect that approach will work well here.
[1] https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_...
[2] https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4