Sure.
And how does that make your day better? I know it does not improve my work in any way shape or form.
I'll take a better coding model that's not multi-modal any time.
If I need an LLM to do images or sound, I'd rather use a dedicated one instead of a jack-of-all-trades-master-of-none model.
Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.
Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.