Sure.
And how does that make your day better? I know it does not improve my work in any way shape or form.
I'll take a better coding model that's not multi-modal any time.
If I need an LLM to do images or sound, I'd rather use a dedicated one instead of a jack-of-all-trades-master-of-none model.
Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.
Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.
You can get decent open-weight models now. That's not difficult. The difficulty is 1) running them and 2) compliance.
My company runs Claude on GCP's Vertex AI solution. We're in the US healthcare IT space, so the models need to be from somewhere that American healthcare agencies and companies have traditionally been okay with sourcing code from - which means the US, Canada, and maybe Europe. The stuff that handles PHI/PII must be in the US. The expense of hosting is more of a PITA than most customers want to go through this early in the technology's lifecycle, and intelligence gains are simply a matter of degree for most business tasks.
In theory, we could find some open-weight model (likely from China) for our development agentic work and host it anywhere you can host AI models. We don't, though, and I think Google, OpenAI/Microsoft, and Anthropic see that as the core of their business.
https://artificialanalysis.ai/#intelligence-comparison-tabs
Differences in token "density" are accounted for by pricing per task