upvote
> It’s multimodal though.

Sure.

And how does that make your day better? I know it does not improve my work in any way shape or form.

I'll take a better coding model that's not multi-modal any time.

If I need an LLM to do images or sound, I'd rather use a dedicated one instead of a jack-of-all-trades-master-of-none model.

reply
Personally, I often paste screenshots into Claude Code of the application it’s working on. And I’ve even had it work autonomously on something and regularly grab its own screenshots.

Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.

Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.

reply
At least 30% of my queries to an LLM include deciphering something from an image. If I’m coding pretty much a hundred percent of bugs include some sort of image.
reply