upvote
Convert it to an image on the fly to feed it into a vision language model and I expect it would work just fine.
reply