upvote
I had a long talk with chat gpt about this today as well. I think its duable and prolly not too hard either, also you could do lotsa funky stuff with stitched frames of a video in one 4x4 grid for example and send that as one image for analysis. that way temporal understanding can be had for fractions of a second by jev... also because vlm works in pixel space you can get around the whole state machine issue as well, so many possibilities...
reply
[flagged]
reply