upvote
I've seen a rather impressive example of what sounds just like this in Qwen recently

https://media.discordapp.net/attachments/1401891025970008154...

I don't know what went into making it, but their twitter is @araminta_k if you're curious

reply
More info, for anyone interested in the above:

* https://alvdansen.github.io/animating-on-twos/

* https://github.com/alvdansen/animating-on-twos

* https://huggingface.co/alvdansen/h3-keyframe-animation

## Quick Start

An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:

https://huggingface.co/alvdansen/h3-keyframe-animation#quick...

reply
deleted
reply
The task you're describing is a video model task, not an image model task. It's inherently temporal.

Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.

reply
Sure and that's what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
reply
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do.
reply
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.

I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.

reply
They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.
reply