upvote
If you think about it there's overlap in what the VAE encodes and main model encodes. Objects at a distance resemble texture and textures zoomed in gain structure. The VAE makes textures more efficiently representable at the cost of reducing the representable space of pixels. So things like tiny text become nonsense scribbles. Working in pixel space, especially with something with recursive or cascaded structure, opens the possibility of using the learnt structure of real writing at a higher level to perfect tiny details that actually cannot be approximated without being obviously wrong.

Time is "just" another dimension. There's temporal continuity between frames, a video VAE would be learning and representing those temporal shifts, but there's nothing to say that e.g. a recursively applied generative model at the pixel level also doesn't learn and represent those things.

As ever, figuring out how to train the thing is the hard bit I expect.

(Handwaving over "textures" here, VAEs encode more like somewhat macro blocks of image whose content is also conditioned on surrounding blocks, rather than tiny patches of patterned pixels.)

(And yes I'm a total imposter layman here, I just see VAEs as seeming to be a crutch that reduce data size - super super helpful of course - but being strictly speaking redundant and inhibiting correct fine detail.)

reply
This is a totally fair point and definitely worth exploring!

The jumping point for this no-VAE work was three-fold:

1) Our goal here is to get 32x32 token reduction to make video training and inference downstream cheaper. To date, the best open-weight Image VAEs like Flux-2 seem to cap out at 16x16 token reduction (8x8 VAE + 2x2 linear patchification). Others like H3 have pushed to 32x32 reduction but requires them swapping out the small VAE decoder with a 2B parameter decoder. So, this is a foray to get 32x32 compression without compromising quality.

2) We believe that end-to-end trained networks will tend to perform better than modularly trained networks (e.g. VAE + DiT). This hypothesis comes from work like REPA-E, where authors are able to get much better results by backpropagating through the pretrained VAE. The latent space for perception / reconstruction seems to have a different "optimal" configuration than a latent space specifically built for generation. That's why we liked the idea of trained E2E here.

3) There's work with VAEs that show that providing additional modalities (e.g. text captions) can help the VAEs improve as well. That's natural to this construction, so we thought language might actually help with the compression, not hurt. To be honest, this this is the most hypothetical of three ideas; and definitely warrants specific ablations.

reply
Also i failed to clarify that the VAE for Linum v2 was the Wan2 VAE
reply