upvote
Yes, one of the main goals of this model was to keep 3D consistency as the camera moves around; it can do this even without relying on an explicit 3D point cloud or Gaussian splat representation of the scene.
reply
Amazing! Thanks for explaining!
reply
I believe so. This is not a model that generates pixels frame by frame from user input like genie 3. Instead, there’s an actual 3D scene / structure generated (point cloud, 3dgs) from the input images.
reply
deleted
reply
Wrong, it does go straight to generating images.
reply
From blog post:

“It generates both image frames from novel views and explicit 3D outputs”

Model can indeed generate novel views but I don’t think in real time (I could be wrong). If you want to navigate a space in real time as user above was asking gotta rely on the 3D output and the explicit representation will provide continuity. User likely was referring to the limitations we’ve seen on video-style world models like genie and precursors.

reply