upvote
The Nvidia video presentation on DLSS 5 says that the model was only trained with various G-buffers as input (including the depth buffer) but during inference, the model only uses the rendered frame. As well as the previous rendered frame reprojected via motion vectors, if I understand correctly, likely to improve temporal stability.
reply
The technical report suggests it does not use the depth buffer: https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf

> The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.

> Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.

reply
Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn't be of benefit to bring as much of that kind of metadata to the model as possible. You'd think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.

I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.

reply
Hmm well the depth buffer only has accurate depth for opaque objects (and even that's not really true). Things like hair wouldn't be in there. Its not ground truth depth.
reply
I think it has more to do with what kind of data they have access to at runtime - IIRC DLSS upscaling has only required the previous frames and motion vectors, so requiring depth buffers would mean it was no longer a "drop-in" replacement

> I'd also be interested in how post-processing fits in with this.

I think screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.

reply
That doesn't make sense; every game has a depth buffer, but not every game has motion vectors.
reply
Because the depth buffer isn't always accurate or accessible.

WoW goes to lengths to hide its depth buffer.

reply