> The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.
> Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields such as depth, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.
I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.
> I'd also be interested in how post-processing fits in with this.
I think screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.
WoW goes to lengths to hide its depth buffer.