This is great until the redundancy no longer exists and you have to rebuild the entire state machine from zero (a humble scene transition or rapidly turning a corner).
Gbuffers have been using motion vectors forever. For realtime global illumination, techniques like ReSTIR already allow for temporal reuse and spatial coherence.
I just don't see the purpose of replacing the traditional rendering pipeline for neural rendering techniques. It would be one thing if we were constrained on the number of triangles we could push per second, but we're not. The bottleneck is usually elsewhere in the pipeline: animating characters/objects, particle systems, physics updates, etc.