Kind of reminds me of this more recent paper:
https://arxiv.org/html/2602.02459v1
Different use case and implementation, but a similar idea. In this case applied to sharing last state(not the whole KV cache) from big brain model running in the cloud with a smaller/dumber model running on-device in a robot, in a latancy-aware way.