upvote
Kind of reminds me of this more recent paper: https://arxiv.org/html/2602.02459v1 Different use case and implementation, but a similar idea. In this case applied to sharing last state(not the whole KV cache) from big brain model running in the cloud with a smaller/dumber model running on-device in a robot, in a latancy-aware way.
reply
You might like The Universal Weight Subspace Hypothesis: https://arxiv.org/abs/2512.05117

Curious to know if anyone is aware of research trying what parent suggested?

reply
Sounds similar to Ramp's Latent Briefing for multi-agent coordination.

https://x.com/RampLabs/status/2042672773747589588

reply