I agree completely - behaviorally the models have changed drastically due to RLHF, RLVR and now maybe even more so due to agentic harnesses. But the mechanism of prediction hasn’t changed, that was all I was clarifying.
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
As you say, that's just an efficiency play and, as I understand it, doesn't change the behavior of the models beyond perhaps a small amount of sampling noise.