upvote
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
reply
deleted
reply
I think you’re missing that MTP can predict more than 1 token in advance.
reply
In fact, isn’t that the “M” in “MTP”?
reply
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
reply
deleted
reply