Hacker News
new
past
comments
ask
show
jobs
points
by
Eisenstein
1 days ago
|
comments
by
girvo
1 days ago
|
next
[-]
It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch
reply
by
13 hours ago
|
parent
|
[-]
deleted
reply
by
medvezhenok
1 days ago
|
prev
|
next
[-]
I think you’re missing that MTP can predict more than 1 token in advance.
reply
by
spider-mario
10 hours ago
|
parent
|
[-]
In fact, isn’t that the “M” in “MTP”?
reply
by
beastman82
1 days ago
|
prev
|
[-]
Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)
reply
by
1 days ago
|
parent
|
[-]
deleted
reply