Hacker News
new
past
comments
ask
show
jobs
points
by
spIrr
7 hours ago
|
comments
by
mordae
7 hours ago
|
[-]
To run at decent speed, all models try hard to use only most likely relevant part of the context and most likely relevant weights (MoE) to predict the next token. Doing the math in full is unfeasible.
reply