Hacker News
new
past
comments
ask
show
jobs
points
by
XCSme
1 hours ago
|
comments
by
4k0hz
6 minutes ago
|
[-]
It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.
reply