upvote
The path to ubiquitous AI (17k tokens/sec) https://news.ycombinator.com/item?id=47086181

You can still try it at https://chatjimmy.ai/, but it's running the rather outdated Llama 3.1 8B

reply
reply
That looks promising! As models become a commodity, this may turn out to be the real AI gold rush.
reply
A question of course would be "is 15,000 tok/s Gemini better than 100 tok/s Opus 5"?
reply
It's better for the billions of free users that Google serves.
reply
Depends what you do. We have certain tasks we spend money on where Gemini 4.6 definitely is better than Opus 5.
reply
Given how fast models are improving, burning the weights into actual ROM is prohibitively expensive if you need (or want) to replace the chip every couple of months.

The alternative is on-TPU flash for storing the weights.

reply
Yes. We're still a ways off from this being ubiquitous, but I am convinced it's coming.
reply