upvote
We're on the cusp of Kimi K3 becoming usable on sub-10k hardware.

https://github.com/gavamedia/deltafin

reply
For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.
reply
14 seconds per token? Not tokens per second. Seconds per token? That’s nowhere near the cusp!
reply
This is a bit of a strawman. The parent comment wasn't suggesting that the inherent defects magically disappear nor was it suggesting that local inference today is sufficient for all tasks.

At the heart of it, self-hosting liberates your use cases from all the horribly opaque configuration, shadow prompting, etc. And local models are only getting better and more diverse every month.

reply
>The parent comment wasn't suggesting that the inherent defects magically disappear

that is almost exactly what they said though..?

"Want it to go away, almost like magic? Local inference. [...] all of the common LLM defects will go away.

reply
> This is a bit of a strawman. The parent comment wasn't suggesting that the inherent defects magically disappear

The parent comment literally said that the common LLM defects would go away.

Direct quote:

> all of the common LLM defects will go away

reply
I get your reading. My interpretation is that they’re claiming enshittification and that the local models don’t have that.

It’s an understandable view, but I’d be astonished by any local model processing very long context better than any frontier model (and now many racks are we talking).

reply
> Want it to go away, almost like magic?

True, almost like magic and magic are not the same thing, but I'd hesitate to call this a strawman.

reply
[dead]
reply
Do you work for a frontier lab by any chance?
reply