The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.
https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...
If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.
I pressed Enter, and the response was instant.
> Generated in 0.037s • 14,205 tok/s
This is unbelievable.
and it gave a very reasonable answer in non-perceptible time.
I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.
"LMS algorithm in bash"
Just barfed it up lol.
Amazing.
This is crazy.
Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen.
and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV.
It seems China is already able to do DUV a lot sooner than others expected.
That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind.
China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction).
The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.
That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame).
How would that disrupt the industry?
High speed SRAM is where the $$$ is
I don't think it would be that difficult to manufacture compared to other process tech. HBM is really hard to do compared to other memory types.
Google is already working on a similar idea but more "flexible".
The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.
But if this is even at 400B size it's insanity those inference prices, maybe 10-20% margins, if it's higher I would like to know is it their own chips or maybe they have accurately sized the model to fit on exactly a B300?
Could be a lot of magical things we can only speculate, but from here there likely isn't another 60-70% margin, like I have heard people claim, I would definitely be willing to bet on that.
Could still be a healthy 10-30% margin. Especially with Terra.
[0] I am constantly surprised how much work pay-as-you-go with DeepSeek / MiMo will get done. I've barely crossed $2 each in a month of use (~200m tokens).
I feel perfectly content in using pay as you go pricing with deepseek. On the other hand, although Anthropic's models used to be my bread and butter for personal work, they are simply too expensive to reach for these days.
Assuming the efficiency gains are real, I feel like something has to give, maybe worse quality due to aggressive quantization/kv cache compression?
Although I'm sure there are some efficiency gains, the technology is too new and labs are scrambling to release too quickly to think that the low-hanging optimization fruit has been picked already.
If you use Codex it's different, the harness has a lot to do with it and there's definitely been changes including recently.
But yeah I do'nt want to know what Kimi 3 is pushing buttons inside Anthropic, OpenAI and Google.
Besides any floor: For every year the tokens get faster and cheaper, we will see new things like properly working AI factories which mimic expert teams. A lot more parallism as well.
MI500 series is supposedly already taping out and they're claiming massive increases (we'll find out end of 2027 prob).
Besides Nvidia Hardware is still sold out and super expensive. Not a single Nvidia consumer GPU got cheaper at all, Nvidia DGX Spark got more expensive too.
It will be swooped of the market the second it hits the market.
Anthropic's big marketing push this year has been entirely focused on getting people to use Opus via a Claude Code subscription, to the point that Sonnet is almost viewed as the poor man's alternative, and from what I've seen, almost nobody uses it.
Actually, here's an interesting project for all the vibe coders looking for their next front page post: scrape a ton of commits from GitHub with Co-Authored-By: Claude and figure out what the percentage split between Opus/Fable/Sonnet is. I'm willing to bet it's less than 10% Sonnet.
This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.
Luna is an extremely strong model.
By benchmarks, which sadly is a poor measure. Yes Luna is a good model under certain circumstances. Whether it is great for general usage is another story. Sonnet is definitely better when prompts are more vague and it needs to decide things. Luna generally sticks to things very strictly and goes off in bad ways.
I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.
That's again not some "intelligence factor" here. Different agents work for different use cases. Luna wins some. Terra wins some. Sonnet wins some. Flash was really good at exploring.
So I'm not sure what your point is? There's a big market for everything. Even within the market you describe it's likely not a Luna-size fits all either.
We are not purely rational creatures, thank God. Sometimes those "limiting factors" you listed -- stress, peer pressure, hormones -- are crucial elements of informing the problem solving process and arriving at a decision or a solution that actually works.
All an LLM can do is fulfill a prompt, no matter how misguided, backwards, or incomplete that prompt actually was.
"Go jump off a bridge." Hmm. Dying makes me stressed out. I'm not gonna do that.
You mean they increased the price and then cut it back and now it is amazing?
Luna had a price hike vs mini (its previous replacement). The cut now just puts it back in that ball park.
Not that this isn't good news, but what's impressive?
I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.
Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.
No it hasn't!
A century ago, some nails cost 2.5% of disposable income, and now the same nails cost 2.3% - only a little cheaper.
The cost of nails has remained remarkably consistent for a century. The problem is that you have ignored the depreciation of money.
Let's assume California prices and income and pick a bigger retail package of nails as you might use for building a house. The numbers used to calculate percentages: in 1926 a 50lb keg of 4" nails was $2.75 and median after tax income might be $108 per month. In 2026 a 50lb carton of 4" nails is $106 and income might be $4,516. Albeit I assume nails are now more readily available and the quality of nails is likely better; and perhaps I should have compared galvinised nail prices.
With an 80% reduction in cost that becomes a ridiculous outlier in efficiency.
This is very likely priced below recovering the cost of the hardware but still above operating expenses.
I have no idea either way but one thing that detracts from these threads is folks claiming things as a fact without evidence.