upvote
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.

Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.

Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.

I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.

reply
> Even if the output is like 5-6 tok/s, that might be usable for some purposes.

You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second.

I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn't be surprised if this is an empty set.

reply
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month.
reply
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.

There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.

Average electric rates by region:

    New England            28.1 cents
    Mid Atlantic           25.1 cents
    East North Central     20.8 cents
    West North Central     14.8 cents
    South Atlantic         16.1 cents
    East South Central     15.5 cents
    Mountain               14.6 cents
    Pacific Contiguous     26.1 cents
    Pacific Noncontiguous  42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph...
reply
I'm actually getting 11 cents in winter, 13 in summer, but my utility company is a co-op. Average for my state is I think 19 cents.

I think you can get down to around 8 if you are signed up for an interruptible load, or a dedicated off peak load, depending on the company, but yeah, standard rates aren't that low.

reply
> Pacific Contiguous 26.1 cents

This is a bit misleading, because it's combining the 50 cents/kWh from California with 15ish cents/kWh in Oregon and Washington. Seattle City Light, for example, charges 13.38 cents/kWh on flat rate pricing, and far less with time-of-use billing (8 cents/kWh on off-peak).

reply
I'm in Arkansas and get rates fairly similar as quoted.

From my last bill

> KWH USAGE 2590 - $183.37

There's a base customer cost of $18 on top of that, but yeah ~$0.077/kWh taxes included.

reply
Is that for a month or a year?
reply
My monthly bills are similar to this, near Arkansas.

Last one was $212 for 2,146 kwh between June 8 - July 6 (28 Days)

reply
If you run off solar with battery backup, you can achieve lower than those rates! Look at Time of Use rates. The super off peak rates instantly become the max price point once you pair TOU with Solar + battery.
reply
Not the parent, but here is one location that has rates in that range in the U.S.

https://casscountyelectric.com/rates

reply
A lot of people quoting low rates are also just referring to their off-peak rate. This is pretty common in EV discussions. It's not exactly a fair argument there, either, because the flip side of having an off-peak rate is that the on-peak rate is usually quite a lot higher. So the true effective rate is a bit higher, somewhere in the middle depending on usage pattern.
reply
deleted
reply
It's easy to have your EV only charge off-peak, though. It's just a setting.
reply
My point is that the tradeoff to get off-peak pricing is that on-peak is way, way more expensive. So you can charge the EV off-peak to maximize the savings, but everything else you do during on-peak time costs way more.

Using myself as an example:

I adjust my A/C to run outside of 5pm-9pm (peak) if at all possible, we try to avoid pointless high-draw usage during that same window, and both of our EVs hold off charging until after 9pm.

My rate from 5pm-9pm is 0.43/kWh. My rate after 9pm is 0.09/kWh. The flat rate alternative, if I did not want to worry about time of day, would be 0.21/kWh. These prices are all-in, including transmission and distribution/whatever.

It would be dishonest to say that my EVs only cost me 0.09/kWh to operate, which on it's face is a claim to paying over 50% less. In reality, time of day pricing typically saves me somewhere between 10% and 15% in an average month compared with flat rate.

reply
If you leave the EV charging out of your consumption, does a time of day plan still save you money on the remaining usage? Or does it cost you? If it saves you money, then it would make sense to be on a ToD plan regardless of EV charging. Which means it makes sense to consider your additional EV draw as costing the marginal off-peak rate. Essentially the EV load has the valuable property of being dispatchable.

You can do the same thought experiment with say a dehumidifier in your basement. It can easily be off during peak usage and still accomplish its job, so its cost of electricity is also the marginal off-peak rate.

reply
> If you leave the EV charging out of your consumption, does a time of day plan still save you money on the remaining usage? Or does it cost you?

It would cost me more (modestly so, less than 10%) to be on TOD without the EVs. This will vary by customer, of course, and I expect that the power company designs TOD to be a wash for the average customer. They even guarantee it won't be more than 10% more expensive over the first year or they will refund the difference.

reply
Fair point.

I guess it depends on if you would be using ToU otherwise.

It looks like about 50% of Californians use ToU plans, but the number is only 10% nation-wide.

reply
deleted
reply
They are most likely not based in the US, but converting to USD to make comparison easier.
reply
I am not in Quebec but Quebec hydro rate D for standard residential would be one example of around what I pay.

https://www.hydroquebec.com/residential/customer-space/rates...

Another example would be Manitoba hydro

All figures in Canadian currency

https://www.hydro.mb.ca/account/rates/residential/

reply
deleted
reply
Specifying USD is indeed often a service usually offered by people born elsewhere for people born elsewhere. Americans seem rarely know about these mysterious places, where bills can come in all sorts of funny sizes and colours. (kind of joking)
reply
> I pay about $0.075 USD per kWh

Around here electricity companies quote prices like yours but that is supply only while transmission, taxes, and fees are again as much on top. Is that really all inclusive?

reply
>and people will compromise speed for data sovereignty

People should always compromise speed for data sovereignty! Who said: that in this digital day and age, information about money is more important than money!

reply
can send safely context if there’s confidential computing ala my site https://trustedrouter.com/
reply
How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces?
reply
Seems clear from their website?

1. Their API server provide an attestation JWT. This JWT is signed by Google's private key. 2. The attestation has details on the running container. I suppose the container host is a Google-provided distro and Google's signer will verify that the OS is theirs and up-to-date. 3. They could've proxy the attestation. To prove this is not the case, the field eat_nonce include the TLS certificate fingerprint, which should match the API server you're connecting to. I suppose you will need to pull their container and verify from the source that the container itself generate the private key, it never leaves the container, and the container has no way to run arbitrary code such as SSH or vulnerabilities.

reply
Is your local compute airgapped?
reply
My local compute is used by me, and I'm accountable to myself whether or not is secure. So to a certain degree, I trust myself and also know what limitations / potential vulnerabilities it might have.
reply
I trust my local compute quite a bit more than a random project that has "trusted" in its name.
reply
Great, so the other member of the set matters for you more than cost.

Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?

reply
Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output.

On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code.

reply
>Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?

Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation.

reply
Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product.
reply
The whole mentality of thinking one knows better than another about what they need causes infinitely more problems than it solves.
reply
As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.
reply
Do you mean by trading dollars for the privacy you need as:

a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place

or

b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place

or

c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds?

Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing.

reply
A) is very doable with e.g. Amazon Bedrock.

They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights.

What kind of privacy needs do you really have beyond that?

reply
It is not my use case but given recent political developments in international relations caused by the executive branch of the US government, off the top of my head, I could think of a lot of European or Canadian firms for which that would not be an option. No matter what they might promise about European sovereignty. For a good 'ol patriotic US domestic company? Sure.
reply
Yes, its the US cloud act risk EU companies run up against on hyperscalers like MS/AWS.

Even for EU companies running open weights on EU stacks LLM inference on the GPU must process plaintext and I can't find any EU provider with NVIDIA H100/H200/Blackwell CC mode plus SEV-SNP or TDX, where you can cryptographically verify the workload ran somewhere the operator cannot inspect.

Personal compute is therefore the only option if you want personal autonomy privacy for IP &c. Maybe another option is to use cloud compute rented to fine tune a personal model that suits your own needs that would help bring the cost down, I don't know enough about this area to know if it kills the "intelligence" of those domains due to limited ?cross-verification within the LLM.

reply
It's also worth considering what you are actually paying for. And it's not keeping the data private, it's taking the blame when there is a breach. Same reason companies hire big consulting firms whenever they need to make an important but possibly risky decision.
reply
>very doable with <US company>

for anyone not US-based, this company is hostile and you have to assume the US government can and will force them to give access to your data.

reply
deleted
reply
There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud.
reply
Having worked in / adjacent several such industries, a lot of the question depends on scale.

A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this.

Worst-case example: Bootstrapped startup working in military.

It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc.

* Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.

reply
> Worst-case example: Bootstrapped startup working in military.

That's the easiest case.

AWS Bedrock models running in AWS Secret Cloud for Industry. (I really have no affiliation with them, I'm just like... this is a completely solved problem, why do people think this is hard and requires on-prem hardware?)

https://www.aboutamazon.com/news/aws/aws-secret-cloud-for-in...

I'm with GP that these are tinfoil hat concerns, when there are solutions to all of these, unless you're perhaps in some country with very specific needs beyond things like European sovereignty or US military secrets (like a non-US defense concern).

reply
You seem to be categorizing everything that considers their data being in the possession of the US an unacceptable risk to be tinfoil hat, which is kind of an insult to a large portion of the world. If you haven't been paying attention to the news in the last 48 months, the political reality has shifted considerably.

Note that the other commenter never said US-based military oriented startup. You just assumed, then jumped to "heck yeah let's use Amazon Secret Cloud for Industry"

Not everyone has or wants an office in Crystal City.

reply
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.

If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised.

I would tend to trust Amazon to avoid being compromised.

reply
At least for regulated applications I worked on, no one cared.

The provider needs to comply with specific rules, have specific certifications, and sign specific agreements. You check the boxes, and you're good to go.

Microsoft does that better than anyone. OpenAI and Anthropic don't do that at all. Google does that rarely and poorly. AWS is not bad, but not as good as Microsoft.

Azure was always my go-to for regulated applications in the cloud. Some do require e.g. on-prem or even air gap, where even Azure is out.

reply
The expectation that one BigTech company has a competent security team while the other doesn't seems entirely baseless?
reply
Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers?

I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems.

reply
It's completely academic. At 5tok/s you can process 13 MTok per month at concurrency 1. I use 5 BILLION tokens per week when coding.
reply
Yeah. At 5 tok/second, you're talking about around $195 worth of output tokens per month. There is no way I can run a usable K3 model for $195 a month of capex, opex, or any-kind-of-ex.

Qwen 3.6 is another matter. Paying provider rates for the amount I run locally would put me in the thousands of dollars. So that's very practical to buy a Macbook instead, plus an RTX card, and so on.

reply
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty.

Are there? At the highest levels of defense and law, AWS and Azure are used.

Having tried selling some of these entities on doing things in-house, there seems to be little interest.

reply
> Are there? At the highest levels of defense and law, AWS and Azure are used.

This is certainly true if the user is an American company. You could look at the European initiatives to run this stuff on hardware they own in facilities they own and control within the borders of Europe for a counter-example.

Such as: https://www.google.com/search?client=firefox-b-d&q=schwarz+s...

https://www.dutchnews.nl/2026/04/government-turns-to-german-...

reply
Yeah, true European cloud providers for these kinds of things seem to be behind, and a lot of the ones offering data compliance at the level of AWS are small enough that it's a bit harder to trust they'll be around and will keep their promises.

Hopefully that changes!

reply
I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower.

I don’t know if that’s 100x more than I’d pay (opex-wise) with an nvidia setup, but I can say the one-time capex is a great deal cheaper. Avoiding VRAM and DDR5 (fast DDR4 should be OK) are the biggest cost savers. ECC RAM is worth the extra price. General datacenter-quality hardware has less price sensitivity, and plenty of bang for your buck.

reply
Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour.

Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the RAM bus speed. An older 2U rack server that is 600W load (typically a 1+1 power supply server when plugged into two kill-a-watt would show 300W on each, equal load balancing) when maxed out with stress-ng might be only 450W total running inference.

If you have 600kWh used in a month by running something 24x7 and your power is $0.15 a kWh, that's more like $90/mo (not counting cooling or any ancillary costs for the environment where it's in).

reply
deleted
reply
If you actually were running this thing at 80% or 100% load, then the first thing you'd want to is get a better PDU and then connect your servers to that (48V DC).
reply
One of the problems in buying used/refurb x86-64 rack servers for test and development/proof of concept environment, is that by volume in the market, there's not that many -48VDC power supplies going around, because maybe 5-10% of enterprise customers buy them. Resulting in many fewer units ending up on the resale market.

The options for AC power supplies for servers with 2 or 4 load sharing redundant power supplies are a lot greater. If you were buying all new hardware and starting from a clean sheet of paper design with lots of money to spend, absolutely. At that point also start looking at higher voltage DC distribution stuff related to open compute platform and 800VDC.

But if I were trying to make the absolute most use of $20,000 to put together a 3TB RAM server (48 x 64GB DIMMs), it would likely end up AC powered.

reply
(Context: Parent comment was edited after I wrote this comment)

Where in the world are you finding that much RAM in a racked server for $200/month?

reply
I think he means electrical bill at his estimated wattage load of the server and his known kWh cost, not rented server/hosting cost.
reply
Aha, right. That makes a lot more sense.
reply
At my house. I have 5Gbps fiber and could pay for 10 or 25 if I need it.
reply
Gotcha. But to be clear, you’re talking only about energy usage, correct?
reply
Yes, what other opex is there? It will have good ventilation, I’m not worried about cooling.
reply
One aspect of this is cyberattack proliferation by way of "Hey boss, I saw this TikTok that says if you let me invest [a tiny piece of the neighborhood's profit|our militia's budget] into some RAM, I could get a fully autonomous cyber operation up and running that pays for itself via ransomware etc. within weeks. You like it, we upgrade to something that can work even faster. We don't need the hacker guy from Swordfish with fifty monitors, we just need my cousin who likes building gaming PCs."

That's a world that I don't think we're ready for.

reply
A similar world is already here.

Young men 14-?? already compromise and attempt to extort organizations daily, sometimes cluelessly from western nations, often not. It doesn’t have to be gangs when the home country doesn’t care / isn’t technologically or culturally developed.

Already seeing AI-written payloads and frameworks in the wild. I think it’ll turn out that AI won’t build you a maintainable ERP but it can create C2 networks, exploit POCs or even 0-days potentially, and let kids make their own ransomware tooling. Then we’re dealing not with a handful of cybercrime tool makers but a generational problem.

reply
I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case.

Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that.

reply
I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully.
reply
Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer.

And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch?

I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :)

reply
You can rent one in the cloud to try it
reply
They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.

So full CPU local AI inference may become viable option in coming years.

reply
the GPU competition is using 16 gpus, so the actual comparison is that the CPU has <1/10th the bandwidth
reply
This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage.
reply
Im wondering if we are finally seeing the end of the "hard disk" era, and are entering a new era of vast instant on systems.
reply
> running it on a no GPU, but tons of RAM server

Or from SSD using something like Colibri[1]. Not going to be quick, but at least runable.

[1]: https://github.com/JustVugg/colibri

reply
It's a great concept but I think it would cross the line from 'very slow' to 'so slow it's unusable' at this size. Even if we say you have an NVME SSD that does 7GB/s reads, that's dramatically slower than being able to hold the whole thing in DRAM. Like the difference between 1.3 tok/s in RAM vs 0.1 tok/s with a colibri-like method.

edit: the results I have seen from people trying colibri with fast consumer grade PCI-E 4.0 NVME SSD are 0.1 tok/s on models that are <700B in size, things that are well under 800GB on disk. With something that's 3T in size it'll probably be a lot slower than hat.

reply
For single stream inference of a MoE model, the size of active sparse parameters will matter a lot more than total parameters. This is generally around half of the reported active parameter count - the other half being a dense subset that can be easily cached in VRAM even on fairly modest consumer setups. So the achievable performance may be quite a bit better than a naïve assessment might suggest.
reply
On a server machine you can have more than 100GB/s of NVMe if you parallelize (RAID 0 and the like). But it's still gonna be noticeably slower.
reply
1536GB of DDR4 ECC server RAM is somewhere between $4000-6000 USD used right now, by the time you put in parallel enough NVME SSD to approach good speeds, you'd be approaching that (and also likely running out of PCI-E bus lanes directly attached to the same motherboard to reasonably do so).
reply
It claims to support using multiple devices RAID-0 style, which should boost performance, but yea probably not very useful for most.

But still fun you can run it at home.

reply
Won't the answer (even for a pretty basic message like "hi") at SSD speeds take like a _entire week_ to _start showing useful output?_ (attempting to do 22k average claude code system prompt + 32k thinking tokens thru 0.1t/s throughput)
reply
As you already went through the thought exercise of laying all this RAM over various slots, then match against the right CPU (which also you'll need multiple) - it becomes clear quite fast that it's trying to mimic the architecture of a GPU except in extremely low fidelity and bandwidth @ a higher energy cost.
reply
> But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.

The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3

Also, their previous models were native INT4, so it would be weird if they went larger now.

reply
Update: Looks like the model is larger after all (1561.44 GB). Only the MoE weights are MXFP4, while the other weights are BF16 (and a few FP32).

* Sparse Experts: 1481.4 GB

* Dense Experts: 1.9 GB

* Self-Attention: 72.4 GB

* LLM Head: 2.4 GB

* Embeddings: 2.4 GB

* Vision Encoder: 0.35 GB (surprisingly small)

plus some miscellaneous parameters.

Most importantly, we now know that the model has 104B active parameters, which is quite a lot and will make it difficult to self-host efficiently.

reply
It will be not 5 tok/sec. More like 0.5 tok per sec with a fast cpu setup.
reply
deleted
reply
> Even if the output is like 5-6 tok/s

On a 3T model I’d imagine you’d be closer to 0.05 tks

reply
Presumably it’s MoE and only needs to read a small fraction of the weights per token. Bonus points if you can get decent speculative decoding without becoming ALU-limited.
reply
Speculative decoding is not really worthwhile for sparsely-loaded models. You end up paying in both memory bandwith and compute (loading experts based on wrongly-predicted tokens) which leaves you worse off overall. It becomes viable (even for sparse MoE) once you're batching so widely that you end up having to load most of your total weights anyway.
reply
> Speculative decoding is not really worthwhile for sparsely-loaded models.

If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens, without too much quality loss.

Another fun idea might be to try to build a model where the router chooses the expert 1-3 tokens in advance.

reply
> If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens

You can!

> AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation.

https://machinelearning.apple.com/research/introducing-third...

reply
Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.
reply
The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).

But the memory bus speed is fully committed when generating tokens or thinking.

reply
Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.
reply
The epyc Venice SP7 socket is apparently 9324 pins

https://x.com/tomshardware/status/2066846693778510331

reply
We are going to need a bigger boat.

https://www.cerebras.ai/

reply
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.

I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.

Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.

reply
I dont quite understand why GGUF is better optimized. Are the performances better for the same amount of VRAM ?
reply
GGUF is at least better than bnb. From what I know, bnb does not yet find a way to quantize MoE with enough accuracy, and maintain the dequant-MoE kernels. In the age of Qwen 3.0, people tried to make some bnb '4-bit' quants of MoE models, but actually the MoE part is not quantized. It's a pity that even Unsloth gave up low-VRAM finetuning with MoE (although they're making their GGUFs for inference), and the world of local training looks stagnated for months.

GGUF is maintained by all the llama.cpp developers. There are many quantization formats and algorithms under this container format, some are optimized for MoE (such as APEX quant), some for CPU and some for GPU, some work surprisingly well below 4-bit (and even near 1-bit). It also supports recent architectures like linear attentions and mHC.

reply
I cannot be sure what the likes of Cursor have done, but I think it's incredibly unlikely that they have trained a QLoRA for Composer.

It's almost certainly full parameter post training of the original model weights.

reply
> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".

No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.

Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

reply
> No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.

Still useful; "are the labs marginally profitable just on the marginal inference costs?" is still a useful question to answer. After all, if they aren't even profitable on inference in isolation, then we can expect to see large price increases.

If they are able to turn a marginal profit on inference alone, then perhaps the price increases won't be so severe (or perhaps they expand the time between generations so that they spend less on training but take longer to complete training).

"Are the labs profitable at all?" is, of course, a much more useful question, but that doesn't mean that the first question is completely useless.

reply
I wasn't saying it isn't useful, I was saying that you cannot infer even the marginal cost of closed models.

The K3 maths can turn true only if the models size is roughly the same and the labs didn't find any better way to run inference at scale.

reply
We know labs make money on inference, and we know they lose a lot of money on inference+training.
reply
> We know labs make money on inference

We don't really know that, for OpenAI and Anthropic. We suspect that, but as far as I know, even they have stopped claiming that they are profitable on inference.

reply
unless you think that Opus is 10T+ params, its pretty much impossible for inference not to be profitable when doing some basic napkin math on other open models, and if Kimi K3 is 3T params with the same performance as Opus then that means that China is actually way more technologically advanced than the American labs.

So which is it?

reply
Just out of curiosity, based on what we know for sure they(OAI+A) make money on pure inference and lose on inference+training?
reply
OpenAI's financials leaked and showed this pretty convincingly.

Anthropic was probably profitable last quarter, without training costs: https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...

reply
> Without training cost you can infer only the marginal cost of serving this kind of models.

Which is by far the most interesting number of the two.

> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

If you get close in output quality, then does that matter?

reply
> If you get close in output quality, then does that matter?

When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?

reply
That only matters if you are an investor not if you are a consumer.
reply
Well, or if you're participating in a discussion on HN where the sub-topic happens to be "if labs are subsidising tokens on API pricing" and literally the cost of serving the tokens is relevant to the sub-topic people are trying to discuss...
reply
Consumers want better models too, of course it matters
reply
> Which is by far the most interesting number of the two.

Only if you don't have to continuously train new models, and you are not at a runway risk.

reply
Training cost is directly impacted by inference cost nowadays. Most of the gains come from RL these days, and that is highly dependant on inference (~7:1 inference:training in units of compute). That's mainly because you want many roll-outs for each training scenario.

Of course inference efficiency is dictated by model architecture, size, etc. You can still guesstimate some of those and have an idea about cost/serve at several size tiers.

reply
That isn't really relevant to GP's point. We still don't know what training costs because we don't know how much RL is done.
reply
deleted
reply
Gross margins are insanely important, possibly the most important single metric if for some reason you were forced to choose one.
reply
Only if you assume that at some point, for any reason, there will be "the model" that doesn't need costly retraining.

I guess this is one of the reason Anthropic i so "active" for asking for a development break.

reply
"I want my competitors to stop competing" is an interesting ask for someone who's currently charging the highest prices in the entire industry.
reply
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.

Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled?

Because Moonshot didn't even incur the majority of the training costs either

reply
> the training cost of the model(s) that Kimi distilled?

The distillation process involves getting conversation traces from the model you are distilling from, and then training your model against them.

You still have to do the training!

reply
I believe they are talking about the closed models' training costs.

I other words, the providers that will be offering K3 inference don't have any training costs to offset, so they are only charging for the inference itself. OAI/Anthropic would need to offset their R&D and training costs in order to not be selling API access at a loss.

reply
It's 3/15 - https://openrouter.ai/moonshotai/kimi-k3

If you're going to open source your model, why would you set your own price high enough that other providers could easily and profitably undercut you?

reply
Since the model is natively MXFP4, I think it'll be even more interesting on the hardware front. It'll comfortably fit on a 8x AMD MI355X node. I suspect that'll drive token prices down, further.
reply
deleted
reply
Say a single Kimi K3 is deployed on 16 x B200s: how many concurrent users can that handle? I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
reply
>I realize the question assumes a major simplification that everyone's prompts/sessions are the same.

well, exactly.

that's tough to answer without just average sampling because some users will ask the model "what's todays date" or "what color is the sky?" and some users will ask "Let's rewrite the linux kernel in brainfuck."

reply
So maybe the better question would be how many concurrent actively-thinking/working agents it could handle?
reply
AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.
reply
This is a very insightful eye opening take. I haven't even thought of it this way. This really is the first open model to be as big as the frontier has been until now.

I think this release is actually great both ways when you think about it. We gonna be able to learn knowledge that labs have been hiding from us (e.g. cost like you mentioned). And labs could learn from whatever optimization techniques people come up with when trying to host this model.

It's honestly just good for everyone in my opinion.

reply
If it is a mixture of experts (MoE) model like the 2.x models, won't this reduce the hardware needed to run the model?

The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. On a single B200 you can have 5-6 experts loaded into VRAM at a time. Realistically that would be 3-4 to account for the context. [!]

[!] With this and other MoE models it looks like an interesting area for research would be to detect or predict which models would be needed ahead of time. That way you could schedule the load into VRAM step before the weights are needed. That way you shouldn't lose much/any performance from offloading the weights to RAM.

reply
You need whole weights in VRAM for optimal performance. Don't be confused by "experts" in the name -- you don't get to load static subset of experts and blast next 100 tokens with them. In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.

> With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. Kimi K2.6 is released as INT4 already.

So 5090 with K2.6 is just gonna sit idle 99% of the time, waiting for next slice of weights to load.

5.6 Sol calculates that single 5090 in raw compute & memory bandwidth can run K2.6 at 35 t/s (256k context depth) -- if it somehow had enough memory to hold whole model in VRAM. Man, I hope HBF succeeds and Nvidia brings it to consumer cards in 5 years..

reply
> In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.

It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.

You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability; though the obvious tradeoff is that having to store KV caches for the wider batches may leave you with less room to cache experts across layers and tokens.

(Plus if you're batching so widely that you end up loading essentially entire model layers, MTP then becomes applicable even for a MoE model. But this typically only applies if you're doing inference on a very large scale, or if your memory bandwidth is so scarce that you have to recover compute intensity by any means feasible.)

reply
> It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.

Caching really has nothing to do with this. With RAM offload you can mostly benefit from:

1) Batching for prefill is a huge win, even with MoE, since the batch sizes can be so large.

2) Keeping non-expert weights in VRAM, so the percentage of weights used per token in VRAM is higher. This benefit reduces with larger models, though.

> You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability;

With MoE it's low probability.

> MTP then becomes applicable even for a MoE model

With MTP it becomes _extremely_ low probability.

reply
> With MoE it's low probability.

For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.

> With MTP it becomes _extremely_ low probability.

This is actually right, MTP is only ever worthwhile in very special cases involving either dense models or extremely wide batching of MoE ones that somehow still leaves unused room for parallelization (which AIUI would involve an assumption of very abundant compute with very limited memory bandwidth).

reply
> For even the sparsest MoE open models, having more than a handful of inferences in the batch is enough to make it more likely than not that you'll get some MoE weight reuse within any given layer. This assumes totally random sampling, ignoring any cross-request correlation that would push that probability even higher in many practical scenarios.

If you tell me the model and the number of parallel streams, I will do the math.

reply
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090

Without leveraging system RAM and/or SSDs, I don't think you can, or how exactly are you running this, if this is something you are doing today? With CPU/expert offloading you could probably do it with a 5090 + 1TB of RAM or something like that, but absolutely not on a single 5090 entirely within VRAM.

reply
Yes hybrid approaches are much better than people realise.

There are a lot of optimisations that are not in the public sphere, source working on start up in this space

reply
> There are a lot of optimisations that are not in the public sphere

Sure, but if we're participating in public discussions, isn't it more fun if we talk about things people can actually read and understand, rather than secret stuff other's can say work, but no can actually validate or know how it works?

It sounds like "hybrid approaches are much better than the public is aware, because everything else is private and secret", but also: ok, so what? No one can run that anyways, (yet?), so why it matters?

reply
Yes, that's what I was saying w.r.t. expert offloading, i.e. ensuring that the GPU could fit the active parameters not all the parameters.
reply
Alright, I guess I misunderstood. To be fair, this part:

> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090.

Is painting a very different perspective, even considering the latter parts it's hard to read that as "Of course offloading everything else that doesn't fit on the GPU itself". But anyways, it's been clarified now so no harm :)

reply
I assume you mean putting only the 32B active parameters on the GPU, and the rest on a bunch of regular server DRAM like on a 768GB to 1024GB RAM server?

Because Kimi K2.6 in Q4 is about 584GB GGUF size on disk and will use slightly more than that in RAM, Q8 is 595GB.

https://huggingface.co/unsloth/Kimi-K2.6-GGUF

reply
You're talking about running this "at home" for 1 user, using a mix of VRAM and RAM (total should be ~1.5TB). That's certainly possible. It'll be slow, especially prompt processing, but doable for single users.

But my comment on running it was more towards serving this profitably at scale. You get much better throughput / unit of compute if you load everything in VRAM and serve many requests at the same time. That's how all inference providers are doing it.

reply
I was talking about running this on a server, hence my comments re 1xB200. Obviously, the more hardware/VRAM you have the better/faster you can run these large models. But if you are a small/medium sized company you could feasibly do it on just one B200. It all depends on how much hardware you can afford to run.
reply
If it takes so much resource to run, how does the sharing of a single llm works? There is some interface that basically submits context/cache plus current promt, from each user, doing basically time-sharing compute?
reply
I believe they batch requests so the same weights in vram are shared across many requests https://www.baseten.co/blog/continuous-vs-dynamic-batching-f...
reply
You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest.

Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.

Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.

reply
I think it's unclear the the DS hosted prices are profitable. AFAIK that haven't claimed that.

OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers

reply
I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :)

But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model.

> DeepSeek, which hosts DS v4, profitably

I specifically mentioned 3rd party providers, because there can be an argument that model creators themselves are subsidising tokens to gather training data for the next model. In fact, ds are public about their gathering of data (at least on openrouter they're marked as such). So that 0.x price point for dsv4-pro is likely subsidised.

reply
Based on the best available information, DeepSeek is pricing the API such that they can repay their infra capex over 10 months, while deprecating/amortizing the cost of said infra over 3 years.

For my product, I run GLM 5.2 and other models myself, in production, on rented hardware. Paying API prices would cost much more.

EDIT: You can now see several other third-party providers for Kimi K3 (Nebius, Fireworks). All charge exactly the same as the first-party. Does that mean that their costs are the same? Seems quite unlikely. It's simply not an efficient market, yet.

reply
They need to get a license from moonshot to provide inference for K3. Probably have to follow the pricing set by moonshot as well.
reply
Xioami's MiMo did match DS-V4's price, although we now know that DeepSeek set their pricing lower than they could have, due to the leaked memo, and simply decided to use "10 months to recover capex" as their yardstick. Interestingly "10 months to recover capex" is also the same price SpaceX is renting space to Anthropic and Google for.
reply
I think cursor will likely do a grok fine tune rather than a kimi one for the next composer.

they noted in their blog post they didn't focus purely on coding for grok 4.5.

reply
Most likely, since they were acquired. But for us outsiders it would be a cool thing, to see if the delta is the same between kimi2.x + cursor data -> kimi3 + cursor data.
reply
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy -- it cannot be compressed further. On all tested large models, it's a ratio of 1.34-1.23 -- and smaller models up to 3.76x. It also increases the effective bandwidth by the same rate.
reply
Exploring compression algorithms for weights is a good idea, and I hope you have a successful product. However, if you can prove this statement:

> reduces it down to its minimum entropy -- it cannot be compressed further.

I think you could make a lot more money elsewhere :-)

https://en.wikipedia.org/wiki/Kolmogorov_complexity#Formal_p...

reply
We're not an AI company ... nor do we have any reason to use it. Just a fun idea that was fruitful.
reply
And to clarify -- it only applies to models, not arbitrary data. So its useful to exactly zero other fields.
reply
That's very interesting. Does that mean you can reduce say, a 30B class Q8 from ~30 GB down to 10 GB or less?
reply
704gb -> 564gb; 358 gb -> 270 gb; 28.79 gb -> 7.65 gb; 439 gb -> 93 gb

It depends on the total entropy of the model. Smaller models have less entropy.

reply
> Smaller models have less entropy.

Interesting. Why is that? I would have expected the opposite, since larger models have to try less hard to fit the training data. Or maybe this leaves more parameters with random initialization, resulting in higher entropy for larger models?

reply
I honestly don't know... I didn't train the models, so I can't tell why the math works out that way. It just does. I suspect it has to do with the fact that all the small models I've tested have been quantized. I don't know of any small model trained from scratch. If you know of any, I'd be happy to encode it and see what it looks like.
reply
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy

LOL

reply
deleted
reply
[flagged]
reply
Thanks for your opinion
reply
deleted
reply
So nowadays the hardware and hosting providers must be in an optimization race, whoever can make the model just a bit smaller or more efficient (to fit on fewer/less powerful cards) will have a huge advantage and can make a lot of money.

I am curios what's the most profitable thing to "plant" (agriculture analogy) on the land (cards) that you have have: web hosting, vps, llms, image/video generation, etc

reply
>realistically you'll need 16x for context / throughput optimisation

Sounds like I'm buying a lottery ticket this week so I can drop $800k on hardware.

reply
> if "labs are subsidising tokens on API pricing"

> SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%

Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...

https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...

https://finance.biggo.com/news/02d45650-b569-4d12-b44d-8d6d8...

reply
Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
reply
even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.
reply
Will the model even be competitive in 10 months though? Seems like models that reach top 20 on OpenRouter see 50% of all token spend by day 80, and 80% by day 180.
reply
As long as the hardware can be used on newer models, hardware costs can be recouped running a future model.

But if they're hoping to recoup non-recurring engineering costs rather than just hardware costs, they do need to consider the useful lifetime of the specific model.

reply
That numner blends in training or no?
reply
> The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models.

Sota closed models don't even answer cybersecurity questions lol.

reply
pardon my ignorance but is fine tuning still considered viable in face of rapid model releases. Is it really worth it?

my friends whove tried in their companies gave up on it.

reply
Anyone who thinks that the labs are not profitable on per token API pricing is delusional and hilariously wrong.
reply
It all depends if you count the fixed cost of training or not. And the cost of the hardware.
reply
Or even the basis of the cost of hardware. There are lease deals, capacity traded for equity, various programs by Nvidia, there's absolutely massive depreciation, etc.
reply
Depreciation is massively overrated by Micheal Berry and his ilk.

Everyone keeps thinking those A100s only have 6 more months of life, and yet they're still going for more than they did per hour in 2024.

Show me evidence that A100 prices have collapsed, and maybe GPU depreciation will be relevant to the market.

reply
[flagged]
reply