RAM was probably the bottleneck for the amount of context they were offering.
I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful
When it first launched on OpenRouter I was getting nearly 70 Tokens/second.
Edit: Ah:
> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
... I'm at a loss for words here. It was being served for free. To the entire world.
> ... I'm at a loss for words here
No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.
Why is Luna not free on OpenRouter? :)