upvote
> We're running Kimi 2.8 on a $107k server

Equipped with what? Is it CPU based inference, a mix...?

reply
Not GP, but my educated guess is that they are running a system with between 4 and 6 MI325 or MI355x or similar AMD GPUs. With the 50k tps as the total figure for all parallel requests. Those cards have a lot of memory for their price, allowing you to push to really high batch sizes while still having a large context size for each request
reply
What is your config?
reply
how many requests per second can the server take?
reply
Are you developing software? Is most of it used on a coding agent? (Like Claude Code or ChatGPT Codex?) If so, what coding agent do you use? If you're not developing software what do you use it for (roughly)?
reply