upvote
Without prompt caching this becomes more expensive than fable 5.1 after turn 50, assuming you start with 40k tokens and add 2k per turn.
reply
> There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds

I'm confused. If it's 1500t/s, isn't that only 90k per minute? How do you hit a 450k/minute limit?

reply
Cached tokens count towards the limit as well. For example, if your context window is 50,000 tokens, it takes 9 requests to reach that limit without generating a single token.
reply
Cached tokens counting toward the limit is ridiculous.
reply
then it's basically useless lol, wtf, this has to be a defect
reply
Could this also be coming from the problem that Qwen3.8-27B's default mode being "extra-high reasoning level"?
reply