At 9000 tokens/s you could interleave a lot of requests so long as pre-fill is also fast.
It really depends on how much you need to keep sessions open to take advantage of KV caching
It depends. If I was running a model locally I would much prefer 9000 t/s. If I was running an inference company, obv 1000 concurrent requests at 150 t/s is preferable.