Do you have $2/hr to rent an RTX 6000 96GB or $5/hr for B200 180GB on the cloud?
You can get lots of tokens per second on the CPU if the entire network fits in L1 cache. Unfortunately the sub 64 kiB model segment isn't looking so hot.
But actually ... 3000? Did GP misplace one or two zeros there?