Hacker News
new
past
comments
ask
show
jobs
points
by
pllbnk
9 hours ago
|
comments
by
jakswa
5 hours ago
|
next
[-]
dang only for certain nvidia GPUs, had my hopes up
reply
by
lowbloodsugar
5 hours ago
|
prev
|
next
[-]
Ok, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request.
reply
by
pllbnk
33 minutes ago
|
parent
|
[-]
Even without ninfer I would get over 80 on LM studio with default settings, so it should be noticeably more on 6000. You might want to try different a different inference engine or settings.
reply
by
beastman82
9 hours ago
|
prev
|
[-]
can't second ninfer enough. amazing tech
reply