Hacker News
new
past
comments
ask
show
jobs
points
by
ThunderSizzle
10 hours ago
|
comments
by
acchow
1 hours ago
|
next
[-]
That's about right. Half the TPS with half the power.
GLM-5.2 will be much more demanding tho
reply
by
syntaxing
7 hours ago
|
prev
|
[-]
Which quants? I get these speed (only 20-30 TPS) at Q4_K_M with MTP for 27B on my framework desktop. I draw sub 130W for the whole machine.
reply