Hacker News
new
past
comments
ask
show
jobs
points
by
why_only_15
11 hours ago
|
comments
by
sroussey
10 hours ago
|
next
[-]
Those machines with GPUs still need RAM of their own, and they generally want large caches to avoid SSD penalties. You even see this spill out in the form of costs for KV cache in <1min, 5m, 1hr rates etc.
reply
by
halJordan
10 hours ago
|
prev
|
[-]
The majority of inference actually does happen in cpu.
reply