Hacker News
new
past
comments
ask
show
jobs
points
by
cobanov
4 hours ago
|
comments
by
adinb
1 hours ago
|
next
[-]
It doesn’t to be a
ton
bigger, 16k and reliable 8k would be a godsend. (I run at 2k)
reply
by
mikodin
3 hours ago
|
prev
|
[-]
What are the models? I am super curious in these as well
reply
by
simcop2387
2 hours ago
|
parent
|
[-]
Probably Kev and/or the decider models. Kev is trained on one of the 4B qwen models, similar for decider but it ranges from 0.8B through to the 35B-A3B model so far I believe.
reply