Hacker News
new
past
comments
ask
show
jobs
points
by
kadushka
7 hours ago
|
comments
by
danielmarkbruce
5 hours ago
|
[-]
I might still be misunderstanding what you are saying, but bitnet also keeps high precision latent weights during training. The optimizer updates those, while the weights used in the forward pass are quantized to ternary values.
reply