upvote
Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally.

That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf

While trying to find the link for that I stumbled upon

https://arxiv.org/pdf/2605.12732

Which also looks pretty interesting

reply
I think this is the main thing I want out of "non-backprop ML" research - figuring out how the whole class of algorithms behaves. And then applying that to figure out how the brain implements its own deep learning.

If we can figure out how the brain's learning dynamics function well enough? We could figure out how to interface with them and extend them.

reply
The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.

"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.

The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?

reply
What about applications like distributed computing using volunteer computers, instead of datacenters? You can't really do that with normal backprop approaches. E.g. for LC0 the training data is generated in a distributed manner, but the neural nets themselves are trained centrally.
reply
I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"
reply
I think both can be true - ie. We can harness energy from the sun at much bigger scale than nature has so far, yet we can still learn a lot from how nature has spent massive parallel search over millions of years on coming up with incredible nano engineering we still have no clue how to do
reply
Even today's LLMs suddenly get power-competitive when you compare by "power per task". Sure, a GPU can draw 1000W under load. But it also works very fast, and doesn't have to spend any time on things like "sleep".

The trick about comparing the two is that different things are expensive to different substrates.

Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent, purpose specific, and fully embedded into the data path, coordinating them can get less natural and, frankly, optional. When wiring is expensive, coordination can become expensive in turn.

Another thing in the same "cheap for GPGPU but expensive for brain" regime is bandwidth. Look no further than optic nerve to see just how hard it is for nerves to push any appreciable amount of data. Another thing is connectivity. For GPGPU, global connectivity is natural - but the brain has to pay in physical wires for all the connectivity it has, and, see "bandwidth": it doesn't have any good wires. Yet another thing is weight reuse: a big part of why humans get "handedness" is that the brain can't just reuse the motion control circuitry for one hand for another nearly identical hand.

And the final thing I can name off the top of my head is memory - but specifically, memory capable of fast R/W. The capacity of human "working memory" is a disgrace, and not because there was no use for more. Humans rapidly lose visual fidelity of representations for objects they aren't directly looking at, and not at all because "being able to check how things looked 2 seconds ago" is useless. Those capabilities were just too expensive for the substrate to afford them easily.

It's why brains, broadly, favor dataflow-like and SSM-like dynamics, with largely fixed asynchronous dataflows and recurrence over updated local information - instead of something that would require a lot of global connectivity and transformer-like many-to-many attention ops. SSM is not necessarily the "best" tool for the job in ML land, for most jobs - but when you struggle to fit "attention" into your connectivity/bandwidth budget, and your memory is extremely expensive but hard-coupled to processing, SSM starts looking very appealing.

Now, something that might be expensive for GPGPU but cheap for brain, for once? Online learning. Maybe it's substrate dependent, or maybe it's going to get cheap in GPU land too once we figure out the trick. But so far? No one figured out how to make it cheap, stable and usable. You'd be lucky to get "pick one".

reply
I mean the reason why models are using less energy is because they are getting smarter per token and also engineering algorithms/chips that make inference cheaper.

If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof

Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.

But coding output/W LLMs have us beat

reply
Coordination is fuck all bits worth of control information broadcast widely. It's very cheap to us.

I frankly don't believe in spiking neural networks giving any advantages over what we have. It's a different way to implement ANNs, but "different" isn't "better". It's how the brain does things, sure, but the answer to "why the brain does what it does" is "workarounds for being made of flesh issues" at least half the time.

I can believe in brain having more capacity than frontier LLMs quite easily. We know a single BNN neuron can have the expressiveness of many ANN neurons. And well leveraged overparametrization + compute overhang could explain a decent chunk of the apparent sample efficiency edge.

But that apparent "extra capacity" could also be tied up in things like neurons having to contend with metabolism, in brain's learning algorithms being noisy, in brain having to use neuron circuits to implement "hot memory", etc - instead of contributing only to performance.

reply
Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server?
reply
No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination
reply
Imagine we did that, split up a model layers as A->B->C. C will need to wait for B to compute a forward pass, which is waiting for A to compute its forward pass. To compute the forward pass, B needs all of the outputs from A, which is an upload and a download (maybe these can be done concurrently).

Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data.

This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it.

reply
Yep. You need to transport all the weights at the boundary regardless of Backprop/NPC.

But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.

An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts

reply
The problem (and contrast with other approaches) is that mat muls requires synchronization. Arranging your networking and training structure to maximize compute and minimize communication is the main craft of ML training infra folks. In your example, yes you can compute layers on different machines (i.e. Tensor Parallelism), but you must be very careful in how you arrange it.
reply
I think Jeff Dean is right in that we will see much more specialised silicon in the future.

If something more bio inspired ie. predictive coding and in-memory compute fundamentally makes continual learning and much lower energy consumption possible there will be specialised hardware for it at some point

FWIW I think the brain has multiple “learning rules” and operates at multiple timescales

reply
Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity?
reply
One aspect of how current training and continual learning are somewhat at odds is that the memory required to train a model is often times 2-3x the memory required to just run it (probably not as bad for PEFT, not sure).

DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.

There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)

reply
Models suffer from "catastrophic forgetting" if you train them on new data.

People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"

reply
Maybe the practical path is to separate fast-changing memory from slow-changing weights. Most things an agent learns during use probably don't need to become parameters immediately.
reply