We introduce PC-ALM, a local alternative to backpropagation. PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics. PC-ALM equips each layer with a feedback control dynamical system that distributes and propagates supervision credit throughout a network.
Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly
There are several reasons the brain can't implement exact backpropagation. One is “phase locking”
In this post, we introduce PC-ALM (Augmented Lagrangian Predictive Coding), a method for training networks that replaces the forward and backward passes of backprop with layer-local dynamical systems. Each layer is coupled only to its neighbors. Instead of forward-then-backward, we run each layer forward in time. When run to convergence, the dynamics of the whole system distribute supervision credit signals quickly and accurately across the entire network.
PC-ALM is an extension of standard predictive coding (PC)
We compare PC-ALM to PC and backprop in a suite of experiments. Local training methods such as PC have historically been difficult to scale. Following the PC literature, we use simple tasks (Fashion-MNIST, CIFAR-10, etc.) and networks such as residual MLPs.
We show that PC-ALM can successfully propagate supervision credit in 1000-layer neural networks, overcoming standard PC's signal decay problem
We focus on deep, small-width networks, a regime in which PC tends to perform poorly.
Ultimately, our motive is to understand how distributed systems (such as the brain) can implement gradient computations without backpropagation. Scientific motivations aside, this research may inform energy-efficient deep learning on neuromorphic hardware, where dynamical systems simulation is cheaper than on GPU
Predictive coding: each layer as a dynamical system
Before explaining PC-ALM, let us first explain PC, interpreting it from a dynamical systems standpoint to emphasize its role as a backprop alternative.
Predictive coding
Predictive coding has its roots in Helmholtz's theories of unconscious perception
Mathematically, predictive coding utilizes a general motif: take a state and update it to reduce a prediction error at the next step,
By applying this update rule to each layer's activation vector (the “state” is the layer's activation
To explain this in more detail, let's write a feedforward network as a constrained optimization problem:
where
This is a quadratic relaxation of the constrained problem.
To train a neural network, PC alternates between inference and learning steps:
Predictive coding
Per mini-batch, a forward pass initializes the activations, followed by
Each
The bottom
Both the inference step and the weight update are layer-local. The weight update is Hebbian-like, in that it multiplies a postsynaptic error by the presynaptic activity (a delta rule), and the dynamics map onto a neural circuit with explicit error neurons
PC trains deep networks, but exhibits signal decay
Since minimizing the free energy with respect to each
Nevertheless, PC has been shown to successfully train networks on simple tasks. For example, MNIST and Fashion-MNIST in 128-layer residual MLPs with wide layer widths (512 neurons per layer)
However, PC struggles on more complex tasks and networks
Each layer adjusts its activity to reduce prediction errors with its neighbors. Supervision enters at the output, but must work its way through this chain of local compromises to influence earlier layers. In deep, narrow networks, the resulting credit signal becomes weak long before it reaches the input.
This leads to a documented signal-decay problem of PC
Our method, PC-ALM, introduces a way to improve signal propagation of PC networks, retaining the layer-local dynamics of PC and keeping inference budget
Augmented Lagrangian Predictive Coding
We propose Augmented Lagrangian Predictive Coding, a variant of PC that uses the augmented Lagrangian (AL)
supervised loss + Lagrangian term + PC energy
In each layer, the augmented Lagrangian introduces a Lagrange multiplier (or dual variable)
The augmented Lagrangian is used extensively in distributed optimization
LeCun (1988)
To use the augmented Lagrangian for training, we make a simple modification to PC:
Augmented Lagrangian predictive coding
Here
By accumulating local prediction errors, the dual variables recover the exact backprop credit signals in a deep linear network. We derive this result in the paper.
Thus, at least for linear networks, PC-ALM gives a method for computing exact supervised loss gradients and distributing them throughout a network, using only layer-local dynamics.
In the experiments below, we test whether this advantage carries over to nonlinear networks.
Mechanistic interpretation
To understand, mechanistically, how PC-ALM works, consider a simple scalar network, with hidden unit
We attach a multiplier
The first step of PC-ALM agrees exactly with a PC step. After that
At convergence the activation
Control theory and credit assignment
PC-ALM offers a control-theoretic perspective on credit assignment. Each layer combines its current prediction error with an accumulated error signal—the proportional and integral terms of a PI feedback controller.
Results
We present below some results for PC-ALM.
PC-ALM: a local learning method for training 1000-layer neural networks
PC-ALM successfully trains 1000-layer MLPs on MNIST. We use the residual MLP setup from Innocenti et al. (2026)
Our results are presented in the figure below; PC-ALM achieves near-BP performance using only layer-local dynamics.

Image classification benchmarks
We tested PC-ALM on a small set of image classification tasks and found that it improves performance over PC on every task, including when training ResNet-18 on CIFAR-10 and Tiny ImageNet.

PC-ALM propagation dynamics
In addition to the performance of PC-ALM across image recognition tasks, we find PC-ALM exhibits surprising dynamical properties. In both PC and PC-ALM inference, movement of the hidden activations
Stable oscillatory transient responses
Individual neurons in PC-ALM also exhibit damped oscillations during inference.
Discussion
This work introduces PC-ALM, a layer-local alternative to backpropagation for training deep networks. To our knowledge, this is the first layer-local method shown to successfully train networks up to 1000 layers.
Summary
Predictive coding remains an attractive candidate for a theory of cortical function. It is rooted in Helmholtz's ideas on unconscious processing. It was later shown that predictive coding can be viewed as a layer-local alternative to backpropagation for training deep networks. PC in its standard form can successfully train networks on simple tasks, but signal decay limits its performance in deep, narrow networks. By introducing Lagrange multipliers and running primal-dual inference on the augmented Lagrangian (vs PC's gradient flow on an energy), PC-ALM lets individual layers compute a gradient signal of a global loss function, using only communication between neighboring layers. The primary motivation of this work is to understand more deeply the mechanisms of credit assignment in real physical systems such as the brain.
Constrained and lifted optimization
PC-ALM relates to a long line of work on constrained optimization approaches to training deep networks. These approaches “lift” training into a larger optimization problem by treating the activations
PC-ALM draws on this optimization lineage to address a question from neuroscience: how can local neural dynamics compute and distribute credit for a global objective?
Prospective-configuration tradeoff
In PC, the settled activations differ from the forward pass. Song et al.
Motivations
This work began with three observations:
1
The Neuro-AI and distributed optimization communities share a concern with locality, but there has been relatively little cross-talk between them.
2
The PC energy is suspiciously similar to the augmented term of the augmented Lagrangian that is commonly used in distributed and constrained/lifted optimization.
3
LeCun (1988) identified the multipliers of the standard Lagrangian with backprop credit signals, suggesting that these multipliers could serve as local credit signals.
Our work on PC-ALM ties these threads together.
Future work
There is a wealth of work that needs to be done. Notably, extending PC-ALM to temporal tasks with temporal credit assignment
We hope this work inspires further connections between augmented Lagrangian methods and biologically plausible credit assignment.
T.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.
Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020.
DOI: 10.1038/s41583-020-0277-3
Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment
A. Ororbia.
arXiv preprint arXiv:2312.09257. 2023.
J. Sacramento, R.P. Costa, Y. Bengio, W. Senn.
Advances in Neural Information Processing Systems, Vol 31. 2018.
‘Backpropagation and the brain’ realized in cortical error neuron microcircuits [link]
K. Max, I. Jaras, A. Granier, K.A. Wilmes, M.A. Petrovici.
PLOS Computational Biology, Vol 22(4), pp. e1014164. 2026.
DOI: 10.1371/journal.pcbi.1014164
Backpropagation through space, time and the brain [link]
B. Ellenberger, P. Haider, F. Benitez, J. Jordan, K. Max, I. Jaras, L. Kriener, M.A. Petrovici.
Nature Communications, Vol 17, pp. 66. 2026.
DOI: 10.1038/s41467-025-66666-z
M. Jaderberg, W.M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, D. Silver, K. Kavukcuoglu.
International Conference on Machine Learning (ICML), Vol 70, pp. 1627—1635. 2017.
Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment
A. Ororbia.
arXiv preprint arXiv:2312.09257. 2023.
T.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.
Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020.
DOI: 10.1038/s41583-020-0277-3
J.C.R. Whittington, R. Bogacz.
Neural Computation, Vol 29(5), pp. 1229—1262. 2017.
DOI: 10.1162/neco_a_00949
Predictive Coding: A Theoretical and Experimental Review
B. Millidge, A. Seth, C.L. Buckley.
arXiv preprint arXiv:2107.12979. 2021.
A survey on neuro-mimetic deep learning via predictive coding [link]
T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.
Neural Networks, Vol 195, pp. 108161. 2026.
DOI: https://doi.org/10.1016/j.neunet.2025.108161
Learning on Arbitrary Graph Topologies via Predictive Coding
T. Salvatori, L. Pinchetti, B. Millidge, Y. Song, T. Bao, R. Bogacz, T. Lukasiewicz.
Advances in Neural Information Processing Systems. 2022.
C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.
Proceedings of the 43rd International Conference on Machine Learning. 2026.
Advancing Neuromorphic Computing With Loihi: A Survey of Results and Outlook
M. Davies, A. Wild, G. Orchard, Y. Sandamirskaya, G.A.F. Guerra, P. Joshi, P. Plank, S.R. Risbud.
Proceedings of the IEEE, Vol 109(5), pp. 911—934. 2021.
Handbuch der physiologischen Optik
H. von Helmholtz.
Leopold Voss. 1867.
R.P. Rao, D.H. Ballard.
Nature Neuroscience, Vol 2, pp. 79—87. 1999.
DOI: 10.1038/4580
B. Millidge, Y. Song, T. Salvatori, T. Lukasiewicz, R. Bogacz.
arXiv preprint arXiv:2207.12316. 2022.
A survey on neuro-mimetic deep learning via predictive coding [link]
T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.
Neural Networks, Vol 195, pp. 108161. 2026.
DOI: https://doi.org/10.1016/j.neunet.2025.108161
R.P. Rao, D.H. Ballard.
Nature Neuroscience, Vol 2, pp. 79—87. 1999.
DOI: 10.1038/4580
An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity
J.C.R. Whittington, R. Bogacz.
Neural Computation, Vol 29(5), pp. 1229—1262. 2017.
DOI: 10.1162/neco_a_00949
{
F. Innocenti, E.M. Achour, C.L. Buckley.
arXiv preprint arXiv:2505.13124. 2025.
Benchmarking Predictive Coding Networks — Made Simple
L. Pinchetti, C. Qi, O. Lokshyn, G. Olivers, C. Emde, M. Tang, A. M’Charrak, S. Frieder, B. Menzat, R. Bogacz, T. Lukasiewicz, T. Salvatori.
arXiv preprint arXiv:2407.01163. 2025.
On the Infinite Width and Depth Limits of Predictive Coding Networks
F. Innocenti, E.M. Achour, R. Bogacz.
arXiv preprint arXiv:2602.07697. 2026.
C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.
Proceedings of the 43rd International Conference on Machine Learning. 2026.
M.R. Hestenes.
Journal of Optimization Theory and Applications, Vol 4, pp. 303—320. 1969.
A Method for Nonlinear Constraints in Minimization Problems
M.J.D. Powell.
Optimization, pp. 283—298. 1969.
Multiplier Methods: A Survey
D.P. Bertsekas.
Automatica, Vol 12(2), pp. 133—145. 1976.
S. Boyd, N. Parikh, E. Chu, B. Peleato, J. Eckstein.
Foundations and Trends in Machine Learning, Vol 3(1), pp. 1—122. 2011.
DOI: 10.1561/2200000016
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016.
On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance
J. Zeng, S. Lin, Y. Yao, D. Zhou.
Journal of Machine Learning Research, Vol 22. 2021.
A Theoretical Framework for Back-Propagation
Y. LeCun. 1988.
On the Infinite Width and Depth Limits of Predictive Coding Networks
F. Innocenti, E.M. Achour, R. Bogacz.
arXiv preprint arXiv:2602.07697. 2026.
Distributed Optimization of Deeply Nested Systems
M.A. Carreira-Perpinan, W. Wang.
Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, PMLR 33. 2014.
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016.
{ADMM} for Efficient Deep Learning with Global Convergence [link]
J. Wang, F. Yu, X. Chen, L. Zhao.
Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery \& Data Mining, pp. 111–119. Association for Computing Machinery. 2019.
DOI: 10.1145/3292500.3330936
On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance
J. Zeng, S. Lin, Y. Yao, D. Zhou.
Journal of Machine Learning Research, Vol 22. 2021.
A. Gotmare, V. Thomas, J. Brea, M. Jaggi.
ICML 2018 Workshop on Credit Assignment in Deep Learning and Deep Reinforcement Learning. 2018.
T. Frerix, T. Mollenhoff, M. Moeller, D. Cremers.
International Conference on Learning Representations. 2018.
A. Askari, G. Negiar, R. Sambharya, L. El Ghaoui.
arXiv preprint arXiv:1805.01532. 2018.
Lifted Proximal Operator Machines
J. Li, C. Fang, Z. Lin.
Proceedings of the AAAI Conference on Artificial Intelligence, Vol 33, pp. 4181—4188. 2019.
Fenchel Lifted Networks: A {L}agrange Relaxation of Neural Network Training
F. Gu, A. Askari, L. El Ghaoui.
Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108. 2020.
Contrastive Learning for Lifted Networks
C. Zach, V. Estellers.
British Machine Vision Conference (BMVC). 2019.
Lifted {B}regman Training of Neural Networks
X. Wang, M. Benning.
Journal of Machine Learning Research, Vol 24. 2023.
A Unified Framework for Lifted Training and Inversion Approaches
X. Wang, A. Valavanis, A. Mahmood, A. Mang, M. Benning, A. Repetti.
arXiv preprint arXiv:2510.09796. 2025.
B. Evens, P. Latafat, A. Themelis, J. Suykens, P. Patrinos.
2021 60th IEEE Conference on Decision and Control (CDC), pp. 5136–5143. IEEE. 2021.
DOI: 10.1109/cdc45484.2021.9682842
An Augmented Lagrangian Method for Training Recurrent Neural Networks [link]
Y. Wang, C. Zhang, X. Chen.
SIAM Journal on Scientific Computing, Vol 47(1), pp. C22-C51. 2025.
DOI: 10.1137/23M1627614
Y. Song, B. Millidge, T. Salvatori, T. Lukasiewicz, Z. Xu, R. Bogacz.
Nature Neuroscience. 2024.
DOI: 10.1038/s41593-023-01514-1
B. Millidge, M. Tang, M. Osanlouy, N.S. Harper, R. Bogacz.
PLOS Computational Biology, Vol 20(4), pp. e1011183. 2024.
DOI: 10.1371/journal.pcbi.1011183
Learning Complex Temporal Dependencies via Local Synaptic Plasticity
J. Ng-Kee-Kwong, M. Tang, T. Akam, R. Bogacz.
bioRxiv preprint. 2026.
DOI: 10.64898/2026.07.09.737423
Blockwise Self-Supervised Learning at Scale
S.A. Siddiqui, D. Krueger, Y. LeCun, S. Deny.
Transactions on Machine Learning Research. 2024.