反向传播的替代方案:增广拉格朗日预测编码
Backprop Alternative: Augmented Lagrangian Predictive Coding

原始链接: https://pub.sakana.ai/pc-alm/

**PC-ALM(增强拉格朗日预测编码)**是一种用于替代反向传播的新型层级局部学习方法,能够训练超深度神经网络。反向传播需要同步的、多阶段的全局过程,这在生物学上是不合理的,而 PC-ALM 则依赖于局部动力系统。每一层都作为一个独立的反馈控制器,仅与相邻层进行通信以分发监督信号。 该方法通过引入拉格朗日乘子扩展了标准的预测编码(PC)。这些对偶变量充当“积分”控制器,能够累积预测误差,从而即使在深而窄的架构(传统 PC 在这种结构下通常会出现信号衰减)中,也能精确恢复反向传播的信用分配信号。 实验表明,PC-ALM 可以成功训练 1000 层的残差多层感知机(MLP),其性能几乎与反向传播相当。从机制上讲,该方法用“弹道式”信用传播取代了反向传播僵化的前向-后向同步,信息以快速移动波阵面的形式在网络中流动。通过架起分布式优化与生物学信用分配之间的桥梁,PC-ALM 为理解大脑学习机制以及开发用于神经形态硬件的高能效训练方法提供了一条有前景的途径。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 反向传播的替代方案:增广拉格朗日预测编码 (sakana.ai) 14 分,发布者:guld,4 小时前 | 隐藏 | 过往 | 收藏 | 2 条评论 帮助 cs702 5 分钟前 | 下一条 [–] MNIST 上约 90% 的准确率。唉。在 CIFAR-10 上表现如何,或者更好一点,在 ImageNet 上表现如何? 回复 guld 4 小时前 | 上一条 [–] Sakana.ai 的新论文 [1] [1]: https://arxiv.org/abs/2605.31022 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

We introduce PC-ALM, a local alternative to backpropagation. PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics. PC-ALM equips each layer with a feedback control dynamical system that distributes and propagates supervision credit throughout a network.

Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly[1, 2]. How the brain solves the multilayer credit assignment problem without explicit use of backprop remains one of the fundamental unsolved problems in neuroscience (though not without progress[3, 4, 5]).

There are several reasons the brain can't implement exact backpropagation. One is “phase locking”[6, 2]. Backpropagation runs in three phases, in strict order: 1) a forward pass, then 2) a backward pass, then 3) a weight update. A weight update is locked until the forward and backward passes have completed—a neuron in an early layer must hold its activation and wait for the error signal to arrive. The brain has no known mechanism that could enforce such strict timing coordination across an entire network[1].

In this post, we introduce PC-ALM (Augmented Lagrangian Predictive Coding), a method for training networks that replaces the forward and backward passes of backprop with layer-local dynamical systems. Each layer is coupled only to its neighbors. Instead of forward-then-backward, we run each layer forward in time. When run to convergence, the dynamics of the whole system distribute supervision credit signals quickly and accurately across the entire network.

PC-ALM is an extension of standard predictive coding (PC)[7, 8, 9, 10]. PC uses diffusive (i.e. energy-based or "heat flow") coupling between layers. Compared to PC, PC-ALM introduces dual neurons (Lagrange multipliers) per layer, making each layer's local recurrence a PI feedback controller. In the limiting case of linear networks, the dual neurons converge to the exact backprop credit signals, despite using only local computation.

We compare PC-ALM to PC and backprop in a suite of experiments. Local training methods such as PC have historically been difficult to scale. Following the PC literature, we use simple tasks (Fashion-MNIST, CIFAR-10, etc.) and networks such as residual MLPs.

We show that PC-ALM can successfully propagate supervision credit in 1000-layer neural networks, overcoming standard PC's signal decay problem[11] while remaining layer-local.

We focus on deep, small-width networks, a regime in which PC tends to perform poorly.

Ultimately, our motive is to understand how distributed systems (such as the brain) can implement gradient computations without backpropagation. Scientific motivations aside, this research may inform energy-efficient deep learning on neuromorphic hardware, where dynamical systems simulation is cheaper than on GPU[12].

Predictive coding: each layer as a dynamical system

Before explaining PC-ALM, let us first explain PC, interpreting it from a dynamical systems standpoint to emphasize its role as a backprop alternative.

Predictive coding

Predictive coding has its roots in Helmholtz's theories of unconscious perception[13]. Rao & Ballard (1999) developed a mathematical framework for PC as a model of visual cortex[14]. The idea of PC is that each layer attempts to model its incoming signals, sending upward only the prediction error (the part that the layer failed to model) to the next layer.

Mathematically, predictive coding utilizes a general motif: take a state and update it to reduce a prediction error at the next step,

statet+1=statetη(statettargett)prediction error

By applying this update rule to each layer's activation vector (the “state” is the layer's activation hi; its “target” is the prediction σ(Wihi1) arriving from the layer below)1, the PC framework effectively sidesteps backprop's need for a synchronized forward and backward pass.

To explain this in more detail, let's write a feedforward network as a constrained optimization problem:

minimizeθ,h12yWLhL12subject tohi=σ(Wihi1),i=1,,L1.

where L is the network depth, h0:=x the input, y the target, θ={Wi} the weights, hi the layer activations, and σ an activation function such as ReLU. Note that each hi is an optimization variable2. We then construct a new loss function that includes the original supervision loss, together with a quadratic penalty for violations of each layer constraint:

FPC(h,θ)=12yWLhL12+12i=1L1hiσ(Wihi1)2.

This is a quadratic relaxation of the constrained problem. FPC is known as the “free energy” of the network[15, 9].

To train a neural network, PC alternates between inference and learning steps:

Predictive coding

inference

for t=1,,T

hihiηhhiFPCfor i=1,,L1

learning

WiWiηθWiFPCfor i=1,,L

Per mini-batch, a forward pass initializes the activations, followed by T inference steps and a single weight update. We set T proportional to network depth; the 1000-layer experiments below use T=2L.

Each hi-update reduces the prediction errors between layers adjacent to i. This is because hiFPC depends only on hi1, hi, and hi+1. Inference requires only nearest-neighbor communication ("message passing") between layers. Explicitly, writing ri=hiσ(Wihi1) for the prediction error between layers i1 and i, the inference update reads3:

hihiηh(rierror belowWi+1(σri+1error above))i=1,,L1

The bottom h0 is “clamped” (fixed) to an input value and the top of the network is clamped to the target y. Running T update steps, the network settles into states hi for each layer, after which a gradient descent step is taken on the same FPC but now with respect to the weights W (given the current remaining prediction errors and the current state activations).

WiWi+ηθ(σri)hi1i=1,,L1

Both the inference step and the weight update are layer-local. The weight update is Hebbian-like, in that it multiplies a postsynaptic error by the presynaptic activity (a delta rule), and the dynamics map onto a neural circuit with explicit error neurons[14, 7].

PC trains deep networks, but exhibits signal decay

PC inference in a 32-layer residual MLP (width 16, ReLU) at weight initialization. Credit = per-layer norm of the prediction error ri. Dashed reference: norm of the backprop adjoint (the loss gradient with respect to that layer's activations).

Since minimizing the free energy with respect to each hi does not enforce the layer-wise constraints to hold exactly, PC results in a different learning trajectory compared to standard backpropagation.

Nevertheless, PC has been shown to successfully train networks on simple tasks. For example, MNIST and Fashion-MNIST in 128-layer residual MLPs with wide layer widths (512 neurons per layer)[16].

However, PC struggles on more complex tasks and networks[17]. Further, PC struggles even on simple networks/tasks if the network width is smaller than its depth[18].

Each layer adjusts its activity to reduce prediction errors with its neighbors. Supervision enters at the output, but must work its way through this chain of local compromises to influence earlier layers. In deep, narrow networks, the resulting credit signal becomes weak long before it reaches the input.

This leads to a documented signal-decay problem of PC[11], as illustrated in the above figure. Increasing T lets credit propagate farther, but requires more computation for each training update.

Our method, PC-ALM, introduces a way to improve signal propagation of PC networks, retaining the layer-local dynamics of PC and keeping inference budget T proportional to network depth.

Augmented Lagrangian Predictive Coding

We propose Augmented Lagrangian Predictive Coding, a variant of PC that uses the augmented Lagrangian (AL)[19, 20, 21] in place of PC's FPC:

L(h,θ,λ)=12yWLhL12+i=1L1λi(hiσ(Wihi1))+12i=1L1hiσ(Wihi1)2

supervised loss  +  Lagrangian term  +  PC energy

In each layer, the augmented Lagrangian introduces a Lagrange multiplier (or dual variable) λiRni of the same dimension as hi.

The augmented Lagrangian is used extensively in distributed optimization[22], and has motivated many distributed methods for training deep networks[23, 24].

LeCun (1988)[25] showed that the Lagrange multipliers of a constrained network encode its backprop credit signals at equilibrium. The augmented Lagrangian combines this classical construction with the quadratic prediction-error penalties already used by PC. This suggests a simple possibility: can PC’s local dynamics recover those credit signals if we add the multipliers?

To use the augmented Lagrangian for training, we make a simple modification to PC:

Augmented Lagrangian predictive coding

inference

for t=1,,T

hihiηhhiLfor i=1,,L1primal

λiλi+α(hiσ(Wihi1))for i=1,,L1dual

learning

WiWiηθWiLfor i=1,,L

Here α is the dual step size. PC-ALM is primal descent, dual ascent on the augmented Lagrangian, versus PC's descent on the energy.

By accumulating local prediction errors, the dual variables recover the exact backprop credit signals in a deep linear network. We derive this result in the paper.

Thus, at least for linear networks, PC-ALM gives a method for computing exact supervised loss gradients and distributing them throughout a network, using only layer-local dynamics.

In the experiments below, we test whether this advantage carries over to nonlinear networks.

Mechanistic interpretation

To understand, mechanistically, how PC-ALM works, consider a simple scalar network, with hidden unit h=w1x and output y^=w2h. We want to propagate the gradient of the supervised loss 12(yy^)2 to w1.

We attach a multiplier λ to the constraint h=w1x, initialize λ=0, and initialize h at its forward-pass value.

The first step of PC-ALM agrees exactly with a PC step. After that λ accumulates the layer's prediction error each step, which feeds back into the h updates. Each h update now takes a gradient step on the energy with prediction targets shifted by the dual.4

At convergence the activation h returns to its forward-pass value (restoring the constraint h=w1x), while λ has accumulated to the backprop credit signal of that layer: i.e. λ=w2(yy^).

PC vs PC-ALM in a two-layer, linear, scalar network, y^=w2w1x, with x,y clamped to data values. Left: training trajectories of BP, PC, and PC-ALM in weight space. In this simple model, both PC and PC-ALM work and converge to the solution manifold. Performance differences between PC and PC-ALM become apparent in deep-narrow networks.

Control theory and credit assignment

PC-ALM offers a control-theoretic perspective on credit assignment. Each layer combines its current prediction error with an accumulated error signal—the proportional and integral terms of a PI feedback controller.5 Global credit assignment emerges from a network of local feedback controllers. We believe this perspective offers a useful principle for designing new local learning algorithms.

Results

We present below some results for PC-ALM.

PC-ALM: a local learning method for training 1000-layer neural networks

PC-ALM successfully trains 1000-layer MLPs on MNIST. We use the residual MLP setup from Innocenti et al. (2026)[18]. We train for five epochs. Our architecture is a simple MLP with residual skip connections, using weight parameterizations that stabilize backprop training at this depth.

Our results are presented in the figure below; PC-ALM achieves near-BP performance using only layer-local dynamics.

MNIST test accuracy versus depth, from shallow networks through 1000 layers, for BP, PC-ALM, and PC. Backprop is global; PC and PC-ALM are local.
MNIST test accuracy vs depth (width N=32, ReLU activation). PC-ALM stays within ~2 percentage points of backprop across the whole range, including at 1000 layers.

Image classification benchmarks

We tested PC-ALM on a small set of image classification tasks and found that it improves performance over PC on every task, including when training ResNet-18 on CIFAR-10 and Tiny ImageNet.

Test accuracy for BP, PC, and PC-ALM across MNIST and Fashion-MNIST at 32, 64, and 128 layers, plus CIFAR-10 and Tiny ImageNet.
Across benchmarks. PC-ALM consistently narrows the gap between standard PC and global backprop. As depth increases, PC’s accuracy falls off much faster than PC-ALM’s.

PC-ALM propagation dynamics

In addition to the performance of PC-ALM across image recognition tasks, we find PC-ALM exhibits surprising dynamical properties. In both PC and PC-ALM inference, movement of the hidden activations hi initially appears only in the final network layers. As inference proceeds, earlier layers receive credit signal, forming a wavefront that advances toward the input. PC-ALM’s primal-dual dynamics drive this wavefront through the network faster than PC. We call this "ballistic" credit propagation to contrast with PC's diffusive, heat-flow-like propagation.

Ballistic vs diffusive credit propagation. Credit magnitude across layers during inference. PC's credit decays with depth; PC-ALM's spreads evenly across the network.

Stable oscillatory transient responses

Individual neurons in PC-ALM also exhibit damped oscillations during inference.

Oscillatory inference dynamics. We consider a deep linear network (σ=identity) with depth L=8 and width N=1. With weights fixed, the coupled inference dynamics across all layers form a linear system, characterized by the eigenvalues shown on the left (Equation (17) of the paper). Right: a neuron’s activity and dual variable during inference. The parameter α is the dual step size. Setting α=0 recovers PC exactly, with all eigenvalues real; increasing α introduces complex eigenvalues and oscillatory dynamics. Increasing α too far eventually destabilizes the network (Equation (14) of the paper).

Discussion

This work introduces PC-ALM, a layer-local alternative to backpropagation for training deep networks. To our knowledge, this is the first layer-local method shown to successfully train networks up to 1000 layers.

Summary

Predictive coding remains an attractive candidate for a theory of cortical function. It is rooted in Helmholtz's ideas on unconscious processing. It was later shown that predictive coding can be viewed as a layer-local alternative to backpropagation for training deep networks. PC in its standard form can successfully train networks on simple tasks, but signal decay limits its performance in deep, narrow networks. By introducing Lagrange multipliers and running primal-dual inference on the augmented Lagrangian (vs PC's gradient flow on an energy), PC-ALM lets individual layers compute a gradient signal of a global loss function, using only communication between neighboring layers. The primary motivation of this work is to understand more deeply the mechanisms of credit assignment in real physical systems such as the brain.

Constrained and lifted optimization

PC-ALM relates to a long line of work on constrained optimization approaches to training deep networks. These approaches “lift” training into a larger optimization problem by treating the activations h as variables alongside the weights W. This includes the method of auxiliary coordinates[26], ADMM training of deep networks[23, 27, 24], BlockProp[28], ProxProp[29], several other variants[30, 31, 32, 33, 34, 35], and augmented Lagrangian methods for network training[36, 37].

PC-ALM draws on this optimization lineage to address a question from neuroscience: how can local neural dynamics compute and distribute credit for a global objective?

Prospective-configuration tradeoff

In PC, the settled activations differ from the forward pass. Song et al.[38] called this prospective configuration and showed that it can improve sample efficiency over backprop. PC-ALM improves credit propagation but gives up prospective configuration at convergence. We suspect a fundamental tradeoff between the two, with intermediate settings of T, α, and dual leak potentially preserving the benefits of prospective configuration while substantially improving credit propagation.

Motivations

This work began with three observations:

1

The Neuro-AI and distributed optimization communities share a concern with locality, but there has been relatively little cross-talk between them.

2

The PC energy is suspiciously similar to the augmented term of the augmented Lagrangian that is commonly used in distributed and constrained/lifted optimization.

3

LeCun (1988) identified the multipliers of the standard Lagrangian with backprop credit signals, suggesting that these multipliers could serve as local credit signals.

Our work on PC-ALM ties these threads together.

Future work

There is a wealth of work that needs to be done. Notably, extending PC-ALM to temporal tasks with temporal credit assignment[39, 40], self-supervised losses[41], and of course larger networks and more difficult tasks.

We hope this work inspires further connections between augmented Lagrangian methods and biologically plausible credit assignment.

Backpropagation and the Brain
T.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.
Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020.
DOI: 10.1038/s41583-020-0277-3

Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment
A. Ororbia.
arXiv preprint arXiv:2312.09257. 2023.

Dendritic cortical microcircuits approximate the backpropagation algorithm[PDF]
J. Sacramento, R.P. Costa, Y. Bengio, W. Senn.
Advances in Neural Information Processing Systems, Vol 31. 2018.

‘Backpropagation and the brain’ realized in cortical error neuron microcircuits[link]
K. Max, I. Jaras, A. Granier, K.A. Wilmes, M.A. Petrovici.
PLOS Computational Biology, Vol 22(4), pp. e1014164. 2026.
DOI: 10.1371/journal.pcbi.1014164

Backpropagation through space, time and the brain[link]
B. Ellenberger, P. Haider, F. Benitez, J. Jordan, K. Max, I. Jaras, L. Kriener, M.A. Petrovici.
Nature Communications, Vol 17, pp. 66. 2026.
DOI: 10.1038/s41467-025-66666-z

Decoupled Neural Interfaces using Synthetic Gradients
M. Jaderberg, W.M. Czarnecki, S. Osindero, O. Vinyals, A. Graves, D. Silver, K. Kavukcuoglu.
International Conference on Machine Learning (ICML), Vol 70, pp. 1627—1635. 2017.

Brain-Inspired Machine Intelligence: A Survey of Neurobiologically-Plausible Credit Assignment
A. Ororbia.
arXiv preprint arXiv:2312.09257. 2023.

Backpropagation and the Brain
T.P. Lillicrap, A. Santoro, L. Marris, C.J. Akerman, G. Hinton.
Nature Reviews Neuroscience, Vol 21(6), pp. 335—346. 2020.
DOI: 10.1038/s41583-020-0277-3
An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity
J.C.R. Whittington, R. Bogacz.
Neural Computation, Vol 29(5), pp. 1229—1262. 2017.
DOI: 10.1162/neco_a_00949

Predictive Coding: A Theoretical and Experimental Review
B. Millidge, A. Seth, C.L. Buckley.
arXiv preprint arXiv:2107.12979. 2021.

A survey on neuro-mimetic deep learning via predictive coding[link]
T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.
Neural Networks, Vol 195, pp. 108161. 2026.
DOI: https://doi.org/10.1016/j.neunet.2025.108161

Learning on Arbitrary Graph Topologies via Predictive Coding
T. Salvatori, L. Pinchetti, B. Millidge, Y. Song, T. Bao, R. Bogacz, T. Lukasiewicz.
Advances in Neural Information Processing Systems. 2022.

{ePC}: Fast and Deep Predictive Coding in Digital Simulation[link]
C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.
Proceedings of the 43rd International Conference on Machine Learning. 2026.

Advancing Neuromorphic Computing With Loihi: A Survey of Results and Outlook
M. Davies, A. Wild, G. Orchard, Y. Sandamirskaya, G.A.F. Guerra, P. Joshi, P. Plank, S.R. Risbud.
Proceedings of the IEEE, Vol 109(5), pp. 911—934. 2021.

Handbuch der physiologischen Optik
H. von Helmholtz.
Leopold Voss. 1867.

Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
R.P. Rao, D.H. Ballard.
Nature Neuroscience, Vol 2, pp. 79—87. 1999.
DOI: 10.1038/4580
A Theoretical Framework for Inference and Learning in Predictive Coding Networks
B. Millidge, Y. Song, T. Salvatori, T. Lukasiewicz, R. Bogacz.
arXiv preprint arXiv:2207.12316. 2022.

A survey on neuro-mimetic deep learning via predictive coding[link]
T. Salvatori, A. Mali, C.L. Buckley, T. Lukasiewicz, R.P. Rao, K. Friston, A. Ororbia.
Neural Networks, Vol 195, pp. 108161. 2026.
DOI: https://doi.org/10.1016/j.neunet.2025.108161

Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects
R.P. Rao, D.H. Ballard.
Nature Neuroscience, Vol 2, pp. 79—87. 1999.
DOI: 10.1038/4580

An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity
J.C.R. Whittington, R. Bogacz.
Neural Computation, Vol 29(5), pp. 1229—1262. 2017.
DOI: 10.1162/neco_a_00949

{μ}{PC}: Scaling Predictive Coding to 100+ Layer Networks
F. Innocenti, E.M. Achour, C.L. Buckley.
arXiv preprint arXiv:2505.13124. 2025.

Benchmarking Predictive Coding Networks — Made Simple
L. Pinchetti, C. Qi, O. Lokshyn, G. Olivers, C. Emde, M. Tang, A. M’Charrak, S. Frieder, B. Menzat, R. Bogacz, T. Lukasiewicz, T. Salvatori.
arXiv preprint arXiv:2407.01163. 2025.

On the Infinite Width and Depth Limits of Predictive Coding Networks
F. Innocenti, E.M. Achour, R. Bogacz.
arXiv preprint arXiv:2602.07697. 2026.

{ePC}: Fast and Deep Predictive Coding in Digital Simulation[link]
C. Goemaere, G. Oliviers, R. Bogacz, T. Demeester.
Proceedings of the 43rd International Conference on Machine Learning. 2026.
Multiplier and Gradient Methods
M.R. Hestenes.
Journal of Optimization Theory and Applications, Vol 4, pp. 303—320. 1969.

A Method for Nonlinear Constraints in Minimization Problems
M.J.D. Powell.
Optimization, pp. 283—298. 1969.

Multiplier Methods: A Survey
D.P. Bertsekas.
Automatica, Vol 12(2), pp. 133—145. 1976.

Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, J. Eckstein.
Foundations and Trends in Machine Learning, Vol 3(1), pp. 1—122. 2011.
DOI: 10.1561/2200000016
Training Neural Networks Without Gradients: A Scalable {ADMM} Approach
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016.

On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance
J. Zeng, S. Lin, Y. Yao, D. Zhou.
Journal of Machine Learning Research, Vol 22. 2021.

A Theoretical Framework for Back-Propagation
Y. LeCun. 1988.

On the Infinite Width and Depth Limits of Predictive Coding Networks
F. Innocenti, E.M. Achour, R. Bogacz.
arXiv preprint arXiv:2602.07697. 2026.

Distributed Optimization of Deeply Nested Systems
M.A. Carreira-Perpinan, W. Wang.
Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, PMLR 33. 2014.

Training Neural Networks Without Gradients: A Scalable {ADMM} Approach
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, T. Goldstein.
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48. 2016.

{ADMM} for Efficient Deep Learning with Global Convergence[link]
J. Wang, F. Yu, X. Chen, L. Zhao.
Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery \& Data Mining, pp. 111–119. Association for Computing Machinery. 2019.
DOI: 10.1145/3292500.3330936

On {ADMM} in Deep Learning: Convergence and Saturation-Avoidance
J. Zeng, S. Lin, Y. Yao, D. Zhou.
Journal of Machine Learning Research, Vol 22. 2021.

Decoupling Backpropagation using Constrained Optimization Methods[link]
A. Gotmare, V. Thomas, J. Brea, M. Jaggi.
ICML 2018 Workshop on Credit Assignment in Deep Learning and Deep Reinforcement Learning. 2018.
Proximal Backpropagation[link]
T. Frerix, T. Mollenhoff, M. Moeller, D. Cremers.
International Conference on Learning Representations. 2018.
Lifted Neural Networks
A. Askari, G. Negiar, R. Sambharya, L. El Ghaoui.
arXiv preprint arXiv:1805.01532. 2018.

Lifted Proximal Operator Machines
J. Li, C. Fang, Z. Lin.
Proceedings of the AAAI Conference on Artificial Intelligence, Vol 33, pp. 4181—4188. 2019.

Fenchel Lifted Networks: A {L}agrange Relaxation of Neural Network Training
F. Gu, A. Askari, L. El Ghaoui.
Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108. 2020.

Contrastive Learning for Lifted Networks
C. Zach, V. Estellers.
British Machine Vision Conference (BMVC). 2019.

Lifted {B}regman Training of Neural Networks
X. Wang, M. Benning.
Journal of Machine Learning Research, Vol 24. 2023.

A Unified Framework for Lifted Training and Inversion Approaches
X. Wang, A. Valavanis, A. Mahmood, A. Mang, M. Benning, A. Repetti.
arXiv preprint arXiv:2510.09796. 2025.

Neural Network Training as an Optimal Control Problem : — An Augmented Lagrangian Approach —[link]
B. Evens, P. Latafat, A. Themelis, J. Suykens, P. Patrinos.
2021 60th IEEE Conference on Decision and Control (CDC), pp. 5136–5143. IEEE. 2021.
DOI: 10.1109/cdc45484.2021.9682842

An Augmented Lagrangian Method for Training Recurrent Neural Networks[link]
Y. Wang, C. Zhang, X. Chen.
SIAM Journal on Scientific Computing, Vol 47(1), pp. C22-C51. 2025.
DOI: 10.1137/23M1627614

Inferring Neural Activity Before Plasticity as a Foundation for Learning Beyond Backpropagation
Y. Song, B. Millidge, T. Salvatori, T. Lukasiewicz, Z. Xu, R. Bogacz.
Nature Neuroscience. 2024.
DOI: 10.1038/s41593-023-01514-1
Predictive Coding Networks for Temporal Prediction
B. Millidge, M. Tang, M. Osanlouy, N.S. Harper, R. Bogacz.
PLOS Computational Biology, Vol 20(4), pp. e1011183. 2024.
DOI: 10.1371/journal.pcbi.1011183

Learning Complex Temporal Dependencies via Local Synaptic Plasticity
J. Ng-Kee-Kwong, M. Tang, T. Akam, R. Bogacz.
bioRxiv preprint. 2026.
DOI: 10.64898/2026.07.09.737423

Blockwise Self-Supervised Learning at Scale
S.A. Siddiqui, D. Krueger, Y. LeCun, S. Deny.
Transactions on Machine Learning Research. 2024.

联系我们 contact @ memedata.com