Back-propagation algorithm
This lecture covers the back-propagation algorithm, which is a method for training fully connected neural networks. This class of neural networks usually consists of multiple layers. The first layer is called the input layer, and the last layer is called the output layer. The layers in between are called hidden layers. Each layer consists of multiple neurons connected to neurons in the previous and next layers. The connections between neurons have weights that determine how much influence a neuron has on the next layer. These neural networks are called multilayer perceptrons (MLPs) or feedforward neural networks.
Training a neural network involves finding the optimal weights that minimize the error between the predicted output and the actual output. The back-propagation algorithm consists of two main steps: the forward pass and the backward pass. In the forward pass, the input data is passed through the network to compute the output using the current weights. Consequently, error signals are computed. In the backward pass, the error signals are propagated back through the network to update the weights.

The forward pass computes the network output and loss; the backward pass propagates derivatives toward the input layers.
Multilayer perceptrons have three main features:
- Every neuron has a nonlinear activation function, which allows the network to learn complex patterns in the data.
- The network consists of multiple hidden layers that are not directly connected to the input or output layers.
- The network has a high degree of connectivity between layers, which allows it to learn a large number of parameters.
Denotation
The back-propagation algorithm can be mathematically described using the following notation:
- Indices , , and denote neurons, where neuron follows neuron , and neuron follows neuron .
- Index denotes the training example.
- represents the sum of squared errors for the -th training example and is called the error energy.
- represents the average error energy over all training examples.
- represents the error signal for neuron for the -th training example.
- represents the output of neuron for the -th training example.
- represents the desired output for neuron for the -th training example.
- represents the weight of the connection from neuron to neuron .
- represents the change in the weight of the connection from neuron to neuron .
- is the induced local field of neuron for the -th training example, i.e., the weighted sum of the inputs to neuron and its bias.
- represents the activation function of neuron .
- is the bias of neuron .
- represents the input to neuron for the -th training example.
- represents the output of neuron for the -th training example.
- is the learning rate, which determines how much the weights are updated during training.
- is the dimension (number of nodes) of the -th layer (); thus, is the number of input nodes, and is the number of output nodes.
- is the depth of the network.
Common schema
The error signal of the output layer is
The error energy of the output layer is
and the average error energy over all training examples is
As the error energy is a function of the weights, the average error energy is also a function of the weights. Thus, the average error energy for the current training set is a cost function that measures training effectiveness.
The induced local field of neuron is defined as the weighted sum of the inputs to neuron and its bias:
where is the number of input nodes to neuron .
The functional signal of neuron is defined as its output:
The change in the weight of the connection from neuron to neuron can be defined using the chain rule as
where is the learning rate, which determines how much the weights are updated during training. The negative sign indicates that we want to minimize the error energy by updating the weights in the direction of the negative gradient.
Partial derivatives in the above equation can be calculated as follows:
Thus, the change in the weight of the connection from neuron to neuron can be calculated as
where
is called the local gradient of neuron for the -th training example. The local gradient is a measure of how much the error energy changes with respect to the induced local field of neuron .
Notice that it is possible to distinguish two cases for the local gradient of neuron :
- If neuron is an output neuron, then the local gradient can be calculated using the corresponding desired output.
- If neuron is a hidden neuron, then the local gradient can be calculated using the local gradients of the neurons in the next layer and the weights of the connections from neuron to those neurons. This is because the error energy is a function of the outputs of all neurons in the network, and the output of neuron affects the outputs of all neurons in the next layer. This credit-assignment problem is solved by the back-propagation algorithm.
The output layer
For the output layer, the local gradient can be calculated using the corresponding desired output as follows:
where is the error signal for neuron for the -th training example, and is the derivative of the activation function for neuron with respect to the induced local field.
The hidden layer
For the hidden layer, denoted by index , the local gradient can be calculated using the local gradients of the neurons in the next layer and the weights of the connections from neuron to the neurons in the next layer.
First, the local gradient of neuron can be calculated using the chain rule as follows:
Let us denote an output neuron by the index . The error energy can then be calculated as follows:
where is the error signal for output neuron for the -th training example.
Thus, the partial derivative of the error energy with respect to the output of neuron can be calculated as follows:
where is the error signal for output neuron for the -th training example. Consequently, the partial derivative of the error signal with respect to the induced local field of output neuron can be calculated as follows:
Also,
Substituting the equations above into the equation for gives
Finally, the local gradient of neuron can be calculated as follows:
The equation above is called the back-propagation formula. It is used to calculate the local gradient of a hidden neuron from the local gradients of the neurons in the next layer and the weights of the connections between them.
Forward pass
The forward pass of the back-propagation algorithm involves passing the input data through the network to compute the output using the current weights.
Assume that the input vector for the -th training example is , where is the number of input nodes. The output of the input layer is simply the input vector itself, i.e., for .
Next, we apply the activation function:
If neuron belongs to the output layer, then the output of the network for the -th training example is for , where is the number of output nodes. The error signal for an output neuron is calculated using the desired output: for .
Backward pass
The backward pass begins at the output layer, presenting it with an error signal. This signal is passed from right to left from layer to layer, with the local gradient calculated in parallel for each neuron. This recursive process involves modifying synaptic weights according to the delta rule. For a neuron located in the output layer, the local gradient is equal to the corresponding error signal multiplied by the first derivative of the nonlinear activation function.
The quantity is then used to calculate changes in the weights associated with the output layer. Knowing the local gradients for all neurons in the output layer, we can calculate the local gradients of all neurons in the previous layer and, therefore, the adjustments to the weights associated with that layer. These calculations are performed for all layers in reverse order.
However, the back-propagation algorithm itself does not specify how to update the weights, and it can be used with different optimization algorithms to find the optimal weights that minimize the error energy. Several optimization algorithms will be covered in the next lecture.
References
- Simon Haykin, Neural Networks and Learning Machines, 3rd ed., Pearson, 2009.
- David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams, “Learning Representations by Back-Propagating Errors,” Nature 323, 533–536, 1986.
- Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep Learning, MIT Press, 2016.