The feedforward neural network is an artificial neural network that operates in a single direction, with data being passed from the input layer through one or multiple hidden layers to the output layer without any loops. Each connection in such an architecture has a corresponding weight. Activation functions are used on the neurons’ input values. The network does not have any memory of past inputs and thus generates the same outputs for the same inputs each time.
The term also includes feed-forward neural network (FFNN), and multilayer perceptron (MLP) if there is at least one hidden layer in it. This tutorial explains the architecture, training, evaluation measures, possible failures, and the position of the architecture within modern large language models.
What Will I Learn?
How Data Moves Through the Network: Input, Hidden, and Output Layers
In a feedforward network, there are three kinds of layers: input layers, hidden layers and output layers. The data starts from the input layer and ends at the output layer after going through all the hidden layers.
The input layer keeps the data value of the features without calculating anything. In case of two features’ dataset, the input layer consists of two neurons.
Each hidden layer computes the data in two ways. First, it calculates the weight of the input along with the bias. Afterward, it passes the sum to an activation function.
The output layer generates the prediction output, and the configuration of the output layer is determined based on the following considerations:
- For regression problems, such as predicting house prices, we have one output node with a linear activation function or no activation function at all.
- For binary classification problems, such as determining whether an email is spam or not, we have one output node with a sigmoid activation function generating values between 0 and 1.
- For multi-class classification problems, such as recognizing hand-written digits between 0 to 9, we have one output node for each class with a softmax activation function.
A worked example
Consider a network with two input features, one hidden layer of two neurons using ReLU activation, and one output neuron using sigmoid activation.
Input values: x1 = 0.5, x2 = 0.8.
Hidden neuron 1 receives weights 0.4 and 0.2, plus a bias of 0.1. Its weighted sum is (0.5 × 0.4) + (0.8 × 0.2) + 0.1 = 0.46. ReLU leaves positive values unchanged, so h1 = 0.46.
Hidden neuron 2 receives weights 0.3 and 0.9, plus a bias of −0.2. Its weighted sum is (0.5 × 0.3) + (0.8 × 0.9) − 0.2 = 0.67. ReLU leaves this value unchanged, so h2 = 0.67.
The output neuron receives weights 0.5 and −0.3 from h1 and h2, plus a bias of 0.2. Its weighted sum is (0.46 × 0.5) + (0.67 × −0.3) + 0.2 = 0.229. Sigmoid converts this to a prediction of 0.557.
This example continues in the training section below, where the network compares 0.557 to a true label and updates its weights.
Feedforward Neural Network vs. Multilayer Perceptron
Feedforward neural network means a network in which data flows only in one direction and there are no loops. Multilayer perceptron is a kind of feedforward network where there is at least one hidden layer of neurons. All multilayer perceptrons are feedforward networks. However, not all feedforward networks are multilayer perceptrons; a simple perceptron with connections between inputs and outputs without hidden layers is a feedforward network but not a multilayer perceptron.
The difference is important in one aspect: a network requires at least one hidden layer and a nonlinear activation function to be able to learn nonlinear relations in the data. Cybenko (1989) formulated this property, called the universal approximation theorem, for sigmoidal activation functions, while Hornik (1991) generalized the theorem to a wider class of activation functions.
Activation Functions: Why the Network Needs Them
An activation function refers to a mathematical function that is performed on a neuron’s weighted sum, bringing about non-linearity to the model. The absence of an activation function implies that increasing the number of layers results in a single linear transformation of the input signal.
Four activation functions are commonly used in feedforward networks:
- Sigmoid returns a value between 0 and 1, such that σ(x) = 1 / (1 + e^−x). Sigmoid is used in the output layer for binary classification tasks.
- Tanh returns a value between -1 and 1, such that tanh(x) = (e^x − e^−x) / (e^x + e^−x). It is used in the hidden layer, when outputs must be centered around zero.
- ReLU (Rectified Linear Unit) returns input values when the values are positive, and 0 otherwise, such that ReLU(x) = max(0, x). It is the most popular activation function in the hidden layer, due to computational simplicity and reduction of vanishing gradients.
- Softmax turns the raw values from a vector into a normalized distribution with a sum of one. It is used in the output layer for multi-class classification tasks.
Training a Feedforward Network: Forward Pass and Backpropagation
Training of a feedforward network involves tuning of its parameters (weights and biases) to make predictions more accurate. The procedure consists of two stages: a forward pass that results in a prediction and backpropagation where error is used to update all weights of the network.
Forward propagation
Forward Propagation is illustrated by the computation described above. Data from input travels through all layers. Every neuron performs calculations based on the weights assigned to it and activation function.
Backpropagation
The backpropagation algorithm determines how much each weight affects the prediction error, and then modifies all weights to reduce the error. The backpropagation algorithm as we know it today was introduced by Rumelhart et al. (1986).
Backpropagation includes the following four steps:
- Evaluate the loss between prediction and label.
- Compute the gradient of the loss with respect to the weights in the output layer.
- Compute the gradient of the loss with respect to all previous layers’ weights, by applying the chain rule.
- Update the weights by subtracting the gradient and multiplying by a learning rate.
mm
Continuing the worked example
As per the worked example, we got a prediction of 0.557 with an actual label of 1. The loss for the same in case of binary cross-entropy as the loss function would be 0.585.
Gradient of the loss function with respect to the weighted sum of the output neuron is the difference between the prediction and the actual value: 0.557 – 1 = -0.443.
Now, let’s perform the weight update for the output layer with a learning rate of 0.1. Weight from h1 will go from 0.5 to 0.520. Similarly, the weight from h2 will shift from −0.3 to −0.270, and the output bias from 0.2 to 0.244.
With the new weights, we obtain a new prediction of 0.575 with a loss of 0.553. In one weight update cycle, the loss decreased from 0.585 to 0.553.
Common loss functions:
- Mean squared error (MSE) for regression tasks.
- Binary cross-entropy for binary classification tasks.
- Categorical cross-entropy for multi-class classification tasks.
Common gradient descent variants:
- Batch gradient descent computes gradients on all the examples before updating the weight vector.
- In stochastic gradient descent, updates happen for each training example individually.
- Mini-batch gradient descent computes the gradient based on a small subset of examples and then performs an update.
Evaluating a Feedforward Network
Accuracy, Precision, Recall, F1 Score, and Confusion Matrix are five criteria used to assess how effectively a trained feed-forward network operates on unseen data.
Imagine that a binary classifier is applied to 100 samples and generates 40 true positives, 10 false positives, 5 false negatives, and 45 true negatives.
| Metric | Formula | Value in this example |
|---|---|---|
| Accuracy | (TP + TN) / total | 0.85 |
| Precision | TP / (TP + FP) | 0.80 |
| Recall | TP / (TP + FN) | 0.89 |
| F1 score | 2 × (Precision × Recall) / (Precision + Recall) | 0.84 |
Precision is defined as the proportion of positive predictions that were true. Recall is the percentage of positive examples that the model captured. The F1 score is the combination of both precision and recall.
Feedforward vs. CNN vs. RNN vs. Transformer
Feed-forward networks process individual inputs without any means of memory or spatial relationship. CNNs introduce the element of spatial relationship. RNNs incorporate the element of memory from past inputs. Transformers process the entire sequence of inputs at once through an attention model.
| Architecture | Data type | Memory of past inputs | Typical use case |
|---|---|---|---|
| Feedforward network (FNN) | Fixed-size vectors | None | Tabular data, simple classification |
| Convolutional network (CNN) | Grid-like data | None | Image classification, object detection |
| Recurrent network (RNN) | Sequential data | Step by step | Time series, short text sequences |
| Transformer | Sequential data | Across the full sequence, via attention | Language modeling, translation, long documents |
Feed-forward neural networks can be used for inputs that have a fixed number of features that do not have any spatial or temporal structure like table rows with no connection. Convolutional neural networks can be used for inputs that are in image format or that have a grid structure. Recurrent neural networks or Transformers can be used when the order of the inputs matter.
Feedforward Layers Inside Modern Language Models
Every Transformer layer within a large language model consists of a feedforward sub-layer, and this sub-layer accounts for two-thirds of the total parameters in the layer.
This parameter breakdown was observed by Geva et al., (2021). In a regular Transformer layer, the attention module has four weight matrices that have a dimension of d × d, where d is the hidden dimension of the model. This results in 4d² parameters for the attention module. On the other hand, the feedforward sub-layer has two weight matrices that are usually d × 4d and 4d × d, resulting in 8d² parameters. When 8 is divided by 8 + 4, we get 66%.
Each token is independently fed through this feedforward layer with the same weights being used at all positions along the sequence. Geva et al. (2021) observed that the feedforward layers act like key-value memories where the information is stored during training and accessed according to the input.
Why Feedforward Networks Fail
There are three reasons behind most of the failures during training of feed-forward networks.
The vanishing gradients phenomenon happens when gradients keep on decreasing while propagating back through several layers and become zero for early layers that do not update any further. Sigmoid and tanh activation functions make the problem even worse due to small derivatives (less than 0.25) throughout their domain.
Dead ReLU neurons happen when a neuron’s net input is negative for all the samples in the training set and the neuron’s output is always 0 because it fails to learn. The common reason for such an effect is an overly high learning rate, which results in a very big bias change update.
Overfitting happens when the network learns training data instead of patterns and has high accuracy on training data and poor accuracy on new data. Overfitting is common in dense feedforward layers since every neuron is connected to every other neuron in the adjacent layer and thus the number of parameters grows significantly compared to the number of examples in many training sets.
Srivastava et al. (2014) developed dropout as a solution to this problem. Dropout works by deactivating a percentage of neurons at random at each iteration. The usual dropout rate for the neurons in hidden layers in feedforward neural networks lies within 0.2 and 0.5.
Real-World Applications
The feedforward neural network is used in four fields:
- Tabular classification: credit score models classify individuals for loans based on structured information including their earnings, credit rating, and current debts.
- Computer vision: feedforward neural networks process basic image classification problems, while convolutional neural networks dominate industrial computer vision since they can better scale to large resolution images.
- Natural Language Processing: feedforward neural networks are the sub-networks in every transformer layer discussed above.
- Time series forecasting: fixed window feedforward neural networks forecast one future value from a fixed number of previous values, but recurrent neural networks and transformers dominate in most sequence prediction problems today.
Building a Feedforward Network in Python
Both examples below build the type of network described in the worked example: an input layer, one hidden layer, and an output layer.
TensorFlow and Keras
This example involves the construction, training, and evaluation of a feed-forward network using the MNIST dataset of handwritten digit images consisting of 70,000 gray-scale images of digits from 0 to 9, each 28×28 pixels (LeCun et al.).
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Flatten
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.losses import SparseCategoricalCrossentropy
from tensorflow.keras.metrics import SparseCategoricalAccuracy
mnist = tf.keras.datasets.mnist
(x_train, y_train), (x_test, y_test) = mnist.load_data()
x_train, x_test = x_train / 255.0, x_test / 255.0
model = Sequential([
Flatten(input_shape=(28, 28)),
Dense(128, activation='relu'),
Dense(10, activation='softmax')
])
model.compile(
optimizer=Adam(),
loss=SparseCategoricalCrossentropy(),
metrics=[SparseCategoricalAccuracy()]
)
model.fit(x_train, y_train, epochs=5)
test_loss, test_acc = model.evaluate(x_test, y_test)
print(f"Test accuracy: {test_acc:.4f}")
This network flattens the 28×28 images into a vector of size 784, feeds it to a hidden layer consisting of 128 neurons with ReLU activation function and gives out 10 probabilities by softmax.
PyTorch
This example builds the exact two-input, two-hidden-neuron, one-output network used in the worked example above.
import torch
import torch.nn as nn
import torch.optim as optim
class FeedForwardNet(nn.Module):
def __init__(self, input_size, hidden_size, output_size):
super().__init__()
self.fc1 = nn.Linear(input_size, hidden_size)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(hidden_size, output_size)
self.sigmoid = nn.Sigmoid()
def forward(self, x):
x = self.fc1(x)
x = self.relu(x)
x = self.fc2(x)
x = self.sigmoid(x)
return x
model = FeedForwardNet(input_size=2, hidden_size=2, output_size=1)
criterion = nn.BCELoss()
optimizer = optim.SGD(model.parameters(), lr=0.1)
X = torch.tensor([[0.5, 0.8]])
y = torch.tensor([[1.0]])
for epoch in range(200):
optimizer.zero_grad()
output = model(X)
loss = criterion(output, y)
loss.backward()
optimizer.step()
print(f"Final loss: {loss.item():.4f}")
Feedforward Neural Network FAQs
Q1. Is a feedforward neural network the same as deep learning?
Ans. No. Feedforward neural network is just one of many architectures. Deep learning means any neural network with several hidden layers, which includes feedforward neural network, convolutional network, recurrent network, and Transformer, among others.
Q2. How many layers and neurons should a feedforward network start with?
Ans. For starters, most people use one or two layers and neuron count from the number of neurons in the input layer to the number of neurons in the output layer, then tweak the parameters according to validation results.
Q3. Why does a feedforward network’s loss stop decreasing during training?
Ans. The three most common reasons are setting the learning rate too low or too high, vanishing gradient in a deep network that uses sigmoid and tanh activation functions, and dead ReLU neurons.
Q4. Are Transformers feedforward networks?
Ans. The transformer network consists of an attention module and feedforward sub-module. The latter is a feedforward network, however, the former is not, because the attention module combines the information from several positions in the sequence.
Q5. What is the difference between a feedforward neural network and logistic regression?
Ans. A logistic regression model is mathematically identical to a feedforward neural network with no hidden layer and sigmoid output. The introduction of one or more hidden layers, along with non-linear activation functions, makes the logistic regression a feedforward neural network, which can learn non-linear relations.
Q6. How does dropout prevent a feedforward network from overfitting?
Ans. The dropout technique randomly turns off a percentage of neurons in each training iteration and thereby avoids the network from becoming dependent on one neuron.