A ReLU activation produces the input value itself whenever the input value is positive and zero whenever the input value is zero or negative. This straightforward operation, described by the formula f(x) = max(0, x), ensures that ReLU serves as the de facto choice of activation function for the hidden layers of convolutional neural networks and multilayer perceptrons in 2026. Language models based on transformer architectures constitute an exception, with most using GELU or SwiGLU in their feed-forward layers.
What Will I Learn?
What Is the ReLU Activation Function?
The Rectified Linear Unit (ReLU) is an activation function employed by artificial neural networks and is piecewise linear. It returns the input value if that value is greater than zero, and returns zero otherwise. The role of an activation function in a neural network layer is to decide if the neuron will propagate its signal at all, and if so, then to what extent. There are various neural network layers that make use of the ReLU activation function, like convolutional layers and fully connected hidden layers.
Agentic AI Course
Average time:6 month(s) + Lifetime Access
Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving
The ReLU Formula, Explained
Imagine ReLU to be a regular light switch with an in-built dimmer such that if the value of x is either zero or negative, then the light switch remains off, while if the value of x becomes positive, then the switch becomes fully open to allow the signal to flow freely without any obstruction, i.e., f(x) = max(0, x), which can also be written as:
- f(x) = x for positive inputs (x > 0)
- f(x) = 0 for zero or negative inputs (x ≤ 0)
Due to the two lines crossing each other at x = 0, the derivative or the rate of change becomes similarly straightforward. The slope remains +1 when the signal is positive, allowing all the feedback to pass through backpropagation, while when the signal is negative, the slope remains 0, stopping any feedback from going through. In the precise point where the lines meet, the slope remains 0 by default.
What ReLU Actually Looks Like
In the graph of the ReLU activation function, there are two straight lines which come together at the origin. In case the input value is below zero, the output line will be a horizontal line at the value of zero. In case the input value is greater than zero, the output line will have a positive slope of one. This one bend is the only place where the graph is not linear.
A Brief, Correct History of ReLU
Threshold-linear units were developed for neuroscience applications. The circuit based on a threshold-linear unit was presented in a paper “Digital Selection and Analog Amplification Coexist in a Cortex-Like Neural Network” authored by Rainer Hahnloser, Rahul Sarpeshkar, Matthew Mahowald, Rodney Douglas, and H. Sebastian Seung and published in Nature in 2000. In 2010, the rectified linear version of the unit was used in restricted Boltzmann machines and introduced to the public as “Rectified Linear Units Improve Restricted Boltzmann Machines” at ICML.
The concept was extended to deep supervised networks in 2011 by Glorot, Bordes, and Bengio. They published their research paper entitled “Deep Sparse Rectifier Neural Networks” in the AISTATS conference. According to their research, neural networks formed using rectified linear neurons can be trained without unsupervised pre-training and form sparse internal representations such that some units of the network will output zero.
Following 2012, the Rectified Linear Unit (ReLU) was the preferred unit for computer vision. This is because Krizhevsky et al. in 2012 adopted it in the architecture of the convolutional neural network known as AlexNet to win the ImageNet Large Scale Visual Recognition Challenge. According to their study, a four-layer network with ReLU units achieved a 25% training error in significantly less time compared to one with tanh units.
Parametric ReLU and He initialization were introduced in 2015 by He et al. through “Delving Deep into Rectifiers” as techniques aimed at stabilizing deep ReLU network training. Since 2017, Transformer architectures stopped using standard ReLU in feed-forward networks, replacing it with either GELU (introduced by Hendrycks et al., 2016), or SwiGLU (introduced by Shazeer, 2020).
Why ReLU Became the Default Choice
ReLU offers four measurable advantages over earlier activation functions such as sigmoid and tanh.
- Computational Simplicity. ReLU requires just one comparison per input. On the other hand, sigmoid and tanh both involve an expensive exponential computation.
- Sparse Activation. All inputs that are below zero will result in an output of exactly zero; thus, a subset of neurons within the network will be idle for a particular input.
- Reduced vanishing-gradient risk. For ReLU, its derivative is equal to 1 for all positive inputs; hence the gradient does not decrease while passing backward through active units. In contrast, the derivatives of sigmoid and tanh both approach zero for larger inputs.
- Faster empirical convergence. According to the AlexNet paper (Krizhevsky et al., 2012), a much faster empirical convergence was observed with ReLU compared to tanh.
ReLU vs. Sigmoid vs. Tanh
| Function | Formula | Output range | Derivative range | Common use |
| ReLU | f(x) = max(0, x) | [0, ∞) | 0 or 1 | CNN and MLP hidden layers |
| Sigmoid | f(x) = 1 / (1 + e^-x) | (0, 1) | (0, 0.25] | Binary classification output layer |
| Tanh | f(x) = (e^x – e^-x) / (e^x + e^-x) | (-1, 1) | (0, 1] | Recurrent network gates (LSTM, GRU) |
The Sigmoid activation function is the choice for binary classifiers since its range (0,1) directly translates to the value of a probability. The Tanh activation function is used inside the LSTM and GRU gates since its zero-centered range of (-1,1) balances the value of the gates around zero.
The Dying ReLU Problem (and How to Fix It)
The problem of dying ReLU takes place when a neuron has an output of zero on all examples in the dataset and does not get any updates while performing backpropagation. It takes place because of a large negative update of weights which makes the pre-activation of a unit go to a negative value for each of the training examples, resulting in 0 values of the gradients and, thus, no further updates of weights. Karpathy states in his CS231n lectures at Stanford University that a large percentage of ReLUs in a neural network, even up to 40%, may become inactive with a very high learning rate.
Four corrective steps address the dying ReLU problem:
- Decrease the learning rate in order to decrease the step size of each weight update.
- Use He initialization (He et al., 2015) to adjust the initial weight variance to match that of the ReLU network.
- Substitute the affected units with either the Leaky ReLU or Parametric ReLU, which produce a non-zero output even if the input is negative.
- Introduce batch normalization prior to applying the activation function in order to keep the values of the pre-activation layer stable and centered around zero.
The Other Limitations of ReLU
- Unbounded output. In contrast with other activation functions, ReLU imposes no ceiling on the outputs for positive numbers. It allows arbitrary increases of output values throughout the network as the network becomes deeper without normalization layers, leading to numerical instability.
- Non-zero mean output. ReLU is always generating a non-negative output, making weight corrections in the next layer biased in one direction only, resulting in slower convergence compared to zero-mean activation function like tanh.
- Loss of information for negative input values. All negative input values are mapped to zero regardless of the size of their input value, making it impossible to use this part of information.
ReLU’s Variants: Leaky ReLU, PReLU, and ELU
Three variants address ReLU’s dying-unit and non-zero-centered output problems by modifying its behavior for negative inputs.
| Variant | Formula | Negative-input behavior | Source |
| Leaky ReLU | f(x) = x if x > 0; αx if x ≤ 0 (α ≈ 0.01) | Small, fixed negative slope | Maas, Hannun, and Ng, 2013 |
| Parametric ReLU (PReLU) | f(x) = x if x > 0; αx if x ≤ 0 (α learned) | Small, learned negative slope | He et al., 2015 |
| Exponential Linear Unit (ELU) | f(x) = x if x > 0; α(e^x − 1) if x ≤ 0 | Smooth curve approaching −α | Clevert, Unterthiner, and Hochreiter, 2015 |
In leaky ReLU and PReLU, a non-zero gradient is maintained when input is less than zero, so that the units don’t become dead. ELU uses an exponential function instead of a straight line at zero to maintain the output values close to zero centered and thus faster convergence, but it has higher computational complexity than ReLU.
Agentic AI Course
Average time:6 month(s) + Lifetime Access
Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving
Is ReLU Still Relevant in the Age of Transformers?
ReLU activation is still being used as the go-to activation function for hidden layers of convolutional neural networks and multilayer perceptrons in 2026, yet GELU and SwiGLU activation functions are mostly preferred by feed-forward networks in language Transformer-based models. The original Transformer model (Vaswani et al., 2017) used ReLU activation in its feed-forward network. Other models gave up on using ReLU: BERT (Devlin et al., 2018) used GELU, while most later language models used SwiGLU (Shazeer, 2020), which is based on the Swish function (Ramachandran, Zoph and Le, 2017).
The reason is that the feedforward neural network in the Transformer needs smooth and non-zero gradient at around x = 0, which GELU and SwiGLU can deliver but not ReLU. However, it does not work for convolutional networks and MLP because they lack this advantage of ReLU.
How to Choose the Right Activation Function
Activation functions depend on the type of layer that is to be created, not on popularity. The standard for convolutional as well as MLP hidden layers is ReLU. Transformer feed forward layers use GELU or SwiGLU. Sigmoid is used in binary classification output layers. Softmax is used in multi-class classification output layers.
How to Implement ReLU in Code
The following five examples compute the same ReLU operation across the frameworks most commonly used in 2026.
NumPy
import numpy as np
def relu(x):
return np.maximum(0, x)
x = np.array([-3, -1, 0, 1, 3])
print(relu(x)) # [0 0 0 1 3]
Native Python
def relu(x):
return max(0, x)
values = [-3, -1, 0, 1, 3]
print([relu(v) for v in values]) # [0, 0, 0, 1, 3]
PyTorch
import torch
import torch.nn as nn
relu = nn.ReLU()
x = torch.tensor([-3.0, -1.0, 0.0, 1.0, 3.0])
print(relu(x)) # tensor([0., 0., 0., 1., 3.])
TensorFlow / Keras
import tensorflow as tf
x = tf.constant([-3.0, -1.0, 0.0, 1.0, 3.0])
print(tf.keras.activations.relu(x)) # [0. 0. 0. 1. 3.]
JAX
import jax.numpy as jnp
from jax.nn import relu
x = jnp.array([-3.0, -1.0, 0.0, 1.0, 3.0])
print(relu(x)) # [0. 0. 0. 1. 3.]
Frequently Asked Questions
Q1. Is ReLU still used in 2026?
Ans. Yes. The Rectified Linear Unit is the default activation function for convolutional neural networks and multilayer perceptrons in 2026. However, language models based on the Transformer model usually do not use the ReLU function.
Q2. What does ReLU stand for?
Ans. The acronym ReLU means Rectified Linear Unit and represents a piecewise linear activation function with the formula f(x) = max(0, x).
Q3. Why is ReLU not differentiable at zero?
Ans. There is no unique value for the derivative at x = 0 because the ReLU function has two different slopes around x = 0 — 0 slope when x < 0 and 1 slope when x > 0. Deep learning libraries provide a fixed subgradient equal to 0 to keep backpropagation uninterrupted.
Q4. What is the difference between ReLU and Leaky ReLU?
Ans. ReLU returns zero for every negative value input; whereas Leaky ReLU returns a small value (for example 0.01 times the input) and not zero, to prevent dying neurons.
Q5. Does batch normalization fix the dying ReLU problem?
Ans. Batch normalization decreases the probability of having dying ReLU neurons, since the distribution of inputs to each layer is standardized.
Q6. Is ReLU used in the output layer of a neural network?
Ans. ReLU is hardly ever applied to output neurons. For binary output, use sigmoid; for multi-class classification, use softmax; and for regression, use linear activation.