Softmax Activation is an activation function that takes in a vector of raw outputs from a neural network, known as logits, and turns them into probabilities. Each individual value ranges from 0 to 1, with all the values adding up to 1. Softmax is used in neural networks as the last layer of multi-class classification tasks, with each neuron representing the probability of a single, unique class. This tutorial explains the formula, provides a numerical example, shows four ways of implementing it using Python, lists common implementation errors, and the use of softmax in Transformer attention layers.
What Will I Learn?
What Is the Softmax Activation Function?
The softmax activation function is a mathematical function that takes an array of numbers and transforms them into probabilities that add up to 1.
While the neural network’s last layer produces logits, softmax assigns a proportion to each logit depending on the sum of all values. Various activation functions such as ReLU, sigmoid, and tanh functions can be used in different neural network layers, but none of them behave like the softmax function. The difference between other activation functions and the softmax function is that the softmax activation function generates the probability distribution for multiple classes instead of one value transformation.
The Softmax Formula, Step by Step
The formula for the softmax function involves dividing the exponential of the logit by the summation of the exponentials of all logits in the vector.
Given logit vector z = [z₁, z₂,…, zₙ], the softmax function is:
This computation is done in four stages.
- Start off with the unprocessed output vector z from the last layer of the network.
- Next, take e (around 2.718) and raise it to the power of each number in z.
- Sum all the resulting exponents to create a sum total.
- Then divide each number by the sum to obtain the probability.
The table below shows this calculation for a four-class image classifier with logits [4.2, 1.3, -0.5, 2.1].
| Class | Logit (z) | e^z | Probability |
|---|---|---|---|
| Cat | 4.2 | 66.69 | 84.3% |
| Dog | 1.3 | 3.67 | 4.6% |
| Bird | -0.5 | 0.61 | 0.8% |
| Fish | 2.1 | 8.17 | 10.3% |
The four probabilities sum to 100%. The class with the highest logit, Cat at 4.2, receives the highest probability, 84.3%.
Softmax in Action — A Visual Walkthrough
Logits do not have any fixed range. They can take on positive, negative, or zero values, and the interpretation of their difference is not straightforward. The softmax outputs range between 0 and 1 and add up to 1 exactly, making them interpretable in a direct manner as the confidence scores. The logit difference of 2.1 between Cat (4.2) and Fish (2.1) results in the 8 times higher probability ratio (84.3% compared to 10.3%) due to the exponential function nature.
Why Neural Networks Use Softmax in the Output Layer
Neural networks utilize the softmax function in the output layer since it gives out probability distribution in mutually exclusive classes and combines well with cross-entropy loss function.
A problem that is described by mutually exclusive classification means that each input is associated with only one class. Classification of images into cats, dogs, birds, and fishes can be considered an example of a mutually exclusive classification problem. The softmax function makes sure that the neural network output is probability values which means that the output values will have range 0 – 1 and will sum up to 1. This enables the neural network to calculate the difference between the predicted probability distribution and actual class using cross-entropy loss during the training process. The formula for categorical cross-entropy loss is: L = −Σᵢ yᵢ log(pᵢ).
Softmax vs. Sigmoid: What’s the Real Difference
The sigmoid function generates a separate probability value for one class, whereas the softmax function generates several interdependent probability values that total to 1.
| Property | Sigmoid | Softmax |
|---|---|---|
| Input | One scalar value | A vector of n values |
| Output | One probability | n probabilities that sum to 1 |
| Class relationship | Independent | Interdependent |
| Typical use | Binary or multi-label classification | Multi-class, mutually exclusive classification |
| Loss function pairing | Binary cross-entropy | Categorical cross-entropy |
There exists a mathematical relationship between sigmoid and softmax. If the number of classes n is equal to 2, the equation for the softmax will simplify completely into that of sigmoid. Sigmoid is a special case of softmax.
How to Implement Softmax in Python
Four Python approaches implement softmax directly: pure Python, NumPy, PyTorch, and TensorFlow. SciPy provides a fifth, ready-made version.
From Scratch With Pure Python and NumPy
The naive version applies the formula directly.
import math
def softmax_naive(logits):
exps = [math.exp(z) for z in logits]
total = sum(exps)
return [e / total for e in exps]
print(softmax_naive([4.2, 1.3, -0.5, 2.1]))
# [0.8428, 0.0464, 0.0077, 0.1032]
Large logit values cause this version to raise an OverflowError, because math.exp() cannot represent numbers above approximately 1.8 × 10³⁰⁸. Subtracting the maximum logit before exponentiating avoids the error without changing the result.
import numpy as np
def softmax_stable(logits):
logits = np.array(logits)
exps = np.exp(logits - np.max(logits))
return exps / exps.sum()
print(softmax_stable([1000.0, 1000.0, 1000.0]))
# [0.3333, 0.3333, 0.3333]
Using PyTorch
import torch
import torch.nn.functional as F
logits = torch.tensor([4.2, 1.3, -0.5, 2.1])
probs = F.softmax(logits, dim=-1)
log_probs = F.log_softmax(logits, dim=-1)
The dim argument sets which axis the probabilities sum across. dim=-1 targets the last axis, which holds the class scores in most model outputs. PyTorch’s CrossEntropyLoss function already applies log_softmax internally, so applying softmax to the model’s output before passing it to that loss function double-counts the operation and produces incorrect gradients.
Using TensorFlow and Keras
import tensorflow as tf
logits = tf.constant([4.2, 1.3, -0.5, 2.1])
probs = tf.nn.softmax(logits, axis=-1)
In a Keras model, softmax is typically set directly as the output layer's activation.
from tensorflow.keras.layers import Dense
output_layer = Dense(4, activation='softmax')
Using SciPy
from scipy.special import softmax
probs = softmax([4.2, 1.3, -0.5, 2.1])
SciPy’s implementation already includes the numerical stability fix, so it produces correct results for both small and large logit values without additional code.
Common Softmax Mistakes (and How to Fix Them)
Four mistakes account for most softmax-related errors in production code.
- Wrong axis on batched data. Applying softmax across the batch axis instead of the class axis produces probabilities that sum to 1 down each column instead of across each row. The fix is setting axis=-1 in TensorFlow or dim=-1 in PyTorch for standard (batch, classes) tensors.
- Applying softmax twice. PyTorch’s CrossEntropyLoss and TensorFlow’s from_logits=True loss functions apply softmax internally. Adding a softmax activation to the model’s output layer as well applies the function twice, which flattens gradients and slows or stalls training.
- Using softmax in a hidden layer. Softmax forces its outputs into a fixed sum of 1, which removes a hidden neuron’s ability to represent independent magnitude and direction. ReLU and its variants are the standard choice for hidden layers instead.
- Mismatching the loss function. Pairing softmax outputs with mean squared error instead of cross-entropy loss is mathematically valid but converges slower, because mean squared error’s gradient with respect to the softmax output is weaker near the correct answer than cross-entropy’s gradient.
The Numerical Stability Problem Nobody Names
Softmax with large logits produces numeric overflow, because the exponential function grows too fast for standard floating-point storage.
A logit of 1000 raises e^1000 to a number with more than 400 digits, which exceeds the maximum value a standard 64-bit float can store and produces an overflow error. Subtracting the maximum logit in the vector from every value before exponentiating solves this problem without changing the final result, since the subtraction cancels out during normalization. This method has a formal name: the log-sum-exp trick, a standard numerical technique for computing log(Σe^x) without overflow. Softmax is a direct application of it. Every major deep learning framework, including PyTorch, TensorFlow, and SciPy, implements this fix internally.
Why Softmax Makes Neural Networks Overconfident
Softmax makes neural networks overconfident because the exponential function amplifies small logit differences into large probability differences, often producing predictions near 100% even when the prediction is wrong.
Guo et al. (2017), in On Calibration of Modern Neural Networks (ICML 2017, pp. 1321–1330), measured this effect directly. The study compared a 5-layer LeNet against a 110-layer ResNet on the CIFAR-100 dataset and found that the deeper, more modern network produced average confidence scores far higher than its actual accuracy justified, while the older, shallower network’s confidence stayed close to its true accuracy. The paper defines this mismatch between predicted confidence and actual correctness as calibration error.
Fixing Overconfidence: Temperature Scaling and Label Smoothing
Temperature scaling and label smoothing are the two most common fixes for softmax overconfidence.
Temperature scaling divides every logit by a single learned value T before applying softmax. A value of T greater than 1 softens the probability distribution, pulling extreme values back toward the center. Guo et al. (2017) found that this single-parameter method calibrated deep networks about as effectively as more complex, multi-parameter alternatives.
Label smoothing changes the training targets instead of the output. Rather than training the model toward a hard target of 1 for the correct class and 0 for every other class, label smoothing uses softened values, such as 0.9 for the correct class and 0.033 for each of three incorrect classes in a four-class problem. This prevents the model from being rewarded for pushing logits to extreme values during training.
The Softmax Bottleneck (and Why It Matters for Language Models)
The softmax bottleneck is a mathematical limit on how much information a single softmax layer can represent, identified by Yang et al. (2018).
Yang, Dai, Salakhutdinov, and Cohen, in Breaking the Softmax Bottleneck: A High-Rank RNN Language Model (ICLR 2018, arXiv:1711.03953), showed that a standard softmax output layer restricts a language model to a low-rank approximation of the true word-probability structure. Natural language requires a high-rank representation to capture context-dependent meaning, which a single softmax layer cannot fully provide. The paper’s proposed fix, Mixture of Softmaxes, combines multiple softmax distributions into one output and improved test perplexity to 47.69 on the Penn Treebank dataset and 40.68 on WikiText-2, outperforming the standard softmax baseline by more than 5.6 points of perplexity on the 1-billion-word dataset.
Softmax Beyond Classification: Attention, Transformers, and LLMs
Every attention layer in a Transformer model applies softmax to convert similarity scores between tokens into attention weights.
Vaswani et al. (2017), in Attention Is All You Need (NeurIPS 2017), defined scaled dot-product attention as:
Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V
Q, K, and V represent the query, key, and value matrices for a sequence of tokens, and dₖ is the dimension of the key vectors. Softmax converts the raw similarity scores between every pair of tokens (QKᵀ) into a set of weights that sum to 1 for each query token. Transformer-based language models, including GPT and Claude, repeat this softmax computation once per attention head, per layer, per token generated. A model with 32 attention heads and 32 layers performs 1,024 separate softmax computations for every single token it processes.
Also Used In: Reinforcement Learning and Model Ensembles
Reinforcement learning policy networks apply softmax to convert raw action-values into a probability distribution, allowing an agent to select actions stochastically rather than always choosing the single highest-valued action. Model ensembling methods use softmax-weighted averaging to combine predictions from multiple models into one final probability distribution.
Softmax Temperature in LLMs vs. Calibration Temperature
Softmax temperature and calibration temperature use the same formula but serve different purposes: one controls a language model’s output randomness, and the other corrects a classifier’s confidence.
Large language model APIs, including those from OpenAI and Anthropic, expose a temperature parameter that divides logits by T before the final softmax step at text-generation time. A temperature below 1 sharpens the distribution toward the single most likely next token, producing more predictable output. A temperature above 1 flattens the distribution, increasing the chance of selecting a less likely token and producing more varied output. This is the identical mathematical operation Guo et al. (2017) used for calibration; only the goal differs. Calibration temperature corrects a trained classifier’s confidence after training. Sampling temperature controls a language model’s output diversity during generation.
Softmax vs. Other Activation Functions at a Glance
| Function | Output range | Typical layer | Typical use |
|---|---|---|---|
| Sigmoid | 0 to 1 | Output | Binary or multi-label classification |
| Softmax | 0 to 1, sums to 1 | Output | Multi-class, mutually exclusive classification |
| ReLU | 0 to infinity | Hidden | Non-linearity between hidden layers |
| Gumbel-Softmax | 0 to 1, sums to 1 (approximate) | Output or sampling | Differentiable sampling from discrete distributions |
Try It Yourself: Interactive Softmax Calculator
Entering a classifier’s own logit values into the calculator applies the same four-step calculation described in the formula section to any set of numbers the reader chooses.
Frequently Asked Questions
Q1. What is the derivative of the softmax function?
Ans. Since each output is dependent on all the inputs, the derivative of the softmax function is the Jacobian matrix, where the partial derivative of the output i w.r.t. the input j is ∂σᵢ/∂zⱼ = σᵢ(δᵢⱼ − σⱼ) and δᵢⱼ is 1 if i = j and 0 otherwise.
Q2. Is softmax used in ChatGPT and Claude?
Ans. Yes. Each of the attention layers in any transformer-based architecture uses softmax for attention weight computation, and the final step of token generation employs the softmax function again.
Q3. Why does PyTorch recommend log_softmax over softmax?
Ans. PyTorch’s log_softmax combines the exponentiation and the log operations into a numerically stable procedure, unlike taking the logarithm of the result from the softmax function, which incurs precision losses.
Q4. Can softmax be implemented without any library?
Ans. Yes. It can be implemented purely in Python with no NumPy, PyTorch, or TensorFlow, as demonstrated in the implementation section above using only the math.exp method.
Q5. Why can’t sigmoid replace softmax in multi-class problems?
Ans. Since each class probability is computed independently by sigmoid, the sum of probabilities will not be equal to 1 and hence it does not represent a mutually exclusive choice among more than two classes.
Bottom Line
The Softmax algorithm has not evolved since its inception by Bridle in 1990, who characterized the technique as “normalized exponential.” What has changed is the application of the formula from one computation in a small recognition network to more than a thousand computations per token in a contemporary Transformer’s attention modules.