A loss function is a mathematical equation to calculate the discrepancy between the prediction of a deep learning algorithm and the actual target value. Training a neural network is about minimizing this number using a mathematical approach referred to as gradient descent. Various types of predictions will require various types of loss functions; regression needs a loss that punishes discrepancies in numerical terms, classification requires a loss that punishes wrong probabilities, and generative requires a loss that punishes unrealistic output. This guide will take you through 20 loss functions used in deep learning, sorted into 6 categories, with formulas, PyTorch and TensorFlow codes, and selection criteria.
What Will I Learn?
What Is a Loss Function? (And How Does It Differ From a Cost Function?)
A loss function computes the error for a single training sample while the cost function computes the error for all samples in the dataset. The two concepts are used interchangeably in literature, but their difference is important in understanding source codes and academic literature.
The cost function is the average of the loss function across n training samples:
For instance, in a data set containing 1,000 houses and their price predictions, the loss function outputs 1,000 errors, one error for each house. The cost function computes the average of the 1,000 errors.
How Loss Functions Drive Model Training
The loss function controls the training process through creating a numerical signal of error which is used by the optimizer to tune the weights of the model in such a way as to minimize errors. The process continues for every training dataset batch until the loss cannot be minimized any further.
Four steps constitute one training process cycle:
- Data is fed into the model to generate a prediction.
- The loss function measures the difference between the prediction and the actual output.
- Gradient with respect to each individual weight in the network is calculated using backpropagation.
- Each individual weight is updated using an optimizer like stochastic gradient descent or Adam based on gradient multiplied by the learning rate.
The process of updating weights was explained in the paper by Rumelhart, Hinton, and Williams “Learning representations by back-propagating errors” in 1986, published in Nature magazine. This process is still the main way to train neural networks.
The Six Categories of Loss Functions Used in Deep Learning
There are six types of loss functions that are applied to deep learning: regression loss functions, classification loss functions, ranking loss functions, image and reconstruction loss functions, adversarial loss functions, and specific loss functions. Each category matches a different type of model output.
| Category | Measures | Example loss functions | Typical task |
|---|---|---|---|
| Regression | Numeric distance between predicted and actual values | MSE, MAE, Huber Loss | House price prediction |
| Classification | Difference between predicted and true class probabilities | Binary Cross-Entropy, Categorical Cross-Entropy, Focal Loss | Spam detection, image classification |
| Ranking | Relative distance or similarity between data points | Contrastive Loss, Triplet Loss, Margin Ranking Loss | Face verification, recommendation systems |
| Image & reconstruction | Pixel-level or structural difference between generated and target images | Dice Loss, Jaccard/IoU Loss, Perceptual Loss | Image segmentation, super-resolution |
| Adversarial | How well a generator output fools a discriminator | GAN Loss, Least Squares GAN Loss, Wasserstein Loss | Image generation |
| Specialized | Sequence alignment, count data, or directional similarity | CTC Loss, Poisson Loss, Cosine Proximity Loss | Speech recognition, demand forecasting |
Regression Loss Functions
The regression loss functions determine the difference between the predicted continuous value and the real value numerically. The three most common regression loss functions used in deep learning are Mean Squared Error, Mean Absolute Error, and Huber Loss.
Mean Squared Error (MSE / L2 Loss)
The Mean Squared Error computes the mean of squared deviations between the predicted value and the actual value.
Squaring eliminates negative signs and puts heavier penalties on larger errors. For example, if the actual values are [3, 5, 2, 8] and predicted values are [2.5, 5.5, 2, 7], the squared errors would be [0.25, 0.25, 0, 1], leading to an MSE of 0.375.
The MSE method works best when the cost of a larger error is not proportional to its magnitude, as in load-bearing calculations.
# PyTorch
import torch
loss_fn = torch.nn.MSELoss()
loss = loss_fn(predictions, targets)
# TensorFlow
import tensorflow as tf
loss_fn = tf.keras.losses.MeanSquaredError()
loss = loss_fn(targets, predictions)
Mean Absolute Error (MAE / L1 Loss)
Mean Absolute Error computes the mean value of the absolute difference between predicted and actual values.
Using the above values, the absolute differences become [0.5, 0.5, 0, 1], resulting in an MAE of 0.5. The formula for MAE does not square the error and hence treats each observation equally based on how far it is from the true value.
# PyTorch
loss_fn = torch.nn.L1Loss()
# TensorFlow
loss_fn = tf.keras.losses.MeanAbsoluteError()
In the table shown above, the number of points used is 7, with 6 having errors that are close to 0.3 while 1 of the points being an outlier whose error is 6.0. The outlier makes up 98% of the MSE and 76% of the MAE. MAE can be used when there are outliers that are noise.
Huber Loss (and the Smooth L1 Variant)
The Huber loss function acts as MSE loss for smaller errors and MAE loss for larger errors. There exists a point at which this function changes from one loss to another. This point is referred to as delta (δ).
This loss function was introduced by Peter J. Huber in his 1964 paper titled “Robust Estimation of a Location Parameter” in the Annals of Mathematical Statistics. The Smooth L1 Loss used in Ross Girshick’s 2015 object detection model, Fast R-CNN, where δ is fixed to 1.0 by default.
# PyTorch
loss_fn = torch.nn.HuberLoss(delta=1.0)
# Fast R-CNN-style variant:
loss_fn_smooth_l1 = torch.nn.SmoothL1Loss(beta=1.0)
# TensorFlow
loss_fn = tf.keras.losses.Huber(delta=1.0)
The chart above illustrates why Huber Loss function is considered a hybrid because below the δ threshold point it traces the curve of MSE, while after this point, it traces the line of MAE.
Classification Loss Functions
The loss function that computes the distance between the predicted probability values and the class labels is called the classification loss function. The six popular loss functions used for classification problems are Binary Cross-Entropy, Categorical Cross-Entropy, Sparse Categorical Cross-Entropy, KL Divergence, Hinge Loss, and Focal Loss.
Binary Cross-Entropy (Log Loss)
The Binary Cross Entropy Loss compares the predicted probability to the real binary value (0 or 1).
If the ground-truth value is 1, then the predicted probability of 0.9 will result in BCE = −log(0.9) = 0.105. However, the same ground-truth value with a predicted probability of 0.2 will result in BCE = −log(0.2) = 1.609. That is, the wrong but sure prediction yields 15 times greater loss
# PyTorch
loss_fn = torch.nn.BCEWithLogitsLoss() # applies sigmoid internally
# TensorFlow
loss_fn = tf.keras.losses.BinaryCrossentropy(from_logits=True)
Both cases for the labels can be found on the graph above. The loss function tends to infinity when a certain prediction approaches the incorrect label, and that is why Binary Cross Entropy punishes certainty more than uncertainty.
Categorical Cross-Entropy
Categorical Cross-Entropy is an extension of Binary Cross-Entropy in cases where we have more than two classes.
In CCE, we need to make sure that our output layer uses softmax activations, and therefore the output sums up to 1. In addition, our target needs to be one-hot encoded — meaning, if we have a classification task of three classes, class 2 will be encoded as [0, 1, 0].
# TensorFlow
loss_fn = tf.keras.losses.CategoricalCrossentropy() # expects one-hot labels
Sparse Categorical Cross-Entropy
The Sparse Categorical Cross-Entropy calculation is identical to the Categorical Cross-Entropy calculation, with the sole difference that Sparse Categorical Cross-Entropy allows for integer class labels rather than one-hot encodings.
In the case of a classification task that could have as many as 10,000 different classes, a one-hot encoding would require a 10,000-length vector whereas a sparse encoding only requires an integer.
# TensorFlow
loss_fn = tf.keras.losses.SparseCategoricalCrossentropy() # expects integer labels
# PyTorch
loss_fn = torch.nn.CrossEntropyLoss() # accepts integer labels natively — no separate "sparse" class exists
PyTorch does not need a separate sparse variant because CrossEntropyLoss already accepts integer class indices and applies log-softmax internally.
Kullback-Leibler (KL) Divergence
The measure KL Divergence quantifies the difference between one probability distribution and a reference probability distribution.
This distance was introduced by Solomon Kullback and Richard Leibler in the 1951 paper “On Information and Sufficiency” published in the Annals of Mathematical Statistics. KL Divergence in deep learning is used in variational autoencoders when it compares the latent distribution that was learned to the standard normal distribution and in knowledge distillation when it compares the output distributions of a small and large model.
# PyTorch
loss_fn = torch.nn.KLDivLoss(reduction="batchmean") # input must be log-probabilities
# TensorFlow
loss_fn = tf.keras.losses.KLDivergence()
Hinge Loss
The Hinge Loss punishes predictions made outside of the correct side of the decision boundary or too near to it; its goal is to widen the distance between different classes as much as possible.
L(y, f(x)) = max(0, 1 − y × f(x)) where y ∈ {−1, 1}
The Hinge Loss is the conventional loss function used by Support Vector Machines.
# PyTorch
loss_fn = torch.nn.MultiMarginLoss() # multi-class hinge loss
# TensorFlow
loss_fn = tf.keras.losses.Hinge() # binary, expects labels of −1 or 1
Focal Loss
Focal Loss introduces a modulating factor into the cross-entropy that reduces the effect of loss from easy examples that have been well-classified, making the network pay attention to difficult examples.
FLoss was created by Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár in “Focal Loss for Dense Object Detection” published in the 2017 IEEE International Conference on Computer Vision (ICCV). The paper trains a one-stage object detector (RetinaNet), discovering that an extreme class imbalance between the foreground and background is the cause why one-stage detectors have always performed worse than two-stage detectors.
By setting γ = 2, a well-classified example with pₜ = 0.9 will have its loss scaled by (1 − 0.9)² = 0.01 – which is 100 times lower. A poorly classified example with pₜ = 0.1 will have its loss scaled by (1 − 0.1)² = 0.81, which is barely reduced. This is the difference that forces the model to continue learning hard examples instead of being dominated by easy examples.
Neither PyTorch nor TensorFlow offers FLoss in the form of a built-in class. The following implementation follows the original paper:
# PyTorch — custom Focal Loss
import torch
import torch.nn as nn
import torch.nn.functional as F
class FocalLoss(nn.Module):
def __init__(self, alpha=0.25, gamma=2.0):
super().__init__()
self.alpha = alpha
self.gamma = gamma
def forward(self, logits, targets):
bce = F.binary_cross_entropy_with_logits(logits, targets, reduction="none")
p_t = torch.exp(-bce)
loss = self.alpha * (1 - p_t) ** self.gamma * bce
return loss.mean()
Ranking Loss Functions
Unlike regression loss, which compares one prediction to a target, ranking loss computes the distance between two or more data points. Some ranking losses used in deep learning include Contrastive Loss, Triplet Loss, and Margin Ranking Loss.
Contrastive Loss
The goal of contrastive loss is to train a model that keeps dissimilar objects distant from each other and similar objects close together in the embedding space.
where dᵢ is the distance between two embeddings, yᵢ is equal to 1 for dissimilar pairs and to 0 for similar pairs, and m is a margin. Contrastive Loss can be used to train siamese neural networks for applications like signature verification.
No PyTorch or TensorFlow core class implements Contrastive Loss directly. It is commonly implemented manually or through the open-source pytorch-metric-learning library.
Triplet Loss
In Triplet Loss, a model is trained with the use of three input images at once: one anchor, one positive image, and one negative image, so that the distance between the anchor and positive becomes smaller while the distance between the anchor and negative increases.
Schroff, Kalenichenko, and Philbin described an application of Triplet Loss in their 2015 paper “FaceNet: A Unified Embedding for Face Recognition and Clustering” at CVPR. With FaceNet, 99.63% accuracy was achieved on Labeled Faces in the Wild dataset, which is 30% better than the previous state-of-the-art.
# PyTorch
loss_fn = torch.nn.TripletMarginLoss(margin=1.0)
Margin Ranking Loss
The Margin Ranking Loss calculates the margin between two examples and punishes the network when the proper order between the two is not maintained by at least the required margin.
Margin Ranking Loss = max(0, −y × (s⁺ − s⁻) + margin)
The function is used in training recommendation systems and learning-to-rank algorithms where it is only the relative ordering that matters.
# PyTorch
loss_fn = torch.nn.MarginRankingLoss(margin=1.0)
Image and Reconstruction Loss Functions
Image and reconstruction loss function classes judge the performance of image generators and segmenters. These loss functions include five popular examples of deep learning such as pixel-wise cross-entropy, dice loss, Jaccard/IoU loss, perceptual loss, and total variation loss.
Pixel-wise Cross-Entropy
The Pixel-wise Cross-Entropy loss uses the Cross-Entropy method individually for each pixel in the segmentation mask and predicts the category for that pixel. The Pixel-wise Cross-Entropy loss method considers each individual pixel as a separate classification problem, making it less effective when the proportion of the target class in the image is low.
Dice Loss
Dice Loss is a measure of overlap between the predicted mask and the ground truth mask, computed as one minus the Dice similarity score.
Dice Loss was first proposed by Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi in the 2016 paper “V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation” at the International Conference on 3D Vision (3DV).
This paper discussed the use cases in medical image processing where the entity to be segmented, like a tumor, occupies only a very small number of pixels in an image — a situation where Pixel-wise Cross-Entropy loss fails.
Neither PyTorch nor TensorFlow core includes Dice Loss. The PyTorch-based medical imaging library MONAI provides a built-in implementation through monai.losses.DiceLoss.
Jaccard Loss (Intersection over Union)
Jaccard Loss is also known as IoU loss, which calculates the ratio of the intersection over union of predicted and ground truth masks, then takes away 1.
Jaccard Loss and Dice Loss are both for segmentation overlap measurement, but Jaccard Loss gives more punishment to one particular error than Dice Loss since the denominator of Jaccard Loss doesn’t include the intersection part counted twice.
Perceptual Loss
Perceptual Loss is the difference between the high-level feature representations of two images and not between their actual pixels.
Perceptual Loss was proposed by Justin Johnson, Alexandre Alahi, and Li Fei-Fei in their paper titled “Perceptual Losses for Real-Time Style Transfer and Super-Resolution” published in ECCV in 2016. The feature representations, φⱼ, are generally obtained using a pre-trained VGG model on ImageNet from an intermediate layer of the model.
Total Variation Loss
Total Variation Loss will minimize the difference between neighboring pixels, which will lead to an increase in the spatial smoothness of a generated image.
Total Variation Loss serves as a regularizer that can be used in conjunction with another loss, like Perceptual Loss.
Adversarial (GAN) Loss Functions
The GAN loss functions are used to train Generative Adversarial Networks (GAN), where there is a combination of a generator and discriminator networks. The three types of adversarial loss functions include GAN loss function, least squares GAN loss function, and Wasserstein loss function.
Standard GAN (Minimax) Loss
GAN loss function trains the generator and discriminator by playing a minimax game, with the discriminator trying to maximize the chance of differentiating between real and generated data, and the generator minimizing the discriminator’s chances of doing so.
This concept was proposed by Ian Goodfellow et al. in the paper titled “Generative Adversarial Networks” published in 2014 in the NeurIPS conference. Both discriminator and generator use Binary Cross Entropy as the loss functions with real data labeled as 1 and generated data as 0.
Least Squares GAN Loss
Least Squares GAN Loss replaces the logarithmic terms in the standard GAN loss with squared-error terms to produce more stable gradients during training.
Xudong Mao and colleagues introduced this variant in the 2017 paper “Least Squares Generative Adversarial Networks,” presented at ICCV.
Wasserstein Loss
Wasserstein Loss measures the Earth Mover’s Distance between the real data distribution and the generated data distribution, replacing the discriminator with a critic that outputs an unbounded score instead of a probability.
Martin Arjovsky, Soumith Chintala, and Léon Bottou introduced Wasserstein GAN in a 2017 paper presented at the International Conference on Machine Learning (ICML). Ishaan Gulrajani and colleagues extended it the same year with “Improved Training of Wasserstein GANs” at NeurIPS, adding a gradient penalty term (WGAN-GP) that removed the need for weight clipping and further stabilized training.
Specialized and Sequence Loss Functions
Specialized loss functions handle tasks that do not fit the regression or classification framework directly. Three specialized loss functions are common in deep learning: CTC Loss, Poisson Loss, and Cosine Proximity Loss.
CTC Loss (Connectionist Temporal Classification)
CTC Loss is used for training sequence models when the alignment of input frames and output labels is unknown.
CTC Loss was proposed by Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber in the paper “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks” published in 2006 in the proceedings of ICML conference. CTC Loss is used for training speech and handwriting recognition models, where the number of input frames or strokes does not correspond to the number of output characters.
# PyTorch
loss_fn = torch.nn.CTCLoss()
Poisson Loss
Models trained with Poisson loss learn to forecast the count data, with the predicted value as the rate parameter of the Poisson distribution.
Poisson Loss = Σ (ŷᵢ − yᵢ × log(ŷᵢ))
Poisson loss can be used in forecasting tasks where the target variable is a count (non-negative integers), like number of customer support tickets per day.
# TensorFlow
loss_fn = tf.keras.losses.Poisson()
# PyTorch
loss_fn = torch.nn.PoissonNLLLoss()
Cosine Proximity Loss
Cosine Proximity Loss measures the cosine of the angle between a predicted vector and a target vector, training a model to align their directions regardless of magnitude.
Cosine Proximity Loss = −(1/N) × Σ (yᵢ · ŷᵢ) / (‖yᵢ‖ × ‖ŷᵢ‖)
# TensorFlow
loss_fn = tf.keras.losses.CosineSimilarity()
# PyTorch
loss_fn = torch.nn.CosineEmbeddingLoss()
Common Training Failures Caused by the Wrong Loss Function
The choice of loss function directly causes several common training failures.
| Symptom | Likely cause | Fix |
|---|---|---|
| Loss becomes NaN | Learning rate too high, causing exploding gradients; or log(0) inside Cross-Entropy when a predicted probability rounds to 0 | Lower the learning rate; add a small epsilon inside the log calculation; apply gradient clipping |
| Loss stays flat and does not decrease | Learning rate too low; loss function mismatched with the output layer’s activation function | Increase the learning rate; match the loss function to the correct output activation |
| Loss decreases, then spikes | Numerical instability in the optimizer; a batch containing extreme outliers | Apply gradient clipping; switch from MSE to Huber Loss for outlier-heavy batches |
| Model ignores the minority class | Standard Cross-Entropy treats every class equally despite class imbalance in the dataset | Switch to Focal Loss or apply per-class weights |
| Training loss is low, validation loss is high | The model has overfit the training data | Add regularization or dropout; apply early stopping; add training data |
How to Choose the Right Loss Function
Choosing a loss function starts with identifying what the model predicts: a continuous number, a class label, a similarity score, or a full image.
Five steps apply to every loss function decision:
- Determine the type of prediction: a number, a category label, a similarity measurement, or an image.
- Examine the presence of outliers among the targets. In case outliers contain signal, choose MSE; if they denote noise, then MAE or Huber Loss.
- Look for class imbalance within the dataset. If one class overwhelms the other classes, choose Focal Loss instead of Cross-Entropy Loss.
- Ensure that the selected loss function is consistent with the type of activation in the output layer; Cross-Entropy can only be used with a softmax or sigmoid output.
- Make sure that the loss function decreases on a subset of the data before using it on all training samples.
# Example: Quick training loop test on a mini-batch
import torch
import torch.nn as nn
model = nn.Linear(10, 1)
criterion = nn.MSELoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
x = torch.randn(8, 10)
y = torch.randn(8, 1)
for step in range(5):
optimizer.zero_grad()
output = model(x)
loss = criterion(output, y)
loss.backward()
optimizer.step()
print(f"Step {step + 1}, Loss: {loss.item():.4f}")
Writing a Custom Loss Function in PyTorch and TensorFlow
A custom loss function is required when a task needs a calculation that no built-in class provides, or when multiple loss functions must be combined into one training signal.
# PyTorch — custom loss via nn.Module
import torch.nn as nn
class CustomLoss(nn.Module):
def __init__(self):
super().__init__()
def forward(self, predictions, targets):
error = predictions - targets
return (error ** 2).mean()
# TensorFlow — custom loss via keras.losses.Loss
import tensorflow as tf
class CustomLoss(tf.keras.losses.Loss):
def call(self, y_true, y_pred):
error = y_pred - y_true
return tf.reduce_mean(tf.square(error))
Multi-task models commonly combine losses through a weighted sum:
Total Loss = w₁ × Loss₁ + w₂ × Loss₂
For example, an object detection model sums a regression loss for bounding-box coordinates and a classification loss for object category, with each term weighted to balance their relative scale.
Loss Function Cheat Sheet
| Loss function | Category | Sensitive to outliers | PyTorch class | TensorFlow class |
|---|---|---|---|---|
| MSE | Regression | High | nn.MSELoss | MeanSquaredError |
| MAE | Regression | Low | nn.L1Loss | MeanAbsoluteError |
| Huber Loss | Regression | Medium | nn.HuberLoss | Huber |
| Binary Cross-Entropy | Classification | N/A | nn.BCEWithLogitsLoss | BinaryCrossentropy |
| Categorical Cross-Entropy | Classification | N/A | nn.CrossEntropyLoss | CategoricalCrossentropy |
| Sparse Categorical Cross-Entropy | Classification | N/A | nn.CrossEntropyLoss | SparseCategoricalCrossentropy |
| KL Divergence | Classification | N/A | nn.KLDivLoss | KLDivergence |
| Hinge Loss | Classification | N/A | nn.MultiMarginLoss | Hinge |
| Focal Loss | Classification | N/A | Custom (no built-in) | Custom (no built-in) |
| Contrastive Loss | Ranking | N/A | Custom / pytorch-metric-learning | Custom |
| Triplet Loss | Ranking | N/A | nn.TripletMarginLoss | Custom |
| Margin Ranking Loss | Ranking | N/A | nn.MarginRankingLoss | Custom |
| Pixel-wise Cross-Entropy | Image/reconstruction | N/A | nn.CrossEntropyLoss (per pixel) | SparseCategoricalCrossentropy (per pixel) |
| Dice Loss | Image/reconstruction | N/A | monai.losses.DiceLoss | Custom |
| Jaccard/IoU Loss | Image/reconstruction | N/A | Custom | Custom |
| Perceptual Loss | Image/reconstruction | N/A | Custom (uses pretrained VGG) | Custom (uses pretrained VGG) |
| Total Variation Loss | Image/reconstruction | N/A | Custom | tf.image.total_variation |
| GAN Loss | Adversarial | N/A | nn.BCEWithLogitsLoss (applied twice) | BinaryCrossentropy (applied twice) |
| Least Squares GAN Loss | Adversarial | N/A | nn.MSELoss (applied twice) | MeanSquaredError (applied twice) |
| Wasserstein Loss | Adversarial | N/A | Custom | Custom |
| CTC Loss | Specialized | N/A | nn.CTCLoss | tf.nn.ctc_loss |
| Poisson Loss | Specialized | N/A | nn.PoissonNLLLoss | Poisson |
| Cosine Proximity Loss | Specialized | N/A | nn.CosineEmbeddingLoss | CosineSimilarity |
Frequently Asked Questions
Q1. Why does my loss become NaN during training?
Ans. The value of a loss will be NaN when the learning rate is too large and results in exploding gradients, or when Cross-Entropy calculates the log of zero. The learning rate and gradients can be reduced to fix this problem.
Q2. Can I combine multiple loss functions in one model?
Ans. Yes. Multi-task networks usually calculate the loss as a weighted sum of several losses, e.g., for coordinate regression and for category classification into one scalar loss.
Q3. What is the difference between Focal Loss and Cross-Entropy?
Ans. Focal Loss introduces a factor (1 − pₜ)^γ which decreases the loss value for easy, correctly classified examples, while regular Cross-Entropy loss treats all the examples equally regardless of how confident the classifier was in its prediction.
Q4. Is a loss function the same as accuracy?
Ans. No. While loss function measures the magnitude of the prediction error on a continuous scale, accuracy calculates the number of examples that were classified correctly; it’s possible to decrease the loss of a network while keeping the same accuracy if predictions approach the decision boundary.
Q5. Can I use different loss functions for the same machine learning problem?
Ans. Yes, people normally try different types of loss functions, like MSE and Huber Loss, in the same regression problem and pick the one which gives the smallest validation loss for their particular data set.
Q6. What loss function do object detection models use?
Ans. The loss functions in object detection problems are a combination of the regression loss, which calculates the bounding box coordinates, and the classification loss, which classifies the object.
Conclusion
Choosing a loss function involves aligning the chosen function with the three characteristics of the task, namely the kind of prediction, whether there are any outliers in the target values, and the class balance in the data. The above table of comparison gives an immediate reference to that end for 20 loss functions. Matching the right loss function with the appropriate output activation function and training curve monitoring is the key to avoiding most of the training errors listed here.