Gradient Clipping: How to Stop Exploding Gradients

|
9 min read
|
15 views
Gradient Clipping

The purpose of gradient clipping is to limit the gradient’s magnitude during backpropagation such that a single update cannot derail the process. Whenever the gradient norm becomes larger than a certain threshold, it is clipped to be within the threshold, but the gradient’s direction remains unchanged. This solves the exploding gradients problem without affecting what the network has to learn.

What Is Gradient Clipping?

Gradient clipping refers to the process whereby the size of gradients is restricted when performing backpropagation in such a way that no single gradient is allowed to become greater than a certain threshold. There are two methods used to clip gradients. One is clipping by value, whereby the individual components of the gradient are clipped individually. The other is clipping by norm.

This process is illustrated in the diagram above. The gradient vector that has not been clipped has a norm value of 38.8 and stretches much further than the boundaries. On the right side, the gradient vector is clipped and shortened, with its norm value being 1.0 and falling on the boundary while preserving its directionality.

Gradient norm clipping was developed in 2013 by Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. They introduced gradient norm clipping as a solution to training instability in recurrent neural networks. This was done through developing gradient norm clipping as a solution to exploding gradients together with a soft constraint on vanishing gradients. In 2026, gradient norm clipping became a default in most deep learning training scripts.

Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Why Exploding Gradients Happen

A gradient shows where each weight needs to move, both the distance and the direction. Backpropagation involves multiplying gradients from layer to layer in which the signal passes through, according to the chain rule. When all these multipliers are greater than 1 for layer-to-layer connections, the result exponentially increases depending on the number of layers.

Three particular models are the most sensitive to this problem: RNNs, LSTMs, and very deep feed-forward networks. In the case of the first two, the example is the most obvious, because in the RNN, one recurrent layer is passed several times, once for each timestep, and thus 50-timestep input passes through 50 recurrent layers with the same weight matrix. When the spectral radius of the recurrent weight matrix is greater than 1, it multiplies the gradient at each time step.

The symptom is visible in the loss curve: values jump to NaN, the loss spikes without warning, or weights grow so large that the model’s output stops changing.

How Gradient Clipping Works (Step by Step)

Norm clipping is performed as a hard-coded routine every training epoch, following the backward pass and preceding weight update.

  1. The gradient of each parameter is computed using the regular backward pass routine.
  2. The global norm is calculated – all gradient tensors are flattened and concatenated into a single vector; the L2 norm of that vector is the global norm.
  3. If the norm is equal to or less than the threshold value, no action is taken.
  4. Norm clipping – if the norm is greater than the threshold, then the gradient is scaled down – each gradient is multiplied by the threshold over the norm, reducing the vector length precisely to the threshold value without changing its direction.

The optimizer updates the (possibly clipped) gradient exactly as it normally does. Norm clipping involves one additional comparison and at most one multiplication per training epoch, so there is virtually no additional cost.

Types of Gradient Clipping

Three distinct strategies fall under the name “gradient clipping.” They differ in what gets measured and what gets scaled.

Clipping by Value

Clipping of the gradients by value clips each individual gradient component based on its upper and lower bound; a gradient component that exceeds the maximum gets reduced to the maximum value, while any that is less than the minimum gets clipped to the minimum value. Clipping results in a change of direction of the gradient vector, since only the oversized components get modified. A two-dimensional gradient vector [3.0, 0.4], when clipped to 1.0, is [1.0, 0.4].

Clipping by Norm (L1 vs. L2)

The clipping method using the norm calculates the length of the full gradient vector with respect to a particular norm, usually the L2 or Euclidean norm, and scales down the full vector by the same amount if its length surpasses the threshold. Since all entries of the vector will be scaled by the same amount, the direction of the gradient vector remains intact. The clip_grad_norm_ function in PyTorch has the default norm type of L2 (norm_type=2.0).

Global-Norm vs. Per-Parameter Clipping

Global norm clipping determines one norm across all the parameters in the entire network, followed by a single scaling factor. Parameter-wise (or layer-wise) clipping finds a norm and a scaling factor for each parameter (or layer). Global norm clipping is a common default option for most training procedures since it maintains the relative scale between different parameters.

Fourth, Adaptive Gradient Clipping (AGC) is a method which does not even use one particular threshold. AGC was proposed for training image classification networks without normalizers in very large batch settings and is characterized by the gradient clipping based on the ratio between the norm of each gradient to the norm of the corresponding parameter, with a threshold that is calculated individually for each parameter.

MethodWhat is measuredWhat gets scaledPreserves directionTypical use
Value clippingEach gradient elementOnly elements past the boundNoRare in modern pipelines
L2 norm clippingLength of the full gradient vectorEvery element, by the same factorYesDefault for RNNs, transformers, LLMs
Per-parameter norm clippingLength of each parameter’s gradientThat parameter’s elements onlyYes, within each groupMixed-scale architectures
Adaptive Gradient ClippingRatio of gradient norm to parameter norm, per unitElements exceeding the unit-wise ratioYes, within each unitNormalizer-free networks, large-batch training

How to Choose a Clip Threshold

The threshold is a hyperparameter, not a fixed constant, but three sources of evidence narrow the search.

The recommended starting point is 1.0, and you can tweak from there. Hugging Face’s Trainer defaults to 1.0 for max_grad_norm, where 0.5 is a more conservative choice and 5.0 is a less aggressive one. In fact, this has converged into the default choice for large-scale training in practice, where the GPT-3 paper mentions clipping it to 1.0 for all model sizes.

See what happens when scaling up. In the DeepMind’s Gopher paper, gradient clipping by global norm was set at a value of 1 for most models, but dropped down to 0.25 for the 7.1 billion parameters model and the whole Gopher model in order to keep stability. A growing model requires to verify clip threshold again, not take it for granted.

Log the pre-clip norm before adjusting the threshold number. Measure the global norm of gradients at each step during a trial run without any clipping. If the norm is constantly below 2.0, a threshold of 1.0 will cause clipping and prevent learning most of the time. If the norm occasionally surges into the hundreds, a threshold of 1.0 will be idle most of the time until the surges happen – which is exactly what it is supposed to do.

Gradient Clipping in Practice: A Worked Example

Gradient Clipping in Practice

Above is plotted an experimental function that was designed specifically for this paper and that has a small two-dimensional loss function surface with a sharp and steep cliff – the very same phenomenon that causes exploding gradients in RNNs, but scaled down so that its behavior could be computed precisely and checked numerically.

At the initial point, the norm of the gradient equals 38.8, 39 times higher than the threshold of 1.0. If clipping does not happen, then after the first update, the value of one parameter jumps from -1.35 to -3.29, overshooting the global minimum (located around -1.83) by more than 1.4. If clipping happens, then after the first step, the value of the parameter changes by precisely 0.05 (learning rate * clipping threshold) to become -1.40.

It does not disappear by itself either. At 80 update steps, the unclipped loss is 0.093, and the clipped loss is 0.037, which is just 2% away from the actual minimum loss of 0.036. The huge step that the unclipped loss had taken landed it in a location where the gradient was too flat for it, and despite the subsequent 80 steps, it remained in this area. The purpose of clipping is not about making any particular step faster; it is to make sure that no step is large enough to push the optimizer into an escapeless area.

Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Gradient Clipping in Transformers and Modern LLM Training

Transformers’ language models carried gradient clipping forward from the RNN era and it seems to have become standard practice despite increasing scale. Global norm clipping with a threshold of 1.0 is employed in the publicly documented training parameters for GPT-3 and Gopher, as cited above, and it is the parameter made available by default in general purpose training libraries like Hugging Face’s Trainer.

In practice, gradient clipping in transformer pretraining is insurance against extreme cases more than an everyday procedure. A properly configured learning rate and schedule will ensure that the gradient norm never exceeds the threshold except in extremely rare cases. These rare cases will be due to either a statistical outlier in the input data, a numerical instability in mixed precision calculations, or an unlucky batch draw.

Why Gradient Clipping Matters (Benefits)

Clipping solves four distinct failure cases which are all caused by the same reason.

  • Clipping solves numerical overflows. Since a gradient is unbound, it could become too big to fit into a 32-bit or 16-bit float value, leading to NaN values, which will corrupt all subsequent calculations.
  • Clipping makes it possible to have stable training at high learning rates, since update size is bounded no matter how big the gradient is.
  • Clipping allows overcoming problems caused by sharp loss surface which recurrent and deep networks generate. In these surfaces, there might be large differences in the magnitude of the gradient of two neighboring points, reaching several orders of magnitude.
  • Clipping eliminates the dependence on activation function type. Saturation activation functions like tanh and sigmoid generate tiny gradients everywhere except for the steepest point of the curve.

Limitations and Common Mistakes

Clipping does not replace proper design of models. There are several common mistakes one should avoid while implementing clipping.

The presence of the clipping may cover up the problem of too high a learning rate. In case the clip threshold triggers almost every update, the actual learning rate has to be lowered rather than the clip threshold itself.

Order of actions is important here. The gradients have to be clipped after the backward pass has been completed but before the step function of the optimizer is run.

Setting the clip threshold too low will slow down the optimization process. The reason is that if the threshold is lower than the actual norm of the gradients generated by the healthy training process, the useful information will get lost at each step.

Clipping cannot solve the problem of vanishing gradients. It helps to reduce too large gradients and does nothing about vanishing ones.

Frequently Asked Questions

Q1 . Does gradient clipping slow down training? 

Ans. Gradient clipping has no significant computational overhead – just one additional norm calculation and multiplication in the worst case per step – and will be slower only when the threshold value is smaller than the gradient norm that can be achieved without causing instability.

Q2. Do I need gradient clipping if I already use a learning rate warmup? 

Ans. The warmup helps decrease the frequency of clipping due to making the early updates small, but it does not eliminate the possibility of getting a very large gradient at a later point during training, so gradient clipping is still advised.

Q3. What clip value should I use for a transformer model? 

Ans. Try starting with 1.0 global norm threshold like in GPT-3 and default Hugging Face Trainer, and lower this value only when you see from the gradient norm log that you need to do so.

Q4. Is gradient clipping still used in modern large language model training? 

Ans. Yes – Global Norm clipping with a cutoff point close to 1.0 is a common practice when it comes to training large language models today.

Q5. How do I know if my model needs gradient clipping? 

Ans. Record the global norm of the gradient during a short training session; if the global norm is small most of the time but occasionally spikes high, it indicates the need for gradient clipping.

Key Takeaways

Gradient clipping involves scaling a large gradient vector to a particular fixed value without altering its direction, and its current form was suggested by Pascanu, Mikolov, and Bengio back in 2013. Value of 1.0, which is applied in GPT-3 and is set as default in the training pipeline created by Hugging Face, is merely a reference based on a decade worth of research into the training of different networks, not an absolute value applicable to all types of architectures at any scale. What should come first when training any network that has not yet been measured for its gradient norms – a test run with the plot of gradient norm after each step.

Gyansetu offers top professional training certification courses designed to enhance your skills and advance your career, providing industry-relevant knowledge and practical expertise.