Each parameter in Adam optimization algorithm is updated with their unique learning rate which is obtained using the estimated mean and variance of gradients. Adam was proposed by Diederik P. Kingma and Jimmy Ba in the paper titled “Adam: A Method for Stochastic Optimization” at ICLR 2015 conference. Adam is still the go-to optimizer for most deep learning use cases until 2026. AdamW, which is an optimization technique that came later on, has taken its place as the default optimizer for training huge language models. In this guide, we will be discussing the exact equations used in Adam formulas, default hyperparameters across frameworks, working code in PyTorch, TensorFlow and NumPy, and the specific conditions under which AdamW or another optimizer performs better.
What Will I Learn?
What Is the Adam Optimizer?
Adam is an algorithm used to optimize the parameters of a neural network by assigning each parameter an adaptive learning rate based on the first and second moments of previous gradient values. The name is not an acronym, Kingma and Ba coined this name based on “adaptive moment estimation.”
The method was developed by Kingma, from the University of Amsterdam, and Ba, from the University of Toronto, and published at ICLR 2015. The original article presents eight reasons for using Adam on non-convex optimization tasks: it is easy to implement, has high computational efficiency, and is memory efficient compared to second-order methods; it is invariant to the scaling of the gradients by a diagonal matrix; it is good for problems with large amounts of data or parameters; and it is suitable for non-stationary and highly noisy/sparse gradients. The authors write that its hyperparameters have “intuitive interpretations and typically require little tuning.”
Adam is a first order optimization technique, which calculates the updates based on the gradient and the square of the gradient. Adam does not calculate the Hessian or any kind of second derivative information, which makes a second-order technique like Newton’s method a true second order technique. In one of the definitions published by ScienceDirect Topics on Adam, Adam has been described to calculate “second order information” – this is confusing second moment (which is basically a running average of square of gradients and is a first-order entity) with second order optimization techniques (which use curvature information).
Adam was evaluated for four tasks by Kingma and Ba including logistic regression on the MNIST digit dataset, logistic regression on the IMDB sentiment analysis dataset, multilayer perceptron on the MNIST dataset and convolutional neural network on the CIFAR-10 image dataset. In all the four cases, Adam was as good or better than AdaGrad, RMSprop, SGD with Nesterov momentum and AdaGrad respectively.
How Adam Optimizer Works
Adam is a combination of two different techniques; one technique is a running average of the gradient that defines the direction of the update, while the other technique is a running average of the squared gradient that defines the step size of the update for every single parameter.

First Moment Estimate: Momentum
The first moment is an exponentially decaying average of past gradients:
g_t is the gradient of the loss with respect to the parameter at step t. β₁ controls how much weight the average gives to past gradients versus the current one; the default value is 0.9. A high β₁ smooths out short-term noise in the gradient and carries the update forward in a consistent direction, which reduces oscillation across steep, narrow loss surfaces.
Second Moment Estimate: RMSprop
The second moment is an exponentially decaying average of the squared gradients:
β₂ defaults to 0.999, giving the squared-gradient average a longer memory than the first moment’s β₁. Dividing the update by the square root of this term later gives parameters with a history of large gradients a smaller step, and parameters with a history of small gradients a larger step. This is the mechanism RMSprop introduced on its own, before Adam combined it with momentum.
Bias Correction
m_t and v_t both start at zero. In the earliest steps, this initialization biases both estimates toward zero, particularly when β₁ and β₂ are close to 1. Adam corrects for this directly:
The correction factor shrinks toward 1 as t increases, so its effect is strongest in the first few steps of training and negligible later.
Final Weight Update
α is the learning rate (default 0.001), and ε is a small constant added to the denominator to prevent division by zero when v̂_t is close to zero (default 1e-8 in the original paper). The diagram below summarizes the full sequence from gradient to parameter update.
Adam Optimizer Hyperparameters
Adam has four hyperparameters. Each one has a default value from the original paper, and two of them differ between the most commonly used deep learning frameworks.
| Hyperparameter | Symbol | Controls | Paper default | PyTorch default | Keras/TensorFlow default |
|---|---|---|---|---|---|
| Learning rate | α | Overall step size | 0.001 | 0.001 | 0.001 |
| Beta1 | β₁ | Decay rate of the first moment (gradient average) | 0.9 | 0.9 | 0.9 |
| Beta2 | β₂ | Decay rate of the second moment (squared-gradient average) | 0.999 | 0.999 | 0.999 |
| Epsilon | ε | Numerical-stability constant | 1e-8 | 1e-8 | 1e-7 |
The epsilon default is the one value that differs by framework. PyTorch’s torch.optim.Adam uses 1e-8, matching the original paper. Keras’s tf.keras.optimizers.Adam uses 1e-7. A model trained with default settings in one framework will not necessarily reproduce the identical training curve in the other, because of this difference alone.
In situations where gradients are sparse, like in natural language processing and certain tasks in computer vision, the original paper suggests using a higher value of β₂ near 1.0 compared to the default value of 0.999 because it will increase the memory of the second moment term.
Adam vs. SGD vs. RMSprop vs. AdamW: Which Should You Use?
Adam vs. SGD
SGD updates the parameters using one constant learning rate. Adam updates every parameter individually, using the learning rate adjusted by the history of gradients for that specific parameter. The difference is evident in the convergence speed on surfaces of functions having valleys that are curved and narrow.
The diagram below provides the benchmark experiment performed for this paper: SGD (one constant learning rate), RMSprop, and Adam optimization algorithms used to minimize the Rosenbrock function, a non-convex function with a curved valley at its minimum. All three methods were initialized at the same starting position and run for 2,000 iterations.

Adam achieves a loss of 0.0003 at iteration 1,000. The corresponding loss value of SGD is 1.20 while that of RMSprop is 0.11. The loss of RMSprop converges to 0.09 from here until the end and fails to match Adam’s final loss. SGD keeps improving and obtains a loss of 0.086 at iteration 2,000, far away from Adam’s almost zero loss.
Convergence speed is not everything. According to Wilson, Roelofs, Stern, Srebro, and Recht in “The Marginal Value of Adaptive Gradient Methods in Machine Learning,” NeurIPS 2017, adaptive methods like Adam, AdaGrad, and RMSProp converge to solutions whose generalization performances were worse than those of SGD with momentum in several deep learning architectures they tested even if the former methods obtained lower loss values.
Adam vs. RMSprop and AdaGrad
The learning rate is adapted for each weight in RMSProp and AdaGrad by calculating the mean of the squares of the gradients over time, just like the second moment in Adam. They do not have any first moment term in their algorithms. Adam uses both of them together. Also, while AdaGrad and RMSProp do not use any form of bias correction, Adam makes use of the bias correction as well. The AdaGrad algorithm calculates the sum of the squares of the gradients throughout the entire training process without any decay, leading to gradual reduction in its learning rate and premature convergence during long training.
Adam vs. AdamW: The Weight Decay Difference
Loshchilov Ilya and Frank Hutter pinpointed a particular problem in the way plain Adam applies L2 regularization in the work “Decoupled Weight Decay Regularization” in 2017 and presented in ICLR 2019. In the usual version of Adam, an L2 penalty term is added to the gradient prior to calculating the momenta. This penalty term is then divided by √v̂_t in the last update step, just like the gradient is. Parameters with a big gradient history get a reduced weight decay penalty compared to the parameters with a smaller gradient history, and that is contrary to the purpose of weight decay.
AdamW takes the penalty out of the gradient and subtracts it from the parameter after the Adam update, making the weight decay independent of the gradient size:
λ is the weight-decay coefficient. PyTorch’s torch.optim.AdamW defaults weight_decay to 0.01; plain torch.optim.Adam defaults it to 0.
| Property | SGD | RMSprop | Adam | AdamW |
|---|---|---|---|---|
| Update direction basis | Raw gradient (momentum optional) | Raw gradient | Momentum (first moment) | Momentum (first moment) |
| Per-parameter adaptive step size | No | Yes | Yes | Yes |
| Bias correction | N/A | No | Yes | Yes |
| Weight decay handling | Coupled to gradient | Coupled to gradient | Coupled to gradient | Decoupled from gradient |
| Typical 2026 use case | Vision models where final generalization matters most | Recurrent networks, non-stationary problems | General-purpose default, rapid prototyping | Transformer and LLM pretraining and fine-tuning |
Advantages of the Adam Optimizer
- Per-parameter adaptive learning rates. Each step size is scaled based on the gradient history of that parameter, requiring less hyperparameter tuning than standard SGD does.
- Bias-corrected moment estimates. The correction step offsets the zero-initialization of m_0 and v_0, preventing artificially small updates in the first few training steps.
- Low memory relative to second-order methods. Adam needs only two additional state tensors per parameter (m and v); true second-order methods that use the Hessian matrix need memory proportional to the square of the parameter count.
- Documented empirical performance. In the original Adam paper, Adam was reported to be equal or superior in performance to AdaGrad, RMSProp, and SGD with Nesterov momentum in the areas of logistic regression, multilayer perceptron, and convolutional neural networks.
- Works with sparse and noisy gradients. The convergence proof and subsequent research prove that Adam can deal with sparse-gradient optimization problems (for example, NLP with a huge embedding table), unlike AdaGrad, which suffers from early learning rate decay.
Limitations of Adam (and When to Avoid It)
- Higher memory use than SGD. Adam stores the parameters plus two moment vectors, m and v, per parameter. That is three times the memory footprint of plain SGD, which stores only the parameters, and 1.5 times the footprint of SGD with momentum, which stores one extra vector.

- A documented generalization gap on some tasks. Wilson et al. (2017), cited above, found Adam-trained models generalized worse than SGD-with-momentum-trained models on several image classification benchmarks, despite reaching lower training loss. This gap does not appear consistently across all task types; the same paper’s construction shows it most clearly on overparameterized, separable classification problems.
- A convergence proof error, later corrected. Reddi, Kale, and Kumar, in “On the Convergence of Adam and Beyond” (ICLR 2018), showed that Adam’s original convergence proof contains an error and that Adam can fail to converge on certain convex optimization problems. They proposed AMSGrad as a fix, which several frameworks, including PyTorch and Keras, now offer as an optional flag on the standard Adam implementation.
- Weight decay requires AdamW, not plain Adam, for reliable regularization. As detailed above, applying an L2 penalty inside plain Adam produces a weight-decay strength that depends on each parameter’s gradient magnitude, an unintended and inconsistent effect.
- Choose SGD with momentum when final generalization matters more than training speed. For computer vision models where test-set accuracy is the primary metric and training-time budget is flexible, SGD with momentum remains a documented alternative worth testing against Adam or AdamW.
Where Adam Is Used in 2026
AdamW is the current optimizer of choice for language model pretraining, as established by an October 2025 large-scale comparison study of pretraining optimizers on OpenReview, which demonstrated that other optimizers boasting 1.4 to 2 times faster training compared to AdamW in reality provide closer to 1.1 time the speedup at the 1.2B parameters scale, when hyperparameters of both optimizers are reasonably tuned.
Some matrix-preconditioned optimizers, such as Muon and SOAP, have been used in pretraining some cutting-edge language models, like Kimi K2/K2.5 and GLM-4.5/4.7, as reported in late 2025. On the other hand, fine-tuning did not follow the trend. According to the findings of the 2026 paper “Can Muon Fine-Tune Adam-Pretrained Models?”, fine-tuning the Adam-pretrained model using Muon yields inferior results compared to continuing with Adam. This phenomenon is called the optimizer mismatch by the authors. Since the majority of the publicly available pretrained models are Adam- or AdamW-pretrained, the mismatch hinders practical implementation of Muon in fine-tuning pipelines.
Adam and AdamW continue to be the norm for finetuning computer vision, training generative adversarial networks, building recommendation system models, and research prototyping outside of language model pretraining because the fundamental assertion made by the original paper—that there is minimal tuning required to converge—continues to apply.
How to Implement Adam Optimizer in Code
Adam Optimizer in PyTorch
import torch
import torch.nn as nn
model = nn.Sequential(
nn.Linear(2, 16), nn.ReLU(),
nn.Linear(16, 8), nn.ReLU(),
nn.Linear(8, 1), nn.Sigmoid()
)
optimizer = torch.optim.Adam(
model.parameters(),
lr=0.001,
betas=(0.9, 0.999),
eps=1e-8
)
criterion = nn.BCELoss()
for epoch in range(50):
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
For weight decay, replace torch.optim.Adam with torch.optim.AdamW and pass weight_decay=0.01 (the PyTorch default for AdamW) or a value tuned to the task.
Adam Optimizer in TensorFlow/Keras
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
from tensorflow.keras.optimizers import Adam
model = Sequential([
Dense(16, activation='relu', input_shape=(2,)),
Dense(8, activation='relu'),
Dense(1, activation='sigmoid')
])
model.compile(
optimizer=Adam(learning_rate=0.001, beta_1=0.9, beta_2=0.999, epsilon=1e-7),
loss='binary_crossentropy',
metrics=['accuracy']
)
model.fit(X_train, y_train, epochs=50, batch_size=32, validation_split=0.2)
Note the explicit epsilon=1e-7 above. Keras’s default differs from PyTorch’s 1e-8, as shown in the hyperparameter table earlier in this guide; setting it explicitly removes that source of cross-framework variation.
Adam From Scratch in NumPy
This implementation applies the exact formulas from the “How Adam Optimizer Works” section, with no framework dependency:
import numpy as np
def adam_step(param, grad, m, v, t, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8):
m = beta1 * m + (1 - beta1) * grad
v = beta2 * v + (1 - beta2) * (grad ** 2)
m_hat = m / (1 - beta1 ** t)
v_hat = v / (1 - beta2 ** t)
param = param - lr * m_hat / (np.sqrt(v_hat) + eps)
return param, m, v
# Example: minimize f(x) = x^2, gradient = 2x
param = np.array([5.0])
m = np.zeros_like(param)
v = np.zeros_like(param)
for t in range(1, 101):
grad = 2 * param
param, m, v = adam_step(param, grad, m, v, t)
print(param) # converges toward 0.0
Common Adam Optimizer Mistakes and Troubleshooting
- Leaving epsilon at its framework default for large vision models. The TensorFlow documentation notes that a value of epsilon = 1e-7 “may not be a good default in general” and that, when training an inception-like model on ImageNet, using epsilon = 1.0 or 0.1 yields better results. Experiment with other epsilon values if your training gets stuck with large image models.
- Using plain Adam’s weight_decay parameter and expecting standard L2 regularization behavior. As described above, this depends on the gradient of individual parameters. Go for AdamW when you need consistent weight decay effects.
- Skipping Learning-rate warmup in Transformer training. Although there is bias correction, the moments calculated in the initial few iterations depend on very small numbers of gradient samples. A fixed base learning rate right from the start might lead to loss spikes; hence, a short warm-up period with linear or cosine decay is usually done while training Transformer models.
- Assuming default hyperparameters transfer identically across frameworks. As you can see in the above hyperparameter table, default values of epsilon in PyTorch and Keras differ. So the model trained in one framework will have to set epsilon explicitly to generate the same training curve in the other framework.
- Not adjusting beta2 for sparse-gradient problems. A beta2 value close to 1.0 was recommended by the authors for NLP and other sparse-gradient applications. Using default beta2 in these kinds of problems may result in noisier effective learning rates for less frequently updated parameters.
FAQs
Q1. Is Adam still the best optimizer in 2026?
Ans. There is no one universal optimizer that would be the best for all tasks. Adam and AdamW continue to be the default starting choice for most of the deep learning problems; AdamW is now used as the default optimizer for language model pretraining, and Muon is the optimizer that has been adopted by some advanced LLM models.
Q2. What is the difference between Adam and AdamW?
Ans. The main difference is that AdamW performs weight decay independently from the gradient-based optimization, while Adam uses weight decay with the gradient.
Q3. What do beta1 and beta2 control in Adam?
Ans. Beta1 is used to control the decay rate of the moving averages of gradients and is set at 0.9, while beta2 is used to control the decay rate of the moving averages of squared gradients and is set at 0.999.
Q4. Do you need a learning rate scheduler with Adam?
Ans. The Adam optimizer does not need a learning rate scheduler to reach convergence in many problems; however, in most large-scale training, especially in the case of Transformer pretraining, Adam or AdamW is always used in conjunction with the warmup-and-decay approach.
Q5. Why isn’t my model converging with Adam?
Ans. The reasons could be due to the following factors: an inappropriate epsilon value for the model’s size, an excessively large learning rate for the problem, the lack of learning rate warm-up for Transformers, and the use of the standard Adam weight decay instead of AdamW decoupled weight decay.