The Q-table contains one entry for every combination of state and action that can arise while an agent acts. In an Atari game where each frame has millions of different pixel combinations, there is no possibility for such a table to exist and learn from experience in memory. The solution to this problem is the introduction of the Deep Q-Network. This single change transformed Q-learning from an algorithm that was applicable to only small-scale discrete problems to the first algorithm able to play Atari games using just pixels.
What Will I Learn?
What Is a Deep Q-Network (DQN)?
Deep Q-Network (DQN) is a type of neural network which uses Q-learning for computing the expected future reward of each action within a certain state, thus being an alternative to Q-table in case of large state space that can’t be stored.
This type of network gets a state as an input (such as four stacked Atari frames), and produces one Q-value per each possible action. Next, the agent chooses the action with the maximum Q-value. The weights of the neural network θ are the parameters which are tuned during training. This specific network was first suggested in 2013 and improved in 2015 by DeepMind; these papers are discussed below.
Why Q-Learning Breaks Down at Scale
The traditional method of Q-learning has Q-value associated with every state action pair, which is stored in a table. This is effective where the number of states is few, say a 4 x 4 grid world having 16 states. However, as soon as the number of states becomes huge in numbers, in the range of millions, this approach falls short.
The Q-Table Problem in Plain Numbers
The DQN model takes Atari data as input, in four consecutive 84×84 grayscale images, where each pixel contains any of 256 intensity levels (Mnih et al., 2015). Such data can generate 256^(84×84×4) different possible states in theory, which cannot be indexed by any table or any machine. The neural network model completely avoids this issue. It does not keep one value for each individual state but rather learns a mapping function from any state (including those not encountered yet) to Q-value.
How a DQN Actually Works
DQNs consist of four elements functioning as a unit: a neural network, an experience replay buffer, a target network, and a loss function which connects all three.
The Neural Network (Function Approximator)
The network represents Q(s,a;θ), which is the expected return when performing action a from state s with current parameters θ. The network in the original Atari version uses the pixel frames as input and processes them using convolutional layers and then fully connected layers, with each output corresponding to the Q-value for every valid action (Mnih et al., 2015). There is no additional step of feature extraction as the features are learned based on rewards.
Experience Replay
The technique of experience replay works by storing every transition (s, a, r, s’) in a fixed-size memory and sampling mini-batches randomly from it when training occurs, rather than training on the transitions that happen sequentially. The technique is designed to eliminate the correlation between subsequent frames that stabilize training and enable each experience to be utilized more than once (Mnih et al., 2015). The technique of experience replay was initially proposed by Long-Ji Lin in 1992.
The Target Network
The target network is a duplicate of the main network, whose role is merely to calculate target Q values when training. The target network parameters, represented by θ⁻, are periodically copied from the main network, as opposed to updating them each time. This duplication prevents the network from trying to chase an ever-changing target, causing it to diverge (Mnih et al., 2015).
The Loss Function and the Bellman Equation
The DQN’s loss function is as follows:
In this equation, r stands for the instant reward, γ for the discount factor, s’ for the subsequent state, and θ⁻ are the parameters of the target network. This is the squared error of the prediction made by the network at the moment and its target defined by the so-called Bellman equation, which is the sum of the instant reward and the discounted future value of the optimal action to be taken from the subsequent state. The Bellman equation was introduced in 1957 by Richard Bellman in dynamic programming.
Where DQN Came From
The architecture of deep Q-network was first introduced by DeepMind in December 2013 in “Playing Atari with Deep Reinforcement Learning” (Mnih et al., 2013). This model was tested for its performance on seven Atari 2600 games in the Arcade Learning Environment. The model achieved better results than any other previous method in six out of seven games and was able to match human expert performance in three.
This extended model was published by DeepMind in Nature in February 2015 in “Human-level control through deep reinforcement learning” (Mnih et al., 2015, Nature 518, pp. 529–533). The model showed improved performance compared to the best previous reinforcement learning techniques on 43 of the 49 Atari 2600 games, while being tested with one fixed network architecture and one set of hyperparameters.
The DQN Training Loop, Step by Step
- Initialize main network (θ), target network (θ⁻), and empty experience replay buffer.
- Observe the current state, and take the action using the ε-greedy approach: randomly selected action with probability ε; action with maximum Q-value otherwise.
- Take the action, get the reward and new state, and add the transition into the experience replay buffer.
- Get a random mini-batch of transitions from the buffer.
- Calculate the target Q-value of every sample from mini-batch based on the target network.
- Update the weights of the main network using minimization of the loss function between the predicted Q-values and target Q-values.
- Copy weights of the main network into the target network after some number of iterations.
- Decrease value of ε over time, moving from exploration to exploitation.
- Repeat steps 2 through 8 until the stable performance of the agent is achieved.
Build a DQN From Scratch: CartPole Walkthrough
CartPole is a classic testbed task wherein the agent controls the pole balance atop a cart by pushing the cart left or right. The state space of CartPole is four-dimensional and it has two possible actions, which allows for quick training.
Setup
Both implementations below use Gymnasium, the maintained fork of OpenAI Gym, for the environment itself.
pip install gymnasium numpy
TensorFlow Version
import gymnasium as gym
import numpy as np
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
env = gym.make("CartPole-v1")
state_size = env.observation_space.shape[0]
action_size = env.action_space.n
model = Sequential([
Dense(24, activation="relu", input_shape=(state_size,)),
Dense(action_size, activation="linear")
])
model.compile(optimizer="adam", loss="mse")
gamma, epsilon, epsilon_decay, epsilon_min = 0.95, 1.0, 0.99, 0.01
for episode in range(50):
state, _ = env.reset()
state = state.reshape(1, state_size)
done = False
while not done:
if np.random.rand() < epsilon:
action = env.action_space.sample()
else:
action = np.argmax(model.predict(state, verbose=0))
next_state, reward, terminated, truncated, _ = env.step(action)
done = terminated or truncated
next_state = next_state.reshape(1, state_size)
target = reward if done else reward + gamma * np.max(model.predict(next_state, verbose=0))
q_values = model.predict(state, verbose=0)
q_values[0][action] = target
model.fit(state, q_values, epochs=1, verbose=0)
state = next_state
epsilon = max(epsilon_min, epsilon * epsilon_decay)
PyTorch Version
import gymnasium as gym
import torch
import torch.nn as nn
import torch.optim as optim
import random
from collections import deque
env = gym.make("CartPole-v1")
state_size = env.observation_space.shape[0]
action_size = env.action_space.n
class DQN(nn.Module):
def __init__(self, in_dim, out_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, 128), nn.ReLU(),
nn.Linear(128, 128), nn.ReLU(),
nn.Linear(128, out_dim)
)
def forward(self, x):
return self.net(x)
policy_net = DQN(state_size, action_size)
target_net = DQN(state_size, action_size)
target_net.load_state_dict(policy_net.state_dict())
optimizer = optim.Adam(policy_net.parameters(), lr=0.001)
memory = deque(maxlen=10000)
gamma, epsilon, epsilon_min, epsilon_decay = 0.99, 1.0, 0.01, 0.995
def select_action(state, epsilon):
if random.random() < epsilon:
return env.action_space.sample()
with torch.no_grad():
return torch.argmax(policy_net(torch.FloatTensor(state))).item()
Both implementations use the same four components covered above: a main network, a target network, a replay buffer, and an ε-greedy policy.
Reading Your Results (What “Reward” Actually Means)
The reward for each episode in CartPole is equivalent to the number of time-steps that the pole remains balanced. For an untrained network, this will generally be less than 20 for each episode. For an adequately trained network in this particular scenario, the score would have increased substantially beyond 50 episodes. However, remember that the maximum time-step for the CartPole-v1 game is set at 500.
DQN vs. Its Successors
DQN’s initial architecture had three known drawbacks: overestimation of Q-values, treating all the experiences equally, and the problem of a single network failing to distinguish “how good is this state” from “how good is this action.” There were three other studies which tackled these issues.
| Algorithm | Paper | Problem addressed | Reported result |
|---|---|---|---|
| Double DQN | van Hasselt, Guez & Silver, 2016 (AAAI) | Overestimation of Q-values from using one network to both select and evaluate actions | Reduced overoptimistic value estimates observed in standard DQN |
| Prioritized Experience Replay | Schaul, Quan, Antonoglou & Silver, 2016 (ICLR) | Uniform sampling treats a rare, high-error transition the same as a routine one | Outperformed uniform-replay DQN on 41 of 49 Atari games |
| Dueling DQN | Wang, Schaul, Hessel, van Hasselt, Lanctot & de Freitas, 2016 (ICML) | A single Q-value output cannot separate state value from action advantage | Improved policy performance over standard DQN across the Atari benchmark suite |
| Rainbow | Hessel et al., 2018 (AAAI) | No single prior variant combined all known DQN improvements | Combined six DQN extensions, including the three above, into one agent |
Real-World Applications of DQN
- Game Playing Agents. DQN was evaluated on 49 Atari 2600 games using the Arcade Learning Environment in which DQN achieved performance comparable to that of humans playing those games professionally (Mnih et al., 2015).
- Robotics. Grasping and placing robotic arm manipulations learn from DQN and its variations without requiring an engineered controller policy.
- Simulation of autonomous vehicles. Driving agents in simulations utilize DQN to choose between discrete maneuvers in constrained tests for driving prior to any actual implementation.
- Finance. Research in algorithmic trading uses DQN to make discrete choices of buying, selling, and holding actions on the basis of state information from the stock market as well as reward in terms of returns of the investment portfolio.
- Healthcare. Treatment scheduling research uses DQN to choose among a pre-defined set of treatment choices depending on state information of patients.
- Telecommunication and networking. Surveys of deep reinforcement learning for the Internet of Things (IoT) and mobile networking applications report studies that used DQN to implement decisions on routing, switching cells, and allocating resources.
- Traffic Signal Control. Traffic control for urban traffic uses DQN to choose traffic timing actions based on traffic states.
- Medical imaging. Deep Q-network (DQN) and other deep reinforcement learning approaches for value function learning have been employed for medical image processing tasks like selecting regions to scan next.
Advantages of DQN
- Manages state spaces that are too big for a Q-table by estimating Q-values using a neural network rather than memorizing Q-values.
- Directly learns from inputs like raw pixels without having to manually engineer features.
- Uses past experience through the replay memory, allowing more efficient learning than from one time through a single transition.
- Increases training stability through the use of the target network, which keeps the target of the update stable across time steps.
Limitations of DQN — And What to Use Instead
- High compute requirements. Training a DQN in any nontrivial environment involves using a GPU and thousands of interactions with the environment.
- Sample inefficiency relative to newer methods. DQN requires a lot of interactions to learn an efficient policy.
- Discrete action spaces only. In standard DQN, actions are selected from a finite list. DQN is incapable of selecting a continuous action, e.g., a steering wheel angle or the motor torque. Some well-known alternatives to select actions for continuous action spaces include Deep Deterministic Policy Gradient (Lillicrap et al., 2016), Proximal Policy Optimization (Schulman et al., 2017), and Soft Actor-Critic (Haarnoja et al., 2018).
- Sensitivity to hyperparameters. Learning rate, buffer size, and frequency of updates influence the convergence of the learning process.
Common DQN Training Problems (And How to Fix Them)
| Problem | Likely cause | Fix |
|---|---|---|
| Reward stays flat across episodes | ε decayed too fast, or replay buffer too small | Slow the ε decay rate; increase buffer size |
| Q-values grow toward very large numbers | Target network updated too frequently, or reward not clipped | Increase target-update interval; clip rewards to a fixed range |
| Training works, then suddenly collapses | Replay buffer too small relative to state diversity | Increase buffer size so recent experience does not overwrite older diverse experience |
| Agent repeats one action regardless of state | ε decayed to near zero too early, network stuck in a local pattern | Reset ε partially; verify state normalization |
Is DQN Still Worth Learning in 2026?
DQN continues to be the default first algorithm that is always covered in reinforcement learning classes, while the Atari test suite that was proposed by DQN is also used as the default benchmark for testing novel algorithms (Mnih et al., 2015). In production systems with continuous actions, approaches such as policy gradients (PPO and SAC) are used more often than DQN. However, for the discrete case and learning the basics behind all future value-based approaches, DQN is the default choice.
FAQ
Q1. What does DQN stand for?
Ans. DQN refers to Deep Q-Network, a neural network that is trained for learning the Q-value function approximation in Q-learning.
Q2. Is DQN the same as Q-learning?
Ans. No – DQN represents Q-learning using a function approximator in the form of a neural network; regular Q-learning keeps Q-values in tables.
Q3. Can DQN handle continuous action spaces?
Ans. No – DQN performs action selection among the available set of actions, while continuous action space problems require approaches like DDPG, PPO, or SAC.
Q3. What is experience replay in DQN?
Ans. No – DQN performs action selection among the available set of actions, while continuous action space problems require approaches like DDPG, PPO, or SAC.
Q4. What is experience replay in DQN?
Ans. The experience replay buffer contains all past transition data, and it uses random sampling of transitions to decouple consecutive experiences in training.
Q5. Why does DQN use a target network?
Ans. A target network stores the delayed copy of the main network’s weight vector, which prevents changing the training target after each iteration and thus oscillations.
Q6. What is the difference between DQN and Double DQN?
Ans. Double DQN consists of two networks, one for action selection and the other for value approximation, and therefore solves the problem of Q-value overestimation in regular DQN (van Hasselt et al., 2016).