Deep Q-Networks (DQN)

|
11 min read
|
29 views
Deep Q-Networks

The Q-table contains one entry for every combination of state and action that can arise while an agent acts. In an Atari game where each frame has millions of different pixel combinations, there is no possibility for such a table to exist and learn from experience in memory. The solution to this problem is the introduction of the Deep Q-Network. This single change transformed Q-learning from an algorithm that was applicable to only small-scale discrete problems to the first algorithm able to play Atari games using just pixels.

Deep Q-Networks

What Is a Deep Q-Network (DQN)?

Deep Q-Network (DQN) is a type of neural network which uses Q-learning for computing the expected future reward of each action within a certain state, thus being an alternative to Q-table in case of large state space that can’t be stored.

This type of network gets a state as an input (such as four stacked Atari frames), and produces one Q-value per each possible action. Next, the agent chooses the action with the maximum Q-value. The weights of the neural network θ are the parameters which are tuned during training. This specific network was first suggested in 2013 and improved in 2015 by DeepMind; these papers are discussed below.

Why Q-Learning Breaks Down at Scale

The traditional method of Q-learning has Q-value associated with every state action pair, which is stored in a table. This is effective where the number of states is few, say a 4 x 4 grid world having 16 states. However, as soon as the number of states becomes huge in numbers, in the range of millions, this approach falls short.

The Q-Table Problem in Plain Numbers

The DQN model takes Atari data as input, in four consecutive 84×84 grayscale images, where each pixel contains any of 256 intensity levels (Mnih et al., 2015). Such data can generate 256^(84×84×4) different possible states in theory, which cannot be indexed by any table or any machine. The neural network model completely avoids this issue. It does not keep one value for each individual state but rather learns a mapping function from any state (including those not encountered yet) to Q-value.

Q-Table

How a DQN Actually Works

DQNs consist of four elements functioning as a unit: a neural network, an experience replay buffer, a target network, and a loss function which connects all three.

The Neural Network (Function Approximator)

The network represents Q(s,a;θ), which is the expected return when performing action a from state s with current parameters θ. The network in the original Atari version uses the pixel frames as input and processes them using convolutional layers and then fully connected layers, with each output corresponding to the Q-value for every valid action (Mnih et al., 2015). There is no additional step of feature extraction as the features are learned based on rewards.

Experience Replay

The technique of experience replay works by storing every transition (s, a, r, s’) in a fixed-size memory and sampling mini-batches randomly from it when training occurs, rather than training on the transitions that happen sequentially. The technique is designed to eliminate the correlation between subsequent frames that stabilize training and enable each experience to be utilized more than once (Mnih et al., 2015). The technique of experience replay was initially proposed by Long-Ji Lin in 1992.

The Target Network

The target network is a duplicate of the main network, whose role is merely to calculate target Q values when training. The target network parameters, represented by θ⁻, are periodically copied from the main network, as opposed to updating them each time. This duplication prevents the network from trying to chase an ever-changing target, causing it to diverge (Mnih et al., 2015).

The Loss Function and the Bellman Equation

The DQN’s loss function is as follows:

L(θ)=E[(r+γ·max(Q(s′,a′;θ−))−Q(s,a;θ))2]L(θ) = E[(r + γ·max(Q(s’, a’; θ⁻)) − Q(s, a; θ))²]

In this equation, r stands for the instant reward, γ for the discount factor, s’ for the subsequent state, and θ⁻ are the parameters of the target network. This is the squared error of the prediction made by the network at the moment and its target defined by the so-called Bellman equation, which is the sum of the instant reward and the discounted future value of the optimal action to be taken from the subsequent state. The Bellman equation was introduced in 1957 by Richard Bellman in dynamic programming.

Deep Q-Networks

Where DQN Came From

The architecture of deep Q-network was first introduced by DeepMind in December 2013 in “Playing Atari with Deep Reinforcement Learning” (Mnih et al., 2013). This model was tested for its performance on seven Atari 2600 games in the Arcade Learning Environment. The model achieved better results than any other previous method in six out of seven games and was able to match human expert performance in three.

This extended model was published by DeepMind in Nature in February 2015 in “Human-level control through deep reinforcement learning” (Mnih et al., 2015, Nature 518, pp. 529–533). The model showed improved performance compared to the best previous reinforcement learning techniques on 43 of the 49 Atari 2600 games, while being tested with one fixed network architecture and one set of hyperparameters.

The DQN Training Loop, Step by Step

  1. Initialize main network (θ), target network (θ⁻), and empty experience replay buffer.
  2. Observe the current state, and take the action using the ε-greedy approach: randomly selected action with probability ε; action with maximum Q-value otherwise.
  3. Take the action, get the reward and new state, and add the transition into the experience replay buffer.
  4. Get a random mini-batch of transitions from the buffer.
  5. Calculate the target Q-value of every sample from mini-batch based on the target network.
  6. Update the weights of the main network using minimization of the loss function between the predicted Q-values and target Q-values.
  7. Copy weights of the main network into the target network after some number of iterations.
  8. Decrease value of ε over time, moving from exploration to exploitation.
  9. Repeat steps 2 through 8 until the stable performance of the agent is achieved.

Build a DQN From Scratch: CartPole Walkthrough

CartPole is a classic testbed task wherein the agent controls the pole balance atop a cart by pushing the cart left or right. The state space of CartPole is four-dimensional and it has two possible actions, which allows for quick training.

Setup

Both implementations below use Gymnasium, the maintained fork of OpenAI Gym, for the environment itself.

pip install gymnasium numpy

TensorFlow Version

import gymnasium as gym

import numpy as np

import tensorflow as tf

from tensorflow.keras.models import Sequential

from tensorflow.keras.layers import Dense

env = gym.make("CartPole-v1")

state_size = env.observation_space.shape[0]

action_size = env.action_space.n

model = Sequential([

    Dense(24, activation="relu", input_shape=(state_size,)),

    Dense(action_size, activation="linear")

])

model.compile(optimizer="adam", loss="mse")

gamma, epsilon, epsilon_decay, epsilon_min = 0.95, 1.0, 0.99, 0.01

for episode in range(50):

    state, _ = env.reset()

    state = state.reshape(1, state_size)

    done = False

    while not done:

        if np.random.rand() < epsilon:

            action = env.action_space.sample()

        else:

            action = np.argmax(model.predict(state, verbose=0))

        next_state, reward, terminated, truncated, _ = env.step(action)

        done = terminated or truncated

        next_state = next_state.reshape(1, state_size)

        target = reward if done else reward + gamma * np.max(model.predict(next_state, verbose=0))

        q_values = model.predict(state, verbose=0)

        q_values[0][action] = target

        model.fit(state, q_values, epochs=1, verbose=0)

        state = next_state

    epsilon = max(epsilon_min, epsilon * epsilon_decay)

PyTorch Version

import gymnasium as gym

import torch

import torch.nn as nn

import torch.optim as optim

import random

from collections import deque

env = gym.make("CartPole-v1")

state_size = env.observation_space.shape[0]

action_size = env.action_space.n

class DQN(nn.Module):

    def __init__(self, in_dim, out_dim):

        super().__init__()

        self.net = nn.Sequential(

            nn.Linear(in_dim, 128), nn.ReLU(),

            nn.Linear(128, 128), nn.ReLU(),

            nn.Linear(128, out_dim)

        )

    def forward(self, x):

        return self.net(x)

policy_net = DQN(state_size, action_size)

target_net = DQN(state_size, action_size)

target_net.load_state_dict(policy_net.state_dict())

optimizer = optim.Adam(policy_net.parameters(), lr=0.001)

memory = deque(maxlen=10000)

gamma, epsilon, epsilon_min, epsilon_decay = 0.99, 1.0, 0.01, 0.995

def select_action(state, epsilon):

    if random.random() < epsilon:

        return env.action_space.sample()

    with torch.no_grad():

        return torch.argmax(policy_net(torch.FloatTensor(state))).item()

Both implementations use the same four components covered above: a main network, a target network, a replay buffer, and an ε-greedy policy.

Reading Your Results (What “Reward” Actually Means)

The reward for each episode in CartPole is equivalent to the number of time-steps that the pole remains balanced. For an untrained network, this will generally be less than 20 for each episode. For an adequately trained network in this particular scenario, the score would have increased substantially beyond 50 episodes. However, remember that the maximum time-step for the CartPole-v1 game is set at 500.

Deep Q-Networks

DQN vs. Its Successors

DQN’s initial architecture had three known drawbacks: overestimation of Q-values, treating all the experiences equally, and the problem of a single network failing to distinguish “how good is this state” from “how good is this action.” There were three other studies which tackled these issues.

AlgorithmPaperProblem addressedReported result
Double DQNvan Hasselt, Guez & Silver, 2016 (AAAI)Overestimation of Q-values from using one network to both select and evaluate actionsReduced overoptimistic value estimates observed in standard DQN
Prioritized Experience ReplaySchaul, Quan, Antonoglou & Silver, 2016 (ICLR)Uniform sampling treats a rare, high-error transition the same as a routine oneOutperformed uniform-replay DQN on 41 of 49 Atari games
Dueling DQNWang, Schaul, Hessel, van Hasselt, Lanctot & de Freitas, 2016 (ICML)A single Q-value output cannot separate state value from action advantageImproved policy performance over standard DQN across the Atari benchmark suite
RainbowHessel et al., 2018 (AAAI)No single prior variant combined all known DQN improvementsCombined six DQN extensions, including the three above, into one agent

Real-World Applications of DQN

  • Game Playing Agents. DQN was evaluated on 49 Atari 2600 games using the Arcade Learning Environment in which DQN achieved performance comparable to that of humans playing those games professionally (Mnih et al., 2015).
  • Robotics. Grasping and placing robotic arm manipulations learn from DQN and its variations without requiring an engineered controller policy.
  • Simulation of autonomous vehicles. Driving agents in simulations utilize DQN to choose between discrete maneuvers in constrained tests for driving prior to any actual implementation.
  • Finance. Research in algorithmic trading uses DQN to make discrete choices of buying, selling, and holding actions on the basis of state information from the stock market as well as reward in terms of returns of the investment portfolio.
  • Healthcare. Treatment scheduling research uses DQN to choose among a pre-defined set of treatment choices depending on state information of patients.
  • Telecommunication and networking. Surveys of deep reinforcement learning for the Internet of Things (IoT) and mobile networking applications report studies that used DQN to implement decisions on routing, switching cells, and allocating resources.
  • Traffic Signal Control. Traffic control for urban traffic uses DQN to choose traffic timing actions based on traffic states.
  • Medical imaging. Deep Q-network (DQN) and other deep reinforcement learning approaches for value function learning have been employed for medical image processing tasks like selecting regions to scan next.
Deep Q-Networks

Advantages of DQN

  • Manages state spaces that are too big for a Q-table by estimating Q-values using a neural network rather than memorizing Q-values.
  • Directly learns from inputs like raw pixels without having to manually engineer features.
  • Uses past experience through the replay memory, allowing more efficient learning than from one time through a single transition.
  • Increases training stability through the use of the target network, which keeps the target of the update stable across time steps.

Limitations of DQN — And What to Use Instead

  • High compute requirements. Training a DQN in any nontrivial environment involves using a GPU and thousands of interactions with the environment.
  • Sample inefficiency relative to newer methods. DQN requires a lot of interactions to learn an efficient policy.
  • Discrete action spaces only. In standard DQN, actions are selected from a finite list. DQN is incapable of selecting a continuous action, e.g., a steering wheel angle or the motor torque. Some well-known alternatives to select actions for continuous action spaces include Deep Deterministic Policy Gradient (Lillicrap et al., 2016), Proximal Policy Optimization (Schulman et al., 2017), and Soft Actor-Critic (Haarnoja et al., 2018).
  • Sensitivity to hyperparameters. Learning rate, buffer size, and frequency of updates influence the convergence of the learning process.
Deep Q-Networks

Common DQN Training Problems (And How to Fix Them)

ProblemLikely causeFix
Reward stays flat across episodesε decayed too fast, or replay buffer too smallSlow the ε decay rate; increase buffer size
Q-values grow toward very large numbersTarget network updated too frequently, or reward not clippedIncrease target-update interval; clip rewards to a fixed range
Training works, then suddenly collapsesReplay buffer too small relative to state diversityIncrease buffer size so recent experience does not overwrite older diverse experience
Agent repeats one action regardless of stateε decayed to near zero too early, network stuck in a local patternReset ε partially; verify state normalization

Is DQN Still Worth Learning in 2026?

DQN continues to be the default first algorithm that is always covered in reinforcement learning classes, while the Atari test suite that was proposed by DQN is also used as the default benchmark for testing novel algorithms (Mnih et al., 2015). In production systems with continuous actions, approaches such as policy gradients (PPO and SAC) are used more often than DQN. However, for the discrete case and learning the basics behind all future value-based approaches, DQN is the default choice.

FAQ

Q1. What does DQN stand for? 

Ans. DQN refers to Deep Q-Network, a neural network that is trained for learning the Q-value function approximation in Q-learning.

Q2. Is DQN the same as Q-learning? 

Ans. No – DQN represents Q-learning using a function approximator in the form of a neural network; regular Q-learning keeps Q-values in tables.

Q3. Can DQN handle continuous action spaces? 

Ans. No – DQN performs action selection among the available set of actions, while continuous action space problems require approaches like DDPG, PPO, or SAC.

Q3. What is experience replay in DQN? 

Ans. No – DQN performs action selection among the available set of actions, while continuous action space problems require approaches like DDPG, PPO, or SAC.

Q4. What is experience replay in DQN? 

Ans. The experience replay buffer contains all past transition data, and it uses random sampling of transitions to decouple consecutive experiences in training.

Q5. Why does DQN use a target network? 

Ans. A target network stores the delayed copy of the main network’s weight vector, which prevents changing the training target after each iteration and thus oscillations.

Q6. What is the difference between DQN and Double DQN? 

Ans. Double DQN consists of two networks, one for action selection and the other for value approximation, and therefore solves the problem of Q-value overestimation in regular DQN (van Hasselt et al., 2016).

Shalki Aggarwal is a Software Engineer II at Microsoft and an AI & Data Science expert specializing in Generative AI, Agentic AI, Python, LangChain, LangGraph, CrewAI, Deep Agents, and Loop Engineering. She is also a corporate trainer for leading organizations including L&T, Bharat Petroleum, Luminous, Denso, and Toshiba Midea, helping teams apply AI and emerging technologies to real-world business challenges.