Deep learning interviews check for three criteria: If the candidate knows the reason behind a particular technique, If the candidate is able to explain the technique without merely stating the definition, and If the candidate knows the appropriate scenario where the technique should be used over another. This guide answers 50 deep learning interview questions classified on the basis of seniority levels – Fresher, Mid-Level, and Senior – and domain within each level.
Each question starts off with the definition followed by the numerical aspect, formula, or comparison that the interviewer expects as a follow-up. A diagram is provided for every architecture. For every question which needs a code to make the understanding easier than the textual explanation, code is provided too.
Fortune Business Insights forecasts that the global deep learning market will expand from $48.03 billion in 2026 to $342.34 billion by 2034, registering a compound annual growth rate of 27.83%. It is due to this reason that there is a continuous trend of deep learning jobs like machine learning engineer, computer vision engineer, NLP engineer, AI research scientist, etc appearing across hiring platforms more often than general software roles.
What Will I Learn?
How This Guide Is Organized
The questions are classified in three levels:
- Beginner-level – neural network basics, training process and all terms which have to be defined properly by the candidate.
- Intermediate-level – specific architecture questions about CNN, RNN, LSTM, regularization for candidates who have 1–4 years of practical experience.
- Advanced-level – questions about transformers, generative architectures and deployment in production for candidates who are supposed to choose the architecture themselves, not implement it.
Every answer contains three parts: definition of the term, explanation of the principle and comparative table/program code when applicable.
Tier 1: Fresher-Level — Neural Network Fundamentals
What is deep learning, and how is it different from machine learning?
Deep Learning is the subfield of Machine Learning that utilizes neural networks comprising three or more layers to directly extract the features out of the raw data without having to manually engineer any features beforehand. In the case of Machine Learning, a person needs to pick and build features on his own.
These differences are exhibited in three ways:
| Attribute | Machine Learning | Deep Learning |
|---|---|---|
| Feature extraction | Manual, done by the data scientist | Automatic, learned during training |
| Data volume needed | Works with hundreds to thousands of rows | Typically needs tens of thousands of examples or more |
| Compute requirement | Runs on a standard CPU | Usually requires a GPU or TPU |
| Common algorithms | Linear regression, decision trees, random forests, SVM | CNNs, RNNs, Transformers, autoencoders |
A random forest predicting house prices from 12 structured columns is machine learning. A convolutional network classifying house conditions from photos is deep learning.
What is a neuron, and how does an artificial neural network model it?
The artificial neuron receives several numeric inputs, multiplies each by a weight and adds a bias to the sum, and then applies the activation function. The output of one neuron serves as the input for neurons in the next layer.
The computation for a single neuron with inputs x₁ and x₂, weights w₁ and w₂, and bias b is:
The value of z is passed to an activation function, which determines the output of the neuron. Such a construction is similar to that of a biological neuron, whose dendrites receive input, the body accumulates it, and the axon passes it forward, except it is a mathematical construct rather than a biological one.
Figure 1: An artificial neural network with an input layer, two hidden layers, and an output layer. Every line represents a weighted connection.
What do the input, hidden, and output layers do?
Input layer feeds data, hidden layers process this data using weights and activation function, and output layer generates the output of this data processing. In order to be considered as a neural network, there must be at least one hidden layer in the network; if there are two or more hidden layers, it is termed deep neural network.
There are three types of layers in all feedforward neural networks:
- Input Layer — one neuron for each of the input features. In the case of an image with 784 pixels, there would be 784 input neurons when we convert it into a single-dimensional array.
- Hidden Layers — every layer is made up of some weights, bias, and an activation function. With the help of more hidden layers, the network can represent more complex functions; however, it will have more parameters.
- Output Layer — the number of neurons and the type of activation function depends on the task. Binary classification requires one neuron with a sigmoid activation function, while multiclass requires one neuron for each class with a softmax function.
What are weights and biases in a neural network?
Weights are numerical quantities that define the effect one neuron’s output has on the following neuron. Bias is another quantity that is included before the application of the activation function. Weights and biases are learned during training using backpropagation.
Weights regulate the steepness of the input-output relation, while the bias allows this relation to be shifted independently of the input. In the absence of bias terms, the output of every neuron would go through the origin point, making the representation too restrictive.
How are weights initialized in a neural network, and why does it matter?
Initial value of weights defines the initial value of weights of a neural network before the training begins, and wrong choice of initialization makes the network fail to converge. There are four types of weight initialization.
- Zero Initialization – the weights are initialized with value zero. This makes the gradient of all neurons in a layer identical, and all neurons learn the same thing. This type of weight initialization is never used in practice.
- Random Weight Initialization – the weights are taken from random distribution. This solves the problem of symmetry, but wrong range selection could make the values vanish or explode.
- Xavier (Glorot) Initialization – the weights are taken from a distribution scaled by the number of input and output neurons in a layer. This was proposed by Glorot and Bengio in 2010 for sigmoid or tanh activation functions.
- He Initialization – it is the same as the Xavier one but scaled for ReLU activation function. This was proposed by He et al. in 2015 due to the nature of ReLU that zeros half of its inputs.
What is an activation function, and which functions are used most often?
Activation function refers to the mathematical function applied to the summation value of a neuron that introduces non-linearity, enabling the network to learn models which are non-linear in nature. The absence of an activation function would lead to the same effect being produced regardless of the number of layers stacked up in the network.
Four activation functions cover most use cases:
Figure 2: Sigmoid, Tanh, ReLU, and Leaky ReLU plotted by output range.
| Function | Output Range | Typical Use |
|---|---|---|
| Sigmoid | 0 to 1 | Output layer of a binary classifier |
| Tanh | -1 to 1 | Hidden layers where zero-centered output helps training |
| ReLU | 0 to input value | Default choice for hidden layers in most networks |
| Softmax | 0 to 1, summing to 1 across outputs | Output layer of a multi-class classifier |
The activation function used in hidden layers by default is ReLU because it is fast and efficient and minimizes the vanishing gradient effect that occurs in sigmoid and tanh functions. The drawback of using ReLU is the “dead ReLU,” which means that if the output of a neuron from the previous layer is negative, then the neuron will not be updated anymore since the gradient of ReLU is 0 for all negative input values.
Sigmoid vs. Softmax — when do you use which?
The sigmoid function generates an independent probability for each class and is applied in binary and multi-label classification tasks, while the softmax function generates a probability distribution among all classes which add up to 1 and is applied in multi-class classification tasks.
| Attribute | Sigmoid | Softmax |
|---|---|---|
| Output interpretation | Independent probability per output | Joint probability distribution across classes |
| Sum of outputs | Not constrained | Always equals 1 |
| Task type | Binary or multi-label classification | Single-label, multi-class classification |
| Paired loss function | Binary cross-entropy | Categorical cross-entropy |
Since the model assigns more than one appropriate tag to an image such as “outdoor,” “daytime,” and “people presence,” the sigmoid function is used here because these labels do not overlap. When a number between 0 and 9 is to be identified, softmax is used.
Tier 1: Fresher-Level — How Networks Learn
What is forward propagation and backward propagation?
In forward propagation, the data is fed into the network to generate a prediction while the backward propagation process computes the gradient of the error w.r.t each weight and then adjusts the weights accordingly to minimize the loss.
The forward propagation process involves three phases which are:
Multiplying the input values by the weights and adding the biases for every neuron in the network, applying the activation function and forwarding the outputs to the next neurons till the output layer generates the prediction.
The backward propagation process involves three phases which are:
Comparing the prediction with the actual value using a loss function, computing the gradients of the loss w.r.t each weight using the chain rule in calculus and updating each weight to minimize the loss.
What is a loss function, and which one should you use?
Loss functions evaluate the distance between the predictions made by the model and the actual target value, and the right one to use varies depending on the task. There are four main loss functions for deep learning tasks.
| Loss Function | Task | Formula Behavior |
|---|---|---|
| Mean Squared Error (MSE) | Regression | Penalizes larger errors more heavily due to squaring |
| Binary Cross-Entropy | Binary classification | Penalizes confident wrong predictions heavily |
| Categorical Cross-Entropy | Multi-class classification, one-hot labels | Compares predicted distribution to one-hot true distribution |
| Sparse Categorical Cross-Entropy | Multi-class classification, integer labels | Same as categorical cross-entropy without requiring one-hot encoding |
The value computed by both categorical cross-entropy and sparse categorical cross-entropy functions is identical; the only difference lies in the labeling method used. Sparse categorical cross-entropy can work with integer labels directly, saving memory in case of having many classes.
What is gradient descent, and what are its three main variants?
Gradient descent is a mathematical technique used to minimize the value of the objective function (loss function) by adjusting the parameters of the machine learning model, as represented in θ = θ − η × ∇L(θ), where η represents the learning rate and ∇L(θ) represents the gradient of the loss. There are three types of gradient descent.
- Batch Gradient Descent – Gradient computation with the complete training data before each weight update. It results in more precise and stable gradient calculations; however, it is computationally expensive for large data sets.
- Stochastic Gradient Descent (SGD) – Gradient computation with one example from the training set in each update. It is faster but results in a lot of noise in update calculations.
- Mini-Batch Gradient Descent – Gradient computation with the subset of training data between 32 and 256 examples in each update. It is a preferred algorithm in practical use since it combines stability of updates with efficiency in computation due to hardware limitations.
What are Adam, RMSProp, and Adagrad, and why is Adam the default choice?
Adam, RMSProp, and Adagrad are optimization algorithms that adjust the learning rate of each weight individually rather than using one learning rate across all weights. Adam has become the default algorithm due to its effectiveness; it works well in the combination of two other algorithms’ properties.
- Adagrad — This algorithm was developed by Duchi et al. in 2011. It adjusts the learning rate depending on the frequency of updates of each parameter: the learning rate is decreased for frequently updated parameters and increased for rarely updated ones. However, the problem of this algorithm is that the learning rate decreases monotonically and can become zero during the training process.
- RMSProp — Developed by Geoffrey Hinton in unpublished lecture notes in 2012. It solves the problem of Adagrad by using the exponentially weighted average of recent squared gradients.
- Adam — This algorithm was developed by Kingma and Ba in 2015. Adam combines the adaptive learning rate from RMSProp and momentum that tracks the exponentially decayed average of previous gradients.
A short PyTorch example showing how the optimizer choice is a one-line change:
import torch.optim as optim
# Stochastic Gradient Descent
optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
# Adam -- the common default
optimizer = optim.Adam(model.parameters(), lr=0.001, betas=(0.9, 0.999))
What is the learning rate, and what happens when it is set incorrectly?
The learning rate is a hyperparameter that controls the size of each weight update during training, and setting it too high or too low prevents the model from training effectively.
Figure 3: Three learning-rate settings and their effect on training loss over time.
A learning rate set too high causes the loss to oscillate or increase over time, because each update overshoots the minimum of the loss function. A learning rate set too low causes training to converge correctly but too slowly, wasting compute time and risking an early stop before the model reaches its best performance. A learning rate finder — which trains for a few hundred steps across a range of learning rates and plots the resulting loss — is a standard method for choosing an initial value.
What is the vanishing gradient problem, and how is it fixed?
The vanishing gradient problem occurs when gradients become extremely small as they are propagated backward through many layers, which causes the earliest layers of a deep network to stop learning. It is most common in deep networks using sigmoid or tanh activations, because both functions have derivatives bounded between 0 and 0.25 (sigmoid) or 0 and 1 (tanh), and multiplying many small derivatives together during backpropagation drives the gradient toward zero.
Four techniques address vanishing gradients:
- Replace sigmoid or tanh with ReLU in hidden layers, since ReLU’s derivative is 1 for all positive inputs.
- Apply batch normalization to keep activations in a stable range throughout training.
- Use residual (skip) connections, which give gradients a direct path backward that bypasses the multiplication chain.
- Initialize weights with Xavier or He initialization to prevent gradients from starting too small.
The exploding gradient problem is the inverse case: gradients grow exponentially during backpropagation, causing unstable, oscillating weight updates. Gradient clipping — capping the gradient at a fixed maximum value before applying the update — is the standard fix.
What is the difference between a parameter and a hyperparameter?
A parameter is a value the model learns during training, such as a weight or bias; a hyperparameter is a value set before training begins that controls how the model learns, such as the learning rate or batch size.
| Parameters | Hyperparameters | |
|---|---|---|
| Set by | The training algorithm | The engineer, before training |
| Examples | Weights, biases | Learning rate, batch size, number of layers, number of epochs |
| Changes during training | Yes | No |
Confusing the two is a common mistake among candidates early in their careers: a learning rate of 0.001 is a hyperparameter choice, while the specific weight values a network converges to after training are parameters.
Tier 2: Mid-Level — Avoiding Overfitting
What is the bias-variance tradeoff, and what do overfitting and underfitting look like?
The bias-variance tradeoff describes the balance between a model that is too simple to capture patterns in the data (high bias, underfitting) and a model that captures noise along with real patterns (high variance, overfitting). A well-trained model sits between the two.
Figure 4: Training error keeps falling as model complexity increases, but validation error rises again once the model starts memorizing noise.
Underfitting shows as poor performance on both training and validation data. Common causes include too few parameters, insufficient training time, or excessive regularization. Overfitting shows as strong training performance paired with weak validation performance. Common causes include a model with more capacity than the dataset requires, training for too many epochs, or too little regularization.
How do L1 and L2 regularization differ?
L1 and L2 regularization both add a penalty term to the loss function to discourage large weights, but L1 can force weights to exactly zero while L2 only shrinks them toward zero.
| Attribute | L1 (Lasso) | L2 (Ridge) |
|---|---|---|
| Penalty term | Sum of absolute weight values | Sum of squared weight values |
| Effect on weights | Can reduce weights to exactly 0 | Shrinks weights toward 0, rarely reaches it |
| Result | Sparse model; performs feature selection | Smoother, more evenly distributed weights |
| Best used when | Only a subset of features matter | Most features contribute some signal |
L1’s tendency to zero out weights makes it useful when the engineer expects only a subset of input features to be relevant. L2 is the more common default because it produces more stable training and rarely underperforms L1 by a meaningful margin.
What is dropout, and how does it prevent overfitting?
Dropout is a regularization technique that randomly disables a fraction of neurons during each training step, which prevents the network from relying too heavily on any single neuron. Srivastava et al. introduced dropout in a 2014 paper published in the Journal of Machine Learning Research.
During training, each neuron in a layer is dropped with a fixed probability, typically between 0.2 and 0.5. During inference, no neurons are dropped, but the layer’s outputs are scaled down to match the expected value seen during training. Dropout is applied to hidden layers and is rarely applied to the input or output layers.
What is batch normalization, and why does it speed up training?
Batch normalization normalizes the inputs to a layer so they have a mean of 0 and a standard deviation of 1 within each mini-batch, which reduces internal covariate shift and allows higher learning rates. Ioffe and Szegedy introduced batch normalization in 2015.
The technique computes the mean and variance of each mini-batch, normalizes the layer’s inputs using those statistics, then applies two learnable parameters — a scale (γ) and a shift (β) — so the network retains the ability to represent the original distribution if that is optimal. Batch normalization also provides a mild regularization effect, which sometimes reduces the need for dropout in the same network.
What is data augmentation, and what techniques does it use?
Data augmentation artificially increases the size and diversity of a training dataset by applying transformations to existing examples, which reduces overfitting when the original dataset is small.
Image data augmentation commonly uses rotation, horizontal or vertical flipping, zooming, translation, and brightness adjustment. Text data augmentation uses synonym replacement, random word insertion or deletion, and back-translation — translating a sentence to another language and back to produce a reworded version. Time-series data augmentation uses jittering, which adds small random noise, and window slicing, which extracts overlapping sub-sequences from a longer sequence.
How do you evaluate a deep learning classification model?
A deep learning classification model is evaluated using accuracy, precision, recall, F1-score, and the ROC-AUC curve, computed from a confusion matrix of true positives, true negatives, false positives, and false negatives.
| Metric | Formula | When It Matters Most |
|---|---|---|
| Accuracy | (TP + TN) / (TP + TN + FP + FN) | Balanced datasets only |
| Precision | TP / (TP + FP) | False positives are costly, e.g., spam detection |
| Recall | TP / (TP + FN) | False negatives are costly, e.g., disease detection |
| F1-score | 2 × (Precision × Recall) / (Precision + Recall) | Imbalanced datasets requiring a single balanced metric |
| ROC-AUC | Area under the true-positive-rate vs. false-positive-rate curve | Comparing models across all classification thresholds |
Accuracy is misleading on imbalanced datasets. A model that predicts “no fraud” for every transaction can score 99% accuracy if only 1% of transactions are fraudulent, while catching zero actual fraud cases. Precision, recall, and F1-score expose this failure; accuracy does not.
Tier 2: Mid-Level — Computer Vision and CNNs
What is a CNN, and why is it used for images instead of a standard neural network?
A Convolutional Neural Network (CNN) is a deep learning architecture that uses convolutional layers to automatically detect spatial features such as edges, textures, and shapes, without manual feature engineering. A standard fully connected network applied directly to an image would require one weight per pixel per neuron, which becomes computationally impractical and ignores the spatial relationship between neighboring pixels.
Figure 5: A CNN processes an image through alternating convolution and pooling layers before a fully connected layer produces the final classification.
A CNN has three core layer types: convolutional layers extract features, pooling layers reduce the spatial size of the feature maps, and fully connected layers perform the final classification using the extracted features.
What does the convolution operation actually compute?
Convolution slides a small matrix called a kernel across an input image, multiplying overlapping values and summing the result to produce a feature map that highlights specific patterns.
Figure 6: A 3×3 kernel sliding over a 5×5 input produces a 3×3 feature map. Each output value is the sum of element-wise multiplication between the kernel and the region it currently covers.
Different kernels detect different patterns. A kernel with a vertical strip of high values detects vertical edges; a kernel with a horizontal strip detects horizontal edges. During training, a CNN learns the kernel values automatically rather than using hand-designed filters.
Stride controls how many pixels the kernel moves between each step. A stride of 1 moves the kernel one pixel at a time, producing a larger output. A stride of 2 skips every other position, producing a smaller output and reducing computation.
Padding adds extra rows and columns, usually zeros, around the input before convolution.
| Padding Type | Effect | When Used |
|---|---|---|
| Valid (no padding) | Output is smaller than input | When shrinking the feature map is acceptable |
| Same (zero padding) | Output size equals input size | When spatial dimensions must be preserved across layers |
| Full padding | Output is larger than input | Rarely used in standard CNN architectures |
What is pooling, and what are the differences between max, average, and global pooling?
Pooling reduces the spatial size of a feature map by aggregating values within a small window, which lowers the parameter count and reduces overfitting.
| Pooling Type | Operation | Effect |
|---|---|---|
| Max pooling | Takes the highest value in each window | Preserves the strongest activation; most commonly used |
| Average pooling | Takes the mean value in each window | Preserves overall signal strength; smoother output |
| Global pooling | Reduces an entire feature map to a single value | Commonly used immediately before the fully connected layer |
Max pooling is the default in most CNN architectures because it retains the most prominent feature in each region, which matters more for classification tasks than the average signal strength.
What is the receptive field of a CNN, and why does it matter?
The receptive field is the region of the original input image that influences a single neuron’s output, and it grows larger with each additional convolutional or pooling layer. In the first convolutional layer, the receptive field equals the kernel size — a 3×3 kernel sees a 3×3 patch of the input. Stacking additional layers increases the receptive field, because each layer aggregates information from a wider region of the previous layer’s output.
A network with a small receptive field captures fine local details such as edges and textures. A network with a large receptive field captures global context needed to recognize whole objects. Architecture depth is chosen based on which type of pattern the task requires.
Object detection vs. image segmentation — what is the difference?
Object detection identifies the presence and location of objects using bounding boxes; image segmentation classifies every individual pixel in the image.
| Attribute | Object Detection | Image Segmentation |
|---|---|---|
| Output | Bounding box coordinates and class label | Pixel-level classification map |
| Precision | Locates approximate object position | Locates exact object boundary and shape |
| Compute cost | Lower | Higher |
| Common models | YOLO, Faster R-CNN, SSD | U-Net, Mask R-CNN, DeepLab |
Segmentation is further split into semantic segmentation, which classifies each pixel by category without distinguishing individual instances, and instance segmentation, which distinguishes separate objects of the same class.
Build a simple CNN image classifier — what does the code look like?
A basic CNN image classifier in PyTorch defines convolutional layers, pooling layers, and a fully connected output layer, then trains using a loss function and an optimizer.
import torch.nn as nn
class SimpleCNN(nn.Module):
def __init__(self, num_classes=10):
super().__init__()
self.conv1 = nn.Conv2d(3, 16, kernel_size=3, padding=1)
self.conv2 = nn.Conv2d(16, 32, kernel_size=3, padding=1)
self.pool = nn.MaxPool2d(2, 2)
self.relu = nn.ReLU()
self.fc = nn.Linear(32 * 8 * 8, num_classes)
def forward(self, x):
x = self.pool(self.relu(self.conv1(x))) # 32x32 -> 16x16
x = self.pool(self.relu(self.conv2(x))) # 16x16 -> 8x8
x = x.view(x.size(0), -1) # flatten
return self.fc(x)
This network applies two convolution-and-pooling blocks, then flattens the resulting feature maps into a fully connected layer that outputs one score per class. A softmax activation, applied either in the loss function or as a final layer, converts these scores into class probabilities.
Tier 2: Mid-Level — Sequence Models
What is an RNN, and what problem does it solve that a standard network cannot?
A Recurrent Neural Network (RNN) processes sequential data by maintaining a hidden state that carries information from previous time steps, which allows it to model dependencies across a sequence. A standard feedforward network has no memory of previous inputs, which makes it unsuitable for tasks such as language modeling or time-series forecasting, where the order of inputs carries meaning.
At each time step, an RNN updates its hidden state using the current input and the previous hidden state: hₜ = f(U·hₜ₋₁ + W·xₜ + b). The same weight matrices U and W are reused at every time step, which allows the network to handle sequences of variable length.
What is LSTM, and how do its gates solve the vanishing gradient problem?
Long Short-Term Memory (LSTM) is a type of RNN that uses three gates to control what information is kept, discarded, or output at each time step, which allows it to retain information across long sequences without the gradient vanishing. Hochreiter and Schmidhuber introduced LSTM in a 1997 paper published in Neural Computation.
Figure 7: The forget, input, and output gates control what enters and leaves the cell state at each time step.
- Forget gate — a sigmoid function decides what proportion of the previous cell state to discard.
- Input gate — a sigmoid and tanh combination decides what new information to add to the cell state.
- Output gate — a sigmoid function decides what portion of the current cell state to output as the hidden state.
The cell state provides a direct path for information to flow across time steps with minimal transformation, which prevents the repeated multiplication that causes vanishing gradients in standard RNNs.
What is GRU, and how does it differ from LSTM?
A Gated Recurrent Unit (GRU) is a simplified version of LSTM that uses two gates instead of three and has no separate cell state, which makes it faster to train while achieving comparable performance on most sequence tasks.
| Attribute | RNN | LSTM | GRU |
|---|---|---|---|
| Gates | None | 3 (forget, input, output) | 2 (reset, update) |
| Separate cell state | No | Yes | No |
| Parameter count | Lowest | Highest | Between RNN and LSTM |
| Handles long sequences | Poorly | Well | Well |
| Training speed | Fastest | Slowest | Faster than LSTM |
| Best used when | Short sequences | Long sequences, larger datasets | Long sequences, smaller datasets or limited compute |
GRU’s reset gate controls how much of the previous hidden state to forget when computing the new candidate state; its update gate controls the balance between the previous hidden state and the new candidate state. With fewer parameters than LSTM, GRU trains faster and is less prone to overfitting on smaller datasets.
Tier 3: Senior-Level — Transformers and NLP
What is the attention mechanism, and why did it replace RNNs for most NLP tasks?
The attention mechanism allows a model to assign a relevance score to every other element in a sequence when processing one element, which captures long-range dependencies without processing the sequence one step at a time. RNNs process sequences sequentially, which prevents parallelization and makes long-range dependencies harder to learn because information must pass through every intermediate time step.
Figure 8: Self-attention lets the word “sat” assign a different weight to every other word in the sentence when building its representation.
Attention computes a score for each pair of positions in a sequence, normalizes those scores with softmax, and uses them as weights to compute a context vector — a weighted sum of all positions relevant to the current one. Because every position can attend to every other position directly, attention removes the sequential bottleneck that limits RNNs.
What is a Transformer, and what are its core components?
A Transformer is a neural network architecture that relies entirely on attention mechanisms to process sequences in parallel, without recurrence. Vaswani et al. introduced the Transformer in a 2017 paper titled “Attention Is All You Need.”
Five components define the architecture:
- Self-attention — each token attends to every other token in the same sequence to build a context-aware representation.
- Multi-head attention — multiple attention computations run in parallel, each learning to focus on different types of relationships between tokens.
- Positional encoding — since Transformers process all tokens simultaneously, sine and cosine functions (or learned embeddings) inject information about token order.
- Feed-forward layers — applied independently to each position after attention, adding non-linear transformation capacity.
- Layer normalization and residual connections — stabilize training and allow gradients to flow through very deep stacks of Transformer blocks.
BERT vs. GPT — what is the actual architectural difference?
BERT uses only the Transformer encoder and processes text bidirectionally; GPT uses only the Transformer decoder and processes text left to right, predicting the next token at each step.
| Attribute | BERT | GPT |
|---|---|---|
| Architecture component | Encoder only | Decoder only |
| Context direction | Bidirectional (reads full sentence) | Unidirectional (left to right) |
| Pre-training objective | Masked language modeling, next sentence prediction | Causal language modeling (next-token prediction) |
| Primary strength | Understanding and classifying existing text | Generating new, coherent text |
| Typical applications | Sentiment analysis, named entity recognition, question answering | Text generation, chatbots, summarization, code generation |
Devlin et al. introduced BERT in 2018. Radford et al. introduced the original GPT architecture the same year. The bidirectional design gives BERT an advantage on tasks that require understanding a complete sentence at once; the autoregressive design gives GPT an advantage on tasks that require producing new text token by token.
How do classic Transformer concepts apply to working with large language models?
Attention, positional encoding, and fine-tuning are the same underlying mechanisms in a large language model as in a standard Transformer; a large language model is a decoder-only Transformer scaled to billions of parameters and trained on a large text corpus. Interviewers in 2026 increasingly ask candidates to connect classic Transformer theory to applied large language model work.
Fine-tuning a large language model uses the same principle as fine-tuning any pre-trained network: unfreeze some or all layers and continue training on a task-specific dataset with a lower learning rate. Retrieval-Augmented Generation (RAG) adds a retrieval step before generation — relevant documents are fetched from an external source and passed into the model’s context window alongside the query — which reduces hallucination without requiring retraining.
Tier 3: Senior-Level — Generative Models
What is a GAN, and how do the generator and discriminator interact?
A Generative Adversarial Network (GAN) consists of two neural networks — a generator that creates synthetic data and a discriminator that classifies data as real or generated — trained simultaneously in competition with each other. Goodfellow et al. introduced GANs in 2014.
Figure 9: The generator produces synthetic samples from random noise; the discriminator scores them against real data, and its feedback trains the generator to improve.
The generator takes a random noise vector as input and outputs synthetic data. The discriminator receives both real and generated samples and predicts which is real. Training alternates between updating the discriminator to better distinguish real from fake, and updating the generator to produce output the discriminator misclassifies as real. Training stops when the generator produces output the discriminator cannot reliably distinguish from real data.
What is mode collapse, and why is GAN training unstable?
Mode collapse occurs when the generator learns to produce only a limited variety of outputs that reliably fool the discriminator, instead of representing the full diversity of the real data distribution. GAN training is inherently unstable for three reasons: the generator and discriminator optimize against each other rather than toward a shared objective, a discriminator that improves too quickly gives the generator a weak gradient signal, and there is no mathematical guarantee that the two networks converge to a stable equilibrium.
Wasserstein GAN, introduced as a variant that replaces the standard GAN loss with the Wasserstein distance, is a common fix for mode collapse because it produces a smoother, more informative gradient signal throughout training.
What is an autoencoder, and what are its main variants?
An autoencoder is an unsupervised neural network that learns to compress input data into a lower-dimensional representation and then reconstruct the original input from that representation. It consists of an encoder, which compresses the input, and a decoder, which reconstructs it.
| Variant | Modification | Primary Use |
|---|---|---|
| Vanilla autoencoder | Standard encoder-decoder | Dimensionality reduction |
| Denoising autoencoder | Trained to reconstruct clean input from corrupted input | Noise removal |
| Sparse autoencoder | Adds a sparsity constraint on the latent representation | Feature extraction |
| Variational autoencoder (VAE) | Models the latent space as a probability distribution | Generating new, realistic samples |
| Convolutional autoencoder | Uses convolutional layers instead of fully connected layers | Image compression and denoising |
A variational autoencoder differs from a standard autoencoder by mapping inputs to a probability distribution rather than a single fixed point in latent space, which allows new samples to be generated by drawing from that distribution. Kingma and Welling introduced the VAE in 2013.
Tier 3: Senior-Level — Transfer Learning and Production
What is transfer learning, and how is it different from fine-tuning?
Transfer learning reuses a pre-trained model’s learned features for a new task by training only a new output layer, while fine-tuning continues training some or all of the pre-trained model’s layers on the new task.
| Attribute | Transfer Learning | Fine-Tuning |
|---|---|---|
| Layers updated | Only the new output layer | Some or all pre-trained layers |
| Training time | Faster | Slower |
| Data required | Works with a small target dataset | Performs better with a larger target dataset |
| Risk | Limited adaptation to the new task | Risk of overfitting on a small target dataset |
A model pre-trained on ImageNet and reused to classify medical images by training only a new final layer is transfer learning. The same model with its last several convolutional blocks unfrozen and retrained at a low learning rate is fine-tuning.
What should you check before deploying a deep learning model to production?
A deep learning model ready for production deployment should be validated for inference latency, memory footprint, and performance drift against real-world data, in addition to standard accuracy metrics. Three specific checks matter most.
- Inference speed — a model that performs well in offline evaluation can be too slow for real-time use. Converting a trained model to ONNX or optimizing it with TensorRT typically reduces inference latency without retraining.
- Model size — mobile and edge deployment often requires quantization or pruning to fit hardware memory constraints.
- Data drift — a model’s accuracy on production data degrades over time as the real-world data distribution shifts away from the training distribution, which requires ongoing monitoring after deployment.
Frequently Asked Questions
Q1. How long does it take to prepare for a deep learning interview?
Ans. Candidates with existing machine learning experience typically need two to four weeks of focused review covering neural network fundamentals, one architecture family in depth (CNN or Transformer, depending on the target role), and hands-on coding practice in PyTorch or TensorFlow.
Q2. Do you need to derive backpropagation math by hand for every interview?
Ans. Most interviews expect a conceptual explanation of backpropagation and the chain rule; a smaller number of research-focused or senior roles require deriving gradients for a specific layer type by hand.
Q3. Is PyTorch or TensorFlow more important to learn first?
Ans. PyTorch is more commonly requested in research and startup job postings as of 2026, while TensorFlow remains common in large enterprise production pipelines; learning one deeply transfers most of the underlying concepts to the other.
Q4. What is asked differently at a startup versus a large technology company?
Ans. Startups more often ask about end-to-end system design and tradeoffs under limited compute, while large technology companies more often ask about optimizing a specific component, such as a custom loss function or a distributed training setup, in depth.