Zero-shot Learning

|
11 min read
|
8 views
Zero-Shot Learning in Deep Learning

Zero shot learning is a technique used in machine learning whereby the model can produce outputs in an unknown category in a way that the model has not been trained on. This is accomplished through the mapping of both known and unknown classes to a common meaning space.

What Will I Learn?

What Is Zero-Shot Learning?

Zero-shot learning (ZSL) is a machine learning approach by which a model learns to recognize, classify, or produce some kind of output for categories that it did not see during its training phase through use of semantic information.

ZSL uses semantic descriptions in place of labeled training data. The semantics can come from text descriptions, attribute sets defined by humans, or word embeddings that relate a novel class category to known categories. It was described in 2009 by Mark Palatucci, Dean Pomerleau, Geoffrey Hinton, and Tom Mitchell in their research on the prediction of semantic output codes as opposed to class labels. The basic principle behind ZSL has not changed since then – a model uses what it already knows about related concepts to make generalizations about a concept it has never seen before.

ZSL is not the same as zero-shot prompting. Zero-shot learning is a training approach while zero-shot prompting is a way to apply an existing large language model by supplying an instruction without any example outputs.

Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Zero-Shot Learning vs. One-Shot, Few-Shot, and Fine-Tuning

Zero-shot learning is based on zero annotated samples of a new category. One-shot learning employs one sample. Few-shot learning employs two to five samples. Fine-tuning fine-tunes the model with hundreds or thousands of annotated samples.

The following table compares all four approaches from the perspectives of what differentiates them from each other.

MethodLabeled examples requiredTypical accuracyBest use case
Zero-shot learning0Lowest of the four; varies by domain gapA new category with no available data
One-shot learning1Low to moderateA new class defined by a single reference sample
Few-shot learning2–5Moderate to highFast adaptation using a small labeled set
Fine-tuningHundreds to thousandsHighestMaximum accuracy on a fixed, well-defined task

Both zero-shot learning and few-shot learning are not done through retraining, but few-shot learning is more accurate where there are at least a few labeled instances available. If there are lots of labeled instances, fine-tuning is always the way to go.

Is Zero-Shot Learning Supervised or Unsupervised?

Zero-shot learning is not a purely supervised or purely unsupervised technique. Supervised learning technique is used for the seen classes which have labels. Then knowledge gained by using this technique is transferred to unseen classes which have no label at all. Such a scenario is considered to be a different problem of transfer learning.

How Does Zero-Shot Learning Work?

Zero shot learning is done through transforming the known and unknown classes into a similar semantic representation and finding the similarity between the new input and the representation of each class.

Semantic Representations and Embeddings

The semantic representation refers to the numeric or textual representation of the class semantics in terms of the attributes or words. There are different types of semantic representations, namely:

  • Attribute-based, which includes attribute values like “has stripes” or “has four legs.”
  • Word embeddings, which include numeric vectors generated by Word2Vec, GloVe, and BERT.
  • Natural language descriptions, which contain sentences describing a class.

The Shared Embedding Space (Where Seen Meets Unseen)

The shared embedding space is a vector space common to both seen and unseen classes and a distance function measures their semantic relatedness in that space. A model maps its input (an image or a sentence, for instance) to this space through an encoder. It maps each class description to the same space via another encoder which is similar to the first one. The model selects the class that has the closest vector to the input vector.

Transfer Learning as the Foundation

Transfer learning provides the generic knowledge that zero-shot learning relies upon, by recycling an existing model that has been trained on a huge amount of data instead of training a new one. A zero-shot text classifier often leverages an existing pre-trained transformer model like BERT on a huge text corpus for generating embeddings. Similarly, a zero-shot image classifier often exploits an existing pre-trained convolutional neural network like ResNet or a vision transformer that has been trained on millions of images.

Worked Example: Classifying an Animal the Model Has Never Seen

An image classifier that has been zero-shot learning for horses, tigers, and other animals can recognize a zebra without any prior knowledge about the zebra because it has not encountered before by just providing a description of “a horse-like animal with black and white stripes found in African grasslands.” This whole process of classification takes four stages:

  1. The classifier converts the input image into an embedded space.
  2. It converts each candidate class description into the same embedded space.
  3. It computes the similarity score of the input image and all candidate classes.
  4. It selects the class with the highest similarity score.

Training Methods for Zero-Shot Learning

Attribute-Based Methods

The zero-shot learning models can be created based on three types of methods: attribute-based method, embedding-based method, or generative-based method. All these approaches link the known and unknown classes in a different way.

Attributive approaches involve training a classifier to recognize manually created attributes like color and shape and then assigning a new class according to the attributes matching its description. However, there is a well-known weakness in the approach. According to Mall, Hariharan, and Bala, some classes cannot be described by a single attribute set. The color and plumage of the American Goldfinch are different depending on the bird’s gender, age, and breeding season.

Embedding-Based Methods

In embedding-based approaches, both classes and instances are converted into vectors and are classified on the basis of the class vectors which are closest to the instances in terms of a distance metric akin to the k-nearest neighbor algorithm. There are two such approaches. The first approach, direct attribute prediction, predicts the attributes of the class and compares them with the closest class. In the second approach, compatibility-based scoring, a mapping function is learned between class and input vectors. CLIP from OpenAI uses this approach for matching images with their textual descriptions.

Generative-Based Methods: VAEs, GANs, and VAEGANs

Generative techniques synthesize training data based on the semantic definitions of new categories and train a regular supervised classifier on these synthetic samples. There are three prevalent architectures for this technique:

  • Variational autoencoders (VAEs) generate samples from a learned probability distribution in latent space. They train stably but sometimes produce blurry output.
  • Generative adversarial networks (GANs) pair a generator against a discriminator in competition. They produce sharper output but train less stably than VAEs.
  • VAEGANs combine both architectures to offset each one’s individual weakness.
Generative-Based Methods

How CLIP and Modern Vision-Language Models Do Zero-Shot Classification

CLIP conducts zero-shot image classification through joint learning of an image encoder and a text encoder using 400 million image-caption pairs obtained from the web, followed by matching the image with the text description having the highest similarity. According to OpenAI, CLIP outperformed in classifying images in 27 different image classification tasks without any further fine-tuning (Radford et al., OpenAI, 2021).

Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Zero-Shot Learning in Large Language Models — the 2026 Reality

In 2026, the most common form of zero-shot learning encountered by people will be a large language model correctly performing a task based on instructions provided to it without any training samples.

Zero-Shot Prompting vs. Zero-Shot Learning: What’s the Difference

Zero-shot learning is about the training process that enables generalization to novel classes by the model. Zero-shot prompting involves using an already trained large language model through which the user provides the instruction without any completion examples.

Where You’re Already Using Zero-Shot Learning Without Knowing It

Zero-shot learning or prompting techniques have found applications in three consumer facing technologies:

  1. Chatbot based on a large language model that is answering a question on a topic it has not been fine-tuned for.
  2. Application that recognizes categories of objects which it has not been taught before to recognize.
  3. Customer support ticket routing to an undefined category with only a textual definition.

The Hugging Face zero-shot-classification pipeline, built on natural language inference models such as facebook/bart-large-mnli, is the standard open-source tool practitioners use to run this task without custom model training (Yin, Hay, and Roth, EMNLP 2019).

Zero-Shot Learning

How Accurate Is Zero-Shot Learning?

Models for zero-shot learning tend to have relatively lower accuracy when compared with models learned from labeled instances of the target classes. This difference depends on both the field and the description quality.

How to Evaluate a Zero-Shot Model

Evaluation of Zero-Shot Models is performed through three types of measurements that are not commonly used in supervised learning: Top-k accuracy, Harmonic Mean, and Generalized Zero-Shot Learning Accuracy.

Top-k Accuracy

Top-k accuracy is defined by the number of times the actual class is part of the k highest ranked outputs from a model, not just the top ranked output being correct. A top-5 accuracy of 80 percent implies that the correct class was among the highest-ranked 5 outputs in 80 percent of cases.

Harmonic Mean (H-Score)

The harmonic mean measures both the accuracy of the model for the seen classes as well as that for the unseen classes in a single number, and punishes a model which does well in either of these categories. The formula for this is: H = 2 × (Aseen × Aunseen) ÷ (Aseen + Aunseen).

Generalized Zero-Shot Learning (GZSL) Accuracy

Generalized zero-shot learning accuracy measures a model’s performance when a test set includes both seen and unseen classes together, a more realistic and more difficult evaluation than testing on unseen classes alone. Xian, Lampert, Schiele, and Akata established this benchmark methodology in a 2018 IEEE Transactions on Pattern Analysis and Machine Intelligence survey.

Real-World Applications of Zero-Shot Learning

Zero-shot learning is used today in text classification, image recognition, product categorization, and medical diagnosis support, in each case replacing the need for labeled examples of every possible category.

Natural Language Processing and Text Classification

Zero-shot learning supports spam detection using category definitions instead of labeled spam examples, sentiment analysis using only the definitions of “positive” and “negative,” and content moderation that identifies harmful text from a written description of the violation category.

Computer Vision and Image Recognition

Zero-shot learning supports object recognition for categories with no training images, satellite image monitoring — for example, detecting deforestation described as “significant canopy loss” — and visual search engines that match a text query to an image with no matching label in their index.

Retail, Recommendations, and the Cold-Start Problem

Retailers use zero-shot learning to classify new inventory into categories directly from text descriptions, and to recommend products to a new user or from a new catalog with no prior interaction history, a scenario known as the cold-start problem.

Healthcare and Rare-Disease Diagnosis

Zero-shot learning offers diagnostic support for rare diseases, where labeled patient data is scarce or unavailable, by linking a new case’s symptom description to related, better-documented conditions. This use case remains largely research-stage rather than standard clinical practice.

Challenges and Limitations of Zero-Shot Learning

Zero-shot learning faces four primary limitations: domain gap, knowledge representation quality, scalability, and interpretability. Each limitation reduces model reliability in a different way.

Domain Gap and Bias Toward Seen Classes

A zero-shot model favors classes it was trained on and performs worse on classes that differ substantially from its training data, a pattern known as bias toward seen classes.

Knowledge Representation Quality

Vague or overlapping attribute definitions cause classification errors. A model may confuse a leopard and a cheetah, since both match the description “spotted big cat,” and that description fails to capture their finer physical differences.

Scalability

As the number of unseen categories grows into the thousands, the shared embedding space becomes harder to search accurately and efficiently, and prediction speed and accuracy both decline.

Interpretability

Zero-shot decisions rely on similarity scores in a high-dimensional embedding space, which is difficult to explain in plain terms. This limitation carries particular weight in high-stakes settings such as medical diagnosis, where a clear explanation for a prediction is often required.

What Happens When Zero-Shot Classification Fails

When a zero-shot model fails, it does not return an error. It returns a confident but incorrect class label, since the model’s architecture always outputs the closest available match, whether or not that match is correct. Production systems mitigate this risk by setting a similarity-score threshold below which a prediction is flagged for human review, rather than accepted automatically.

Should You Use Zero-Shot Learning? A Decision Framework

Use zero-shot learning when no labeled examples exist for a target category and some accuracy loss is acceptable. Use few-shot learning or fine-tuning instead when a labeled dataset is available and higher accuracy is required. Apply the following steps to decide:

  1. Count the labeled examples available for the target class. Zero examples points toward zero-shot learning.
  2. Set the accuracy threshold the task requires. A lower threshold supports zero-shot learning; a higher threshold supports few-shot learning or fine-tuning.
  3. Estimate the cost and time required to label new data. High labeling cost favors zero-shot learning, even at some accuracy cost.
  4. Check the domain gap between the target class and the classes in the pretrained model’s original training data. A large gap reduces zero-shot accuracy and may require few-shot examples instead.
Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Frequently Asked Questions About Zero-Shot Learning

Q1. Is zero-shot learning the same as zero-shot prompting?

Ans. No. Zero-shot learning is a training paradigm that teaches a model to generalize to unseen classes using semantic information. Zero-shot prompting is a technique for querying an already-trained large language model with an instruction and no example answers. The two terms describe different stages of a model’s lifecycle.

Q2. Is zero-shot learning supervised or unsupervised?

Ans. Neither. Zero-shot learning applies supervised learning to labeled seen classes during training, then transfers that knowledge to unseen classes with zero additional labels — a setup classified separately as a transfer learning problem.

Q3. How accurate is zero-shot learning compared to supervised learning?

Ans. Zero-shot learning typically scores lower than supervised learning on the same task, since supervised models train directly on labeled examples of the target class. The exact gap depends on the domain and the quality of the semantic descriptions used.

Q4. What is the difference between one-shot and zero-shot learning?

Ans. One-shot learning trains a model using exactly one labeled example of a new class. Zero-shot learning trains a model using zero labeled examples, relying entirely on semantic descriptions or attributes instead.

Q5. Do I need to write code to use zero-shot learning?

Ans. Not always. Tools such as the Hugging Face zero-shot-classification pipeline, and large language models accessed through a chat interface, let a user run zero-shot classification with a plain-language description or prompt instead of custom code.

Q6. What is generalized zero-shot learning (GZSL)?

Ans. Generalized zero-shot learning is a test setup where a model must classify inputs that could belong to either seen or unseen classes, without knowing in advance which group a given input belongs to. It is a harder and more realistic test than standard zero-shot learning.

Key Takeaways

Zero-shot learning accuracy continues to depend on the size and quality of the pretrained foundation model behind it. Teams evaluating a zero-shot approach should apply the decision framework above before committing engineering time to a labeled-data collection effort that a zero-shot model might make unnecessary.

Gyansetu offers top professional training certification courses designed to enhance your skills and advance your career, providing industry-relevant knowledge and practical expertise.