Few-shot learning (FSL) is a machine learning approach in which a model learns to perform a task, usually classification, from only a small number of labeled examples per class. Instead of the hundreds or thousands of labeled samples conventional supervised learning needs for each category, a few-shot model works from a handful, often between three and 10. That difference matters most where labeled data is rare, expensive, or slow to collect: for example, a hospital with only a few confirmed images of a rare condition, or a factory line that sees a specific defect once a month.
The idea moved from research into everyday practice at the start of the decade. The 2020 paper that introduced GPT-3, titled "Language Models Are Few-Shot Learners," showed that a large pretrained model with 175 billion parameters could take on new language tasks from just a few examples placed in the prompt, with no extra training. That result reframed how many teams think about adapting AI to a new problem.
This article explains how few-shot learning works, the main technical approaches, where it fits against one-shot and zero-shot learning, its benefits and limits, and the data infrastructure that lets a promising few-shot model scale into production.
How few-shot learning works
Few-shot learning starts from a different premise than standard supervised learning. Rather than teaching a model one task with a large data set, it trains a model to become good at learning new tasks quickly. The mechanics come down to three ideas: how tasks are framed, what data the model sees, and how the model is trained to generalize.
The N-way, K-shot framework
Most few-shot tasks are described as N-way, K-shot problems. N is the number of classes the model has to distinguish, and K is the number of labeled examples available for each class. A 5-way, 3-shot task asks the model to distinguish five categories using three examples of each. The framework gives researchers and practitioners a consistent way to state how hard a given few-shot problem is: A smaller K and a larger N both make the task harder.
Support sets and query sets
Few-shot learning splits its data into two groups that mirror how the model gets used in practice. The support set holds the few labeled examples the model learns from, typically the K examples for each of the N classes. The query set holds new, unlabeled examples that the model has to classify based on what the support set taught it.
This split maps cleanly to real-world tasks. A manufacturer, for example, might have three verified images of a specific solder defect (the support set) and need to flag that defect across thousands of boards coming off the line (the query set).
Meta-learning: Learning to learn
The engine behind most few-shot methods is meta-learning, often described as "learning to learn." Instead of memorizing one data set, the model trains across many small, varied tasks so it can generalize to a new task drawn from a similar distribution. Over many of these episodes, it picks up transferable structure, such as what makes two inputs similar, rather than facts about any single set of classes.
That transferable structure is what replaces raw data volume. Conventional supervised learning leans on scale, since a large labeled data set is what lets a model generalize. Few-shot learning trades some of that scale for prior knowledge built up across tasks, which is why it can work when only a few labeled examples exist.
Types of few-shot learning
Few-shot methods generally fall into a few families. They are not mutually exclusive, and production systems often combine them.
Metric-based methods
Metric-based approaches learn a similarity function that compares a new input to the labeled examples in the support set, then assigns the class of whatever it most resembles. The lineage traces back through several landmark methods: Siamese networks learn to score whether two inputs belong to the same class, matching networks act as a differentiable nearest-neighbor classifier, and prototypical networks compute a single representative vector, or prototype, for each class and classify by distance to those prototypes. These methods suit cases where the goal is to recognize new classes without retraining the whole model.
Optimization-based meta-learning
Optimization-based methods tune how a model learns so it can adapt in only a few gradient steps. The best-known example is model-agnostic meta-learning (MAML), introduced in 2017. MAML learns a parameter initialization that sits in a useful place for many tasks, so that a handful of updates on a new task's support set moves the model to strong performance. Because it makes no assumptions about the network, it works with any model trained by gradient descent.
Transfer learning and data-level methods
Transfer learning adapts a model already trained on a large data set to a new task with limited data, often by fine-tuning only the final layers or a small set of components. Data-level methods take a different angle: When the support set is too thin, they generate additional training examples through augmentation or generative models to add variety. Both approaches place real demands on the underlying infrastructure since they move and process far more data than the small support set alone suggests.
In-context few-shot learning in large language models
Large language models introduced a form of few-shot learning that needs no additional training. In in-context learning, a few worked examples, or "shots," go into the prompt alongside the instruction, and the model generates an answer using patterns it learned during pretraining. The GPT-3 result showed this at scale, and it’s now a standard way to steer generative AI systems. It differs from the methods above in one important way: The model's weights do not change. The adaptation happens entirely in the prompt.
Few-shot vs. one-shot vs. zero-shot learning
Few-shot learning sits on a spectrum defined by how many labeled examples the model gets. One-shot learning is the special case where K equals one, so the model has exactly one example per class to work from. Zero-shot learning is the limiting case that uses no labeled examples of the target class at all, relying instead on a description of the class or on knowledge transferred from related classes.