In-context learning (ICL) is a large language model’s ability to pick up a task temporarily from the prompt itself, typically from a few input-output examples (few-shot prompting), with no change to the model’s weights. The effect lasts only for that conversation; nothing is permanently learned, unlike low-rank-adaptation or other fine-tuning. Wikipedia calls it an emergent ability of large models whose effectiveness grows with scale, and notes that training for it can be seen as a form of meta-learning (“learning to learn”). See large-language-model and generative-pretrained-transformer.

Variants

  • Zero-shot: instruction only. One-shot / few-shot: one to a handful of worked examples in the prompt.
  • Many-shot: hundreds or thousands of examples, made practical by long context windows. Agarwal et al. (Google, NeurIPS 2024, arXiv 2404.11018) report significant gains across tasks, that many-shot learning can override pre-training biases and sometimes matches fine-tuning, and that inference cost grows linearly with the number of examples. They also propose Reinforced ICL (model-generated reasoning chains) and Unsupervised ICL (domain questions only).

Key papers

  • Brown et al., “Language Models are Few-Shot Learners” (OpenAI, 2020, arXiv 2005.14165): GPT-3 (175B parameters) showed that scaling greatly improves task-agnostic few-shot performance without fine-tuning.
  • Garg et al., “What Can Transformers Learn In-Context? A Case Study of Simple Function Classes” (2022, arXiv 2208.01066): transformers trained from scratch can learn unseen linear functions from in-context examples with performance comparable to the optimal least-squares estimator, and also sparse linear functions, two-layer networks and decision trees.
  • Olsson et al., “In-context Learning and Induction Heads” (Anthropic, 2022, arXiv 2209.11895): attention heads that complete patterns like [A][B] … [A] → [B] emerge at the moment in-context learning improves sharply; causal evidence in small attention-only models, correlational evidence in larger ones.

Limits and practice

  • Results are sensitive to the order of examples, the quality of demonstration labels and small wording changes (Wikipedia, “Prompt engineering”).
  • The examples consume context-window tokens on every request; for stable behaviour at scale, teams move examples into fine-tuning, retrieval (retrieval-augmented-generation) or reusable instructions. See prompt-injection-and-agent-security for the risk of untrusted text in context.

The earlier version of this note paired ICL with Meta’s “Large Context Models” (LCM); that link was never verified and has been dropped.

Sources (accessed 2026-10-02)