Tokenization

Tokenization splits text into the integer units (tokens) a language model reads and writes. Modern LLMs use subword schemes, most commonly byte-pair encoding (BPE), adapted to NLP by Sennrich et al. (2015) to handle rare words, with variants such as WordPiece and SentencePiece; byte-level BPE guarantees any string can be encoded. A token is often a word fragment: roughly a few characters of English, but more for non-English text, code and numbers.

Why it matters

  • Cost and limits: APIs bill and cap context in tokens, so the same text can cost more in some languages. See context-rot-and-long-context and prompt-caching.
  • Odd failures: counting letters, arithmetic and exact string manipulation are harder because the model sees tokens, not characters.
  • Model-specific: each family has its own vocabulary; embeddings, prompts and token counts are not portable between them (vector-embeddings).
  • Multimodal models add image/audio tokens (multimodal-models).

Related: large-language-model, generative-pretrained-transformer, inference-in-generative-ai.

Sources