Tokenization
Tokenization splits text into the integer units (tokens) a language model reads and writes. Modern LLMs use subword schemes, most commonly byte-pair encoding (BPE), adapted to NLP by Sennrich et al. (2015) to handle rare words, with variants such as WordPiece and SentencePiece; byte-level BPE guarantees any string can be encoded. A token is often a word fragment: roughly a few characters of English, but more for non-English text, code and numbers.
Why it matters
- Cost and limits: APIs bill and cap context in tokens, so the same text can cost more in some languages. See context-rot-and-long-context and prompt-caching.
- Odd failures: counting letters, arithmetic and exact string manipulation are harder because the model sees tokens, not characters.
- Model-specific: each family has its own vocabulary; embeddings, prompts and token counts are not portable between them (vector-embeddings).
- Multimodal models add image/audio tokens (multimodal-models).
Related: large-language-model, generative-pretrained-transformer, inference-in-generative-ai.
Sources
- https://arxiv.org/abs/1508.07909 (Neural Machine Translation of Rare Words with Subword Units, 2015)
- https://huggingface.co/docs/transformers/tokenizer_summary (Hugging Face docs, reachable 2026-10-07)
- arXiv API 2026-10-07: Sennrich et al., submitted 2015-08-31.