A glossary of AI terms, rewritten 2026-10-05: classical ML first, then modern LLM, agent, training, safety and infrastructure vocabulary, then abbreviations (A-Z). One term per line; a term links to a vault note where one exists. Definitions are neutral and deliberately free of model names, prices and benchmark scores, which go stale. They are working definitions summarised for orientation; the attributions and dates cited in them come from the papers and pages named in Sources, and fast-moving agent terms are not standardised.
1. Classical ML and AI terms
- Activation Function= A function applied to a neuron’s weighted input to produce its output, introducing the non-linearity that lets networks learn complex patterns (e.g. ReLU, sigmoid).
- Agent (classical)= An entity that perceives an environment and acts on it to achieve goals; in the LLM sense see AI Agent below.
- artificial-intelligence (AI)= The simulation of human intelligence processes by machines, especially computer systems. These processes include learning, reasoning, and self-correction.
- Algorithm= A set of rules to be followed in calculations or other problem-solving operations, especially by a computer.
- Algorithm Bias= Systematic and repeatable errors in a computer system that create unfair outcomes such as privileging one arbitrary group over others.
- Artificial Neural Network= A computing model of layered, connected units (neurons) whose weights are learned from data.
- Attention Mechanisms= The part of a Transformer that weights how relevant each token in the input is to every other token; introduced for translation and central to large-language-models.
- Autoencoder= A neural network trained to compress its input into a small latent representation and reconstruct it, used for dimensionality reduction and generative modelling.
- Autonomous Systems= Systems capable of performing tasks without human intervention.
- autonomous-vehicles= Vehicles that drive with little or no human input; see the note.
- Backpropagation= The algorithm that computes the gradient of a loss with respect to a network’s weights, layer by layer, so gradient descent can train it.
- Bias-Variance Tradeoff= The tension between a model too simple to fit the data (high bias) and one so flexible it fits noise (high variance).
- Big Data= Data sets too large or complex for traditional processing tools, commonly characterised by volume, velocity and variety.
- Binary Classification= A supervised task that assigns each input to one of two classes (e.g. spam or not spam).
- Chatbot= A program that converses with users in natural language; modern ones are built on large-language-models.
- Classification= A supervised task that assigns each input to one of a set of discrete classes.
- Clustering= An unsupervised task that groups similar data points without predefined labels.
- Cognitive Computing= Systems that are designed to simulate human thought processes in a computerized model.
- computer-vision= An interdisciplinary field that deals with how computers can gain high-level understanding from digital images or videos. It seeks to automate tasks that the human visual system can do.
- Convolutional Neural Network (CNN)= A type of deep neural network often used in image recognition and processing that is specifically designed to process pixel data.
- Cross-Validation= Evaluating a model by repeatedly training on part of the data and testing on the held-out remainder.
- Data Mining= The process of discovering patterns in large data sets using methods from statistics, machine learning and databases.
- Data Science= An interdisciplinary field that uses statistics, programming and domain knowledge to extract insight from data.
- Decision Tree= A model that predicts an outcome by following a tree of feature-based yes/no splits from root to leaf.
- Deep Learning= A subset of machine learning that uses neural networks with many layers (deep neural networks) to analyze various factors in large amounts of data.
- Dimensionality Reduction= Mapping data to fewer variables while keeping most of its structure (e.g. PCA).
- Ensemble Learning= Combining several models (e.g. bagging, boosting) to get better predictions than any single one.
- Epoch= One full pass of the training algorithm over the entire training data set.
- Evolutionary Algorithm= An optimisation method inspired by biological evolution (selection, mutation, recombination of candidate solutions).
- Expert System= An artificial intelligence program that has expert-level knowledge about a particular domain and can emulate expert decision-making abilities.
- Explainable AI (XAI)= An emerging field in machine learning and artificial intelligence aiming at making black box systems transparent, so users can understand how models make decisions.
- Feature Engineering= Selecting, transforming or creating input variables to make a machine-learning model work better.
- Feature Extraction= Deriving a reduced, informative set of features from raw data.
- Federated Learning= Training a shared model across many devices or servers that keep their data local and only share model updates.
- Fuzzy Logic= A form of many-valued logic which deals with reasoning that is approximate rather than fixed and exact.
- Generative Adversarial Networks (GANs)= A class of machine learning frameworks designed by pitting two neural networks against each other in order to generate new, synthetic instances of data that can pass for real data.
- Genetic Algorithm= Search heuristics that mimic the process of natural selection to generate useful solutions to optimization and search problems.
- Gradient Descent= An iterative optimisation method that adjusts parameters in the direction that most reduces the loss.
- Heuristic= A practical rule of thumb that finds a good-enough solution quickly without guaranteeing the optimum.
- Hyperparameter= A setting chosen before training (e.g. learning rate, number of layers) rather than learned from data.
- Image Recognition= The task of identifying what an image contains (an object, scene, face or text), typically by classifying it with a computer-vision model such as a CNN.
- machine-learning (ML)= A subset of AI that involves the creation of algorithms that can modify themselves without human intervention to produce desired outputs by feeding on data inputs.
- Loss Function= A measure of how far a model’s predictions are from the targets, which training tries to minimise.
- Natural Language Processing (NLP)= A field at the intersection of computer science, artificial intelligence, and computational linguistics which focuses on interactions between computers and human languages.
- Neural Network= A series of algorithms that attempt to recognize underlying relationships in a set of data through a process that mimics the way the human brain operates. See neural-networks and Artificial Neural Network.
- Object Detection= A computer-vision task that locates objects in an image and labels each with a class.
- Optimization Algorithms= Procedures that search for the parameter values that minimise or maximise an objective function; in machine learning, gradient descent and its variants (SGD, Adam) are the common ones.
- Overfitting= When a model fits its training data, including noise, so well that it generalises poorly to new data.
- Pattern Recognition= The automated detection of regularities in data and their use to classify or act on it; the older name for much of what is now called machine learning.
- Perceptron= The simplest neural network unit: a weighted sum of inputs passed through a threshold, used as a linear binary classifier.
- Precision and Recall= Classification metrics: precision is the share of predicted positives that are correct, recall the share of actual positives that were found.
- Predictive Analytics= Using historical data, statistics and machine learning to estimate the likelihood of future outcomes.
- Recommender System= A system that predicts which items (products, videos, articles) a user is likely to prefer, using collaborative filtering, content features or both.
- Recurrent Neural Network (RNN)= A type of neural network where connections between nodes form a directed graph along a temporal sequence, allowing it to exhibit temporal dynamic behavior. Used for tasks like language modeling and translation.
- Regression= A supervised task that predicts a continuous numeric value.
- Regularization= Techniques (e.g. weight decay, dropout) that discourage overfitting by constraining the model.
- Reinforcement Learning= An area of machine learning concerned with how software agents ought to take actions in an environment so as to maximize some notion of cumulative reward.
- Robotics= The branch of technology that deals with the design, construction, operation, and application of robots.
- Self-supervised Learning= Learning from labels derived from the data itself (e.g. predicting the next word or a masked patch); the basis of pre-training.
- Semantic Analysis= Interpreting the meaning of text (as opposed to its form), e.g. word senses and relations.
- Semi-supervised Learning= An approach to machine learning that combines a small amount of labeled data with a large amount of unlabeled data during training. It falls between supervised learning (fully labeled data) and unsupervised learning (no labels).
- Sentiment Analysis= An NLP task that classifies the opinion or emotion expressed in text (e.g. positive, negative, neutral).
- Sequential Model= A model that processes ordered data such as text, audio or time series, where earlier elements affect later ones (RNNs, Transformers); in Keras, also a plain stack of layers.
- Supervised Learning= A type of machine learning algorithm that uses a known dataset called the training dataset, which includes input data and response values (outputs) to train a model which makes predictions or decisions without human intervention.
- Support Vector Machine= A supervised model that separates classes with the maximum-margin boundary.
- TensorFlow= An open-source machine-learning framework from Google for building and training neural networks; see tensorflow-ecosystem-note and jax for the alternative.
- Transfer Learning= Reusing a model trained on one task as the starting point for another task.
- Unsupervised Learning= Machine learning algorithms used when the information used to train is neither classified nor labeled. It studies how systems can infer a function to describe hidden structure from unlabeled data.
- Virtual Assistant= Software that takes spoken or written requests and carries out tasks or answers questions for a user; older ones were rule-based, current ones are built on large-language-models. See also AI agent.
- Neurosymbolic AI= Approaches that combine neural networks with symbolic reasoning or knowledge representation.
- ontology= A formal, shared specification of the concepts in a domain and the relations between them; see resource-description-framework and owl.
2. Modern LLM terms
- Large Language Model= A neural network, almost always a Transformer, trained on very large text corpora to predict the next token, which can then follow instructions and generate text.
- Transformer= The neural architecture (Vaswani et al., 2017) built on self-attention rather than recurrence; the basis of modern LLMs. See transformer-models-2023.
- Generative Pretrained Transformer= A decoder-only Transformer pre-trained to predict the next token, then adapted for tasks.
- Tokenizer= The component that splits text into tokens and maps them to integer IDs, usually with a learned subword vocabulary.
- BPE (Byte-Pair Encoding)= A tokenizer algorithm that builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pairs.
- Token= The unit a language model reads and writes: text is split into tokens, each mapped to an integer index in the model’s vocabulary. Billing and limits are usually per token.
- Context Window= The span of tokens a model can attend to in one request (prompt plus output). See in-context-learning.
- Context Rot= The degradation in accuracy and instruction-following as the context fills with more, and often irrelevant, tokens, even within the advertised window.
- Context Compaction= Summarising or trimming earlier conversation so a long-running session can continue within the context window.
- context-engineering= Designing everything that goes into the model’s context (instructions, retrieved data, tools, memory, history) so it can do the task reliably; a broader successor to prompt writing.
- Prompt Engineering= Crafting the instructions and examples given to a model to get better outputs; see chain-of-thought-prompting and meta-prompting.
- System Prompt= Developer-set instructions placed ahead of the conversation that define a model’s role, rules and style.
- In-Context Learning= A model’s ability to pick up a task from examples or instructions in the prompt without any weight update.
- Few-shot / Zero-shot Prompting= Giving a model a few worked examples (few-shot) or none (zero-shot) in the prompt.
- Chain of Thought (CoT)= Having a model write out intermediate reasoning steps before the answer, which improves accuracy on multi-step problems.
- tree-of-thought-prompting= A prompting method that explores and evaluates several reasoning branches instead of one chain.
- Reasoning Model= An LLM trained to produce extended step-by-step reasoning before its final answer, for harder tasks.
- Test-Time Compute= Spending more computation at inference (longer reasoning, sampling several answers, search) to improve an answer, instead of only scaling training.
- Prompt Caching= Reusing the already-computed internal state for a repeated prompt prefix so later requests are faster and cheaper.
- inference-and-serving= Stored attention keys and values for tokens already processed, so each new token does not recompute the whole sequence; it grows with context length.
- Structured Outputs= Constraining a model to emit output that conforms to a schema (e.g. JSON), so software can parse it reliably.
- Function Calling= The model emits structured calls to external functions or APIs, which the application executes and returns; the open standard for exposing tools is model-context-protocol.
- retrieval-augmented-generation= Fetching relevant documents at query time and putting them in the prompt so the model answers from them; see retrieval-augmented-generation-overview.
- GraphRAG= RAG that retrieves over a knowledge graph or graph-derived summaries, not just text chunks, to answer questions needing relations across documents.
- Embedding= A numeric vector that represents a token, word, passage or image so that similar items are close together; used for search and RAG.
- Vector Database= A database that stores embeddings and returns the nearest ones to a query vector, usually via approximate indexes (see ann-index-algorithms).
- Chunking= Splitting documents into pieces (by size, structure or meaning) before embedding them for retrieval.
- Reranking= A second retrieval stage in which a more accurate model rescores the top candidates from a fast first-stage search.
- Hybrid Search= Combining keyword (lexical) and vector search and merging the rankings.
- Semantic Search= Retrieval by meaning (embedding similarity) rather than by exact keyword match.
- Temperature= A sampling setting that controls randomness: low gives more deterministic output, high more varied.
- Sampling (top-k / top-p)= Methods that restrict next-token choice to the k most likely tokens or the smallest set whose probability mass reaches p.
- Logits= The raw, unnormalised scores a model produces for each vocabulary token before softmax.
- Pre-training= The first, large-scale training phase on broad data (usually next-token prediction) that gives a model general capability.
- Post-training= Everything after pre-training that shapes behaviour: instruction tuning, preference optimisation, reinforcement learning.
- Instruction Tuning= Fine-tuning on instruction-and-response pairs so a base model follows requests.
- Hallucination= Fluent, plausible output that is factually wrong or unsupported by the source.
- Grounding= Tying a model’s answer to verifiable sources (retrieved documents, tool results) so claims can be checked.
- Multimodal Model= A model that handles more than one data type, such as text, images, audio and video.
- vision-models= A model that takes images (and text) as input and answers or reasons about them in language.
- Diffusion Model= A generative model that learns to reverse a gradual noising process, producing images, video or audio by iterative denoising.
- World Model= A model that learns how an environment evolves and responds to actions, so it can simulate or plan; also used for generative video that simulates scenes.
- Mixture of Experts (MoE)= An architecture in which a gating network routes each input to a few specialised sub-networks, so only part of the model’s parameters is used per token.
- State-Space Model (SSM)= A sequence model built on linear recurrent dynamics that scales linearly with sequence length, an alternative to attention.
- Mamba= A state-space architecture (Gu and Dao, 2023) with input-dependent (“selective”) state updates for efficient long sequences.
- Hybrid Architecture= A model mixing layer types, typically attention layers with state-space or linear-attention layers (or dense with MoE), to balance quality and speed.
- Dense Model= A model that uses all its parameters for every token, as opposed to an MoE.
- Multi-Token Prediction= Training or decoding a model to predict several future tokens at once; see multi-token-prediction.
- Encoder / Decoder= Transformer halves: an encoder builds representations of input, a decoder generates output token by token; most LLMs are decoder-only.
- Latent Space= The internal vector space in which a model represents data in compressed form.
3. Agents and agentic engineering
- AI Agent= A software system that uses a model (usually an LLM) to pursue a goal by choosing and executing actions, often via tools, in a loop.
- Agentic AI= The approach or class of systems in which AI plans, acts with tools, observes results and iterates with limited human steering, as opposed to single-turn question answering.
- Agent Harness= The code and configuration around a model that turns it into an agent: the loop, tools, context management, permissions, memory and hooks.
- Subagent= An agent started by another agent to handle a delegated subtask in its own context, returning only a result to the parent.
- Skill (agent skill)= A packaged folder of instructions, scripts and resources that an agent loads on demand when a task needs it.
- MCP (Model Context Protocol)= An open protocol for connecting AI applications to tools, data and prompts through servers, so one integration works across clients.
- A2A (Agent2Agent protocol)= An open protocol for communication and task hand-off between independent agents, possibly from different vendors; see also acp-agent-communication-protocol.
- AGENTS.md= A plain-markdown file in a repository that gives coding agents project instructions (build, test, conventions); see also claude-md-and-agent-instructions.
- Agentic AI Foundation= A neutral foundation hosting open agent standards; see the note for current members.
- Orchestrator= An agent or program that splits work, assigns it to worker agents or tools, and combines results; see ai-code-orchestrators-2026.
- Multi-Agent System= Several cooperating agents, each with a role, working on one task; see multi-agent-systems.
- Workflow (agentic)= A fixed, developer-defined sequence of model and tool steps, as opposed to an agent that decides its own steps.
- ReAct= A pattern that interleaves reasoning text with tool actions and observations (Yao et al., 2022).
- Planning= An agent step that breaks a goal into ordered subtasks before or while executing them; see task-decomposition-patterns.
- Memory (short-term)= State an agent keeps within a session: the current context, scratchpad and recent history.
- Memory (long-term)= Information persisted across sessions (facts, preferences, past outcomes) and retrieved when relevant.
- Computer Use= An agent controlling a graphical desktop by viewing screenshots and issuing mouse and keyboard actions.
- ai-browsers= An agent that navigates and acts in a web browser to complete tasks for the user.
- agent-sandboxes= An isolated execution environment (container, VM, restricted process) where an agent can run code and commands without risking the host.
- Worktree (git worktree)= A git feature giving a repository several working directories on different branches; coding-agent tools use it to let parallel agents edit without colliding.
- Human-in-the-Loop (HITL)= Design in which a person approves, corrects or steers the agent at defined points.
- Vibe Coding= Building software by describing intent to an AI and accepting its code with little or no review; the term is attributed to Andrej Karpathy (a post on X, dated 2025-02-02 per secondary reports; not rechecked).
- Spec-Driven Development= Writing a detailed specification first and having AI agents implement, test and revise code against it.
- Agentic Coding= Using AI agents that read the codebase, edit files, run tests and iterate to complete software tasks.
- Loop Engineering= Designing the repeated act-observe-verify cycle that an agent runs; see the note.
- agent-evals-and-observability= Tracing an agent’s model calls, tool calls and decisions so failures can be debugged and measured.
- Agent Gateway= A proxy layer that mediates agent-to-tool and agent-to-model traffic with routing, auth and policy; see mcp-security-with-gateway.
- Agent Payments Protocols= Standards that let agents pay or purchase on a user’s behalf with authorisation controls.
- Agent SDK= A library for building agents with the harness features (loop, tools, context handling) already provided.
- Context Graph= A graph of entities, relations and decision history used to give agents structured context.
- Software Factory= A setup where agents carry out most steps of software production under human specification and review.
- Compound Engineering= A workflow in which each completed task feeds lessons back into the agent’s instructions and tooling so later tasks improve.
- WebMCP= A proposal for web pages to expose tools to in-browser agents via MCP-style interfaces.
4. Training, inference and efficiency
- Fine-tuning= Further training of a pre-trained model on narrower data to change its behaviour; see LoRA below for a parameter-efficient method.
- LoRA (Low-Rank Adaptation)= A parameter-efficient fine-tuning method that trains small low-rank matrices added to frozen weights.
- QLoRA= LoRA applied on top of a quantised, frozen base model so large models can be fine-tuned on modest hardware.
- PEFT= Parameter-efficient fine-tuning: methods (LoRA, adapters, prompt tuning) that train few parameters instead of the whole model.
- Distillation= Training a smaller “student” model to imitate a larger “teacher” model’s outputs.
- Quantization= Storing weights (and sometimes activations) at lower numeric precision to cut memory and speed up inference, with some accuracy cost.
- Pruning= Removing weights, neurons or layers that contribute little, to shrink a model.
- Speculative Decoding= A small draft model proposes several tokens that the large model verifies in one pass, speeding generation without changing the output distribution.
- Continuous Batching= A serving technique that adds and removes requests from the running batch at each step instead of waiting for a whole batch to finish; see vllm.
- inference-and-serving= Running a trained model to produce outputs, as opposed to training it.
- inference-and-serving= The two inference phases: prefill processes the whole prompt in parallel; decode generates output tokens one at a time.
- inference-and-serving= A KV-cache memory-management method that stores the cache in non-contiguous blocks to reduce waste (associated with vLLM).
- Batch Size= The number of examples (or requests) processed together in one step.
- Throughput vs Latency= Throughput is total work per unit time; latency is the time one request takes. Serving trades one against the other.
- Scaling Laws= Empirical relationships showing how model loss falls predictably with more parameters, data and compute.
- RLHF= Reinforcement learning from human feedback: a reward model learns human preferences and the language model is optimised against it.
- RLVR= Reinforcement learning with verifiable rewards: the reward comes from an automatic check (e.g. tests pass, answer matches), used to train reasoning.
- DPO (Direct Preference Optimization)= Optimising a model directly on preference pairs with a classification-style loss, with no separate reward model or RL loop.
- GRPO= Group Relative Policy Optimization: an RL method that scores several sampled answers to a prompt against each other, avoiding a separate value model.
- Reward Model= A model trained to score outputs by human preference, used as the optimisation target in RLHF.
- Synthetic Data= Training or evaluation data generated by models or programs rather than collected from people.
- Model Collapse= Degradation when models are trained repeatedly on model-generated data, losing the diversity of the original distribution.
- Data Contamination= Test or benchmark material leaking into training data; see Benchmark Contamination below.
- Checkpoint= A saved snapshot of a model’s weights during or after training.
- Epoch and Step= An epoch is a pass over the data set; a step is one optimiser update.
- Mixed Precision= Training with a blend of 16-bit and 32-bit numbers to save memory and time.
- Parallelism (data, tensor, pipeline)= Ways to split training or serving across many accelerators: by data batches, by splitting matrices, or by splitting layers.
- Compute-Optimal Training= Balancing model size and training tokens for a fixed compute budget (the “Chinchilla” finding, 2022); see Scaling Laws.
- Edge / On-Device Inference= Running models locally on phones or PCs instead of in a data centre; see ollama.
- Model Routing= Choosing, per request, which model answers (a smaller, cheaper one for easy requests, a larger one for hard ones), usually in a gateway or router layer in front of several models.
5. Safety, security and evaluation
- Alignment= Making AI systems pursue the goals and values their developers and users intend.
- Sycophancy= A model’s tendency to agree with or flatter the user instead of giving accurate answers.
- Reward Hacking= An RL-trained model exploiting flaws in the reward signal to score high without doing the intended task.
- Constitutional AI= Anthropic’s technique (paper, 2022; not re-fetched) in which a model critiques and revises its own outputs against a written set of principles, using AI feedback to train for harmlessness.
- Interpretability= Research into understanding what happens inside a model, from circuits and features to causes of behaviour; “mechanistic interpretability” reverse-engineers the internals.
- Guardrails= Checks and filters around a model’s input, output or actions that block unsafe or out-of-policy behaviour.
- Red Teaming= Adversarially probing a system to find failures and vulnerabilities before attackers or users do.
- Jailbreak= A prompt or technique that gets a model to bypass its safety training or rules.
- Prompt Injection= Hidden or untrusted text that hijacks an agent’s instructions; “indirect” injection arrives via web pages, files or tool outputs.
- Data Exfiltration= An attacker getting private data out of a system, a typical goal of prompt injection against agents.
- Excessive Agency= A security risk where an agent holds more permissions or autonomy than a task needs.
- Evals= Systematic tests that measure how well a model or application performs a task.
- Benchmark= A standard test set and scoring method used to compare models.
- Benchmark Contamination= Benchmark questions or answers appearing in training data, inflating scores.
- LLM-as-Judge= Using a language model to grade another model’s outputs against criteria.
- Eval Harness= Software that runs a model through benchmarks or tests in a repeatable way and records scores.
- MMLU= Massive Multitask Language Understanding: a multiple-choice benchmark across 57 subjects (Hendrycks et al., 2020).
- SWE-bench= A benchmark that asks a model to resolve real GitHub issues in open-source repositories, checked by the repository’s tests.
- GPQA= Graduate-level Google-Proof Q&A: hard multiple-choice science questions written by domain experts (2023).
- HLE (Humanity’s Last Exam)= A very hard, expert-written, closed-ended question set built to stay challenging as models improve (2025).
- ARC-AGI= A benchmark of novel visual grid puzzles (Chollet’s Abstraction and Reasoning Corpus) meant to test skill acquisition on unseen tasks.
- Model Card= A document describing a model’s intended use, training data, performance and limits (Mitchell et al., 2019).
- System Card= A document covering a deployed AI system, including safety evaluations, mitigations and risks, beyond the model alone.
- AI Safety= The field aimed at preventing harm from AI systems, from misuse to loss of control.
- Watermarking= Embedding a hidden statistical signal in generated text, images or audio so it can later be detected as AI-made.
- Content Provenance (C2PA)= A standard for attaching signed metadata to media recording how it was made or edited.
- Deepfake= Synthetic or manipulated audio, image or video that realistically depicts someone saying or doing something they did not.
- EU AI Act= The European Union’s risk-based AI regulation (Regulation 2024/1689), in force from August 2024 (per Wikipedia, secondary) with obligations phased in over following years; see EU AI Act for amended dates.
- GPAI (General-Purpose AI)= The EU AI Act’s term for models with broad capability across tasks; their providers have transparency and copyright duties, with extra duties for models posing systemic risk.
- AI Ethics= The study of fairness, accountability, transparency and societal impact of AI.
- Data Privacy (AI)= Protecting personal data used in training and prompts.
- Differential Privacy= A mathematical guarantee that a result barely depends on any single individual’s data.
- Human Oversight= Keeping people able to monitor, intervene in and override an AI system.
- Model Welfare= Research on whether AI systems could have experiences or interests that matter morally, and what low-cost precautions follow if so; a contested, early-stage field.
6. Models, products and infrastructure vocabulary
- Foundation Model= A large model trained on broad data that can be adapted to many downstream tasks (Stanford CRFM, 2021).
- Frontier Model= One of the most capable general-purpose models at a given time, at the edge of current capability.
- Open-Weight Model= A model whose trained weights are downloadable, usually under a licence that may limit use; training data and code are often not released.
- Open Source AI= AI released with source, weights and enough information to study, use, modify and share; the Open Source Initiative has a formal definition; many models called “open” release weights only.
- language-models= A compact language model designed to run cheaply or on-device.
- Base Model vs Instruct Model= A base model only predicts text; an instruct (chat) model has been post-trained to follow instructions.
- Model Family and Series= Related models released under a shared name and generation; see language-models.
- Distilled Model= A smaller model produced by distillation.
- Sovereign AI= National or regional control over AI compute, data, models and talent, so a country does not depend on foreign providers.
- Neocloud= A specialised cloud provider built around GPU capacity for AI workloads, as opposed to the general hyperscalers.
- Hyperscaler= A very large cloud provider such as AWS, Azure or Google Cloud.
- inference-and-serving= A company that hosts models and sells API access to run them, often including open-weight models.
- Router (LLM)= A service that fronts several model providers behind one API with routing, fallback, caching and cost tracking; see PortKey and claude-code-router.
- API Platform= A vendor service exposing models through an API with keys, quotas and tooling.
- Ollama= A tool for running open-weight models locally.
- vLLM= An open-source high-throughput LLM inference and serving engine.
- Aider= Command-line AI pair-programming tool.
- autogen= Microsoft’s multi-agent framework (described in the vault audit KB-AI-AUDIT-2026-09 as maintenance-only, converging into microsoft-agent-framework; status not rechecked).
- LangGraph= A framework for building stateful agent workflows as graphs.
- CrewAI= A framework for role-based multi-agent teams.
- AI Slop= Low-quality, mass-produced AI-generated content published with little care or review.
- Copilot= A product pattern where an AI assists a user inside an application, as opposed to acting alone.
- Rate Limit= A cap on requests or tokens per time period imposed by an API provider.
- Speech Recognition (ASR)= Converting speech audio into text; see whisper-and-asr-models.
- Text-to-Speech (TTS)= Generating spoken audio from text; see also speech-synthesis.
- Voice Activity Detection (VAD)= Detecting whether an audio segment contains speech, used to decide when a speaker starts and stops.
- Voice Cloning= Synthesising speech that imitates a specific person’s voice; see voice-cloning.
- GPU= A graphics processing unit; the parallel processor that dominates AI training and inference.
- TPU= Tensor Processing Unit, Google’s custom AI accelerator.
- NPU= Neural processing unit: an on-chip accelerator for AI inference in phones and PCs.
- HBM= High Bandwidth Memory: stacked memory placed next to AI accelerators; its capacity and bandwidth often limit inference speed.
- Data Center= A facility housing servers, networking and power for large-scale computing.
- AI Accelerator= Any chip designed to speed up machine-learning workloads (GPU, TPU, NPU and others).
- Hugging Face= A hub for sharing models, datasets and libraries, widely used to distribute open-weight models.
- AI Research Lab= An organisation, commercial or academic, doing frontier AI research; see the note.
7. Abbreviations (A-Z)
- AGI= Artificial General Intelligence: a hypothetical AI matching human-level ability across most cognitive tasks; definitions vary and are disputed.
- AI= Artificial Intelligence.
- ANN= Artificial Neural Network; also Approximate Nearest Neighbour search, see ann-index-algorithms.
- API= Application Programming Interface: a defined way for programs to call a service.
- ASI= Artificial Superintelligence: a hypothetical AI far surpassing human ability in virtually every domain.
- ASR= Automatic Speech Recognition (speech-to-text).
- A2A= Agent2Agent protocol, see A2A (Agent2Agent protocol).
- BERT= Bidirectional Encoder Representations from Transformers, an encoder-only model (2018).
- BPE= Byte-Pair Encoding, see Tokenizer.
- CNN= Convolutional Neural Network.
- CoT= Chain of Thought, see chain-of-thought-prompting.
- CPU= Central Processing Unit.
- DL= Deep Learning.
- DM= Data Mining.
- DPO= Direct Preference Optimization, see RLHF.
- FLOPs= Floating-point operations: a count of arithmetic work, used to measure training compute (FLOPS with a capital S means operations per second).
- GAN= Generative Adversarial Network.
- GPAI= General-Purpose AI (EU AI Act term).
- GPQA= Graduate-Level Google-Proof Q&A, a benchmark.
- GPT= Generative Pretrained Transformer, see Generative Pretrained Transformer.
- GPU= Graphics Processing Unit.
- GRPO= Group Relative Policy Optimization, see RLHF.
- HBM= High Bandwidth Memory.
- HITL= Human-in-the-Loop.
- HLE= Humanity’s Last Exam, a benchmark.
- IP= Internet Protocol, or Intellectual Property depending on context.
- KV= Key-Value (as in KV cache), see inference-and-serving.
- LLM= Large Language Model, see large-language-model.
- LMM= Large Multimodal Model, see Multimodal Model.
- LoRA= Low-Rank Adaptation.
- LVM= Large Vision Model.
- MCP= Model Context Protocol, see model-context-protocol.
- MFU= Model FLOPs Utilization: the share of an accelerator’s peak compute that training or inference actually achieves.
- ML= Machine Learning, see machine-learning.
- MMLU= Massive Multitask Language Understanding, see MMLU.
- MoE= Mixture of Experts.
- NLP= Natural Language Processing.
- NN= Neural Network, see neural-networks.
- NPU= Neural Processing Unit.
- PEFT= Parameter-Efficient Fine-Tuning.
- QLoRA= Quantized LoRA.
- RAG= Retrieval-Augmented Generation, see retrieval-augmented-generation.
- RL= Reinforcement Learning.
- RLHF= Reinforcement Learning from Human Feedback, see RLHF.
- RLVR= Reinforcement Learning with Verifiable Rewards, see RLHF.
- RNN= Recurrent Neural Network.
- SFT= Supervised Fine-Tuning: fine-tuning on labelled input-output examples.
- SLM= Small Language Model.
- SSM= State-Space Model.
- STT= Speech-To-Text.
- SWE-bench= Software Engineering benchmark, see SWE-bench.
- TPOT= Time Per Output Token: the average time between generated tokens, a serving-latency metric.
- TPU= Tensor Processing Unit.
- TTFT= Time To First Token: delay from sending a request to receiving the first output token.
- TTS= Text-To-Speech.
- VAD= Voice Activity Detection.
- VLM= Vision-Language Model.
- XAI= Explainable AI.
8. Further reading (Wikipedia, reference pages)
Wikipedia: Algorithm https://en.wikipedia.org/wiki/Algorithm
Wikipedia:Artificial Intelligence https://en.wikipedia.org/wiki/Artificial_intelligence
Wikipedia: Artificial Neural Network https://en.wikipedia.org/wiki/Artificial_neural_network
Wikipedia: Autonomous Robot https://en.wikipedia.org/wiki/Autonomous_robot
Wikipedia: Cognitive Computing https://en.wikipedia.org/wiki/Cognitive_computing
Wikipedia: Computer Vision https://en.wikipedia.org/wiki/Computer_vision
Wikipedia: Data Mining https://en.wikipedia.org/wiki/Data_mining
Wikipedia: Deep Learning https://en.wikipedia.org/wiki/Deep_learning
Wikipedia: Expert System https://en.wikipedia.org/wiki/Expert_system
Wikipedia: Fuzzy Logic https://en.wikipedia.org/wiki/Fuzzy_logic
Wikipedia: Genetic Algorithm https://en.wikipedia.org/wiki/Genetic_algorithm
Wikipedia: Machine Learning https://en.wikipedia.org/wiki/Machine_learning
Wikipedia: Natural Language Processing https://en.wikipedia.org/wiki/Natural_language_processing
Wikipedia: Reinforcement Learning https://en.wikipedia.org/wiki/Reinforcement_learning
Wikipedia: Robotics https://en.wikipedia.org/wiki/Robotics
Wikipedia: Supervised Learning https://en.wikipedia.org/wiki/Supervised_learning
Wikipedia: Unsupervised Learning https://en.wikipedia.org/wiki/Unsupervised_learning
Wikipedia: Bias in artificial intelligence https://en.wikipedia.org/wiki/Bias_in_artificial_intelligence
Wikipedia: Explainable artificial intelligence https://en.wikipedia.org/wiki/Explainable_artificial_intelligence
Wikipedia: Glossary of artificial intelligence https://en.wikipedia.org/wiki/Glossary_of_artificial_intelligence
Wikipedia: Large language model https://en.wikipedia.org/wiki/Large_language_model
Wikipedia: Regulation (EU) 2024/1689 (AI Act) https://en.wikipedia.org/wiki/Regulation_(EU)_2024/1689
Sources (accessed 2026-10-05, earlier pass 2026-10-02)
- https://en.wikipedia.org/wiki/Glossary_of_artificial_intelligence (fetched 2026-10-05; only partly covers the 8 newly defined classical terms)
- https://en.wikipedia.org/wiki/Regulation_(EU)_2024/1689 (fetched 2026-10-05: GPAI obligations, 12 months after entry into force)
- https://en.wikipedia.org/wiki/Large_language_model, /Fuzzy_logic, /Federated_learning (2026-10-02)
- Vault notes linked above (concept notes in KB-AI/ai-concepts, ai-prompt-engineering/context-engineering, ai-memory, ai-agentic-systems, ai-infrastructure, ai-coding) were used for terminology consistency.
- Stable definitions rely on standard textbook definitions and primary papers (Vaswani et al. 2017 Transformer; Gu and Dao 2023 Mamba; Hendrycks et al. 2020 MMLU; Yao et al. 2022 ReAct; Mitchell et al. 2019 model cards; Anthropic 2022 Constitutional AI); these papers were not re-fetched. Definitions of fast-moving agent terms (subagent, harness, skill, worktree usage, A2A, AGENTS.md) are working definitions from the vault notes, not re-verified against primary docs.
- Audit note: removed the abbreviations “SAM = Small Action Model” and “LCM = Large Context Model” (no source found). AutoGen status: KB-AI-AUDIT-2026-09.md Appendix A.