Multimodal Models

Multimodal models process and/or generate more than one modality: text, images, audio, video, sometimes actions. Common designs encode images or audio into tokens that a language model reads alongside text. CLIP (2021) aligned images and text in a shared embedding space; Visual Instruction Tuning (LLaVA, 2023) attached a vision encoder to an LLM and tuned it on instructions; GPT-4 (2023) was announced as accepting image and text input.

Kinds

Related: large-language-model, tokenization, computer-vision.

Sources

Open items

  • Current model-by-modality capability table not included; see the model notes.