Embeddings and Vector Search
An embedding maps text, images or other data to a vector so that similar items sit close together. Vector search finds nearest neighbours of a query vector, using cosine similarity or dot product, usually through approximate-nearest-neighbour indexes (HNSW, IVF, product quantization). It powers semantic search and the retrieval step of RAG.
Practice
- Embed documents once, store in a vector database or search engine; embed the query at run time with the same model. Switching model means re-embedding.
- Chunking, hybrid keyword + vector search and reranking matter as much as the model: hybrid-search-and-rank-fusion, rerankers-and-cross-encoders, query-understanding.
- Multi-vector approaches such as ColBERT: exploring-colbert-with-ragatouille. Graph-based alternatives: graphrag.
- Evaluate with ir-ranking-metrics and rag-evaluation.
Vault notes
Models: embedding-models, embeddinggemma. Index algorithms: ann-index-algorithms, turbovec. Databases: sql-vector-databases-transforming-genai-applications. Related concepts: tokenization, agent-memory-systems, multimodal-models.
Sources
- https://en.wikipedia.org/wiki/Word_embedding (secondary)
- https://arxiv.org/abs/2005.11401 (Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, submitted 2020-05-22; checked via arXiv API 2026-10-07)