TurboVec
by community / open source (MIT license)
A Rust-based vector indexing library with Python bindings for fast, memory-efficient local semantic search — built on Google Research’s TurboQuant algorithm.
Features
- Rust implementation with Python bindings and pip install support
- 4-bit quantization to reduce memory usage while maintaining fast and accurate search
- Data-oblivious quantization: no training phase, online ingestion, incremental saves, runtime filtering; repo claims a 10M-document corpus fits in 4 GB instead of 31 GB and search faster than FAISS (vendor claim, unverified)
- Build once, save, and reuse the index unless underlying documents change
- Integrates with BGE-M3 via Ollama for embedding
- Designed for local document Q&A / RAG workflows over PDFs and text files
- Semantic search over keyword search — retrieved context fed strictly to the LLM
- Hallucination mitigation: strict prompting forces answers from retrieved context only
Superpowers
TurboVec is the missing piece for engineers who want a fast, self-hosted RAG stack without cloud dependencies. Built in Rust with 4-bit quantization, it delivers fast vector search at a fraction of the memory cost of other local indexing solutions. The typical pipeline: load BGE-M3 via Ollama, chunk documents, embed, index with TurboVec, save the index, and reuse it until documents change. At query time, embed the query, retrieve top-k chunks, dedupe, and send to the LLM with a strict prompt that only uses retrieved context. This is the stack Gao Dalie demonstrated alongside DeepSeek V4 as a fully local, open-weight RAG solution approachable for developers using Python and Ollama.
Sources
- https://github.com/RyanCodrai/turbovec (accessed 2026-09-30)
- Workflow description (BGE-M3 via Ollama, DeepSeek V4 demo) carried over from the original note; not re-verified.