TurboVec

by community / open source (MIT license)

A Rust-based vector indexing library with Python bindings for fast, memory-efficient local semantic search — built on Google Research’s TurboQuant algorithm.

Features

  • Rust implementation with Python bindings and pip install support
  • 4-bit quantization to reduce memory usage while maintaining fast and accurate search
  • Data-oblivious quantization: no training phase, online ingestion, incremental saves, runtime filtering; repo claims a 10M-document corpus fits in 4 GB instead of 31 GB and search faster than FAISS (vendor claim, unverified)
  • Build once, save, and reuse the index unless underlying documents change
  • Integrates with BGE-M3 via Ollama for embedding
  • Designed for local document Q&A / RAG workflows over PDFs and text files
  • Semantic search over keyword search — retrieved context fed strictly to the LLM
  • Hallucination mitigation: strict prompting forces answers from retrieved context only

Superpowers

TurboVec is the missing piece for engineers who want a fast, self-hosted RAG stack without cloud dependencies. Built in Rust with 4-bit quantization, it delivers fast vector search at a fraction of the memory cost of other local indexing solutions. The typical pipeline: load BGE-M3 via Ollama, chunk documents, embed, index with TurboVec, save the index, and reuse it until documents change. At query time, embed the query, retrieve top-k chunks, dedupe, and send to the LLM with a strict prompt that only uses retrieved context. This is the stack Gao Dalie demonstrated alongside DeepSeek V4 as a fully local, open-weight RAG solution approachable for developers using Python and Ollama.

Sources