I wrote an X thread on 2026‑10‑04 about the major vector‑search indexing approaches used by Google and Microsoft, summarising their trade‑offs and when I choose each in practice.
A basic nearest‑neighbor query computes the distance between the query and every stored vector.
binary quantization can make RAG search 32x more memory-efficient while querying over 36 million vectors in under 30 ms.
- Flat index – exact search, linear cost, best for small datasets.
I pick Flat when exactness matters, HNSW for latency‑critical workloads, IVF when I need probe‑level control, and IVF‑PQ or ScaNN when memory and compute must be minimised.
- IVF – cluster‑based, probe parameter controls workload.
- HNSW – multilayer graph, low latency, extra memory.
- IVF‑PQ – combines IVF with product quantization for memory/compute savings.
- ScaNN – partition‑select‑re‑rank pipeline for large‑scale high‑throughput retrieval.