
Approximate Nearest-Neighbor Search: Recall, Memory, and Latency Tradeoffs
A retrieval system answers in four milliseconds over a million vectors. Your first instinct is that the index must be computing similarity faster than…
Read tutorialFinding and ranking information by meaning-related representations, including vector, hybrid, and approximate-neighbor retrieval methods.
Tagged articles
18 articles in this tag.

A retrieval system answers in four milliseconds over a million vectors. Your first instinct is that the index must be computing similarity faster than…
Read tutorial
A clean pipeline that answers confidently and wrongly is not a mystery. It is a trace you have not read yet.
Read tutorial
A retrieval system returns a top result with a similarity score of 0.83. Most people read that number as "83% relevant" or "83% likely to be correct." It…
Read tutorial
The reranker "feels" better. The top result looks more relevant than it did before. And yet you cannot say whether the stage earned its latency, because…
Read tutorial
Your RAG system just produced a confident, polished, and completely wrong answer. The natural instinct is to blame the model—swap the prompt, change the…
Read tutorial
Two retrievers return two different top-5 lists for the same query. You merge them, and the merged order looks like it was decided by a coin flip. It…
Read tutorial
Most teams add hybrid retrieval because a vendor page promised better accuracy, then never check whether it helped their own corpus. This exercise forces…
Read tutorial
A query returns a sentence the source no longer contains. The team blames the model. The model is innocent — the index is still holding a chunk that was…
Read tutorial
You can explain RAG. You can sketch an agent loop. But when a real question lands on your desk, do you know whether to retrieve, route, or just answer?…
Read tutorial
You know RAG retrieves documents and feeds them to a language model. But if someone asked you what actually happens between a source PDF and a grounded…
Read tutorial
Your knowledge base has the answer. You know it does. You wrote the documents, chunked them carefully, embedded them, and verified the chunks look right.…
Read tutorial
Your retriever returns ten chunks. They all look topically related to your question. None of them actually answer it. The LLM dutifully composes a response…
Read tutorial
Two retrievers return the same relevant chunk. One puts it at rank 1; the other buries it at rank 4. A single "accuracy" number calls them identical. The…
Read tutorial
Reading about RAG is easy. Building one is where the real learning happens—because the first pipeline you run will almost certainly fail in ways no diagram…
Read tutorial
You swap in a stronger reranker. The top three results look sharper, the ordering feels smarter, and the demo lands. Then you measure recall and it barely…
Read tutorial
A filter can be correct and still delete your best evidence. Here is how to watch it happen in twenty lines of NumPy.
Read tutorial
A vector database is not a fancy search box or a generic "database for AI." It is the retrieval engine that decides which evidence earns a seat on the…
Read tutorial
You search your company's help docs for "how do I reset my password?" and get nothing useful. You know the answer is in there—you've read the page…
Read tutorial