Shashank
← All work

Hybrid retrieval engine

Answering questions over a private corpus with citations you can check and retrieval quality you can measure. Three channels fused in a single database query, reranked, and shipped with the evaluation harness most RAG systems never get.

The problem

Status
In production
Channels
Dense, lexical, exact match
Fusion
Reciprocal rank fusion
Role
Sole engineer

A vector database and a prompt produce something that looks like a working product. The trouble arrives later. Retrieval quietly returns the wrong passages, the model fills the gap with plausible invention, and nobody notices because no number is tracking whether answers got worse.

The failure is specific and predictable: vector search is unreliable on exact strings. Part numbers, policy references, API names and product codes are precisely the things people type, and precisely what semantic similarity handles badly.

Architecture

QUERY RETRIEVE FUSE REFINE ANSWER CHECK rejects unverifiable claims Query Dense vectors pgvector HNSW Full-text search tsvector Exact match pg_trgm Rank fusion RRF Rerank cross-encoder + MMR Generate schema-constrained Verify
Per-stage tracing · Golden datasets · recall@k · nDCG@k · MRR · Ablation runs

All three retrieval channels run inside one database query and are combined by rank rather than by score, so no per-corpus weight tuning is needed.

Decisions worth defending

Fusion by rank, not by score

Cosine similarity, a full-text relevance score and a trigram hit are on incomparable scales. Blending them means retuning weights every time the data changes. Rank is the only thing the three channels agree on, and fusing on it needs no tuning at all.

Citations are verified, not trusted

The model returns quotes attached to source references, and those quotes are checked programmatically against what it was actually shown. Names and identifiers appearing in an answer are checked against the index. A model that cites a source for a claim it invented is the most common failure in this category, and it is entirely detectable.

Context is packed, never truncated

Sources are admitted whole under a real token budget, in rank order. A passage that does not fit is dropped and reported rather than cut in half. Slicing the finished prompt to length removes the evidence first, which is exactly backwards.

Degradation is reported, not absorbed

When the embedding provider is unavailable, the lexical channels keep serving and the response says it is degraded. Silently falling back to whatever can be found is how a broken retrieval path hides for months behind answers that merely look plausible.

The harness comes first

A golden dataset and ablation runs exist before tuning starts. On the first run the harness caught a default setting that was making results worse, which is the entire argument for building it early rather than after launch.

Stack

Storage

  • PostgreSQL with pgvector
  • HNSW with iterative scan
  • Generated tsvector column
  • Trigram index

Retrieval

  • Reciprocal rank fusion
  • Cross-encoder reranking
  • Maximal marginal relevance
  • Deterministic entity extraction

Service

  • Python, FastAPI, asyncpg
  • Pydantic throughout
  • Typer command line
  • Docker

Evaluation

  • Golden datasets
  • recall@k, nDCG@k, MRR
  • Stage ablations
  • Latency percentiles
← VectorLead All work →