When vector search needs a keyword retrieval branch
A RAG system retrieves passages before asking a model to answer. When the right passage never reaches the model, changing the prompt may leave the underlying problem untouched. Start by inspecting the retrieved evidence for a small set of failed queries.
Consider two illustrative queries: 'How do I restore a failed backup?' and 'ERR_BACKUP_214'. The first asks about a concept; the second contains a specific identifier. Evaluate both styles against your corpus instead of assuming one retrieval method handles them equally well. These are proposed test cases, not measured failures of a particular embedding model.
Run a vector-only baseline and a keyword-only baseline, then compare which relevant passages each misses. Adding a second branch is worth investigating when it finds useful evidence the first branch misses. If both fail because the document is missing, outdated or incorrectly chunked, repair ingestion before tuning fusion.
PostgreSQL full-text search is not automatically BM25
BM25 and PostgreSQL's built-in text ranking are different choices for lexical retrieval. PostgreSQL 16 provides ts_rank, based on matching lexeme frequency, and ts_rank_cd, based on cover density including proximity. Calling ts_rank_cd 'BM25-style' hides a material implementation difference. A BM25 design needs an implementation that explicitly provides BM25.
PostgreSQL text search transforms text into lexemes using a configured parser and dictionaries. It does not promise a literal substring match for every identifier. plainto_tsquery joins surviving terms with AND; this can be restrictive for a long question. Inspect how your actual queries are parsed.
For identifiers that must match exactly, consider a separate normalized identifier field and an equality lookup. Decide explicitly whether case, hyphens and leading zeros are significant. Test that policy with domain examples; do not make access or billing decisions from an approximate semantic match.
References: PostgreSQL 16 — Controlling Text Search
Why combine ranks instead of raw scores?
Vector and lexical rankers use different scoring systems. Directly adding their raw outputs makes the result depend on their scales. A calibrated weighted combination can be a valid alternative, but its calibration needs evaluation on representative queries.
RRF offers a rank-based baseline: sort each branch, assign positions starting at one, and add reciprocal rank contributions for each candidate. It discards score magnitudes, so it also discards information about how far apart two candidates scored. Compare that tradeoff rather than assuming fusion is always better.
Use one candidate identity throughout the pipeline, such as a chunk ID tied to a document version. If one branch returns documents and the other returns chunks, define the mapping before fusion. Otherwise duplicates can consume the context budget or receive inconsistent credit.
A worked Reciprocal Rank Fusion example
For candidate d, RRF(d) = sum of 1 / (k + rank(d)) over the lists containing d. A missing candidate contributes zero for that list. k is the smoothing constant; it is separate from the number of candidates you retrieve.
The table uses k = 60 and fictional ranks. Candidate A appears in both lists and outranks candidates appearing in only one. The arithmetic demonstrates the method; it says nothing about whether A contains a correct answer.
The original paper evaluated k = 60. Treat it as a starting configuration, not a universal optimum. Corpus changes can change ranks and therefore fused results. RRF is not immune to those changes and its score is not a probability of relevance.
| Candidate | Vector rank | Keyword rank | RRF score (k = 60) |
|---|---|---|---|
| A | 2 | 3 | 1/62 + 1/63 = 0.032002 |
| B | 1 | Absent | 1/61 = 0.016393 |
| C | Absent | 1 | 1/61 = 0.016393 |
One PostgreSQL store or separate search services?
pgvector documents combining vector search with PostgreSQL full-text search. For a team already operating PostgreSQL, that is a reasonable baseline to evaluate. Separate services may fit different indexing or scaling requirements. Choose from measured requirements and operational capacity, not a universal document-count cutoff.
The following questions are our proposed architecture review. Keeping records together can simplify the design, but it does not automatically make asynchronous embeddings current or authorization correct. Separate stores require explicit coordination of updates, deletions and permissions.
| Concern | Single PostgreSQL store | Separate search services |
|---|---|---|
| Freshness | How are document and embedding versions published together? | How are indexing lag and deletion propagation detected? |
| Authorization | Do both branches enforce the same tenant and access rules? | How are access changes propagated and rechecked? |
| Capacity | Can search share resources with transactional traffic? | Which workloads need independent scaling? |
| Recovery | Can backups restore documents, vectors and index configuration? | Can failed synchronization be replayed without resurrecting deleted content? |
| Cost and latency | Measure under representative filters and concurrency. | Include network calls, synchronization and operating effort. |
References: pgvector — Open-Source Vector Similarity Search for PostgreSQL
A PostgreSQL implementation plan to validate
This is a proposed workflow, not a tested SQL recipe. Pin the PostgreSQL and pgvector versions, embedding model and dimensions, text-search configuration and index settings before collecting results. Record those choices alongside your evaluation dataset.
pgvector's approximate indexes can return fewer matches under selective filters because filtering is applied after the index scan. Its documentation describes tuning and iterative scans for supported versions. Compare against exact search when diagnosing missing neighbors; do not remove authorization filters to fill the result list.
- Store chunk ID, document version, tenant, access metadata, content and embedding. Track ingestion failures and missing vectors explicitly.
- Generate the text-search representation with a documented language configuration. Verify identifier handling with representative documents and queries.
- Retrieve a bounded set from each branch using identical eligibility rules. Keep vector distance ordering separate from descending lexical relevance ordering.
- Assign deterministic ranks to the retrieved candidates, then merge their union by chunk identity. Deduplicate each branch and use a documented tie policy.
- Apply RRF and select passages within the answer context budget. Preserve source references and verify that the user can still access the selected document versions.
- Inspect query plans and measure latency under concurrency. CTE syntax alone does not establish parallel execution or a latency target.
References: pgvector — Open-Source Vector Similarity Search for PostgreSQL
Build an evaluation dataset before claiming improvement
Create a versioned set of questions with judged relevant passages. Include paraphrases, exact identifiers, ambiguous questions, outdated documents and questions with no answer in the corpus. Add access-control fixtures for users who must not receive otherwise relevant material. Use permitted, sanitized examples rather than exporting private queries into a public benchmark.
Compare keyword-only, vector-only and hybrid retrieval on the same corpus snapshot and access rules. Keep a held-out set separate from the queries used to tune candidate counts and fusion settings. Report results by query type so a gain on common questions does not hide regressions on identifiers.
Recall@K measures the fraction of judged relevant items retrieved in the first K positions. MRR averages the reciprocal position of the first relevant result, with zero when none is found within the evaluated window. State that window and how you handle questions with no relevant answer. Incompletely judged corpora limit what these metrics establish.
Also measure latency, failures, cost per query and downstream answer support. Check whether cited passages actually justify the generated answer and whether the system abstains when evidence is insufficient. This article reports no measured recall improvement, customer outcome or reduction in hallucinations.
Define failure handling before connecting the answer model
A retrieval pipeline needs an explicit policy for partial failure. Decide in advance whether a vector-branch timeout permits a lexical-only response, a limited result with a clear status, or no answer. Record which path ran so degraded behavior is visible in your evaluation and monitoring.
Treat authorization failures differently from a relevance shortfall. Do not retry by dropping tenant or permission constraints. Test timeouts, unavailable embedding services, malformed inputs, stale vectors and deletion races with the same care as successful queries.
Keep retries bounded within the request deadline. Log operational signals without copying private passages or raw user queries into unrestricted logs. Include a request identifier, branch status, candidate count, duration and configuration version where appropriate.
For an initial project discussion, describe the corpus, a few sanitized examples of failed searches, your access restrictions and the latency requirement. These inputs help decide whether the next step is dataset repair, a retrieval evaluation or an architecture change. Use the project form below to discuss your RAG retrieval problem.
