Hybrid Search RAG Over Internal Docs: Production Guide
Hybrid search RAG over internal docs combines vector and keyword search to improve retrieval accuracy. I’ve used this in production for internal knowledge bases and support ticket systems, and it cuts down on missed relevant chunks by 30-40% compared to pure vector or keyword alone.
What is hybrid search in RAG?
Hybrid search in RAG means running both dense vector retrieval (like embeddings) and sparse keyword search (like BM25) in parallel, then fusing the results. You get the semantic understanding of vectors plus the exact-match precision of keywords. This matters when your internal docs have jargon, acronyms, or exact phrases that embeddings might miss.
Combining vector and keyword search for internal documents
Internal docs are messy. They have code snippets, product names like “WidgetX Pro”, and internal abbreviations. Pure vector search might return semantically similar but irrelevant chunks. Pure keyword search misses synonyms or paraphrased answers. Hybrid fixes this by giving you both. At LogicLoop, we saw a 22% increase in MRR@10 after switching from vector-only to hybrid on our engineering wiki.
Implementing hybrid search with Elasticsearch and pgvector
Here’s how we wired it up in Python with FastAPI. First, store your chunks in both systems:
# Store in Elasticsearch (keyword/BM25)from elasticsearch import Elasticsearches = Elasticsearch("http://localhost:9200")es.index(index="docs", id chunk_id, body={"text": chunk_text, "metadata": metadata})
# Store in pgvector (dense vectors)import asyncpgconn = await asyncpg.connect(dsn=DB_DSN)await conn.execute( "INSERT INTO doc_chunks (id, embedding, text) VALUES ($1, $2, $3)", chunk_id, embedding.tolist(), chunk_text)At query time, run both searches and fuse scores:
def hybrid_search(query_text, query_vector, k=10): # Keyword search via ES es_results = es.search( index="docs", body={ "query": {"match": {"text": query_text}}, "size": k * 2 # get more to fuse later } ) keyword_hits = [(hit["_id"], hit["_score"]) for hit in es_results["hits"]["hits"]]
# Vector search via pgvector vector_results = await conn.fetch( "SELECT id, text, embedding <=> $1 AS distance FROM doc_chunks ORDER BY distance LIMIT $2", query_vector, k * 2 ) vector_hits = [(row["id"], 1 - row["distance"]) for row in vector_results] # convert to similarity
# Simple score fusion: normalize and add all_scores = {} max_es = max([s for _, s in keyword_hits]) if keyword_hits else 1 max_vec = max([s for _, s in vector_hits]) if vector_hits else 1
for doc_id, score in keyword_hits: all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_es for doc_id, score in vector_hits: all_scores[doc_id] = all_scores.get(doc_id, 0) + score / max_vec
# Return top k fused sorted_ids = sorted(all_scores.items(), key=lambda x: x[1], reverse=True)[:k] return [doc_id for doc_id, _ in sorted_ids]We normalize scores to [0,1] per system before adding. This prevents one system from dominating due to scale differences.
Evaluating hybrid search performance on enterprise data
We tested on 500k internal Confluence pages and Jira tickets. Metrics: MRR@10 and recall@10. Baseline: pure vector (pgvector only).
| Method | MRR@10 | Recall@10 | Latency (p95) |
|---|---|---|---|
| Vector only | 0.41 | 0.68 | 120ms |
| Keyword only | 0.33 | 0.52 | 80ms |
| Hybrid | 0.50 | 0.79 | 150ms |
Hybrid added ~30ms latency but gained significant quality. The cost? Running two indexes doubles storage and write overhead. We mitigated this by using Elasticsearch’s snapshot lifecycle and pgvector’s partitioning by doc source.
Failure mode we hit: when query text is very short (1-2 words), keyword search dominates and vector adds noise. We now route sub-3-word queries to keyword-only via a simple length check.
Hybrid search RAG architecture diagram
See our full RAG pipeline diagram for context, but here’s the hybrid-specific flow:
- User query enters FastAPI endpoint
- Query split:
- Text → Elasticsearch (BM25)
- Text → Embedding model (e.g., bge-small-en) → pgvector
- Results fused via score normalization and weighted sum
- Top-k chunks passed to LLM (e.g., Mistral 7B) for generation
- Response returned
We keep the fusion logic in a dedicated service so we can swap fusion algorithms (like RRF or learning-to-rank) without touching the LLM layer.
Tools and libraries for hybrid search RAG
- Elasticsearch: Mature, battle-tested for keyword. Use the
knnandmatchcombo in one query if on 8.0+. - pgvector: Simple POSTGRES extension. Great if you’re already on PG.
- SentenceTransformers: For generating embeddings (we use
BAAI/bge-small-en-v1.5). - Redis: Optional, for caching frequent query→results pairs.
- Avoid: Over-engineering with complex rerankers early on. Start simple: normalize + add.
We tried using Weaviate’s hybrid search but switched to ES+pgvector because we needed fine-grained control over indexing policies and couldn’t justify the ops overhead of a new system.
When NOT to use hybrid search
- If your docs are mostly natural language with little jargon (e.g., public blog posts), pure vector may suffice.
- If latency is critical (<50ms) and you can’t afford the extra 30ms, test keyword-only first.
- If your team doesn’t have Elasticsearch expertise, consider pgvector-only with good chunking and embedding model tuning.
FAQ
Does hybrid search always beat vector-only?
No. On clean, well-written documentation with minimal acronyms, vector-only can match hybrid. Test on your data.
How do I tune the weight between vector and keyword?
Start with equal weight (normalize then add). If you have relevance labels, use a validation set to sweep weights from 0.0 to 1.0 for vector (1-weight for keyword).
Can I use hybrid search with LLMs that have long context?
Yes. Hybrid improves the quality of chunks fed into the LLM, which matters more than raw context length when dealing with noisy internal data.
What embedding model works best for hybrid?
We use BAAI/bge-small-en-v1.5. It’s small, fast, and works well with keyword fusion. Larger models help marginally but add latency.
Key Takeaways
- Hybrid search RAG over internal docs improves retrieval by combining semantic and exact-match strengths.
- Implement with Elasticsearch (keyword) and pgvector (vector) for control and production maturity.
- Fuse scores via normalization - don’t just add raw scores.
- Expect ~30ms latency increase; validate gains on your data with MRR@10 and recall@10.
- Route very short queries to keyword-only to avoid vector noise.
- Start simple: equal-weight fusion. Tune only if you have labeled data.
- Storage and write costs double - plan for it.
Hybrid Search RAG Implementation: A Practical Guide
RAG Chunking Best Practices for Production Systems
RAG Chunking Evaluation: Metrics, Trade-offs, and Production Lessons
Working on something similar?
If you're building backend or AI systems and want a second set of senior eyes, let's talk.