Exclusive Discount Deal
Upto 50% OFF
Offer ends in:
23 DAYS
|
21 HOURS
|
28 MINS
|
56 SECS
Home / Blog / Architecting Production RAG Pipelines with Hybrid Re-Ranking
AI & Emerging Tech • Oct 8, 2026

Architecting Production RAG Pipelines with Hybrid Re-Ranking

A deep dive into building enterprise-grade Retrieval-Augmented Generation pipelines combining BM25 sparse keyword search, dense vector embeddings, and Cohere re-ranking algorithms.

UPTO 50% OFF
Trending:
BrickTry

Requirement Scope

AI is analyzing your requirement...

Generating custom modules, implementation options, and dynamic clarification questions.

Add Custom Requirement or Module

Add your own specific features, integrations, or components. AI will incorporate them to dynamically generate the next relevant options.

1. Progressive Clarifications

Click to expand & answer

2. Scope Modules & Features (/ Selected)

Click row to expand details · Customize options
✓
✕
Completeness:

Retrieval-Augmented Generation (RAG) systems fail in production not because large language models lack reasoning capacity, but because the retrieval layer fails to surface relevant context. Standard semantic search using dense vector embeddings excels at capturing conceptual similarity, yet it frequently misses exact keyword matches, SKU numbers, legal citations, and domain-specific acronyms. Conversely, classical sparse lexical search (BM25) captures exact tokens brilliantly while remaining entirely blind to semantic nuance.

Building an enterprise-grade RAG pipeline requires merging these two paradigms into a hybrid retrieval engine, followed by a neural cross-encoder re-ranking stage to filter out false positives before context hits the LLM context window.


The Hybrid RAG Topology

A production-grade retrieval pipeline demands a multi-stage architecture to balance low latency with high precision.

[ User Query ]
       │
       ├─────────────────────────────────┐
       ▼                                 ▼
[ Sparse Search: BM25 ]        [ Dense Search: HNSW Vector ]
       │                                 │
       └──────────────┬──────────────────┘
                      ▼
       [ Reciprocal Rank Fusion (RRF) ]
                      │
                      ▼
       [ Cross-Encoder Re-Ranking (Cohere / BGE) ]
                      │
                      ▼
       [ Top-K Context Window -> LLM ]

1. Dual-Path Retrieval

Incoming queries bypass naive vector searches. Instead, the query fans out concurrently:

  • Sparse Path: Inverted index matching exact term frequencies (BM25) over tokenized enterprise documents.
  • Dense Path: Approximate Nearest Neighbor (ANN) vector search (e.g., pgvector with HNSW index) mapping embeddings generated via models like text-embedding-3-large or open-source equivalents.

2. Reciprocal Rank Fusion (RRF)

Merging two independent scoring distributions (cosine similarity and BM25 score) requires normalizing disparate scales. Reciprocal Rank Fusion avoids normalization pitfalls by operating purely on the ordinal ranks assigned by each retriever:

$$RRF_Score(d \in D) = \sum_{m \in M} \frac{1}{k + r_m(d)}$$

Where $M$ represents the set of retrieval methods, $r_m(d)$ is the rank of document $d$ in method $m$, and $k$ is a constant smoothing factor (typically set to 60).

3. Neural Re-Ranking

The top 50 candidates returned by RRF undergo cross-encoder re-ranking. Bi-encoders embed queries and documents independently; cross-encoders process the query and document simultaneously through self-attention layers, capturing deep semantic interactions at the cost of higher compute latency. Models like cohere-rerank-v3.5 or bge-reranker-large compress the candidate pool down to the top 4–6 highly relevant chunks.


Architectural Configuration Trade-Offs

Component Strategy A (Optimized for Speed) Strategy B (Optimized for Precision) Production Recommendation
Vector Index IVF-Flat (Inverted File) HNSW (Hierarchical Navigable Small World) HNSW (m=16, ef_construction=64) for sub-50ms recall >98%.
Sparse Engine PostgreSQL Full-Text Search (tsvector) Dedicated OpenSearch / Elasticsearch Cluster pgvector + BM25 Extension for unified transactional and search storage.
Re-Ranking Single-stage Bi-Encoder cosine cut-off Two-stage Cross-Encoder (Top 50 $\rightarrow$ Top 5) Two-stage Cross-Encoder; latency overhead (~30ms) is dwarfed by token savings.
Chunking Strategy Fixed-size (512 tokens with 50 overlap) Semantic boundary-aware chunking (Markdown/AST) Semantic/Markdown Splitting to prevent context fragmentation.

Implementation: Python Hybrid Retrieval & RRF Pipeline

The following production-ready Python snippet implements concurrent hybrid retrieval via PostgreSQL (pgvector and tsvector) followed by Reciprocal Rank Fusion.

import os
from typing import List, Dict, Any
import psycopg2
from psycopg2.extras import RealDictCursor
from openai import OpenAI

client = OpenAI(api_key=os.environ.get("OPENAI_API_KEY"))

def get_embedding(text: str) -> List[float]:
    response = client.embeddings.create(
        input=[text],
        model="text-embedding-3-small"
    )
    return response.data[0].embedding

def hybrid_search(query_text: str, limit: int = 20) -> List[Dict[str, Any]]:
    query_vector = get_embedding(query_text)

    conn = psycopg2.connect(os.environ.get("DATABASE_URL"))
    try:
        with conn.cursor(cursor_factory=RealDictCursor) as cursor:
            # Execute parallel sparse and dense queries
            sql = """
            WITH sparse AS (
                SELECT id, content, metadata,
                       ts_rank_cd(to_tsvector('english', content), plainto_tsquery('english', %s)) as score
                FROM documents
                WHERE to_tsvector('english', content) @@ plainto_tsquery('english', %s)
                ORDER BY score DESC
                LIMIT %s
            ),
            dense AS (
                SELECT id, content, metadata,
                       1 - (embedding <=> %s::vector) as score
                FROM documents
                ORDER BY embedding <=> %s::vector
                LIMIT %s
            ),
            rrf AS (
                SELECT
                    COALESCE(s.id, d.id) as id,
                    COALESCE(s.content, d.content) as content,
                    COALESCE(s.metadata, d.metadata) as metadata,
                    COALESCE(1.0 / (60 + s.row_num), 0.0) +
                    COALESCE(1.0 / (60 + d.row_num), 0.0) as rrf_score
                FROM (SELECT id, content, metadata, ROW_NUMBER() OVER (ORDER BY score DESC) as row_num FROM sparse) s
                FULL OUTER JOIN (SELECT id, content, metadata, ROW_NUMBER() OVER (ORDER BY score DESC) as row_num FROM dense) d
                ON s.id = d.id
            )
            SELECT id, content, metadata, rrf_score
            FROM rrf
            ORDER BY rrf_score DESC
            LIMIT %s;
            """
            cursor.execute(sql, (query_text, query_text, limit, query_vector, query_vector, limit, limit))
            return cursor.fetchall()
    finally:
        conn.close()

TypeScript Integration & Re-Ranking

Once RRF yields a narrowed candidate array, apply a cross-ranking model before passing payloads to your LLM generator. Here is a TypeScript service layer handling API orchestration for re-ranking.

import { CohereClient } from "cohere-ai";

const cohere = new CohereClient({
  token: process.env.COHERE_API_KEY,
});

export interface DocumentCandidate {
  id: string;
  content: string;
  metadata: Record<string, any>;
  rrfScore: number;
}

export async function rerankDocuments(
  query: string,
  candidates: DocumentCandidate[],
  topN: number = 4
): Promise<DocumentCandidate[]> {
  if (candidates.length === 0) return [];

  const documents = candidates.map((c) => c.content);

  const response = await cohere.v2.rerank({
    model: "rerank-v3.5",
    query: query,
    documents: documents,
    topN: topN,
    returnDocuments: false,
  });

  // Map re-ordered indices back to original candidate objects
  return response.results.map((result) => {
    const original = candidates[result.index];
    return {
      ...original,
      rrfScore: result.relevanceScore, // Overwrite with cross-encoder score
    };
  });
}

How BrickTry Accelerates & Powers This

Architecting, benchmarking, and maintaining hybrid RAG pipelines involves significant infrastructure overhead: configuring Postgres pgvector extensions, tuning HNSW parameters, synchronizing embedding chunk pipelines, and securing API keys. BrickTry streamlines this workflow from prototype to production:

  • BrickTry Lab Sandbox (/lab): Spin up instant, zero-setup in-browser Node.js and Python container runtimes to prototype vector embeddings and test RRF algorithms live without local environment configuration friction.
  • AI-Human Dev Pairing: Autonomous AI scaffolding instantly generates your initial database migration scripts (pgvector), vector chunking utilities, and API wrappers, while dedicated senior full-stack engineering pods review your query latency, indexing strategies, and context-window memory usage.
  • Interactive Scoping Engine: Translates raw product specifications into modular system milestones, breaking down document ingestion pipelines, chunking strategies, and caching layers into manageable deployment tasks.
  • Unified Importer: Seamlessly ingest existing repositories or legacy monolithic search scripts from GitHub or CodeCanyon into BrickTry's modern TypeScript/Python architecture with 1-click refactoring.
  • 100% Source Code Ownership: Retain complete ownership of your GitHub repositories, Docker orchestration files, and database schemas with zero vendor lock-in, ensuring enterprise-grade compliance and portability.

Build, Test, and Scale This on BrickTry

BrickTry pairs you with autonomous AI scaffolding supervised by dedicated senior full-stack software engineers in an interactive in-browser development sandbox. Test, build, and deploy production-grade software with 100% source code ownership and zero vendor lock-in.

Launch Interactive Requirement Builder →

❤️

Support BrickTry Platform & Engineering Development

Help us build, maintain, and advance our AI engineering platform. Every donation fuels open-source tooling, infrastructure, and continuous improvements.

$
Donor Details
Promote Your Brand / Link Wall

UPI / Credit & Debit Cards / Netbanking
Razorpay
Secure 256-bit encrypted checkout
View Leaderboard & Wall

Hey!

Welcome, Let's chat —
start a new conversation
below.

Recent conversations
See all

Hi ,We’d like to inform you that the Integ...

Abhishek A Agrawal • 1d ago

Abhishek A Agrawal

Back in a few hours