B Ben Moataz
Pillar
All guides Work with me Hybrid Search & RAG

Hybrid search and RAG retrieval that actually returns the right thing

How I design hybrid search and RAG retrieval systems — combining keyword and vector search, reranking, and evaluation so retrieval is precise on identifiers and strong on meaning.

8

in-depth guides in this pillar

2

supporting essays in the same cluster

2

capability lanes this work connects to

Why this pillar

Most retrieval systems in production are quietly worse than their demo. Pure vector search understands meaning but fumbles the exact identifiers — names, codes, versions — that real queries hinge on, and teams discover it only after shipping. Hybrid retrieval, fused and reranked, is the engineering response to that gap.

These guides are how I actually build it: combining lexical and semantic retrieval, fusing the rankings deliberately, reranking for real relevance, and — the part most teams skip — evaluating against the queries you actually get, so search is a system you improve on purpose rather than one you ship once and hope holds.

Guides

In-depth, code-backed guides.

Guide

Metadata Filtering in Vector Search: Pre-Filter, Post-Filter, and the Recall Cliff

How metadata filtering actually behaves in vector search: why post-filtering breaks tenant isolation, when selective filters collapse ANN recall, and what to do.

search
Read the guide →
Guide

Embedding Model Selection for Retrieval: How to Choose Without Trusting a Leaderboard

How I pick an embedding model for retrieval: the constraints that decide it before quality does, a bake-off you can run, and the re-index nobody prices.

search
Read the guide →
Guide

How to Evaluate RAG Retrieval: Eval Sets, Metrics, and Ship Gates

Retrieval eval is 20% metrics and 80% eval set. How I build one that survives re-indexing, which number to read at which k, and how to gate a change.

search
Read the guide →
Guide

RAG Chunking Strategy: How to Split Documents So Retrieval Works

Chunking decides what your retriever can find. Here's how I split real documents — structure-first boundaries, parent-child units, and how to prove it worked.

search
Read the guide →
Guide

Why Is My RAG Retrieval Bad? A Diagnostic Order of Operations

Bad RAG retrieval is four or five distinct failures wearing one costume. Here's how I localize which one you have before changing anything.

search
Read the guide →
Guide

Reranking in RAG: How a Cross-Encoder Fixes Retrieval Quality

A cross-encoder reranker re-scores your top candidates by reading query and document together. Here's how I add one, size it, and prove it worked.

search
Read the guide →
Guide

How to Build Hybrid Search with pgvector and BM25 in Postgres

Build hybrid search in Postgres with pgvector, tsvector, and RRF in one SQL query — the schema, index tuning, and when you actually need real BM25.

search
Read the guide →
Guide

Hybrid Search vs Vector Search: Why RAG Retrieval Needs Both

Vector search understands meaning but fumbles exact identifiers; keyword search is the opposite. Here's how I build hybrid retrieval that does both — with fusion, reranking, and a pgvector setup.

search
Read the guide →
Related essays

Field notes and opinionated takes in the same cluster.

Capabilities

The delivery lanes this work maps to.

Work with me

Retrieval returning the wrong things in production?

This is exactly the kind of system I get brought in to design, audit, and make dependable. If that's where you are, let's talk.

Subscribe

New guides in this pillar, by email

I publish an in-depth engineering guide most weeks. Drop your email and I'll send new ones as they land — no noise.

Subscribe via RSS → Email capture isn't wired up yet — the RSS feed is live now.