B Ben Moataz
Guides
Hybrid search & RAGReliable systems

Deep engineering guides on the systems I actually build.

Two pillars: hybrid search and RAG retrieval, and reliable queues, workers, and monitoring. Long-form, code-backed, and opinionated — the real tradeoffs, not a definitions dump.

2

engineering pillars

12

in-depth guides published so far

Code-backed

practitioner writing, not vendor marketing

Pillars

Pick the cluster you're working in.

Latest guides

Everything published across both pillars.

Guide

Metadata Filtering in Vector Search: Pre-Filter, Post-Filter, and the Recall Cliff

How metadata filtering actually behaves in vector search: why post-filtering breaks tenant isolation, when selective filters collapse ANN recall, and what to do.

search
Read the guide →
Guide

Embedding Model Selection for Retrieval: How to Choose Without Trusting a Leaderboard

How I pick an embedding model for retrieval: the constraints that decide it before quality does, a bake-off you can run, and the re-index nobody prices.

search
Read the guide →
Guide

Worker Fleet Architecture at Scale: Pools, Scaling Signals, and Draining

Scaling a worker fleet isn't adding workers to one queue. Here's how I partition pools, pick the scaling signal, and drain workers without losing jobs.

operationsautomation
Read the guide →
Guide

How to Evaluate RAG Retrieval: Eval Sets, Metrics, and Ship Gates

Retrieval eval is 20% metrics and 80% eval set. How I build one that survives re-indexing, which number to read at which k, and how to gate a change.

search
Read the guide →
Guide

RAG Chunking Strategy: How to Split Documents So Retrieval Works

Chunking decides what your retriever can find. Here's how I split real documents — structure-first boundaries, parent-child units, and how to prove it worked.

search
Read the guide →
Guide

Why Is My RAG Retrieval Bad? A Diagnostic Order of Operations

Bad RAG retrieval is four or five distinct failures wearing one costume. Here's how I localize which one you have before changing anything.

search
Read the guide →
Guide

Retry with Exponential Backoff and Jitter (and the Retry Budget Nobody Sets)

Exponential backoff caps how fast a client retries; jitter stops every client retrying together. Here's the retry logic I actually ship, and the two limits most teams miss.

operations
Read the guide →
Guide

Reranking in RAG: How a Cross-Encoder Fixes Retrieval Quality

A cross-encoder reranker re-scores your top candidates by reading query and document together. Here's how I add one, size it, and prove it worked.

search
Read the guide →
Guide

Dead Letter Queue Design Patterns (Routing, Envelopes, and Redrive)

A DLQ is the giving-up mechanism, and most teams build it wrong. The routing, envelope, isolation, and redrive patterns I use to make failed messages recoverable.

operations
Read the guide →
Guide

How to Build Hybrid Search with pgvector and BM25 in Postgres

Build hybrid search in Postgres with pgvector, tsvector, and RRF in one SQL query — the schema, index tuning, and when you actually need real BM25.

search
Read the guide →
Guide

Hybrid Search vs Vector Search: Why RAG Retrieval Needs Both

Vector search understands meaning but fumbles exact identifiers; keyword search is the opposite. Here's how I build hybrid retrieval that does both — with fusion, reranking, and a pgvector setup.

search
Read the guide →
Guide

How to Design an Idempotent Job Queue (Retries, Backoff, and Dead Letters)

At-least-once delivery makes idempotency mandatory, not optional. Here's how I design job queues that retry safely, back off with jitter, and dead-letter poison messages — with code.

operations
Read the guide →
Subscribe

Get new engineering guides by email

In-depth, code-backed guides on hybrid search, RAG, queues, and reliability — sent as they publish. No noise.

Subscribe via RSS → Email capture isn't wired up yet — the RSS feed is live now.