B Ben Moataz
← Back to writing search

Metadata Filtering in Vector Search: Pre-Filter, Post-Filter, and the Recall Cliff

How metadata filtering actually behaves in vector search: why post-filtering breaks tenant isolation, when selective filters collapse ANN recall, and what to do.

Professional headshot of Ben Moataz Ben Moataz · September 15, 2026 · 15 min read · Updated Sep 15, 2026

Metadata filtering is where a vector search stops being a similarity problem and starts being a correctness problem. The mechanics are simple — attach structured fields to each vector, constrain them at query time — but the order of operations decides everything: filter before the search and you get correct, complete results at the cost of index support; filter after it and a selective filter quietly returns three results instead of twenty, or worse, silently leaks another tenant’s documents into the candidate set before the filter runs.

This guide covers the part the vendor docs skip. Not “here is how to pass a where clause” — that’s an API detail you can read in five minutes — but which filters belong in the index versus the query, why an ANN index falls off a recall cliff exactly when your filter gets useful, and how to tell whether the thing you’re calling a ranking problem is actually a filter that deleted the answer.

The two-sentence version of pre- vs post-filtering

Post-filtering runs the approximate nearest-neighbour search over the whole index, gets k candidates back, then drops the ones that don’t match the filter. It works on any index, needs no special support, and is what you get by default in a lot of naive implementations. It’s also arithmetic you can’t win: if your filter matches 2% of the corpus and you retrieve 50 candidates, you should expect roughly one survivor.

Pre-filtering restricts the search to the matching subset first, then searches within it. You always get k results if k matching documents exist. The catch is that a graph-based index like HNSW isn’t a set you can subset for free — the edges were built over the whole corpus, so restricting the walk to matching nodes can disconnect the graph and strand the search in a region it can’t escape.

Modern engines mostly implement a third thing — filtered search, or in-filter traversal — where the filter is evaluated during the index walk rather than before or after it, sometimes with extra edges or a fallback to brute force when selectivity crosses a threshold. pgvector calls it iterative scan; Qdrant, Weaviate, Pinecone and the rest each have their own machinery. The names differ and the shape is the same: the engine is trying to avoid the two bad ends of the tradeoff by deciding per query.

The practical consequence for you: you need to know which of the three your store is doing, and you need to know what it does when the filter is very selective. That is the entire decision. Everything below is detail on top of it.

Correctness filters are not the same object as relevance filters

The single most useful distinction I draw in a retrieval design review is this one, because the two categories have different failure modes and different acceptable implementations.

A correctness filter defines what this user is allowed to see. Tenant id. Workspace. ACL group. Document classification. Deleted-at is null. If one of these fails open, you have a data leak, not a bad search result — and in a RAG system the leak is compounded, because the wrong document doesn’t just appear in a list, it gets read by a model and paraphrased into a confident answer with no citation trail the user will check.

A relevance filter narrows the search to be more useful. Date range. Document type. Language. Source system. Status is active. If one of these misbehaves, the user gets worse results and eventually complains.

Three rules follow from the split:

Correctness filters never post-filter. Not because post-filtering returns the wrong answer at the end — it does drop the non-matching rows — but because every layer between the ANN search and the filter is now handling data that user isn’t cleared for. If your candidate set gets logged, cached, traced, or fed into a reranker running on a third-party API, the isolation boundary was already crossed before the filter ran. Push tenancy as far down as it goes: a separate collection or namespace per tenant if the store supports it, a partitioned index if it doesn’t, a pre-filter at absolute minimum.

Correctness filters should be impossible to forget. Not a parameter the caller passes. A retrieval client that takes a user context object and constructs the filter itself, with no code path that reaches the index without it. I’ve seen the failure twice and both times the cause was the same: a new endpoint, written in a hurry, that called the vector store directly because the wrapper was inconvenient.

Relevance filters can be soft. A date range that returns nothing is a chance to widen and tell the user, not an empty page. I’ll come back to this in the fallback section, because “no results” is usually a filter bug, not a corpus gap.

This is the failure mode that costs teams the most time, because it produces no error and no obviously wrong output. The search just gets quietly worse, in a way that correlates with how narrow the filter is — which means it’s worst exactly for your most specific, highest-intent queries.

The mechanism, for a graph index like HNSW: the index is a navigable small-world graph whose edges were chosen to make greedy traversal converge on nearest neighbours across the full corpus. Restrict traversal to nodes matching a filter and you’re walking a subgraph that was never designed to be connected. Below some selectivity, the matching nodes are sparse enough that the walk hits a local region with no matching neighbours to expand into, terminates early, and reports the results it found. Those results are real vectors that really match your filter — they’re just not the nearest ones. Recall degrades; latency may even improve, because the search gave up sooner, which is how this hides on a dashboard.

Post-filtering has a different but equally silent version: you asked for 20, you got 4, and unless somebody checks result counts against k nobody notices that the answer was in position 63 of the unfiltered ranking.

The thresholds are engine-specific and move between versions, so I won’t quote numbers as though they’re universal. What’s stable is the shape:

  • Broad filters (say, more than a few percent of the corpus matching) are mostly fine on any strategy. A filter that keeps half your documents costs you almost nothing.
  • Narrow filters are where strategies diverge sharply, and where post-filtering stops being viable at all.
  • Very narrow filters — a single tenant out of ten thousand, one document id, one conversation thread — are often best served by not using the ANN index at all. If the filter leaves a few thousand vectors, an exact brute-force scan over those is both faster and perfectly accurate. Several engines will make this switch for you; if yours won’t, make it yourself.

That last point is the one I most want people to take away. The fastest filtered vector search is frequently the one that skips the vector index. Teams resist it because the index feels like the point of the system, but exact search over 5,000 vectors is a few milliseconds of tight loop, and it has recall of 1.0 by construction.

Measure it before you argue about it

You don’t have to reason about your engine’s thresholds from first principles. You can measure the cliff in about twenty minutes, and the measurement is the same one regardless of store: compare filtered ANN results against exact filtered results on the same queries.

import numpy as np

def exact_filtered_topk(query_vec, vectors, ids, matches_filter, k):
    """Ground truth: brute-force cosine over the filtered subset."""
    keep = np.array([matches_filter(i) for i in ids])
    subset, subset_ids = vectors[keep], np.asarray(ids)[keep]
    if len(subset_ids) == 0:
        return []
    # vectors assumed L2-normalized, so dot product == cosine similarity
    scores = subset @ query_vec
    top = np.argsort(-scores)[:k]
    return list(subset_ids[top])


def filter_recall_curve(queries, store, vectors, ids, filters, k=20):
    """For each filter, report selectivity, recall@k, and how many rows came back."""
    rows = []
    for name, (predicate, store_filter) in filters.items():
        selectivity = sum(predicate(i) for i in ids) / len(ids)
        recalls, returned = [], []
        for q in queries:
            truth = set(exact_filtered_topk(q.vector, vectors, ids, predicate, k))
            if not truth:
                continue
            got = store.search(q.vector, k=k, filter=store_filter)
            got_ids = [r.id for r in got]
            returned.append(len(got_ids))
            recalls.append(len(truth & set(got_ids)) / len(truth))
        rows.append({
            "filter": name,
            "selectivity": round(selectivity, 4),
            "recall@k": round(float(np.mean(recalls)), 3),
            "avg_returned": round(float(np.mean(returned)), 1),
            "k": k,
        })
    return rows

Run that across filters spanning three or four orders of magnitude of selectivity — everything, one language, one month, one customer — and you get a table that ends the debate. Two columns matter. recall@k falling as selectivity drops is the cliff. avg_returned falling below k is post-filtering starving, and it is a separate bug from the cliff even though they co-occur.

Keep this table. Re-run it when you upgrade the store, when the corpus doubles, and when someone adds a new filterable field — the thresholds move with all three. It slots directly into the harness from how I evaluate RAG retrieval, as a per-filter breakdown rather than one aggregate number.

Design the metadata at index time, not at query time

Almost every painful filtering problem I’ve been called into was created months earlier, at ingestion, by treating metadata as a junk drawer. Some discipline that pays for itself:

Filterable fields are a schema, not a blob. Decide the small set of fields you will actually filter on, type them, and index them. source, tenant_id, doc_type, language, created_at, status covers most systems. Everything else — the original file path, the extraction confidence, the ingest job id — is payload you carry for display and debugging, not a search dimension. The distinction matters because filterable fields need an index and a cardinality budget, and payload doesn’t.

Normalize at write time. EN, en, en-US, and English in the same field is a filter that silently misses a quarter of the corpus, and it will be blamed on the embedding model. Same for timezone-naive timestamps, trailing whitespace in source names, and enum values that drifted when a second ingestion path was added. Validate on ingest and reject rather than coerce quietly.

Denormalize what you filter on. If the filter is “documents belonging to projects this user can access,” and access lives in another service, you cannot evaluate that inside the vector store. Either materialize the resolved access list onto the vector’s metadata at index time and accept the update cost, or resolve the allowed set first and pass it as an explicit id filter. What you must not do is post-filter it in application code — see the correctness section.

Watch the cardinality of the field you filter hardest on. A field with three values partitions the corpus into three big chunks and is cheap on any index. A field with a million values — one per user — is a different data structure problem entirely, and usually means you want per-tenant namespaces rather than a filter.

Denormalize time into buckets if you filter on ranges constantly. created_at >= X on a timestamp is fine, but if 90% of your queries are “last 30 days,” an indexed recency_bucket field is a cheaper, more selective predicate that the engine can use without a range scan.

The chunk-level consequence is worth stating: filters apply to whatever you indexed, so if you index chunks, every chunk needs its parent document’s metadata copied onto it. That’s more duplication than people expect, and it’s why the metadata contract belongs in the same design conversation as your chunking strategy rather than being bolted on after.

Push the filter into both arms of a hybrid retriever

If you run hybrid search, the filter has to go into the lexical arm and the semantic arm, before fusion. This sounds obvious and is routinely gotten wrong, because the two arms are often written by different people at different times.

Filtering after fusion wastes candidate budget on rows you’re about to discard: each arm returns 50, you fuse 100, the filter keeps 12, and you’re reranking a set that was effectively retrieved at depth 6. Worse, the two arms starve unevenly — a filter correlated with document type will gut the lexical arm while barely touching the semantic one, and your carefully tuned fusion weights now describe a system you’re no longer running.

The pgvector and BM25 build shows the shape in SQL: the predicate lives inside both CTEs, so each arm independently returns 50 matching candidates and fusion sees a full candidate set. The same rule holds with a dedicated search cluster and a separate vector store — pass the filter to both, at the same depth, and verify the depths are actually equal rather than assuming.

There’s a reason this matters more in hybrid than in pure dense retrieval, and it’s the argument in hybrid search vs vector search: the lexical arm is usually the one carrying exact identifiers, codes and names, which are exactly the queries a user pairs with a narrow filter. Starve that arm with a badly-placed predicate and you lose the queries the filter existed to serve.

Empty results are a design decision, not an error

When a filtered search returns nothing, you have three options, and the default — showing an empty page — is almost always the worst one.

Widen and say so. Drop the least important relevance filter and re-run, labelling the results: “No results in the last 30 days. Showing results from the last year.” This is nearly always right for date ranges, which users pick arbitrarily.

Fall back to unfiltered and mark it. Useful when the filter was inferred rather than chosen — if a query router guessed “the user probably means invoices” and that produced nothing, the guess was wrong and should not survive contact with an empty result set.

Return empty honestly. Correct when the filter was explicit and the absence is itself the answer: no documents for this customer, no incidents this quarter.

Never widen a correctness filter. If tenant scoping produces zero results, zero is the answer.

The instrumentation that makes this tractable: log the filter alongside the result count on every search, and alert on the rate of zero-result filtered queries. A jump in that rate is one of the earliest signals that an ingestion path started writing a metadata value in a new shape — usually days before anyone reports bad search quality.

What this looks like when it goes wrong

Three patterns, all of which arrive described as “our retrieval is bad”:

The filter deleted the answer. The document exists, the embedding is fine, ranking is fine, and a doc_type value changed in an upstream system six weeks ago so the predicate no longer matches it. The tell: retrieve the same query with no filter and the gold document comes back at rank 2. This is the first check in the diagnostic order I walk through in why is my RAG retrieval bad, and it’s first because it’s cheap and it’s common.

The filter starved the candidate set. Recall looks acceptable in aggregate and terrible on the subset of traffic that uses filters. The tell: avg_returned below k, or a recall number that splits sharply when you group queries by whether a filter was applied. Aggregate metrics hide this completely, which is why the filter breakdown belongs in your eval harness rather than in an ad-hoc notebook.

The filter was correct and the reranker never saw enough. Filtered retrieval returned 20 good candidates, the reranker was configured for a 50-candidate input, and the top-5 it produces are the best of a thin set. The tell is a reranker that stopped helping on filtered queries specifically. The fix is to raise retrieval depth when a filter is present, not to change the reranker.

When you don’t need this

Plenty of systems should not build any of this. If your corpus is small enough to scan exactly — and “small enough” is larger than most people assume, comfortably into the tens of thousands of vectors for a single-user tool — skip the ANN index and its filtering machinery entirely. Exact search with a plain predicate has perfect recall, zero tuning, and no cliff.

If you’re single-tenant with one document type and no time dimension, you have no correctness filters and one or two relevance filters, and the default behaviour of any vector store will serve you fine. Adding a filter schema before you have queries that need it is how you end up maintaining six indexed fields where two would do.

And if you’re filtering primarily to compensate for bad ranking — narrowing to a document type because the right documents keep losing to the wrong ones — the filter is a symptom. Fix the ranking. A filter that exists to hide a retrieval problem will eventually hide the answer too.

If you’re working through retrieval quality more broadly, the hybrid search and RAG guides cover the stages around this one, and correlation scoring is where filtering meets entity-level matching in the systems I build.

FAQ

Is pre-filtering always better than post-filtering? No — it’s better whenever the filter is selective, which is most of the time, but it isn’t free. Pre-filtering needs index support and can cost latency on broad filters where post-filtering is essentially free. The honest default is: use your engine’s native filtered search, and post-filter only for predicates that match nearly everything and that can’t leak anything.

How selective does a filter have to be before recall degrades? Engine- and version-specific, and it moves with corpus size and index parameters, so measure rather than trusting a number from a blog post — including this one. The measurement takes twenty minutes with the recall-curve code above, and the answer is worth more than any published threshold because it’s your index.

Should tenant isolation be a filter or separate collections? Separate collections or namespaces when the tenant count is manageable and tenants are large — it makes leaks structurally impossible and keeps each index small. A filter when you have many small tenants and per-collection overhead would dominate. If you use a filter, it must be a pre-filter, and it must be injected by shared code rather than passed by each caller.

Does metadata filtering slow down vector search? Sometimes it speeds it up — a narrow filter can turn an index walk into a small exact scan. What costs you is the middle ground: a filter selective enough to fight the graph but broad enough that brute force is still expensive. That regime is where latency and recall both suffer, and where it’s worth checking whether your engine exposes a threshold you can tune.

Can I filter on a field I didn’t index? Only in application code, after results come back, which means post-filtering with all its problems — and if it’s a correctness filter, you’ve moved the isolation boundary into your app. Adding a filterable field usually means re-indexing, so it’s worth deciding the filter schema before the first large ingest rather than after.

Professional headshot of Ben Moataz
Written by
Ben Moataz

Systems Architect, Consultant, and Product Builder

This article is grounded in hands-on work across Correlation and scoring, including systems such as SOVRINT, TraxinteL, and Viralink.

I write from hands-on work across product systems, evidence pipelines, ranking layers, monitoring surfaces, and automation runtimes that have to stay reliable under operational pressure.

  • → Years spent building product systems, automation infrastructure, and operator-facing platforms.
  • → Project records and case studies tied directly to the same capability lanes discussed in the writing.
  • → A public archive designed to connect essays back to real systems, delivery constraints, and consulting work.
Relevant work

Expertise and case studies tied to this article.

Related reading

More writing on adjacent systems problems.

Next article

Embedding Model Selection for Retrieval: How to Choose Without Trusting a Leaderboard

How I pick an embedding model for retrieval: the constraints that decide it before quality does, a bake-off you can run, and the re-index nobody prices.

Work with me

Building or fixing a system like this?

This is exactly the kind of work I get brought in for. Teams unsure whether a system, architecture, or workflow will hold up under real load and scrutiny.

System Audit Start here · fixed scope
  • → A focused review of the system, architecture, or codebase in question.
  • → A clear map of the risks, bottlenecks, and failure modes that matter.
  • → A prioritized roadmap — what to fix first, and what to leave alone.
Subscribe

Get new essays by email

Field notes on intelligence systems, evidence engineering, and automation that survives reality. No noise.

Subscribe via RSS → Email capture isn't wired up yet — the RSS feed is live now.