B Ben Moataz
← Back to writing search

Embedding Model Selection for Retrieval: How to Choose Without Trusting a Leaderboard

How I pick an embedding model for retrieval: the constraints that decide it before quality does, a bake-off you can run, and the re-index nobody prices.

Professional headshot of Ben Moataz Ben Moataz · September 8, 2026 · 13 min read · Updated Sep 08, 2026

Choose an embedding model by ranking your hard constraints first — maximum sequence length against your chunk size, where the data is allowed to go, the latency you can afford at query time, and the storage and memory cost of the vector dimension — and only then run a bake-off between the two or three survivors on your own labelled queries. A public leaderboard tells you how a model ranks on somebody else’s corpus with somebody else’s query distribution, which is a weak predictor of how it will rank on yours. And because switching models means re-embedding everything, this is one of the few retrieval decisions you can’t cheaply undo, so it’s worth an afternoon of measurement instead of an hour of reading comparison posts.

Why the leaderboard can’t make this call for you

Every “best embedding models for RAG” post is a leaderboard reading with a paragraph of commentary per row. The leaderboard is a real, useful artifact — it just doesn’t answer the question people bring to it.

Three reasons it under-delivers:

It averages over tasks you don’t have. A general embedding benchmark scores classification, clustering, semantic similarity, summarization and retrieval, and the headline number is an average across all of it. You care about one column. A model that wins overall can sit mid-table on retrieval, and a model tuned for symmetric similarity (is sentence A like sentence B?) can be mediocre at the asymmetric job you actually run (does this short question match this long passage?).

Its corpora aren’t yours. Benchmark retrieval sets skew toward web text, Wikipedia, and academic abstracts, with fluent natural-language queries. If your corpus is support macros, insurance policies, court filings, German-language service manuals or a codebase, and your queries are six-word fragments typed by someone in a hurry, you’re extrapolating from a different distribution in both directions at once. Relative model ranking is what’s fragile here — the gap between two models can shrink, vanish, or reverse when the domain changes.

Popular benchmarks leak into training. Once a benchmark is the scoreboard everyone optimizes against, some of its signal converts into overfitting, and some fraction of test data ends up in training corpora scraped from the open web. This doesn’t make scores fake; it makes small differences between adjacent models less meaningful than they look.

My working rule: use the leaderboard to build a shortlist of three, never to pick one. Retrieval column only, filtered by the constraints below before you even read the scores.

Rank the constraints before you rank the models

Most model choices are decided by constraints, not quality, and it saves an enormous amount of time to apply them first. Four that actually eliminate candidates:

1. Maximum sequence length versus your chunk size. A model with a 512-token limit silently truncates anything longer. It doesn’t error — it embeds the first part of your chunk and throws the rest away, which produces a vector that confidently represents the wrong thing. If you’re already chunking to ~300 tokens, a 512-token model is fine and the 8k-context models buy you nothing. If you’re indexing whole sections or using a small-to-big scheme, check the number. This interacts directly with your chunking strategy, and it’s the constraint I most often find quietly violated in a pipeline someone inherited.

2. Where the data is allowed to go. If the corpus contains PHI, client-confidential material, or anything covered by a data-residency clause, a hosted API may be off the table entirely, or may need a specific region and a signed agreement. This kills more shortlists than latency does. Decide it before you fall in love with a model you can’t legally call.

3. Query-time latency, not indexing latency. Indexing is a batch job; if it takes six hours, you run it overnight. The embedding call in the request path is what your users feel, and it’s on the critical path ahead of vector search, reranking, and generation. A hosted API adds a network round trip on every query. A large self-hosted model without a GPU can be slower than that round trip. If you’re already spending most of your latency budget on a cross-encoder reranker, the embedding step’s share matters more than the model’s benchmark score.

4. Dimension, because dimension is a bill. Vector dimension drives index memory, storage, and search speed roughly linearly. Going from 1024 to 3072 dimensions triples the size of your HNSW graph’s vector payload for the same corpus — which for a large index is the difference between “fits in RAM” and “doesn’t,” and that cliff costs more retrieval quality than the model upgrade gained. Several modern models are trained so their vectors can be truncated to a shorter prefix with graceful degradation, which lets you buy back dimension cheaply, but you have to measure the truncated version rather than assume it.

Only after these four do I look at quality. Usually two or three candidates are left, and they’re often closer to each other than the marketing suggests.

The asymmetry nobody prices: switching means re-indexing

Here’s the structural fact that makes this decision different from every other retrieval knob.

Vectors from two different models are not comparable. There is no migration, no partial rollout, no gradual cutover within a single index — a query embedded with model B and compared against documents embedded with model A produces garbage that looks like plausible search results. So changing the embedding model means embedding your entire corpus again, rebuilding the index, and cutting over.

That has three practical consequences:

  • You can’t A/B test it in production the way you’d A/B a reranker or a top_k value. You need a second full index to compare against, which for a large corpus is a real cost in embedding spend and infrastructure. Most teams can afford this once on a sample, not repeatedly on production scale.
  • The cost is recurring in a way people forget. Your first re-index is the visible cost. Then someone changes the chunking, and you pay again. Budget embedding cost as “corpus size × re-embed frequency,” not as a one-time line item.
  • It should therefore be an early, deliberate decision. This is the argument for spending a day on it up front. It’s also the argument for not chasing every new model release: the honest question isn’t “is this model better?” but “is it enough better on my queries to justify a full re-index and the cutover risk?”

When I do switch, the shape is always the same: build the new index alongside the old one, run the same eval set against both, compare per-query rather than in aggregate, and keep the old index until the new one has served live traffic. The cutover is a flag on which index the retriever reads.

Run a bake-off on your own queries

The measurement is much less work than people expect, and it’s the only part of this that produces a defensible answer. You need a labelled eval set — real queries paired with the documents that answer them — which is the same asset you should already have for measuring retrieval quality. Label documents, not chunks, precisely so the labels survive this experiment: a model swap re-chunks nothing, but a chunking change later will invalidate chunk-level labels.

Then embed a corpus sample with each candidate and read recall at your candidate depth — the depth you retrieve to before reranking. That’s the number the embedding model is actually responsible for. Precision at the top of the list is largely the reranker’s job.

"""Compare embedding models on your own labelled queries.

eval_set: [{"query": str, "relevant_doc_ids": [str, ...]}, ...]
corpus:   [{"doc_id": str, "chunk_id": str, "text": str}, ...]

`embed_fn(texts, kind)` is whatever the candidate model needs. `kind` is
"query" or "document" -- several model families require different prefixes
or instructions for each side, and getting that wrong is the single most
common reason a strong model tests as mediocre.
"""

import numpy as np


def build_index(corpus, embed_fn, batch_size=64):
    vectors = []
    for start in range(0, len(corpus), batch_size):
        batch = corpus[start : start + batch_size]
        vectors.extend(embed_fn([c["text"] for c in batch], kind="document"))
    matrix = np.asarray(vectors, dtype=np.float32)
    # Normalize once so a dot product is cosine similarity.
    matrix /= np.linalg.norm(matrix, axis=1, keepdims=True) + 1e-12
    return matrix


def recall_at_k(eval_set, corpus, matrix, embed_fn, k=50):
    doc_ids = np.array([c["doc_id"] for c in corpus])
    queries = [row["query"] for row in eval_set]
    q = np.asarray(embed_fn(queries, kind="query"), dtype=np.float32)
    q /= np.linalg.norm(q, axis=1, keepdims=True) + 1e-12

    scores = q @ matrix.T                      # (n_queries, n_chunks)
    top = np.argpartition(-scores, kth=k - 1, axis=1)[:, :k]

    per_query = []
    for i, row in enumerate(eval_set):
        # Rank the k candidates properly; argpartition does not order them.
        ranked = top[i][np.argsort(-scores[i, top[i]])]
        retrieved_docs = set(doc_ids[ranked].tolist())
        relevant = set(row["relevant_doc_ids"])
        hit = len(retrieved_docs & relevant) / max(len(relevant), 1)
        per_query.append({"query": row["query"], "recall": hit})

    mean = sum(r["recall"] for r in per_query) / len(per_query)
    misses = [r["query"] for r in per_query if r["recall"] == 0.0]
    return {"recall_at_k": mean, "k": k, "zero_recall": misses}


# for name, embed_fn in candidates.items():
#     matrix = build_index(corpus, embed_fn)
#     print(name, recall_at_k(eval_set, corpus, matrix, embed_fn, k=50))

Read the output in this order:

  1. The zero-recall list, not the mean. Two models with the same average can fail on completely different queries. If model A misses your identifier lookups and model B misses your conceptual questions, that’s a decision, and the mean hides it.
  2. The per-query differences. If the winner is ahead by a hair and the queries it wins are noise, you’ve measured nothing — treat it as a tie and pick on cost, latency, or hosting.
  3. The mean, last. It’s a summary, not evidence.

Brute-force cosine over a sample is deliberate here: you’re measuring the model, so you don’t want your ANN index’s recall loss confounding the comparison. Ten to twenty thousand chunks is plenty and runs in seconds on a laptop.

The prefix footgun that fakes a bad model

If you take one operational detail from this guide, take this one.

Retrieval is asymmetric: a short question has to match a long passage, and those are different kinds of text. Several open model families encode that asymmetry explicitly by requiring you to prefix or instruct each side — a query prefix for questions and a passage prefix for documents, or a task instruction attached only to the query. Some families want an instruction on the query and nothing on the document. Some want nothing at all, and adding a prefix hurts.

Every one of these fails silently. You get vectors, the index builds, search returns results, and quality is quietly worse than the model can do. I have seen a model dismissed as “not better than what we have” three separate times when the actual finding was that it had been called with the wrong convention — usually because the evaluation harness embedded queries and documents through the same code path.

So: read the model card, use its exact strings, and make kind="query" versus kind="document" an explicit argument in your embedding wrapper rather than an implicit assumption. If a model tests surprisingly badly, check this before you conclude anything.

Where the embedding model matters less than you think

There’s a reason “swap the embedding model” is the most common first move and rarely the most valuable one. It’s a single dial that changes everything at once, at maximum cost, and it only fixes one class of failure.

If your retriever is already hybrid — lexical plus semantic, fused — the lexical arm covers exactly the queries embeddings are worst at: SKUs, error codes, ticket numbers, surnames, rare tokens. A better embedding model improves the semantic arm’s contribution to a list that already had those queries handled. And if you rerank the fused candidates, a cross-encoder re-scores the top of the list with far more signal than any bi-encoder has, which compresses the difference between two decent embedding models further.

The stack matters more than the model. A mid-tier embedding model inside hybrid retrieval with a reranker will beat a leaderboard-topping model doing pure dense search on most real corpora, and it costs less to run. Before you spend the re-index, confirm you’re not misdiagnosing which stage is actually failing — the answer being absent from the index, split across a chunk boundary, or removed by a metadata filter all look identical to “we need a better model” from the outside.

When a domain-specific or fine-tuned model earns its place

Two situations genuinely justify going past a good general model:

Real vocabulary mismatch. Your users and your documents use different words for the same thing, and a general model doesn’t place them close enough to win against thousands of competing chunks — clinical shorthand against formal diagnostic language, internal product codenames against public names, or a domain where ordinary words carry a specific technical meaning. You can see this in the bake-off: the misses cluster by terminology rather than by topic.

Enough labelled pairs to fine-tune. Fine-tuning an embedding model on your own query-document pairs is the highest-ceiling option and it needs real supervision — thousands of pairs, ideally mined from click or feedback data rather than written by hand. If you have that data, this beats model shopping. If you don’t, it’s a project to acquire the data, not a model choice.

Two cheaper things to try first, because they often close the gap: put the document’s heading path or title into the embedded text so an orphaned paragraph carries its context, and lean harder on the lexical arm for the vocabulary you can enumerate. Both are edits to your pipeline, not a re-index of your corpus.

When you don’t need this

I’d skip the whole exercise and take a sensible default when:

  • The corpus is small and homogeneous. A few thousand chunks of ordinary prose. Any current general-purpose model will retrieve them well; spend the day on chunking or on getting a reranker in.
  • You have an unfixed structural problem. Broken PDF extraction, a document class that was never ingested, a filter that removes the answer. Comparing models on a corpus that’s missing the data is measuring the wrong thing precisely.
  • You have no eval set and no traffic yet. Without labelled queries the bake-off has no verdict, and inventing queries to justify a model choice mostly measures your imagination. Ship a default, watch what people ask, then come back.
  • You’re already hybrid and reranked and the misses aren’t representational. If your failures are boundary splits and missing documents, a new model will move nothing and cost a re-index.

Default when you skip it: a current general-purpose model from a provider you can already legally call, at the smallest dimension that clears your quality bar, with hybrid retrieval in front of it.

FAQ

Is a bigger embedding dimension better? Not reliably, and it’s never free. Dimension costs memory, storage and search time on every query forever. Where a model offers shorter truncations of the same vector, test the short version against your own eval set — often the quality difference is small enough that the operational saving wins outright.

Should I use an open model or a hosted API? Decide it on data residency and operations, not quality — the top of both categories is close enough that the constraints should choose. Hosted means no GPUs, no version management, and a per-query network hop plus a per-token bill. Self-hosted means the data never leaves, a fixed cost you control, and an inference service that is now yours to run, scale and keep alive.

Do I have to re-embed everything to change models? Yes. Vectors from different models don’t share a space, so there is no partial migration. Build the new index alongside the old, compare on the same eval set, cut over with a flag, and keep the old index until the new one has proven itself on live traffic.

How often should I re-evaluate my choice? Once or twice a year, or when something changes on your side — a new document type, a new language, a noticeable shift in what users ask. Chasing releases costs a re-index each time and usually buys less than the reranking or chunking work you’ve been postponing.

Does the same model handle multiple languages? Only if it was trained for it. A strong English model can degrade sharply on other languages, and a multilingual model buys cross-lingual matching (an English question retrieving a German document) at some cost to peak English performance. If your corpus is multilingual, make sure your eval set is too — otherwise you’re validating on the one language that works.

Can I mix models, one per document type? Not inside a single index — the vectors aren’t comparable. You can run separate indexes per collection and fuse the result lists by rank, which is the same machinery hybrid search already uses, but it’s real operational complexity and I’d want a measured reason before taking it on.


The pattern here is the one that runs through all of this work: constraints first, your own data second, leaderboards last. An embedding model is one component in a retrieval stack whose other parts — chunk boundaries, the lexical arm, the reranker, the eval loop — are cheaper to change and usually further from optimal. Pick a model that clears your constraints, prove it on fifty real queries, and go spend the rest of the week on the parts you can still iterate on.

If you’re weighing a retrieval change and can’t yet tell what it would buy you, that’s the gap worth closing first. I write about the rest of the pipeline in the hybrid search and RAG guides, and about building ranking and relevance layers that hold up in correlation and scoring.

Professional headshot of Ben Moataz
Written by
Ben Moataz

Systems Architect, Consultant, and Product Builder

This article is grounded in hands-on work across Correlation and scoring, including systems such as SOVRINT, TraxinteL, and Viralink.

I write from hands-on work across product systems, evidence pipelines, ranking layers, monitoring surfaces, and automation runtimes that have to stay reliable under operational pressure.

  • → Years spent building product systems, automation infrastructure, and operator-facing platforms.
  • → Project records and case studies tied directly to the same capability lanes discussed in the writing.
  • → A public archive designed to connect essays back to real systems, delivery constraints, and consulting work.
Relevant work

Expertise and case studies tied to this article.

Related reading

More writing on adjacent systems problems.

Next article

Worker Fleet Architecture at Scale: Pools, Scaling Signals, and Draining

Scaling a worker fleet isn't adding workers to one queue. Here's how I partition pools, pick the scaling signal, and drain workers without losing jobs.

Work with me

Building or fixing a system like this?

This is exactly the kind of work I get brought in for. Teams unsure whether a system, architecture, or workflow will hold up under real load and scrutiny.

System Audit Start here · fixed scope
  • → A focused review of the system, architecture, or codebase in question.
  • → A clear map of the risks, bottlenecks, and failure modes that matter.
  • → A prioritized roadmap — what to fix first, and what to leave alone.
Subscribe

Get new essays by email

Field notes on intelligence systems, evidence engineering, and automation that survives reality. No noise.

Subscribe via RSS → Email capture isn't wired up yet — the RSS feed is live now.