Search

Vector Search Doctor

Diagnose vector search quality on your own data

Evaluate embedding models with exact search, test your production setup with approximate search and make data-driven decisions that improve retrieval quality.

IS IT THE MODEL OR THE SEARCH CONFIGURATION?

Poor vector search results can come from different parts of the pipeline. 
Vector Search Doctor separates these problems and makes them measurable.

EMBEDDING MODEL

Defines the semantic representation of your data.

EXACT SEARCH BASELINE

Estabilishes the best retrieval quality your model can achieve.

APPROXIMATE SEARCH IN PRODUCTION

Measures how much quality is preserved with ANN search.

Two evaluators. One complete diagnosis.

01

Embedding Model Evaluator

Test the model on your own data

Determine whether your embedding model excels at semantic natural language queries searches or struggles with them on your custom domain data to establish your true retrieval baseline.

Don't have an evaluation dataset ready?

No problem. Use our Dataset Generator to automatically generate queries, assign relevance ratings, and export MTEB-compatible datasets (corpus, queries, candidates) in minutes.

02

Approximate search evaluator

Measure what happens in production

Run the same queries against your real search engine configuration and compare the results with the exact-search baseline. Understand how much retrieval quality is preserved when ANN search is introduced.

From model to Production

1

Evaluate the model

Test your Hugging Face embedding model on custom domain data to isolate its true semantic capabilities across keyword and natural language queries.

2

Establish the baseline

Compute brute-force exact vector search to set the maximum retrieval quality ceiling (e.g., nDCG@10) achievable by your model on your dataset.

3

Test your setup

Test approximate search against your search engine using saved embeddings, evaluating dozens of configurations without recomputing vectors or wasting GPU resources.

4

Measure & Optimise

Quantify exact-to-approximate quality loss, tune HNSW parameters (ef_search, M), and select the optimal latency-accuracy balance backed by evidence.

MEASURE THE QUALITY YOU KEEP

Illustrative evaluation based on different HNSW configurations.

EXACT SEARCH (BASELINE)

nDCG@10

0.850

Best achievable quality with this model

APPROXIMATE SEARCH

nDCG@10

0.820

Quality achieved with your ANN configuration

Quality Retained

96.5%

Percentage of baseline quality retained

FAQ

Turn Search Quality into Evidence

Everything you need to understand where your vector search is losing quality.

How do I know if the embedding model is the problem?

Run the Embedding Model Evaluator using exact vector search.

Because exact search compares every query against every document in brute-force mode, it establishes the maximum retrieval ceiling your model can achieve on your domain. If scores (such as nDCG@10) are low during exact search, the problem lies in the embedding model itself—meaning it is not a good fit for your specific data, terminology, or query patterns.

Exact vector search isolates the quality of the embedding model. Approximate Nearest Neighbor (ANN) search accelerates retrieval in production but trades off a small amount of accuracy.

Comparing both reveals the Exact-to-Approximate gap: it tells you precisely what percentage of retrieval quality your production engine preserves and whether quality loss stems from the model or an overly aggressive ANN configuration.

Yes. Unlike generic public leaderboards (like MTEB), Vector Search Doctor is designed specifically for custom domain data.

You can evaluate using your own:

  • Document corpus & queries

  • Relevance judgments (ground truth)

  • Hugging Face embedding models

  • Search engine index/collection

Don’t have a test dataset yet? You can use our open-source Dataset Generator to automatically generate queries and relevance ratings.

The Embedding Model Evaluator supports two main tasks:

  • Retrieval: Evaluated primarily using nDCG@10.

  • Reranking: Evaluated primarily using MAP.

Both binary (0: not relevant, 1: relevant) and graded (0: not relevant, 1: acceptable, 2: highly relevant) rating scales are supported.

The Approximate Search Evaluator includes out-of-the-box support for:

  • Apache Solr

  • Elasticsearch

  • OpenSearch

  • Vespa

Yes. It turns HNSW parameter tuning into an evidence-based process. You can run benchmarks across different index configurations to measure how adjustments impact retrieval accuracy:

  • ef_search: Controls search-time neighbor exploration depth (latency vs. accuracy trade-off).

  • ef_construction: Controls graph construction depth during indexing.

  • M: Sets the maximum connection links per graph node (memory vs. accuracy trade-off).

The tool computes standard Information Retrieval (IR) metrics across your top-$k$ results (e.g., top-10):

  • nDCG (Normalized Discounted Cumulative Gain)

  • MAP (Mean Average Precision)

  • MRR (Mean Reciprocal Rank)

  • Precision & Recall

Quality Retained is the percentage of exact-search baseline quality preserved by your production ANN setup.

  • Example:

    • Exact Search nDCG@10 = 0.850

    • Production ANN nDCG@10 = 0.820

    • Quality Retained: 0.820 / 0.850 = 96.5%

This single metric proves whether your production search engine is operating near optimal performance or sacrificing too much accuracy for speed.

Its primary purpose is to measure retrieval quality. However, by pairing quality report metrics with your search engine’s query latency, CPU, and memory profiles, you can objectively select the best production configuration for your latency budget.

1. For Model Evaluation:

  • A Hugging Face model ID

  • A simple YAML config file

  • Three JSONL files: corpus.jsonl, queries.jsonl, candidates.jsonl (or generated via Dataset Generator)

2. For Approximate Search Evaluation:

  • An indexed search engine collection containing document embeddings

  • Query embeddings saved from Model Evaluation

  • A native JSON query template with the $vector placeholder

  • Search engine endpoint & connection details

Search Quality Ecosystem

Vector Search Doctor integrates naturally with other Sease open source projects to help you build, evaluate and improve search quality systems end to end.

OPEN SOURCE PROJECT BUILT BY SEARCH EXPERTS

RRE

Evaluate and compare search relevance over time.

Dataset Generator

Build high quality relevance datasets for search evaluation.