Vector Search Doctor
Diagnose vector search quality on your own data
Evaluate embedding models with exact search, test your production setup with approximate search and make data-driven decisions that improve retrieval quality.
IS IT THE MODEL OR THE SEARCH CONFIGURATION?
Poor vector search results can come from different parts of the pipeline.
Vector Search Doctor separates these problems and makes them measurable.
EMBEDDING MODEL
Defines the semantic representation of your data.
EXACT SEARCH BASELINE
Estabilishes the best retrieval quality your model can achieve.
APPROXIMATE SEARCH IN PRODUCTION
Measures how much quality is preserved with ANN search.
Two evaluators. One complete diagnosis.
01
Embedding Model Evaluator
Test the model on your own data
Determine whether your embedding model excels at semantic natural language queries searches or struggles with them on your custom domain data to establish your true retrieval baseline.
- Retrieval & Reranking
- Exact Vector Search
- nDCG@10, MAP, MRR
- Binary or Graded Relevance
- Embeddings Export (documents & queries)
Don't have an evaluation dataset ready?
No problem. Use our Dataset Generator to automatically generate queries, assign relevance ratings, and export MTEB-compatible datasets (corpus, queries, candidates) in minutes.
02
Approximate search evaluator
Measure what happens in production
Run the same queries against your real search engine configuration and compare the results with the exact-search baseline. Understand how much retrieval quality is preserved when ANN search is introduced.
- Exact vs. Approximate Quality: Measure the true quality preserved by your ANN setup.
- Native JSON Templates: Plug in your query templates using the $vector placeholder.
- Search Engine Native: Full out-of-the-box support for Solr, Elasticsearch, OpenSearch and Vespa.
- HNSW & Metric Optimization: Fine-tune ef_search, M, nDCG@10, MAP, and MRR.
From model to Production
1
Evaluate the model
Test your Hugging Face embedding model on custom domain data to isolate its true semantic capabilities across keyword and natural language queries.
2
Establish the baseline
Compute brute-force exact vector search to set the maximum retrieval quality ceiling (e.g., nDCG@10) achievable by your model on your dataset.
3
Test your setup
Test approximate search against your search engine using saved embeddings, evaluating dozens of configurations without recomputing vectors or wasting GPU resources.
4
Measure & Optimise
Quantify exact-to-approximate quality loss, tune HNSW parameters (ef_search, M), and select the optimal latency-accuracy balance backed by evidence.
MEASURE THE QUALITY YOU KEEP
Illustrative evaluation based on different HNSW configurations.
EXACT SEARCH (BASELINE)
nDCG@10
0.850
Best achievable quality with this model
APPROXIMATE SEARCH
nDCG@10
0.820
Quality achieved with your ANN configuration
Quality Retained
96.5%
Percentage of baseline quality retained
FAQ
Turn Search Quality into Evidence
Everything you need to understand where your vector search is losing quality.
How do I know if the embedding model is the problem?
Run the Embedding Model Evaluator using exact vector search.
Because exact search compares every query against every document in brute-force mode, it establishes the maximum retrieval ceiling your model can achieve on your domain. If scores (such as nDCG@10) are low during exact search, the problem lies in the embedding model itself—meaning it is not a good fit for your specific data, terminology, or query patterns.
Why compare exact and approximate vector search?
Exact vector search isolates the quality of the embedding model. Approximate Nearest Neighbor (ANN) search accelerates retrieval in production but trades off a small amount of accuracy.
Comparing both reveals the Exact-to-Approximate gap: it tells you precisely what percentage of retrieval quality your production engine preserves and whether quality loss stems from the model or an overly aggressive ANN configuration.
Can I evaluate the tool on my own data?
Yes. Unlike generic public leaderboards (like MTEB), Vector Search Doctor is designed specifically for custom domain data.
You can evaluate using your own:
Document corpus & queries
Relevance judgments (ground truth)
Hugging Face embedding models
Search engine index/collection
Don’t have a test dataset yet? You can use our open-source Dataset Generator to automatically generate queries and relevance ratings.
What evaluation tasks and relevance scales are supported?
The Embedding Model Evaluator supports two main tasks:
Retrieval: Evaluated primarily using nDCG@10.
Reranking: Evaluated primarily using MAP.
Both binary (0: not relevant, 1: relevant) and graded (0: not relevant, 1: acceptable, 2: highly relevant) rating scales are supported.
What search engines are supported?
The Approximate Search Evaluator includes out-of-the-box support for:
Apache Solr
Elasticsearch
OpenSearch
Vespa
Can Vector Search Doctor help me tune HNSW?
Yes. It turns HNSW parameter tuning into an evidence-based process. You can run benchmarks across different index configurations to measure how adjustments impact retrieval accuracy:
ef_search: Controls search-time neighbor exploration depth (latency vs. accuracy trade-off).ef_construction: Controls graph construction depth during indexing.M: Sets the maximum connection links per graph node (memory vs. accuracy trade-off).
What metrics does Vector Search Doctor calculate?
The tool computes standard Information Retrieval (IR) metrics across your top-$k$ results (e.g., top-10):
nDCG (Normalized Discounted Cumulative Gain)
MAP (Mean Average Precision)
MRR (Mean Reciprocal Rank)
Precision & Recall
What does “quality retained” mean?
Quality Retained is the percentage of exact-search baseline quality preserved by your production ANN setup.
Example:
Exact Search nDCG@10 =
0.850Production ANN nDCG@10 =
0.820Quality Retained:
0.820 / 0.850 = 96.5%
This single metric proves whether your production search engine is operating near optimal performance or sacrificing too much accuracy for speed.
Does Vector Search Doctor measure latency?
Its primary purpose is to measure retrieval quality. However, by pairing quality report metrics with your search engine’s query latency, CPU, and memory profiles, you can objectively select the best production configuration for your latency budget.
What do I need to get started?
1. For Model Evaluation:
A Hugging Face model ID
A simple YAML config file
Three JSONL files:
corpus.jsonl,queries.jsonl,candidates.jsonl(or generated via Dataset Generator)
2. For Approximate Search Evaluation:
An indexed search engine collection containing document embeddings
Query embeddings saved from Model Evaluation
A native JSON query template with the
$vectorplaceholderSearch engine endpoint & connection details
Search Quality Ecosystem
Vector Search Doctor integrates naturally with other Sease open source projects to help you build, evaluate and improve search quality systems end to end.
OPEN SOURCE PROJECT BUILT BY SEARCH EXPERTS