Search

Main blog

Welcome to our Main Blog, the cornerstone of our exploration into information retrieval. This dedicated space serves as a comprehensive repository where we delve into our research, findings, and various topics predominantly centered around information retrieval.

Understanding Embeddings in the Italian Language – Part 3

A study to assess the effectiveness of multilingual embedding models in handling Italian language, with an investigation on fine-tuning.

Understanding Embeddings in the Italian Language – Part 2

A study to assess the effectiveness of multilingual embedding models in handling Italian language, with an investigation on fine-tuning.

Understanding Embeddings in the Italian Language – Part 1

A study to assess the effectiveness of multilingual embedding models in handling Italian language, with an investigation on fine-tuning.

Databases vs. Search Engines: Has the Gap Finally Closed?

Sease Ltd. explores whether modern databases like PostgreSQL and MongoDB can replace dedicated search engines such as Apache Solr and OpenSearch. While databases have improved their search capabilities, they still lack the advanced features, scalability, and performance of search engines for complex applications. The choice depends on the scale and architecture of the project.

Boosted K-Nearest Neighbor Search

Is it possible to integrate eDisMax-like boosting in Approximate Nearest Neighbor search for Solr and Lucene?
Bridging the Gap Between Theory and Practice in Vector Search

Vector Search Doctor (Part 2): Bridging the Gap Between Theory and Practice in Vector Search

This blogpost introduces Approximate Search Evaluator: a tool to measure vector search performance, addressing the accuracy/speed trade-off.

Vector Search Doctor (Part 1): Beyond the MTEB Leaderboard for Custom Datasets

Embedding Model Evaluator is an MTEB benchmark designed to evaluate embedding models on user-provided datasets on retrieval and reranking tasks.

Search Quality Evaluation with LLMs: the Dataset Generator

Dataset Generator automates the creation of relevance datasets for search evaluation, generating queries and relevance ratings with LLMs.

The AI side of the Vespa Search Engine

Vespa implements several useful features for customizing and improving Vector Search. Here, we will go into detail of each of them.

Benchmarking JSON Facet Methods in Apache Solr

This blogpost explores the performance impact of DocValues vs. Inverted Index for Apache Solr facets done through JSON facet API.
Colbert Comes to Apache Solr: Implementation and Tutorial of Late Interaction Model Reranking

Colbert Comes to Apache Solr: Implementation and Tutorial of Late Interaction Model Reranking

Discover late interaction in Apache Solr: how to implement ColBERT-style neural reranking to boost search accuracy.

Late Interaction Comes to Solr: Neural Reranking Introduction

Discover late interaction in Apache Solr: how to implement ColBERT-style neural reranking to boost search accuracy.
Apache Solr 10 – What Is New for Vector Search and LTR

Apache Solr 10 – What Is New for Vector Search and LTR

This blog summarises the main new features introduced in Apache Solr 10.0.0, focusing on Vector Search and Learning to Rank (LTR).
Apache Solr New Learning To Rank Cache

Apache Solr New Learning To Rank Cache

A comparison between Solr's current LTR cache and a new implementation that works not only for logging features, but also for reranking.
New in Apache Solr 10 Improved Filtering In Vector Search with ACORN

New in Apache Solr 10: Improved Filtering In Vector Search with ACORN

Hybrid Search Using a Custom Algorithm in Apache Solr

Hybrid Search Using a Custom Algorithm in Apache Solr

This blog post explores the Combined Query Feature using a custom algorithm in Apache Solr with a hands-on approach.
Hybrid Search with Reciprocal Rank Fusion in Apache Solr

Hybrid Search with Reciprocal Rank Fusion in Apache Solr

We explore the Combined Query Feature using Reciprocal Rank Fusion in Apache Solr with a hands-on approach.