Skip to main content
Learn how to use the TF-IDF variant for keyword-based search without external dependencies.
In this guide you’ll learn when to use TF-IDF instead of semantic search, how to configure and query it, and understand its limitations.

When to Use TF-IDF

Basic Usage

src/tfidf-basic.ts

Key Differences from VectoriaDB

Important: Reindexing

TF-IDF requires reindexing after document changes to update IDF (Inverse Document Frequency) values:
src/tfidf-reindex.ts
Forgetting to call reindex() after changes will result in incorrect search results.

Configuration Options

src/tfidf-config.ts

Search Options

src/tfidf-search.ts

TF-IDF Algorithm

TF-IDF (Term Frequency-Inverse Document Frequency) works by:
  1. Term Frequency (TF): How often a term appears in a document
  2. Inverse Document Frequency (IDF): How rare a term is across all documents
  3. TF-IDF Score: TF x IDF - terms that are frequent in a document but rare overall get high scores
This means:
  • Common words like “the”, “is”, “a” get low scores (low IDF)
  • Unique terms specific to a document get high scores
  • The query is matched against TF-IDF vectors using cosine similarity

Example: Tool Discovery

src/tfidf-tool-discovery.ts

Limitations

  1. No semantic understanding - “car” won’t match “automobile”
  2. Reindex requirement - Must call reindex() after changes
  3. Limited to keywords - Misspellings and synonyms aren’t handled
  4. Memory for large vocabularies - IDF tables grow with vocabulary size

Hybrid Approach

For best of both worlds, you can use TF-IDF as a pre-filter before semantic search:
src/tfidf-hybrid.ts

Welcome

Getting started

Search

Semantic search options

Storage

Storage adapters