When to Use TF-IDF
Basic Usage
Key Differences from VectoriaDB
Important: Reindexing
TF-IDF requires reindexing after document changes to update IDF (Inverse Document Frequency) values:Configuration Options
Search Options
TF-IDF Algorithm
TF-IDF (Term Frequency-Inverse Document Frequency) works by:- Term Frequency (TF): How often a term appears in a document
- Inverse Document Frequency (IDF): How rare a term is across all documents
- TF-IDF Score: TF x IDF - terms that are frequent in a document but rare overall get high scores
- Common words like “the”, “is”, “a” get low scores (low IDF)
- Unique terms specific to a document get high scores
- The query is matched against TF-IDF vectors using cosine similarity
Example: Tool Discovery
Limitations
- No semantic understanding - “car” won’t match “automobile”
- Reindex requirement - Must call
reindex()after changes - Limited to keywords - Misspellings and synonyms aren’t handled
- Memory for large vocabularies - IDF tables grow with vocabulary size
Hybrid Approach
For best of both worlds, you can use TF-IDF as a pre-filter before semantic search:Related
Overview
Getting started
Search
Semantic search options
Persistence
Storage adapters