Skip to main content
By default, ParadeDB uses the same tokenizer at both index time and search time. This makes sense for most cases — you want queries tokenized the same way the data was indexed. But sometimes you need different tokenizers. The classic example is autocomplete:
  • Index timeedge_ngram: "shoes"s, sh, sho, shoe, shoes
  • Search timeunicode_words: "sho"sho
If you used edge_ngram at search time too, typing "sho" would produce s, sh, sho — matching far too many documents.

Usage

Set search_tokenizer as a WITH option on the index to define a default search-time tokenizer for all text and JSON fields:
With this configuration:
  • Index time: title is tokenized into edge ngrams
  • Search time: queries against title automatically use the unicode_words tokenizer
The search_tokenizer value can include parameters, e.g. search_tokenizer='simple(lowercase=false)'. Because search_tokenizer only affects query-time behavior, you can change it without reindexing:

Example

Without search_tokenizer, the query 'sho' would be tokenized into the edge ngrams s, sh, sho and match every title starting with s — not just those starting with sho.

Overriding at Query Time

You can still override the search tokenizer for a specific query by casting the query string:

Priority

When resolving which tokenizer to use at search time, ParadeDB checks in this order:
  1. Query-level cast — e.g. 'sho'::pdb.edge_ngram(...) (highest priority)
  2. Index-level WITH option — e.g. WITH (search_tokenizer='unicode_words')
  3. Index-time tokenizer — the tokenizer used to build the index (fallback)

Supported Tokenizers

Any available tokenizer can be used as a search_tokenizer: unicode_words, simple, whitespace, ngram, edge_ngram, regex_pattern, literal, literal_normalized, chinese_compatible, lindera, icu, jieba, source_code.