unicode_words tokenizer splits text according to word boundaries defined by the Unicode Standard Annex #29
rules. All characters are lowercased by default.
This tokenizer is the default text tokenizer. If no tokenizer is specified for a text field, the unicode_words tokenizer will be used
(unless the text field is the key field, in which case the text is not tokenized).
Expected Response
Remove Emojis
By default, emojis in the source text are preserved. To remove emojis, setremove_emojis to true.
Expected Response