Analysis internalsmedium3-5 years

A search team wants to try 'better normalization' and proposes swapping the English stemmer for a lemmatizer, expecting mostly the same results with cleaner output. What's actually different about what each one does to a word, and what would change?

A stemmer strips suffixes by pattern — fast, no dictionary needed — and it doesn't care what the word means or even whether the result is a real word: consuming, consumer, and consumes all get chopped down to consum, which isn't an English word at all. A lemmatizer instead looks the word up against a vocabulary and returns its actual dictionary form (ran → run, better → good), often needing the word's part of speech to disambiguate a word like saw. Swapping to a lemmatizer wouldn't just be "cleaner" — it changes what matches what, needs language resources a suffix-stripper doesn't, costs more per document, and Elasticsearch doesn't ship one out of the box the way it ships stemmers for many languages.

The lesson behind it →
More on Analysis internals