Analysis internalsmedium3-5 years
A search team wants to try 'better normalization' and proposes swapping the English stemmer for a lemmatizer, expecting mostly the same results with cleaner output. What's actually different about what each one does to a word, and what would change?
A stemmer strips suffixes by pattern — fast, no dictionary needed — and it doesn't care what the word means or even whether the result is a real word: consuming, consumer, and consumes all get chopped down to consum, which isn't an English word at all. A lemmatizer instead looks the word up against a vocabulary and returns its actual dictionary form (ran → run, better → good), often needing the word's part of speech to disambiguate a word like saw. Swapping to a lemmatizer wouldn't just be "cleaner" — it changes what matches what, needs language resources a suffix-stripper doesn't, costs more per document, and Elasticsearch doesn't ship one out of the box the way it ships stemmers for many languages.
PreviousA custom analyser is defined with token filters in the order `["stemmer", "lowercase", "stop"]`. A document containing "The Consumers Are Running" is indexed through it, and later a lowercase query for "consumers running" fails to match. Trace the chain and explain why.Next Two documents both mention `kafka`: one five times, one fifty times. Under BM25's default `k1 = 1.2`, would the fifty-mention document score roughly ten times higher on that term? Walk through why or why not.