Analysis internalsmedium3-5 years

A custom analyser is defined with token filters in the order `["stemmer", "lowercase", "stop"]`. A document containing "The Consumers Are Running" is indexed through it, and later a lowercase query for "consumers running" fails to match. Trace the chain and explain why.

The chain runs in a fixed order — character filters, then the tokenizer, then the token filters in the order they're declared — and each stage only sees what the previous stage produced. Here stemmer runs before lowercase, so it receives the raw-cased tokens The, Consumers, Are, Running straight from the tokenizer. The English stemmer's suffix rules are written for lowercase input, so a capitalized word like Consumers and Running doesn't match the suffix patterns the stemmer expects and passes through unstemmed, while stop (also misplaced after stemming here) doesn't remove The/Are because the stemmer already mangled the casing question moot in a different way — the net effect is the indexed terms don't match what a correctly-ordered analyser at query time would produce for consumers running.

The lesson behind it →
More on Analysis internals