Relevance and BM25hard5-8 years
Work through the lesson's own comparison: document A is 50 words (corpus average 200) and mentions `kafka` twice; document B is 800 words and mentions `kafka` six times. With `k1 = 1.2`, `b = 0.75`, which scores higher on that term, and why does the shorter document win despite fewer mentions?
Document A, despite mentioning kafka only twice against B's six, scores higher on that term — about 1.74 versus about 1.33 in the lesson's own worked arithmetic. IDF(kafka) is identical for both, since it only depends on the corpus, not on which document is being scored, so the entire difference comes from the length-adjusted term-frequency fraction: A is a quarter of the average document length, so a term appearing in it at all counts as A being disproportionately "about" that term, where B's six mentions are diluted across four times the average length and count for comparatively less. b (default 0.75) is the knob controlling how strongly relative length discounts the raw mention count this way.
PreviousTwo documents both mention `kafka`: one five times, one fifty times. Under BM25's default `k1 = 1.2`, would the fifty-mention document score roughly ten times higher on that term? Walk through why or why not.Next A product search puts `{ "term": { "status": "PUBLISHED" } }` inside a `bool` query's `must` clause alongside the user's search text. What's wasted by doing that, and what's the fix?