Inverted index fundamentalseasy0-2 years

`WHERE title LIKE '%kafka%'` scans every row. A search engine answers the same kind of query in milliseconds over millions of documents. What structural difference makes that possible, and what's the cost paid for it?

A LIKE '%kafka%' scan has to look inside every row's text at query time because nothing was precomputed. A search engine instead does the work at index time: it builds an inverted index that maps each term to the list of documents containing it (the postings list). A query for kafka becomes a lookup of one postings list, and a multi-term query intersects a few short lists — work proportional to how many documents actually contain those terms, not the size of the whole corpus. The cost is that this precomputed structure is expensive to update: Lucene segments are write-once, so an update is really a delete plus a fresh insert, and a write isn't searchable until the next periodic refresh, not instantly.

The lesson behind it →