mirror of
https://github.com/zotero/zotero.git
synced 2026-10-11 03:38:25 +00:00
Add the query-to-score flow to Zotero.Lexical, built on the content and
item-text indexes:
- Term statistics span both corpora: document frequency sums MATCH
counts over fulltextContent and fulltextItemText, corpus size sums
the state tables, so a term's rarity is a property of the library --
and libraries with few attachments still get real weights
- Units too common to matter are cut relative to the query's best unit
(INFORMATIVE_WEIGHT_FRACTION), so "of" is dropped next to "communism"
while an all-common query keeps its best word
- Matchers are index probes: matchContent() against the content index,
matchFields()/matchNotes()/matchAnnotations() against the item-text
columns. Notes fetch text only for probe matches plus stale/unindexed
notes (getStaleOrUnindexedNoteIDs); quoted phrases verify literally
against stored text everywhere (whitespace/hyphen runs interchange,
other punctuation must match)
- scoreItemIDs() assembles the score (see below), with a floor and
cancellation
Ranking algorithm, per query:
1. Parse into units (words, quoted phrases, CJK runs); trailing
mid-word token matches as a prefix
2. Weigh each unit by smoothed BM25 IDF from the combined corpora;
keep the informative ones
3. Match: presence (1) in titles, abstracts, annotations; saturated,
length-normalized term frequency for notes (computed from text)
and documents (recovered as rank ratios per unit -- for a one-unit
query, ranks compare documents exactly, and the strongest match
anchors 1)
4. Score = sum over units of weight x best boosted evidence across
sources (title x2, abstract x1.3; max, so one word never counts
twice), normalized against the query's ceiling: 1 = full-strength
match on everything asked; below SCORE_FLOOR is no match
|
||
|---|---|---|
| .. | ||
| components | ||
| content | ||
| resource | ||
| tests | ||
| chrome.manifest | ||
| runtests.sh | ||