zotero/test
Bogdan Abaev 6dec06f18c Ranked lexical scoring over the full-text indexes
Add the query-to-score flow to Zotero.Lexical, built on the content and
item-text indexes:

- Term statistics span both corpora: document frequency sums MATCH
  counts over fulltextContent and fulltextItemText, corpus size sums
  the state tables, so a term's rarity is a property of the library --
  and libraries with few attachments still get real weights
- Units too common to matter are cut relative to the query's best unit
  (INFORMATIVE_WEIGHT_FRACTION), so "of" is dropped next to "communism"
  while an all-common query keeps its best word
- Matchers are index probes: matchContent() against the content index,
  matchFields()/matchNotes()/matchAnnotations() against the item-text
  columns. Notes fetch text only for probe matches plus stale/unindexed
  notes (getStaleOrUnindexedNoteIDs); quoted phrases verify literally
  against stored text everywhere (whitespace/hyphen runs interchange,
  other punctuation must match)
- scoreItemIDs() assembles the score (see below), with a floor and
  cancellation

Ranking algorithm, per query:
  1. Parse into units (words, quoted phrases, CJK runs); trailing
     mid-word token matches as a prefix
  2. Weigh each unit by smoothed BM25 IDF from the combined corpora;
     keep the informative ones
  3. Match: presence (1) in titles, abstracts, annotations; saturated,
     length-normalized term frequency for notes (computed from text)
     and documents (recovered as rank ratios per unit -- for a one-unit
     query, ranks compare documents exactly, and the strongest match
     anchors 1)
  4. Score = sum over units of weight x best boosted evidence across
     sources (title x2, abstract x1.3; max, so one word never counts
     twice), normalized against the query's ceiling: 1 = full-strength
     match on everything asked; below SCORE_FLOOR is no match
2026-08-18 18:00:55 -07:00
..
components fx140: Asyncify/ESMify tests 2025-07-30 22:30:53 -04:00
content fx153: ownerGlobal -> documentGlobal 2026-08-03 11:48:40 -04:00
resource Remove Chai as Promised 2025-07-30 22:31:08 -04:00
tests Ranked lexical scoring over the full-text indexes 2026-08-18 18:00:55 -07:00
chrome.manifest fx115: Restore test runner 2024-03-30 00:58:53 -04:00
runtests.sh runtests.sh: Add -p option to run a shard of the test files (e.g., -p 2/4) 2026-07-08 10:50:32 -04:00