mirror of
https://github.com/zotero/zotero.git
synced 2026-10-07 02:58:09 +00:00
Mirror of https://github.com/zotero/zotero.git
Add the query-to-score flow to Zotero.Lexical, built on the content and
item-text indexes:
- Term statistics span both corpora: document frequency sums MATCH
counts over fulltextContent and fulltextItemText, corpus size sums
the state tables, so a term's rarity is a property of the library --
and libraries with few attachments still get real weights
- Units too common to matter are cut relative to the query's best unit
(INFORMATIVE_WEIGHT_FRACTION), so "of" is dropped next to "communism"
while an all-common query keeps its best word
- Matchers are index probes: matchContent() against the content index,
matchFields()/matchNotes()/matchAnnotations() against the item-text
columns. Notes fetch text only for probe matches plus stale/unindexed
notes (getStaleOrUnindexedNoteIDs); quoted phrases verify literally
against stored text everywhere (whitespace/hyphen runs interchange,
other punctuation must match)
- scoreItemIDs() assembles the score (see below), with a floor and
cancellation
Ranking algorithm, per query:
1. Parse into units (words, quoted phrases, CJK runs); trailing
mid-word token matches as a prefix
2. Weigh each unit by smoothed BM25 IDF from the combined corpora;
keep the informative ones
3. Match: presence (1) in titles, abstracts, annotations; saturated,
length-normalized term frequency for notes (computed from text)
and documents (recovered as rank ratios per unit -- for a one-unit
query, ranks compare documents exactly, and the strongest match
anchors 1)
4. Score = sum over units of weight x best boosted evidence across
sources (title x2, abstract x1.3; max, so one word never counts
twice), normalized against the query's ceiling: 1 = full-strength
match on everything asked; below SCORE_FLOOR is no match
|
||
|---|---|---|
| .github/workflows | ||
| app | ||
| chrome | ||
| defaults/preferences | ||
| document-worker@6d0c0ce45d | ||
| js-build | ||
| note-editor@acec74d09b | ||
| reader@c6edbfff75 | ||
| resource | ||
| scripts | ||
| scss | ||
| styles@dff7452b24 | ||
| test | ||
| translators@60f2d542ff | ||
| types/gecko | ||
| .babelrc | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| chrome.manifest | ||
| CLAUDE.md | ||
| CONTRIBUTING.md | ||
| COPYING | ||
| eslint.config.mjs | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| update.rdf | ||
| version | ||
Zotero
Zotero is a free, easy-to-use tool to help you collect, organize, cite, and share your research sources.
Please post feature requests or bug reports to the Zotero Forums. If you're having trouble with Zotero, see Getting Help.
For more information on how to use this source code, see the Zotero documentation.