From fff76fcbfc7449fe6406f5282065300adb9c93c9 Mon Sep 17 00:00:00 2001 From: Bogdan Abaev Date: Wed, 7 Oct 2026 15:10:16 -0700 Subject: [PATCH] Best Match: hybrid lexical and semantic search MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Best Match becomes a ranked search over everything in a library: its items' metadata, notes and annotations, and the full text of its attachments. A lexical engine (Zotero.Lexical) always ranks; with semantic search enabled, the embedding engine ranks too and the two are fused, so an item can match by its words, by its meaning, or -- ranking highest -- by both. Matched passages are shown under their items, read in the item pane, and opened in the reader at the passage. Attachment vectors can come from the dataserver instead of being computed here. 1. Lexical engine (Zotero.Lexical, fulltext.js) A word-level FTS5 index, ftindex.fulltextItemText (plus a CJK 2-gram twin and a state table), holds each item's searchable text in per-type columns: title and abstract for regular items, note text for notes, the marked passage and comment for annotations. Regular items and annotations index inline in their save transaction; notes are written by the existing stale-flag queue alongside the trigram tables; a backfill queue covers pre-existing items and joins the startup and background drains. The trigram index answers substrings and inflates counts, so it can't tell "fall" from "rainfall" or weigh a word by its rarity; word-level FTS answers both with index probes. Ranking is FTS5's BM25 over that index and the existing content index: a query parses into word, phrase and CJK-run terms joined by OR, so a document missing a word still ranks below one that has them all; term weight comes from rarity in the user's own library, with no stoplist. Consecutive words are added as phrase terms, item-text columns are weighted (title 6, abstract 4, annotation 2, note 1), and an attachment's full text counts only where enough of the query's terms occur within 200 tokens of each other (FTS5 NEAR; 75% of them, so all of them up to three), so one rare word can't carry a long document to the top. Scores are divided by the most the expression could earn, so both indexes report the 0-1 share of the query a document carries. 2. Fusion (Zotero.BestMatch) Both engines score every candidate and their rankings are fused with Reciprocal Rank Fusion; results are the union of the engines' matches. Each engine's tail is cut against its own strongest match before fusion (search.bestMatchMargin, percent, default 50, in the Advanced pane): a library on one subject needs a tighter margin to separate a specific answer from the field, a varied one hardly needs it. The Relevance bar shows the item's strongest single piece of evidence, whichever engine found it. While the semantic index is still building, or semantic search is off, Best Match degrades to lexical ranking rather than showing nothing, so the quick-search mode is always offered and the "index not ready" state is gone. Advanced Search's bestMatch condition scores through the same facade. A temporary pref, search.bestMatchEngine (hybrid, lexical, semantic), selects the engine for testing. 3. Attachments, structured text and chunks (Zotero.SDT) Attachments are indexed on the structured text the document worker extracts (PDF, EPUB, snapshot), cached as a pack per attachment. How a pack is cut into chunks is moved to the structured-document-text module. The chunker lives there so that the client, the dataserver's indexer and anything else that embeds a document cut it the same way and produce rows that can be exchanged. Each chunk carries an anchor -- page rects for PDFs, selectors for EPUBs and snapshots -- that locates its text in the file across extractor versions, so a row cut elsewhere stays usable when a local re-cut would block differently. Anchors are stored as deflated JSON; chunk text is not stored, it's read back from the pack. The document-worker build bundles the chunker next to the pack reader (structured-document-text.js and structured-document-text-chunker.js), and Zotero requires both; on the main thread the chunker only cuts plain text and reports its version. A large book takes hundreds of ms to inflate and cut, so cutting (sdt.getChunks) and reading anchors back with their reader positions (sdt.readAnchors) run in the document worker, with no fallback here. A worker failure, or a chunker version that disagrees with the bundled one, comes back as a failed cut rather than a verdict on the attachment. On a worker error, the document worker manager fails every pending request and starts a fresh worker for the next one, instead of leaving the queue stuck behind an unanswered request. MIN_CHUNKER_VERSION forces a re-cut of attachments cut by an older chunker. 4. Model, vectors and calibration One model, bekko-embedding-v1-a25m: multilingual, 8192-token window, Matryoshka-trained so its 384-dimensional output is cut to 256. Two runtime tokenizer defects are corrected -- the leading word marker the runtime fails to add, and whitespace runs tokenized unlike the reference; the model `revision` tracks changes that alter vectors. Stored vectors are centered on the model's mean and quantized to int8 (the shared mean would otherwise spend most of the 8 bits), 256 bytes a row, and scored by cosine in SQL. The mean, the score floor and the bar's ceiling are measured by Calibration.record() over a corpus of triples -- a query, a passage that answers it, a near miss from the same field (resource/embeddings-calibration-corpus.json) -- and pasted into the model config; the floor sits at the 95th percentile of near-miss scores, the ceiling at the median of matches. 5. Indexing runs (Zotero.Embeddings.Indexing) Nothing is queued. An item change or a finished sync kicks a run, and a kick that lands during a run makes it go again. A run first indexes locally the items, notes and annotations saved since the index last looked at them -- a stamp other than the item's clientDateModified, or a save in the last 5 s -- reading their text from the tables, then takes every eligible attachment through five steps. Each step is one pass over the outstanding attachments, 50 at a time by itemID as full-text sync pages, and finds its own work in the index, so a run cut short resumes where it was: - reconcile: what's stored still holds. A file that's gone, or rows cut by a chunker before MIN_CHUNKER_VERSION, lose everything; - extract: every attachment's text cached as a pack, the server's included, so no preview waits on an extraction. - fetch: the server is asked for everything that's its. - cut: this client's attachments are cut into pending rows with anchors. An attachment the worker couldn't cut is left as it is for a later run, since a plain-text fallback would stand in for the file for good; only an extraction failure falls back to plain text. Rows adopted for a file that has since arrived wait the same way. - embed: every pending row gets a vector, pooled across attachments and pages. The attachment work gives way to a kick, so a just-edited item is searchable without waiting behind the library's documents, and to a sync in progress. Items are loaded without caching and never more than a page at a time. Startup and Resume clear the attachment stamps for a full pass. Text too short to say anything is skipped: under two words of title and abstract, under three of anything else. The index is one table of rows (itemID, chunkIndex, embedding, anchor) and one of per-item state (itemIndexState: sourceKey, contentHash, extractor, clientDateModified, and the item's standing with the server); counts are derived from rows. What a run is built on is declared beside it: Sources (what each kind of item offers and what text it yields), Progress (what the preferences pane reads, recomputed on a clock), Store (the only code that writes the tables), Runtime (the machine's say: a token budget per engine call halved under memory pressure, a memory floor to start at all, the engine restarted to give memory back, half the optimal thread count unless the prefs pane is open or the system idle, and main-thread work paced against idle time), and Diagnostics (rates and shape summaries of what Indexing records). The engine, which the runtime terminates when idle, is checked and recreated before each use. 6. Sync client (Zotero.Embeddings.Sync) Available when the account syncs (pref embeddings.sync.enabled is an override) and no endpoint is active -- configuring one says to embed there instead. The server is asked about a stored file in a library syncing with Zotero Storage, in sync or to download, that it hasn't declined. Rows arriving before their file are adopted when it does. GET users/{userID}/embeddings?model=M&itemKey=K1,K2,... returns { model, items: [ { key, status: 'success', contentHash, chunks, version, rows: [{ chunkIndex, embedding, anchor }] }, { key, status: 'declined' }, { key, status: 'pending' } ] } - success: rows replace what's stored, if contentHash matches the file here (or its plain text), rows number exactly `chunks`, and each row is well formed; otherwise declined. Rows already held from `version` are kept. - declined: cut and embedded locally in the same run. - pending, or a key missing from the reply: left as it is and asked again next run. - A reply in another model declines the batch. A failed request stops the pipeline; the run goes back to the server after 5 minutes, or at once on Resume. GET users/{userID}/embeddings?format=versions&model=M&since=V (with If-Modified-Since-Version) returns { version, items: { key: version }, models }. After every sync, checkLibrary() asks for what changed since the version it recorded and marks every key whose rows aren't from the version named to be fetched again, declined ones included. The library's version is recorded last, so a failure repeats the delta. Server is expected to provide raw vector (not centered, not quantized) so it does not have to worry about mean vector config. 7. Remote endpoint (Zotero.Embeddings.Endpoint) Passages can be embedded by a server serving the same model -- a local llama.cpp server in the OpenAI format, or Text Embeddings Inference -- configured from Settings -> Advanced. The server is trusted only because its vectors match the local model's, never by name: verify() embeds fixed texts both ways and stores a typed verdict (ok, unreachable, unauthorized, not-embeddings, width-mismatch, low-agreement, context-too-small) keyed to the URL, model version and format. Every batch carries a sentinel text whose local vector is cached, so a server switched to another model or pooling is caught on that batch; three consecutive failures skip the endpoint for the rest of the run. Queries always embed locally. The Configure Endpoint dialog shows the facts the server must match and a copyable llama.cpp command. 8. Previews (Zotero.BestMatch.Session, item tree, item pane) A Session scores a query and owns the passages its results matched in. The chunk is the unit of a match for both engines: passages come from the index's own rows read back by anchor, or, for an unindexed item, from the same cut applied to its structured text (only where already extracted; generating it costs seconds) or its plain text. An item returns at most three quoted matches, each blending the model's score with how much of the query the passage's own words carry. The quoted line is chosen after ranking: the sentence the lexical engine picks where the passage says the query's words; where it only means them, the sentence a static multilingual model (potion-multilingual-128M, via the runtime's static-embeddings backend) finds closest -- weaker than a dense model, far faster, and never stored. The item tree shows matches as two-line child rows -- where the passage is (section path, page) and the line worth reading, with the query's words marked -- and a new query shows its results from the top. The ten best-ranked items' previews are derived before scoring resolves; the rest are derived while the main thread is idle, paced against what each costs, and arrive in batches. A changed item's previews are invalidated and re-derived; the tree is no longer refreshed as embedding progresses, which kept freezing it mid-search. Selecting an item shows every passage it matched in a Search Results section of the item pane; selecting match rows shows the passages themselves, grouped by attachment, in a pane of their own. Double-click or Enter opens the attachment at the passage, a PDF scrolled to and highlighting the anchor's rects. 9. Preferences The model menu is replaced by one switch, search.bestMatch.enableSemantic (off: lexical ranking only, nothing indexed, what's indexed kept), and the mode-change confirmations go with it. The pane shows two progress bars -- metadata, notes and annotations; attachments -- with the step under way ("Preparing documents", "Syncing semantic data… X / Y", "Generating semantic data locally for N items"), the endpoint's status and its Configure dialog, the quality cutoff, and a diagnostics panel, hidden by default: throughput, inference speed, padding efficiency, batches, engine threads and restarts, process memory and CPU, chunk size distributions. 10. Build and dependencies The document-worker submodule gains the sdt.getChunks and sdt.readAnchors actions and builds the chunker as a second bundle next to the pack reader. Zotero.ML allows the static-embeddings backend and Mozilla's model hub. The embeddings database is at version 14 and is rebuilt on upgrade. --- .../content/zotero/collectionViewItemTree.jsx | 309 +- .../zotero/components/windowed-list.js | 13 +- chrome/content/zotero/customElements.js | 3 + .../zotero/elements/collapsibleSection.js | 2 +- chrome/content/zotero/elements/itemDetails.js | 3 + chrome/content/zotero/elements/itemPane.js | 31 +- .../zotero/elements/itemPaneSidenav.js | 2 +- .../zotero/elements/quickSearchTextbox.js | 33 +- .../zotero/elements/searchResultRow.js | 172 + .../zotero/elements/searchResultsBox.js | 171 + .../zotero/elements/searchResultsPane.js | 132 + .../content/zotero/elements/zoteroSearch.js | 9 +- chrome/content/zotero/itemTree.jsx | 152 +- chrome/content/zotero/itemTreeColumns.jsx | 18 +- chrome/content/zotero/itemTreeRow.js | 249 +- .../preferences/embeddingsEndpoint.xhtml | 182 + .../preferences/preferences_advanced.js | 294 +- .../preferences/preferences_advanced.xhtml | 73 +- chrome/content/zotero/xpcom/bestMatch.js | 978 +++ .../content/zotero/xpcom/collectionTreeRow.js | 4 +- chrome/content/zotero/xpcom/data/item.js | 18 +- chrome/content/zotero/xpcom/data/search.js | 33 +- chrome/content/zotero/xpcom/embeddings.js | 5234 +++++++++++++---- chrome/content/zotero/xpcom/fulltext.js | 407 +- chrome/content/zotero/xpcom/lexical.js | 1138 ++++ chrome/content/zotero/xpcom/ml.js | 54 +- chrome/content/zotero/xpcom/notifier.js | 2 +- .../content/zotero/xpcom/pdfWorker/manager.js | 62 + chrome/content/zotero/xpcom/sdt.js | 202 +- chrome/content/zotero/xpcom/searchQuery.js | 4 +- .../zotero/xpcom/sync/syncAPIClient.js | 91 +- .../content/zotero/xpcom/sync/syncRunner.js | 21 + .../zotero/xpcom/utilities_internal.js | 179 +- chrome/content/zotero/zotero.mjs | 2 + chrome/content/zotero/zoteroPane.js | 69 +- chrome/locale/en-US/zotero/preferences.ftl | 101 +- chrome/locale/en-US/zotero/zotero.ftl | 13 + defaults/preferences/zotero.js | 22 +- js-build/document-worker.js | 2 +- resource/embeddings-calibration-corpus.json | 348 ++ resource/schema/userdata.sql | 2 +- scss/_zotero.scss | 2 + scss/abstracts/_variables.scss | 1 + scss/components/_item-tree.scss | 62 +- scss/elements/_annotationRow.scss | 6 +- scss/elements/_itemPaneSearchResults.scss | 48 + scss/elements/_searchResultsBox.scss | 85 + scss/preferences/_advanced.scss | 99 +- test/tests/advancedSearchTest.js | 23 + test/tests/bestMatchTest.js | 1121 ++++ test/tests/collectionViewItemTreeTest.js | 699 ++- test/tests/embeddingsTest.js | 3661 +++++++++++- test/tests/fulltextTest.js | 87 + test/tests/lexicalTest.js | 671 +++ test/tests/preferences_advancedTest.js | 49 +- test/tests/sdtTest.js | 315 +- test/tests/searchQueryTest.js | 5 +- test/tests/searchTest.js | 9 +- test/tests/utilities_internalTest.js | 17 +- 59 files changed, 15983 insertions(+), 1811 deletions(-) create mode 100644 chrome/content/zotero/elements/searchResultRow.js create mode 100644 chrome/content/zotero/elements/searchResultsBox.js create mode 100644 chrome/content/zotero/elements/searchResultsPane.js create mode 100644 chrome/content/zotero/preferences/embeddingsEndpoint.xhtml create mode 100644 chrome/content/zotero/xpcom/bestMatch.js create mode 100644 chrome/content/zotero/xpcom/lexical.js create mode 100644 resource/embeddings-calibration-corpus.json create mode 100644 scss/elements/_itemPaneSearchResults.scss create mode 100644 scss/elements/_searchResultsBox.scss create mode 100644 test/tests/bestMatchTest.js create mode 100644 test/tests/lexicalTest.js diff --git a/chrome/content/zotero/collectionViewItemTree.jsx b/chrome/content/zotero/collectionViewItemTree.jsx index 495027819e..bd9d148bf3 100644 --- a/chrome/content/zotero/collectionViewItemTree.jsx +++ b/chrome/content/zotero/collectionViewItemTree.jsx @@ -43,7 +43,7 @@ const React = require('react'); const ReactDOM = require('react-dom'); const ItemTree = require('zotero/itemTree'); const { ItemTreeRowProvider } = ItemTree; -const { LibraryHeaderItemTreeRow, SpacerItemTreeRow } = require('zotero/itemTreeRow'); +const { LibraryHeaderItemTreeRow, SpacerItemTreeRow, SearchMatch } = require('zotero/itemTreeRow'); const { OS } = ChromeUtils.importESModule("chrome://zotero/content/osfile.mjs"); const { ZOTERO_CONFIG } = ChromeUtils.importESModule('resource://zotero/config.mjs'); @@ -171,23 +171,24 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { } /** - * Best-match ranks for the Relevance column, computed over the merged - * result set in _refresh() while a best-match search is active + * Best-match ranks for the Relevance column while a best-match search is + * active, from the session's last scoring pass * * @returns {Map} - treeViewID -> 1-based rank (1 = most similar) */ getBestMatchRanks() { - return this._bestMatchRanks || new Map(); + return this._bestMatchSession?.ranks ?? new Map(); } /** * Score fractions for the Relevance column's bars, computed alongside the - * ranks (see Zotero.Embeddings.getScoreFraction()) + * ranks (see Zotero.BestMatch.Session#barFractions), so the bars always + * agree with the ranking * - * @returns {Map} - treeViewID -> 0-1 fraction of the model's display range + * @returns {Map} - treeViewID -> 0-1 fraction for the bar */ getBestMatchBarFractions() { - return this._bestMatchBarFractions || new Map(); + return this._bestMatchSession?.barFractions ?? new Map(); } /** @@ -202,62 +203,25 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { } /** - * Compute the current index coverage across the selected rows' libraries. - * Never throws -- the banner is informational and shouldn't break a - * refresh. + * Compute the current index coverage (see Zotero.BestMatch.getIndexState()) * * @return {Promise} */ async _getBestMatchIndexState() { - try { - let status = Zotero.Embeddings.Indexing.getStatus(); - if (!status.enabled) { - return null; - } - // Counts aren't populated until the indexer runs in this session - if (!status.libraries.length) { - status = await Zotero.Embeddings.Indexing.refreshStatus(); - } - let libraryIDs = new Set( - this.collectionTreeRows - .map(row => row.ref?.libraryID) - .filter(id => id !== undefined) - ); - let libraries = status.libraries - .filter(lib => !libraryIDs.size || libraryIDs.has(lib.libraryID)); - let indexed = libraries.reduce((sum, lib) => sum + lib.indexed, 0); - let total = libraries.reduce((sum, lib) => sum + lib.eligible, 0); - if (indexed >= total) { - return null; - } - // Only an explicit pause reports as paused. Anything else -- - // between runs (startup, the pre-run debounce) or after an error - // (detailed in the preferences) -- reports as indexing, since the - // banner explains the incomplete coverage, not the indexer state - return { - type: status.paused ? 'paused' : 'indexing', - indexed, - total - }; - } - catch (e) { - Zotero.logError(e); - return null; - } + return Zotero.BestMatch.getIndexState(); } /** - * The semantic stage of a best-match search: score the merged, + * The ranking stage of a best-match search: score the merged, * deduplicated results from all selected rows against the query in a - * single call, and keep the scoreable items ranked globally across the - * selection. Child items (attachments, notes, annotations) are scored via - * their top-level item, so result sets at other levels (e.g. a saved - * search returning annotations) rank by their parent item. Equal scores - * get equal ranks, so tied rows (including a child and its parent) order - * deterministically via the secondary sort fields. + * single session call, which also derives the match previews and + * computes the ranks and bar fractions the Relevance column reads (see + * Zotero.BestMatch.Session#score()). An item is kept when it or + * anything beneath it matched, so a strongly matching annotation keeps + * its attachment and its paper in the results. * * @param {Zotero.Item[]} items - Merged results from all selected rows - * @return {Promise} - The scoreable items + * @return {Promise} - The matching items */ async _applyBestMatch(items) { // With multiple selected rows carrying different best-match sources, @@ -265,10 +229,13 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { let queryRow = this.collectionTreeRows.find(rowIsBestMatchSearch); let query = queryRow.getBestMatchQuery(); let source = queryRow.getBestMatchSource(); - // A top-K cutoff is reapplied to the merged candidates below only when - // the source is the transient Advanced Search, which applies uniformly - // to every selected row. A saved search's cutoff is part of that row's - // own membership and must not trim other selected rows' results. + // Each selected row's search applies a top-K cutoff to its own scope, + // so K is reapplied to the merged candidates (see Session#score()) -- + // a multi-row selection returns K members total rather than K per row. + // Only when the source is the transient Advanced Search, which applies + // uniformly to every selected row: a saved search's cutoff is part of + // that row's own membership and must not trim other selected rows' + // results. let topK = queryRow.advancedSearch && source ? source.getBestMatchQuery().topK : false; @@ -277,83 +244,121 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { // searches, so keep unscoreable items -- they sort after the ranked // ones. (A uniform top-K set contains no unscoreable items anyway.) let keepUnscored = !!source; - // Map each item to the item whose embedding scores it - let sourceIDByItem = new Map(); - for (let item of items) { - if (!(item instanceof Zotero.Item)) { - continue; - } - let source = item.isRegularItem() ? item : item.topLevelItem; - if (source) { - sourceIDByItem.set(item, source.id); - } - } - let scores; + let candidateIDs = items + .filter(item => item instanceof Zotero.Item) + .map(item => item.id); let generation = this._bestMatchGeneration; + // The session scores the query and owns the match previews the tree + // shows as child rows. A new query gets a fresh session -- the old + // one must derive nothing more -- while a re-score of the same query + // (an item edit, an index update) keeps it, so already-derived + // previews survive; the previews of the items that actually changed + // are invalidated in notify(). + let session = this._bestMatchSession; + let newQuery = !session || session.queryText !== query; + if (newQuery) { + session?.dispose(); + session = Zotero.BestMatch.createSession(query); + session.onPreviewsFilled = itemIDs => this._showFilledPreviews(session, itemIDs); + this._bestMatchSession = session; + } try { - scores = await Zotero.Embeddings.scoreItemIDs(query, [...new Set(sourceIDByItem.values())], { + // Scoring derives the best-scored items' previews before it + // resolves; the rest arrive through onPreviewsFilled above + await session.score(candidateIDs, { + topK, // A newer filter (e.g. more typed search text) makes this // query obsolete -- stop scoring and let its refresh take over shouldCancel: () => generation !== this._bestMatchGeneration }); } catch (e) { - if (e instanceof Zotero.Embeddings.ScoringCancelledError) { + if (e instanceof Zotero.BestMatch.ScoringCancelledError) { throw e; } - // Scoring can fail while the model is still downloading or the - // index is being rebuilt - if (e instanceof Zotero.Embeddings.IndexNotReadyError) { - Zotero.debug("Embeddings: index not ready for best-match search"); + Zotero.logError(e); + session.dispose(); + if (this._bestMatchSession == session) { + this._bestMatchSession = null; } - else { - Zotero.logError(e); - } - this._bestMatchRanks = new Map(); this._bestMatchIndexState = await this._getBestMatchIndexState(); - // A rank-only search's membership doesn't depend on the index, so + // A rank-only search's membership doesn't depend on scoring, so // show its results unranked; anything else shows no results rather // than an unranked scope return keepUnscored ? items : []; } - // Each selected row's search applies a top-K cutoff to its own scope, - // so trim the merged candidates to K again here, with the same - // deterministic order as search(), so a multi-row selection returns K - // members total rather than K per row - if (topK) { - scores = new Map( - [...scores.entries()] - .sort((a, b) => (b[1] - a[1]) || (a[0] - b[0])) - .slice(0, topK) - ); + // A cancellation that lands after the last derivation resolves + // score() normally, so check once more before building the view state + if (generation !== this._bestMatchGeneration) { + throw new Zotero.BestMatch.ScoringCancelledError(); } - let rankOfScore = new Map( - [...new Set(scores.values())].sort((a, b) => b - a).map((score, i) => [score, i + 1]) - ); + // A new query's results are shown from the top (see _refresh()) + if (newQuery) { + this._scrollToTopOnUpdate = true; + } + // The session's ranks cover every row with a match anywhere beneath + // it, so they say which items stay in the results let kept = []; - let ranks = new Map(); - let fractions = new Map(); for (let item of items) { - let sourceID = sourceIDByItem.get(item); - if (sourceID === undefined || !scores.has(sourceID)) { + if (!(item instanceof Zotero.Item) || !session.ranks.has(item.treeViewID)) { if (keepUnscored) { kept.push(item); } continue; } kept.push(item); - ranks.set(item.treeViewID, rankOfScore.get(scores.get(sourceID))); - fractions.set( - item.treeViewID, - Zotero.Embeddings.getScoreFraction(scores.get(sourceID)) - ); } - this._bestMatchRanks = ranks; - this._bestMatchBarFractions = fractions; this._bestMatchIndexState = await this._getBestMatchIndexState(); return kept; } + /** + * Show the match rows of previews derived after the search resolved (see + * Zotero.BestMatch.Session#score()), by reopening each item -- the same + * path that builds children for an expansion the user asks for. + * + * @param {Zotero.BestMatch.Session} session - Ignored once it isn't the + * session the tree is showing + * @param {Number[]} itemIDs + */ + _showFilledPreviews(session, itemIDs) { + if (this._bestMatchSession !== session) { + return; + } + let shown = []; + for (let itemID of itemIDs) { + // Looked up per item, since reopening one shifts the rows below it + let item = Zotero.Items.get(itemID); + let index = item ? this._rowMap[item.treeViewID] : undefined; + if (index === undefined || !this.isContainer(index)) { + continue; + } + if (this.isContainerOpen(index)) { + this._toggleOpenState(index); + } + this._toggleOpenState(index); + shown.push(itemID); + } + // Redrawing is the expensive part, so a batch with no rows in the + // tree (under a collapsed parent, say) costs nothing + if (!shown.length) { + return; + } + // The twisty appears with the preview, so the rows redraw too + this.itemTree.invalidateRowCache(shown); + this.runListeners('update', true, { restoreSelection: true, restoreScroll: true }); + } + + /** + * The session holding the passages of the active best-match search, or + * null when no such search is running + * + * @return {Zotero.BestMatch.Session|null} + */ + get bestMatchSession() { + return this._bestMatchSession ?? null; + } + /** * When showing multiple libraries, group rows by library in collections-list * order -- independent of the active sort direction @@ -585,14 +590,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { this.itemTree._refreshPromise = deferred.promise; try { - // A best-match rerank after an embeddings change reuses the cached - // search results -- only the ranking depends on the embeddings, so - // there's no need to re-run the underlying search - if (!options.reuseSearchResults) { - this.collectionTreeRows.forEach(row => row.clearCache()); - } - this._bestMatchRanks = null; - this._bestMatchBarFractions = null; + this.collectionTreeRows.forEach(row => row.clearCache()); this._bestMatchIndexState = null; // Get the full set of items we want to show, merged across all selected rows let newSearchItemSet = new Set(); @@ -656,7 +654,12 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { || item.isRegularItem(); }); } - // The semantic stage: one scoring pass over the merged results + // The ranking stage: one scoring pass over the merged results + if (!bestMatchSearch && this._bestMatchSession) { + // Leaving best-match search: the previews go with it + this._bestMatchSession.dispose(); + this._bestMatchSession = null; + } if (bestMatchSearch) { try { newSearchItems = await this._applyBestMatch(newSearchItems); @@ -665,7 +668,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { // A newer filter superseded this one mid-scoring -- leave the // rows as they are and let the newer filter's refresh replace // them - if (e instanceof Zotero.Embeddings.ScoringCancelledError) { + if (e instanceof Zotero.BestMatch.ScoringCancelledError) { deferred.resolve(); return; } @@ -712,6 +715,11 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { if (!row.isObjectRow) { continue; } + // Don't copy search-match rows -- they're rebuilt from the new + // query's previews when their container reopens + if (row.ref instanceof SearchMatch) { + continue; + } // Top-level items if (row.level == 0) { // A top-level attachment moved into a parent. Don't copy, it will be added @@ -794,6 +802,14 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { // applied to rows carried over from the previous view this._sort(forceSortAll || this._groupedByLibrary ? null : [...addedItemIDs]); + // Set before the containers below are rebuilt, since their children + // are filtered against these (see ItemTreeRow#getChildItems(), + // which hides the annotations a search didn't match): rebuilding + // them against the previous search's state leaves rows this one + // excludes + this._searchMode = newSearchMode; + this._searchItemIDs = newSearchItemIDs; // items matching the search + // Toggle all open containers closed and open to refresh child items var t = new Date(); for (let i = this.rows.length - 1; i >= 0; i--) { @@ -809,8 +825,6 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { this.refreshRowMap(); } - this._searchMode = newSearchMode; - this._searchItemIDs = newSearchItemIDs; // items matching the search this.itemTree.invalidateRowCache(true); if (this.viewMode != 'publications') { @@ -838,11 +852,19 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { if (!this.isContainer(i) || this.isContainerOpen(i)) { continue; } - let item = this.getRow(i).ref; + let row = this.getRow(i); + if (!(row.ref instanceof Zotero.Item)) { + continue; + } + let item = row.ref; let attachments = item.isRegularItem() ? item.getAttachments() : []; // expand item row if it is a parent of a match // OR if it has a child that is a parent of a match - let shouldBeOpened = searchParentIDs.has(item.id) || attachments.some(id => searchParentIDs.has(id)); + // OR if it has best-match preview rows to show -- one still + // deriving has none, and opens in _showFilledPreviews() instead + let shouldBeOpened = searchParentIDs.has(item.id) + || attachments.some(id => searchParentIDs.has(id)) + || this._bestMatchSession?.getPreviews(item.id)?.state == 'filled'; if (shouldBeOpened) { this._toggleOpenState(i, true); } @@ -866,6 +888,12 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { try { await this._refresh(options); + // A new best-match query shows its results from the top (see + // _applyBestMatch()), wherever the previous ones were scrolled to + if (this._scrollToTopOnUpdate) { + this._scrollToTopOnUpdate = false; + options = { ...options, scrollToTop: true }; + } this.runListeners('update', true, options); await this.itemTree.waitForLoad(); this.itemTree.runListeners('refresh'); @@ -900,12 +928,30 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { await this.itemTree._refreshPromise; const cachedSelection = this.itemTree._cachedSelection; - const collectionTreeRows = this.collectionTreeRows; + const collectionTreeRows = this.collectionTreeRows; + + // Opening an attachment writes its lastRead, which arrives here as an + // ordinary modify and triggers best match rerun if it is active. + // For now, do nothing since best match searches are costly. + if (type == 'item' && action == 'modify' && ids.length + && collectionTreeRows.some(rowIsBestMatchSearch) + && ids.every((id) => { + let item = Zotero.Items.get(id); + return item && item.isAttachment(); + })) { + return; + } + + // A changed item's derived match previews are stale: back to pending, + // re-derived by the refresh the change triggers below, which keeps + // every other item's derived text. + if (type == 'item' && ['modify', 'refresh'].includes(action) && this._bestMatchSession) { + this._bestMatchSession.invalidate(ids.map(id => parseInt(id))); + } var initialRowCount = this.getRowCount(); var madeChanges = false; var refresh = false; - var reuseSearchResults = false; var sort = false; // Selection strategy @@ -956,18 +1002,10 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { if (items.length == 0) return; } - if (action == 'refresh' && type == 'item' && extraData && extraData.embeddingsUpdate - && collectionTreeRows.some(rowIsBestMatchSearch)) { - // The background indexer committed new or changed embeddings, so - // rerun the active best-match search to update the scores and ranks. - // Only the ranking depends on the embeddings, so unless a selected - // row trims membership by score (a top-K cutoff), the underlying - // search results are unchanged and can be reused while just the - // ranking is recomputed. - reuseSearchResults = collectionTreeRows.every((row) => { - return typeof row.hasBestMatchCutoff == 'function' - && !row.hasBestMatchCutoff(); - }); + if (action == 'refresh' && type == 'item' && this._bestMatchSession + && ids.some(id => this._bestMatchSession.getPreviews(parseInt(id)))) { + // The invalidation above reset these items' previews, and only a + // scoring pass derives previews, so re-run the search this.itemTree.invalidateRowCache(ids); refresh = true; madeChanges = true; @@ -1248,7 +1286,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider { } if (refresh) { - await this._refresh({ reuseSearchResults }); + await this._refresh(); } if (sort) { await this.itemTree._ensureSortContextReady(); @@ -1456,9 +1494,8 @@ class CollectionViewItemTree extends ItemTree { /** * While a best-match search runs against a partially built embeddings * index, show a banner above the items list with the indexing progress, - * so incomplete results aren't mistaken for a complete ranking. Updates - * arrive with the periodic refreshes the indexer triggers as it fills - * the index. + * so incomplete results aren't mistaken for a complete ranking. The + * counts are read when the search refreshes. */ _renderTablePrologue() { let state = this.rowProvider.getBestMatchIndexState(); diff --git a/chrome/content/zotero/components/windowed-list.js b/chrome/content/zotero/components/windowed-list.js index cd264e6851..9f820b0cfe 100644 --- a/chrome/content/zotero/components/windowed-list.js +++ b/chrome/content/zotero/components/windowed-list.js @@ -159,11 +159,7 @@ module.exports = class { Object.assign(this, options); const { itemHeight, targetElement, innerElem } = this; const itemCount = this._getItemCount(); - const [offsetIdx, offset] = this._rowOffsets.at(-1); - const listHeight = offset + (itemCount - offsetIdx) * this.itemHeight; - innerElem.style.position = 'relative'; - innerElem.style.height = `${listHeight}px`; - + // Recalculate custom row height offsets this._rowOffsets = [[0, 0]]; let previousRowOffset = 0; @@ -176,6 +172,13 @@ module.exports = class { previousRowOffset = offset; } + // From the offsets just recalculated: the list is as tall as the rows + // it now has, not as the rows it had + const [offsetIdx, offset] = this._rowOffsets.at(-1); + const listHeight = offset + (itemCount - offsetIdx) * itemHeight; + innerElem.style.position = 'relative'; + innerElem.style.height = `${listHeight}px`; + this.scrollDirection = 0; this.scrollOffset = targetElement.scrollTop; } diff --git a/chrome/content/zotero/customElements.js b/chrome/content/zotero/customElements.js index 29fd6a0432..e2a75165c5 100644 --- a/chrome/content/zotero/customElements.js +++ b/chrome/content/zotero/customElements.js @@ -75,6 +75,9 @@ Services.scriptloader.loadSubScript('chrome://zotero/content/elements/itemTreeMe ['attachment-row', 'chrome://zotero/content/elements/attachmentRow.js'], ['attachment-annotations-box', 'chrome://zotero/content/elements/attachmentAnnotationsBox.js'], ['annotation-row', 'chrome://zotero/content/elements/annotationRow.js'], + ['search-results-box', 'chrome://zotero/content/elements/searchResultsBox.js'], + ['search-result-row', 'chrome://zotero/content/elements/searchResultRow.js'], + ['search-results-pane', 'chrome://zotero/content/elements/searchResultsPane.js'], ['annotation-items-pane', 'chrome://zotero/content/elements/annotationItemsPane.js'], ['context-notes-list', 'chrome://zotero/content/elements/contextNotesList.js'], ['note-row', 'chrome://zotero/content/elements/noteRow.js'], diff --git a/chrome/content/zotero/elements/collapsibleSection.js b/chrome/content/zotero/elements/collapsibleSection.js index c509da76db..1ae9ef089a 100644 --- a/chrome/content/zotero/elements/collapsibleSection.js +++ b/chrome/content/zotero/elements/collapsibleSection.js @@ -393,7 +393,7 @@ } get _disableSavingOpenState() { - return !!this.closest('merge-pane, scaffold-item-preview, annotation-items-pane'); + return !!this.closest('merge-pane, scaffold-item-preview, annotation-items-pane, search-results-pane'); } get _disableContextMenu() { diff --git a/chrome/content/zotero/elements/itemDetails.js b/chrome/content/zotero/elements/itemDetails.js index 2e9a705f0a..67f9a8784d 100644 --- a/chrome/content/zotero/elements/itemDetails.js +++ b/chrome/content/zotero/elements/itemDetails.js @@ -55,6 +55,9 @@ + +