Rerunning itemTree refresh on embedding progress takes
a good amount of work on the main thread, which can
lead to frozen UI if there is an active search when
embedding runs.
Refreshing the itemTree during a search is likely not
necesary because it is a bit jarring to see the row order
change all of a sudden as you scroll down the itemTree, and
lexical search is a decent fallback until indexing
is ready.
BestMatch session extracts only a small portion of
previews during scoring to show immediately. The rest
is derived when the browser is idle to not freeze the UI
and sent to the itemTree via onPreviewsFilled callback.
Similar approach to earlier placeholder rows - but
better performing.
Instead of by global frequency. So if you search "size of a dinosaur"
chunks returned from a paper about dinosaurs use the weights
of "size" and "dinosaur" from within the paper, where "size"
is much rarer and more useful. Otherwise, chunks that
meantion "dinosaur" a lot are returned.
Firstly, it did not always perform well with itemTree scroll.
More importantly, semantic search does take some time -
can be 10+ seconds on a large library. A few extra seconds
to fetch all snippets should not be a big problem. And if
the tail of search results is so long that it becomes -
the fix is to trim the number of search results by
stricter relevance criteria.
Refactor of best match module to contain all best match-related
logic from collectionViewItemTree, so itemTree just
calls relevant methods when needed.
One drawback is that we can't pick the best sentence from
semantic chunk to use as a snippet because that would
mean re-embedding every chunk's sentences on search. So
instead just show the first sentence if no lexical chunks
are available.
Drop the wildcard from the last search term.
It was added for search-as-you-type and probably not
needed now.
Disabled context menu options on search result rows
fix irrelevant annotations not getting hidden
Added search results pane that is shown when a search
result row is selected. It displays the entire chunk
that the snippet is based on. When an attachment is
selected, its search results section lists all search
results
Render search snippets in itemTree lazily, as the user
scrolls to them. Fulltext table is contentless, so we cannot
fetch snippet() for each search match. For embeddings,
we need to fetch the structured-text to locate the right
block. Both of these operations can take a long time
when done to a lot of items in _refresh before rendering,
which is why search snippets are extracted on demand.
BestMatch.Session is a new object to wrap the interaction
between the item tree and the search engines. BestMatch.Session.score
returns the search results with an indication which
of them should have search snippets. Not all search
results do - purely semantic matches on abstracts or
notes, as well as all matches on annotations get a snippet.
ItemTree renders a placeholder child row for items that will
have snippets.
Based on the matches flag above, the itemTree renders
placeholder rows. When the placeholder row is rendered,
onSearchMatchRendered is called to tell BestMatch.Session
which attachment's snippets need to be shown. BestMatch.Session
maintains a queue and handles extracting of snippets
when the browser is free to avoid freezing the main thread.
When the snippets are extracted, the placeholder row is
replaced with rows of search matches.
BestMatch.Session maintains the state of what snippets were already
extracted.
Drop search result itemPane componenets, on a new search
scroll the itemTree to the top to see the most relevant results.
With semantic search disabled, Best Match runs a purely lexical ranking:
query terms are weighted by their rarity in the user's own library, any term
can match (OR semantics), and items are scored by how much of the query they
cover and where — titles count most, then abstracts, then fulltext, notes, and annotations.
Consecutive query words also count as a unit: an item containing "special education"
as a phrase ranks above one containing the words scattered, with the phrase's
importance again measured by its rarity. Items whose evidence is a single matched
unit that is common are not included (e.g. "fall of communism" -> do no list all
items that just have "fall").
With semantic search enabled, both engines score every candidate and the two rankings
are fused with Reciprocal Rank Fusion: an item can match by its words, by its meaning,
or — ranking highest — by both. Results are the union of the two engines' matches, so
a literal match the model doesn't understand and a paraphrase the words don't catch both surface.
The relevance bar shows the item's strongest single piece of evidence, whichever engine found it.
While the semantic index is still building or switching models, searches degrade
to lexical ranking instead of showing nothing; the "not ready" error state is gone.
Previews (the Search Results section) now explain the actual ranking for any matched item,
not just attachments: excerpts of the item's own text around the literal matches,
with the matched words highlighted. When semantic chunks are available they're shown too — with
literal matches highlighted inside them — and merged with the lexical excerpts by evidence strength,
deduplicating excerpts that show the same passage.
Notes are indexed on their text and annotations on the passage they mark
together with their comment, so both match on what they actually say
rather than on their parent's title and abstract.
A note can hold more text than a model's context window, so
Zotero.Embeddings.Chunking splits long text using the selected model's
tokenizer. Paragraphs are the topic units: two never share a chunk
unless one is too small to embed on its own, in which case small
paragraphs are combined, and a paragraph over the window is split at
sentence boundaries into even pieces. Stored embeddings are keyed by item
and chunk, so the index is rebuilt on upgrade.
An item scores as its best chunk, so a long note
that addresses a query in one paragraph isn't diluted by the rest, and
its rank reflects the best match anywhere beneath it: a strongly matching
annotation lifts its attachment and its paper. The relevance bar reports
only the row's own score, so a paper ranked by its annotation shows a
high rank over an empty bar rather than claiming to be a match it isn't.
List order and bar fill deliberately disagree in that case.
Annotation rows render the relevance cell, and the
Relevance column moves to the far right while a best-match search is
active so the bars line up across item, note, attachment and annotation
rows.
A best-match rerank re-ran the underlying search on every embeddings
update, even though only the ranking depends on the embeddings. Reuse
the cached search results and recompute just the ranking, except when a
top-K cutoff makes membership depend on the scores.
The item tree reran the active best-match search on every 'refresh' item
event, so unrelated bursts (e.g., full-text indexing) triggered a full
re-search and re-score. Flag the embeddings indexer's own notifications
and rerank only for those.
While a best-match search runs against a partially built embeddings
index, results cover only the indexed items and can look arbitrary.
Show the indexing progress in a banner above the items list, updating
as the index fills, so incomplete results aren't mistaken for a
complete ranking.
Covers saved-search ranking and quick-search precedence, mixed
top-K/collection selections, rank-only behavior with an unavailable
index, reranking on indexer refresh events, scoped 'Match any' saves
keeping the marker at the root, numeric-operator round trips, and
concurrent query embeds sharing one worker call.
Semantic similarity has no natural relevance threshold, so instead of
asking the user to pick an arbitrary result count, show every scored
item and surface the ranking directly: a Relevance column appears and
becomes the sort while a best-match search is active, and the previous
sort and columns return when it clears. The merged results are scored
in a single pass in the row provider, so ranks are global across a
multi-collection selection, child items (attachments, notes,
annotations) rank via their top-level item, and equal scores get equal
ranks that order deterministically via the secondary sort fields. Items
without a stored embedding are filtered out.
Each cell renders the score's position within the model's display range
as a bar, so relevant results read as full and the irrelevant tail
reads as empty. The ranges are provisional per-model display constants.
Sorting uses the ranks, which are also exposed to assistive technology
and as the cell tooltip. On a focused selected row the bar
switches to white so the fill doesn't vanish into the accent selection
background.
Closes#5974.
Additional fixes for broken item tree behaviour when multiple items are
selected, and changing focus with ctrl/cmd-arrow keys.
Aligned Collection Tree/Virtualized Tree collapse/expand behaviour when
multiple containers are selected, one of them is focused, and arrow key
left-right is pressed, to the behaviour in Item Tree - now all of them
are collapsed/expanded.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
The view-wide branches (trash, duplicates, feeds, Recently Read) read
the first selected row, which multi-collection selection preserved with
a getter rather than updating. setCollectionTreeRows() now derives the
kind of view the selection adds up to, throwing if the rows disagree,
and those branches test it.
Multi-collection selection left .collectionTreeRow and similar in place
to reduce breakage, but that would just leave plugins and other callers
potentially broken when multiple rows were selected. All singular
getters now throw and say what to use instead. getSelectedLibraryIDs()
was added to replace getSelectedLibraryID().
Collections and saved searches can be selected together, so a
collection-item change called getDescendents() on rows that don't have
it and compared search IDs against collection IDs.
Recently Read rows can span libraries, but read attachments were marked
as matches for the first row's library only, leaving the rest as grayed
context rows.
Code walking the items list assumed every row was an object, so the
headers and spacers shown whenever more than one row is selected got
picked up as items: getSortedItems() passed them to export and report
generation, restoring from the trash called item methods on them, and
the item pane counted them in "N items in this view".
With "Hide Non-Matching Annotations" enabled, an attachment displayed as
empty and non-expandable if no annotations matched the search, so
searching by any non-annotation condition made it impossible to expand
attachments to browse their annotations. Now only hide the non-matching
annotations when the attachment actually has a matching one.
https://forums.zotero.org/discussion/132519/beta-advanced-search-cannot-expand-annotations-of-search-results
When viewing the trash, trashed collections and saved searches were appended
to the items list unconditionally, so every advanced search (and quick search)
in the trash matched all of them. Now they're filtered by name during a quick
search and excluded entirely when an advanced search or tag filter is active,
since they can't match item-level conditions or tags.
Fixes#5957
When multiple collection-list rows are selected, group the combined
items by library under sticky headers (e.g., "My Library", "Group X (2
collections selected)"), or show a single summary header for a multi-row
selection within one library, with blank spacer rows separating
libraries.
When the items list contains items from more than one library, group
them by library -- in collections-list order, independent of the active
sort -- with a section heading above each library's items.
Grouping is triggered automatically by an items list spanning more than
one library, not the kind of selection behind it, so any future source
of multi-library items would be separated the same way. Today the
cross-library collection selection is the only such source.
Allow selecting multiple collections, saved searches, or library roots in the
collection tree -- within a library or across libraries -- and show the union
of their items. The selection is threaded through the pane as an array
(getCollectionTreeRows(), changeCollectionTreeRows(), etc.); the item pane, tag
selector, reports, and export operate on all selected rows.
Adding items (new items and notes, drag-and-drop, the attachment dialog,
import, Add by Identifier) targets every selected collection. Only rows that can
share an items view may be combined: collections, saved searches, and library
roots mix freely, and multiple Recently Read rows can be combined across
libraries, but other special views (Trash, Duplicates, etc.) and rows from
different visibility groups can't be shown together, so a selection mixing them
keeps only the focused row. In-window advanced search runs across all selected
collections.
Advanced search value autocomplete is now scoped to the searched library, fixing
a long-standing TODO where suggestions were drawn from all libraries regardless
of the search scope; for a cross-library selection it spans the selected
libraries.
For a cross-library selection, the tag selector shows the union of tags
(colored tags only when a single library is in scope, since colors are
per-library), and deleting a tag spans all selected libraries while
rename/color/split are disabled. If advanced search is open, the
collection and saved-search conditions are omitted, since each is scoped
to a single library.
In other words, don't show the children that make the parent "recently
read" as context rows.
Children that don't match the Recently Read condition (such as notes,
as well as other child attachments that weren't read since the cutoff)
are still shown as context rows.
And add tests for the new behavior, and clean up an unused local var.
The test opened an advanced search window but never closed it, which
likely caused the intermittent failures in the ZoteroPane focus() test
"should shift-tab across the zotero pane".
Columns now have properties: `enabledIn`, `disabledIn` and `defaultIn`,
corresponding to column picker availability and default visibility. The
properties now filter based on attached collection view type instead of
visibilityGroup.
Visibility groups are for views where we want distinct column sets to
persist, like the feeds view.
Collection type properties are used to specify which columns are
available for a given type, regardless of whether it's in a different
visibility group or not.
Split ItemTree megaclass into:
- ItemTree - concerned with drawing the virtualized table container and
column interaction
- ItemTreeRowProvider - provides rows and issues notifications for
render updates
- ItemTreeRow and subclasses - contains row-specific data and rendering
logic
- CollectionViewItemTree and its accompanying classes - a version of
ItemTree that renders items attached to a given Collection or
CollectionView (CollectionTreeRow).
Various improvements in logic and rendering, separation of concerns.
2026-04-27 14:44:39 -04:00
Renamed from test/tests/itemTreeTest.js (Browse further)