Added search results pane that is shown when a search
result row is selected. It displays the entire chunk
that the snippet is based on. When an attachment is
selected, its search results section lists all search
results
Render search snippets in itemTree lazily, as the user
scrolls to them. Fulltext table is contentless, so we cannot
fetch snippet() for each search match. For embeddings,
we need to fetch the structured-text to locate the right
block. Both of these operations can take a long time
when done to a lot of items in _refresh before rendering,
which is why search snippets are extracted on demand.
BestMatch.Session is a new object to wrap the interaction
between the item tree and the search engines. BestMatch.Session.score
returns the search results with an indication which
of them should have search snippets. Not all search
results do - purely semantic matches on abstracts or
notes, as well as all matches on annotations get a snippet.
ItemTree renders a placeholder child row for items that will
have snippets.
Based on the matches flag above, the itemTree renders
placeholder rows. When the placeholder row is rendered,
onSearchMatchRendered is called to tell BestMatch.Session
which attachment's snippets need to be shown. BestMatch.Session
maintains a queue and handles extracting of snippets
when the browser is free to avoid freezing the main thread.
When the snippets are extracted, the placeholder row is
replaced with rows of search matches.
BestMatch.Session maintains the state of what snippets were already
extracted.
Drop search result itemPane componenets, on a new search
scroll the itemTree to the top to see the most relevant results.
With semantic search disabled, Best Match runs a purely lexical ranking:
query terms are weighted by their rarity in the user's own library, any term
can match (OR semantics), and items are scored by how much of the query they
cover and where — titles count most, then abstracts, then fulltext, notes, and annotations.
Consecutive query words also count as a unit: an item containing "special education"
as a phrase ranks above one containing the words scattered, with the phrase's
importance again measured by its rarity. Items whose evidence is a single matched
unit that is common are not included (e.g. "fall of communism" -> do no list all
items that just have "fall").
With semantic search enabled, both engines score every candidate and the two rankings
are fused with Reciprocal Rank Fusion: an item can match by its words, by its meaning,
or — ranking highest — by both. Results are the union of the two engines' matches, so
a literal match the model doesn't understand and a paraphrase the words don't catch both surface.
The relevance bar shows the item's strongest single piece of evidence, whichever engine found it.
While the semantic index is still building or switching models, searches degrade
to lexical ranking instead of showing nothing; the "not ready" error state is gone.
Previews (the Search Results section) now explain the actual ranking for any matched item,
not just attachments: excerpts of the item's own text around the literal matches,
with the matched words highlighted. When semantic chunks are available they're shown too — with
literal matches highlighted inside them — and merged with the lexical excerpts by evidence strength,
deduplicating excerpts that show the same passage.
Add optional pref to index and search fulltext of attachments.
When enabled, attachment IDs are enqueue after regular items,
notes, and annotations.
Added helpers to extract outline and sections from structured
text module. During indexing, the sections of the attachment
are extracted, large sections are broken into chunks, and
small sections are combined to fit into the context window
of the model. Then, each chunk is embedded with its outline
path as the prefix and added to embeddings table.
Each row now contains full text of the chunk
and it's path - a good amount of duplication needed to
ensure that we can reliably connect the embedding of
the chunk to its text for a preview.
On search in Best Match mode, attachment rows get the score
of the highest ranking chunk, so if there is a very relevant
chunk in an attachment, the regular item with a non-relevant
abstract will still rank highly.
When the attachment row is selected, top 5 matching chunks
appear in the new search results collapsible-section of
the item pane, so one can examine matching chunks
without opening the actual reader.
A root-level 'bestMatch' condition -- serialized as a marker like
joinMode and resultLevel -- ranks results with the Relevance column via
a "Sort results by best match for" field below the root group's
conditions, composing semantic ranking with Boolean filtering. A
best-match quick search converts to it via the Advanced Search button,
and a selected saved search's own marker activates ranking too,
overridden by an active best-match quick search.
Without a cutoff the condition is rank-only and membership is untouched
-- including while the index is unavailable -- so a saved search acts
identically as a source, and unscoreable items just sort last. With the
optional "keeping top" cutoff, carried in the marker's operator,
membership becomes the K most similar results, applied in search() so
scopes, counts, and the API see the same set. A transient search's
cutoff, which applies uniformly to every selected row, is reapplied
over the merged results so a multi-collection selection returns K
members total; a saved search's cutoff is part of its own membership
and never trims other selected rows. The query embedding is cached
in-flight, so the per-row membership passes and the merged ranking
share one worker embed.
Semantic similarity has no natural relevance threshold, so instead of
asking the user to pick an arbitrary result count, show every scored
item and surface the ranking directly: a Relevance column appears and
becomes the sort while a best-match search is active, and the previous
sort and columns return when it clears. The merged results are scored
in a single pass in the row provider, so ranks are global across a
multi-collection selection, child items (attachments, notes,
annotations) rank via their top-level item, and equal scores get equal
ranks that order deterministically via the secondary sort fields. Items
without a stored embedding are filtered out.
Each cell renders the score's position within the model's display range
as a bar, so relevant results read as full and the irrelevant tail
reads as empty. The ranges are provisional per-model display constants.
Sorting uses the ranks, which are also exposed to assistive technology
and as the cell tooltip. On a focused selected row the bar
switches to white so the fill doesn't vanish into the accent selection
background.
- added environment to run embedding models locally
(transformers.js, ONNX Runtime WASM binary, etc.). The actual
inference execution happens in a separate worker environment (worker.js)
- added local itemEmbeddings table to store embeddings locally
- in advanced preferences, one can select two options for
semantic search model: english and multilingual. English model
(bge-small-en-v1.5) is better for english-only corpus
but multilingual (multilingual-e5-small) is necessary to handle
abstracts with any other language than english. We can add
more language-specific models as needed.
- when the model is selected, Zotero.Embeddings.download
will download the model (quantized ~100mb) and store it locally.
- Zotero.Embeddings.Indexing will start a process to
index all regular items with title+abstract. It happens in batches
and takes some time. The progress will appear in the
advanced preferences pane. Embeddings are inserted
into itemEmbeddings SQL table. For now, the table is local
only, no syncing is involved.
- when embedding model pref is set to "Disabled", the model
is deleted and embeddings table is cleared.
- when an embedding model is selected, quick search dropdown
has a new "Similarity" mode, which will run semantic search
on the current scope of items.
- semantic search does not clearly define what counts
as "relevant" and what is "not relevant". In addition,
it will change depending on the library and query. So
we cannot semantically filter out items the way
it is done via SQL. Semantic search returns the ranking
but items cannot be sorted because it is done by the itemTree
based on column selection.
So in "similarity" quicksearch mode, there is also a dropdown
to select how many top relevant items to keep (top 5 - top 100).
It allows the user to keep the most relevant items depending
on the context, without conflicting with itemTree sorting.
- semantic search happens in-memory. On a large 5K library
it's fast, but we could consider sqlite-vec extension if
needed.
Bug 2008041 made disabled, checked, hidden, collapsed, and selected
boolean attributes, so their value is empty and [disabled="true"] no
longer matches. Nothing sets any of them to "false", so matching on
presence alone is equivalent.
Bug 2008041's change to boolean attributes also covers hidden and
collapsed, whose UA selectors became [hidden] and [collapsed], so
setAttribute('hidden', false) now hides the element. Switch the setters
that can be passed a falsy value to toggleAttribute(), read them with
hasAttribute(), and match the [collapsed=true] selectors in our own
stylesheets to the new presence-only form.
- Fix multiple potential scenarios causing a template engine crash
- Add support for specifying string literals in the template engine
- Validate `if/else/elseif` order and matching clause closures
- Validate to ensure every `{{` is properly closed with a matching `}}`
- When a template is invalid, display a warning, do not offer batch-renaming tools, do not update synced setting
- When a template is invalid, prompt the user to fix or reset the template when
Closes#5965
Instead of just prefixing the labels with "-", indent the whole row,
including the icon. This looks better and fixes FAYT on subcollection
names.
https://forums.zotero.org/discussion/132561/
When the item pane is dragged wide, the items pane is squeezed and the
quick search wrapper kept its intrinsic width and overflowed, pushing the
trailing Advanced Search button out under the item pane. Let the wrapper
shrink so the button stays within the pane.
Fixes#5982
- Reword the header as one sentence with a result-level menu ("Find
[attachments] matching [all] of the following:")
- Provide a per-group menu to bind the group's descendant conditions to
the same attachment, note, or annotation (e.g., one annotation that is
both red and contains a given word, not two different ones)
- Show a hint that offers to group ungrouped sibling conditions (e.g.,
two annotation conditions at the top level, to bind them to one
annotation)
- Show a warning when conditions can't combine at the chosen result
level (e.g., an annotation condition with a note result level)
- Remove the two legacy checkboxes:
- "Show top-level items" becomes result level = top-level item and is
migrated on save
- "Include parent and child items", which has no result-level
equivalent, keeps working, stays editable, and round-trips on
searches that already have it, but it isn't offered on new searches
and is removed on save if unchecked
Render the search as a tree of groups: a root group plus nested
search-condition-group elements, each with its own join-mode menu and a
remove control. Each condition row gets a "( )" button that wraps it in
a new group in place, so further conditions can be added to combine with
it under a separate join mode. Switch the builder to rebuild-from-tree --
the DOM is the source of truth and the search's flat conditions (with
groupStart/joinMode/groupEnd markers) are regenerated on each edit, so
the old conditionID-as-index tracking is gone.
macOS-normalize-controls zeros margins on inputs and checkboxes but not
menulists or buttons, so their native platform margins threw off the
spacing. Zero them and restore spacing via the containers' gaps.
https://github.com/zotero/zotero/pull/5658#issuecomment-4696398750
The inner <deck> stacks both panes in one grid cell, so the area was
always sized to the taller (saved-search) pane and never shrank back.
The non-selected pane is now removed from layout entirely.
https://github.com/zotero/zotero/pull/5658#issuecomment-4696398750
After the first bubble is added, the focused input gets a placeholder
indicating that typing a number will add it as a page to the just-added
bubble. The placeholder is truncated if it's too close to the edge in
multi-item citations.
Also add a tip to the item details popup explaining that locators can
be typed into the main input field, with a link to the documentation.
The tip stops appearing once a typed locator has been used.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
Plus guidance-panel changes:
- Fix description not updating when multiple panels exist
in the document
- Fix nonfunctional noautohide attribute
- Show "Got It" button for noautohide with no navigation
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
* Remove `overflow: visible` in RTF Scan to prevent richlistbox from expanding the width and causing clipping
* Introduce small margins as an alternative to prevent focus rings from being clipped in the `wizard`
* Fix "Display as" alignment on Windows
- after a new bubble is added to the citation, it is recorded
as a just-added bubble. The next locator typed without a search query
will go to that item instead of going to the item before where the locator was typed.
Same logic applies when multiple bubbles are added at once.
- the record of just-added bubbles is cleared on focusout
or keypress of an arrow key. That way, it's discarded if the user
is almost certainly not intending to immediately type a locator.
- added a special case to recognize a numeric value as a page locator
if it is typed when just-added bubble is recorded. That special locator
will be added to the just-added bubble as one is typing without
pressing Enter after debounce. Enter will immediately add the locator
without waiting for debounce.
- if a just-added bubble is recorded, cmd-z will clear
whatever numeric locator may have been typed and place
it back into the input, in case one meant to type an
actual search query
- added a special case to recognize ":<number>" as a page locator
in the same circumstances that "page <number>" is currently recognized
- do not use year extraction (SearchHandler._cleanYear)
when parsing input. It strips the first number from
a range of numbers and conflicts with the new locator
logic.
- ensure a bubble with a very long locator does not overflow
- Replace polling with a `ready` promise for initialization
- Use fixed height for the rich list and re-order initialization to avoid layout shifting
- Ensure the "accept" button is disabled until initialization completes