To override default hybrid search mode of best match
in case one wants to run just semantic or just
lexical search.
It replaces old temp pref set in settings panel.
Best Match becomes a ranked search over everything in a library: its
items' metadata, notes and annotations, and the full text of its
attachments. A lexical engine (Zotero.Lexical) always ranks; with
semantic search enabled, the embedding engine ranks too and the two are
fused, so an item can match by its words, by its meaning, or -- ranking
highest -- by both. Matched passages are shown under their items, read
in the item pane, and opened in the reader at the passage. Attachment
vectors can come from the dataserver instead of being computed here.
1. Lexical engine (Zotero.Lexical, fulltext.js)
A word-level FTS5 index, ftindex.fulltextItemText (plus a CJK 2-gram
twin and a state table), holds each item's searchable text in per-type
columns: title and abstract for regular items, note text for notes, the
marked passage and comment for annotations. Regular items and
annotations index inline in their save transaction; notes are written
by the existing stale-flag queue alongside the trigram tables; a
backfill queue covers pre-existing items and joins the startup and
background drains. The trigram index answers substrings and inflates
counts, so it can't tell "fall" from "rainfall" or weigh a word by its
rarity; word-level FTS answers both with index probes.
Ranking is FTS5's BM25 over that index and the existing content index:
a query parses into word, phrase and CJK-run terms joined by OR, so a
document missing a word still ranks below one that has them all; term
weight comes from rarity in the user's own library, with no stoplist.
Consecutive words are added as phrase terms, item-text columns are
weighted (title 6, abstract 4, annotation 2, note 1), and an
attachment's full text counts only where enough of the query's terms
occur within 200 tokens of each other (FTS5 NEAR; 75% of them, so all
of them up to three), so one rare word can't carry a long document to
the top. Scores are divided by the most the expression could earn, so
both indexes report the 0-1 share of the query a document carries.
2. Fusion (Zotero.BestMatch)
Both engines score every candidate and their rankings are fused with
Reciprocal Rank Fusion; results are the union of the engines' matches.
Each engine's tail is cut against its own strongest match before fusion
(search.bestMatchMargin, percent, default 50, in the Advanced pane): a
library on one subject needs a tighter margin to separate a specific
answer from the field, a varied one hardly needs it. The Relevance bar
shows the item's strongest single piece of evidence, whichever engine
found it.
While the semantic index is still building, or semantic search is off,
Best Match degrades to lexical ranking rather than showing nothing, so
the quick-search mode is always offered and the "index not ready" state
is gone. Advanced Search's bestMatch condition scores through the same
facade. A temporary pref, search.bestMatchEngine (hybrid, lexical,
semantic), selects the engine for testing.
3. Attachments, structured text and chunks (Zotero.SDT)
Attachments are indexed on the structured text the document worker
extracts (PDF, EPUB, snapshot), cached as a pack per attachment. How a
pack is cut into chunks is moved to the structured-document-text
module. The chunker lives there so that the client, the dataserver's
indexer and anything else that embeds a document cut it the same way
and produce rows that can be exchanged. Each chunk carries an anchor --
page rects for PDFs, selectors for EPUBs and snapshots -- that locates
its text in the file across extractor versions, so a row cut elsewhere
stays usable when a local re-cut would block differently. Anchors are
stored as deflated JSON; chunk text is not stored, it's read back from
the pack.
The document-worker build bundles the chunker next to the pack reader
(structured-document-text.js and structured-document-text-chunker.js),
and Zotero requires both; on the main thread the chunker only cuts
plain text and reports its version. A large book takes hundreds of ms
to inflate and cut, so cutting (sdt.getChunks) and reading anchors back
with their reader positions (sdt.readAnchors) run in the document
worker, with no fallback here. A worker failure, or a chunker version
that disagrees with the bundled one, comes back as a failed cut rather
than a verdict on the attachment. On a worker error, the document
worker manager fails every pending request and starts a fresh worker
for the next one, instead of leaving the queue stuck behind an
unanswered request. MIN_CHUNKER_VERSION forces a re-cut of attachments
cut by an older chunker.
4. Model, vectors and calibration
One model, bekko-embedding-v1-a25m: multilingual, 8192-token window,
Matryoshka-trained so its 384-dimensional output is cut to 256. Two
runtime tokenizer defects are corrected -- the leading word marker the
runtime fails to add, and whitespace runs tokenized unlike the
reference; the model `revision` tracks changes that alter vectors.
Stored vectors are centered on the model's mean and quantized to int8
(the shared mean would otherwise spend most of the 8 bits), 256 bytes a
row, and scored by cosine in SQL. The mean, the score floor and the
bar's ceiling are measured by Calibration.record() over a corpus of
triples -- a query, a passage that answers it, a near miss from the same
field (resource/embeddings-calibration-corpus.json) -- and pasted into
the model config; the floor sits at the 95th percentile of near-miss
scores, the ceiling at the median of matches.
5. Indexing runs (Zotero.Embeddings.Indexing)
Nothing is queued. An item change or a finished sync kicks a run, and a
kick that lands during a run makes it go again. A run first indexes
locally the items, notes and annotations saved since the index last
looked at them -- a stamp other than the item's clientDateModified, or
a save in the last 5 s -- reading their text from the tables, then
takes every eligible attachment through five steps. Each step is one
pass over the outstanding attachments, 50 at a time by itemID as
full-text sync pages, and finds its own work in the index, so a run cut
short resumes where it was:
- reconcile: what's stored still holds. A file that's gone, or rows cut
by a chunker before MIN_CHUNKER_VERSION, lose everything;
- extract: every attachment's text cached as a pack, the server's
included, so no preview waits on an extraction.
- fetch: the server is asked for everything that's its.
- cut: this client's attachments are cut into pending rows with anchors.
An attachment the worker couldn't cut is left as it is for a later
run, since a plain-text fallback would stand in for the file for good;
only an extraction failure falls back to plain text. Rows adopted for
a file that has since arrived wait the same way.
- embed: every pending row gets a vector, pooled across attachments and
pages.
The attachment work gives way to a kick, so a just-edited item is
searchable without waiting behind the library's documents, and to a
sync in progress. Items are loaded without caching and never more than
a page at a time. Startup and Resume clear the attachment stamps for a
full pass. Text too short to say anything is skipped: under two words
of title and abstract, under three of anything else.
The index is one table of rows (itemID, chunkIndex, embedding, anchor)
and one of per-item state (itemIndexState: sourceKey, contentHash,
extractor, clientDateModified, and the item's standing with the server);
counts are derived from rows. What a run is built on is declared beside
it: Sources (what each kind of item offers and what text it yields),
Progress (what the preferences pane reads, recomputed on a clock),
Store (the only code that writes the tables), Runtime (the machine's
say: a token budget per engine call halved under memory pressure, a
memory floor to start at all, the engine restarted to give memory back,
half the optimal thread count unless the prefs pane is open or the
system idle, and main-thread work paced against idle time), and
Diagnostics (rates and shape summaries of what Indexing records). The
engine, which the runtime terminates when idle, is checked and
recreated before each use.
6. Sync client (Zotero.Embeddings.Sync)
Available when the account syncs (pref embeddings.sync.enabled is an
override) and no endpoint is active -- configuring one says to embed
there instead. The server is asked about a stored file in a library
syncing with Zotero Storage, in sync or to download, that it hasn't
declined. Rows arriving before their file are adopted when it does.
GET users/{userID}/embeddings?model=M&itemKey=K1,K2,... returns
{ model, items: [
{ key, status: 'success', contentHash, chunks, version,
rows: [{ chunkIndex, embedding, anchor }] },
{ key, status: 'declined' },
{ key, status: 'pending' } ] }
- success: rows replace what's stored, if contentHash matches the file
here (or its plain text), rows number exactly `chunks`, and each row
is well formed; otherwise declined. Rows already held from `version`
are kept.
- declined: cut and embedded locally in the same run.
- pending, or a key missing from the reply: left as it is and asked
again next run.
- A reply in another model declines the batch. A failed request stops
the pipeline; the run goes back to the server after 5 minutes, or at
once on Resume.
GET users/{userID}/embeddings?format=versions&model=M&since=V (with
If-Modified-Since-Version) returns { version, items: { key: version },
models }. After every sync, checkLibrary() asks for what changed since
the version it recorded and marks every key whose rows aren't from the
version named to be fetched again, declined ones included. The
library's version is recorded last, so a failure repeats the delta.
Server is expected to provide raw vector (not centered, not quantized)
so it does not have to worry about mean vector config.
7. Remote endpoint (Zotero.Embeddings.Endpoint)
Passages can be embedded by a server serving the same model -- a local
llama.cpp server in the OpenAI format, or Text Embeddings Inference --
configured from Settings -> Advanced. The server is trusted only
because its vectors match the local model's, never by name: verify()
embeds fixed texts both ways and stores a typed verdict (ok,
unreachable, unauthorized, not-embeddings, width-mismatch,
low-agreement, context-too-small) keyed to the URL, model version and
format. Every batch carries a sentinel text whose local vector is
cached, so a server switched to another model or pooling is caught on
that batch; three consecutive failures skip the endpoint for the rest
of the run. Queries always embed locally. The Configure Endpoint dialog
shows the facts the server must match and a copyable llama.cpp command.
8. Previews (Zotero.BestMatch.Session, item tree, item pane)
A Session scores a query and owns the passages its results matched in.
The chunk is the unit of a match for both engines: passages come from
the index's own rows read back by anchor, or, for an unindexed item,
from the same cut applied to its structured text (only where already
extracted; generating it costs seconds) or its plain text. An item
returns at most three quoted matches, each blending the model's score
with how much of the query the passage's own words carry. The quoted
line is chosen after ranking: the sentence the lexical engine picks
where the passage says the query's words; where it only means them,
the sentence a static multilingual model (potion-multilingual-128M, via
the runtime's static-embeddings backend) finds closest -- weaker than a
dense model, far faster, and never stored.
The item tree shows matches as two-line child rows -- where the passage
is (section path, page) and the line worth reading, with the query's
words marked -- and a new query shows its results from the top. The ten
best-ranked items' previews are derived before scoring resolves; the
rest are derived while the main thread is idle, paced against what each
costs, and arrive in batches. A changed item's previews are invalidated
and re-derived; the tree is no longer refreshed as embedding progresses,
which kept freezing it mid-search. Selecting an item shows every
passage it matched in a Search Results section of the item pane;
selecting match rows shows the passages themselves, grouped by
attachment, in a pane of their own. Double-click or Enter opens the
attachment at the passage, a PDF scrolled to and highlighting the
anchor's rects.
9. Preferences
The model menu is replaced by one switch, search.bestMatch.enableSemantic
(off: lexical ranking only, nothing indexed, what's indexed kept), and
the mode-change confirmations go with it. The pane shows two progress
bars -- metadata, notes and annotations; attachments -- with the step
under way ("Preparing documents", "Syncing semantic data… X / Y",
"Generating semantic data locally for N items"), the endpoint's status
and its Configure dialog, the quality cutoff, and a diagnostics panel,
hidden by default: throughput, inference speed, padding efficiency,
batches, engine threads and restarts, process memory and CPU, chunk
size distributions.
10. Build and dependencies
The document-worker submodule gains the sdt.getChunks and
sdt.readAnchors actions and builds the chunker as a second bundle next
to the pack reader. Zotero.ML allows the static-embeddings backend and
Mozilla's model hub. The embeddings database is at version 14 and is
rebuilt on upgrade.
Semantic similarity has no natural relevance threshold, so instead of
asking the user to pick an arbitrary result count, show every scored
item and surface the ranking directly: a Relevance column appears and
becomes the sort while a best-match search is active, and the previous
sort and columns return when it clears. The merged results are scored
in a single pass in the row provider, so ranks are global across a
multi-collection selection, child items (attachments, notes,
annotations) rank via their top-level item, and equal scores get equal
ranks that order deterministically via the secondary sort fields. Items
without a stored embedding are filtered out.
Each cell renders the score's position within the model's display range
as a bar, so relevant results read as full and the irrelevant tail
reads as empty. The ranges are provisional per-model display constants.
Sorting uses the ranks, which are also exposed to assistive technology
and as the cell tooltip. On a focused selected row the bar
switches to white so the fill doesn't vanish into the accent selection
background.
The embeddings are a local, rebuildable, model-specific index, so they
don't belong in the main database or its backups. Follow the full-text
content index pattern: a lazily attached embeddings.sqlite versioned via
PRAGMA user_version, tied to the main database by localUserKey, with
corruption recovery and idle-maintenance vacuuming via the DBConnection
hooks. Since a cross-database foreign key isn't possible, item deletions
now clear embeddings via the notifier, and the indexed-model identity
moves from a pref into the database's meta table.
- added environment to run embedding models locally
(transformers.js, ONNX Runtime WASM binary, etc.). The actual
inference execution happens in a separate worker environment (worker.js)
- added local itemEmbeddings table to store embeddings locally
- in advanced preferences, one can select two options for
semantic search model: english and multilingual. English model
(bge-small-en-v1.5) is better for english-only corpus
but multilingual (multilingual-e5-small) is necessary to handle
abstracts with any other language than english. We can add
more language-specific models as needed.
- when the model is selected, Zotero.Embeddings.download
will download the model (quantized ~100mb) and store it locally.
- Zotero.Embeddings.Indexing will start a process to
index all regular items with title+abstract. It happens in batches
and takes some time. The progress will appear in the
advanced preferences pane. Embeddings are inserted
into itemEmbeddings SQL table. For now, the table is local
only, no syncing is involved.
- when embedding model pref is set to "Disabled", the model
is deleted and embeddings table is cleared.
- when an embedding model is selected, quick search dropdown
has a new "Similarity" mode, which will run semantic search
on the current scope of items.
- semantic search does not clearly define what counts
as "relevant" and what is "not relevant". In addition,
it will change depending on the library and query. So
we cannot semantically filter out items the way
it is done via SQL. Semantic search returns the ranking
but items cannot be sorted because it is done by the itemTree
based on column selection.
So in "similarity" quicksearch mode, there is also a dropdown
to select how many top relevant items to keep (top 5 - top 100).
It allows the user to keep the most relevant items depending
on the context, without conflicting with itemTree sorting.
- semantic search happens in-memory. On a large 5K library
it's fast, but we could consider sqlite-vec extension if
needed.
citeproc-rs is no longer maintained, and people who had enabled the
hidden pref were hitting errors.
This also drops the free() call on CSL engines, which only existed to
free the citeproc-rs wasm driver.
https://forums.zotero.org/discussion/133515/
Translators are evaluated with the system principal, which worked until
now only because nsContentSecurityUtils::IsEvalAllowed() exempted any
profile with "JS hacks" present -- which for us meant
xpinstall.signatures.required being false. Bug 2038660 dropped that
exemption, so ask for eval explicitly.
Bug 2043845 put the FORCE_ALLOWED_DTD content policy type behind a pref
that's off for everything but Thunderbird, so NS_NewChannel() refused the
load with NS_ERROR_CONTENT_BLOCKED and the XML parser silently skipped the
entity. Custom elements whose markup uses DTD entities then failed to
parse with "not well-formed XML".
Bug 2038660 applies "script-src chrome: resource: moz-src:" to every
chrome: document, which blocks inline <script>s and inline event
handlers, so the main window loaded but ran none of its scripts.
We should move our inline scripts and event handlers into separate files
so that we can drop the pref.
After the first bubble is added, the focused input gets a placeholder
indicating that typing a number will add it as a page to the just-added
bubble. The placeholder is truncated if it's too close to the edge in
multi-item citations.
Also add a tip to the item details popup explaining that locators can
be typed into the main input field, with a link to the documentation.
The tip stops appearing once a typed locator has been used.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
In Add/Edit Citation mode, display a preview of the citation
in the bottom section. The section can be hidden/displayed
via the toggle in the right corner.
Remove io.preview from editor instance, so that citation
dialog knows not to show the preview even if the preference
is set.
A minor refactor to have resizeWindow() resolve
when the animation is fully over, and clear minHeight
on window in list mode before resizing, restoring
it when resizing animation is done, same as in library mode.
It allows us to fully expand the window in list mode before
showing the preview.
Fixes: zotero#5910
- Switch journal mode from DELETE to WAL for better write performance.
With EXCLUSIVE locking mode, SQLite uses heap memory for the WAL
index, avoiding an -shm file. Set synchronous=NORMAL (matching what
Mozilla uses for Places). Checkpoint WAL on database close so the
.sqlite file has all data (for copies or backups).
- Add periodic database compaction on idle (after DB backup) using
VACUUM INTO and do an atomic file swap back to zotero.sqlite if no
writes occurred during the operation. Check if vacuuming is needed
based on time interval (default 14 days) and freelist ratio (default
10% threshold).
- Disable auto_vacuum, which causes fragmentation and is unnecessary
with periodic VACUUM
- Remove the VACUUM call from the integrity check, which was always just
an awkward hack to let people trigger a VACUUM without having an
explicit button
Closes#652
Plus guidance-panel changes:
- Fix description not updating when multiple panels exist
in the document
- Fix nonfunctional noautohide attribute
- Show "Got It" button for noautohide with no navigation
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
- last-closed -> last-used for citationDialogMode pref value
- pref label citationDialogLastClosedMode -> citationDialogLastUsedMode
- update Zotero.Prefs version. While migrating, if
integration.useClassicAddCitationDialog is true, set
integration.citationDialogLastUsedMode to 'library'
so the citation dialog opens in library mode on the first run.
If citationDialogLastClosedMode pref exists, it is
migrated to citationDialogLastUsedMode and cleared
- redone logic of extracting the locator to allow for better
detection of locators and to not rely on long regular
expressions that are hard to troubleshoot. Added tests for it.
These are the new rules for extracting locators:
- locator value can be either a set of numbers potentially
with some punctuation in between OR text surrounded with
single or double quotes
- locator labels can be typed in any of 3 formats:
full (chapter), short (chap.), or short with no
punctuation (chap)
- locator labels go before the locator values and there
may or may not be a space in between. For example,
chap.10 is the same as chap. 10.
Fixes: #5092Fixes: #5093
Implemented redesigned citation dialog with library
and list modes one can switch between. This dialog is
a direct replacement of quickFormat and classic citation
dialogs.
In Library Mode:
- items table has a new + button column to add items from the
citation and rows of selected items are highlighted
- open, cited and selected items appear in a section
between the items table and bubbleInput. When there are no
matches, a message is shown.
- only top-level items are shown when citing
items and only notes/notes' parents - when adding a note
- selected items are gathered in a collapsible deck to save space.
Click on the deck will expand it. All selected items can be
added via "Add all" button.
- when an item is added, bubble-input may increase in height and
push itemTree lower. To try to preserve relative positioning
of the mouse, itemTree will scroll to be over the row that was just clicked
In List mode:
- arrow up/down from the input will change the selected item
with the focus remaining in the input
- selected items are a collapsible list section
Other behaviors and fixes:
- one can add any locator (not just pages) by typing its full or short name
in an input and pressing Enter (e.g. line 10, or l. 10, or chap. "test chapter")
- Added a new preference to select if the citation dialog
should always open in list mode, library mode or in the last
mode that was used
- bubbles whose items are selected in library or list mode
are highlighted
- arrow up/down from a bubble will focus the bubble
above/below it for easier navigation across bubble-input
- multi-select is supported on items in list mode or item cards in
library mode via Shift-arrow or Cmd+click to select multiple items
- Cmd/Ctrl + Enter will always accept the dialog no matter what
is focused
- when there are no bubbles, accept button is disabled
- after an item is added, bubble-input is always refocused
- added suppressed property to itemTreeMenubar to be able to
hide if in list mode, where it is not relevant. Fixed menubar
getting stuck or re-appearing on Alt keypress on Linux by setting
height: 0 vs hiding it via hidden.
- added initialFolder and onActivate prop to collectionTree to set which collection
should be reopened when dialog opens and to be able to set a custom
onActivate handler.
- added getExtraField prop to itemTree to get data on if
an item is in a citation or not
Implementation details:
- citationDialog.js is the main file. It relies on a number of
helper files in citationDialog/* directory to keep the main
file less cluttered. popupHandler.js contains the logic of
opening/closing the item details popup to add locator/prefix/suffix/etc.
keyboardHandler.js is responsible for overall keyboard navigation
throughout the dialog. searchHandler.js contains the logic
for running the search based on user's query. Finally, Helpers.js
has general helper functions that don't handle any actual logic.
- SearchHandler is set to run search in two ways: for cited/selected/open
items (which are cached and do not involve any actual SQL query) and
general search for items across all libraries. When layout.search
runs, firstly cited/selected/open items are updated after which the
SQL search runs.
- bubbleInput.js is a customElement responsible for bubbles interface.
bubbleInput.refresh takes a list of items and handles adding/removing/reordering
of bubbles as needed. It always has two inputs on each side of
a bubble, as opposed to having inputs inserted dynamically. It
allows keyboardHandler.js to handle navigation with arrows within
bubbleInput.
- when user interacts with bubbleInput, it emits custom events that are
handled by IOManager singleton in citationDialog.js, which
may update the items information and pass them back to bubbleInput.refresh
to have the list of bubbles updated
- CitationDataManager singleton is responsible for storing
items added into the citation in CitationDataManager.items
in an object with both Zotero.Item and the actual citation item.
It is easier to pass both items to other components
and helpers, as opposed to sharing functions to convert
items back and forth.
- Linked remaining found inputs/menulists to their labels.
- Made the "Choose resolver" a proper visible
label, so that it does not oddly disappear if you click
on the dropdown with "Custom" showing after even if the
selection did not change. For the purpose of VPAT, a
visible label is always good for success criteria 3.3.2
https://www.w3.org/WAI/WCAG21/Understanding/labels-or-instructions
- Added explicit names to +/- buttons
- aria-labelledBy for inputs surrounded by text
- fluent strings for resolver preferences
- remove openURL.version preference, hardcode "1.0"