A root-level 'bestMatch' condition -- serialized as a marker like
joinMode and resultLevel -- ranks results with the Relevance column via
a "Sort results by best match for" field below the root group's
conditions, composing semantic ranking with Boolean filtering. A
best-match quick search converts to it via the Advanced Search button,
and a selected saved search's own marker activates ranking too,
overridden by an active best-match quick search.
Without a cutoff the condition is rank-only and membership is untouched
-- including while the index is unavailable -- so a saved search acts
identically as a source, and unscoreable items just sort last. With the
optional "keeping top" cutoff, carried in the marker's operator,
membership becomes the K most similar results, applied in search() so
scopes, counts, and the API see the same set. A transient search's
cutoff, which applies uniformly to every selected row, is reapplied
over the merged results so a multi-collection selection returns K
members total; a saved search's cutoff is part of its own membership
and never trims other selected rows. The query embedding is cached
in-flight, so the per-row membership passes and the merged ranking
share one worker embed.
Rows kept identifying the active mode as bestMatch after the model was
disabled in the preferences, leaving the active filter returning nothing.
Observe the model pref, switch to Fields & Tags, and rerun any active
search.
download() skipped every existing file, so after a revision bump the old
files were kept and then marked as the new revision. Stamp the directory
with the revision being downloaded and clear it on mismatch, so bumps
replace the files while interrupted downloads of the same revision still
resume.
Refreshes are serialized, so with scoring slower than the quick-search
debounce, intermediate queries queued instead of becoming obsolete. Bump
a generation counter when a filter or the selected rows change and check
it between scoring chunks, abandoning the stale pass and leaving the
rows for the newer refresh to replace.
New, changed, or removed vectors now rerank an active best-match
search: the indexer announces committed batches -- and index clears from
disabling or a model switch -- with a coalesced 'refresh' item event,
and the items view reruns the search for it. Batch writes also skip
items deleted while the batch was embedding, so a delete can't be undone
by an in-flight batch, and the delete notifier is awaited.
A search could embed the query with one model's prefix on a worker
initialized for another and compare it against vectors from a third.
Tag the worker with the model version it was initialized for, make
scoring wait out an in-progress model switch, refuse to score an index
that wasn't stamped by the active model, and discard results if the
model changes mid-scoring. The items view treats a not-ready index as
an empty result rather than showing an unranked scope.
Semantic similarity has no natural relevance threshold, so instead of
asking the user to pick an arbitrary result count, show every scored
item and surface the ranking directly: a Relevance column appears and
becomes the sort while a best-match search is active, and the previous
sort and columns return when it clears. The merged results are scored
in a single pass in the row provider, so ranks are global across a
multi-collection selection, child items (attachments, notes,
annotations) rank via their top-level item, and equal scores get equal
ranks that order deterministically via the secondary sort fields. Items
without a stored embedding are filtered out.
Each cell renders the score's position within the model's display range
as a bar, so relevant results read as full and the irrelevant tail
reads as empty. The ranges are provisional per-model display constants.
Sorting uses the ranks, which are also exposed to assistive technology
and as the cell tooltip. On a focused selected row the bar
switches to white so the fill doesn't vanish into the accent selection
background.
The embeddings are a local, rebuildable, model-specific index, so they
don't belong in the main database or its backups. Follow the full-text
content index pattern: a lazily attached embeddings.sqlite versioned via
PRAGMA user_version, tied to the main database by localUserKey, with
corruption recovery and idle-maintenance vacuuming via the DBConnection
hooks. Since a cross-database foreign key isn't possible, item deletions
now clear embeddings via the notifier, and the indexed-model identity
moves from a pref into the database's meta table.
"bestMatch" names what the condition does -- rank results by how well
they match the text -- rather than the current scoring mechanism, so
the name can account for later changes to how it works. Rename the
condition-facing identifiers with it; the embeddings engine keeps its
similarity vocabulary.
- added environment to run embedding models locally
(transformers.js, ONNX Runtime WASM binary, etc.). The actual
inference execution happens in a separate worker environment (worker.js)
- added local itemEmbeddings table to store embeddings locally
- in advanced preferences, one can select two options for
semantic search model: english and multilingual. English model
(bge-small-en-v1.5) is better for english-only corpus
but multilingual (multilingual-e5-small) is necessary to handle
abstracts with any other language than english. We can add
more language-specific models as needed.
- when the model is selected, Zotero.Embeddings.download
will download the model (quantized ~100mb) and store it locally.
- Zotero.Embeddings.Indexing will start a process to
index all regular items with title+abstract. It happens in batches
and takes some time. The progress will appear in the
advanced preferences pane. Embeddings are inserted
into itemEmbeddings SQL table. For now, the table is local
only, no syncing is involved.
- when embedding model pref is set to "Disabled", the model
is deleted and embeddings table is cleared.
- when an embedding model is selected, quick search dropdown
has a new "Similarity" mode, which will run semantic search
on the current scope of items.
- semantic search does not clearly define what counts
as "relevant" and what is "not relevant". In addition,
it will change depending on the library and query. So
we cannot semantically filter out items the way
it is done via SQL. Semantic search returns the ranking
but items cannot be sorted because it is done by the itemTree
based on column selection.
So in "similarity" quicksearch mode, there is also a dropdown
to select how many top relevant items to keep (top 5 - top 100).
It allows the user to keep the most relevant items depending
on the context, without conflicting with itemTree sorting.
- semantic search happens in-memory. On a large 5K library
it's fast, but we could consider sqlite-vec extension if
needed.
MDPI serves a JS proof-of-work interstitial in place of the article
page, so Find Full Text found nothing. Run the challenge in a hidden
browser when a page's meta refresh points to a registered challenge
host, then retry the page.
Also match meta refresh URL= case-insensitively, since MDPI's is
uppercase, and bound meta refreshes by the redirect limit.
https://forums.zotero.org/discussion/132837/https://forums.zotero.org/discussion/134079/
An hour of retries makes sense for syncing, which runs automatically and
can just spin during server maintenance rather than showing errors that
send people to the forums, but it was also inherited by foreground
requests, where it stalled operations the user was waiting on. Sync and
file syncing now ask for the long window explicitly.
A file URL returning a server error was retried for up to an hour inside
the download, and a 429 or Retry-After was waited out there too. Since
the queue processes one item at a time, that blocked every other
selected item. Throttling now goes to Find Full Text's own per-domain
handling, as it did before 10.0, which also needed to read Retry-After
from a fetch Response and parse HTTP-date values.
https://forums.zotero.org/discussion/133703/
Retry-After was honored unconditionally with no attempt limit, so a server
returning 429 or 503 with the header on every request retried forever.
Retry-After waits and backoff intervals now share the errorDelayMax
budget, with each Retry-After counted as at least a second so that a
value of 0 can't loop forever.
This runs during Find Full Text and connector saves, where the default
30-second timeout and hour of 5xx retries are far longer than a user
wants to wait.
A custom resolver returning a 429/5xx inherited Zotero.HTTP's default
retry policy, which could cause the whole queue to stall for up to an
hour -- or indefinitely for a Retry-After -- on a dead resolver, despite
the 5-second timeout on the request.
https://forums.zotero.org/discussion/133642/
Since the switch to fetch() in 0fe31b0f04 (Zotero 10.0.0), download
requests didn't use cookies, and the Referer header was silently dropped
as a forbidden header. Sites that check either -- e.g., IEEE Xplore,
which returns a 502 -- failed during Find Full Text.
https://forums.zotero.org/discussion/133703/
As of Firefox 153.3.0esr, Services.scriptloader refuses jar:file: and file:
URIs unless allowUnsafeURL is passed (Mozilla bug 1974213), so no plugin
loads: bootstrap.js fails with "Trying to load untrusted URI", then
"Plugin ... is missing bootstrap method 'startup'".
Pass allowUnsafeURL when loading a plugin's bootstrap.js and prefs.js,
and preference pane scripts, and temporarily enable
security.allow_unsafe_subscript_loads, which covers loads from plugin
code outside the bootstrap scope.
Also add a test that installs a fixture plugin that loads a script from its XPI
and sets a default pref.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
Switching to a view that has the same tags but with different types
(manual vs. automatic) kept the previous list, so with "Show Automatic"
off a tag could show or hide incorrectly.
If a view contained manual and automatic tags with the same name, only
one was kept, and it would disappear with "Show Automatic" off.
https://forums.zotero.org/discussion/133869/
As of Firefox 140.17/153.4 (Bug 2068406), setting nsIFilePicker's
displayDirectory to a directory that doesn't exist or isn't readable
throws instead of being ignored. This broke callers that pass a saved
path that may be stale (e.g., choosing a PDF/EPUB handler after the old
app's folder was removed, or a missing Scaffold translators directory).
The action column's width was given as "32px", which VirtualizedTable
now turns into an invalid CSS width, collapsing the column and hiding
the buttons needed to resolve unmapped and ambiguous citations. Also
restore the accept-match icon on ambiguous-citation candidates.
https://forums.zotero.org/discussion/133740/
A condition that matched a child item in the trash rolled up to its
parent when the search had a top-level (or other ancestor) result level,
or when "Include parent and child items of matching items" was checked.
https://forums.zotero.org/discussion/133836/
A standalone PDF recognized by ISBN wasn't moved under the new parent
item, because the move was attempted before the item had been saved
and had an ID.
https://forums.zotero.org/discussion/133957/
A userdata upgrade that took more than 5 minutes (e.g., dropping a large
legacy word index) was rolled back by Sqlite.sys.mjs, while its remaining
statements autocommitted and marked the database as upgraded, leaving
steps 126 and 127 missing.
Disable the transaction timeout and replace steps 122-129 (added in
Zotero 7, 9, and 10) with step 130, which checks for each change before
making it. To speed up the upgrade, drop the legacy word index without
overwriting the freed pages, since the full text already exists on disk
and we're just rewriting it in fulltext.sqlite.
https://forums.zotero.org/discussion/133859/zotero-connector-in-chrome-not-workinghttps://forums.zotero.org/discussion/133965/citation-saving-does-not-work-with-zotero-connector-firefox
Firefox has AppKit draw the menulist's capsule and chevron, but it sizes
the box and positions the label itself, using a fixed dropdown border
and the label's CSS margins. On macOS 27 (and probably 26), that left the
label closer to the edges and the chevron than in a native NSPopUpButton.
It also made the box 26.5px tall. Firefox assumes regular pop-up buttons
are 22px tall and only draws them unscaled in boxes up to 2px taller, so
it drew the control at 22px and scaled the image up, enlarging the
capsule and chevron.
Add padding to match native label insets, and trim the label's vertical
margins to keep the box at 24px, so Firefox draws the control unscaled.
The result measures the same as a native NSPopUpButton.
Build the launcher with the macOS 26.5 SDK. AppKit chooses control metrics
based on the main executable's SDK, so with the old 15.5 launcher, native
buttons and menulists used pre-Tahoe metrics with built-in margins while
Mozilla's XUL (built with 26.5) stopped compensating for them on macOS 26+
(bug 1992898), leaving labels cramped and controls indented.
Also add custom libmozglue.dylib (#6056) and set source info in mozconfig,
which official builds require and which Mozilla only detects automatically
from Mercurial checkouts.
Mozilla's helper apps (GPU, content, etc.) use the hardened runtime and
Mozilla's Team ID, so in unsigned builds they couldn't load our custom
libmozglue.dylib and failed to launch. Re-sign them ad hoc.
Firefox stopped inflating native controls on Tahoe (bug 1992898), but
still uses the old widget border sizes, so labels nearly touch the edges
of the new capsule-shaped buttons and run into menulist arrows.
curl is already required, so drop wget as a build requirement.
build_autoupdate.sh now checks the HTTP status code for its ETag cache
instead of relying on wget not creating the file on a 304.
If NOTARIZATION_PROFILE is set, notarize_mac_app and notarization_info
authenticate with that stored notarytool profile (e.g., an App Store
Connect API key) instead of an Apple ID and app-specific password,
unlocking the keychain first if needed.
153.3.0esr updated the Chromium sandbox, changing the TargetConfig
interface that xul.dll calls into the launcher's sandbox broker
through. The stubs were still built from 153.0esr, so the 153.3.0esr
xul.dll called the wrong methods and crashed on startup with
"config->SetProcessMitigations(initialMitigations) failed".
The release script never set SAFARI_APP_EXTENSION, so 10.0 through
10.0.3 shipped with only the web extension, which doesn't run on Big Sur
or Monterey. Only beta builds have been including both.
https://forums.zotero.org/discussion/133871/
This showed up as devtools failing to start in a -d build on some machines, but any Subprocess.call() could fail because of it.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>