The embeddings are a local, rebuildable, model-specific index, so they
don't belong in the main database or its backups. Follow the full-text
content index pattern: a lazily attached embeddings.sqlite versioned via
PRAGMA user_version, tied to the main database by localUserKey, with
corruption recovery and idle-maintenance vacuuming via the DBConnection
hooks. Since a cross-database foreign key isn't possible, item deletions
now clear embeddings via the notifier, and the indexed-model identity
moves from a pref into the database's meta table.
- added environment to run embedding models locally
(transformers.js, ONNX Runtime WASM binary, etc.). The actual
inference execution happens in a separate worker environment (worker.js)
- added local itemEmbeddings table to store embeddings locally
- in advanced preferences, one can select two options for
semantic search model: english and multilingual. English model
(bge-small-en-v1.5) is better for english-only corpus
but multilingual (multilingual-e5-small) is necessary to handle
abstracts with any other language than english. We can add
more language-specific models as needed.
- when the model is selected, Zotero.Embeddings.download
will download the model (quantized ~100mb) and store it locally.
- Zotero.Embeddings.Indexing will start a process to
index all regular items with title+abstract. It happens in batches
and takes some time. The progress will appear in the
advanced preferences pane. Embeddings are inserted
into itemEmbeddings SQL table. For now, the table is local
only, no syncing is involved.
- when embedding model pref is set to "Disabled", the model
is deleted and embeddings table is cleared.
- when an embedding model is selected, quick search dropdown
has a new "Similarity" mode, which will run semantic search
on the current scope of items.
- semantic search does not clearly define what counts
as "relevant" and what is "not relevant". In addition,
it will change depending on the library and query. So
we cannot semantically filter out items the way
it is done via SQL. Semantic search returns the ranking
but items cannot be sorted because it is done by the itemTree
based on column selection.
So in "similarity" quicksearch mode, there is also a dropdown
to select how many top relevant items to keep (top 5 - top 100).
It allows the user to keep the most relevant items depending
on the context, without conflicting with itemTree sorting.
- semantic search happens in-memory. On a large 5K library
it's fast, but we could consider sqlite-vec extension if
needed.
Add a clientVersion column for items, collections, searches, and
libraries, incremented once per library per transaction on every
object save or deletion. The local API reports these versions instead
of synced versions -- in object JSON, format=versions, since=
filtering, and Last-Modified-Version -- since synced versions don't
reflect local changes and are 0 for unsynced objects. Group metadata
responses keep reporting the synced group version, which has no local
counterpart.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
Someone ended up (via a plugin, presumably) with stored-file attachments
with a full path after 'storage:', which broke file syncing. Throw when
setting a stored-file path containing a slash, and strip paths from
existing filenames in a schema update step. No particular reason to
think that the file with that basename will exist in the storage dir,
but at least it will be looking for the right file and not be totally
broken.
Separately, the dataserver will clean up filenames with full paths and
block going forward.
https://forums.zotero.org/discussion/132822/reference-sychronization-error
Accented stop-words weren't matched by the existing unaccented
entries (e.g., "fur"), so automatic journal abbreviation kept and
capitalized them: "Jahrbuch für Heimatkunde" became "Jahrb. Für
Heimatkunde" instead of "Jahrb. Heimatkunde".
https://groups.google.com/g/zotero-dev/c/uP18QEKe2JU/m/AsoGWxd9AwAJ
The condition `required` flag was removed in #5962, but the column was
kept so older clients could still read the database. The full-text
search changes bump the userdata compatibility version, locking out
those clients, so the column can now be dropped.
Index attachment content into a contentless trigram FTS5 table in a
separate, attached fulltext.sqlite, normalized so matching is accent-
and case-insensitive. For content containing CJK characters, a companion
'ascii'-tokenized table holds bigrams so 1-2 character CJK queries, which
the trigram tokenizer can't match, still work. The extracted text still
lives in the .zotero-ft-cache files, so the index is fully derived and
rebuildable.
Use the FTS index for the fulltextContent condition, falling back to the
cached-text scan for queries too short to index, and point quick
search's content matching at the FTS index in place of the now-removed
word index. (One side effect: quick search now matches attachment
content by substring rather than by word.)
Already-extracted content is migrated into the index at startup, slowing
down on active usage. A background queue then extracts not-yet-indexed
attachments gradually when Zotero is idle. Attachments with no local
file or full-text content are recorded as missing. Content downloaded
via sync is processed into the index immediately when the sync finishes,
rather than waiting for idle like it did before, so it's searchable
immediately in on-demand file-download mode.
The index DB is tied to the main DB via the local user key and rebuilt
if they don't match (e.g., after a delete-and-resync). We compact it by
running FTS5's 'optimize' command once the indexing queue drains, and we
vacuum the attached database when necessary to reclaim disk space.
Closes#2038, #2044
Addresses #1595
Search now ignores accents, so "seance" matches "séance" and vice versa.
Text is normalized with Unicode NFKD compatibility decomposition (which
also handles typographic ligatures, superscripts, full-width forms,
etc.) plus a small map for letters NFKD leaves alone (ø, œ, æ, ß, ...)
and the fraction slash, via Z.Utilities.Internal.normalizeForSearch().
The HTML tags we support in item fields are stripped, so markup isn't
matched (#81). Typographic quotes (#29, #1876) and dashes are folded to
ASCII.
Each searchable column gets a normalized shadow column --
itemDataValues.valueNormalized, tags.nameNormalized,
creators.firstNameNormalized/lastNameNormalized, and
itemAnnotations.textNormalized/commentNormalized -- populated at write
time and matched via COALESCE(normalized, raw) LIKE. NULL is stored when
normalizing only changes case, so plain-ASCII values are only stored
once. This covers the contains/doesNotContain/beginsWith operators in
both quick search and Advanced Search.
The new columns are local-only derived data and aren't synced. Older
clients will ignore them, so this doesn't break DB compatibility.
Existing rows are backfilled after the startup sync by
Zotero.Schema.populateNormalizedSearchColumns(), which should only take
a few seconds on most databases.
Closes#29, #81, #1300, #1876
example:
"Atmospheric Chemistry and Physics" is currently incorrectly abbreviated as "Atmospheric Chem. Phys." should be abbreviated "Atmos. Chem. Phys."
Track when attachments are last opened or read, storing a `lastRead` Unix timestamp on the attachment. For user library items, `lastRead` syncs as an attachment property in item JSON. For group library items, it syncs via a per-user synced setting (like `lastPageIndex`).
- Add `lastRead` column to `itemAttachments`
- Add `AttachmentReadObserver` to update `lastRead` on file open and page change (throttled to 5 min for page changes)
- Add "Recently Read" virtual collection (items read in last 14 days, sorted by `lastRead` descending)
- Add `lastRead` search condition with date operators
- Add `lastRead` item tree column with new `dependsOnChildren` property for parent item aggregation
- Add `getItemLastRead()` to return max `lastRead` across child attachments
Also:
- Generalize collection tree SCSS to support universal (context-fill) icons alongside themed icons
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>
After citation-style-language/styles#7928 renamed Vancouver styles to
NLM terminology, Zotero installations end up with both vancouver.csl
and nlm-citation-sequence.csl. To fix, on init, delete any installed style
whose ID appears in the renamed-styles mapping if the target style
also exists.
---------
Co-authored-by: Dan Stillman <dstillman@zotero.org>