Commit graph

38 commits

Author SHA1 Message Date
Bogdan Abaev
fff76fcbfc Best Match: hybrid lexical and semantic search
Best Match becomes a ranked search over everything in a library: its
items' metadata, notes and annotations, and the full text of its
attachments. A lexical engine (Zotero.Lexical) always ranks; with
semantic search enabled, the embedding engine ranks too and the two are
fused, so an item can match by its words, by its meaning, or -- ranking
highest -- by both. Matched passages are shown under their items, read
in the item pane, and opened in the reader at the passage. Attachment
vectors can come from the dataserver instead of being computed here.

1. Lexical engine (Zotero.Lexical, fulltext.js)

A word-level FTS5 index, ftindex.fulltextItemText (plus a CJK 2-gram
twin and a state table), holds each item's searchable text in per-type
columns: title and abstract for regular items, note text for notes, the
marked passage and comment for annotations. Regular items and
annotations index inline in their save transaction; notes are written
by the existing stale-flag queue alongside the trigram tables; a
backfill queue covers pre-existing items and joins the startup and
background drains. The trigram index answers substrings and inflates
counts, so it can't tell "fall" from "rainfall" or weigh a word by its
rarity; word-level FTS answers both with index probes.

Ranking is FTS5's BM25 over that index and the existing content index:
a query parses into word, phrase and CJK-run terms joined by OR, so a
document missing a word still ranks below one that has them all; term
weight comes from rarity in the user's own library, with no stoplist.
Consecutive words are added as phrase terms, item-text columns are
weighted (title 6, abstract 4, annotation 2, note 1), and an
attachment's full text counts only where enough of the query's terms
occur within 200 tokens of each other (FTS5 NEAR; 75% of them, so all
of them up to three), so one rare word can't carry a long document to
the top. Scores are divided by the most the expression could earn, so
both indexes report the 0-1 share of the query a document carries.

2. Fusion (Zotero.BestMatch)

Both engines score every candidate and their rankings are fused with
Reciprocal Rank Fusion; results are the union of the engines' matches.
Each engine's tail is cut against its own strongest match before fusion
(search.bestMatchMargin, percent, default 50, in the Advanced pane): a
library on one subject needs a tighter margin to separate a specific
answer from the field, a varied one hardly needs it. The Relevance bar
shows the item's strongest single piece of evidence, whichever engine
found it.

While the semantic index is still building, or semantic search is off,
Best Match degrades to lexical ranking rather than showing nothing, so
the quick-search mode is always offered and the "index not ready" state
is gone. Advanced Search's bestMatch condition scores through the same
facade. A temporary pref, search.bestMatchEngine (hybrid, lexical,
semantic), selects the engine for testing.

3. Attachments, structured text and chunks (Zotero.SDT)

Attachments are indexed on the structured text the document worker
extracts (PDF, EPUB, snapshot), cached as a pack per attachment. How a
pack is cut into chunks is moved to the structured-document-text
module. The chunker lives there so that the client, the dataserver's
indexer and anything else that embeds a document cut it the same way
and produce rows that can be exchanged. Each chunk carries an anchor --
page rects for PDFs, selectors for EPUBs and snapshots -- that locates
its text in the file across extractor versions, so a row cut elsewhere
stays usable when a local re-cut would block differently. Anchors are
stored as deflated JSON; chunk text is not stored, it's read back from
the pack.

The document-worker build bundles the chunker next to the pack reader
(structured-document-text.js and structured-document-text-chunker.js),
and Zotero requires both; on the main thread the chunker only cuts
plain text and reports its version. A large book takes hundreds of ms
to inflate and cut, so cutting (sdt.getChunks) and reading anchors back
with their reader positions (sdt.readAnchors) run in the document
worker, with no fallback here. A worker failure, or a chunker version
that disagrees with the bundled one, comes back as a failed cut rather
than a verdict on the attachment. On a worker error, the document
worker manager fails every pending request and starts a fresh worker
for the next one, instead of leaving the queue stuck behind an
unanswered request. MIN_CHUNKER_VERSION forces a re-cut of attachments
cut by an older chunker.

4. Model, vectors and calibration

One model, bekko-embedding-v1-a25m: multilingual, 8192-token window,
Matryoshka-trained so its 384-dimensional output is cut to 256. Two
runtime tokenizer defects are corrected -- the leading word marker the
runtime fails to add, and whitespace runs tokenized unlike the
reference; the model `revision` tracks changes that alter vectors.

Stored vectors are centered on the model's mean and quantized to int8
(the shared mean would otherwise spend most of the 8 bits), 256 bytes a
row, and scored by cosine in SQL. The mean, the score floor and the
bar's ceiling are measured by Calibration.record() over a corpus of
triples -- a query, a passage that answers it, a near miss from the same
field (resource/embeddings-calibration-corpus.json) -- and pasted into
the model config; the floor sits at the 95th percentile of near-miss
scores, the ceiling at the median of matches.

5. Indexing runs (Zotero.Embeddings.Indexing)

Nothing is queued. An item change or a finished sync kicks a run, and a
kick that lands during a run makes it go again. A run first indexes
locally the items, notes and annotations saved since the index last
looked at them -- a stamp other than the item's clientDateModified, or
a save in the last 5 s -- reading their text from the tables, then
takes every eligible attachment through five steps. Each step is one
pass over the outstanding attachments, 50 at a time by itemID as
full-text sync pages, and finds its own work in the index, so a run cut
short resumes where it was:

- reconcile: what's stored still holds. A file that's gone, or rows cut
  by a chunker before MIN_CHUNKER_VERSION, lose everything;
- extract: every attachment's text cached as a pack, the server's
  included, so no preview waits on an extraction.
- fetch: the server is asked for everything that's its.
- cut: this client's attachments are cut into pending rows with anchors.
  An attachment the worker couldn't cut is left as it is for a later
  run, since a plain-text fallback would stand in for the file for good;
  only an extraction failure falls back to plain text. Rows adopted for
  a file that has since arrived wait the same way.
- embed: every pending row gets a vector, pooled across attachments and
  pages.

The attachment work gives way to a kick, so a just-edited item is
searchable without waiting behind the library's documents, and to a
sync in progress. Items are loaded without caching and never more than
a page at a time. Startup and Resume clear the attachment stamps for a
full pass. Text too short to say anything is skipped: under two words
of title and abstract, under three of anything else.

The index is one table of rows (itemID, chunkIndex, embedding, anchor)
and one of per-item state (itemIndexState: sourceKey, contentHash,
extractor, clientDateModified, and the item's standing with the server);
counts are derived from rows. What a run is built on is declared beside
it: Sources (what each kind of item offers and what text it yields),
Progress (what the preferences pane reads, recomputed on a clock),
Store (the only code that writes the tables), Runtime (the machine's
say: a token budget per engine call halved under memory pressure, a
memory floor to start at all, the engine restarted to give memory back,
half the optimal thread count unless the prefs pane is open or the
system idle, and main-thread work paced against idle time), and
Diagnostics (rates and shape summaries of what Indexing records). The
engine, which the runtime terminates when idle, is checked and
recreated before each use.

6. Sync client (Zotero.Embeddings.Sync)

Available when the account syncs (pref embeddings.sync.enabled is an
override) and no endpoint is active -- configuring one says to embed
there instead. The server is asked about a stored file in a library
syncing with Zotero Storage, in sync or to download, that it hasn't
declined. Rows arriving before their file are adopted when it does.

GET users/{userID}/embeddings?model=M&itemKey=K1,K2,... returns
  { model, items: [
      { key, status: 'success', contentHash, chunks, version,
        rows: [{ chunkIndex, embedding, anchor }] },
      { key, status: 'declined' },
      { key, status: 'pending' } ] }

- success: rows replace what's stored, if contentHash matches the file
  here (or its plain text), rows number exactly `chunks`, and each row
  is well formed; otherwise declined. Rows already held from `version`
  are kept.
- declined: cut and embedded locally in the same run.
- pending, or a key missing from the reply: left as it is and asked
  again next run.
- A reply in another model declines the batch. A failed request stops
  the pipeline; the run goes back to the server after 5 minutes, or at
  once on Resume.

GET users/{userID}/embeddings?format=versions&model=M&since=V (with
If-Modified-Since-Version) returns { version, items: { key: version },
models }. After every sync, checkLibrary() asks for what changed since
the version it recorded and marks every key whose rows aren't from the
version named to be fetched again, declined ones included. The
library's version is recorded last, so a failure repeats the delta.

Server is expected to provide raw vector (not centered, not quantized)
so it does not have to worry about mean vector config.

7. Remote endpoint (Zotero.Embeddings.Endpoint)

Passages can be embedded by a server serving the same model -- a local
llama.cpp server in the OpenAI format, or Text Embeddings Inference --
configured from Settings -> Advanced. The server is trusted only
because its vectors match the local model's, never by name: verify()
embeds fixed texts both ways and stores a typed verdict (ok,
unreachable, unauthorized, not-embeddings, width-mismatch,
low-agreement, context-too-small) keyed to the URL, model version and
format. Every batch carries a sentinel text whose local vector is
cached, so a server switched to another model or pooling is caught on
that batch; three consecutive failures skip the endpoint for the rest
of the run. Queries always embed locally. The Configure Endpoint dialog
shows the facts the server must match and a copyable llama.cpp command.

8. Previews (Zotero.BestMatch.Session, item tree, item pane)

A Session scores a query and owns the passages its results matched in.
The chunk is the unit of a match for both engines: passages come from
the index's own rows read back by anchor, or, for an unindexed item,
from the same cut applied to its structured text (only where already
extracted; generating it costs seconds) or its plain text. An item
returns at most three quoted matches, each blending the model's score
with how much of the query the passage's own words carry. The quoted
line is chosen after ranking: the sentence the lexical engine picks
where the passage says the query's words; where it only means them,
the sentence a static multilingual model (potion-multilingual-128M, via
the runtime's static-embeddings backend) finds closest -- weaker than a
dense model, far faster, and never stored.

The item tree shows matches as two-line child rows -- where the passage
is (section path, page) and the line worth reading, with the query's
words marked -- and a new query shows its results from the top. The ten
best-ranked items' previews are derived before scoring resolves; the
rest are derived while the main thread is idle, paced against what each
costs, and arrive in batches. A changed item's previews are invalidated
and re-derived; the tree is no longer refreshed as embedding progresses,
which kept freezing it mid-search. Selecting an item shows every
passage it matched in a Search Results section of the item pane;
selecting match rows shows the passages themselves, grouped by
attachment, in a pane of their own. Double-click or Enter opens the
attachment at the passage, a PDF scrolled to and highlighting the
anchor's rects.

9. Preferences

The model menu is replaced by one switch, search.bestMatch.enableSemantic
(off: lexical ranking only, nothing indexed, what's indexed kept), and
the mode-change confirmations go with it. The pane shows two progress
bars -- metadata, notes and annotations; attachments -- with the step
under way ("Preparing documents", "Syncing semantic data… X / Y",
"Generating semantic data locally for N items"), the endpoint's status
and its Configure dialog, the quality cutoff, and a diagnostics panel,
hidden by default: throughput, inference speed, padding efficiency,
batches, engine threads and restarts, process memory and CPU, chunk
size distributions.

10. Build and dependencies

The document-worker submodule gains the sdt.getChunks and
sdt.readAnchors actions and builds the chunker as a second bundle next
to the pack reader. Zotero.ML allows the static-embeddings backend and
Mozilla's model hub. The embeddings database is at version 14 and is
rebuilt on upgrade.
2026-10-08 15:10:26 -07:00
Dan Stillman
e679473db1 Remove citeproc-rs
citeproc-rs is no longer maintained, and people who had enabled the
hidden pref were hitting errors.

This also drops the free() call on CSL engines, which only existed to
free the citeproc-rs wasm driver.

https://forums.zotero.org/discussion/133515/
2026-08-31 12:38:05 -04:00
Dan Stillman
c7cffdd26c Move prebuilt reader/note-editor files without the shell
The shell glob in the mv breaks on Windows paths, so every Windows
build silently fell back to building the submodules from source.
2026-08-20 12:34:40 -04:00
Martynas Bagdonas
98d82d2909 Add SDT support 2026-06-12 13:47:08 +03:00
Dan Stillman
a487dd9eef Rename pdf-worker submodule to document-worker
Some checks are pending
CI / Build, Upload, Test (push) Waiting to run
Repo moved to zotero/document-worker on GitHub

After pulling, run:

    git submodule sync
    git submodule update --init document-worker

If an old `pdf-worker/` directory is left behind, it can be removed manually.
2026-04-22 15:50:02 -04:00
Dan Stillman
0ff3ec6a1b Move file renaming strings to fileRenaming.ftl
Move file-renaming-* and rename-files-preview-* strings from zotero.ftl
to a dedicated fileRenaming.ftl. Rename rename-files-preview-* keys to
file-renaming-preview-window-* for consistency.
2026-04-02 15:48:14 -04:00
Abe Jellinek
b6ffa16893 Add failsafe to ensure pdf.js build worked 2026-03-02 17:46:05 -05:00
Abe Jellinek
b0b2378866
Add support code for Read Aloud (#5355) 2026-03-02 13:00:23 -05:00
Dan Stillman
9e23214c77 Add English strings for new item fields and creator types 2025-12-19 21:57:44 -05:00
Dan Stillman
9d22c5406f Add special handling for general-sentence-separator string 2025-10-09 15:32:04 -04:00
Dan Stillman
f81763b173 Remove Chai as Promised
A test using `assert.eventually` was failing after the Bluebird removal
but worked with just `await`, and since there hasn't really been much
point to Chai as Promised since the introduction of `async`/`await` ages
ago, just remove the library instead of figuring out why.
2025-07-30 22:31:08 -04:00
Martynas Bagdonas
36d979f71e Update build scripts for note-editor, pdf-worker, and reader:
- Finalize renaming of pdf-reader to reader
- Remove client- prefixes from URL paths
- Update pdf-worker to use the new document-worker path in preparation for repository rename #5052
- Ensure note-editor properly uses the ZIP build instead of silently falling back to a local rebuild
- Log errors to console
2025-04-30 13:46:32 +03:00
abaevbog
dc084b6961
Move note-editor popup a11y strings to Fluent (#5212)
Per https://github.com/zotero/zotero/pull/5205#issuecomment-2809055908
2025-04-18 00:14:26 -04:00
Abe Jellinek
650be7d350
Scaffold code completion (#5171) 2025-04-09 00:36:22 -04:00
Bogdan Abaev
6d52656a92 integration.ftl file for "integration-*" strings
And a few minor .ftl string edits
2025-03-10 22:40:43 -04:00
Abe Jellinek
eb30c23a0d
Scaffold: Infer types of parameters based on naming conventions (#5038) 2025-02-18 05:07:51 -05:00
Martynas Bagdonas
b84f1a70f0 Update pdf-worker submodule and extract only worker.js from zip build
The pdf-worker zip builds now include cmaps and standard_fonts but in Zotero client we reuse them from Zotero Reader path
2025-01-20 19:12:50 +02:00
Tom Najdek
7c1eb0c3f1
Fix localize-ftl crashing in certain cases #4773 (#4775)
* Tweak the script to use `en-us` `.ftl` files as the source of truth.
* Update `ftl-tx` to a version that can handle referencing terms with arguments.
2024-10-23 01:30:58 -04:00
abaevbog
cde21ac9f2
reader.ftl file for a11y strings in the reader (#4752)
Per: https://github.com/zotero/reader/pull/142#issuecomment-2410574199
2024-10-15 01:35:07 -04:00
Bogdan Abaev
8de5eb360d vpat 43: add label to scaffold translator output (#4018)
In a new scaffold.ftl file.

Vpat 43 also indicates that the read-only state of the text field
does not get announced - it is a VoiceOver-specific issue on
Firefox, so it is not addressed here as it should be fixed
globally.
2024-09-03 07:35:24 -04:00
Tom Najdek
144f2caed8
Use production builds of react libraries (#4482)
And remove patching that doesn't seem to be required anymore
2024-08-02 03:43:02 -04:00
abaevbog
f387c67fbc
remove unneeded .trim from react-virt patch (#4379)
Followup to #4370, per https://github.com/zotero/zotero/pull/4370#discussion_r1674336402
2024-07-12 00:21:45 -04:00
abaevbog
a07117c938
Patch react-virtualized to fix react 18 scroll lag (#4370)
React 18 introduced automatic batching which tries to avoid unnecessary
re-rendering (https://react.dev/blog/2022/03/08/react-18-upgrade-guide#automatic-batching).
In some components based on react-virtualized (e.g. tag selector),
it leads to visible lagginess on scroll via keypress or
mouse wheel (not trackpad). This patch wraps setState of the scroll handler
of react-virtualized with ReactDOM.flushSync, which opts out
of automatic batching.

Fixes: #4368
2024-07-11 01:18:25 -04:00
Tom Najdek
76caebdd8a
Show current locale in error message in localize-ftl. Fix #4302 2024-07-01 11:08:41 +02:00
Tom Najdek
b9f0d26cee
Improve ftl localization scripts 2024-06-21 15:40:10 +02:00
Tom Najdek
b7244998a1
Few fixes to ftl-to-json and localize-ftl scripts (#3707)
* Omit msg-ref-only strings from Transifex JSON
* Fix msg-ref-only strings not included in translated .ftl files
* Update ftl-tx. Simplify localize-ftl script.
* Tweak FTL -> JSON conversion to produce a single file
2024-06-18 06:34:17 -04:00
Tom Najdek
0f6cae891d
JS Build: Fix watch exits if omni update fails 2024-04-02 17:38:15 +02:00
Tom Najdek
95cfc4be13
JS Build: Fix watch exits with error in some scenarios 2024-03-13 11:34:53 +01:00
Tom Najdek
e7f899c09d
JS Build: Fix watch exits when encountering an error 2024-03-12 17:04:47 +01:00
Tom Najdek
0478e66a47
Improve build process. Fix #3758 (#3809)
- Remove directories that should no longer be present in the build.
- Add watching for new files.
- Add debounce and batching to reduce verbosity and avoid needless cleanup and "add to omni" steps
  when entire directories are affected.
2024-03-08 01:26:24 -05:00
Tom Najdek
a68424658a Fix a bug in the build process
Fixes TypeError: Cannot read properties of undefined (reading 'toFixed')
2024-01-24 04:03:12 -05:00
Dan Stillman
838a3090fd Move pdf-reader submodule to reader 2023-08-08 01:40:24 -04:00
Dan Stillman
086399da33 Handle multiple Fluent source files
With list specified in js-build/config.js
2023-05-29 22:46:24 -04:00
Dan Stillman
0de1305665 Tweaks to Fluent/Transifex JSON processing scripts (#3058)
- Use locale directories for JSON files, since that's where the
  Transifex client will interact with them
- Skip non-locale directories (e.g., don't create an .ftl file for
  .DS_Store)
- Other minor simplification
2023-05-29 22:46:24 -04:00
Tom Najdek
afaf0b4968 Add scripts to convert ftl to/from Transifex JSON (#3058) 2023-05-29 22:46:24 -04:00
Dan Stillman
c326a6c971 Fix more files for combined repos 2023-04-29 07:50:54 -04:00
Dan Stillman
c55ef8714b Update app build scripts for new combined repo
Also:

- Replace `install.rdf` with `version` file
- Remove lots of obsolete logic in `prepare_build` (formerly
  `build_xpi`)
2023-04-26 04:40:22 -04:00
Dan Stillman
ae0091fbae Rename scripts folder to js-build 2023-04-26 04:40:22 -04:00