Commit graph

16713 commits

Author SHA1 Message Date
Bogdan Abaev
3ecb7ac198 char budget -> token budget for embedding indexing
Use token budget instead of character budget to enforce
the limit on the memory usage during indexing. Tokens are
the proper measure of embedding work, and different
languages use different amount of tokens. e.g. English
is ~4 characters per token and Chinese is ~1-2, which
means that character budget can mean significantly
more indexing work depending on user's library.
2026-08-20 15:54:09 -07:00
Bogdan Abaev
b67a97ed4c bump weights of title and abstract 2026-08-20 13:50:20 -07:00
Bogdan Abaev
06028bf887 simplify lexical search
Replace per-term scoring model with standard bm25
search which is much simpler and less finicky.
2026-08-20 13:14:37 -07:00
Bogdan Abaev
3f79ce4cd3 lexical test fix
Some checks are pending
CI / Test (shard 1) (push) Waiting to run
CI / Test (shard 2) (push) Waiting to run
CI / Test (shard 3) (push) Waiting to run
CI / Test (shard 4) (push) Waiting to run
CI / Utilities Tests (push) Waiting to run
CI / Build, Upload (push) Waiting to run
2026-08-19 09:55:26 -07:00
Bogdan Abaev
4f782d83bc temp: pref to select hybrid/semantic/lexical mode
Some checks are pending
CI / Test (shard 1) (push) Waiting to run
CI / Test (shard 2) (push) Waiting to run
CI / Test (shard 3) (push) Waiting to run
CI / Test (shard 4) (push) Waiting to run
CI / Utilities Tests (push) Waiting to run
CI / Build, Upload (push) Waiting to run
To make it easier to experiment
2026-08-18 18:01:04 -07:00
Bogdan Abaev
3a1a582cca best match mode: fused bm25 and semantic search
With semantic search disabled, Best Match runs a purely lexical ranking:
query terms are weighted by their rarity in the user's own library, any term
can match (OR semantics), and items are scored by how much of the query they
cover and where — titles count most, then abstracts, then fulltext, notes, and annotations.
Consecutive query words also count as a unit: an item containing "special education"
as a phrase ranks above one containing the words scattered, with the phrase's
importance again measured by its rarity. Items whose evidence is a single matched
unit that is common are not included (e.g. "fall of communism" -> do no list all
items that just have "fall").

With semantic search enabled, both engines score every candidate and the two rankings
are fused with Reciprocal Rank Fusion: an item can match by its words, by its meaning,
or — ranking highest — by both. Results are the union of the two engines' matches, so
a literal match the model doesn't understand and a paraphrase the words don't catch both surface.
The relevance bar shows the item's strongest single piece of evidence, whichever engine found it.

While the semantic index is still building or switching models, searches degrade
to lexical ranking instead of showing nothing; the "not ready" error state is gone.

Previews (the Search Results section) now explain the actual ranking for any matched item,
not just attachments: excerpts of the item's own text around the literal matches,
with the matched words highlighted. When semantic chunks are available they're shown too — with
literal matches highlighted inside them — and merged with the lexical excerpts by evidence strength,
deduplicating excerpts that show the same passage.
2026-08-18 18:01:04 -07:00
Bogdan Abaev
6dec06f18c Ranked lexical scoring over the full-text indexes
Add the query-to-score flow to Zotero.Lexical, built on the content and
item-text indexes:

- Term statistics span both corpora: document frequency sums MATCH
  counts over fulltextContent and fulltextItemText, corpus size sums
  the state tables, so a term's rarity is a property of the library --
  and libraries with few attachments still get real weights
- Units too common to matter are cut relative to the query's best unit
  (INFORMATIVE_WEIGHT_FRACTION), so "of" is dropped next to "communism"
  while an all-common query keeps its best word
- Matchers are index probes: matchContent() against the content index,
  matchFields()/matchNotes()/matchAnnotations() against the item-text
  columns. Notes fetch text only for probe matches plus stale/unindexed
  notes (getStaleOrUnindexedNoteIDs); quoted phrases verify literally
  against stored text everywhere (whitespace/hyphen runs interchange,
  other punctuation must match)
- scoreItemIDs() assembles the score (see below), with a floor and
  cancellation

Ranking algorithm, per query:
  1. Parse into units (words, quoted phrases, CJK runs); trailing
     mid-word token matches as a prefix
  2. Weigh each unit by smoothed BM25 IDF from the combined corpora;
     keep the informative ones
  3. Match: presence (1) in titles, abstracts, annotations; saturated,
     length-normalized term frequency for notes (computed from text)
     and documents (recovered as rank ratios per unit -- for a one-unit
     query, ranks compare documents exactly, and the strongest match
     anchors 1)
  4. Score = sum over units of weight x best boosted evidence across
     sources (title x2, abstract x1.3; max, so one word never counts
     twice), normalized against the query's ceiling: 1 = full-strength
     match on everything asked; below SCORE_FLOOR is no match
2026-08-18 18:00:55 -07:00
Bogdan Abaev
628fd23069 Word-level index over item text for ranked search
Add ftindex.fulltextItemText, an FTS5 table (plus CJK 2-gram twin and
state table) holding each item's searchable text in per-type columns:
title and abstract for regular items, note text for notes, passage and
comment for annotations. One row per itemID; text stored normalized.

- Regular items and annotations index inline in their save transaction;
  notes are written by the existing stale-flag queue in the same step
  that feeds the trigram tables. Erase clears the entries.
- Backfill queue (processItemTextIndexQueue) covers pre-existing items
  after the index rebuild, wired into the startup and background
  drains; feed libraries are excluded. _indexDBVersion bumped to 3.
- Advanced prefs: "Items and annotations indexed" line in Index
  Statistics; the item-text queue joins the progress/up-to-date logic
  and drains while the pane is open.

Why an index: ranked search needs two things per query term that no
existing structure can answer -- whole-word membership ("fall" must not
match "rainfall") and per-word document counts, which drive term
weighting ("communism" outranks "fall" by rarity). The trigram note
index answers substrings, not words, and inflates counts; scanning
item text in JS costs a pass over the whole library per query and
can't prefilter without dropping diacritic matches. Word-level FTS
answers both with index probes, normalized at write time.
2026-08-18 17:36:48 -07:00
Bogdan Abaev
eaf922a02d test fix + schema bump
Some checks failed
CI / Test (shard 1) (push) Has been cancelled
CI / Test (shard 2) (push) Has been cancelled
CI / Test (shard 3) (push) Has been cancelled
CI / Test (shard 4) (push) Has been cancelled
CI / Utilities Tests (push) Has been cancelled
CI / Build, Upload (push) Has been cancelled
2026-08-13 14:29:40 -07:00
Bogdan Abaev
2568b5ccc4 Semantic search on fulltext of attachments
Add optional pref to index and search fulltext of attachments.
When enabled, attachment IDs are enqueue after regular items,
notes, and annotations.

Added helpers to extract outline and sections from structured
text module. During indexing, the sections of the attachment
are extracted, large sections are broken into chunks, and
small sections are combined to fit into the context window
of the model. Then, each chunk is embedded with its outline
path as the prefix and added to embeddings table.
Each row now contains full text of the chunk
and it's path - a good amount of duplication needed to
ensure that we can reliably connect the embedding of
the chunk to its text for a preview.

On search in Best Match mode, attachment rows get the score
of the highest ranking chunk, so if there is a very relevant
chunk in an attachment, the regular item with a non-relevant
abstract will still rank highly.

When the attachment row is selected, top 5 matching chunks
appear in the new search results collapsible-section of
the item pane, so one can examine matching chunks
without opening the actual reader.
2026-08-13 13:39:30 -07:00
Bogdan Abaev
0d52d7098c Fix test failures from dangling model switches
Some checks failed
CI / Test (shard 2) (push) Has been cancelled
CI / Test (shard 3) (push) Has been cancelled
CI / Test (shard 4) (push) Has been cancelled
CI / Test (shard 1) (push) Has been cancelled
CI / Utilities Tests (push) Has been cancelled
CI / Build, Upload (push) Has been cancelled
The pruneModels test set the model pref without stubbing startIndexing,
so its switch chains outlived the test: a delayed disabled-model prune
deleted the shared test calibration, and a later indexing run re-measured
the real model under the fake test-model key, breaking centering in
scoreItemIDs. Only reproducible with a window loaded (full suite / CI),
since ModelHub is inert without one. Stub startIndexing and wait out
switches in the pruneModels test, and stub ensureCalibration wherever
tests run a real indexing pass.
2026-08-11 14:55:27 -07:00
Bogdan Abaev
7997d55cef Remove prefs setting from embeddings tests
To avoid triggering pref-related notifiers which
cause race conditions when tests are run.
2026-08-11 14:39:08 -07:00
Bogdan Abaev
d7d2e5a870 tighten what is skipped for embedding
Very short notes one or two non-descriptive words (e.g. "test")
tend to rank highly for very many irrelevant queries,
so raise the bar for indexing to have 2 words for
reguar items' title and abstract, and 3 words for
everything else.
2026-08-11 14:23:23 -07:00
Bogdan Abaev
b3f3d86c10 fix broken inference calls after inactivity
Firefox ML engine self terminates when idle. Its status
needs to be verified in _getEngine and the old engine needs
to be shut down. Then, a new engine is created. Otherwise,
all subsequent inference calls will fail.
2026-08-11 11:27:10 -07:00
Bogdan Abaev
958153ff26 streamline changing embed models
Move the generation of mean vectors into its own
Zotero.Embeddings.Calibration object from tests.

Now a test run is not required to get mean vector
for a newly added model. Instead, the corpus on which
we generate embeddings to calculate mean vector, as well
as the logic for actual mean vector calculation, lives
in Zotero.Embeddings.Calibration so it can be done on
demand - at the start of the first indexing pass for a
model that hasn't been measured yet.

Each passage from Zotero.Embeddings.Calibration corpus
also has a matching query, to which the passage is
an expected search result. Those queries are also embedded
during calibration, which is used to calculate the appropriate
min score cutoff and max display score for relevance bar.
Score floor is computed by embedding all queries and passages,
finding the similarity between all pairs and locating the cutoff
that only 1% of non-matching query/passage pairs exceed.
Max display score is the median of the matching pairs, so
half of genuine matches fill the bar completely.

Calculated score cutoff, max display value and mean vector
are stored in the embeddings database to be reused later.

During calibration of language-specific models, passages/queries
for languages that the model cannot handle are excluded
(e.g drop russian/chinese/spanish that english models are not meant to handle).
For this, model config expects a new optional language field.
A model that omits it is measured against the whole corpus.

Chunk size is now bounded by CHUNK_MAX_TOKENS rather than by
the model's context window, which still caps it. A model with
a large window would otherwise stop notes being chunked at all,
putting a whole note into a single vector and defeating scoring
an item by its best chunk.

With the mean vector and both score bounds measured rather than
configured, each model's hardcoded meanVector, minScore and
maxDisplayScore are gone from MODELS, along with an unused files
array. A model entry is now only facts that can be read off its
model card.

Added bge-small-zh-v1.5 as a model specifically for chinese,
as well as a number of multilingual and english models
for testing.
2026-08-11 11:27:10 -07:00
Bogdan Abaev
c00062edf8 Search note and annotation content
Notes are indexed on their text and annotations on the passage they mark
together with their comment, so both match on what they actually say
rather than on their parent's title and abstract.

A note can hold more text than a model's context window, so
Zotero.Embeddings.Chunking splits long text using the selected model's
tokenizer. Paragraphs are the topic units: two never share a chunk
unless one is too small to embed on its own, in which case small
paragraphs are combined, and a paragraph over the window is split at
sentence boundaries into even pieces. Stored embeddings are keyed by item
and chunk, so the index is rebuilt on upgrade.

An item scores as its best chunk, so a long note
that addresses a query in one paragraph isn't diluted by the rest, and
its rank reflects the best match anywhere beneath it: a strongly matching
annotation lifts its attachment and its paper. The relevance bar reports
only the row's own score, so a paper ranked by its annotation shows a
high rank over an empty bar rather than claiming to be a match it isn't.
List order and bar fill deliberately disagree in that case.

Annotation rows render the relevance cell, and the
Relevance column moves to the far right while a best-match search is
active so the bars line up across item, note, attachment and annotation
rows.
2026-08-11 11:27:04 -07:00
Dan Stillman
cf58a148ab Fall back to text matches when nothing clears the minimum score
A query naming a word an item's title or abstract uses shouldn't come
back empty because the model scored the item below the floor. Text
matches are used only when the model matched nothing, ranked by their
own scores.
2026-08-05 15:41:42 -04:00
Dan Stillman
82efc8a486 Don't index items with too little text to say anything
A one-character title carries no signal but still scores as a moderate
match against any query. A single ideograph can be a whole word, so
those are kept.
2026-08-05 15:41:42 -04:00
Dan Stillman
4d8b983dda Don't rank items that aren't matches in a best-match search
Centered scores make the scale meaningful, so a per-model minimum can
drop items a query doesn't match rather than ordering noise above real
results.
2026-08-05 15:41:42 -04:00
Dan Stillman
a5fd13e2a4 Subtract each model's mean vector before comparing embeddings
Every embedding a model produces shares a large common direction that
says nothing about the text, so items with little content scored as
moderate matches against everything: for the query "sun", an item titled
"C" outscored a paper about a coronal mass ejection. Removing that
direction spreads the scores out, so no relevance reads as no score.

The mean is a constant per model, computed over titles and abstracts
across fields and languages, and applies to stored vectors as they're
compared, so the index doesn't change. Display ranges are refit to the
scores that result.
2026-08-05 15:41:42 -04:00
Dan Stillman
4a70708e07 Wait for an open transaction before vacuuming the embeddings database
SQLite can't vacuum from within a transaction, and the idle handler that
vacuums the attached database can run while other work has one open.
2026-08-05 15:41:42 -04:00
Dan Stillman
b45f1c751e Read stored embedding hashes in one query per chunk
Indexing re-enqueues every eligible item on each start to find what
changed, and checked each one's stored hash with its own query, so a
large library ran thousands of queries at startup.
2026-08-05 15:41:41 -04:00
Dan Stillman
07efc4af83 Generate embeddings with Firefox's inference runtime
Replace the bundled transformers.js/ONNX Runtime worker with a Zotero.ML
engine on the native ONNX backend, which embeds a batch of abstracts
several times faster and runs the model outside the main process. The
runtime downloads the model files and caches them in the profile
directory, so drop the model directory in the data directory along with
the code that filled it.

Size batches by the amount of text they hold rather than by a fixed item
count, since a batch of long abstracts needs far more memory than the
same number of short ones, and shrink the budget further while the
system is under memory pressure.
2026-08-05 15:41:41 -04:00
Dan Stillman
5eb71b9d10 Add Zotero.ML for Firefox's machine-learning runtime
The runtime runs models in a separate, memory-gated inference process
using the native ONNX Runtime and llama.cpp libraries Firefox ships. It
reads its model configuration, runtime configuration, and model-host
policy from Remote Settings, which our build omits, so loading the empty
module throws on the first collection lookup. Supply those collections
directly instead.

Allow only the two backends that use those native libraries, onnx-native
and llama.cpp. The rest either load a WebAssembly runtime from a Remote
Settings attachment, whose record lookup also needs the Translations
actors our build omits (onnx, wllama, and the best-* backends that fall
back to them), or aren't local at all (openai).
2026-08-05 15:41:41 -04:00
Dan Stillman
5550fe20ca Reuse search results when reranking a best-match search
A best-match rerank re-ran the underlying search on every embeddings
update, even though only the ranking depends on the embeddings. Reuse
the cached search results and recompute just the ranking, except when a
top-K cutoff makes membership depend on the scores.
2026-08-05 15:41:41 -04:00
Dan Stillman
2112da0766 Rerank best-match only on embeddings-index updates
The item tree reran the active best-match search on every 'refresh' item
event, so unrelated bursts (e.g., full-text indexing) triggered a full
re-search and re-score. Flag the embeddings indexer's own notifications
and rerank only for those.
2026-08-05 15:41:41 -04:00
Dan Stillman
24524e1279 Confirm before wiping the index on a Best-Match mode change
Changing the mode -- including disabling -- silently dropped the
stored embeddings and any other downloaded model, costing a full
reindex. Prompt first, since the menu click gives no hint of the cost.
2026-08-05 15:41:41 -04:00
Dan Stillman
e750daffc7 Say "Mode" rather than "Model" in the Best-Match preferences
"Model" and "Downloading model…" are machine-learning terms. The menu
reads as a mode choice: disabled, English, or multilingual.
2026-08-05 15:41:41 -04:00
Dan Stillman
1d17876fe4 Explain the semantic search model choice in the preferences
The model names alone don't say how to choose. Add a caption under the
menu with the trade-off, since switching later means a full reindex.
2026-08-05 15:41:40 -04:00
Dan Stillman
268d59d61f Call semantic search "Best Match" in the UI 2026-08-05 15:41:35 -04:00
Dan Stillman
bb528d278a Show an indexing-progress banner during best-match searches
While a best-match search runs against a partially built embeddings
index, results cover only the indexed items and can look arbitrary.
Show the indexing progress in a banner above the items list, updating
as the index fills, so incomplete results aren't mistaken for a
complete ranking.
2026-08-05 15:27:48 -04:00
Dan Stillman
c5d41229bb Normalize best-match queries
The query is trimmed and a single pair of wrapping quotes is stripped
-- they carry no phrase semantics, since the whole query embeds as one
string. A query that normalizes to nothing (e.g., just quotes) is
treated as no search at all rather than being scored against noise.
2026-08-05 15:27:48 -04:00
Dan Stillman
eb0684d231 Keep embedding blobs out of debug output 2026-08-05 15:27:48 -04:00
Dan Stillman
92853d6d90 Only show the download status when the model actually needs downloading 2026-08-05 15:27:48 -04:00
Dan Stillman
80d9c05648 Show a Stopping state while indexing finishes its current batch
A requested stop only takes effect between batches, so the Stop button
looked unresponsive. Report the stop request in the status line and
disable the button until the batch finishes.
2026-08-05 15:27:47 -04:00
Dan Stillman
5d639d2242 Localize the semantic search index counts 2026-08-05 15:27:47 -04:00
Dan Stillman
4fe312bc65 Add tests for best-match search interactions
Covers saved-search ranking and quick-search precedence, mixed
top-K/collection selections, rank-only behavior with an unavailable
index, reranking on indexer refresh events, scoped 'Match any' saves
keeping the marker at the root, numeric-operator round trips, and
concurrent query embeds sharing one worker call.
2026-08-05 15:27:47 -04:00
Dan Stillman
486bac342f Add semantic ranking to Advanced Search
A root-level 'bestMatch' condition -- serialized as a marker like
joinMode and resultLevel -- ranks results with the Relevance column via
a "Sort results by best match for" field below the root group's
conditions, composing semantic ranking with Boolean filtering. A
best-match quick search converts to it via the Advanced Search button,
and a selected saved search's own marker activates ranking too,
overridden by an active best-match quick search.

Without a cutoff the condition is rank-only and membership is untouched
-- including while the index is unavailable -- so a saved search acts
identically as a source, and unscoreable items just sort last. With the
optional "keeping top" cutoff, carried in the marker's operator,
membership becomes the K most similar results, applied in search() so
scopes, counts, and the API see the same set. A transient search's
cutoff, which applies uniformly to every selected row, is reapplied
over the merged results so a multi-collection selection returns K
members total; a saved search's cutoff is part of its own membership
and never trims other selected rows. The query embedding is cached
in-flight, so the per-row membership passes and the merged ranking
share one worker embed.
2026-08-05 15:27:46 -04:00
Dan Stillman
772189a71e Fall back from best-match mode when semantic search is disabled
Rows kept identifying the active mode as bestMatch after the model was
disabled in the preferences, leaving the active filter returning nothing.
Observe the model pref, switch to Fields & Tags, and rerun any active
search.
2026-08-05 15:26:56 -04:00
Dan Stillman
68431dee46 Replace stale model files when the model revision changes
download() skipped every existing file, so after a revision bump the old
files were kept and then marked as the new revision. Stamp the directory
with the revision being downloaded and clear it on mismatch, so bumps
replace the files while interrupted downloads of the same revision still
resume.
2026-08-05 15:26:56 -04:00
Dan Stillman
356d36e4aa Stop best-match scoring for superseded queries
Refreshes are serialized, so with scoring slower than the quick-search
debounce, intermediate queries queued instead of becoming obsolete. Bump
a generation counter when a filter or the selected rows change and check
it between scoring chunks, abandoning the stale pass and leaving the
rows for the newer refresh to replace.
2026-08-05 15:26:56 -04:00
Dan Stillman
132820a374 Update active best-match searches as embeddings change
New, changed, or removed vectors now rerank an active best-match
search: the indexer announces committed batches -- and index clears from
disabling or a model switch -- with a coalesced 'refresh' item event,
and the items view reruns the search for it. Batch writes also skip
items deleted while the batch was embedding, so a delete can't be undone
by an in-flight batch, and the delete notifier is awaited.
2026-08-05 15:26:56 -04:00
Dan Stillman
771d17adca Guard best-match scoring against model changes
A search could embed the query with one model's prefix on a worker
initialized for another and compare it against vectors from a third.
Tag the worker with the model version it was initialized for, make
scoring wait out an in-progress model switch, refuse to score an index
that wasn't stamped by the active model, and discard results if the
model changes mid-scoring. The items view treats a not-ready index as
an empty result rather than showing an unranked scope.
2026-08-05 15:26:56 -04:00
Dan Stillman
bc73c713d7 Replace the best-match top-K cutoff with a ranked Relevance column
Semantic similarity has no natural relevance threshold, so instead of
asking the user to pick an arbitrary result count, show every scored
item and surface the ranking directly: a Relevance column appears and
becomes the sort while a best-match search is active, and the previous
sort and columns return when it clears. The merged results are scored
in a single pass in the row provider, so ranks are global across a
multi-collection selection, child items (attachments, notes,
annotations) rank via their top-level item, and equal scores get equal
ranks that order deterministically via the secondary sort fields. Items
without a stored embedding are filtered out.

Each cell renders the score's position within the model's display range
as a bar, so relevant results read as full and the irrelevant tail
reads as empty. The ranges are provisional per-model display constants.
Sorting uses the ranks, which are also exposed to assistive technology
and as the cell tooltip. On a focused selected row the bar
switches to white so the fill doesn't vanish into the accent selection
background.
2026-08-05 15:26:55 -04:00
Dan Stillman
89e0f7201f Store embeddings in an attached database instead of zotero.sqlite
The embeddings are a local, rebuildable, model-specific index, so they
don't belong in the main database or its backups. Follow the full-text
content index pattern: a lazily attached embeddings.sqlite versioned via
PRAGMA user_version, tied to the main database by localUserKey, with
corruption recovery and idle-maintenance vacuuming via the DBConnection
hooks. Since a cross-database foreign key isn't possible, item deletions
now clear embeddings via the notifier, and the indexed-model identity
moves from a pref into the database's meta table.
2026-08-05 15:26:55 -04:00
Dan Stillman
b201be5d5a Rename the similarity condition and quick-search mode to bestMatch
"bestMatch" names what the condition does -- rank results by how well
they match the text -- rather than the current scoring mechanism, so
the name can account for later changes to how it works. Rename the
condition-facing identifiers with it; the embeddings engine keeps its
similarity vocabulary.
2026-08-05 15:26:55 -04:00
Bogdan Abaev
29bbaf9e6f local semantic search on abstracts
- added environment to run embedding models locally
(transformers.js, ONNX Runtime WASM binary, etc.). The actual
inference execution happens in a separate worker environment (worker.js)
- added local itemEmbeddings table to store embeddings locally
- in advanced preferences, one can select two options for
semantic search model: english and multilingual. English model
(bge-small-en-v1.5) is better for english-only corpus
but multilingual (multilingual-e5-small) is necessary to handle
abstracts with any other language than english. We can add
more language-specific models as needed.
- when the model is selected, Zotero.Embeddings.download
will download the model (quantized ~100mb) and store it locally.
- Zotero.Embeddings.Indexing will start a process to
index all regular items with title+abstract. It happens in batches
and takes some time. The progress will appear in the
advanced preferences pane. Embeddings are inserted
into itemEmbeddings SQL table. For now, the table is local
only, no syncing is involved.
- when embedding model pref is set to "Disabled", the model
is deleted and embeddings table is cleared.
- when an embedding model is selected, quick search dropdown
has a new "Similarity" mode, which will run semantic search
on the current scope of items.
- semantic search does not clearly define what counts
as "relevant" and what is "not relevant". In addition,
it will change depending on the library and query. So
we cannot semantically filter out items the way
it is done via SQL. Semantic search returns the ranking
but items cannot be sorted because it is done by the itemTree
based on column selection.
So in "similarity" quicksearch mode, there is also a dropdown
to select how many top relevant items to keep (top 5 - top 100).
It allows the user to keep the most relevant items depending
on the context, without conflicting with itemTree sorting.
- semantic search happens in-memory. On a large 5K library
it's fast, but we could consider sqlite-vec extension if
needed.
2026-08-05 15:26:55 -04:00
Dan Stillman
3af8cea1af Name annotation types by type alone in the Advanced Search
Some checks are pending
CI / Test (shard 1) (push) Waiting to run
CI / Test (shard 2) (push) Waiting to run
CI / Test (shard 3) (push) Waiting to run
CI / Test (shard 4) (push) Waiting to run
CI / Utilities Tests (push) Waiting to run
CI / Build, Upload (push) Waiting to run
The menu labeled Annotation Type listed "Highlight annotation" and "Image
Annotation", from the strings the reader announces annotations with. Use
the short names, which existed for two of the six.
2026-08-05 15:02:21 -04:00
windingwind
09249bfb73
fx153: Interface should load with nsIID instead of string (#6008) 2026-08-05 14:18:32 -04:00
Dan Stillman
8fb2870f93 Restore dxcompiler.dll on Windows for WebGPU and leave WebGPU pref on
mozinference (which is currently CPU-only) is better for many tasks, but
plugins may want WebGPU for some features (e.g., chatbots), so just
follow Firefox, which currently enables it by default for Windows and
Apple Silicon macOS. Adds 5.8 MB compressed to the Windows installer.
2026-08-05 14:10:04 -04:00