Every engine returned hundreds of items barely above its floor for a
query with a few strong answers. Items now stay only within a margin
of the engine's top score (semantic: display-band fraction; lexical:
share), applied per engine before fusion and in single-engine modes.
A lone result is never cut, and match sets follow.
Margin is search.bestMatchMargin, in percent, exposed in the Advanced
pane. Default 50%.
The setting trades list length for precision, and how much it bites
depends on the library. Where everything is on one subject, the best
match sits close above its neighbours, so a lower value is what
separates a specific answer from the rest of the field, at the cost
of cutting broad questions short. Where subjects are many, the tail
falls away on its own and the value matters little.
Grop the grid calculation og query-passage couples.
Instead, each query has a perfectly matching passage
as well as a passage from roughly the same field but
one that should still not match the result. The floor
is calculated from the scores of query similarities to
those near-match passages. This is significantly simpler
and more explicit.
Passages can be embedded by a local llama.cpp server serving the same
model, configured from Settings -> Advanced. The server is trusted only
because its vectors match the local model's, never by name:
- Zotero.Embeddings.Endpoint: verify() embeds calibration passages
locally and remotely (plus one chunk-length text) and stores a typed
verdict (ok, unreachable, unauthorized, not-embeddings, width-mismatch,
low-agreement, context-too-small) keyed to the URL and model version.
Passages route to the endpoint only while a matching verdict says ok;
queries always embed locally.
- Every batch carries a sentinel text whose local vector is cached, so a
server switched to another model or pooling is caught on that batch and
nothing wrong is stored. Three consecutive failures skip the endpoint
for the rest of the run; a run start rechecks a verified endpoint.
- MODELS gains `serving` (GGUF repo and quant) for both bekko models;
getCommand() renders the llama.cpp command with a 4096 context.
- Preferences: status line and Configure Endpoint dialog with the URL,
the facts the server must match, a copyable command, typed errors, and
a Stop Using Endpoint button.
Also fixes two tokenizer defects in the runtime's transformers.js that
the endpoint comparison exposed, both affecting every locally embedded
bekko vector:
- Metaspace's leading word marker was never prepended (the runtime needs
the legacy add_prefix_space flag). `leadingSpace: true` in MODELS adds
it in embedMany(). Short texts had drifted to cosine 0.83-0.95 against
the reference.
- Newlines and repeated spaces were tokenized differently from the
reference. Whitespace runs now collapse to one space in both embedding
paths.
List bekko-embedding-v1-a8m and bekko-embedding-v1-a25m as multilingual options.
bekko-embedding-v1-a8m is just as good as bge-en-v1.5 and multilingual-e5-small but smaller,
twice as fast, and allows us to store 258 out of 368 dimensions to save on storage.
In addition, bekko-embedding-v1-a8m has readily available GGUF
files to run the model in llama or ollama as a possible
sidecar for faster embedding.
Note: the only english model that performs significantly better on english
benchmarks while being fast is mdbr-leaf-ir. But it does not
have standard GGUF available, so leaving it off for now
Rerunning itemTree refresh on embedding progress takes
a good amount of work on the main thread, which can
lead to frozen UI if there is an active search when
embedding runs.
Refreshing the itemTree during a search is likely not
necesary because it is a bit jarring to see the row order
change all of a sudden as you scroll down the itemTree, and
lexical search is a decent fallback until indexing
is ready.
Store each chunk's embedding as int8 instead of float32, cutting the
vector column from 1536 to 384 bytes per row
at about the same retrieval quality on the benchmark.
Per-vector absmax scaling: the largest component becomes 127.
Vectors are centered on the model's calibration mean before
quantizing. The shared mean dominates a raw embedding, so quantizing
raw would spend most of the 8 bits on it; centering first gives the
int8 range to the informative residual, and drops the per-row vec_sub
from the scoring SQL. The mean is now baked into stored rows, so a
changed mean means a re-index, which the version-keyed calibration
already implied.
Quantized vectors have arbitrary norms, so scoring uses cosine rather
than the dot product, in JS and in SQL (vec_distance_cosine over
vec_int8(embedding) with the query quantized the same way). Calibration
thresholds are measured on centered, quantized vectors so they
describe what search computes.
Instead of breaking up section into pieces of at least
X size (high number of smallersections), try to split section
into chunks of similar sizes, meaning there is less chunks
but they are larger. Overall, it is better for retrival
quality and, importantly, for the amount of storage used
About a half of them do not carry any useful information
(e.g. citations, table-of-content, legend abbreviations)
so they don't quite deserve to be indexed as a separate chunk.
For easier evaluation of the progress and troubleshooting.
ETA: seconds to embed the remaining chunks at the current window rate
Throughput (2 min window): chunks/s and estimated tokens/s, wall-clock
Inference speed (run): chunks/s and tokens/s counting engine time only
Padding efficiency: real tokens over padded tokens, window and run
Batches (run): count, mean chunks and mean tokens per batch
Token budget and number of memory-pressure events this run
Engine threads: in use, runtime optimum, active boost reasons
Engine restarts (run): by cause, memory cap or thread change
Inference and main process: memory and CPU from the 10 s sample
Available system memory
Slice: chunks embedded out of the current slice
Chunks stored: count, total tokens, mean and median size
Chunk sizes: counts in five buckets bounded at 120, 384, 640 and 768 tokens
Split sections: share of chunks from split sections, mean parts per split
Chunks per document: document count, mean, median, max
Documents by chunk count: buckets at 0, 1–10, 11–50, 51–200, over 200
Instead of per-item. It gives a more accurante
representation of the progress since chunks are
units of work and items will have varying amount of
text to embed.
Updated the advanced prefs pane with loading bars
instead of per-library breakdown to reflect that.
Increase the number of threads used by ml engine up
to the "optimal" count when the advanced prefs window
is open (meaning, the user may be actively looking at
indexing progress) or when the system is idle (meaning
the higher cpu usage would not inconvenience the user)
1. Restart the engine every several batches to free up
memory. Without shutting down the engine, ONNX never
releases the memory back to the OS, so the memory
footprint just keeps growing.
2. Set numThreads at half of the optimal concurency
to keep CPU usage low enough and not heat up the device
too much.
With this, the intention is that indexing will indeed
be a background process that would not be noticeable
for the user and not interfere with daily usage.
Using tokenizer to determine exact token size of a chunk
puts a lot of work on the main thread during indexing
only to find if a chunk is over the min limit and is
under the max. There is not much value in precision
so we estimate the token value per chunk with a language
table (e.g. english ~4 chars/token, CJK ~1 char/token, etc.)
If needed precision ends up being necessary, we can
run a tokenizer on a small subset of the text in document
and derive chars/token ration from that instead of
actually tokenizing the entire document.
BestMatch session extracts only a small portion of
previews during scoring to show immediately. The rest
is derived when the browser is idle to not freeze the UI
and sent to the itemTree via onPreviewsFilled callback.
Similar approach to earlier placeholder rows - but
better performing.
Lexical matches from a longer natural language query add a good
amount of noise to search results, so hybrid mode is
sometimes even worse than just semantic. However, lexical
matches are necessary to find particular keywords -
if you search for X, the item with title X or the only
item stating what X is better show up. To reconcile
these two conflicting requirements, fuse scores
in the following way:
- if there are > 2 search terms,
the user is likely asking for a conceptual query that should
be able to produce an informative vector to find
decent matches. Run lexical search on ONLY title + abstract
to make sure that if a rare term is mentioned in metadata,
the item can surface. No fulltext lexical to not add
noise.
- if there are <= 2 search terms, it's likely that
particular terms are being looked up, so run semantic
with fulltext lexical searches
Instead of by global frequency. So if you search "size of a dinosaur"
chunks returned from a paper about dinosaurs use the weights
of "size" and "dinosaur" from within the paper, where "size"
is much rarer and more useful. Otherwise, chunks that
meantion "dinosaur" a lot are returned.
Firstly, it did not always perform well with itemTree scroll.
More importantly, semantic search does take some time -
can be 10+ seconds on a large library. A few extra seconds
to fetch all snippets should not be a big problem. And if
the tail of search results is so long that it becomes -
the fix is to trim the number of search results by
stricter relevance criteria.
Refactor of best match module to contain all best match-related
logic from collectionViewItemTree, so itemTree just
calls relevant methods when needed.
One drawback is that we can't pick the best sentence from
semantic chunk to use as a snippet because that would
mean re-embedding every chunk's sentences on search. So
instead just show the first sentence if no lexical chunks
are available.
Drop the wildcard from the last search term.
It was added for search-as-you-type and probably not
needed now.
Disabled context menu options on search result rows
fix irrelevant annotations not getting hidden
Added search results pane that is shown when a search
result row is selected. It displays the entire chunk
that the snippet is based on. When an attachment is
selected, its search results section lists all search
results
Make the chunk the single unit of a search match. Both
engines now score the same passages instead of returning
two incompatible things that had to be merged: lexical
windows cut around a matched word, and semantic chunks.
A window cut around a word doesn't know where it sits in
the document, so it could not carry a section or page, and
overlap between the two kinds had to be guessed at with a
substring probe. A chunk knows its location, which is what
a snippet header and navigating to the match both need.
Chunks come from the best source an item has: the ones the
semantic index already holds, ranked by the model or not,
and for an unindexed item the same chunking applied to its
structured text or, failing that, its flat text. Everything
above the ladder is indifferent to which rung ran.
An item returns at most three matches, blending the model's
score with how much of the query the chunk's own words
carry. Each entry keeps the whole chunk plus the extent of
the one line worth quoting, so the tree can show a line and
the item pane can read the passage without deriving twice.
Choosing that line is the expensive half -- a chunk that
only means the query has to be read by the model -- so it
happens after ranking and only for the three that survive,
asking about an item's lines in a single call.
Structured text is only used when already extracted.
Generating it costs seconds, which is not a price a preview
can charge.
Fix scrolling while matches fill in. A row asked for its
matches when painted, but a paint pass only touches rows
newly in view and each request replaces the last, so rows
still waiting were dropped and re-requested; the request now
states everything still on screen. Restoring scroll position
also snapped to a row boundary, discarding how far into the
top row the view was scrolled -- unnoticeable at 24px rows,
a visible jump backwards at the height of a match row.
Move the chunking from Zotero.Embedding.Chunking to
Zotero.Utilities.Internal.Chunking so that it can be
shared by Zotero.Embeddings and Zotero.Lexical.
Lexical would need to chunk a document to fetch search
snippets on a per-chunk basis.
Also, helper lexical functions to score how much of a
query a piece of text carries, bm25 style.
Render search snippets in itemTree lazily, as the user
scrolls to them. Fulltext table is contentless, so we cannot
fetch snippet() for each search match. For embeddings,
we need to fetch the structured-text to locate the right
block. Both of these operations can take a long time
when done to a lot of items in _refresh before rendering,
which is why search snippets are extracted on demand.
BestMatch.Session is a new object to wrap the interaction
between the item tree and the search engines. BestMatch.Session.score
returns the search results with an indication which
of them should have search snippets. Not all search
results do - purely semantic matches on abstracts or
notes, as well as all matches on annotations get a snippet.
ItemTree renders a placeholder child row for items that will
have snippets.
Based on the matches flag above, the itemTree renders
placeholder rows. When the placeholder row is rendered,
onSearchMatchRendered is called to tell BestMatch.Session
which attachment's snippets need to be shown. BestMatch.Session
maintains a queue and handles extracting of snippets
when the browser is free to avoid freezing the main thread.
When the snippets are extracted, the placeholder row is
replaced with rows of search matches.
BestMatch.Session maintains the state of what snippets were already
extracted.
Drop search result itemPane componenets, on a new search
scroll the itemTree to the top to see the most relevant results.
Do not store the text of each embedded chunk, just store
the block start/end and character offset into the block
as indications where the chunk is in structured text.
When we need to show that snippet, re-derive it from
structured text. This way, we don't store another copy
of fulltext and save on database space. In addition,
better indication of chunk offset allows us to navigate
to matching piece of text from the snippet more accurately.
Use token budget instead of character budget to enforce
the limit on the memory usage during indexing. Tokens are
the proper measure of embedding work, and different
languages use different amount of tokens. e.g. English
is ~4 characters per token and Chinese is ~1-2, which
means that character budget can mean significantly
more indexing work depending on user's library.
With semantic search disabled, Best Match runs a purely lexical ranking:
query terms are weighted by their rarity in the user's own library, any term
can match (OR semantics), and items are scored by how much of the query they
cover and where — titles count most, then abstracts, then fulltext, notes, and annotations.
Consecutive query words also count as a unit: an item containing "special education"
as a phrase ranks above one containing the words scattered, with the phrase's
importance again measured by its rarity. Items whose evidence is a single matched
unit that is common are not included (e.g. "fall of communism" -> do no list all
items that just have "fall").
With semantic search enabled, both engines score every candidate and the two rankings
are fused with Reciprocal Rank Fusion: an item can match by its words, by its meaning,
or — ranking highest — by both. Results are the union of the two engines' matches, so
a literal match the model doesn't understand and a paraphrase the words don't catch both surface.
The relevance bar shows the item's strongest single piece of evidence, whichever engine found it.
While the semantic index is still building or switching models, searches degrade
to lexical ranking instead of showing nothing; the "not ready" error state is gone.
Previews (the Search Results section) now explain the actual ranking for any matched item,
not just attachments: excerpts of the item's own text around the literal matches,
with the matched words highlighted. When semantic chunks are available they're shown too — with
literal matches highlighted inside them — and merged with the lexical excerpts by evidence strength,
deduplicating excerpts that show the same passage.
Add the query-to-score flow to Zotero.Lexical, built on the content and
item-text indexes:
- Term statistics span both corpora: document frequency sums MATCH
counts over fulltextContent and fulltextItemText, corpus size sums
the state tables, so a term's rarity is a property of the library --
and libraries with few attachments still get real weights
- Units too common to matter are cut relative to the query's best unit
(INFORMATIVE_WEIGHT_FRACTION), so "of" is dropped next to "communism"
while an all-common query keeps its best word
- Matchers are index probes: matchContent() against the content index,
matchFields()/matchNotes()/matchAnnotations() against the item-text
columns. Notes fetch text only for probe matches plus stale/unindexed
notes (getStaleOrUnindexedNoteIDs); quoted phrases verify literally
against stored text everywhere (whitespace/hyphen runs interchange,
other punctuation must match)
- scoreItemIDs() assembles the score (see below), with a floor and
cancellation
Ranking algorithm, per query:
1. Parse into units (words, quoted phrases, CJK runs); trailing
mid-word token matches as a prefix
2. Weigh each unit by smoothed BM25 IDF from the combined corpora;
keep the informative ones
3. Match: presence (1) in titles, abstracts, annotations; saturated,
length-normalized term frequency for notes (computed from text)
and documents (recovered as rank ratios per unit -- for a one-unit
query, ranks compare documents exactly, and the strongest match
anchors 1)
4. Score = sum over units of weight x best boosted evidence across
sources (title x2, abstract x1.3; max, so one word never counts
twice), normalized against the query's ceiling: 1 = full-strength
match on everything asked; below SCORE_FLOOR is no match
Add ftindex.fulltextItemText, an FTS5 table (plus CJK 2-gram twin and
state table) holding each item's searchable text in per-type columns:
title and abstract for regular items, note text for notes, passage and
comment for annotations. One row per itemID; text stored normalized.
- Regular items and annotations index inline in their save transaction;
notes are written by the existing stale-flag queue in the same step
that feeds the trigram tables. Erase clears the entries.
- Backfill queue (processItemTextIndexQueue) covers pre-existing items
after the index rebuild, wired into the startup and background
drains; feed libraries are excluded. _indexDBVersion bumped to 3.
- Advanced prefs: "Items and annotations indexed" line in Index
Statistics; the item-text queue joins the progress/up-to-date logic
and drains while the pane is open.
Why an index: ranked search needs two things per query term that no
existing structure can answer -- whole-word membership ("fall" must not
match "rainfall") and per-word document counts, which drive term
weighting ("communism" outranks "fall" by rarity). The trigram note
index answers substrings, not words, and inflates counts; scanning
item text in JS costs a pass over the whole library per query and
can't prefilter without dropping diacritic matches. Word-level FTS
answers both with index probes, normalized at write time.
Add optional pref to index and search fulltext of attachments.
When enabled, attachment IDs are enqueue after regular items,
notes, and annotations.
Added helpers to extract outline and sections from structured
text module. During indexing, the sections of the attachment
are extracted, large sections are broken into chunks, and
small sections are combined to fit into the context window
of the model. Then, each chunk is embedded with its outline
path as the prefix and added to embeddings table.
Each row now contains full text of the chunk
and it's path - a good amount of duplication needed to
ensure that we can reliably connect the embedding of
the chunk to its text for a preview.
On search in Best Match mode, attachment rows get the score
of the highest ranking chunk, so if there is a very relevant
chunk in an attachment, the regular item with a non-relevant
abstract will still rank highly.
When the attachment row is selected, top 5 matching chunks
appear in the new search results collapsible-section of
the item pane, so one can examine matching chunks
without opening the actual reader.
The pruneModels test set the model pref without stubbing startIndexing,
so its switch chains outlived the test: a delayed disabled-model prune
deleted the shared test calibration, and a later indexing run re-measured
the real model under the fake test-model key, breaking centering in
scoreItemIDs. Only reproducible with a window loaded (full suite / CI),
since ModelHub is inert without one. Stub startIndexing and wait out
switches in the pruneModels test, and stub ensureCalibration wherever
tests run a real indexing pass.