Best Match: hybrid lexical and semantic search

Best Match becomes a ranked search over everything in a library: its
items' metadata, notes and annotations, and the full text of its
attachments. A lexical engine (Zotero.Lexical) always ranks; with
semantic search enabled, the embedding engine ranks too and the two are
fused, so an item can match by its words, by its meaning, or -- ranking
highest -- by both. Matched passages are shown under their items, read
in the item pane, and opened in the reader at the passage. Attachment
vectors can come from the dataserver instead of being computed here.

1. Lexical engine (Zotero.Lexical, fulltext.js)

A word-level FTS5 index, ftindex.fulltextItemText (plus a CJK 2-gram
twin and a state table), holds each item's searchable text in per-type
columns: title and abstract for regular items, note text for notes, the
marked passage and comment for annotations. Regular items and
annotations index inline in their save transaction; notes are written
by the existing stale-flag queue alongside the trigram tables; a
backfill queue covers pre-existing items and joins the startup and
background drains. The trigram index answers substrings and inflates
counts, so it can't tell "fall" from "rainfall" or weigh a word by its
rarity; word-level FTS answers both with index probes.

Ranking is FTS5's BM25 over that index and the existing content index:
a query parses into word, phrase and CJK-run terms joined by OR, so a
document missing a word still ranks below one that has them all; term
weight comes from rarity in the user's own library, with no stoplist.
Consecutive words are added as phrase terms, item-text columns are
weighted (title 6, abstract 4, annotation 2, note 1), and an
attachment's full text counts only where enough of the query's terms
occur within 200 tokens of each other (FTS5 NEAR; 75% of them, so all
of them up to three), so one rare word can't carry a long document to
the top. Scores are divided by the most the expression could earn, so
both indexes report the 0-1 share of the query a document carries.

2. Fusion (Zotero.BestMatch)

Both engines score every candidate and their rankings are fused with
Reciprocal Rank Fusion; results are the union of the engines' matches.
Each engine's tail is cut against its own strongest match before fusion
(search.bestMatchMargin, percent, default 50, in the Advanced pane): a
library on one subject needs a tighter margin to separate a specific
answer from the field, a varied one hardly needs it. The Relevance bar
shows the item's strongest single piece of evidence, whichever engine
found it.

While the semantic index is still building, or semantic search is off,
Best Match degrades to lexical ranking rather than showing nothing, so
the quick-search mode is always offered and the "index not ready" state
is gone. Advanced Search's bestMatch condition scores through the same
facade. A temporary pref, search.bestMatchEngine (hybrid, lexical,
semantic), selects the engine for testing.

3. Attachments, structured text and chunks (Zotero.SDT)

Attachments are indexed on the structured text the document worker
extracts (PDF, EPUB, snapshot), cached as a pack per attachment. How a
pack is cut into chunks is moved to the structured-document-text
module. The chunker lives there so that the client, the dataserver's
indexer and anything else that embeds a document cut it the same way
and produce rows that can be exchanged. Each chunk carries an anchor --
page rects for PDFs, selectors for EPUBs and snapshots -- that locates
its text in the file across extractor versions, so a row cut elsewhere
stays usable when a local re-cut would block differently. Anchors are
stored as deflated JSON; chunk text is not stored, it's read back from
the pack.

The document-worker build bundles the chunker next to the pack reader
(structured-document-text.js and structured-document-text-chunker.js),
and Zotero requires both; on the main thread the chunker only cuts
plain text and reports its version. A large book takes hundreds of ms
to inflate and cut, so cutting (sdt.getChunks) and reading anchors back
with their reader positions (sdt.readAnchors) run in the document
worker, with no fallback here. A worker failure, or a chunker version
that disagrees with the bundled one, comes back as a failed cut rather
than a verdict on the attachment. On a worker error, the document
worker manager fails every pending request and starts a fresh worker
for the next one, instead of leaving the queue stuck behind an
unanswered request. MIN_CHUNKER_VERSION forces a re-cut of attachments
cut by an older chunker.

4. Model, vectors and calibration

One model, bekko-embedding-v1-a25m: multilingual, 8192-token window,
Matryoshka-trained so its 384-dimensional output is cut to 256. Two
runtime tokenizer defects are corrected -- the leading word marker the
runtime fails to add, and whitespace runs tokenized unlike the
reference; the model `revision` tracks changes that alter vectors.

Stored vectors are centered on the model's mean and quantized to int8
(the shared mean would otherwise spend most of the 8 bits), 256 bytes a
row, and scored by cosine in SQL. The mean, the score floor and the
bar's ceiling are measured by Calibration.record() over a corpus of
triples -- a query, a passage that answers it, a near miss from the same
field (resource/embeddings-calibration-corpus.json) -- and pasted into
the model config; the floor sits at the 95th percentile of near-miss
scores, the ceiling at the median of matches.

5. Indexing runs (Zotero.Embeddings.Indexing)

Nothing is queued. An item change or a finished sync kicks a run, and a
kick that lands during a run makes it go again. A run first indexes
locally the items, notes and annotations saved since the index last
looked at them -- a stamp other than the item's clientDateModified, or
a save in the last 5 s -- reading their text from the tables, then
takes every eligible attachment through five steps. Each step is one
pass over the outstanding attachments, 50 at a time by itemID as
full-text sync pages, and finds its own work in the index, so a run cut
short resumes where it was:

- reconcile: what's stored still holds. A file that's gone, or rows cut
  by a chunker before MIN_CHUNKER_VERSION, lose everything;
- extract: every attachment's text cached as a pack, the server's
  included, so no preview waits on an extraction.
- fetch: the server is asked for everything that's its.
- cut: this client's attachments are cut into pending rows with anchors.
  An attachment the worker couldn't cut is left as it is for a later
  run, since a plain-text fallback would stand in for the file for good;
  only an extraction failure falls back to plain text. Rows adopted for
  a file that has since arrived wait the same way.
- embed: every pending row gets a vector, pooled across attachments and
  pages.

The attachment work gives way to a kick, so a just-edited item is
searchable without waiting behind the library's documents, and to a
sync in progress. Items are loaded without caching and never more than
a page at a time. Startup and Resume clear the attachment stamps for a
full pass. Text too short to say anything is skipped: under two words
of title and abstract, under three of anything else.

The index is one table of rows (itemID, chunkIndex, embedding, anchor)
and one of per-item state (itemIndexState: sourceKey, contentHash,
extractor, clientDateModified, and the item's standing with the server);
counts are derived from rows. What a run is built on is declared beside
it: Sources (what each kind of item offers and what text it yields),
Progress (what the preferences pane reads, recomputed on a clock),
Store (the only code that writes the tables), Runtime (the machine's
say: a token budget per engine call halved under memory pressure, a
memory floor to start at all, the engine restarted to give memory back,
half the optimal thread count unless the prefs pane is open or the
system idle, and main-thread work paced against idle time), and
Diagnostics (rates and shape summaries of what Indexing records). The
engine, which the runtime terminates when idle, is checked and
recreated before each use.

6. Sync client (Zotero.Embeddings.Sync)

Available when the account syncs (pref embeddings.sync.enabled is an
override) and no endpoint is active -- configuring one says to embed
there instead. The server is asked about a stored file in a library
syncing with Zotero Storage, in sync or to download, that it hasn't
declined. Rows arriving before their file are adopted when it does.

GET users/{userID}/embeddings?model=M&itemKey=K1,K2,... returns
  { model, items: [
      { key, status: 'success', contentHash, chunks, version,
        rows: [{ chunkIndex, embedding, anchor }] },
      { key, status: 'declined' },
      { key, status: 'pending' } ] }

- success: rows replace what's stored, if contentHash matches the file
  here (or its plain text), rows number exactly `chunks`, and each row
  is well formed; otherwise declined. Rows already held from `version`
  are kept.
- declined: cut and embedded locally in the same run.
- pending, or a key missing from the reply: left as it is and asked
  again next run.
- A reply in another model declines the batch. A failed request stops
  the pipeline; the run goes back to the server after 5 minutes, or at
  once on Resume.

GET users/{userID}/embeddings?format=versions&model=M&since=V (with
If-Modified-Since-Version) returns { version, items: { key: version },
models }. After every sync, checkLibrary() asks for what changed since
the version it recorded and marks every key whose rows aren't from the
version named to be fetched again, declined ones included. The
library's version is recorded last, so a failure repeats the delta.

Server is expected to provide raw vector (not centered, not quantized)
so it does not have to worry about mean vector config.

7. Remote endpoint (Zotero.Embeddings.Endpoint)

Passages can be embedded by a server serving the same model -- a local
llama.cpp server in the OpenAI format, or Text Embeddings Inference --
configured from Settings -> Advanced. The server is trusted only
because its vectors match the local model's, never by name: verify()
embeds fixed texts both ways and stores a typed verdict (ok,
unreachable, unauthorized, not-embeddings, width-mismatch,
low-agreement, context-too-small) keyed to the URL, model version and
format. Every batch carries a sentinel text whose local vector is
cached, so a server switched to another model or pooling is caught on
that batch; three consecutive failures skip the endpoint for the rest
of the run. Queries always embed locally. The Configure Endpoint dialog
shows the facts the server must match and a copyable llama.cpp command.

8. Previews (Zotero.BestMatch.Session, item tree, item pane)

A Session scores a query and owns the passages its results matched in.
The chunk is the unit of a match for both engines: passages come from
the index's own rows read back by anchor, or, for an unindexed item,
from the same cut applied to its structured text (only where already
extracted; generating it costs seconds) or its plain text. An item
returns at most three quoted matches, each blending the model's score
with how much of the query the passage's own words carry. The quoted
line is chosen after ranking: the sentence the lexical engine picks
where the passage says the query's words; where it only means them,
the sentence a static multilingual model (potion-multilingual-128M, via
the runtime's static-embeddings backend) finds closest -- weaker than a
dense model, far faster, and never stored.

The item tree shows matches as two-line child rows -- where the passage
is (section path, page) and the line worth reading, with the query's
words marked -- and a new query shows its results from the top. The ten
best-ranked items' previews are derived before scoring resolves; the
rest are derived while the main thread is idle, paced against what each
costs, and arrive in batches. A changed item's previews are invalidated
and re-derived; the tree is no longer refreshed as embedding progresses,
which kept freezing it mid-search. Selecting an item shows every
passage it matched in a Search Results section of the item pane;
selecting match rows shows the passages themselves, grouped by
attachment, in a pane of their own. Double-click or Enter opens the
attachment at the passage, a PDF scrolled to and highlighting the
anchor's rects.

9. Preferences

The model menu is replaced by one switch, search.bestMatch.enableSemantic
(off: lexical ranking only, nothing indexed, what's indexed kept), and
the mode-change confirmations go with it. The pane shows two progress
bars -- metadata, notes and annotations; attachments -- with the step
under way ("Preparing documents", "Syncing semantic data… X / Y",
"Generating semantic data locally for N items"), the endpoint's status
and its Configure dialog, the quality cutoff, and a diagnostics panel,
hidden by default: throughput, inference speed, padding efficiency,
batches, engine threads and restarts, process memory and CPU, chunk
size distributions.

10. Build and dependencies

The document-worker submodule gains the sdt.getChunks and
sdt.readAnchors actions and builds the chunker as a second bundle next
to the pack reader. Zotero.ML allows the static-embeddings backend and
Mozilla's model hub. The embeddings database is at version 14 and is
rebuilt on upgrade.
This commit is contained in:
Bogdan Abaev 2026-10-07 15:10:16 -07:00
parent 15f7841570
commit fff76fcbfc
59 changed files with 15983 additions and 1811 deletions

View file

@ -43,7 +43,7 @@ const React = require('react');
const ReactDOM = require('react-dom');
const ItemTree = require('zotero/itemTree');
const { ItemTreeRowProvider } = ItemTree;
const { LibraryHeaderItemTreeRow, SpacerItemTreeRow } = require('zotero/itemTreeRow');
const { LibraryHeaderItemTreeRow, SpacerItemTreeRow, SearchMatch } = require('zotero/itemTreeRow');
const { OS } = ChromeUtils.importESModule("chrome://zotero/content/osfile.mjs");
const { ZOTERO_CONFIG } = ChromeUtils.importESModule('resource://zotero/config.mjs');
@ -171,23 +171,24 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
}
/**
* Best-match ranks for the Relevance column, computed over the merged
* result set in _refresh() while a best-match search is active
* Best-match ranks for the Relevance column while a best-match search is
* active, from the session's last scoring pass
*
* @returns {Map} - treeViewID -> 1-based rank (1 = most similar)
*/
getBestMatchRanks() {
return this._bestMatchRanks || new Map();
return this._bestMatchSession?.ranks ?? new Map();
}
/**
* Score fractions for the Relevance column's bars, computed alongside the
* ranks (see Zotero.Embeddings.getScoreFraction())
* ranks (see Zotero.BestMatch.Session#barFractions), so the bars always
* agree with the ranking
*
* @returns {Map} - treeViewID -> 0-1 fraction of the model's display range
* @returns {Map} - treeViewID -> 0-1 fraction for the bar
*/
getBestMatchBarFractions() {
return this._bestMatchBarFractions || new Map();
return this._bestMatchSession?.barFractions ?? new Map();
}
/**
@ -202,62 +203,25 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
}
/**
* Compute the current index coverage across the selected rows' libraries.
* Never throws -- the banner is informational and shouldn't break a
* refresh.
* Compute the current index coverage (see Zotero.BestMatch.getIndexState())
*
* @return {Promise<Object|null>}
*/
async _getBestMatchIndexState() {
try {
let status = Zotero.Embeddings.Indexing.getStatus();
if (!status.enabled) {
return null;
}
// Counts aren't populated until the indexer runs in this session
if (!status.libraries.length) {
status = await Zotero.Embeddings.Indexing.refreshStatus();
}
let libraryIDs = new Set(
this.collectionTreeRows
.map(row => row.ref?.libraryID)
.filter(id => id !== undefined)
);
let libraries = status.libraries
.filter(lib => !libraryIDs.size || libraryIDs.has(lib.libraryID));
let indexed = libraries.reduce((sum, lib) => sum + lib.indexed, 0);
let total = libraries.reduce((sum, lib) => sum + lib.eligible, 0);
if (indexed >= total) {
return null;
}
// Only an explicit pause reports as paused. Anything else --
// between runs (startup, the pre-run debounce) or after an error
// (detailed in the preferences) -- reports as indexing, since the
// banner explains the incomplete coverage, not the indexer state
return {
type: status.paused ? 'paused' : 'indexing',
indexed,
total
};
}
catch (e) {
Zotero.logError(e);
return null;
}
return Zotero.BestMatch.getIndexState();
}
/**
* The semantic stage of a best-match search: score the merged,
* The ranking stage of a best-match search: score the merged,
* deduplicated results from all selected rows against the query in a
* single call, and keep the scoreable items ranked globally across the
* selection. Child items (attachments, notes, annotations) are scored via
* their top-level item, so result sets at other levels (e.g. a saved
* search returning annotations) rank by their parent item. Equal scores
* get equal ranks, so tied rows (including a child and its parent) order
* deterministically via the secondary sort fields.
* single session call, which also derives the match previews and
* computes the ranks and bar fractions the Relevance column reads (see
* Zotero.BestMatch.Session#score()). An item is kept when it or
* anything beneath it matched, so a strongly matching annotation keeps
* its attachment and its paper in the results.
*
* @param {Zotero.Item[]} items - Merged results from all selected rows
* @return {Promise<Zotero.Item[]>} - The scoreable items
* @return {Promise<Zotero.Item[]>} - The matching items
*/
async _applyBestMatch(items) {
// With multiple selected rows carrying different best-match sources,
@ -265,10 +229,13 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
let queryRow = this.collectionTreeRows.find(rowIsBestMatchSearch);
let query = queryRow.getBestMatchQuery();
let source = queryRow.getBestMatchSource();
// A top-K cutoff is reapplied to the merged candidates below only when
// the source is the transient Advanced Search, which applies uniformly
// to every selected row. A saved search's cutoff is part of that row's
// own membership and must not trim other selected rows' results.
// Each selected row's search applies a top-K cutoff to its own scope,
// so K is reapplied to the merged candidates (see Session#score()) --
// a multi-row selection returns K members total rather than K per row.
// Only when the source is the transient Advanced Search, which applies
// uniformly to every selected row: a saved search's cutoff is part of
// that row's own membership and must not trim other selected rows'
// results.
let topK = queryRow.advancedSearch && source
? source.getBestMatchQuery().topK
: false;
@ -277,83 +244,121 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
// searches, so keep unscoreable items -- they sort after the ranked
// ones. (A uniform top-K set contains no unscoreable items anyway.)
let keepUnscored = !!source;
// Map each item to the item whose embedding scores it
let sourceIDByItem = new Map();
for (let item of items) {
if (!(item instanceof Zotero.Item)) {
continue;
}
let source = item.isRegularItem() ? item : item.topLevelItem;
if (source) {
sourceIDByItem.set(item, source.id);
}
}
let scores;
let candidateIDs = items
.filter(item => item instanceof Zotero.Item)
.map(item => item.id);
let generation = this._bestMatchGeneration;
// The session scores the query and owns the match previews the tree
// shows as child rows. A new query gets a fresh session -- the old
// one must derive nothing more -- while a re-score of the same query
// (an item edit, an index update) keeps it, so already-derived
// previews survive; the previews of the items that actually changed
// are invalidated in notify().
let session = this._bestMatchSession;
let newQuery = !session || session.queryText !== query;
if (newQuery) {
session?.dispose();
session = Zotero.BestMatch.createSession(query);
session.onPreviewsFilled = itemIDs => this._showFilledPreviews(session, itemIDs);
this._bestMatchSession = session;
}
try {
scores = await Zotero.Embeddings.scoreItemIDs(query, [...new Set(sourceIDByItem.values())], {
// Scoring derives the best-scored items' previews before it
// resolves; the rest arrive through onPreviewsFilled above
await session.score(candidateIDs, {
topK,
// A newer filter (e.g. more typed search text) makes this
// query obsolete -- stop scoring and let its refresh take over
shouldCancel: () => generation !== this._bestMatchGeneration
});
}
catch (e) {
if (e instanceof Zotero.Embeddings.ScoringCancelledError) {
if (e instanceof Zotero.BestMatch.ScoringCancelledError) {
throw e;
}
// Scoring can fail while the model is still downloading or the
// index is being rebuilt
if (e instanceof Zotero.Embeddings.IndexNotReadyError) {
Zotero.debug("Embeddings: index not ready for best-match search");
Zotero.logError(e);
session.dispose();
if (this._bestMatchSession == session) {
this._bestMatchSession = null;
}
else {
Zotero.logError(e);
}
this._bestMatchRanks = new Map();
this._bestMatchIndexState = await this._getBestMatchIndexState();
// A rank-only search's membership doesn't depend on the index, so
// A rank-only search's membership doesn't depend on scoring, so
// show its results unranked; anything else shows no results rather
// than an unranked scope
return keepUnscored ? items : [];
}
// Each selected row's search applies a top-K cutoff to its own scope,
// so trim the merged candidates to K again here, with the same
// deterministic order as search(), so a multi-row selection returns K
// members total rather than K per row
if (topK) {
scores = new Map(
[...scores.entries()]
.sort((a, b) => (b[1] - a[1]) || (a[0] - b[0]))
.slice(0, topK)
);
// A cancellation that lands after the last derivation resolves
// score() normally, so check once more before building the view state
if (generation !== this._bestMatchGeneration) {
throw new Zotero.BestMatch.ScoringCancelledError();
}
let rankOfScore = new Map(
[...new Set(scores.values())].sort((a, b) => b - a).map((score, i) => [score, i + 1])
);
// A new query's results are shown from the top (see _refresh())
if (newQuery) {
this._scrollToTopOnUpdate = true;
}
// The session's ranks cover every row with a match anywhere beneath
// it, so they say which items stay in the results
let kept = [];
let ranks = new Map();
let fractions = new Map();
for (let item of items) {
let sourceID = sourceIDByItem.get(item);
if (sourceID === undefined || !scores.has(sourceID)) {
if (!(item instanceof Zotero.Item) || !session.ranks.has(item.treeViewID)) {
if (keepUnscored) {
kept.push(item);
}
continue;
}
kept.push(item);
ranks.set(item.treeViewID, rankOfScore.get(scores.get(sourceID)));
fractions.set(
item.treeViewID,
Zotero.Embeddings.getScoreFraction(scores.get(sourceID))
);
}
this._bestMatchRanks = ranks;
this._bestMatchBarFractions = fractions;
this._bestMatchIndexState = await this._getBestMatchIndexState();
return kept;
}
/**
* Show the match rows of previews derived after the search resolved (see
* Zotero.BestMatch.Session#score()), by reopening each item -- the same
* path that builds children for an expansion the user asks for.
*
* @param {Zotero.BestMatch.Session} session - Ignored once it isn't the
* session the tree is showing
* @param {Number[]} itemIDs
*/
_showFilledPreviews(session, itemIDs) {
if (this._bestMatchSession !== session) {
return;
}
let shown = [];
for (let itemID of itemIDs) {
// Looked up per item, since reopening one shifts the rows below it
let item = Zotero.Items.get(itemID);
let index = item ? this._rowMap[item.treeViewID] : undefined;
if (index === undefined || !this.isContainer(index)) {
continue;
}
if (this.isContainerOpen(index)) {
this._toggleOpenState(index);
}
this._toggleOpenState(index);
shown.push(itemID);
}
// Redrawing is the expensive part, so a batch with no rows in the
// tree (under a collapsed parent, say) costs nothing
if (!shown.length) {
return;
}
// The twisty appears with the preview, so the rows redraw too
this.itemTree.invalidateRowCache(shown);
this.runListeners('update', true, { restoreSelection: true, restoreScroll: true });
}
/**
* The session holding the passages of the active best-match search, or
* null when no such search is running
*
* @return {Zotero.BestMatch.Session|null}
*/
get bestMatchSession() {
return this._bestMatchSession ?? null;
}
/**
* When showing multiple libraries, group rows by library in collections-list
* order -- independent of the active sort direction
@ -585,14 +590,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
this.itemTree._refreshPromise = deferred.promise;
try {
// A best-match rerank after an embeddings change reuses the cached
// search results -- only the ranking depends on the embeddings, so
// there's no need to re-run the underlying search
if (!options.reuseSearchResults) {
this.collectionTreeRows.forEach(row => row.clearCache());
}
this._bestMatchRanks = null;
this._bestMatchBarFractions = null;
this.collectionTreeRows.forEach(row => row.clearCache());
this._bestMatchIndexState = null;
// Get the full set of items we want to show, merged across all selected rows
let newSearchItemSet = new Set();
@ -656,7 +654,12 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
|| item.isRegularItem();
});
}
// The semantic stage: one scoring pass over the merged results
// The ranking stage: one scoring pass over the merged results
if (!bestMatchSearch && this._bestMatchSession) {
// Leaving best-match search: the previews go with it
this._bestMatchSession.dispose();
this._bestMatchSession = null;
}
if (bestMatchSearch) {
try {
newSearchItems = await this._applyBestMatch(newSearchItems);
@ -665,7 +668,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
// A newer filter superseded this one mid-scoring -- leave the
// rows as they are and let the newer filter's refresh replace
// them
if (e instanceof Zotero.Embeddings.ScoringCancelledError) {
if (e instanceof Zotero.BestMatch.ScoringCancelledError) {
deferred.resolve();
return;
}
@ -712,6 +715,11 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
if (!row.isObjectRow) {
continue;
}
// Don't copy search-match rows -- they're rebuilt from the new
// query's previews when their container reopens
if (row.ref instanceof SearchMatch) {
continue;
}
// Top-level items
if (row.level == 0) {
// A top-level attachment moved into a parent. Don't copy, it will be added
@ -794,6 +802,14 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
// applied to rows carried over from the previous view
this._sort(forceSortAll || this._groupedByLibrary ? null : [...addedItemIDs]);
// Set before the containers below are rebuilt, since their children
// are filtered against these (see ItemTreeRow#getChildItems(),
// which hides the annotations a search didn't match): rebuilding
// them against the previous search's state leaves rows this one
// excludes
this._searchMode = newSearchMode;
this._searchItemIDs = newSearchItemIDs; // items matching the search
// Toggle all open containers closed and open to refresh child items
var t = new Date();
for (let i = this.rows.length - 1; i >= 0; i--) {
@ -809,8 +825,6 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
this.refreshRowMap();
}
this._searchMode = newSearchMode;
this._searchItemIDs = newSearchItemIDs; // items matching the search
this.itemTree.invalidateRowCache(true);
if (this.viewMode != 'publications') {
@ -838,11 +852,19 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
if (!this.isContainer(i) || this.isContainerOpen(i)) {
continue;
}
let item = this.getRow(i).ref;
let row = this.getRow(i);
if (!(row.ref instanceof Zotero.Item)) {
continue;
}
let item = row.ref;
let attachments = item.isRegularItem() ? item.getAttachments() : [];
// expand item row if it is a parent of a match
// OR if it has a child that is a parent of a match
let shouldBeOpened = searchParentIDs.has(item.id) || attachments.some(id => searchParentIDs.has(id));
// OR if it has best-match preview rows to show -- one still
// deriving has none, and opens in _showFilledPreviews() instead
let shouldBeOpened = searchParentIDs.has(item.id)
|| attachments.some(id => searchParentIDs.has(id))
|| this._bestMatchSession?.getPreviews(item.id)?.state == 'filled';
if (shouldBeOpened) {
this._toggleOpenState(i, true);
}
@ -866,6 +888,12 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
try {
await this._refresh(options);
// A new best-match query shows its results from the top (see
// _applyBestMatch()), wherever the previous ones were scrolled to
if (this._scrollToTopOnUpdate) {
this._scrollToTopOnUpdate = false;
options = { ...options, scrollToTop: true };
}
this.runListeners('update', true, options);
await this.itemTree.waitForLoad();
this.itemTree.runListeners('refresh');
@ -900,12 +928,30 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
await this.itemTree._refreshPromise;
const cachedSelection = this.itemTree._cachedSelection;
const collectionTreeRows = this.collectionTreeRows;
const collectionTreeRows = this.collectionTreeRows;
// Opening an attachment writes its lastRead, which arrives here as an
// ordinary modify and triggers best match rerun if it is active.
// For now, do nothing since best match searches are costly.
if (type == 'item' && action == 'modify' && ids.length
&& collectionTreeRows.some(rowIsBestMatchSearch)
&& ids.every((id) => {
let item = Zotero.Items.get(id);
return item && item.isAttachment();
})) {
return;
}
// A changed item's derived match previews are stale: back to pending,
// re-derived by the refresh the change triggers below, which keeps
// every other item's derived text.
if (type == 'item' && ['modify', 'refresh'].includes(action) && this._bestMatchSession) {
this._bestMatchSession.invalidate(ids.map(id => parseInt(id)));
}
var initialRowCount = this.getRowCount();
var madeChanges = false;
var refresh = false;
var reuseSearchResults = false;
var sort = false;
// Selection strategy
@ -956,18 +1002,10 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
if (items.length == 0) return;
}
if (action == 'refresh' && type == 'item' && extraData && extraData.embeddingsUpdate
&& collectionTreeRows.some(rowIsBestMatchSearch)) {
// The background indexer committed new or changed embeddings, so
// rerun the active best-match search to update the scores and ranks.
// Only the ranking depends on the embeddings, so unless a selected
// row trims membership by score (a top-K cutoff), the underlying
// search results are unchanged and can be reused while just the
// ranking is recomputed.
reuseSearchResults = collectionTreeRows.every((row) => {
return typeof row.hasBestMatchCutoff == 'function'
&& !row.hasBestMatchCutoff();
});
if (action == 'refresh' && type == 'item' && this._bestMatchSession
&& ids.some(id => this._bestMatchSession.getPreviews(parseInt(id)))) {
// The invalidation above reset these items' previews, and only a
// scoring pass derives previews, so re-run the search
this.itemTree.invalidateRowCache(ids);
refresh = true;
madeChanges = true;
@ -1248,7 +1286,7 @@ class CollectionViewItemTreeRowProvider extends ItemTreeRowProvider {
}
if (refresh) {
await this._refresh({ reuseSearchResults });
await this._refresh();
}
if (sort) {
await this.itemTree._ensureSortContextReady();
@ -1456,9 +1494,8 @@ class CollectionViewItemTree extends ItemTree {
/**
* While a best-match search runs against a partially built embeddings
* index, show a banner above the items list with the indexing progress,
* so incomplete results aren't mistaken for a complete ranking. Updates
* arrive with the periodic refreshes the indexer triggers as it fills
* the index.
* so incomplete results aren't mistaken for a complete ranking. The
* counts are read when the search refreshes.
*/
_renderTablePrologue() {
let state = this.rowProvider.getBestMatchIndexState();

View file

@ -159,11 +159,7 @@ module.exports = class {
Object.assign(this, options);
const { itemHeight, targetElement, innerElem } = this;
const itemCount = this._getItemCount();
const [offsetIdx, offset] = this._rowOffsets.at(-1);
const listHeight = offset + (itemCount - offsetIdx) * this.itemHeight;
innerElem.style.position = 'relative';
innerElem.style.height = `${listHeight}px`;
// Recalculate custom row height offsets
this._rowOffsets = [[0, 0]];
let previousRowOffset = 0;
@ -176,6 +172,13 @@ module.exports = class {
previousRowOffset = offset;
}
// From the offsets just recalculated: the list is as tall as the rows
// it now has, not as the rows it had
const [offsetIdx, offset] = this._rowOffsets.at(-1);
const listHeight = offset + (itemCount - offsetIdx) * itemHeight;
innerElem.style.position = 'relative';
innerElem.style.height = `${listHeight}px`;
this.scrollDirection = 0;
this.scrollOffset = targetElement.scrollTop;
}

View file

@ -75,6 +75,9 @@ Services.scriptloader.loadSubScript('chrome://zotero/content/elements/itemTreeMe
['attachment-row', 'chrome://zotero/content/elements/attachmentRow.js'],
['attachment-annotations-box', 'chrome://zotero/content/elements/attachmentAnnotationsBox.js'],
['annotation-row', 'chrome://zotero/content/elements/annotationRow.js'],
['search-results-box', 'chrome://zotero/content/elements/searchResultsBox.js'],
['search-result-row', 'chrome://zotero/content/elements/searchResultRow.js'],
['search-results-pane', 'chrome://zotero/content/elements/searchResultsPane.js'],
['annotation-items-pane', 'chrome://zotero/content/elements/annotationItemsPane.js'],
['context-notes-list', 'chrome://zotero/content/elements/contextNotesList.js'],
['note-row', 'chrome://zotero/content/elements/noteRow.js'],

View file

@ -393,7 +393,7 @@
}
get _disableSavingOpenState() {
return !!this.closest('merge-pane, scaffold-item-preview, annotation-items-pane');
return !!this.closest('merge-pane, scaffold-item-preview, annotation-items-pane, search-results-pane');
}
get _disableContextMenu() {

View file

@ -55,6 +55,9 @@
<tags-box id="zotero-editpane-tags" class="zotero-editpane-tags" data-pane="tags"/>
<related-box id="zotero-editpane-related" class="zotero-editpane-related" data-pane="related"/>
<search-results-box id="zotero-editpane-search-results" data-pane="search-results" hidden="true"/>
</html:div>
</html:div>
</hbox>

View file

@ -45,6 +45,7 @@
<description id="batch-edit-prompt-message" />
<button id="batch-edit-prompt-enable" data-l10n-id="item-pane-batch-editing-enable" />
</groupbox>
<search-results-pane id="zotero-search-results-pane" />
</deck>
<item-pane-sidenav id="zotero-view-item-sidenav" no-context-notes="true" class="zotero-view-item-sidenav"/>
`);
@ -55,6 +56,7 @@
this._duplicatesPane = this.querySelector("#zotero-duplicates-merge-pane");
this._messagePane = this.querySelector("#zotero-item-message");
this._annotationsPane = this.querySelector("#zotero-annotations-pane");
this._searchResultsPane = this.querySelector("#zotero-search-results-pane");
this._batchEditEnableBtn = this.querySelector("#batch-edit-prompt button");
this._batchEditPromptMessage = this.querySelector("#batch-edit-prompt-message");
this._sidenav = this.querySelector("#zotero-view-item-sidenav");
@ -113,12 +115,12 @@
}
get mode() {
return ["message", "item", "note", "duplicates", "annotations", "batch-edit-prompt"][this._deck.selectedIndex];
return ["message", "item", "note", "duplicates", "annotations", "batch-edit-prompt", "search-results"][this._deck.selectedIndex];
}
/**
* Set mode of item pane
* @param {"message" | "item" | "note" | "duplicates" | "annotations" | "batch-edit-prompt"} type view type
* @param {"message" | "item" | "note" | "duplicates" | "annotations" | "batch-edit-prompt" | "search-results"} type view type
*/
set mode(type) {
this.setAttribute("view-type", type);
@ -133,6 +135,11 @@
}
render() {
// Passages of a search match, rather than items: nothing an item
// pane shows describes one, so the passages are all there is
if (this.searchMatches?.length) {
return this.renderSearchResults(this.searchMatches);
}
if (!this.data) return false;
let renderStatus = false;
// Only annotations selected
@ -193,6 +200,13 @@
return true;
}
renderSearchResults(matches) {
this.mode = "search-results";
this._searchResultsPane.matches = matches;
this._searchResultsPane.render();
return true;
}
renderNoteEditor(item) {
this.mode = "note";
@ -627,8 +641,12 @@
getCurrentPane(mode = undefined) {
if (!mode) {
// Guess a mode from the current data
// Passages of a search match, which aren't items at all
if (this.searchMatches?.length) {
mode = "search-results";
}
// Only annotation items selected
if (this.data.length > 0 && this.data.every(item => item.isAnnotation())) {
else if (this.data.length > 0 && this.data.every(item => item.isAnnotation())) {
mode = "annotations";
}
// No/multiple objects are selected OR selected object is a trashed collection/search
@ -648,7 +666,8 @@
item: "_itemDetails",
note: "_noteEditor",
duplicates: "_duplicatesPane",
annotations: "_annotationsPane"
annotations: "_annotationsPane",
"search-results": "_searchResultsPane"
};
return this[map[mode]];
}
@ -736,6 +755,10 @@
this._deck.selectedIndex = 5;
break;
}
case "search-results": {
this._deck.selectedIndex = 6;
break;
}
}
let isViewingItem = type == "item";
let isViewingDuplicates = type == "duplicates";

View file

@ -102,7 +102,7 @@
}
get _builtInPanes() {
return ["info", "abstract", "attachments", "notes", "note-info", "attachment-info", "attachment-annotations", "libraries-collections", "tags", "related"];
return ["info", "abstract", "attachments", "notes", "note-info", "attachment-info", "attachment-annotations", "libraries-collections", "tags", "related", "search-results"];
}
get container() {

View file

@ -46,15 +46,14 @@
}
get _searchModes() {
let modes = {
// Best Match ranks with the semantic engine when one is enabled
// and with the lexical engine otherwise, so it's always offered
return {
titleCreatorYear: Zotero.getString('quickSearch.mode.titleCreatorYear'),
fields: Zotero.getString('quickSearch.mode.fieldsAndTags'),
everything: Zotero.getString('quickSearch.mode.everything')
everything: Zotero.getString('quickSearch.mode.everything'),
bestMatch: Zotero.getString('quickSearch-mode-best-match')
};
if (Zotero.Embeddings.isEnabled()) {
modes.bestMatch = Zotero.getString('quickSearch-mode-best-match');
}
return modes;
}
_searchModePopup = null;
@ -68,9 +67,9 @@
}
disconnectedCallback() {
if (this._modelPrefObserverID) {
Zotero.Prefs.unregisterObserver(this._modelPrefObserverID);
this._modelPrefObserverID = null;
if (this._semanticSearchObserverID) {
Zotero.Prefs.unregisterObserver(this._semanticSearchObserverID);
this._semanticSearchObserverID = null;
}
}
@ -138,17 +137,11 @@
this._advancedButton = advancedButton;
}
// Disabling semantic search in the preferences invalidates an active
// best-match mode: fall back to Fields & Tags and rerun any active
// search under the new mode
this._modelPrefObserverID = Zotero.Prefs.registerObserver('embeddings.model', () => {
if (Zotero.Prefs.get('search.quicksearch-mode') === 'bestMatch'
&& !Zotero.Embeddings.isEnabled()) {
Zotero.Prefs.set('search.quicksearch-mode', 'fields');
this.updateMode();
if (this.value) {
this.dispatchEvent(new Event('command'));
}
// Turning semantic ranking on or off changes what an active
// best-match search ranks with: rerun it under the new engine
this._semanticSearchObserverID = Zotero.Prefs.registerObserver('search.bestMatch.enableSemantic', () => {
if (Zotero.Prefs.get('search.quicksearch-mode') === 'bestMatch' && this.value) {
this.dispatchEvent(new Event('command'));
}
});

View file

@ -0,0 +1,172 @@
/*
***** BEGIN LICENSE BLOCK *****
Copyright © 2026 Corporation for Digital Scholarship
Vienna, Virginia, USA
https://www.zotero.org
This file is part of Zotero.
Zotero is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Zotero is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with Zotero. If not, see <http://www.gnu.org/licenses/>.
***** END LICENSE BLOCK *****
*/
"use strict";
{
// A best-match search result card: one passage of an item's text that the
// query matched (see Zotero.BestMatch.Session#getPreviews()) -- where the
// passage sits in the document as the head, the passage itself as the
// quote with the query's words marked in it, presented like an
// annotation-row (the two share their styling, see
// scss/elements/_annotationRow.scss).
//
// The whole passage is quoted, not the line the tree shows: the card is
// where a match is read rather than scanned. A quote too tall for the
// card is clamped, with a toggle to see the rest.
class SearchResultRow extends XULElementBase {
content = MozXULElement.parseXULToFragment(`
<html:div class="head">
<html:div class="title">
<html:span class="path"/>
</html:div>
<html:div class="location"/>
</html:div>
<html:div class="body">
<html:div class="quote"/>
<html:button class="show-more" data-l10n-id="search-result-row-show-more" hidden="true"/>
</html:div>
`);
_result = null;
// Whether to mark the line of the passage the tree quotes -- its
// snippet -- where the passage is shown in place of that row
highlightSnippet = false;
get result() {
return this._result;
}
set result(result) {
this._result = result;
this.render();
}
init() {
this._path = this.querySelector('.path');
this._location = this.querySelector('.location');
this._quote = this.querySelector('.quote');
this._showMore = this.querySelector('.show-more');
this._showMore.addEventListener('click', (event) => {
// The card's activation (open the attachment) shouldn't fire
// for the toggle
event.stopPropagation();
this._toggleExpanded();
});
this.render();
}
render() {
if (!this.initialized || !this._result) return;
// Where the passage sits: the headings it falls under, or the
// generic fulltext label for a passage from a document with no
// outline to read
if (this._result.outlinePath) {
this._path.removeAttribute('data-l10n-id');
this._path.textContent = this._result.outlinePath;
}
else {
document.l10n.setAttributes(this._path, 'search-result-row-fulltext');
}
// The page the chunk's section starts on, labeled the way
// annotation rows label theirs
this._location.hidden = !this._result.pageLabel;
if (this._result.pageLabel) {
this._location.textContent
= Zotero.getString('pdfReader.page') + ' ' + this._result.pageLabel;
}
this._renderQuote();
// Offer "Show More" only when the quote is actually clamped,
// which is only measurable once the card has a layout
this.classList.remove('expanded');
this._showMore.hidden = true;
requestAnimationFrame(() => {
this._showMore.hidden
= this._quote.scrollHeight <= this._quote.clientHeight;
});
// A11y - make focusable and describe the card
this.setAttribute('tabindex', 0);
this.setAttribute('aria-label', [
this._result.outlinePath,
this._location.hidden ? '' : this._location.textContent,
this._result.text
].filter(Boolean).join('. '));
}
// The passage's text, with any matched ranges wrapped for highlighting,
// and its snippet wrapped too when asked to mark it
_renderQuote() {
let text = this._result.text || '';
let snippet = this.highlightSnippet ? this._result.snippet : null;
this._quote.replaceChildren();
if (!snippet) {
this._appendWithMatches(this._quote, text, 0, text.length);
return;
}
let marked = document.createElement('span');
marked.className = 'snippet';
this._appendWithMatches(this._quote, text, 0, snippet.start);
this._appendWithMatches(marked, text, snippet.start, snippet.end);
this._quote.append(marked);
this._appendWithMatches(this._quote, text, snippet.end, text.length);
}
// Append text[start, end) to a node, the matched ranges in it wrapped
_appendWithMatches(node, text, start, end) {
let position = start;
for (let [rangeStart, rangeEnd] of this._result.ranges || []) {
rangeStart = Math.max(rangeStart, start);
rangeEnd = Math.min(rangeEnd, end);
if (rangeStart >= rangeEnd) {
continue;
}
if (rangeStart > position) {
node.append(text.slice(position, rangeStart));
}
let match = document.createElement('span');
match.className = 'match';
match.textContent = text.slice(rangeStart, rangeEnd);
node.append(match);
position = rangeEnd;
}
if (position < end) {
node.append(text.slice(position, end));
}
}
_toggleExpanded() {
let expanded = this.classList.toggle('expanded');
document.l10n.setAttributes(this._showMore,
expanded ? 'search-result-row-show-less' : 'search-result-row-show-more');
}
}
customElements.define('search-result-row', SearchResultRow);
}

View file

@ -0,0 +1,171 @@
/*
***** BEGIN LICENSE BLOCK *****
Copyright © 2026 Corporation for Digital Scholarship
Vienna, Virginia, USA
https://www.zotero.org
This file is part of Zotero.
Zotero is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Zotero is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with Zotero. If not, see <http://www.gnu.org/licenses/>.
***** END LICENSE BLOCK *****
*/
{
const { ItemPaneSectionElementBase } = ChromeUtils.importESModule(
"chrome://zotero/content/elements/itemPaneSectionElementBase.mjs",
{ global: "current" }
);
// Why the selected item matched the active best-match search: a card per
// passage of the item's own text that the search matched (see
// Zotero.BestMatch.Session#getPreviews()), each carrying the whole
// passage rather than the line the tree quotes, so a match can be read
// without opening anything.
//
// Shows every passage the selected item matched in. A single passage
// selected on its own is a search-results-pane, not an item with a
// section.
class SearchResultsBox extends ItemPaneSectionElementBase {
content = MozXULElement.parseXULToFragment(`
<collapsible-section data-l10n-id="section-search-results" data-pane="search-results">
<html:div class="body">
</html:div>
</collapsible-section>
`);
get item() {
return this._item;
}
set item(item) {
super.item = item instanceof Zotero.Item ? item : null;
// A new item's emptiness isn't known until asyncRender scores it
this._count = undefined;
}
get collectionTreeRows() {
return super.collectionTreeRows;
}
// The item pane sets collectionTreeRows after item, so this is where
// everything visibility depends on is finally known
set collectionTreeRows(collectionTreeRows) {
super.collectionTreeRows = collectionTreeRows;
this._updateHidden();
}
init() {
this.initCollapsibleSection();
this._body = this.querySelector('.body');
// The header's count placeholder needs a value before the first
// async render fills in the real one
this._section.setCount(0);
// Double-click (or Enter on a focused card) opens the attachment
// at the chunk
this._body.addEventListener('dblclick', this._handleActivate);
this._body.addEventListener('keydown', (event) => {
if (event.key == 'Enter') {
this._handleActivate(event);
}
});
}
get _itemsView() {
return this.closest('item-pane')?.itemsView ?? null;
}
// The session holding the passages of the active best-match search,
// or null when no such search is running. It derives a preview once
// per item and keeps it, so reading one costs nothing after the first
// time.
get _session() {
return this._itemsView?.bestMatchSession ?? null;
}
// A new search re-renders even when the item didn't change
get _renderDependencies() {
return [...super._renderDependencies, this._session];
}
render() {}
async asyncRender() {
if (!this.initialized) return;
if (this._isAlreadyRendered("async")) return;
let item = this.item;
let session = this._session;
this._body.replaceChildren();
if (!item || !session) {
this._count = 0;
this._updateHidden();
return;
}
let preview = session.getPreviews(item.id);
if (preview?.state == 'pending') {
// Previews are derived before the tree's rows appear, but a
// selection can still land mid-re-score, so settle this one
// outright
await session.fill([item.id]);
// The selection, or the search, may have moved on while
// deriving
if (this.item !== item || this._session !== session) {
return;
}
preview = session.getPreviews(item.id);
}
let entries = preview?.state == 'filled' ? preview.entries : [];
this._count = entries.length;
this._section.setCount(entries.length);
this._updateHidden();
// Left in the order the preview holds them, strongest match
// first: with only a handful of cards shown, the best one earning
// the top slot matters more than reading them in document order
for (let entry of entries) {
let row = document.createXULElement('search-result-row');
row.result = entry;
this._body.append(row);
}
}
// Open the activated card's attachment at its passage
_handleActivate = (event) => {
let row = event.target.closest('search-result-row');
// The Show More toggle isn't an activation
if (!row || !this.item || event.target.closest('.show-more')) {
return;
}
if (typeof ZoteroPane == 'undefined') {
return;
}
ZoteroPane.viewSearchMatch(this.item.id, row.result, event)
.catch(e => Zotero.logError(e));
};
_updateHidden() {
// Visible only during a best-match search; asyncRender hides it
// again when nothing matched. Deciding emptiness needs the async
// derivation, so unlike the annotations section this one can't
// know its final state synchronously -- it appears, then empties
// out, rather than flickering in late.
this.hidden = !this.item || !this._session || this.tabType == 'reader'
|| this._count === 0;
}
}
customElements.define("search-results-box", SearchResultsBox);
}

View file

@ -0,0 +1,132 @@
/*
***** BEGIN LICENSE BLOCK *****
Copyright © 2026 Corporation for Digital Scholarship
Vienna, Virginia, USA
https://www.zotero.org
This file is part of Zotero.
Zotero is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Zotero is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with Zotero. If not, see <http://www.gnu.org/licenses/>.
***** END LICENSE BLOCK *****
*/
"use strict";
{
// The pane shown when what's selected is search matches rather than
// items: a card per selected passage (see
// Zotero.BestMatch.Session#getPreviews()), grouped under the attachment
// each came from.
//
// A passage isn't an item, so nothing an item pane says about one -- its
// fields, its attachments, its tags -- has anything to describe. What
// there is to show is the passage itself.
class SearchResultsPane extends XULElementBase {
content = MozXULElement.parseXULToFragment(`
<html:div class="custom-head"></html:div>
<html:div class="body zotero-view-item"></html:div>
`);
_matches = [];
// @param {Object[]} matches - { itemID, entry }, in the order they're shown
set matches(matches) {
this._matches = matches || [];
}
get matches() {
return this._matches;
}
init() {
this._body = this.querySelector('.body');
// Double-click, or Enter on a focused card, opens the attachment
// at the passage
this._body.addEventListener('dblclick', this._handleActivate);
this._body.addEventListener('keydown', (event) => {
if (event.key == 'Enter') {
this._handleActivate(event);
}
});
}
render() {
if (!this.initialized) return;
this._body.replaceChildren();
// Grouped by attachment, in the order the matches arrive, so the
// pane reads in the order the rows do
let byItem = new Map();
for (let match of this._matches) {
if (!byItem.has(match.itemID)) {
byItem.set(match.itemID, []);
}
byItem.get(match.itemID).push(match.entry);
}
for (let [itemID, entries] of byItem) {
let item = Zotero.Items.get(itemID);
let section = document.createXULElement('collapsible-section');
section.dataset.l10nId = 'section-search-results';
section.dataset.pane = `search-results-${itemID}`;
section.summary = item ? item.getDisplayTitle() : '';
document.l10n.setArgs(section, { count: entries.length });
let body = document.createElement('div');
body.className = 'body';
section.append(body);
this._body.append(section);
for (let entry of entries) {
let row = document.createXULElement('search-result-row');
// Shown for the row it quotes, so that line is marked
row.highlightSnippet = true;
row.result = entry;
row.dataset.itemId = itemID;
body.append(row);
}
}
}
// The buttons the pane's host puts above the cards, if any
renderCustomHead(callback) {
let customHead = this.querySelector(".custom-head");
customHead.replaceChildren();
if (callback) {
callback({
doc: document,
append: (...args) => customHead.append(...args),
});
}
}
// Open the activated card's attachment at its passage
_handleActivate = (event) => {
let row = event.target.closest('search-result-row');
// The Show More toggle isn't an activation
if (!row || event.target.closest('.show-more')) {
return;
}
if (typeof ZoteroPane == 'undefined') {
return;
}
ZoteroPane.viewSearchMatch(parseInt(row.dataset.itemId), row.result, event)
.catch(e => Zotero.logError(e));
};
}
customElements.define("search-results-pane", SearchResultsPane);
}

View file

@ -639,13 +639,10 @@
this._lastTopKValue = this.bestMatchTopKInput.value;
}
// The best-match field is a root-level modifier, offered only for
// top-level item results and only when semantic search is enabled (or
// the search already carries a query, so it stays editable)
// The best-match field is a root-level modifier, offered for
// top-level item results
updateBestMatchRow() {
this.bestMatchRow.hidden = !this.isRoot
|| this.resultLevel != 'item'
|| !(Zotero.Embeddings.isEnabled() || this.bestMatchInput.value);
this.bestMatchRow.hidden = !this.isRoot || this.resultLevel != 'item';
}
set resultLevel(val) {

View file

@ -31,7 +31,7 @@ const LibraryTree = require('./libraryTree');
const VirtualizedTable = require('components/virtualized-table');
const { VirtualizedTree, formatColumnName } = VirtualizedTable;
const { COLUMNS } = require("zotero/itemTreeColumns");
const { ItemTreeRow } = require('zotero/itemTreeRow');
const { ItemTreeRow, SearchMatch } = require('zotero/itemTreeRow');
const { OS } = ChromeUtils.importESModule("chrome://zotero/content/osfile.mjs");
const { ZOTERO_CONFIG } = ChromeUtils.importESModule('resource://zotero/config.mjs');
@ -107,6 +107,9 @@ class ItemTreeRowProvider {
this._searchItemIDs = new Set();
this._searchParentIDs = new Set();
this._includeTrashed = false;
// The best-match search session whose previews rows show as match
// children (see SearchMatch in itemTreeRow.js), while one is active
this._bestMatchSession = null;
this.onUpdate = this.createEventBinding('update');
}
@ -231,6 +234,7 @@ class ItemTreeRowProvider {
}
return row.isContainerEmpty({
includeTrashed: this._includeTrashed,
getMatchPreviews: this._bestMatchSession?.getPreviews,
});
}
@ -248,8 +252,22 @@ class ItemTreeRowProvider {
_refreshContainer(index, skipRowMapRefresh = false) {
if (!this.isContainer(index)) return;
// Reopening recreates child rows closed, so remember which
// descendants were open and reopen them afterward
let level = this.getLevel(index);
let openDescendantIDs = [];
for (let i = index + 1; i < this._rows.length && this.getLevel(i) > level; i++) {
if (this.isContainer(i) && this.isContainerOpen(i)) {
openDescendantIDs.push(this.getRow(i).id);
}
}
this._closeContainer(index, true);
this._openContainer(index, true);
if (openDescendantIDs.length) {
// _restoreOpenState() looks rows up by id
this.refreshRowMap();
this._restoreOpenState(openDescendantIDs);
}
if (!skipRowMapRefresh) {
this.refreshRowMap();
}
@ -282,6 +300,7 @@ class ItemTreeRowProvider {
searchItemIDs: this._searchItemIDs,
includeTrashed: this._includeTrashed,
filterChildItems: this.itemTree.props.filterChildItems,
getMatchPreviews: this._bestMatchSession?.getPreviews,
});
let childRows = childRefs.map(ref => this.createRow(ref, level + 1, false));
@ -1270,6 +1289,57 @@ var ItemTree = class ItemTree extends LibraryTree {
this._cachedScrollPosition = this._saveScrollPosition();
}
/**
* A search-match row shows two lines -- where its passage is, and the
* line of it worth reading -- so it stands a line taller than the rows
* around it. Neither line wraps, so the height is the same for every one
* of them and can be told without measuring anything.
*
* @return {Number}
*/
_getSearchMatchRowHeight() {
let textHeight = this.tree._renderedTextHeight
* (this.tree.props.disableFontSizeScaling ? 1 : Zotero.Prefs.get('fontSize'));
// Two lines, and half the room a one-line row leaves around its text:
// the second line gives the row enough weight without it
let padding = (this.tree._rowHeight - textHeight) / 2;
return Math.round(textHeight * 2 + padding);
}
/**
* Tell the table which rows are taller than the rest.
*
* The heights are keyed by row index, so they mean something different
* after every insertion, removal and sort -- which is why this runs from
* handleRowModelUpdate(), where all of those end up, rather than from the
* places that change rows.
*/
_updateSearchMatchRowHeights() {
if (!this.tree?._jsWindow) {
return;
}
let height = null;
let heights = [];
let indexes = [];
for (let i = 0, count = this.getRowCount(); i < count; i++) {
let type = this.getRow(i)?.type;
if (type == 'search-match') {
height ??= this._getSearchMatchRowHeight();
heights.push([i, height]);
indexes.push(i);
}
}
// Most updates leave the tall rows exactly where they were. Telling
// the table again would have it rebuild its offsets and forget which
// way the view was moving for nothing.
let signature = height + '|' + indexes.join(',');
if (signature === this._searchMatchRowHeights) {
return;
}
this._searchMatchRowHeights = signature;
this.tree.updateCustomRowHeights(heights);
}
/**
* NOTE: This method must not trigger further update events (e.g. by calling
* sort() or refresh()) to avoid recursive update loops and UI flashing.
@ -1281,6 +1351,8 @@ var ItemTree = class ItemTree extends LibraryTree {
* @param {boolean} options.restoreSelection - Whether to restore the cached selection.
* @param {boolean} options.ensureRowsAreVisible - Whether to ensure selected rows are visible.
* @param {boolean} options.restoreScroll - Whether to restore the cached scroll position.
* @param {boolean} options.scrollToTop - Whether to show the list from the top, ignoring
* the cached scroll position.
* @param {boolean} options.loading - Whether to show loading state (hides tree, shows message).
* @param {string} options.message - Optional message to display (for loading, errors, intro text).
*/
@ -1313,6 +1385,10 @@ var ItemTree = class ItemTree extends LibraryTree {
this._treebox && this._treebox.scrollTo(0);
}
// Before anything is drawn or scrolled to: the rows just changed, and
// heights are what say where each one sits
this._updateSearchMatchRowHeights();
if (rows === true) {
if (this.tree) {
this.tree.invalidate();
@ -1339,7 +1415,8 @@ var ItemTree = class ItemTree extends LibraryTree {
const itemsViewInActiveWindow = Zotero.getActiveZoteroPane()?.itemsView == this;
const prioritizeRestore = !(options.selectInActiveWindow && itemsViewInActiveWindow);
const ensureVisible = options.restoreScroll ? false : options.ensureRowsAreVisible;
const ensureVisible = options.restoreScroll || options.scrollToTop
? false : options.ensureRowsAreVisible;
if (prioritizeRestore && options.restoreSelection) {
this._restoreSelection(null, options.expandCollapsedParents, ensureVisible);
@ -1353,7 +1430,10 @@ var ItemTree = class ItemTree extends LibraryTree {
}
}
if (options.restoreScroll) {
if (options.scrollToTop) {
this._treebox?.scrollTo(0);
}
else if (options.restoreScroll) {
this._restoreScrollPosition();
}
@ -1896,6 +1976,31 @@ var ItemTree = class ItemTree extends LibraryTree {
}
}
/**
* The session holding the passages of the active best-match search, when
* the view's rows come from one
*
* @return {Zotero.BestMatch.Session|null}
*/
get bestMatchSession() {
return this.rowProvider?.bestMatchSession ?? null;
}
/**
* The passages the selection names, when search-match rows are all it
* holds. Empty for any selection with something else in it, so a caller
* can tell "these are passages" from "these are items".
*
* @return {Object[]} - { itemID, entry } per selected passage
*/
getSelectedSearchMatches() {
let selected = this.getSelectedObjects();
if (!selected.length || !selected.every(ref => ref instanceof SearchMatch)) {
return [];
}
return selected.map(ref => ({ itemID: ref.itemID, entry: ref.entry }));
}
/**
* Get selected items, omitting collections and searches in the trash
*/
@ -2330,6 +2435,7 @@ var ItemTree = class ItemTree extends LibraryTree {
div.classList.toggle('first-highlighted', this._highlightedRows.has(rowData.id) && !this._highlightedRows.has(prevRowID));
div.classList.toggle('last-highlighted', this._highlightedRows.has(rowData.id) && !this._highlightedRows.has(nextRowID));
div.classList.toggle('annotation-row', row.type === 'annotation');
div.classList.toggle('search-match-row', row.type === 'search-match');
div.classList.toggle('library-header-row', row.type === 'library-header');
div.classList.toggle('spacer-row', row.type === 'spacer');
if (row.type !== 'annotation') {
@ -2406,8 +2512,10 @@ var ItemTree = class ItemTree extends LibraryTree {
}
if (isFirstColumn) {
// A row with no icon of its own gets none: the indent and twisty
// the tree adds don't depend on one
const icon = row.getIcon();
icon.classList.add('cell-icon', 'item-icon');
icon?.classList.add('cell-icon', 'item-icon');
if (cell.querySelector('.cell-text') === null) {
let textSpan = document.createElement('span');
@ -2417,7 +2525,9 @@ var ItemTree = class ItemTree extends LibraryTree {
cell.append(textSpan);
}
cell.prepend(icon);
if (icon) {
cell.prepend(icon);
}
cell.classList.add('first-column');
}
@ -2775,6 +2885,10 @@ var ItemTree = class ItemTree extends LibraryTree {
}
col.hidden = false;
col.sortDirection = -1;
// Far right, so the bar lines up across item, note, attachment,
// and annotation rows -- annotation rows use a custom layout
// that puts their bar at the row's end
col.ordinal = Math.max(...this._columns.map(c => c.ordinal ?? 0)) + 1;
this._sortedColumn = col;
}
}
@ -2823,13 +2937,25 @@ var ItemTree = class ItemTree extends LibraryTree {
if (row === undefined) {
return;
}
this._treebox.scrollToRow(Math.max(row - scrollPosition.offset, 0), true);
var topRow = Math.max(row - scrollPosition.offset, 0);
// scrollToRow() aligns a row's top with the viewport's, which throws
// away however far into that row the view had been scrolled. Rows are
// tall enough now for that to read as the list jumping backwards, so
// restore the exact pixel when we know it.
if (scrollPosition.pixelOffset !== undefined) {
this._treebox.scrollTo(
this._treebox._getItemPosition(topRow) + scrollPosition.pixelOffset);
return;
}
this._treebox.scrollToRow(topRow, true);
}
/**
* Return an object describing the current scroll position to restore after changes
*
* @return {Object|Boolean} - Object with .id (a treeViewID) and .offset, or false if no rows
* @return {Object|Boolean} - Object with .id (a treeViewID), .offset (rows between the
* anchor and the top of the view) and .pixelOffset (how far into the top row the
* view is scrolled), or false if no rows
*/
_saveScrollPosition() {
if (!this._treebox) return false;
@ -2838,6 +2964,12 @@ var ItemTree = class ItemTree extends LibraryTree {
if (first === undefined || first === null) {
return false;
}
// How far into the first visible row the view is scrolled. Measured
// against the same offset getFirstVisibleRow() reads, so the two
// always describe the same position.
var pixelOffset = typeof treebox._getItemPosition == 'function'
? treebox.scrollOffset - treebox._getItemPosition(first)
: undefined;
var last = treebox.getLastVisibleRow();
for (let i = first; i <= last; i++) {
// If an object is selected, keep the first selected one in position
@ -2846,7 +2978,8 @@ var ItemTree = class ItemTree extends LibraryTree {
if (!row) return false;
return {
id: row.ref.treeViewID,
offset: i - first
offset: i - first,
pixelOffset
};
}
}
@ -2864,7 +2997,8 @@ var ItemTree = class ItemTree extends LibraryTree {
if (!row) return false;
return {
id: row.ref.treeViewID,
offset: 0
offset: 0,
pixelOffset
};
}

View file

@ -397,9 +397,13 @@ const COLUMNS = [
renderCell(index, data, column, isFirstColumn, doc) {
let cell = doc.createElement('span');
cell.className = `cell ${column.className}`;
let fraction = this.rowProvider.getBestMatchBarFractions()
.get(this.getRow(index).id);
if (fraction !== undefined) {
let row = this.getRow(index);
// Rows that carry their own relevance (e.g. a search-match row,
// showing the strength of the evidence it displays) report it
// themselves; every other row's bar comes from the view's scores
let fraction = this.rowProvider.getBestMatchBarFractions().get(row.id)
?? row.getRelevanceFraction();
if (fraction !== null && fraction !== undefined) {
let bar = doc.createElement('span');
bar.className = 'relevance-bar';
let fill = doc.createElement('span');
@ -408,9 +412,11 @@ const COLUMNS = [
bar.append(fill);
cell.append(bar);
// The rank reaches assistive technology via the row label; show
// it visually as a tooltip
doc.l10n.formatValue('items-column-relevance-rank', { rank: data })
.then(label => cell.title = label);
// it visually as a tooltip. Match rows carry no rank of their own.
if (data) {
doc.l10n.formatValue('items-column-relevance-rank', { rank: data })
.then(label => cell.title = label);
}
}
return cell;
}

View file

@ -12,6 +12,10 @@ XPCOMUtils.defineLazyPreferenceGetter(
const ATTACHMENT_STATE_LOAD_DELAY = 150;
// Headings named above a search match. The path from the document's root can
// be several levels deep, and the deepest are the ones that place the passage.
const LOCATION_HEADINGS = 2;
/**
* Base row in an ItemTree.
*
@ -103,6 +107,17 @@ class ItemTreeRow {
return getCSSItemTypeIcon('document');
}
/**
* The 0-1 fraction the Relevance column's bar shows for this row on its
* own, for rows carrying their own relevance rather than taking it from
* the view's best-match scores, or null for rows that don't
*
* @return {Number|null}
*/
getRelevanceFraction() {
return null;
}
renderRow(div, index, columns, rowData, renderCtx) {
for (let column of columns) {
if (column.hidden) continue;
@ -461,23 +476,30 @@ class FileItemTreeRow extends ZoteroItemTreeRow {
return true;
}
isContainerEmpty() {
isContainerEmpty({ getMatchPreviews } = {}) {
// An attachment with search matches to show can be expanded even
// with no annotations of its own
if (getMatchPreviews?.(this.ref.id)?.state == 'filled') {
return false;
}
return this.ref.numAnnotations() == 0;
}
getChildItems({ searchMode, searchItemIDs } = {}) {
getChildItems({ searchMode, searchItemIDs, getMatchPreviews } = {}) {
let annotations = this.ref.getAnnotations();
// With "Hide Non-Matching Annotations" enabled, if any of the attachment's
// annotations match a search, show only those and hide the rest. If none match,
// show them all, since otherwise the attachment couldn't be expanded to browse its
// annotations at all.
let matchRows = SearchMatch.forItem(this.ref, getMatchPreviews);
// With "Hide Non-Matching Annotations" enabled, show only the annotations
// that matched a search. When none of them did, they're shown anyway rather
// than leaving an attachment that can't be expanded to browse them at all --
// unless it has match rows, which are already something to expand to.
if (searchMode && Zotero.Prefs.get("hideContextAnnotationRows")) {
let matches = annotations.filter(annotation => searchItemIDs.has(annotation.id));
if (matches.length) {
if (matches.length || matchRows.length) {
annotations = matches;
}
}
return annotations;
// Fulltext match rows come after the annotations
return [...annotations, ...matchRows];
}
_supportsBestAttachmentState() {
@ -562,6 +584,214 @@ class AnnotationItemTreeRow extends ZoteroItemTreeRow {
div.append(cell);
}
}
// The relevance bar while a best-match search shows the Relevance column
let relevanceColumn = columns.find(column => column.dataKey == 'relevance');
if (relevanceColumn && !relevanceColumn.hidden) {
let cell = renderCtx.renderCell(index, rowData?.relevance, relevanceColumn, false);
if (cell) {
div.append(cell);
}
}
}
}
/**
* The reference a search-match row wraps: one place a best-match search
* matched inside an item.
*
* Item tree rows normally wrap data objects. A preview isn't a stored
* object, so this stands in as the tree's reference to one.
*/
class SearchMatch {
constructor(itemID, entry) {
this.itemID = itemID;
// A preview entry (see Zotero.BestMatch.Session#getPreviews())
this.entry = entry;
this.treeViewID = 'SM' + itemID + '-' + entry.key;
this.id = this.treeViewID;
}
/**
* The search-match refs to materialize under an item, from its
* best-match preview: one ref per quoted entry, and nothing when the
* item has no filled preview or its preview derived nothing.
*
* A preview holds every passage the item matched in; the tree shows the
* strongest few, which are the ones with a line quoted. The rest are
* read whole in the item pane.
*
* @param {Zotero.Item} item
* @param {Function} [getMatchPreviews] - itemID -> preview accessor (see
* Zotero.BestMatch.Session#getPreviews()), passed by the row
* provider while a best-match search is active
* @return {SearchMatch[]}
*/
static forItem(item, getMatchPreviews) {
let preview = getMatchPreviews?.(item.id);
if (preview?.state != 'filled') {
return [];
}
return preview.entries
.slice(0, Zotero.BestMatch.MAX_QUOTED_PASSAGES)
.map(entry => new SearchMatch(item.id, entry));
}
}
/**
* Row showing one place a best-match search matched inside its parent row's
* item: a derived excerpt with its matches highlighted. The ref is a
* SearchMatch carrying the preview entry it shows.
*/
class SearchMatchItemTreeRow extends ItemTreeRow {
get type() {
return 'search-match';
}
/**
* The line of the passage this row shows: its snippet, with ellipses
* where it cuts -- none before a line opening on a sentence -- and the
* query's matches located within it. The whole passage stays on the
* entry.
*
* @return {Object} - { text, ranges }
*/
getQuotedLine() {
let { text, ranges, snippet } = this.ref.entry;
let start = snippet ? snippet.start : 0;
let end = snippet ? snippet.end : text.length;
let prefix = start > 0 && !snippet?.startsSentence ? '…' : '';
let quoted = prefix + text.slice(start, end) + (end < text.length ? '…' : '');
let quotedRanges = [];
for (let [rangeStart, rangeEnd] of ranges || []) {
let from = Math.max(rangeStart, start);
let to = Math.min(rangeEnd, end);
if (from >= to) {
continue;
}
quotedRanges.push([
from - start + prefix.length,
to - start + prefix.length
]);
}
return { text: quoted, ranges: quotedRanges };
}
getDisplayTitle() {
return this.getQuotedLine().text;
}
getField(field) {
if (field == 'title') {
return this.getDisplayTitle();
}
return super.getField(field);
}
/**
* Where in the document this row's passage sits, for the line above the
* quote: the headings it falls under and the page it starts on. Only the
* deepest headings are named -- a full outline path is longer than the
* line, and the leaf is what says where you'd land.
*
* Empty for a passage that knows neither, which is what a document with
* no structured text to read gives.
*
* @return {String}
*/
getLocationLabel() {
let { outlinePath, pageLabel } = this.ref.entry;
let parts = [];
if (outlinePath) {
parts.push(outlinePath.split(' > ').slice(-LOCATION_HEADINGS).join(' › '));
}
if (pageLabel) {
parts.push(Zotero.ftl.formatValueSync(
'items-search-match-page', { page: pageLabel }));
}
return parts.join(' · ');
}
/**
* A match row's bar shows the strength of the evidence it displays,
* rather than its item's relevance
*/
getRelevanceFraction() {
return this.ref.entry?.strength ?? null;
}
/**
* No icon: every match row would carry the same one, which would say
* nothing while taking room from the quote
*/
getIcon() {
return null;
}
/**
* A match row spans the tree's whole width with one cell. It shows a
* passage rather than an item, so the columns describe nothing about it
* -- including the relevance bar, which would rank passages against each
* other where the eye is meant to be reading them.
*/
renderRow(div, index, columns, rowData, renderCtx) {
let titleColumn = Object.assign(
{},
columns.find(column => column.dataKey == 'title'),
{ className: 'title' }
);
div.appendChild(renderCtx.renderCell(index, rowData.title, titleColumn, true));
}
/**
* Stack the row's lines beside the tree's indent and twisty, which are
* added to the first cell of every row and would otherwise be stacked
* along with them
*
* @param {...Element} lines
* @return {Element}
*/
_renderLines(...lines) {
let wrapper = document.createElement('span');
wrapper.className = 'search-match-lines';
wrapper.append(...lines);
return wrapper;
}
/**
* Two lines: where the passage is, and the line of it worth reading.
* Neither wraps, so every match row is the same height and the tree can
* tell what that height is without measuring (see
* ItemTree#_getSearchMatchRowHeight()).
*/
renderPrimaryCell(index, data, column) {
let span = document.createElement('span');
span.className = `cell ${column.className} primary`;
let locationSpan = document.createElement('span');
locationSpan.className = 'search-match-location';
locationSpan.textContent = this.getLocationLabel();
let textSpan = document.createElement('span');
textSpan.className = 'cell-text';
let { text, ranges } = this.getQuotedLine();
let last = 0;
for (let [start, end] of ranges || []) {
if (start > last) {
textSpan.append(text.slice(last, start));
}
let mark = document.createElement('span');
mark.className = 'search-match-highlight';
mark.textContent = text.slice(start, end);
textSpan.append(mark);
last = end;
}
if (last < text.length) {
textSpan.append(text.slice(last));
}
span.append(this._renderLines(locationSpan, textSpan));
return span;
}
}
@ -742,6 +972,7 @@ class SpacerItemTreeRow extends ItemTreeRow {
ItemTreeRow.create = function (ref, level, isOpen) {
if (ref instanceof Zotero.Collection) return new CollectionItemTreeRow(ref, level, isOpen);
if (ref instanceof Zotero.Search) return new SearchItemTreeRow(ref, level, isOpen);
if (ref instanceof SearchMatch) return new SearchMatchItemTreeRow(ref, level, isOpen);
if (ref.isAnnotation?.()) return new AnnotationItemTreeRow(ref, level, isOpen);
if (ref.isFileAttachment?.()) return new FileItemTreeRow(ref, level, isOpen);
return new ZoteroItemTreeRow(ref, level, isOpen);
@ -752,6 +983,8 @@ module.exports.ItemTreeRow = ItemTreeRow;
module.exports.ZoteroItemTreeRow = ZoteroItemTreeRow;
module.exports.FileItemTreeRow = FileItemTreeRow;
module.exports.AnnotationItemTreeRow = AnnotationItemTreeRow;
module.exports.SearchMatchItemTreeRow = SearchMatchItemTreeRow;
module.exports.SearchMatch = SearchMatch;
module.exports.CollectionItemTreeRow = CollectionItemTreeRow;
module.exports.SearchItemTreeRow = SearchItemTreeRow;
module.exports.SpacerItemTreeRow = SpacerItemTreeRow;

View file

@ -0,0 +1,182 @@
<?xml version="1.0"?>
<!--
***** BEGIN LICENSE BLOCK *****
Copyright © 2026 Corporation for Digital Scholarship
Vienna, Virginia, USA
https://www.zotero.org
This file is part of Zotero.
Zotero is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Zotero is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with Zotero. If not, see <http://www.gnu.org/licenses/>.
***** END LICENSE BLOCK *****
-->
<?xml-stylesheet href="chrome://global/skin/" type="text/css"?>
<?xml-stylesheet href="chrome://zotero/skin/zotero.css" type="text/css"?>
<?xml-stylesheet href="chrome://zotero/skin/preferences.css"?>
<?xml-stylesheet href="chrome://zotero-platform/content/zotero.css"?>
<window
xmlns="http://www.mozilla.org/keymaster/gatekeeper/there.is.only.xul"
xmlns:html="http://www.w3.org/1999/xhtml"
data-l10n-id="preferences-advanced-semantic-search-endpoint-dialog"
onload="init()">
<dialog
id="zotero-embeddings-endpoint"
orient="vertical"
buttons="cancel,accept,extra2"
style="padding: 20px 15px; min-width: 620px">
<script src="chrome://global/content/globalOverlay.js"/>
<script src="chrome://zotero/content/include.js"/>
<html:link rel="localization" href="preferences.ftl"/>
<vbox id="zotero-embeddings-endpoint-container">
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-intro"/>
<separator class="thin"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-url" control="endpoint-url"/>
<html:input type="text" id="endpoint-url"/>
<description class="endpoint-note"
data-l10n-id="preferences-advanced-semantic-search-endpoint-privacy"/>
<separator class="thin"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-requirements"/>
<html:div class="endpoint-facts">
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-model"/>
<label id="endpoint-model"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-pooling"/>
<label id="endpoint-pooling"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-file"/>
<label id="endpoint-file"/>
</html:div>
<separator class="thin"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-llama"/>
<hbox align="center">
<html:textarea id="endpoint-command" class="endpoint-command" readonly="readonly" rows="2"/>
<button id="endpoint-copy" data-l10n-id="preferences-advanced-semantic-search-endpoint-copy"/>
</hbox>
<label id="endpoint-command-url" class="endpoint-note"/>
<separator class="thin"/>
<hbox id="endpoint-progress" align="center" hidden="true">
<image class="zotero-spinner-16"/>
<label data-l10n-id="preferences-advanced-semantic-search-endpoint-verifying"/>
</hbox>
<description id="endpoint-error" class="endpoint-error" hidden="true"/>
</vbox>
<script>
let io;
let urlInput;
let dialog;
function init() {
io = window.arguments[0];
urlInput = document.getElementById('endpoint-url');
dialog = document.getElementById('zotero-embeddings-endpoint');
let { command, url } = Zotero.Embeddings.Endpoint.getCommand();
let stored = Zotero.Prefs.get('embeddings.endpoint');
// Where the suggested command serves, so a server started as
// suggested is accepted as is
urlInput.value = stored || url;
document.l10n.setAttributes(dialog.getButton('accept'),
'preferences-advanced-semantic-search-endpoint-accept');
// Offered only while there is an endpoint to remove
let remove = dialog.getButton('extra2');
document.l10n.setAttributes(remove, 'preferences-advanced-semantic-search-endpoint-remove');
remove.hidden = !stored;
document.addEventListener('dialogextra2', removeEndpoint);
let serving = Zotero.Embeddings.getServing();
document.getElementById('endpoint-model').value = Zotero.Embeddings.getModelName();
document.getElementById('endpoint-pooling').value = Zotero.Embeddings.getPooling();
document.getElementById('endpoint-file').value = `${serving.gguf} (${serving.quant})`;
document.getElementById('endpoint-command').value = command;
document.l10n.setAttributes(document.getElementById('endpoint-command-url'),
'preferences-advanced-semantic-search-endpoint-then-use', { url });
document.getElementById('endpoint-copy').addEventListener('command', () => {
Zotero.Utilities.Internal.copyTextToClipboard(command);
});
// Verification is asynchronous: hold the dialog open until it
// has an answer
document.addEventListener('dialogaccept', (event) => {
event.preventDefault();
accept();
});
urlInput.addEventListener('input', () => showError(null));
urlInput.focus();
}
// Stop using the endpoint. Nothing stored is affected: the vectors
// it produced match the model's, and a later endpoint verifies anew.
function removeEndpoint() {
Zotero.Prefs.clear('embeddings.endpoint');
io.ok = true;
window.close();
}
async function accept() {
let url = urlInput.value.trim();
// An emptied URL turns the endpoint off; nothing to verify
if (!url) {
removeEndpoint();
return;
}
setBusy(true);
showError(null);
let verdict;
try {
verdict = await Zotero.Embeddings.Endpoint.verify(url);
}
catch (e) {
Zotero.logError(e);
verdict = { state: 'unreachable' };
}
setBusy(false);
if (verdict.state != 'ok') {
showError(verdict);
return;
}
Zotero.Prefs.set('embeddings.endpoint', url);
io.ok = true;
window.close();
}
function setBusy(busy) {
document.getElementById('endpoint-progress').hidden = !busy;
dialog.getButton('accept').disabled = busy;
dialog.getButton('cancel').disabled = busy;
dialog.getButton('extra2').disabled = busy;
urlInput.disabled = busy;
}
function showError(verdict) {
let box = document.getElementById('endpoint-error');
box.hidden = !verdict;
if (!verdict) {
return;
}
document.l10n.setAttributes(box,
`preferences-advanced-semantic-search-endpoint-error-${verdict.state}`,
{ pooling: Zotero.Embeddings.getPooling() });
}
</script>
</dialog>
</window>

View file

@ -63,33 +63,31 @@ Zotero_Preferences.Advanced = {
initSemanticSearch: function () {
// Populate the model menu from the model registry. The menu isn't
// bound to the pref, since a mode change needs a confirmation prompt
// before the pref -- and with it the stored index -- is touched.
let modelMenu = document.getElementById('semantic-search-model');
let modelPopup = modelMenu.querySelector('menupopup');
for (let { name, l10nID } of Zotero.Embeddings.getAvailableModels()) {
let menuitem = document.createXULElement('menuitem');
menuitem.setAttribute('value', name);
document.l10n.setAttributes(menuitem, l10nID);
modelPopup.append(menuitem);
}
modelMenu.value = Zotero.Embeddings.getModelName();
modelMenu.addEventListener('command', () => this.handleSemanticSearchModeChange());
// Bound by hand rather than through the `preference` attribute, which
// would write the menu's string value to a boolean preference
let enableMenu = document.getElementById('semantic-search-enable');
enableMenu.value = String(Zotero.Embeddings.isEnabled());
enableMenu.addEventListener('command', () => {
Zotero.Prefs.set('search.bestMatch.enableSemantic', enableMenu.value === 'true');
});
// Live progress updates from the background indexer
this._semanticSearchListener = status => this.updateSemanticSearchUI(status);
Zotero.Embeddings.Indexing.addProgressListener(this._semanticSearchListener);
let modelPrefObserverID = Zotero.Prefs.registerObserver(
'embeddings.model',
// Full engine threads while the user is watching indexing progress
Zotero.Embeddings.setThreadBoost('prefs-open', true);
let enablePrefObserverID = Zotero.Prefs.registerObserver(
'search.bestMatch.enableSemantic',
() => {
modelMenu.value = Zotero.Embeddings.getModelName();
enableMenu.value = String(Zotero.Embeddings.isEnabled());
this.updateSemanticSearchUI(Zotero.Embeddings.Indexing.getStatus());
}
);
document.getElementById('zotero-prefpane-advanced').addEventListener('unload', () => {
Zotero.Embeddings.Indexing.removeProgressListener(this._semanticSearchListener);
Zotero.Prefs.unregisterObserver(modelPrefObserverID);
Zotero.Embeddings.setThreadBoost('prefs-open', false);
Zotero.Prefs.unregisterObserver(enablePrefObserverID);
});
document.getElementById('semantic-search-resume').addEventListener('command', () => {
@ -100,58 +98,44 @@ Zotero_Preferences.Advanced = {
Zotero.Embeddings.Indexing.stopIndexing();
});
document.getElementById('semantic-search-endpoint-configure').addEventListener('command', () => {
this.openSemanticSearchEndpointDialog();
});
let diagnostics = document.getElementById('semantic-search-diagnostics');
let toggle = document.getElementById('semantic-search-diagnostics-toggle');
toggle.addEventListener('command', () => {
diagnostics.hidden = !diagnostics.hidden;
document.l10n.setAttributes(toggle, diagnostics.hidden
? 'preferences-advanced-semantic-search-diagnostics-show'
: 'preferences-advanced-semantic-search-diagnostics-hide');
});
// The endpoint's stored verdict is read lazily; have it in memory
// before the status line first renders
Zotero.Embeddings.Endpoint.load().then(() => {
this.updateSemanticSearchUI(Zotero.Embeddings.Indexing.getStatus());
});
// Render current state, then compute up-to-date per-library counts
this.updateSemanticSearchUI(Zotero.Embeddings.Indexing.getStatus());
Zotero.Embeddings.Indexing.refreshStatus();
},
handleSemanticSearchModeChange: async function () {
let modelMenu = document.getElementById('semantic-search-model');
let oldValue = Zotero.Embeddings.getModelName();
let newValue = modelMenu.value;
if (newValue === oldValue) {
return;
}
// Leaving an enabled mode wipes the stored index (and, when disabling,
// the downloaded data), so confirm first. Enabling from Disabled just
// starts indexing.
if (Zotero.Embeddings.isEnabled()) {
let [title, text, button] = await document.l10n.formatValues(
(newValue
? ['switch-title', 'switch-text', 'switch-button']
: ['disable-title', 'disable-text', 'disable-button'])
.map(id => ({ id: `preferences-advanced-semantic-search-${id}` }))
);
let ps = Services.prompt;
let index = ps.confirmEx(
window,
title,
text,
ps.BUTTON_POS_0 * ps.BUTTON_TITLE_IS_STRING
+ ps.BUTTON_POS_1 * ps.BUTTON_TITLE_CANCEL
+ ps.BUTTON_POS_1_DEFAULT,
button, null, null, null, {}
);
if (index !== 0) {
modelMenu.value = oldValue;
return;
}
}
Zotero.Prefs.set('embeddings.model', newValue);
},
updateSemanticSearchUI: function (status) {
let statusBox = document.getElementById('semantic-search-status');
statusBox.hidden = !status.enabled;
this.updateSemanticSearchEndpointUI(status.endpoint, status.enabled);
if (!status.enabled) {
return;
}
// Phase / status message
let phaseLabel = document.getElementById('semantic-search-phase');
let hasRemaining = status.libraries.some(lib => lib.indexed < lib.eligible);
// Attachments rather than chunks: the chunk counts cover only the
// documents cut here, so they can't say what's left
let hasRemaining = status.items.done < status.items.total
|| status.attachments.done < status.attachments.total;
if (status.error) {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-error', { error: status.error });
}
@ -171,9 +155,55 @@ Zotero_Preferences.Advanced = {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-downloading');
}
}
// Preparing shows how far through the library's attachments its pass
// has got, since nothing else reports that step
else if (status.phase === 'preparing') {
if (status.sweep.total) {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-preparing-progress', {
done: status.sweep.done,
total: status.sweep.total
});
}
else {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-preparing');
}
}
// The server's attachments are asked for, with how far through the
// library the step has got
else if (status.phase === 'fetching-documents') {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-fetching-documents', {
done: status.sweep.done,
total: status.sweep.total
});
}
// Then the rest are embedded here -- "here" being worth saying only
// when there's a server that might have done it, and whose refusals
// are the reason for some of it
else if (status.phase === 'indexing-documents') {
if (status.server && status.embedWork) {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-indexing-documents-locally', {
count: status.embedWork.own,
declined: status.embedWork.declined
});
}
else {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-indexing-documents');
}
}
else if (status.phase === 'indexing') {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-indexing');
}
// The server couldn't be reached: its attachments wait for the
// retry, which Resume brings forward
else if (status.serverUnreachable && !status.indexing && !status.paused) {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-server-unreachable', {
minutes: Math.max(1, Math.round((status.serverUnreachable.retryAt - Date.now()) / 60000))
});
}
// Between runs, with attachments still to come from the server
else if (status.attachments.awaiting && !status.paused) {
document.l10n.setAttributes(phaseLabel, 'preferences-advanced-semantic-search-awaiting');
}
else {
document.l10n.setAttributes(phaseLabel,
(status.paused || hasRemaining)
@ -182,10 +212,10 @@ Zotero_Preferences.Advanced = {
}
// Offer a manual restart when enabled but not currently indexing and
// indexing is stopped or there's outstanding work (or the last run
// errored out)
// indexing is stopped, there's outstanding work, the server couldn't
// be reached, or the last run errored out
document.getElementById('semantic-search-resume').hidden
= status.indexing || !(status.error || status.paused || hasRemaining);
= status.indexing || !(status.error || status.paused || hasRemaining || status.serverUnreachable);
// Offer to stop indexing while it's running; a requested stop takes
// effect once the current batch finishes, so disable the button in
@ -194,24 +224,122 @@ Zotero_Preferences.Advanced = {
stopButton.hidden = !status.indexing;
stopButton.disabled = !!status.stopping;
// Per-library "indexed / total" counts
let grid = document.getElementById('semantic-search-libraries');
if (grid.childElementCount !== status.libraries.length * 2) {
grid.textContent = '';
for (let i = 0; i < status.libraries.length; i++) {
grid.append(document.createXULElement('label'), document.createXULElement('label'));
}
// One bar for items, notes and annotations, one for attachment full
// text
document.getElementById('semantic-search-items-row').hidden = !status.items.total;
this._updateSemanticSearchBar('items', status.items);
document.getElementById('semantic-search-attachments-row').hidden = !status.attachments.total;
this._updateSemanticSearchBar('attachments', status.attachments);
this._updateSemanticSearchDiagnostics(status);
},
// Whether the model can be served at all, and how the configured
// server stands
updateSemanticSearchEndpointUI: function (endpoint, enabled) {
let row = document.getElementById('semantic-search-endpoint-row');
// Shown only while semantic search is on and the model can be served
row.hidden = !enabled || !Zotero.Embeddings.Endpoint.isSupported();
if (row.hidden) {
return;
}
status.libraries.forEach((lib, i) => {
grid.children[i * 2].setAttribute(
'value',
Zotero.Utilities.Internal.stringWithColon(lib.name)
);
grid.children[i * 2 + 1].setAttribute(
'value',
`${lib.indexed.toLocaleString()} / ${lib.eligible.toLocaleString()}`
);
let label = document.getElementById('semantic-search-endpoint-status');
let id = {
off: 'off',
unknown: null,
unverified: 'unverified',
ok: 'valid',
unreachable: 'unreachable'
}[endpoint.state];
if (id === null) {
label.removeAttribute('data-l10n-id');
label.value = '';
return;
}
document.l10n.setAttributes(label, `preferences-advanced-semantic-search-endpoint-${id || 'invalid'}`);
},
openSemanticSearchEndpointDialog: function () {
let io = { ok: false };
window.openDialog('chrome://zotero/content/preferences/embeddingsEndpoint.xhtml',
'zotero-preferences-embeddingsEndpoint', 'chrome,modal,centerscreen', io);
this.updateSemanticSearchUI(Zotero.Embeddings.Indexing.getStatus());
},
// Key/value rows of pipeline diagnostics for developers, so the labels
// are plain English rather than localized
_updateSemanticSearchDiagnostics: function ({ items, chunks: indexed, diagnostics, eta }) {
let n = (value, digits = 0) => (value ?? 0).toLocaleString(undefined, {
maximumFractionDigits: digits, minimumFractionDigits: digits
});
let pct = value => Math.round((value || 0) * 100) + '%';
let mb = bytes => n(bytes / 1024 / 1024) + ' MB';
let speed = rate => `${n(rate.chunksPerSecond, 1)} chunks/s, ${n(rate.tokensPerSecond)} tokens/s`;
let bucketList = (buckets, total, unit) => buckets.map(({ from, to, count }) => {
let range = from === null ? `< ${n(to)}` : (to === null ? `≥ ${n(from)}` : `${n(from)}–${n(to - 1)}`);
return `${range}${unit}: ${n(count)} (${pct(total ? count / total : 0)})`;
}).join(' · ');
let proc = p => (p ? `${mb(p.memory)}${p.cpu === null ? '' : `, CPU ${p.cpu}%`}` : '—');
let duration = (seconds) => {
let h = Math.floor(seconds / 3600);
let m = Math.floor(seconds % 3600 / 60);
return h ? `${h} h ${m} m` : `${m} m ${Math.floor(seconds % 60)} s`;
};
let rows = [];
let { window, run, engine, processes, slice, chunks } = diagnostics;
rows.push(['Items indexed', `${n(items.done)} / ${n(items.total)}`]);
rows.push(['Chunks for local inference', `${n(indexed.done)} / ${n(indexed.total)}`]);
rows.push(['Local inference ETA', eta === null ? '—' : duration(eta)]);
rows.push(['Throughput (2 min)', window ? speed(window) : '—']);
rows.push(['Inference speed (run)', run ? speed(run) : '—']);
rows.push(['Padding efficiency', window || run
? `${window ? pct(window.paddingEfficiency) : '—'} (2 min), ${run ? pct(run.paddingEfficiency) : '—'} (run)`
: '—']);
rows.push(['Batches (run)', run
? `${n(run.batches)} · ${n(run.chunksPerBatch, 1)} chunks · ${n(run.tokensPerBatch)} tokens avg`
: '—']);
rows.push(['Token budget', `${n(diagnostics.tokenBudget)} · ${n(diagnostics.pressureEvents)} memory-pressure events`]);
rows.push(['Engine threads', `${engine.threads} of ${engine.optimalThreads}`
+ (engine.boosts.length ? ` (boost: ${engine.boosts.join(', ')})` : '')]);
rows.push(['Engine restarts (run)', `memory ${diagnostics.restarts.memory}, threads ${diagnostics.restarts.threads}`]);
rows.push(['Inference process', proc(processes?.inference)]);
rows.push(['Main process', proc(processes?.main)]);
rows.push(['Available memory', processes?.available ? mb(processes.available) : '—']);
rows.push(['Slice', slice ? `${n(slice.done)} / ${n(slice.total)} chunks` : '—']);
if (chunks) {
let { perDocument } = chunks;
rows.push(['Chunks per document', `${n(perDocument.count)} documents · mean ${n(perDocument.mean, 1)} · median ${n(perDocument.median)} · max ${n(perDocument.max)}`]);
rows.push(['Documents by chunks', bucketList(perDocument.buckets, perDocument.count, '')]);
}
let box = document.getElementById('semantic-search-diagnostics');
while (box.childElementCount < rows.length * 2) {
box.append(document.createXULElement('label'), document.createXULElement('label'));
}
while (box.childElementCount > rows.length * 2) {
box.lastElementChild.remove();
}
rows.forEach(([label, value], i) => {
box.children[i * 2].value = label;
box.children[i * 2 + 1].value = value;
});
},
// Fill one progress bar. The percentage is rounded down, so it reads
// 100% only when everything is done.
_updateSemanticSearchBar: function (name, { done, total }) {
let bar = document.getElementById(`semantic-search-${name}-progress`);
bar.max = Math.max(total, 1);
bar.value = done;
document.l10n.setAttributes(
document.getElementById(`semantic-search-${name}-value`),
'preferences-advanced-semantic-search-progress-value',
{ percent: total ? Math.floor(done / total * 100) : 0 }
);
},
@ -551,15 +679,16 @@ Zotero_Preferences.Advanced = {
document.getElementById('fulltext-stats-notes').textContent = stats.notesIndexed.toLocaleString();
document.getElementById('fulltext-stats-not-available').textContent = stats.notAvailable.toLocaleString();
// Indexed + Partial + indexed notes are already what's in the search index, so they're the
// bar's numerator. Pending work across the auto-draining queues: extracted content not yet
// in the index (remaining), indexable attachments not yet extracted (unindexedQueue), and
// notes not yet indexed or edited since their last index update (noteQueue). Show the bar
// while anything's pending; otherwise "up to date". Items with no local file or content
// aren't counted here -- nothing local can index them -- so "up to date" can sit next to a
// nonzero "not available".
let inIndex = stats.indexed + stats.partial + stats.notesIndexed;
let pending = stats.remaining + stats.unindexedQueue + stats.noteQueue;
// Indexed + Partial + indexed notes + indexed item text are already what's in the search
// index, so they're the bar's numerator. Pending work across the auto-draining queues:
// extracted content not yet in the index (remaining), indexable attachments not yet
// extracted (unindexedQueue), notes not yet indexed or edited since their last index update
// (noteQueue), and items and annotations not yet in the item-text index (itemTextQueue).
// Show the bar while anything's pending; otherwise "up to date". Items with no local file
// or content aren't counted here -- nothing local can index them -- so "up to date" can sit
// next to a nonzero "not available".
let inIndex = stats.indexed + stats.partial + stats.notesIndexed + stats.itemTextIndexed;
let pending = stats.remaining + stats.unindexedQueue + stats.noteQueue + stats.itemTextQueue;
let total = inIndex + pending;
let complete = document.getElementById('fulltext-stats-complete');
if (pending > 0 && total > 0) {
@ -578,6 +707,7 @@ Zotero_Preferences.Advanced = {
await Zotero.FullText.processAttachmentIndexQueue({ maxTime: 500 });
await Zotero.FullText.processAttachmentExtractionQueue({ maxTime: 500 });
await Zotero.FullText.processNoteIndexQueue({ maxTime: 500 });
await Zotero.FullText.processItemTextIndexQueue({ maxTime: 500 });
this._indexStatsTimeoutID = setTimeout(() => this.updateIndexStats(), 250);
}
else {

View file

@ -291,23 +291,47 @@
<label data-l10n-id="fulltext-stats-notes-indexed"/>
<label id="fulltext-stats-notes"/>
</hbox>
<hbox>
<label data-l10n-id="fulltext-stats-items-indexed"/>
<label id="fulltext-stats-items"/>
</hbox>
</vbox>
</groupbox>
<groupbox id="semantic-search">
<label><html:h2 data-l10n-id="preferences-advanced-semantic-search-title"/></label>
<hbox align="center">
<label data-l10n-id="preferences-advanced-semantic-search-model" control="semantic-search-model"/>
<menulist id="semantic-search-model" native="true">
<label data-l10n-id="preferences-advanced-semantic-search-enable"
control="semantic-search-enable"/>
<!-- Bound in initSemanticSearch(): the `preference` attribute
writes the menu's string value, which a boolean pref can't
take -->
<menulist id="semantic-search-enable" native="true">
<menupopup>
<menuitem value="" data-l10n-id="preferences-advanced-semantic-search-disabled"/>
<!-- Model items are populated from Zotero.Embeddings.getAvailableModels() -->
<menuitem value="true"
data-l10n-id="preferences-advanced-semantic-search-enabled"/>
<menuitem value="false"
data-l10n-id="preferences-advanced-semantic-search-disabled"/>
</menupopup>
</menulist>
</hbox>
<description id="semantic-search-model-description"
data-l10n-id="preferences-advanced-semantic-search-model-description"/>
<hbox align="center">
<label data-l10n-id="preferences-advanced-best-match-margin"
control="best-match-margin"/>
<html:input id="best-match-margin" class="html-input" type="number"
min="0" max="100" step="1" size="3"
data-preference="search.bestMatchMargin"/>
</hbox>
<description class="semantic-search-hint"
data-l10n-id="preferences-advanced-best-match-margin-description"/>
<hbox id="semantic-search-endpoint-row" align="center" hidden="true">
<label id="semantic-search-endpoint-status"/>
<button id="semantic-search-endpoint-configure"
data-l10n-id="preferences-advanced-semantic-search-endpoint-configure"/>
</hbox>
<vbox id="semantic-search-status" hidden="true">
<separator class="thin"/>
<hbox align="center">
<label data-l10n-id="preferences-advanced-semantic-search-status"/>
<label id="semantic-search-phase"/>
@ -318,7 +342,42 @@
data-l10n-id="preferences-advanced-semantic-search-stop"
hidden="true"/>
</hbox>
<vbox class="form-grid" id="semantic-search-libraries"/>
<html:div id="semantic-search-progress">
<html:div id="semantic-search-items-row" class="semantic-search-progress-row">
<label data-l10n-id="preferences-advanced-semantic-search-items"/>
<html:progress id="semantic-search-items-progress" value="0" max="1"/>
<label id="semantic-search-items-value"/>
</html:div>
<html:div id="semantic-search-attachments-row" class="semantic-search-progress-row">
<label data-l10n-id="preferences-advanced-semantic-search-attachments"/>
<html:progress id="semantic-search-attachments-progress" value="0" max="1"/>
<label id="semantic-search-attachments-value"/>
</html:div>
</html:div>
<button id="semantic-search-diagnostics-toggle"
data-l10n-id="preferences-advanced-semantic-search-diagnostics-show"/>
<html:div id="semantic-search-diagnostics" hidden="true"/>
</vbox>
<!-- Temporary, for testing: pin best-match search to one engine -->
<vbox>
<separator class="thin"/>
<hbox align="center">
<label data-l10n-id="preferences-advanced-best-match-engine"
control="best-match-engine"/>
<menulist id="best-match-engine"
preference="extensions.zotero.search.bestMatchEngine"
native="true">
<menupopup>
<menuitem value="hybrid"
data-l10n-id="preferences-advanced-best-match-engine-hybrid"/>
<menuitem value="lexical"
data-l10n-id="preferences-advanced-best-match-engine-lexical"/>
<menuitem value="semantic"
data-l10n-id="preferences-advanced-best-match-engine-semantic"/>
</menupopup>
</menulist>
</hbox>
</vbox>
</groupbox>
</vbox>

View file

@ -0,0 +1,978 @@
/*
***** BEGIN LICENSE BLOCK *****
Copyright © 2026 Corporation for Digital Scholarship
Vienna, Virginia, USA
https://www.zotero.org
This file is part of Zotero.
Zotero is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Zotero is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with Zotero. If not, see <http://www.gnu.org/licenses/>.
***** END LICENSE BLOCK *****
*/
/**
* Zotero.BestMatch -- the engine behind best-match search: scoring a set of
* items by relevance to a query. The lexical engine (Zotero.Lexical) always
* scores; when a semantic model is enabled (Zotero.Embeddings), both engines
* score and their rankings are fused with Reciprocal Rank Fusion, so an item
* can match by its words, by its meaning, or -- ranking highest -- by both.
* The facade owns everything a consumer would otherwise need engine
* knowledge for: what counts as an empty query, how results map onto the
* relevance bar, and what the failure modes are.
*/
Zotero.BestMatch = new function () {
// The constant in a Reciprocal Rank Fusion contribution, 1 / (RRF_K +
// rank): high enough that a handful of rank positions in one engine
// can't drown out the other engine's opinion entirely
const RRF_K = 60;
// How a passage's two kinds of evidence weigh against each other. The
// model's reading leads: it is the calibrated signal, and it chose which
// passages are worth showing. Saying the query's own words lifts a
// passage above an equally similar one that only paraphrases them.
const SEMANTIC_WEIGHT = 0.7;
const LEXICAL_WEIGHT = 0.3;
// About a line: what a passage is quoted down to for a one-line preview
const SNIPPET_CHARS = 150;
// A sentence longer than this is quoted as a line of its words rather
// than whole
const LONG_SENTENCE_CHARS = 200;
// The shortest stretch of sentences the static model weighs on its own
const MIN_UNIT_CHARS = 60;
// Most passages quoted for one item. The strongest few already say what
// the item has to offer at a glance; the rest are still derived --
// they're read whole rather than quoted, which needs no line chosen.
const MAX_QUOTED_PASSAGES = 3;
// Previews derived before scoring resolves, in screen order: enough to
// cover the top of the results. The rest follow in the background, since
// reading every matched item's text takes far longer than the ranking.
const PRELOADED_MATCH_PREVIEWS = 10;
// How long background-derived previews accumulate before the consumer
// hears about them, so a pass redraws the view about twice a second
const PREVIEW_BATCH_INTERVAL = 500;
// How long a background derivation waits for an idle main thread before
// running anyway
const PREVIEW_IDLE_TIMEOUT = 1000;
// How long the pass rests after deriving a preview, as a multiple of what
// that preview cost. Deriving reads the item's text; run flat out, that
// work lands inside the frames of whatever the user is doing and
// scrolling stutters.
const PREVIEW_PAUSE_RATIO = 3;
const PREVIEW_MAX_PAUSE = 250;
this.MAX_QUOTED_PASSAGES = MAX_QUOTED_PASSAGES;
this.PRELOADED_MATCH_PREVIEWS = PRELOADED_MATCH_PREVIEWS;
let _sentenceSegmenter = null;
//
// Errors
//
/**
* Thrown when scoring is abandoned via the shouldCancel callback -- e.g.
* because a newer query superseded the one being scored
*/
this.ScoringCancelledError = class extends Error {
constructor(message = 'Scoring cancelled') {
super(message);
this.name = 'BestMatchScoringCancelledError';
}
};
// The semantic engine joins the ranking whenever a model is enabled; the
// lexical engine always ranks
function _useSemantic() {
return Zotero.Embeddings.isEnabled();
}
function _hasPreviews(itemID) {
return !!Zotero.Items.get(itemID)?.isFileAttachment?.();
}
// Resolves the next time the main thread is idle, or after
// PREVIEW_IDLE_TIMEOUT regardless
function _idle() {
return new Promise(resolve => requestIdleCallback(
resolve, { timeout: PREVIEW_IDLE_TIMEOUT }));
}
/**
* Whether a query has anything for best-match search to rank by. The
* lexical engine needs at least one scoring unit; failing that, the
* semantic engine can embed any text its normalization leaves standing,
* when a model is enabled. Callers treat a query that fails this as no
* active search.
*
* @param {String} queryText
* @return {Boolean}
*/
this.isSearchableQuery = function (queryText) {
if (Zotero.Lexical.parseQuery(queryText || '').length) {
return true;
}
return _useSemantic() && !!Zotero.Embeddings.normalizeQuery(queryText || '');
};
/**
* Embedding-index coverage, for banners explaining incomplete best-match
* results. Null when the semantic engine is disabled or everything
* eligible is indexed. Never throws -- the state is informational and
* shouldn't break a search.
*
* @return {Promise<Object|null>} - { type: 'indexing'|'paused', indexed,
* total }, counted in items, attachments included
*/
this.getIndexState = async function () {
try {
let status = Zotero.Embeddings.Indexing.getStatus();
if (!status.enabled) {
return null;
}
// Counts aren't populated until the indexer runs in this session
if (!status.items.total && !status.attachments.total) {
status = await Zotero.Embeddings.Indexing.refreshStatus();
}
// Coverage is coverage: attachment fulltext is reported separately
// in the preferences, but an incomplete index is incomplete
// whichever part of it is still filling in
let indexed = status.items.done + status.attachments.done;
let total = status.items.total + status.attachments.total;
if (indexed >= total) {
return null;
}
// Only an explicit pause reports as paused. Anything else --
// between runs (startup, the pre-run debounce) or after an error
// (detailed in the preferences) -- reports as indexing, since the
// state explains the incomplete coverage, not the indexer.
return {
type: status.paused ? 'paused' : 'indexing',
indexed,
total
};
}
catch (e) {
Zotero.logError(e);
return null;
}
};
/**
* Score a given set of items by relevance to a query. Items that aren't
* matches by any active engine's standards aren't returned. Scores are
* 0-1 against what a perfect answer to the query could show, so one
* number both orders the items and sizes their relevance bars -- an item
* never displays as more relevant than one ranked above it.
*
* With a semantic model enabled, both engines score concurrently and
* their results fuse (see _fuse()). A semantic index that isn't ready --
* mid-build or mid-model-switch -- drops the semantic engine from the
* query instead of failing it, leaving the lexical scores alone.
*
* @param {String} queryText
* @param {Number[]} itemIDs - Candidate item IDs to score
* @param {Object} [options]
* @param {Function} [options.shouldCancel] - Checked between scoring
* stages; return true to abandon scoring with a ScoringCancelledError
* @return {Promise<Object>} - { scores, matches }: scores maps
* itemID -> score (0-1, higher is more relevant); matches says which
* items each engine can show match excerpts in, as { lexical,
* semantic } Sets of itemIDs (see getMatchingExcerpts()). Every
* lexical match has excerpts to show; a semantic match does only
* when a previewable chunk carries it (see
* Zotero.Embeddings.scoreItemIDs()). An engine that didn't rank
* contributes an empty Set.
* @throws {Zotero.BestMatch.ScoringCancelledError}
*/
this.scoreItemIDs = async function (queryText, itemIDs, options = {}) {
try {
// Temporary, for testing: the bestMatchEngine pref pins scoring to
// one engine instead of the hybrid default
let engine = Zotero.Prefs.get('search.bestMatchEngine');
if (engine == 'semantic') {
let semantic = await Zotero.Embeddings.scoreItemIDs(queryText, itemIDs, options);
let kept = _nearTop(semantic.scores, _semanticFraction);
// On the model's display band, so scores are 0-1 like the other
// modes' -- but unclamped and rescaled by the strongest score,
// so a top tier past the band's ceiling keeps its ordering
// instead of flattening into a tie
let fractions = new Map([...kept].map(
([itemID, score]) => [itemID, _semanticFraction(score)]
));
let scale = Math.max(1, ...fractions.values());
return {
scores: new Map([...fractions].map(
([itemID, fraction]) => [itemID, fraction / scale]
)),
matches: {
lexical: new Set(),
semantic: new Set([...semantic.previewableIDs].filter(id => kept.has(id)))
}
};
}
// A query the semantic engine can't embed ranks lexically alone
if (engine == 'lexical' || !_useSemantic()
|| !Zotero.Embeddings.normalizeQuery(queryText || '')) {
let scores = _nearTop(
await Zotero.Lexical.scoreItemIDs(queryText, itemIDs, options), share => share
);
return {
scores,
matches: { lexical: new Set(scores.keys()), semantic: new Set() }
};
}
// Both engines score the same candidates concurrently. allSettled
// rather than all, so one engine's failure still leaves the
// other's rejection observed rather than unhandled
let [lexical, semantic] = await Promise.allSettled([
Zotero.Lexical.scoreItemIDs(queryText, itemIDs, options),
Zotero.Embeddings.scoreItemIDs(queryText, itemIDs, options)
]);
if (lexical.status == 'rejected') {
throw lexical.reason;
}
if (semantic.status == 'rejected') {
if (semantic.reason instanceof Zotero.Embeddings.IndexNotReadyError) {
Zotero.debug("Semantic index not ready -- ranking lexically: "
+ semantic.reason.message);
let scores = _nearTop(lexical.value, share => share);
return {
scores,
matches: {
lexical: new Set(scores.keys()),
semantic: new Set()
}
};
}
throw semantic.reason;
}
// Each engine's tail is cut against its own strongest match, so
// an item enters fusion only with the evidence that stood up
let lexicalScores = _nearTop(lexical.value, share => share);
let semanticScores = _nearTop(semantic.value.scores, _semanticFraction);
return {
scores: _fuse(lexicalScores, semanticScores),
matches: {
lexical: new Set(lexicalScores.keys()),
semantic: new Set(
[...semantic.value.previewableIDs].filter(id => semanticScores.has(id))
)
}
};
}
catch (e) {
if (e instanceof Zotero.Embeddings.ScoringCancelledError
|| e instanceof Zotero.Lexical.ScoringCancelledError) {
throw new this.ScoringCancelledError(e.message);
}
throw e;
}
};
/**
* A best-match search session: one query's scoring pass plus the
* previews explaining its matches.
*
* score() ranks candidates and derives the first few previews on screen
* (PRELOADED_MATCH_PREVIEWS) before it resolves. The rest derive in the
* background, in the same order and paced to stay out of the user's way
* (see PREVIEW_PAUSE_RATIO), reported through onPreviewsFilled and
* awaitable through previewsSettled. An item's entries hold both
* engines' evidence, merged, deduplicated and ordered by strength (see
* getMatchingExcerpts()). A re-score keeps previews already derived;
* fill() re-derives ones invalidate() dropped back to pending. A
* disposed session derives nothing.
*/
this.Session = class {
constructor(queryText) {
this._queryText = queryText;
this._previews = new Map();
this._effectiveScores = new Map();
this._disposed = false;
// Bumped per background pass, so a re-score abandons the last one
this._derivation = 0;
this._previewsSettled = Promise.resolve();
// Set by a consumer showing previews as they arrive: called with
// the itemIDs filled since the last call (see
// PREVIEW_BATCH_INTERVAL), never for what score() derived itself
this.onPreviewsFilled = null;
}
get queryText() {
return this._queryText;
}
/**
* Score candidates for this session's query (see
* Zotero.BestMatch.scoreItemIDs()), rebuild the preview set from the
* engines' match sets, recompute ranks and barFractions, and derive
* the first pending previews on screen before resolving. Items still
* matched keep their settled previews -- a re-score doesn't re-derive
* kept text -- and items no longer matched lose theirs.
*
* The previews past PRELOADED_MATCH_PREVIEWS derive after this
* resolves (see onPreviewsFilled, previewsSettled). Ranks and
* barFractions are complete either way -- neither depends on a
* preview.
*
* @param {Number[]} itemIDs - Candidate item IDs to score
* @param {Object} [options] - Passed through to scoreItemIDs()
* @param {Number} [options.topK] - Keep only the K best-scored items,
* with a deterministic tiebreak, so equal scores keep a stable
* membership; previews are only built and derived for the kept
* items
* @param {Function} [options.shouldCancel] - Also checked between
* preview derivations
* @return {Promise<Map>} - itemID -> score, as scoreItemIDs() returns
* @throws {Zotero.BestMatch.ScoringCancelledError}
*/
async score(itemIDs, options = {}) {
let { scores, matches } = await Zotero.BestMatch.scoreItemIDs(
this._queryText, itemIDs, options);
if (this._disposed) {
return scores;
}
if (options.topK) {
scores = new Map(
[...scores.entries()]
.sort((a, b) => (b[1] - a[1]) || (a[0] - b[0]))
.slice(0, options.topK)
);
}
let previews = new Map();
for (let itemID of new Set([...matches.lexical, ...matches.semantic])) {
if (!scores.has(itemID) || !_hasPreviews(itemID)) {
continue;
}
let existing = this._previews.get(itemID);
if (existing && existing.state != 'pending') {
previews.set(itemID, existing);
continue;
}
previews.set(itemID, {
state: 'pending',
entries: [],
lexical: matches.lexical.has(itemID),
semantic: matches.semantic.has(itemID)
});
}
this._previews = previews;
this._rank(scores);
// Derive in the order rows appear on screen: an item's preview
// rows render under its top-level ancestor, which is ranked by
// the best match anywhere beneath it -- an item's own score can
// sit far below its position
let pending = [...scores.keys()]
.filter(id => previews.get(id)?.state == 'pending')
.map(id => [id, this._topLevelScore(id), this._effectiveScores.get(id) ?? 0])
.sort((a, b) => (b[1] - a[1]) || (b[2] - a[2]) || (a[0] - b[0]))
.map(([id]) => id);
for (let itemID of pending.slice(0, PRELOADED_MATCH_PREVIEWS)) {
if (this._disposed) {
return scores;
}
if (options.shouldCancel?.()) {
throw new Zotero.BestMatch.ScoringCancelledError();
}
await this._derive(itemID);
}
this._previewsSettled = this._deriveRest(
pending.slice(PRELOADED_MATCH_PREVIEWS));
return scores;
}
/**
* Resolves once the previews score() didn't wait for have derived, or
* once the session is disposed
*
* @return {Promise}
*/
get previewsSettled() {
return this._previewsSettled;
}
// Derive the previews score() left behind, in screen order (see
// score()), reporting them in batches (see onPreviewsFilled). Left
// unawaited by score(), so a consumer draws the ranking while the
// explanations behind it fill in. A newer pass -- another score() --
// or dispose() abandons this one.
async _deriveRest(itemIDs) {
let derivation = ++this._derivation;
let filled = [];
let reportedAt = Date.now();
let report = () => {
if (!filled.length) {
return;
}
let reported = filled;
filled = [];
reportedAt = Date.now();
// A consumer that throws shouldn't strand the rest
try {
this.onPreviewsFilled?.(reported);
}
catch (e) {
Zotero.logError(e);
}
};
for (let itemID of itemIDs) {
await _idle();
if (this._disposed || derivation !== this._derivation) {
return;
}
let started = Date.now();
await this._derive(itemID);
await Zotero.Promise.delay(Math.min(PREVIEW_MAX_PAUSE,
(Date.now() - started) * PREVIEW_PAUSE_RATIO));
// A preview that derived nothing has no rows to redraw
if (this._previews.get(itemID)?.state == 'filled') {
filled.push(itemID);
}
if (Date.now() - reportedAt >= PREVIEW_BATCH_INTERVAL) {
report();
}
}
report();
}
/**
* Ranks from this session's last scoring pass: 1-based, tied
* effective scores share a rank, and every row with a match anywhere
* beneath it is covered (see _rank()). Empty before the first pass.
*
* @return {Map} - treeViewID -> rank (1 = most relevant)
*/
get ranks() {
return this._ranks ?? new Map();
}
/**
* Score fractions for the relevance bars, keyed like ranks: each
* row's own score alone, so a row that only inherited its rank from
* a descendant shows its rank over an empty bar
*
* @return {Map} - treeViewID -> 0-1 fraction
*/
get barFractions() {
return this._barFractions ?? new Map();
}
// Rank the scored items, lifting each item's score onto its ancestors
// (annotation -> attachment -> top-level item) first, so an item's
// effective score -- and so its rank -- is the best match anywhere
// beneath it. Equal effective scores get equal ranks, so tied rows
// (including a child and its parent) order deterministically via the
// consumer's secondary sort fields.
_rank(scores) {
let effectiveScores = new Map(scores);
for (let [itemID, score] of scores) {
let parentItemID = Zotero.Items.get(itemID)?.parentItemID;
while (parentItemID) {
let current = effectiveScores.get(parentItemID);
if (current === undefined || score > current) {
effectiveScores.set(parentItemID, score);
}
parentItemID = Zotero.Items.get(parentItemID)?.parentItemID;
}
}
let rankOfScore = new Map(
[...new Set(effectiveScores.values())].sort((a, b) => b - a)
.map((score, i) => [score, i + 1])
);
let ranks = new Map();
let fractions = new Map();
for (let [itemID, score] of effectiveScores) {
let item = Zotero.Items.get(itemID);
if (!item) {
continue;
}
ranks.set(item.treeViewID, rankOfScore.get(score));
fractions.set(item.treeViewID, scores.get(itemID) || 0);
}
this._ranks = ranks;
this._barFractions = fractions;
this._effectiveScores = effectiveScores;
}
// The effective score of an item's top-level ancestor -- what places
// the row subtree the item's preview rows render in
_topLevelScore(itemID) {
let id = itemID;
let parentItemID;
while ((parentItemID = Zotero.Items.get(id)?.parentItemID)) {
id = parentItemID;
}
return this._effectiveScores.get(id) ?? 0;
}
/**
* The preview to show for an item, or null when there's nothing to
* show: no preview for it (see _hasPreviews()), or one that derived
* nothing after all. Passed to consumers as a bare function, so it's
* bound to its session.
*
* @param {Number} itemID
* @return {Object|null} - { state, entries }: state is 'pending'
* (not yet derived) or 'filled'; entries are the derived entries
* (see getMatchingExcerpts()), each with a `key` unique within
* the preview and stable for as long as the preview stays filled
*/
getPreviews = (itemID) => {
let preview = this._previews.get(itemID);
return preview && preview.state != 'empty' ? preview : null;
};
/**
* Derive the given items' previews, in order, for previews put back
* to pending after scoring -- see invalidate(). Items already settled
* are skipped, so filling again is free.
*
* @param {Number[]} itemIDs
*/
async fill(itemIDs) {
for (let itemID of itemIDs) {
if (this._disposed) {
return;
}
await this._derive(itemID);
}
}
/**
* Drop the given items' previews back to pending, for items whose
* content changed and made derived text stale
*
* @param {Number[]} itemIDs
*/
invalidate(itemIDs) {
for (let itemID of itemIDs) {
let preview = this._previews.get(itemID);
if (!preview) {
continue;
}
// A fresh object, so a derivation of the old one that's still
// in flight can't settle it (see _derive())
this._previews.set(itemID, {
...preview,
state: 'pending',
entries: []
});
}
}
/**
* End the session: abandon in-flight derivation. A disposed session
* derives nothing.
*/
dispose() {
this._disposed = true;
}
// Derive one pending item's preview and settle it with the result --
// its entries, all at once, each keyed for row identity. An item
// already settled is left alone, and a preview replaced while
// deriving (see invalidate()) is left to its next derivation. A
// derivation that failed would fail again, so it settles for showing
// nothing rather than being retried.
async _derive(itemID) {
let preview = this._previews.get(itemID);
if (!preview || preview.state != 'pending') {
return;
}
try {
let entries = await this.getMatchingExcerpts(itemID);
if (this._disposed || this._previews.get(itemID) != preview) {
return;
}
preview.entries = entries.map((entry, i) => ({ key: i, ...entry }));
preview.state = entries.length ? 'filled' : 'empty';
}
catch (e) {
Zotero.logError(e);
preview.state = 'empty';
}
}
/**
* Every passage explaining why an item matched this session's query,
* strongest first: the chunks the model matched, or else the item's
* text cut the same way. Only the engines that matched the item are
* asked. The strongest MAX_QUOTED_PASSAGES also carry a `snippet`, the
* one line that best shows the query.
*
* @param {Number} itemID
* @return {Promise<Object[]>} - Entries with `text`, `ranges` and
* `strength`, plus location fields where the passage knows them,
* strongest first; the first MAX_QUOTED_PASSAGES also have
* `snippet`
*/
async getMatchingExcerpts(itemID) {
let queryText = this._queryText;
let passages = await this._getPassages(itemID);
if (!passages.length) {
return [];
}
let texts = passages.map(passage => passage.text);
// Locating the query's words in texts already in hand is cheap,
// unlike scanning a document, so it isn't gated on the item
// having matched lexically -- only on the lexical engine being
// one this session listens to at all
let [ranges, lexical] = await Promise.all([
this._lexicalEnabled()
? Zotero.Lexical.findMatchRanges(queryText, texts)
: texts.map(() => []),
this._lexicalApplies(itemID) ? Zotero.Lexical.scoreTexts(queryText, texts) : null
]);
let entries = [];
for (let i = 0; i < passages.length; i++) {
let passage = passages[i];
let share = lexical ? lexical[i] : 0;
// A passage the model never weighed has only its words to
// recommend it, so one that says nothing of the query isn't a
// match at all
if (passage.score === undefined && !share) {
continue;
}
// Over the engines that spoke for this item, so a strength is
// the same 0-1 fraction whether one weighed the passage or
// both did
let weighed = passage.score !== undefined;
let semanticWeight = weighed ? SEMANTIC_WEIGHT : 0;
let lexicalWeight = lexical ? LEXICAL_WEIGHT : 0;
let fraction = weighed ? Zotero.Embeddings.getScoreFraction(passage.score) : 0;
entries.push({
...passage,
ranges: ranges[i],
strength: (semanticWeight * fraction + lexicalWeight * share)
/ (semanticWeight + lexicalWeight)
});
}
entries.sort((a, b) => b.strength - a.strength);
// Only the strongest few passages are shown as rows in the tree,
// so only they get a line chosen to quote
await this._pickSnippets(entries.slice(0, MAX_QUOTED_PASSAGES));
return entries;
}
/**
* The passages of an item to weigh against the query: the chunks the
* model matched, or else the item's text cut into chunks -- along its
* structure where it's been extracted, which knows where each passage
* sits, and flat otherwise.
*
* @param {Number} itemID
* @return {Promise<Object[]>} - Passages, each with `text`, location
* fields where known, and a `score` from the model
*/
async _getPassages(itemID) {
let queryText = this._queryText;
if (this._semanticApplies(itemID)) {
try {
let chunks = await Zotero.Embeddings.getMatchingChunks(
queryText, itemID, { limit: Infinity });
// Only fulltext chunks carry their own text; item-level
// matches have nothing to excerpt
chunks = chunks.filter(chunk => chunk.text);
if (chunks.length) {
return chunks;
}
}
catch (e) {
if (!(e instanceof Zotero.Embeddings.IndexNotReadyError)) {
throw e;
}
}
}
if (!this._lexicalApplies(itemID)) {
return [];
}
return this._cutPassages(itemID);
}
// Cut an item's text into passages the size the index's are: along its
// outline where it has structured text, so each passage knows its
// section, page and position, and along its flat text where it doesn't
async _cutPassages(itemID) {
let item = await Zotero.Items.getAsync(itemID);
if (!item) {
return [];
}
// Only structure already extracted: generating it costs seconds,
// which is not a price a preview may charge. Cut and placed in the
// document worker, ahead of what's queued there.
let cut = await Zotero.SDT.getItemChunks(itemID,
{ cachedOnly: true, isPriority: true, positions: true });
if (cut.ok && cut.chunks.length) {
return cut.chunks;
}
let text = await item.attachmentText;
if (!text) {
return [];
}
return Zotero.SDT.getPlainTextChunks(text);
}
/**
* Choose where in each passage to quote from: where it says the query's
* words, the sentence or line the lexical engine picks; where it only
* means the query, the sentence the static model finds closest;
* otherwise its opening. The index's own model isn't asked, to keep
* model calls off the pass that derives every preview.
*
* @param {Object[]} entries - Set in place
*/
async _pickSnippets(entries) {
let meant = [];
for (let entry of entries) {
let sentences = Zotero.BestMatch.splitSentences(entry.text);
let pick = entry.ranges.length
? await Zotero.Lexical.pickQuote(this._queryText, entry.text, sentences,
{ longSentence: LONG_SENTENCE_CHARS, width: SNIPPET_CHARS })
: null;
// Says the query's words: quoted where the lexical engine picks
if (pick) {
entry.snippet = pick.line || _quoteFrom(sentences, pick.sentence.start, entry.text);
}
// Means it without saying it: quoted where the static model picks
else if (entry.score !== undefined) {
meant.push({ entry, sentences });
}
else {
entry.snippet = _quoteFrom(sentences, 0, entry.text);
}
}
await this._pickMeantSentences(meant);
}
// Quote each passage from its sentence closest to the query, as the
// static model weighs them in one call, a short sentence weighed with
// the next. Quoted from the opening when there's nothing to choose
// between, or the model isn't downloaded or fails.
async _pickMeantSentences(passages) {
for (let { entry, sentences } of passages) {
entry.snippet = _quoteFrom(sentences, 0, entry.text);
}
passages = passages
.map(passage => ({ ...passage, units: _sentenceUnits(passage.entry.text, passage.sentences) }))
.filter(passage => passage.units.length > 1);
if (!passages.length || !await Zotero.Embeddings.Static.isDownloaded()) {
return;
}
let scores;
try {
scores = await Zotero.Embeddings.Static.similarities(this._queryText,
passages.flatMap(passage => passage.units.map(unit => unit.text)));
}
catch (e) {
Zotero.logError(e);
return;
}
let offset = 0;
for (let { entry, sentences, units } of passages) {
let best = null;
for (let unit of units) {
let score = scores[offset++];
if (!best || score > best.score) {
best = { unit, score };
}
}
entry.snippet = _quoteFrom(sentences, best.unit.start, entry.text);
}
}
// Whether this session's query reaches each engine for the item being
// derived: the engine has to be one the session listens to, and to
// have found something in the item worth speaking about.
_semanticApplies(itemID) {
return this._previews.get(itemID)?.semantic !== false
&& this._modelApplies();
}
_lexicalApplies(itemID) {
return this._previews.get(itemID)?.lexical !== false
&& this._lexicalEnabled();
}
// Whether each engine reaches this session at all, apart from what it
// made of any one item. The bestMatchEngine pref is temporary, for
// testing: it pins a session to a single engine, which then decides
// not only what matched but how a match is quoted.
_modelApplies() {
return Zotero.Prefs.get('search.bestMatchEngine') != 'lexical'
&& _useSemantic()
&& !!Zotero.Embeddings.normalizeQuery(this._queryText || '');
}
_lexicalEnabled() {
return Zotero.Prefs.get('search.bestMatchEngine') != 'semantic';
}
};
/**
* Start a search session for a query (see Zotero.BestMatch.Session)
*
* @param {String} queryText
* @return {Zotero.BestMatch.Session}
*/
this.createSession = function (queryText) {
return new this.Session(queryText);
};
/**
* The sentences of a text, whole and trimmed, each as its extent
* (`start`, `end`) -- the units a snippet is quoted in
*
* @param {String} text
* @return {Object[]}
*/
this.splitSentences = function (text) {
if (!_sentenceSegmenter) {
// Sentence rules don't vary by locale, but the default locale
// varies by machine, so pin one
_sentenceSegmenter = new Intl.Segmenter('en', { granularity: 'sentence' });
}
let sentences = [];
for (let { segment, index } of _sentenceSegmenter.segment(text)) {
let trimmed = segment.trim();
if (!trimmed) {
continue;
}
let start = index + (segment.length - segment.trimStart().length);
sentences.push({ start, end: start + trimmed.length });
}
return sentences;
};
// The extent to quote from the sentence starting at `from`: sentences are
// added until the quote passes SNIPPET_CHARS, the one crossing it taken
// whole, since half a sentence reads as a truncation and the row clips
// what doesn't fit. A text with no sentences is quoted from its start.
function _quoteFrom(sentences, from, text) {
sentences = sentences.filter(sentence => sentence.start >= from);
if (!sentences.length) {
return { start: 0, end: Math.min(text.length, SNIPPET_CHARS) };
}
let { start, end } = sentences[0];
for (let i = 1; i < sentences.length && end - start < SNIPPET_CHARS; i++) {
end = sentences[i].end;
}
return { start, end, startsSentence: true };
}
// A passage's sentences as the units its closest sentence is chosen from:
// a sentence shorter than MIN_UNIT_CHARS is joined to the next, and a
// short last one to the unit before it
//
// @return {Object[]} - [{ start, end, text }]
function _sentenceUnits(text, sentences) {
let units = [];
let current = null;
for (let sentence of sentences) {
if (current) {
current.end = sentence.end;
}
else {
current = { start: sentence.start, end: sentence.end };
}
if (current.end - current.start >= MIN_UNIT_CHARS) {
units.push(current);
current = null;
}
}
if (current) {
if (units.length) {
units[units.length - 1].end = current.end;
}
else {
units.push(current);
}
}
return units.map(unit => ({ ...unit, text: text.slice(unit.start, unit.end) }));
}
// The semantic engine's raw score as an unclamped fraction of the display
// band (see Zotero.Embeddings.getScoreFraction()): the strength fusion
// and the margin read
function _semanticFraction(score) {
return Zotero.Embeddings.getScoreFraction(score, { clamped: false });
}
// An engine's results that stand within the margin of its strongest: the
// items whose fraction is at least (1 - margin) of the top fraction.
// Every engine returns a tail of items barely above its floor -- for a
// query with a few strong answers, hundreds of them -- that aren't
// matches for that query in any sense a reader would accept. Measuring
// the cut from the top rather than by count lets a broad query keep
// hundreds of comparable results while a specific one keeps a handful.
// A lone result is never cut. The margin is a pref, in percent.
function _nearTop(scores, toFraction) {
let margin = Zotero.Prefs.get('search.bestMatchMargin') / 100;
if (scores.size < 2 || !(margin < 1)) {
return scores;
}
let fractions = new Map([...scores].map(([itemID, score]) => [itemID, toFraction(score)]));
let cutoff = Math.max(...fractions.values()) * (1 - margin);
return new Map([...scores].filter(([itemID]) => fractions.get(itemID) >= cutoff));
}
// Fuse the two engines' scores with strength-weighted Reciprocal Rank
// Fusion: an item's fused score sums fraction / (RRF_K + rank) over the
// engines that matched it, where fraction is that engine's own 0-1
// measure of the evidence -- the lexical score, or the semantic score on
// the model's display band. Rank rewards agreement between the engines
// without calibrating their scales against each other; the fraction
// keeps the reward proportionate to what each engine actually found.
// Pure reciprocal ranks would be blind to that magnitude in both
// directions: a pair of barely-above-floor matches would buy the full
// agreement bonus, and a strong match only one engine can see -- a
// paraphrase without the query's words, say -- would cap at half however
// good it is.
//
// The semantic fractions enter unclamped: a strong query's whole top
// tier can sit past the display band's ceiling, where clamping would
// flatten the model's ordering into a tie -- and a tie decided by which
// items the lexical engine also happened to match, rather than by what
// the model actually read.
//
// Fused scores are normalized against the best possible sum
// (full-strength evidence at rank 1 in both engines) -- or against the
// strongest sum where unclamped fractions push past it -- to keep them
// 0-1. Within an engine, tied scores share a rank, so fusion is
// deterministic however a map orders its entries.
function _fuse(lexicalScores, semanticScores) {
let engines = [
[lexicalScores, score => Math.min(1, Math.max(0, score))],
[semanticScores, _semanticFraction]
];
let scores = new Map();
for (let [engineScores, toFraction] of engines) {
let rankOfScore = new Map(
[...new Set(engineScores.values())]
.sort((a, b) => b - a)
.map((score, i) => [score, i + 1])
);
for (let [itemID, score] of engineScores) {
let sum = scores.get(itemID) || 0;
scores.set(itemID,
sum + toFraction(score) / (RRF_K + rankOfScore.get(score)));
}
}
let ceiling = Math.max(2 / (RRF_K + 1), ...scores.values());
for (let [itemID, sum] of scores) {
scores.set(itemID, sum / ceiling);
}
return scores;
}
};

View file

@ -718,7 +718,7 @@ Zotero.CollectionTreeRow.prototype.getBestMatchSource = function () {
}
// An active best-match quick search overrides a selected saved search's
// own marker, so the user's typed query wins
if (this.searchText && Zotero.Embeddings.normalizeQuery(this.searchText)
if (this.searchText && Zotero.BestMatch.isSearchableQuery(this.searchText)
&& (this.searchMode || Zotero.Prefs.get('search.quicksearch-mode')) === 'bestMatch') {
return false;
}
@ -742,7 +742,7 @@ Zotero.CollectionTreeRow.prototype.isBestMatchSearch = function () {
if (this.getBestMatchSource()) {
return true;
}
return !!(this.searchText && Zotero.Embeddings.normalizeQuery(this.searchText))
return !!(this.searchText && Zotero.BestMatch.isSearchableQuery(this.searchText))
&& (this.searchMode || Zotero.Prefs.get('search.quicksearch-mode')) === 'bestMatch';
};

View file

@ -1789,7 +1789,13 @@ Zotero.Item.prototype._saveData = async function (env) {
await Zotero.DB.queryAsync(sql, [itemID].concat(del));
}
}
// Keep the item-text search index current with the item's title and abstract. Indexed
// inline -- the text is tiny -- so the item is searchable the moment the save commits.
if ((isNew || this._changed.itemData) && this.isRegularItem() && !this.isFeedItem) {
await Zotero.FullText.indexItemText(this);
}
//
// Creators
//
@ -2297,6 +2303,9 @@ Zotero.Item.prototype._saveData = async function (env) {
// Clear cached child items of the parent attachment
reloadParentChildItems[parentItemID] = true;
// Keep the item-text search index current with the annotation's passage and comment
await Zotero.FullText.indexItemText(this);
// Reload display title of annotations
if (this.isAnnotation()) {
this.updateDisplayTitle();
@ -5628,10 +5637,15 @@ Zotero.Item.prototype._eraseData = async function (env) {
await Zotero.Fulltext.clearItemWords(this.id);
//Zotero.Fulltext.clearItemContent(this.id);
}
// Clear the note content index (notes and attachments can both have itemNotes rows)
// Clear the note content index (notes and attachments can both have itemNotes rows);
// this also clears the note's item-text index entries
if (this.isNote() || this.isAttachment()) {
await Zotero.FullText.clearNoteIndex(this.id);
}
// Clear the item-text index for everything else (regular items, annotations)
else {
await Zotero.FullText.clearItemTextIndex(this.id);
}
await Zotero.DB.queryAsync('DELETE FROM items WHERE itemID=?', this.id);

View file

@ -831,29 +831,18 @@ Zotero.Search.prototype.search = async function (asTempTable) {
//Zotero.debug(ids);
// A root-level 'bestMatch' condition with a top-K cutoff makes membership
// semantic: only the K results most similar to the query match, so the
// saved search returns the same set when used as a source (scopes, counts,
// the API). Without a cutoff, best match is only a ranking in the items
// list and membership is untouched. If the index isn't usable (no model,
// or mid-switch), a cutoff search matches nothing rather than an arbitrary
// set.
// relevance-based: only the K results most relevant to the query match,
// so the saved search returns the same set when used as a source (scopes,
// counts, the API). Without a cutoff, best match is only a ranking in the
// items list and membership is untouched.
let bestMatch = this.getBestMatchQuery();
if (ids && ids.length && bestMatch && bestMatch.topK) {
try {
let scores = await Zotero.Embeddings.scoreItemIDs(bestMatch.query, ids);
ids = [...scores.entries()]
// Deterministic order: by score, then by itemID for equal scores
.sort((a, b) => (b[1] - a[1]) || (a[0] - b[0]))
.slice(0, bestMatch.topK)
.map(([itemID]) => itemID);
}
catch (e) {
if (!(e instanceof Zotero.Embeddings.IndexNotReadyError)) {
throw e;
}
Zotero.debug("Embeddings index not ready -- best-match cutoff search matches nothing");
ids = [];
}
let { scores } = await Zotero.BestMatch.scoreItemIDs(bestMatch.query, ids);
ids = [...scores.entries()]
// Deterministic order: by score, then by itemID for equal scores
.sort((a, b) => (b[1] - a[1]) || (a[0] - b[0]))
.slice(0, bestMatch.topK)
.map(([itemID]) => itemID);
}
if (!ids || !ids.length) {
@ -883,7 +872,7 @@ Zotero.Search.prototype.getBestMatchQuery = function () {
depth--;
}
else if (depth == 0 && condition.condition == 'bestMatch' && condition.value
&& Zotero.Embeddings.normalizeQuery(condition.value)) {
&& Zotero.BestMatch.isSearchableQuery(condition.value)) {
return {
query: condition.value,
topK: parseInt(condition.operator) || false

File diff suppressed because it is too large Load diff

View file

@ -44,7 +44,7 @@ Zotero.Fulltext = Zotero.FullText = new function () {
// content and note tables. The tables are only created when this is bumped (setUpContentDB()
// drops and recreates everything), so any schema change -- a new table or a changed FTS5
// table definition, including its tokenizer -- needs a bump.
const _indexDBVersion = 2;
const _indexDBVersion = 3;
// Version of the index format. Bump to force a rebuild of the index from the cached text
// (e.g., after a normalization change) -- items recorded at a lower version in
// fulltextIndexState/fulltextNoteIndexState reenter the queue and are re-indexed at startup,
@ -97,6 +97,7 @@ Zotero.Fulltext = Zotero.FullText = new function () {
var _attachmentIndexQueueCursor = 0;
var _attachmentExtractionQueueCursor = 0;
var _noteIndexQueueCursor = 0;
var _itemTextIndexQueueCursor = 0;
// Flush a content batch to the index once it reaches this many characters, so peak memory stays
// bounded by the budget rather than by item size (each up to fulltext.textMaxLength)
const _maxContentBatchChars = 2000000;
@ -150,6 +151,9 @@ Zotero.Fulltext = Zotero.FullText = new function () {
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextIndexState");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextNoteIndexState");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.noteText");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextItemText");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextItemTextCJK");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextItemTextState");
await Zotero.DB.queryAsync("DROP TABLE IF EXISTS ftindex.fulltextIndexMeta");
// Non-CJK scripts: word index over the normalized text. Searches match words by
// prefix (see getWordMatchClause); multi-token terms match adjacent tokens as
@ -216,6 +220,37 @@ Zotero.Fulltext = Zotero.FullText = new function () {
+ " text TEXT\n"
+ ")"
);
// Word-level index over the items' own searchable text, one row per item
// (rowid = itemID), each row filling only the columns its item type has:
// title and abstract for regular items, note for note items (the plain text
// the note tables index), annotation for annotation items (the marked
// passage plus the comment). Item types never share an itemID, so one table
// holds them all, and column filters ('title: owl') tell the sources apart.
// Unlike the trigram note tables, which answer substring searches, this one
// answers word searches: whole-token and prefix matching, and per-word
// document counts.
await Zotero.DB.queryAsync(
"CREATE VIRTUAL TABLE ftindex.fulltextItemText USING fts5("
+ "title, abstract, note, annotation, "
+ "tokenize='unicode61', content='', contentless_delete=1)"
);
// CJK 2-grams of the same columns (see fulltextContentCJK)
await Zotero.DB.queryAsync(
"CREATE VIRTUAL TABLE ftindex.fulltextItemTextCJK USING fts5("
+ "title, abstract, note, annotation, "
+ "tokenize='ascii', content='', contentless_delete=1)"
);
// Which regular items and annotations are in the item-text index, and at
// what format version, making their backfill queue a countable, resumable
// query. Notes aren't tracked here: their item-text row is written in the
// same step that indexes them into the note tables, so the note index state
// already covers it.
await Zotero.DB.queryAsync(
"CREATE TABLE ftindex.fulltextItemTextState (\n"
+ " itemID INTEGER PRIMARY KEY,\n"
+ " version INT NOT NULL\n"
+ ")"
);
// Index metadata, including the localUserKey the index was built against (above)
await Zotero.DB.queryAsync(
"CREATE TABLE ftindex.fulltextIndexMeta (\n"
@ -1626,6 +1661,146 @@ Zotero.Fulltext = Zotero.FullText = new function () {
};
/**
* Number of regular items and annotations not yet in the item-text index at the current
* format version -- the item-text backfill queue. Ordinary saves index inline (see
* indexItemText()), so this queue only holds pre-existing items after an index rebuild or
* a format-version bump. Feed items aren't indexed: their libraries churn wholesale and
* aren't part of item search.
*
* @return {Promise<Integer>}
*/
this.getItemTextIndexQueueCount = async function () {
return Zotero.DB.valueQueryAsync(
"SELECT COUNT(*) FROM items I "
+ "JOIN libraries L USING (libraryID) "
+ "LEFT JOIN ftindex.fulltextItemTextState S USING (itemID) "
+ "WHERE I.itemTypeID NOT IN (?, ?) "
+ "AND L.type IN ('user', 'group') "
+ "AND (S.itemID IS NULL OR S.version<?)",
[
Zotero.ItemTypes.getID('attachment'),
Zotero.ItemTypes.getID('note'),
_contentIndexVersion
]
);
};
/**
* Process the item-text index queue (see getItemTextIndexQueueCount()), indexing the
* searchable text of regular items and annotations that aren't in the item-text index yet.
*
* @param {Object} [options]
* @param {Integer} [options.maxTime] - Stop after roughly this many ms
* @param {Boolean} [options.checkIdle] - Stop as soon as the user is no longer idle
* @return {Promise<Integer>} The number of items processed
*/
this.processItemTextIndexQueue = async function ({ maxTime = null, checkIdle = false, onProgress = null } = {}) {
if (_indexingInProgress) {
return 0;
}
_indexingInProgress = true;
let start = Date.now();
let processed = 0;
let attachmentTypeID = Zotero.ItemTypes.getID('attachment');
let noteTypeID = Zotero.ItemTypes.getID('note');
let titleFieldIDs = new Set([
Zotero.ItemFields.getID('title'),
...Zotero.ItemFields.getTypeFieldsFromBase('title')
]);
let abstractFieldID = Zotero.ItemFields.getID('abstractNote');
try {
while (true) {
if (Zotero.Sync.Runner.syncInProgress) {
return processed;
}
if (checkIdle && !_canDrainIndex()) {
return processed;
}
let rows = await Zotero.DB.queryAsync(
"SELECT I.itemID FROM items I "
+ "JOIN libraries L USING (libraryID) "
+ "LEFT JOIN ftindex.fulltextItemTextState S USING (itemID) "
+ "WHERE I.itemTypeID NOT IN (?, ?) "
+ "AND L.type IN ('user', 'group') "
+ "AND (S.itemID IS NULL OR S.version<?) AND I.itemID>? "
+ "ORDER BY I.itemID LIMIT 100",
[attachmentTypeID, noteTypeID, _contentIndexVersion, _itemTextIndexQueueCursor]
);
if (!rows.length) {
// Nothing left above the cursor; restart from the beginning to catch anything
// re-queued below it
if (_itemTextIndexQueueCursor) {
_itemTextIndexQueueCursor = 0;
continue;
}
Zotero.debug("No queued items to index for item-text search");
return processed;
}
Zotero.debug("Indexing item text for " + rows.length + " "
+ Zotero.Utilities.pluralize(rows.length, 'item'));
// Read and index each item in one transaction, so a save can't land between
// reading its text and marking it indexed at the current version
await Zotero.DB.executeTransaction(async function () {
let itemIDs = rows.map(row => row.itemID);
let placeholders = itemIDs.join(',');
let fieldRows = await Zotero.DB.queryAsync(
"SELECT itemID, fieldID, value FROM itemData "
+ "JOIN itemDataValues USING (valueID) "
+ "WHERE itemID IN (" + placeholders + ") "
+ "AND fieldID IN ("
+ [...titleFieldIDs, abstractFieldID].join(',') + ")"
);
let fieldsByItem = new Map();
for (let row of fieldRows) {
let fields = fieldsByItem.get(row.itemID) || {};
let column = titleFieldIDs.has(row.fieldID) ? 'title' : 'abstract';
fields[column] = fields[column]
? fields[column] + ' ' + row.value
: row.value;
fieldsByItem.set(row.itemID, fields);
}
let annotationRows = await Zotero.DB.queryAsync(
"SELECT itemID, text, comment FROM itemAnnotations "
+ "WHERE itemID IN (" + placeholders + ")"
);
for (let row of annotationRows) {
fieldsByItem.set(row.itemID, {
annotation: [row.text, row.comment].filter(Boolean).join(' ')
});
}
// Confirm the items still exist inside the transaction, so one deleted
// since it was queued is cleared rather than re-recorded
let liveIDs = new Set(await Zotero.DB.columnQueryAsync(
"SELECT itemID FROM items WHERE itemID IN (" + placeholders + ")"
));
for (let row of rows) {
if (!liveIDs.has(row.itemID)) {
await clearItemTextEntries(row.itemID);
continue;
}
// An item with no text at all still gets a state row, so the
// queue converges
await setItemTextIndex(row.itemID, fieldsByItem.get(row.itemID) || {});
}
});
processed += rows.length;
_itemTextIndexQueueCursor = rows[rows.length - 1].itemID;
if (onProgress) {
onProgress(processed);
}
if (maxTime && (Date.now() - start) >= maxTime) {
return processed;
}
}
}
finally {
_indexingInProgress = false;
}
};
/**
* Process the note index queue, indexing the text content of notes that aren't in the note
* index yet or were edited since their last index update. Note saves only mark the note stale
@ -1756,7 +1931,9 @@ Zotero.Fulltext = Zotero.FullText = new function () {
// search is incomplete. Pause for syncs so it doesn't compete with full-text content
// downloads.
let queueCount = async () => {
return (await this.getAttachmentIndexQueueCount()) + (await this.getNoteIndexQueueCount());
return (await this.getAttachmentIndexQueueCount())
+ (await this.getNoteIndexQueueCount())
+ (await this.getItemTextIndexQueueCount());
};
let total = await queueCount();
// Items drained so far, tracked across processor calls so the bar can advance after each
@ -1827,6 +2004,10 @@ Zotero.Fulltext = Zotero.FullText = new function () {
done = base + await this.processNoteIndexQueue(
{ maxTime: 1000, onProgress: n => updateProgress(base + n) }
);
base = done;
done = base + await this.processItemTextIndexQueue(
{ maxTime: 1000, onProgress: n => updateProgress(base + n) }
);
}
if (itemProgress) {
if (stalled) {
@ -1894,7 +2075,8 @@ Zotero.Fulltext = Zotero.FullText = new function () {
let steps = 0;
let capped = false;
try {
let tables = [_contentTables, _noteTables].flatMap(({ main, cjk }) => [main, cjk]);
let itemTextTables = { main: 'fulltextItemText', cjk: 'fulltextItemTextCJK' };
let tables = [_contentTables, _noteTables, itemTextTables].flatMap(({ main, cjk }) => [main, cjk]);
for (let table of tables) {
// The first merge of a table takes a negative page count, which gathers its
// segments onto one level and disregards usermerge so that all of them merge. The
@ -2023,7 +2205,8 @@ Zotero.Fulltext = Zotero.FullText = new function () {
}
let remaining = (await self.getAttachmentIndexQueueCount())
+ (await self.getAttachmentExtractionQueueCount())
+ (await self.getNoteIndexQueueCount());
+ (await self.getNoteIndexQueueCount())
+ (await self.getItemTextIndexQueueCount());
if (remaining == 0) {
// Drained -- compact the index and stop. A pass that ran out of steps leaves
// the indexed count standing, so keep going until the merge is finished.
@ -2051,6 +2234,7 @@ Zotero.Fulltext = Zotero.FullText = new function () {
await self.processAttachmentIndexQueue({ maxTime: 1000, checkIdle: true });
await self.processAttachmentExtractionQueue({ maxTime: 1000, checkIdle: true });
await self.processNoteIndexQueue({ maxTime: 1000, checkIdle: true });
await self.processItemTextIndexQueue({ maxTime: 1000, checkIdle: true });
// Back off after a run that made no progress, so contention (e.g., another
// caller holding _indexingInProgress) can clear before the next check
_scheduleQueueDrain(_queueDrainStalledRuns ? _drainIdleDelay * 1000 : 250);
@ -2314,11 +2498,120 @@ Zotero.Fulltext = Zotero.FullText = new function () {
async function setNoteIndex(itemID, text) {
await setIndexEntries(itemID, text, _noteTables);
// The note's text is also word-indexed. Its currency is tracked by the note index
// state written above, so no item-text state row of its own.
await setItemTextIndex(itemID, { note: text }, { recordState: false });
// Now indexed at the current version, so it's no longer stale
_staleNoteText.delete(itemID);
}
// The item-text index's columns, in table order (see setUpContentDB)
const _itemTextColumns = ['title', 'abstract', 'note', 'annotation'];
/**
* Index an item's searchable text into the item-text FTS5 tables, replacing any existing
* entries. `fields` maps column names (see _itemTextColumns) to raw text, normalized here;
* missing and empty columns are indexed as nothing. Recording the item in the item-text
* state table is the caller's choice: regular items and annotations are tracked there,
* while a note's currency is tracked by the note index state it shares with the note
* tables.
*
* Must be called within a transaction.
*/
async function setItemTextIndex(itemID, fields, { recordState = true } = {}) {
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemText WHERE rowid=?", itemID);
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemTextCJK WHERE rowid=?", itemID);
let values = [];
let cjkValues = [];
for (let column of _itemTextColumns) {
let normalized = Zotero.Utilities.Internal.normalizeForSearch(fields[column] || '') || '';
values.push(normalized);
cjkValues.push(normalized ? getCJKBigrams(normalized) : '');
}
let columnSQL = _itemTextColumns.join(', ');
let placeholders = _itemTextColumns.map(() => '?').join(', ');
if (values.some(Boolean)) {
await Zotero.DB.queryAsync(
"INSERT INTO ftindex.fulltextItemText (rowid, " + columnSQL + ") "
+ "VALUES (?, " + placeholders + ")",
[itemID, ...values],
{ debugParams: false }
);
}
if (cjkValues.some(Boolean)) {
await Zotero.DB.queryAsync(
"INSERT INTO ftindex.fulltextItemTextCJK (rowid, " + columnSQL + ") "
+ "VALUES (?, " + placeholders + ")",
[itemID, ...cjkValues],
{ debugParams: false }
);
}
if (recordState) {
await Zotero.DB.queryAsync(
"REPLACE INTO ftindex.fulltextItemTextState (itemID, version) VALUES (?, ?)",
[itemID, _contentIndexVersion]
);
}
}
// Remove an item's entries from the item-text tables and its state row (a no-op for
// notes, which have none).
//
// Must be called within a transaction.
async function clearItemTextEntries(itemID) {
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemText WHERE rowid=?", itemID);
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemTextCJK WHERE rowid=?", itemID);
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemTextState WHERE itemID=?", itemID);
}
/**
* Index an item's searchable text into the item-text index: title fields (including
* type-specific ones) and abstract for a regular item, the marked passage and comment for
* an annotation. Called from the item's own save, within its transaction, so the index is
* current the moment the save commits -- these texts are small enough to index inline,
* unlike note content, which is flagged and indexed in the background (see
* flagNoteStale()).
*
* @param {Zotero.Item} item
* @return {Promise}
*/
this.indexItemText = async function (item) {
if (item.isAnnotation()) {
await setItemTextIndex(item.id, {
annotation: [item.annotationText, item.annotationComment]
.filter(Boolean).join(' ')
});
return;
}
await setItemTextIndex(item.id, {
title: item.getField('title', false, true),
abstract: item.getField('abstractNote')
});
};
/**
* Remove an item's entries from the item-text index, e.g., when the item is erased (the
* index database can't FK-cascade to zotero.sqlite).
*
* Must be called within a transaction.
*
* @param {Number} itemID
* @return {Promise}
*/
this.clearItemTextIndex = async function (itemID) {
await clearItemTextEntries(itemID);
};
/**
* Flag a note as edited since its last index update, without re-indexing it. Called on every
* note save, so it must stay cheap: it writes a one-row version-0 marker (instead of rebuilding
@ -2383,6 +2676,8 @@ Zotero.Fulltext = Zotero.FullText = new function () {
*/
this.clearNoteIndex = async function (itemID) {
await clearIndexEntries(itemID, _noteTables);
// The note's word-indexed text lives in the item-text tables (see setNoteIndex())
await clearItemTextEntries(itemID);
_staleNoteText.delete(itemID);
};
@ -2606,6 +2901,94 @@ Zotero.Fulltext = Zotero.FullText = new function () {
};
/**
* IDs of the given notes whose note-index entries can't be relied on to reflect their
* current text: notes edited since their last index update (see flagNoteStale()) and notes
* not in the index at the current format version (e.g., mid-backfill). A caller matching
* against the index should read these notes' current text instead (see
* getNoteSearchTexts()).
*
* @param {Integer[]} itemIDs
* @return {Promise<Integer[]>}
*/
this.getStaleOrUnindexedNoteIDs = async function (itemIDs) {
let result = [];
let chunkSize = 500;
for (let i = 0; i < itemIDs.length; i += chunkSize) {
let chunk = itemIDs.slice(i, i + chunkSize);
result.push(...await Zotero.DB.columnQueryAsync(
"SELECT N.itemID FROM itemNotes N "
+ "LEFT JOIN ftindex.fulltextNoteIndexState S USING (itemID) "
+ "WHERE N.itemID IN (" + chunk.map(() => '?').join(',') + ") "
+ "AND (S.itemID IS NULL OR S.version<?)",
[...chunk, _contentIndexVersion]
));
}
return result;
};
/**
* Normalized searchable plain text of the given notes, keyed by itemID --
* the same text note searching matches against: the note index's stored
* plain text where the index is current, the in-memory text of notes
* edited since their last index update (see flagNoteStale()), and a fresh
* extraction for notes the index doesn't hold yet (e.g., mid-backfill).
* IDs that aren't notes are simply absent from the result.
*
* @param {Integer[]} itemIDs
* @return {Promise<Map>} - itemID -> text
*/
this.getNoteSearchTexts = async function (itemIDs) {
let texts = new Map();
let chunkSize = 500;
for (let i = 0; i < itemIDs.length; i += chunkSize) {
let chunk = itemIDs.slice(i, i + chunkSize);
let placeholders = chunk.map(() => '?').join(',');
let rows = await Zotero.DB.queryAsync(
"SELECT itemID, text FROM ftindex.noteText WHERE itemID IN (" + placeholders + ")",
chunk
);
for (let row of rows) {
texts.set(row.itemID, row.text);
}
// A note edited since its last index update still holds its
// pre-edit text in noteText -- its current text is what counts
let staleIDs = new Set(await Zotero.DB.columnQueryAsync(
"SELECT itemID FROM ftindex.fulltextNoteIndexState "
+ "WHERE version=0 AND itemID IN (" + placeholders + ")",
chunk
));
// Stale notes not seen since their edit (e.g., flagged in a
// previous session) and notes with no index entry at all are both
// extracted from the stored note
let fetchIDs = chunk.filter((id) => {
return staleIDs.has(id) ? !_staleNoteText.has(id) : !texts.has(id);
});
if (fetchIDs.length) {
let noteRows = await Zotero.DB.queryAsync(
"SELECT itemID, note FROM itemNotes WHERE itemID IN ("
+ fetchIDs.map(() => '?').join(',') + ")",
fetchIDs
);
for (let row of noteRows) {
let text = _normalizeNoteText(row.note);
texts.set(row.itemID, text);
if (staleIDs.has(row.itemID)) {
_staleNoteText.set(row.itemID, text);
}
}
}
for (let id of staleIDs) {
if (_staleNoteText.has(id)) {
texts.set(id, _staleNoteText.get(id));
}
}
}
return texts;
};
/**
* Return the ids of attachment items in the given library whose full-text content matches
* `searchText` (see getWordMatchClause for the matching semantics). This is the non-regexp
@ -2981,14 +3364,24 @@ Zotero.Fulltext = Zotero.FullText = new function () {
+ "WHERE S.version >= ?";
var notesIndexed = await Zotero.DB.valueQueryAsync(sql, _contentIndexVersion);
// Regular items and annotations in the item-text index at the current version (their
// titles, abstracts, and annotation text -- notes are counted above)
var sql = "SELECT COUNT(*) FROM ftindex.fulltextItemTextState WHERE version >= ?";
var itemTextIndexed = await Zotero.DB.valueQueryAsync(sql, _contentIndexVersion);
// Pending work, all auto-draining: items with extracted content not yet in the search index
// (remaining), indexable attachments not yet extracted (unindexedQueue), and notes not yet
// indexed or edited since their last index update (noteQueue)
// (remaining), indexable attachments not yet extracted (unindexedQueue), notes not yet
// indexed or edited since their last index update (noteQueue), and items and annotations
// not yet in the item-text index (itemTextQueue)
var remaining = await this.getAttachmentIndexQueueCount();
var unindexedQueue = await this.getAttachmentExtractionQueueCount();
var noteQueue = await this.getNoteIndexQueueCount();
var itemTextQueue = await this.getItemTextIndexQueueCount();
return { indexed, partial, notAvailable, notesIndexed, remaining, unindexedQueue, noteQueue };
return {
indexed, partial, notAvailable, notesIndexed, itemTextIndexed,
remaining, unindexedQueue, noteQueue, itemTextQueue
};
};

File diff suppressed because it is too large Load diff

View file

@ -27,7 +27,7 @@
* Zotero.ML -- access to Firefox's machine-learning runtime
* (toolkit/components/ml), which runs models in a separate, memory-gated
* inference process using the native ONNX Runtime and llama.cpp libraries
* Firefox ships.
* Firefox ships, or for static embeddings, plain JS over a token table.
*
* Engines are created through createEngine(), which supplies the runtime
* configuration Zotero's build requires. Callers pass the model and task.
@ -37,14 +37,17 @@ Zotero.ML = new function () {
// The runtime additionally always allows chrome://, resource://, and
// localhost.
const ALLOWED_MODEL_HOSTS = [
{ filter: 'ALLOW', urlPrefix: 'https://huggingface.co/' }
{ filter: 'ALLOW', urlPrefix: 'https://huggingface.co/' },
// Mozilla's hub, the one host serving static-embeddings tables
{ filter: 'ALLOW', urlPrefix: 'https://model-hub.mozilla.org/' }
];
// The backends that run on the native libraries Firefox ships. The rest
// either load a WebAssembly runtime from a Remote Settings attachment
// The backends that need no runtime this build lacks: the native libraries
// Firefox ships, and 'static-embeddings', plain JS over a token table. The
// rest either load a WebAssembly runtime from a Remote Settings attachment
// ('onnx' and 'wllama', plus the 'best-*' backends that fall back to them)
// or aren't local at all ('openai').
const NATIVE_BACKENDS = ['onnx-native', 'llama.cpp'];
const NATIVE_BACKENDS = ['onnx-native', 'llama.cpp', 'static-embeddings'];
// Default host and URL layout for models, for operations that address the
// model cache without creating an engine
@ -154,15 +157,15 @@ Zotero.ML = new function () {
* @return {Promise<Object[]>} - [{ taskName, name, modelId, revision }]
*/
this.listModels = async function ({ taskName } = {}) {
let host = new URL(MODEL_HUB_ROOT_URL).host + '/';
let hosts = ALLOWED_MODEL_HOSTS.map(({ urlPrefix }) => new URL(urlPrefix).host + '/');
let models = await _getModelHub().listModels();
if (taskName) {
models = models.filter(model => model.taskName === taskName);
}
return models.map(model => ({
...model,
modelId: model.name.startsWith(host) ? model.name.slice(host.length) : model.name
}));
return models.map((model) => {
let host = hosts.find(prefix => model.name.startsWith(prefix));
return { ...model, modelId: host ? model.name.slice(host.length) : model.name };
});
};
/**
@ -179,13 +182,42 @@ Zotero.ML = new function () {
await _getModelHub().deleteModels({ taskName, model, revision, deletedBy: 'zotero' });
};
/**
* Read a single file of a model (e.g. its tokenizer) from the runtime's
* model cache, fetching it from the model hub if it isn't cached yet.
*
* @param {Object} options
* @param {String} options.taskName - Task the model is cached for, as
* passed to createEngine() -- the cache registers every file under it
* @param {String} options.modelId - The model id, as passed to createEngine()
* @param {String} options.file - File path within the model repository
* (e.g. 'tokenizer.json')
* @param {String} [options.engineId]
* @param {String} [options.revision='main']
* @return {Promise<ArrayBuffer>}
*/
this.getModelFile = async function ({ taskName, modelId, file, engineId, revision = 'main' }) {
let [buffer] = await _getModelHub().getModelFileAsArrayBuffer({
engineId,
taskName,
model: modelId,
revision,
file
});
return buffer;
};
function _getModelHub() {
let { ModelHub } = ChromeUtils.importESModule(
"chrome://global/content/ml/ModelHub.sys.mjs"
);
return new ModelHub({
rootUrl: MODEL_HUB_ROOT_URL,
urlTemplate: MODEL_HUB_URL_TEMPLATE
urlTemplate: MODEL_HUB_URL_TEMPLATE,
// A hub constructed without a list denies every external host --
// the engine's own hub gets this same policy from the Remote
// Settings mock (see _configureRuntime())
allowDenyList: ALLOWED_MODEL_HOSTS
});
}

View file

@ -27,7 +27,7 @@
Zotero.Notifier = new function () {
// Options that apply to an entire event, not a specific object
this.EVENT_LEVEL_OPTIONS = ['autoSyncDelay', 'skipAutoSync', 'embeddingsUpdate'];
this.EVENT_LEVEL_OPTIONS = ['autoSyncDelay', 'skipAutoSync'];
var _observers = {};
var _types = [

View file

@ -79,6 +79,8 @@ class PDFWorker {
}
async _query(action, data, transfer, options = {}) {
// The worker is dropped after an error, so there may be none yet
this._init();
return new Promise((resolve, reject) => {
this._lastPromiseID++;
this._waitingPromises[this._lastPromiseID] = {
@ -179,8 +181,22 @@ class PDFWorker {
});
this._worker.addEventListener('error', (event) => {
Zotero.logError(`Document worker error (${event.filename}:${event.lineno}): ${event.message}`);
// An error outside an action's own handling leaves its request
// unanswered, so every pending request fails and the next one
// starts a fresh worker
this._reset(new Error(`Document worker error: ${event.message}`));
});
}
_reset(error) {
let pending = Object.values(this._waitingPromises);
this._waitingPromises = {};
this._worker.terminate();
this._worker = null;
for (let { reject } of pending) {
reject(error);
}
}
canImport(item) {
if (item.isPDFAttachment()) {
@ -694,6 +710,52 @@ class PDFWorker {
}, !!isPriority);
}
/**
* Cut a structured document text pack into chunks for embedding. The
* buffer is transferred to the worker, so the caller's copy is detached.
*
* @param {ArrayBuffer} buf - The pack, as getStructuredDocumentText() returns it
* @param {Object} [options]
* @param {Boolean} [options.isPriority]
* @param {Boolean} [options.positions] - Give each chunk the reader
* `positions` its anchor resolves to
* @returns {Promise<Object>} - { chunks, sourceHash, chunkerVersion }
*/
async getStructuredDocumentTextChunks(buf, { isPriority = false, positions = false } = {}) {
return this._enqueue(async () => {
try {
var result = await this._query('sdt.getChunks', { buf, positions }, [buf]);
}
catch (e) {
this._throwWorkerError('sdt.getChunks', e);
}
return result;
}, !!isPriority);
}
/**
* Read chunk anchors back from a structured document text pack. The
* buffer is transferred to the worker, so the caller's copy is detached.
*
* @param {ArrayBuffer} buf - The pack, as getStructuredDocumentText() returns it
* @param {Object[]} anchors - A chunk's `anchor` per entry
* @param {Object} [options]
* @param {Boolean} [options.isPriority]
* @returns {Promise<Object[]>} - Per anchor, { text, outlinePath,
* pageLabel, positions } or null where it doesn't resolve
*/
async readStructuredDocumentTextAnchors(buf, anchors, { isPriority = false } = {}) {
return this._enqueue(async () => {
try {
var result = await this._query('sdt.readAnchors', { buf, anchors }, [buf]);
}
catch (e) {
this._throwWorkerError('sdt.readAnchors', e);
}
return result;
}, !!isPriority);
}
/**
* Get data for recognizer-server
*

View file

@ -39,14 +39,17 @@ Zotero.SDT = new function () {
let _documentWorkerMetadata = null;
let _documentWorkerMetadataErrorLogged = false;
// Load the bundled SDT module lazily, so that a missing or broken
// resource degrades to an 'unavailable' result instead of breaking
// Zotero startup. Load failures aren't cached, so a load is retried
// on the next call (require() itself caches successful loads)
// Load the bundled SDT modules -- the pack reader and the chunker --
// lazily, so a missing or broken resource degrades to 'unavailable'
// rather than breaking startup; a failed load is retried on the next
// call (require() itself caches successful loads)
function _getModule() {
if (!_module) {
try {
_module = require('resource://zotero/document-worker/structured-document-text.js');
_module = {
...require('resource://zotero/document-worker/structured-document-text.js'),
...require('resource://zotero/document-worker/structured-document-text-chunker.js')
};
}
catch (e) {
if (!_moduleErrorLogged) {
@ -61,22 +64,23 @@ Zotero.SDT = new function () {
/**
* Get the structured document text pack for a PDF, EPUB, or snapshot
* attachment, generating and caching it if necessary. The returned bytes
* are owned by the caller -- a later regeneration can't affect them.
*
* If the cached pack was produced by an older (but still readable)
* processor version, it's returned as is and a regeneration is started in
* the background, so that processor bumps don't block consumers.
* attachment, generating and caching it if necessary. A pack from an
* older but readable processor version is returned as is and regenerated
* in the background, unless allowStale is false.
*
* @param {Integer} itemID
* @param {Object} [options]
* @param {Boolean} [options.isPriority] - Put a needed extraction at the
* front of the worker queue (for user-initiated requests)
* @param {Boolean} [options.allowStale=true] - Whether a cached pack from
* an older processor version may be returned
* @param {Boolean} [options.cachedOnly] - Return 'not-cached' rather than
* extracting the document when no pack is cached
* @param {Function} [options.onProgress] - Called with SDT generation
* progress from 0 to 100 when generation is needed
* @returns {Promise<Object>} { ok: true, bytes: ArrayBuffer, packVersion,
* schemaMajorVersion }, or { ok: false, reason: 'unavailable' |
* 'password-required' | 'failed' }
* 'password-required' | 'not-cached' | 'failed' }
*/
this.getPack = async function (itemID, options = {}) {
try {
@ -87,13 +91,20 @@ Zotero.SDT = new function () {
if (!context.ok) {
return { ok: false, reason: context.reason };
}
let cache = await _readValidCache(context, { allowStaleProcessorVersion: true });
let cache = await _readValidCache(context, {
allowStaleProcessorVersion: options.allowStale !== false
});
if (cache.ok) {
if (cache.staleProcessorVersion) {
_generate(context, {}).catch(e => Zotero.logError(e));
}
return _makeResult(cache);
}
// Extracting a document costs seconds; a caller that only wants
// structure if it's already there says so rather than waiting
if (options.cachedOnly) {
return { ok: false, reason: 'not-cached' };
}
return await _generate(context, options);
}
catch (e) {
@ -103,13 +114,9 @@ Zotero.SDT = new function () {
};
/**
* Ensure that a current pack is cached for an attachment, generating or
* regenerating it if necessary, without returning it. For warming up the
* cache (e.g., at import time), so that later getPack() calls are hits.
*
* Unlike getPack(), which returns a stale-processor pack immediately and
* regenerates in the background, this resolves only once the cache is
* fully current.
* Ensure a current pack is cached for an attachment, generating or
* regenerating it if necessary, without returning it. Unlike getPack(),
* this resolves only once the cache is current.
*
* @param {Integer} itemID
* @param {Object} [options] - See getPack()
@ -138,20 +145,159 @@ Zotero.SDT = new function () {
};
/**
* Get a parsed pack reader for in-process consumers
* Whether a pack is cached for an attachment, by the cache file's
* presence alone -- nothing read or validated, for deciding cheaply
* whether to extract ahead of time
*
* @param {Integer} itemID
* @param {Object} [options] - See getPack()
* @returns {Promise<Object|null>}
* @returns {Promise<Boolean>}
*/
this.getReader = async function (itemID, options = {}) {
this.isCached = async function (itemID) {
let item = await Zotero.Items.getAsync(itemID, { noCache: true });
if (!item || !item.isAttachment() || !_getProcessorType(item)) {
return false;
}
return IOUtils.exists(_getCachePath(item));
};
/**
* The chunks of an attachment's structured text, for embedding, cut in
* the document worker so that neither inflating the pack nor cutting it
* blocks this thread. Each carries its plain `text`, the `embedText`
* with outline context woven in, its estimated `tokens`, `outlinePath`,
* `pageLabel` and `anchor` -- where in the file its text is, null when
* the document has no geometry for it.
*
* @param {Integer} itemID
* @param {Object} [options] - getPack()'s options, `isPriority` putting
* the cut at the front of the worker queue too
* @param {Boolean} [options.positions] - Give each chunk the reader
* `position` its anchor opens at: the first of the positions its
* text is shown at, null when the anchor doesn't resolve
* @returns {Promise<Object>} - { ok: true, chunks, contentHash }, the
* hash being the source file's as the pack records it, or
* { ok: false, reason }: getPack()'s reasons, or 'cut-failed' when
* the pack is there but the worker couldn't cut it
*/
this.getItemChunks = async function (itemID, { positions = false, ...options } = {}) {
let result = await this.getPack(itemID, options);
if (!result.ok) {
return { ok: false, reason: result.reason };
}
let cut;
try {
cut = await Zotero.PDFWorker.getStructuredDocumentTextChunks(
result.bytes, { isPriority: options.isPriority, positions });
}
catch {
// Logged by the worker manager
return { ok: false, reason: 'cut-failed' };
}
// The worker and the bundled chunker come from the same build, so a
// mismatch is a build defect, not something to work around
let version = _getModule().CHUNKER_VERSION;
if (cut.chunkerVersion !== version) {
Zotero.logError(new Error(`Document worker chunker version ${cut.chunkerVersion} `
+ `doesn't match the bundled chunker's ${version}`));
return { ok: false, reason: 'cut-failed' };
}
let chunks = positions ? cut.chunks.map(_withFirstPosition) : cut.chunks;
return { ok: true, chunks, contentHash: cut.sourceHash };
};
/**
* The chunks of a plain text with no structure, paragraphs separated by
* blank lines, as getItemChunks() gives them but without an anchor
*
* @param {String} text
* @param {Object} [options] - The module's chunking options
* @returns {Object[]}
*/
this.getPlainTextChunks = function (text, options = {}) {
return _getModule().getPlainTextChunks(text, options);
};
/**
* Read chunks back from an attachment's structured text by their stored
* anchors, in the document worker so that neither inflating the pack
* nor reading it blocks this thread. An anchor on nothing the current
* extraction has, as after the file changed, reads back as null.
*
* @param {Integer} itemID
* @param {Object[]} anchors - A chunk's `anchor` per entry
* @param {Object} [options] - getPack()'s options, `isPriority` putting
* the read at the front of the worker queue too
* @returns {Promise<Object>} { ok: true, chunks: [{ text, outlinePath,
* pageLabel, position } | null] }, `position` the reader position
* the chunk opens at, or { ok: false, reason }: getPack()'s
* reasons, or 'failed' when the worker couldn't read the pack
*/
this.readAnchors = async function (itemID, anchors, options = {}) {
let result = await this.getPack(itemID, options);
if (!result.ok) {
return { ok: false, reason: result.reason };
}
try {
let chunks = await Zotero.PDFWorker.readStructuredDocumentTextAnchors(
result.bytes, anchors, { isPriority: options.isPriority });
return { ok: true, chunks: chunks.map(chunk => chunk && _withFirstPosition(chunk)) };
}
catch {
// Logged by the worker manager
return { ok: false, reason: 'failed' };
}
};
/**
* The identity of the extraction and division this module would produce
* for an attachment right now: its processor type and version, the pack
* schema's major version and the chunker's version. Null when the
* attachment isn't a supported type or the extraction module isn't
* available.
*
* @param {Zotero.Item} item
* @returns {Promise<String|null>} - e.g. 'pdf/3/1/1'
*/
this.getProcessorVersion = async function (item) {
let processorType = _getProcessorType(item);
if (!processorType) {
return null;
}
return _openPack(new Uint8Array(result.bytes));
let metadata = await _getDocumentWorkerMetadata();
if (!metadata) {
return null;
}
return _getProcessorVersion(metadata, processorType);
};
/**
* The identities getProcessorVersion() gives right now, one per
* processor type. Empty when the extraction module isn't available.
*
* @returns {Promise<String[]>}
*/
this.getProcessorVersions = async function () {
let metadata = await _getDocumentWorkerMetadata();
if (!metadata) {
return [];
}
return Object.keys(metadata.SDT_PROCESSOR_VERSIONS)
.map(processorType => _getProcessorVersion(metadata, processorType));
};
function _getProcessorVersion(metadata, processorType) {
return processorType
+ '/' + metadata.SDT_PROCESSOR_VERSIONS[processorType]
+ '/' + _getSchemaMajorVersion(metadata.SDT_SCHEMA_VERSION)
+ '/' + _getModule().CHUNKER_VERSION;
}
// A chunk from the worker with only the first of its positions, the one
// the reader opens at
function _withFirstPosition({ positions, ...chunk }) {
return { ...chunk, position: positions?.[0] || null };
}
async function _readValidCache({ sourceHash, cachePath, processorType }, options) {
let bytes;
try {
@ -378,8 +524,10 @@ Zotero.SDT = new function () {
}
async function _getAttachmentContext(itemID) {
// getAsync() returns false, not null, for a nonexistent item
let item = await Zotero.Items.getAsync(itemID);
// getAsync() returns false, not null, for a nonexistent item. The
// item is read, never shown or saved, so it isn't kept in the cache
// when it wasn't there already.
let item = await Zotero.Items.getAsync(itemID, { noCache: true });
if (!item || !item.isAttachment()) {
return { ok: false, reason: 'unavailable' };
}

View file

@ -1503,7 +1503,9 @@ Zotero.SearchQuery = new function () {
? { joinMode: 'all', children: [tree] }
: tree);
if (text) {
if (mode === 'bestMatch' && Zotero.Embeddings?.isEnabled()) {
// Best Match ranks by the text -- lexically when no semantic
// index is available -- rather than filtering by it
if (mode === 'bestMatch') {
search.addCondition('bestMatch', 'contains', text);
}
else {

View file

@ -536,6 +536,94 @@ Zotero.Sync.APIClient.prototype = {
},
/**
* Which of a library's embeddings of a model have changed since a
* version: the library's current version and, per item key, the version
* its rows last changed at, with the models the server embeds the
* library with when it says.
*
* @param {String} model - Zotero.Embeddings.getModelVersion()
* @param {Integer} [since]
* @return {Promise<Object|false>} - { version, items: { key: version },
* models }, or false when nothing has changed since `since`
*/
getEmbeddingVersions: async function (libraryType, libraryTypeID, model, since) {
var params = {
libraryType: libraryType,
libraryTypeID: libraryTypeID,
target: "embeddings",
format: "versions",
model
};
if (since) {
params.since = since;
}
var uri = this.buildRequestURI(params);
var options = {
successCodes: [200, 304, 404]
};
if (since) {
options.headers = {
"If-Modified-Since-Version": since
};
}
var xmlhttp = await this.makeRequest("GET", uri, options);
if (xmlhttp.status == 304) {
return false;
}
if (xmlhttp.status == 404) {
return { version: 0, items: {}, models: [] };
}
var json = this._parseJSON(xmlhttp.responseText);
return {
version: json.version || 0,
items: json.items || {},
models: json.models || null
};
},
/**
* The server's answer for each of the given items, in a model: its
* embeddings, each vector base64 of its stored int8 form, and the
* version they're from; a refusal; or word that the server is still
* working on it. The model is echoed.
*
* @param {String} model - Zotero.Embeddings.getModelVersion()
* @param {String[]} itemKeys
* @return {Promise<Object>} - { model, dtype, items: [
* { key, status: 'success', contentHash, chunks, version,
* rows: [{ chunkIndex, embedding, anchor }] }
* | { key, status: 'declined' }
* | { key, status: 'pending' } ] },
* each embedding base64 in the encoding `dtype` names; model and dtype
* null and items empty when the server has no embeddings for the
* library
*/
getEmbeddings: async function (libraryType, libraryTypeID, model, itemKeys = []) {
var params = {
libraryType: libraryType,
libraryTypeID: libraryTypeID,
target: "embeddings",
model
};
if (itemKeys.length) {
params.itemKey = itemKeys.join(',');
}
var uri = this.buildRequestURI(params);
var xmlhttp = await this.makeRequest("GET", uri, { successCodes: [200, 404] });
if (xmlhttp.status == 404) {
return { model: null, dtype: null, items: [] };
}
var json = this._parseJSON(xmlhttp.responseText);
return {
model: json.model || null,
dtype: json.dtype || null,
items: json.items || []
};
},
createAPIKeyFromCredentials: async function (username, password) {
var body = JSON.stringify({
username,
@ -825,7 +913,8 @@ Zotero.Sync.APIClient.prototype = {
'sort',
'direction',
'since',
'sincetime'
'sincetime',
'model'
];
queryParams = {};

View file

@ -304,6 +304,10 @@ Zotero.Sync.Runner_Module = function (options = {}) {
attempt++;
continue;
}
_stopCheck();
await _doEmbeddingsCheck(client, [...successfulLibraries]);
break;
}
}
@ -831,6 +835,23 @@ Zotero.Sync.Runner_Module = function (options = {}) {
}
return resyncLibraries;
}.bind(this);
/**
* Have each library's embeddings fetched again if the server has embedded
* it anew.
*/
var _doEmbeddingsCheck = async function (client, libraries) {
for (let libraryID of libraries) {
_stopCheck();
try {
await Zotero.Embeddings.Sync.checkLibrary(client, libraryID);
}
catch (e) {
Zotero.logError(e);
}
}
};
/**

View file

@ -1620,7 +1620,184 @@ Zotero.Utilities.Internal = {
stringWithColon: function (str) {
return Zotero.getString('punctuation.colon.withString', str);
},
// Section titles of a reference list, lowercase and without punctuation.
// The titles are the PDF extractor's (document-worker's
// src/pdf/structure/reference/titles.js).
_referenceHeadings: new Set([
// English
'references',
'bibliography',
'references and notes',
'works cited',
'literature cited',
'cited references',
'list of references',
'selected references',
'sources',
'citations',
// French
'références',
'bibliographie',
'références bibliographiques',
// German
'literaturverzeichnis',
'quellenverzeichnis',
'referenzen',
'literatur',
'literaturangaben',
// Spanish, Galician
'referencias',
'bibliografía',
'referencias bibliográficas',
// Portuguese
'referências',
'bibliografia',
'referências bibliográficas',
// Italian
'riferimenti',
'riferimenti bibliografici',
// Dutch
'referenties',
'literatuurlijst',
'bibliografie',
'bronnen',
// Swedish
'referenser',
'referenslista',
'litteraturförteckning',
'källförteckning',
// Norwegian, Danish
'referanser',
'referencer',
'litteraturliste',
'kildeliste',
// Finnish
'lähteet',
'viitteet',
'kirjallisuusluettelo',
// Polish, Czech, Slovak
'literatura',
'piśmiennictwo',
'wykaz literatury',
'seznam literatury',
'literatúra',
'zoznam literatúry',
// Russian
'список литературы',
'литература',
'библиография',
'источники',
// Ukrainian
'список літератури',
'література',
'бібліографія',
'джерела',
// Romanian
'referințe',
'lista de referințe',
// Turkish
'kaynaklar',
'kaynakça',
// Greek
'βιβλιογραφία',
'αναφορές',
'παραπομπές',
// Arabic
'المراجع',
'المصادر',
'قائمة المراجع',
// Persian
'منابع',
'مراجع',
'کتابنامه',
// Hebrew
'ביבליוגרפיה',
'מקורות',
'רשימת מקורות',
// Chinese, Japanese
'参考文献',
'参考资料',
'文献',
// Korean
'참고문헌',
// Hindi
'संदर्भ',
'संदर्भ सूची',
'ग्रंथ सूची',
// Bengali
'তথ্যসূত্র',
'গ্রন্থপঞ্জি',
// Indonesian, Malay
'daftar pustaka',
'referensi',
'rujukan',
'senarai rujukan',
// Vietnamese
'tài liệu tham khảo',
'tài liệu',
// Thai
'บรรณานุกรม',
'เอกสารอ้างอิง',
// Tagalog
'mga sanggunian',
'mga talaakdaan',
// Catalan
'referències',
// Basque
'erreferentziak',
// Hungarian
'irodalomjegyzék',
'hivatkozások',
// Serbian, Croatian, Bosnian, Slovenian
'popis literature',
'bibliografija',
'reference',
'viri',
// Lithuanian, Latvian
'literatūra',
'šaltiniai',
'atsauces',
'bibliogrāfija',
// Estonian
'kirjandus',
'viited',
'bibliograafia',
// Bulgarian, Macedonian
'използвана литература',
'референци',
'библиографија',
// Albanian
'referencat',
'bibliografi',
// Georgian
'ლიტერატურა',
'წყაროები',
'ბიბლიოგრაფია',
// Armenian
'գրականություն',
'աղբյուրներ',
'մատենագիտություն'
]),
/**
* Whether a string is likely a reference list's title: without
* punctuation and in lowercase, it's one of _referenceHeadings
*
* @param {String} str
* @return {Boolean}
*/
isLikelyReference: function (str) {
let normalized = str
.normalize('NFKC')
.replace(/[^\p{L}\p{M}\p{N}]+/gu, ' ')
.trim()
.toLowerCase();
return Zotero.Utilities.Internal._referenceHeadings.has(normalized);
},
/**
* Resolve `locale` to the best-fit entry in `locales`.

View file

@ -108,6 +108,8 @@ const xpcomFilesLocal = [
'httpIntegrationClient',
'id',
'integration',
'lexical',
'bestMatch',
'locale',
'locateManager',
'mime',

View file

@ -2082,7 +2082,8 @@ var ZoteroPane = new function () {
// Reproduce the quick search mode as editable conditions, one per word joined
// with "all": Title/Creator/Year and All Fields & Tags each map to a single
// condition, Everything to an "any" group of Any Field plus full-text.
// condition, Everything to an "any" group of Any Field plus full-text, and
// Best Match to the ranking field with the text as typed.
// Title/Creator/Year matches only top-level items, so set the result level to item.
if (tree) {
// The words are joined to the conditions with "all", so an "any"
@ -2094,8 +2095,14 @@ var ZoteroPane = new function () {
if (mode === 'titleCreatorYear') {
search.addCondition('resultLevel', 'item');
}
// Best Match ranks by the free text rather than filtering by its
// words. The pane shows the ranking field for top-level item results.
if (mode === 'bestMatch') {
search.addCondition('bestMatch', 'contains', searchText.trim());
search.addCondition('resultLevel', 'item');
if (text) {
search.addCondition('bestMatch', 'contains', text);
}
parts = [];
}
for (let part of parts) {
if (mode === 'everything') {
@ -2272,11 +2279,21 @@ var ZoteroPane = new function () {
return false;
}
var selectedItems = this.itemsView.getSelectedObjects();
var selectedObjects = this.itemsView.getSelectedObjects();
// A selection of search-match rows names passages rather than
// items, and the pane shows those instead
var searchMatches = this.itemsView.getSelectedSearchMatches();
// The pane shows data objects; rows standing in for something else
// (a search-match row, a library header) have none to show, so a
// selection of only those reads as an empty one
var selectedItems = selectedObjects.filter(o => o instanceof Zotero.DataObject);
// Display buttons at top of item pane depending on context. This needs to run even if the
// selection hasn't changed, because the selected items might have been modified.
this.itemPane.data = selectedItems;
this.itemPane.searchMatches = searchMatches;
this.itemPane.collectionTreeRows = collectionTreeRows;
this.itemPane.itemsView = this.itemsView;
this.itemPane.editable = this.collectionsView.editable;
@ -2291,7 +2308,10 @@ var ZoteroPane = new function () {
// Check if selection has actually changed. The onselect event that calls this
// can be called in various situations where the selection didn't actually change,
// such as whenever selectEventsSuppressed is set to false.
var ids = selectedItems.map(item => item.treeViewID);
// Keyed on what's selected rather than what the pane shows, so
// moving between two passages of the same attachment still
// counts as a change
var ids = selectedObjects.map(o => o.treeViewID);
ids.sort();
if (ids.length && Zotero.Utilities.arrayEquals(_lastSelectedItems, ids)) {
return false;
@ -4571,10 +4591,11 @@ var ZoteroPane = new function () {
show.add(m.showInLibrary);
show.add(m.sep1);
}
[
m.showInLibrary,
m.duplicateItem,
m.changeParentItem,
m.addToCollection,
m.removeItems,
m.moveToTrash,
m.deleteFromLibrary,
@ -4582,7 +4603,7 @@ var ZoteroPane = new function () {
m.createBib,
m.loadReport
].forEach(x => disable.add(x));
}
// Show "Export Note…" if all notes or attachments
@ -5588,6 +5609,14 @@ var ZoteroPane = new function () {
let { noLocateOnMissing } = options;
for (let i = 0; i < items.length; i++) {
let item = items[i];
// A search-match row stands in for a passage of its attachment
// rather than for an item, so it opens at that passage
if (!(item instanceof Zotero.Item)) {
if (item.itemID && item.entry) {
await this.viewSearchMatch(item.itemID, item.entry, event);
}
continue;
}
if (item.isRegularItem()) {
// Prefer local file attachments
let attachment = await item.getBestAttachment();
@ -5858,6 +5887,32 @@ var ZoteroPane = new function () {
};
/**
* Open the attachment a best-match search passage came from, at the
* passage.
*
* A PDF passage carries the page geometry the reader scrolls to and
* highlights. An EPUB or snapshot passage carries none -- their views
* navigate by DOM selector, which a chunk doesn't know -- so those open
* where the attachment was left.
*
* @param {Number} itemID - The attachment the passage belongs to
* @param {Object} entry - A preview entry (see
* Zotero.BestMatch.Session#getPreviews())
* @param {Event} [event]
* @return {Promise}
*/
this.viewSearchMatch = async function (itemID, entry, event = null) {
let item = Zotero.Items.get(itemID);
if (!item || !item.isFileAttachment()) {
return;
}
let position = entry?.position;
await this.viewAttachment(itemID, event, false,
position ? { location: { position } } : undefined);
};
/**
* Update the parent of the selected items
*

View file

@ -92,17 +92,47 @@ preferences-advanced-server-disabled = The { -app-name } HTTP server is disabled
preferences-advanced-server-enable-and-restart =
.label = Enable and Restart
preferences-advanced-semantic-search-title = Best-Match Search
preferences-advanced-semantic-search-model = Mode:
preferences-advanced-semantic-search-status = Status:
preferences-advanced-semantic-search-enable = Enable search by semantic meaning:
preferences-advanced-semantic-search-enabled =
.label = Enabled
preferences-advanced-semantic-search-disabled =
.label = Disabled
preferences-advanced-semantic-search-english =
.label = English
preferences-advanced-semantic-search-multilingual =
.label = Multilingual
preferences-advanced-semantic-search-model-description = “{ preferences-advanced-semantic-search-english.label }” gives the best results for libraries with English-language content. “{ preferences-advanced-semantic-search-multilingual.label }” supports searching in and across many languages.
preferences-advanced-best-match-engine = Engine (temporary, for testing):
preferences-advanced-best-match-engine-hybrid =
.label = Hybrid
preferences-advanced-best-match-engine-lexical =
.label = Lexical
preferences-advanced-best-match-engine-semantic =
.label = Semantic
preferences-advanced-best-match-margin = Quality cutoff below top (%):
preferences-advanced-best-match-margin-description = At 30%, results scoring up to 30% below the best one are kept; 100% keeps every result, 0 drops everything but the top match.
preferences-advanced-semantic-search-downloading = Downloading…
preferences-advanced-semantic-search-downloading-progress = Downloading… { $percent }%
preferences-advanced-semantic-search-preparing = Preparing documents…
preferences-advanced-semantic-search-preparing-progress = Preparing documents… { $done } / { $total }
preferences-advanced-semantic-search-indexing-documents = Indexing documents…
preferences-advanced-semantic-search-fetching-documents = Syncing semantic data… { $done } / { $total }
preferences-advanced-semantic-search-indexing-documents-locally =
{ $declined ->
[0] Generating semantic data locally for { $count ->
[one] { $count } item
*[other] { $count } items
}…
*[other] Generating semantic data locally for { $count ->
[one] { $count } item
*[other] { $count } items
}… ({ $declined ->
[one] { $declined } not available
*[other] { $declined } not available
} from the server)
}
preferences-advanced-semantic-search-awaiting = Waiting for embeddings from the server…
preferences-advanced-semantic-search-server-unreachable =
{ $minutes ->
[one] The server couldn’t be reached. Trying again in { $minutes } minute…
*[other] The server couldn’t be reached. Trying again in { $minutes } minutes…
}
preferences-advanced-semantic-search-indexing = Indexing…
preferences-advanced-semantic-search-stopping = Stopping…
preferences-advanced-semantic-search-idle = Up to date
@ -112,12 +142,58 @@ preferences-advanced-semantic-search-resume =
.label = Resume
preferences-advanced-semantic-search-stop =
.label = Stop
preferences-advanced-semantic-search-switch-title = Change Mode?
preferences-advanced-semantic-search-switch-text = Changing the mode will rebuild the search index, which can take a long time for large libraries.
preferences-advanced-semantic-search-switch-button = Change Mode
preferences-advanced-semantic-search-disable-title = Disable Best-Match Search?
preferences-advanced-semantic-search-disable-text = Disabling will delete the search index and downloaded data.
preferences-advanced-semantic-search-disable-button = Disable
preferences-advanced-semantic-search-items = Metadata, notes, annotations
preferences-advanced-semantic-search-attachments = Attachments
preferences-advanced-semantic-search-progress-value = { $percent }%
preferences-advanced-semantic-search-diagnostics-show =
.label = Show Diagnostics
preferences-advanced-semantic-search-diagnostics-hide =
.label = Hide Diagnostics
preferences-advanced-semantic-search-endpoint-configure =
.label = Configure Endpoint…
preferences-advanced-semantic-search-endpoint-off =
.value = Embedding endpoint: not configured
preferences-advanced-semantic-search-endpoint-unverified =
.value = Embedding endpoint: not verified
preferences-advanced-semantic-search-endpoint-valid =
.value = Embedding endpoint: valid
preferences-advanced-semantic-search-endpoint-unreachable =
.value = Embedding endpoint: unreachable
preferences-advanced-semantic-search-endpoint-invalid =
.value = Embedding endpoint: not valid
preferences-advanced-semantic-search-endpoint-dialog =
.title = Configure Embedding Endpoint
preferences-advanced-semantic-search-endpoint-intro =
.value = Send the embedding workload to a local or remote server:
preferences-advanced-semantic-search-endpoint-url =
.value = Server URL:
preferences-advanced-semantic-search-endpoint-privacy = A remote server receives the text of your library.
preferences-advanced-semantic-search-endpoint-requirements =
.value = The server must serve the model Zotero is using:
preferences-advanced-semantic-search-endpoint-model =
.value = Model:
preferences-advanced-semantic-search-endpoint-pooling =
.value = Pooling:
preferences-advanced-semantic-search-endpoint-file =
.value = File:
preferences-advanced-semantic-search-endpoint-llama =
.value = Run with llama.cpp:
preferences-advanced-semantic-search-endpoint-then-use =
.value = Then use { $url }
preferences-advanced-semantic-search-endpoint-copy =
.label = Copy
preferences-advanced-semantic-search-endpoint-verifying =
.value = Checking that the server’s vectors match the model’s…
preferences-advanced-semantic-search-endpoint-accept =
.label = Verify and Use
preferences-advanced-semantic-search-endpoint-remove =
.label = Stop Using Endpoint
preferences-advanced-semantic-search-endpoint-error-unreachable = The server didn’t respond.
preferences-advanced-semantic-search-endpoint-error-unauthorized = The server requires authentication, which isn’t supported.
preferences-advanced-semantic-search-endpoint-error-not-embeddings = That URL isn’t an embeddings endpoint Zotero can use (OpenAI-style or Text Embeddings Inference).
preferences-advanced-semantic-search-endpoint-error-width-mismatch = The server is serving a different model.
preferences-advanced-semantic-search-endpoint-error-low-agreement = The server’s vectors don’t match this model’s. Check that it serves the model above with { $pooling } pooling.
preferences-advanced-semantic-search-endpoint-error-context-too-small = The server can’t embed long passages. Start it with a context of at least 4096 tokens.
preferences-advanced-language-and-region-title = Language and Region
preferences-advanced-enable-bidi-ui =
.label = Enable bidirectional text editing utilities
@ -190,3 +266,4 @@ fulltext-stats-attachments-indexed = Attachments indexed:
fulltext-stats-partially-indexed = Partially indexed:
fulltext-stats-not-available = Full-text content or file not available:
fulltext-stats-notes-indexed = Notes indexed:
fulltext-stats-items-indexed = Items and annotations indexed:

View file

@ -466,6 +466,8 @@ items-column-relevance-rank = Rank { $rank }
items-best-match-indexing = Indexing in progress — { $indexed } of { $total } items indexed
items-best-match-indexing-paused = Indexing is paused — { $indexed } of { $total } items indexed
# $page (String) - a page label, e.g. "12" or "ix"
items-search-match-page = p. { $page }
report-error =
.label = Report Error…
@ -651,6 +653,7 @@ pane-related = Related
pane-attachment-info = Attachment Info
pane-attachment-preview = Preview
pane-attachment-annotations = Annotations
pane-search-results = Search Results
pane-header-attachment-associated =
.label = Rename associated file
@ -690,6 +693,14 @@ section-related =
.label = { $count } Related
section-attachment-info =
.label = { pane-attachment-info }
section-search-results =
.label = { $count ->
[one] { $count } Search Result
*[other] { $count } Search Results
}
search-result-row-fulltext = Full Text
search-result-row-show-more = Show More
search-result-row-show-less = Show Less
section-button-remove =
.tooltiptext = { general-remove }
@ -728,6 +739,8 @@ sidenav-attachment-preview =
.tooltiptext = { pane-attachment-preview }
sidenav-attachment-annotations =
.tooltiptext = { pane-attachment-annotations }
sidenav-search-results =
.tooltiptext = { pane-search-results }
sidenav-libraries-collections =
.tooltiptext = { pane-libraries-collections }
sidenav-tags =

View file

@ -108,16 +108,32 @@ pref("extensions.zotero.keys.toggleRead", "`");
pref("extensions.zotero.keys.showTabsMenu", ";");
pref("extensions.zotero.search.quicksearch-mode", "fields");
// Temporary, for testing: which engine best-match search runs -- 'lexical',
// 'semantic', or 'hybrid' (both, fused)
pref("extensions.zotero.search.bestMatchEngine", "hybrid");
// How far below an engine's strongest match an item may fall and still count
// as one of its results, as a percentage of that strongest score -- 10 keeps
// items within 10 percent of the top, 100 keeps everything above the floor
pref("extensions.zotero.search.bestMatchMargin", 50);
// Whether best-match search ranks semantically. Off, nothing is indexed and
// ranking is lexical only; what's already indexed is kept either way.
pref("extensions.zotero.search.bestMatch.enableSemantic", false);
// Fulltext indexing
pref("extensions.zotero.fulltext.textMaxLength", 500000);
pref("extensions.zotero.fulltext.pdfMaxPages", 100);
pref("extensions.zotero.search.useLeftBound", true);
// Semantic search embeddings - disabled when "";
pref("extensions.zotero.embeddings.model", "");
// Set when the user stops indexing; nothing is indexed until indexing is started again
// Set when the user stops indexing or turns semantic search off; nothing is indexed until indexing is started again
pref("extensions.zotero.embeddings.indexingPaused", false);
// An OpenAI-style /v1/embeddings URL to send indexing to, used only once
// verified to serve the active model (see Zotero.Embeddings.Endpoint)
pref("extensions.zotero.embeddings.endpoint", "");
// Whether attachments kept in Zotero Storage are left for the server to
// embed, their rows fetched as indexing reaches them. Not shown in the UI:
// an override for testing, since a syncing account uses the server as a
// matter of course.
pref("extensions.zotero.embeddings.sync.enabled", true);
// Notes
pref("extensions.zotero.note.fontFamily", "-apple-system, BlinkMacSystemFont, \"Segoe UI\", \"Helvetica Neue\", Helvetica, Arial, sans-serif");

View file

@ -8,7 +8,7 @@ const { getSignatures, writeSignatures, onSuccess, onError } = require('./utils'
const { buildsURL } = require('./config');
const sharedAssetDirs = ['cmaps', 'standard_fonts'];
const requiredFiles = ['worker.js', 'metadata.json', 'structured-document-text.js'];
const requiredFiles = ['worker.js', 'metadata.json'];
async function getDocumentWorker(signatures) {
const t1 = Date.now();

View file

@ -0,0 +1,348 @@
{
"en": [
{
"query": "apical epithelial cap",
"passage": "Following partial amputation of the caudal fin at the mid-ray level, a wound epidermis formed within six hours and thickened into an apical epithelial cap by eighteen hours post-injury. Beneath this cap, mesenchymal cells derived from mature osteoblasts underwent dedifferentiation, downregulating osterix and re-entering the cell cycle to establish a proliferative blastema by forty-eight hours. Bromodeoxyuridine incorporation peaked at day three, with labeling indices reaching 34 percent in the distal blastema compared to 6 percent in adjacent unamputated tissue. Pharmacological inhibition of Wnt/beta-catenin signaling with a small-molecule antagonist applied continuously from day one reduced outgrowth length at day fourteen by roughly 40 percent relative to vehicle-treated controls, and this effect was reversible when the inhibitor was withdrawn after day seven. Aged fish, defined here as those older than twenty-four months, regenerated fin rays at approximately 0.18 millimeters per day, compared with 0.31 millimeters per day in six-month-old siblings, a difference attributable in part to a delayed onset of blastemal proliferation rather than a slower maximal rate once proliferation began. Lineage tracing using a photoconvertible reporter confined to osteoblast precursors showed that fin ray segments regenerated almost exclusively from local positional memory, with fewer than 3 percent of labeled cells migrating more than two segments from their origin. Denervation of the fin prior to amputation blunted blastema growth by half, consistent with a permissive role for peripheral nerve-derived signals in sustaining early proliferation, though the specific secreted factor responsible was not isolated in this cohort.",
"nearMiss": "Adult zebrafish carrying a temperature-sensitive allele affecting xanthophore survival were shifted from 25 to 31 degrees Celsius at four weeks of age, and the body stripe pattern was photographed every third day for the following five weeks to follow how melanophore stripes reorganize when the interleaving yellow cells are lost. Within twelve days of the shift, dark stripes on the flank widened by an average of 38 percent and their borders lost the sharp edge seen in siblings kept at the permissive temperature, and by day thirty melanophores had spread into a nearly uniform field over the posterior trunk. Returning fish to 25 degrees allowed xanthophores to repopulate from a small reserve of precursor cells along the myosepta, and stripes reformed over the following month in positions that differed from the original pattern in eleven of fourteen animals, indicating that stripe placement is set by ongoing cell interactions rather than by a fixed positional memory. Time-lapse imaging of the interstripe region during recovery showed melanophores retreating from arriving xanthophores at contact, consistent with the short-range repulsion proposed in earlier modeling work."
},
{
"query": "COX-deficient fibers and mitochondrial deletion load",
"passage": "Vastus lateralis biopsies were obtained from a pilot cohort of sixty-one adults spanning ages forty to eighty-nine, and mitochondrial DNA heteroplasmy at seven previously characterized deletion breakpoints was quantified using digital droplet PCR with allele-specific probes. Heteroplasmy levels were low and relatively uniform across whole-muscle homogenates, averaging 2.1 percent, but single-fiber dissection revealed a markedly different picture: a subset of individual fibers, rising from under 1 percent in donors under fifty to nearly 14 percent in donors over eighty, carried deletion loads exceeding the 60 percent threshold associated with a biochemical defect. These high-heteroplasmy fibers stained negative for cytochrome c oxidase while retaining strong succinate dehydrogenase activity, the classic mosaic pattern of segmental respiratory chain deficiency, and their cross-sectional area was on average 22 percent smaller than neighboring COX-positive fibers within the same section. Clonal expansion of a single deletion species within a fiber, rather than accumulation of many distinct low-level deletions, accounted for the majority of high-heteroplasmy fibers examined by long-range sequencing, supporting a model in which stochastic replicative advantage of deleted genomes during satellite-cell-mediated fiber repair drives focal energetic failure. Grip strength correlated inversely with the proportion of COX-deficient fibers per cross-section (r = -0.46), though this relationship was substantially weaker than the correlation with total fiber number, suggesting that mosaic mitochondrial dysfunction is one contributor among several to age-related strength decline. Notably, six donors over seventy showed almost no COX-deficient fibers despite unremarkable physical activity histories, indicating considerable interindividual variability that body mass index, sex, and reported exercise frequency did not explain in this sample.",
"nearMiss": "Twelve competitive cyclists completed two glycogen-depleting rides separated by three weeks, each followed by a six-hour recovery period in which carbohydrate was given either as a single large bolus at the end of exercise or as six smaller feedings at hourly intervals, with total intake matched at 1.2 grams per kilogram of body mass per hour. Muscle glycogen was measured in biopsies taken from the thigh immediately after exercise and again at the end of the recovery window, using an enzymatic assay on freeze-dried tissue. Resynthesis over the six hours averaged 41 millimoles per kilogram dry mass under the hourly feeding schedule and 35 under the bolus condition, a difference that reached significance despite the small sample, and blood glucose remained within a narrower range under hourly feeding, with fewer readings above 8 millimoles per liter. Plasma insulin peaked higher after the bolus but had returned toward baseline by the third hour, whereas hourly feeding held it at a moderate plateau for the whole window. The authors suggest that the sustained insulin signal, rather than total carbohydrate delivered, explains the faster restoration, and note that riders reported less gastrointestinal discomfort on the divided schedule."
},
{
"query": "How much does the SLCO1B1 variant raise myopathy risk on simvastatin?",
"passage": "In a retrospective analysis of eight hundred and thirty patients initiated on simvastatin forty milligrams daily, carriers of the SLCO1B1 c.521C variant in homozygous form experienced myopathy, defined as unexplained muscle pain accompanied by creatine phosphokinase elevation above five times the upper reference limit, at a rate of 18 percent over twelve months, compared with 3 percent among those homozygous for the reference allele. Heterozygous carriers fell at an intermediate 7 percent, consistent with a gene-dose relationship. Plasma simvastatin acid concentrations, measured at steady state in a subset of 140 participants, were on average 2.4-fold higher in variant homozygotes, attributable to reduced hepatic uptake transporter activity and consequently prolonged systemic exposure rather than altered hepatic metabolism, since concentrations of the parent lactone were comparable across genotypes. Switching affected patients to a lower simvastatin dose or to a statin less dependent on this transporter, such as fluvastatin, resolved symptoms within six weeks in 81 percent of cases without recurrence over the following year. Discontinuation of statin therapy altogether occurred in 11 percent of the variant homozygote group versus 2 percent of the reference group, a difference with direct implications for long-term cardiovascular risk reduction given that nonadherence following an adverse event is rarely revisited in routine care. Genotyping cost, estimated at eighteen dollars per patient in this health system, was recovered within the first year through avoided emergency visits for suspected rhabdomyolysis in the small number of variant carriers who would otherwise have continued full-dose therapy unmonitored.",
"nearMiss": "A pharmacist-led review of long-term proton pump inhibitor prescriptions was run across nine primary care practices over eighteen months, identifying 1,146 adults who had been on daily therapy for more than a year without a documented indication that required indefinite treatment. Each patient received a letter explaining the review, followed by a structured telephone consultation in which the pharmacist proposed either a step-down to alternate-day dosing or a supervised stop with an on-demand antacid for breakthrough symptoms. At six months, 512 patients had discontinued the drug entirely and a further 203 remained on a reduced dose, while 187 had resumed daily therapy, most within the first eight weeks and most citing the return of reflux symptoms at night. Practices that scheduled a follow-up call at two weeks retained more patients off therapy than those relying on the patient to make contact, and patients over seventy-five were the least likely to restart. Prescribing costs across the nine practices fell by an estimated 31 percent for this drug class over the study period, and no admissions for gastrointestinal bleeding were recorded among participants during follow-up."
},
{
"query": "cryptochrome 4 expression",
"passage": "Captive European robins tested in Emlen funnels during the autumn migratory period consistently oriented their nocturnal restlessness toward the southwest, matching the expected seasonal heading, under the ambient geomagnetic field of approximately 46 microtesla. Application of a broadband radiofrequency field between 0.1 and 10 megahertz at intensities as low as 15 nanotesla abolished this directional preference, with birds instead distributing their hops randomly around the funnel, an effect that reversed within twenty minutes of switching off the field. Retinal expression of cryptochrome 4, assayed by quantitative immunohistochemistry, was substantially higher in the double cone and long-wavelength single cone photoreceptors of migratory-phase birds than in non-migratory-phase birds sampled six weeks apart, and this seasonal upregulation tracked closely with the onset of migratory restlessness in the same individuals. Site-directed mutagenesis of the terminal tryptophan in the electron transfer chain of purified robin cryptochrome 4, expressed heterologously in insect cells, extended the lifetime of the photoinduced radical pair from roughly one microsecond to over six microseconds, a change predicted by radical-pair models to increase magnetic sensitivity substantially. Birds with one eye covered retained compass orientation, ruling out a strictly monocular mechanism in this species, though response latency to a shifted magnetic field increased modestly under monocular occlusion. Disruption of the trigeminal nerve, by contrast, did not affect directional choice in the funnel assay, arguing against a role for the previously proposed magnetite-based receptor in this particular orientation behavior, at least as tested here.",
"nearMiss": "Blackcaps trapped at a coastal ringing station on the Baltic during three autumn seasons were weighed, scored for visible subcutaneous fat on a nine-point scale, and fitted with numbered rings so that individuals recaptured on later days could be followed through their stopover. Of 2,340 birds ringed, 411 were caught at least twice, and their mass gain averaged 0.42 grams per day, with the highest rates recorded in birds arriving lean in the first week of September when elderberry and dogwood fruit were most abundant along the station's hedgerows. Birds arriving already carrying substantial fat departed after a median of two nights, while lean arrivals stayed a median of six, and stopover length lengthened by roughly a day for every two days of northerly wind recorded at the station's mast. Faecal samples collected from a subset of 180 birds showed a shift from insect remains to almost exclusively fruit skins and seeds over the course of each season. The authors estimate that a bird leaving the site with the observed median fat load could cover the roughly 1,100 kilometers to the next major staging region without feeding, and that the fruit resource at this single hedgerow system supports several thousand migrants each autumn."
},
{
"query": "elk browsing decline after wolf reintroduction",
"passage": "Riparian transects established along a twelve-kilometer stretch of the valley floor were resurveyed annually for eighteen years following the reintroduction of a small wolf population into the surrounding highlands. Elk browsing intensity on willow shoots, scored on a standardized five-point index, declined from an average of 3.8 in the five years preceding reintroduction to 1.6 within a decade afterward, and willow stem height in browsed patches increased from a median of 34 centimeters to 210 centimeters over the same period. This recovery was spatially uneven: willow stands within 200 meters of steep-sided ravines, where elk vigilance costs were presumably highest, recovered roughly three years earlier than stands on open floodplain terraces more than a kilometer from cover, a pattern consistent with a landscape of fear effect operating independently of any measured decline in elk population density, which fell by only 12 percent over the study period. Beaver dam counts along the same transects rose from four active dams at the outset to twenty-nine by year fifteen, tracking the recovery of willow as a construction material and food source with an approximate five-year lag. Stream cross-sections adjacent to new beaver complexes narrowed and deepened relative to control reaches without dams, and summer water temperatures measured at these sites were on average 1.3 degrees Celsius cooler, a change plausibly linked to increased shading from taller riparian vegetation. Coyote scat frequency declined by roughly half following wolf establishment, and a corresponding rise in ground-nesting bird detections was recorded, though causal attribution rests on correlational survey data rather than controlled removal experiments.",
"nearMiss": "Occupancy surveys of forty-one temporary ponds scattered across a sandy coastal plain were repeated three times each spring over eleven consecutive years, with each pond checked for egg masses, dip-netted for larvae, and monitored for the date on which standing water disappeared. Pond hydroperiod, measured as the number of days between spring filling and drying, ranged from 38 to 190 days across sites and varied by as much as 60 days at the same pond between wet and dry years. Two salamander species showed sharply different responses: the larger species bred successfully only in ponds holding water for at least 110 days, and its occupancy across the plain fell from 27 ponds to 12 over the three driest years before recovering, whereas the smaller species completed metamorphosis in as little as 55 days and maintained near-constant occupancy throughout. Fish were recorded in only four ponds, all connected to a drainage ditch during floods, and no amphibian larvae were ever detected in those four. Occupancy models fitted to the eleven-year record identified hydroperiod and distance to the nearest occupied pond as the two strongest predictors of colonization, with canopy cover having no detectable effect."
},
{
"query": "How much did fetal hemoglobin increase after BCL11A base editing?",
"passage": "Eleven participants with severe sickle cell disease received a single infusion of autologous CD34-positive cells edited ex vivo with an adenine base editor targeting the erythroid-specific enhancer of BCL11A, following myeloablative conditioning with busulfan. Editing efficiency in the infused cell product averaged 74 percent of alleles across the eleven manufacturing runs, and engraftment, defined as neutrophil recovery above 500 cells per microliter, occurred at a median of twenty-two days post-infusion. Fetal hemoglobin as a proportion of total hemoglobin rose from a pretreatment baseline below 1 percent to a plateau averaging 42 percent by month six, with pancellular distribution confirmed by flow cytometry showing fetal hemoglobin present in more than 90 percent of red cells rather than concentrated in a minority subpopulation, a distinction thought relevant to preventing sickling under hypoxic stress. Vaso-occlusive crises requiring hospitalization fell from a median of 3.9 per year in the two years before treatment to zero in ten of eleven participants over a median follow-up of twenty months; the remaining participant had one crisis at month fourteen coinciding with a documented respiratory infection. Off-target editing, assessed at forty-one computationally predicted sites plus sites identified by an unbiased genome-wide assay in the manufacturing product, was not detected above the 0.1 percent limit of quantification at any tested locus. Two participants experienced prolonged but self-resolving thrombocytopenia attributed to conditioning rather than the editing procedure itself, and durability beyond twenty months remains to be established in this small cohort.",
"nearMiss": "Apheresis platelet units collected from 96 volunteer donors were split into paired bags, one stored under standard agitation at 22 degrees Celsius for the licensed five days and the other held for seven days under the same conditions, to characterize the additional deterioration a two-day extension would introduce. Daily sampling showed pH falling from a mean of 7.2 at collection to 6.9 by day five and 6.7 by day seven, with three of the 96 extended units dropping below the 6.4 threshold at which a unit is discarded. Aggregation in response to collagen fell to 71 percent of the day-one value by day five and to 58 percent by day seven, while surface exposure of the activation marker P-selectin rose steadily across the storage period, reaching a median of 44 percent of platelets by the final day. Lactate accumulation accelerated after day five in units with the highest initial cell counts, suggesting that glucose depletion rather than time alone drives the late decline. Bacterial culture of every unit at the end of storage was negative. The authors conclude that most seven-day units remain within release specifications but that the tail of poorly performing units widens, and recommend a second quality check before issue for any unit stored past day five."
},
{
"query": "FokI polymorphism",
"passage": "Postmenopausal women enrolled in a bone health registry, numbering four hundred and seventeen, were genotyped for the FokI polymorphism in the vitamin D receptor gene and stratified into ff, Ff, and FF genotype groups. Lumbar spine bone mineral density, measured by dual-energy X-ray absorptiometry, was on average 0.045 grams per square centimeter lower in ff homozygotes than in FF homozygotes after adjustment for age, body mass index, and years since menopause, a difference roughly equivalent to four years of additional bone loss at the observed rate of decline in this cohort. Radiolabeled calcium absorption, assessed using a dual-isotope oral tracer test in a subgroup of ninety-two women, was 19 percent lower in ff homozygotes than in FF homozygotes when dietary calcium intake was held constant across groups by a standardized test meal, suggesting a functional difference in intestinal calcium handling rather than a purely statistical association. Serum 25-hydroxyvitamin D concentrations did not differ significantly across genotype groups, indicating that the effect was not simply a consequence of differing vitamin D status. Over a mean follow-up of 6.2 years, incident vertebral fracture occurred in 14 percent of ff homozygotes compared with 7 percent of FF homozygotes, though this difference did not reach significance after adjustment for baseline bone density itself, implying that most of the genotype's fracture-relevant effect operates through its influence on bone mass rather than through an independent pathway. Calcium supplementation at 1200 milligrams daily over two years narrowed the density gap between ff and FF groups by about a third.",
"nearMiss": "Records from 384 men aged sixty-five and older admitted to a district hospital with a fractured hip over a three-year period were reviewed to relate the timing of first weight-bearing after surgical fixation to length of stay and to discharge destination. Surgeons at the hospital differed in their standard instructions, with roughly half permitting full weight-bearing on the first postoperative day and the remainder restricting patients to toe-touch loading for two to four weeks depending on fracture type and implant. Patients cleared for early loading left hospital a median of four days sooner and were nearly twice as likely to return directly home rather than to a rehabilitation ward, after accounting for age, pre-injury mobility and the presence of dementia. Rates of implant failure requiring revision within a year were similar in the two groups, at 4.1 and 3.6 percent. Delirium during the admission was recorded in 29 percent of the restricted group and 18 percent of the early loading group, a difference the authors attribute partly to longer periods of immobility. They recommend that departments standardize on early loading wherever the fixation permits it, and note that a physiotherapy session on the day of surgery itself was the single scheduling change most strongly associated with shorter stays."
},
{
"query": "sanitizer effectiveness against established Listeria biofilms",
"passage": "Stainless steel coupons inoculated with three Listeria monocytogenes isolates recovered from a prior facility investigation were incubated at 4, 12, and 22 degrees Celsius for up to ten days to characterize biofilm development under conditions representative of a chilled food processing environment. Viable cell counts recovered from coupon surfaces by sonication reached 10^6 colony-forming units per square centimeter by day seven at 12 degrees, a temperature commonly encountered on equipment surfaces during processing shifts, whereas counts at 4 degrees plateaued roughly one log lower over the same period. Scanning electron microscopy showed a patchy monolayer at day three progressing to a structured biofilm with visible channels and an extracellular matrix by day seven at both 12 and 22 degrees. Exposure to a quaternary ammonium sanitizer at the manufacturer-recommended working concentration reduced planktonic cultures of all three isolates by more than five logs within the specified contact time, but the same treatment applied to seven-day biofilms achieved only a 1.2 to 2.1 log reduction, and surviving cells recovered to pre-treatment density within twenty-four hours under continued cold storage. One isolate, carrying a previously described efflux-associated gene cluster, tolerated twice the sanitizer concentration in planktonic testing without any reduction in growth rate, though its biofilm was not disproportionately more resistant than the other two isolates once matrix formation was established, suggesting the matrix itself was the dominant protective factor. Alternating sanitizer classes across consecutive cleaning cycles reduced the fraction of surviving biofilm cells by an additional order of magnitude over a simulated four-week cleaning schedule.",
"nearMiss": "Fresh-cut romaine lettuce processed at a commercial facility was packaged under three modified atmospheres, air, 5 percent oxygen with 10 percent carbon dioxide, and 2 percent oxygen with 15 percent carbon dioxide, and held at 4 degrees Celsius for fourteen days to relate spoilage organism growth to the sensory changes that determine shelf life in retail practice. Lactic acid bacteria and pseudomonad counts were taken every second day, alongside a panel assessment of off-odor, leaf-edge browning and slime. In air, pseudomonads dominated and reached 8 log colony-forming units per gram by day eight, coinciding with the point at which panelists rejected the product for browning and a sour smell, while under the higher carbon dioxide atmosphere pseudomonad growth was suppressed and lactic acid bacteria took over, reaching similar counts only by day twelve. Headspace measurements showed carbon dioxide continuing to rise inside the low-oxygen packs as the tissue respired, and by day ten this had produced a fermented off-odor in a third of packs that panelists judged as objectionable as the browning it had prevented. A predictive model fitted to the count data estimated shelf life at 7, 10 and 11 days for the three atmospheres, and the authors recommend the intermediate atmosphere as the best balance between microbial suppression and physiological injury to the leaf."
},
{
"query": "Are texture change and color pattern controlled by the same circuit in octopus camouflage?",
"passage": "Electrical stimulation of discrete sites within the optic lobe of anesthetized common octopuses elicited reproducible, site-specific chromatophore patterns on the contralateral mantle, ranging from uniform darkening to sharply defined mottled patterns resembling those produced spontaneously during substrate matching. Mapping across twenty-two stimulation sites in six animals revealed a coarse somatotopic organization, with adjacent electrode positions producing patterns affecting overlapping but distinct skin regions, though the map was not strictly point-to-point and considerable overlap existed between sites more than two millimeters apart. Simultaneously recorded papillae, the small dermal projections responsible for three-dimensional texture change, activated independently of chromatophore pattern in roughly a third of trials, indicating that texture and color pattern generation are not obligatorily coupled at this level of the circuit despite typically co-occurring during natural camouflage on textured substrates. Animals placed on a checkerboard substrate with tile sizes varied across trials matched pattern granularity to tile size within a narrow range, producing disruptive patterning on tiles between one and four centimeters and more uniform stippled patterning on both smaller and larger tiles, a transition that occurred within the first two seconds of substrate exposure in most trials. Lesioning a discrete subregion of the optic lobe abolished disruptive patterning specifically while leaving uniform darkening and stippling intact, suggesting at least partial anatomical separation of the circuitry generating distinct pattern classes. Animals rendered blind in one eye by an opaque contact lens continued to produce substrate-appropriate patterning using visual input from the remaining eye alone, with no measurable difference in matching accuracy.",
"nearMiss": "High-speed video at 2,000 frames per second was used to record 214 prey-capture strikes by juvenile cuttlefish presented with live shrimp in a filming arena, in order to describe the kinematics of the two feeding tentacles from the moment of ejection to contact. The tentacles extended over a mean of 28 milliseconds, reaching peak velocities near 2.3 meters per second, with the terminal clubs rotating outward during the final third of extension so that the sucker-bearing surfaces faced the prey at contact. Strike distance was remarkably constant at about 0.6 mantle lengths regardless of shrimp size, and animals that positioned themselves outside this range invariably made a short approach swim before ejecting the tentacles rather than compensating with a longer strike. Misses, which made up 17 percent of attempts, were associated with prey that moved laterally in the 40 milliseconds before ejection, and frame-by-frame analysis showed no evidence of in-flight trajectory correction, indicating that the strike is ballistic once launched. Kinematic profiles were indistinguishable between individuals reared on live and on dead prey, suggesting that the motor program does not require experience to develop."
},
{
"query": "myofibroblast density",
"passage": "In a rat coronary ligation model, the infarct border zone showed peak myofibroblast density at day seven post-occlusion, identified by alpha-smooth muscle actin colocalization with vimentin, before declining gradually over the following three weeks as scar maturation proceeded. Collagen type I to type III ratio within the scar rose from near parity at day seven to approximately 4:1 by day twenty-eight, consistent with a shift from an early, more compliant matrix toward a stiffer, more cross-linked scar. Daily administration of an experimental antifibrotic compound targeting the transforming growth factor beta pathway, started on day three post-ligation, reduced peak myofibroblast density by 37 percent and lowered scar collagen content by 21 percent at day twenty-eight relative to vehicle, but left infarct size itself, measured by triphenyltetrazolium staining at day one, unchanged, confirming the drug acted on remodeling rather than acute injury. Left ventricular ejection fraction, tracked by echocardiography, was modestly better preserved in treated animals at day twenty-eight, 41 percent versus 34 percent in controls, though end-diastolic volume was similar between groups, suggesting the functional benefit derived more from preserved contractile geometry in viable myocardium than from a smaller absolute scar. Treatment delayed beyond day seven produced no detectable benefit on any measured endpoint, indicating a narrow therapeutic window tied to the peak of myofibroblast activity. Border zone arrhythmia inducibility, tested by programmed electrical stimulation at day twenty-eight, was lower in treated animals, an effect plausibly linked to reduced fibrotic content interrupting conduction pathways, though this mechanistic link was not directly tested through conduction mapping.",
"nearMiss": "Left ventricular tissue from mice subjected to eight weeks of transverse aortic constriction was compared with sham-operated controls to characterize the shift in titin isoform expression that accompanies pressure-overload hypertrophy. Protein gels resolved the two principal cardiac isoforms, and the proportion of the more compliant N2BA form rose from 11 percent of total titin in sham hearts to 27 percent in the constricted animals, a change that was already detectable at two weeks and plateaued by week six. Passive stiffness of isolated cardiomyocytes, measured by stretching single cells with a force transducer, fell in proportion to the isoform shift, whereas stiffness measured in intact muscle strips rose, indicating that changes in the extracellular matrix more than offset the cellular adaptation at the tissue level. Echocardiography showed wall thickness increasing by 34 percent with preserved ejection fraction but a prolonged early filling deceleration time consistent with impaired relaxation. Phosphorylation of the N2B spring element by protein kinase A was reduced by roughly half in the overloaded hearts, and in vitro treatment of skinned cells with the kinase restored passive tension toward sham values, pointing to a reversible contribution that could be targeted pharmacologically."
},
{
"query": "nitrogenase activity decline during drought stress",
"passage": "Root nodules were harvested from field-grown soybean plants at the R2 reproductive stage to examine how progressive soil drying affects symbiotic nitrogen fixation. Infection thread formation, visualized by clearing and staining a subset of root segments, proceeded normally under well-watered conditions, with rhizobial cells released into cortical cell cytoplasm and packaged into symbiosomes within approximately five days of initial root hair curling. Nitrogenase activity, measured by the acetylene reduction assay on excised nodules, declined by 68 percent within four days of withholding irrigation, well before any measurable drop in nodule leghemoglobin concentration, which fell by only 9 percent over the same interval. This sequencing suggested that the initial decline in fixation was not driven by loss of the oxygen-buffering capacity that leghemoglobin provides but rather by a rapid, reversible downregulation of nitrogenase gene expression, confirmed by a corresponding 71 percent drop in nifH transcript abundance measured by quantitative reverse transcription PCR over the same four-day window. Rewatering after six days of drought restored acetylene reduction activity to within 15 percent of pre-drought levels within forty-eight hours, considerably faster than the recovery of overall nodule fresh weight, which remained depressed for at least two weeks. Nodule oxygen permeability, estimated from the kinetics of an infiltrated oxygen-sensitive dye, decreased under drought stress, an adaptation that would be expected to further limit nitrogenase activity given the enzyme's extreme oxygen sensitivity, and this permeability change appeared to be an actively regulated response of the nodule cortex rather than a passive consequence of tissue dehydration.",
"nearMiss": "Soybean cultivars from three maturity groups were sown at four dates spaced two weeks apart across two seasons at a research farm in the central plains to quantify how planting date shifts seed composition, which sets the price a crop receives at the crushing plant. Seed protein concentration rose from a mean of 38.9 percent at the earliest sowing to 41.2 percent at the latest, while oil concentration moved in the opposite direction, from 21.4 to 19.6 percent, so that the sum of the two remained nearly constant across treatments. Cooler temperatures during the seed-filling period at the later sowings accounted for most of the oil decrease in a regression across site-years, and cultivars from the latest maturity group were the most sensitive. Yield fell by about 4 percent for each two-week delay after the first sowing date, so the protein gain came at a cost that the premium paid for high-protein grain did not fully recover under the price schedule in force during the study. Seed size and germination percentage were unaffected by sowing date. The authors conclude that growers targeting protein premiums should select cultivars with inherently high protein rather than delay planting, and provide a composition table for the twelve cultivars tested."
},
{
"query": "What causes the relapsing fever pattern in Borrelia hermsii infection?",
"passage": "A cluster investigation following six cases of relapsing fever among guests who had stayed in the same group of rustic cabins over a four-month period identified Borrelia hermsii by blood smear and PCR in five of the six patients, with the sixth diagnosed serologically after an initial negative smear during an afebrile interval. All affected cabins had rodent nesting material in wall cavities or attic spaces, and Ornithodoros soft ticks, the vector responsible for this spirochete, were recovered from four of five cabins inspected, typically from crevices near sleeping areas rather than from more exposed locations. Patients presented with a characteristic pattern of recurring fever, with an initial febrile episode lasting a median of three days followed by an afebrile interval of about seven days before a second, generally milder, episode, a relapsing course that had led two patients to be initially misdiagnosed with a viral syndrome before recurrence prompted further testing. Peak spirochetemia during febrile episodes, estimated by dark-field microscopy cell counts, coincided with maximal antigenic variation detectable by variant-specific serology, consistent with the established model in which each relapse is driven by outgrowth of a spirochete subpopulation expressing a novel surface protein variant evading the prior antibody response. All patients responded to a ten-day course of oral doxycycline with resolution of fever within forty-eight hours, though one patient developed a transient Jarisch-Herxheimer-like reaction, with rigors and a brief blood pressure drop, within two hours of the first dose, a recognized risk of rapid spirochete lysis that resolved without intervention. Trapping efforts removed an estimated 80 percent of rodents from the affected structures, and no further cases were reported over the subsequent two years.",
"nearMiss": "Bloodstream isolates of Candida species collected from 41 hospitals over a five-year surveillance period, numbering 3,218, were speciated by mass spectrometry and tested for susceptibility to fluconazole, the echinocandins and amphotericin B by broth microdilution. Candida albicans accounted for 46 percent of isolates at the start of the period but only 37 percent by its end, with the decline matched by a rise in Candida glabrata and, in the final two years, the first appearances of Candida auris at three hospitals in one metropolitan area. Fluconazole resistance among glabrata isolates increased from 9 to 15 percent over the period, and echinocandin resistance, rare at the outset, was detected in 3.8 percent of glabrata isolates by the final year, concentrated in patients with prior echinocandin exposure. Hospitals with the highest antifungal consumption per thousand patient-days did not show higher resistance rates, whereas prior exposure at the individual patient level was strongly predictive. Sequencing of resistant isolates identified hot-spot mutations in the target gene in most cases. The report recommends species-level identification for every bloodstream isolate and routine susceptibility testing for glabrata, and notes that the shift in species distribution alone would reduce the expected success of empirical fluconazole from about 85 to about 75 percent."
},
{
"query": "AS160 phosphorylation",
"passage": "Rodents maintained on a sixty percent fat diet for eight weeks developed fasting hyperinsulinemia and, when tested by hyperinsulinemic-euglycemic clamp, required a glucose infusion rate 44 percent lower than chow-fed controls to maintain euglycemia, confirming whole-body insulin resistance. In isolated soleus muscle from these animals, insulin-stimulated GLUT4 translocation to the plasma membrane, quantified by subcellular fractionation and immunoblotting, was reduced by roughly half relative to controls despite total cellular GLUT4 content being unchanged, indicating a defect in trafficking rather than transporter abundance. Phosphorylation of AS160 at the Akt substrate site, a step normally required to relieve inhibition of GLUT4 vesicle exocytosis, was blunted in high-fat-fed muscle following an insulin bolus, and this blunting correlated with elevated diacylglycerol content and increased membrane-associated protein kinase C theta activity, consistent with a lipid-mediated interruption of proximal insulin signaling rather than a defect downstream at the vesicle fusion machinery itself. A four-week treadmill exercise intervention begun after the eight-week diet period restored insulin-stimulated GLUT4 translocation to within 12 percent of chow-fed control values without any change in body weight or fat mass, dissociating the metabolic benefit from weight loss per se. This restoration was accompanied by normalized AS160 phosphorylation and a partial, though incomplete, reduction in intramuscular diacylglycerol content, suggesting exercise acts at least partly through the same lipid-signaling node disrupted by chronic fat feeding. Muscle glycogen content, depleted by exercise, recovered to baseline within twenty-four hours on a standard chow refeed.",
"nearMiss": "Fourteen healthy adults attended three morning sessions in random order, consuming a 400 kilocalorie mixed meal alone or preceded thirty minutes earlier by a small drink containing 15 or 30 grams of whey protein, with blood sampled every fifteen minutes for three hours to follow gut hormone release and glucose handling. The 30 gram preload roughly doubled the plasma concentration of glucagon-like peptide 1 measured at the time the meal was eaten and blunted the subsequent glucose peak by 1.4 millimoles per liter relative to the meal alone, while the 15 gram preload produced intermediate effects. Gastric emptying, assessed by the appearance of a paracetamol tracer added to the meal, was slowed by about 20 minutes after the larger preload, and participants reported lower hunger ratings in the third hour. Insulin secretion was higher after the preloads in the first thirty minutes following the meal but total three-hour insulin output did not differ, indicating that the earlier timing rather than the amount of insulin accounted for the flatter glucose curve. The authors propose that a modest protein preload is a practical way to reduce postprandial glucose excursions without altering meal composition, and are testing it in people with impaired glucose tolerance."
},
{
"query": "heritability of hip joint laxity in dogs",
"passage": "A working dog breeding program spanning eleven generations used the distraction index, a radiographic measure of passive hip joint laxity obtained under standardized sedation and positioning, to screen breeding candidates before each mating decision. Distraction index values across the founding population ranged widely, from 0.28 to 0.81, with a heritability estimate of 0.35 derived from an animal model incorporating the full pedigree, indicating substantial but incomplete genetic control over this trait alongside a sizable environmental or developmental contribution. Selecting breeding pairs whose average distraction index fell in the lowest quartile of each generation produced a cumulative reduction in mean population distraction index from 0.52 in the founding generation to 0.34 by generation eleven, a pace of improvement roughly consistent with predictions from the estimated heritability under the applied selection intensity. Radiographic hip scores assigned separately by a masked evaluator using a standard categorical grading scheme correlated with distraction index at r = 0.61, a moderate but imperfect relationship that led the program to retain both measures rather than substituting one for the other. Clinical osteoarthritis, assessed at five years of age by a combination of gait scoring and radiographic osteophyte grading, was present in 9 percent of dogs from the final three breeding generations compared with an estimated 38 percent in the founding generation based on historical records, though differences in diagnostic criteria over the eleven-year span make this comparison only approximate. A small number of sires with low individual distraction index nonetheless produced offspring with above-average values, illustrating that individual phenotype alone was an imperfect predictor of breeding value.",
"nearMiss": "Foot injuries recorded by veterinarians attending a long-distance sled race over six consecutive years were matched to trail surface and weather logs kept by race officials, covering 4,860 dog starts and 1,207 treated pad injuries. Abrasions and split pads made up 71 percent of injuries and were concentrated on stretches of the trail where hard, crusted snow alternated with exposed gravel, while cracked pads clustered in years with prolonged cold below minus 30 degrees Celsius and low humidity. Teams whose handlers reported using protective booties on every leg had an injury rate of 14 per hundred starts compared with 31 for teams using booties only on the sections mushers judged hazardous, and the difference persisted after adjusting for team speed and for the dogs' age. Wrist and shoulder soft-tissue injuries, by contrast, showed no association with surface or bootie use and rose with team speed on downhill sections. A trial of a pre-race wax applied to the pads of 60 dogs on four teams reduced abrasions on the gravel sections by about a third relative to untreated teammates. The authors recommend mandatory booties on the sections identified and suggest that race routing away from the two worst gravel stretches would remove a large fraction of treated injuries."
},
{
"query": "How does oxygen tension affect telomere shortening in cultured fibroblasts?",
"passage": "Primary human dermal fibroblasts, serially passaged until proliferation ceased, underwent a median of 52 population doublings before entering a stable growth-arrested state characterized by flattened, enlarged morphology and positive staining for senescence-associated beta-galactosidase in more than 80 percent of cells. Mean terminal restriction fragment length, measured by Southern blot as a proxy for average telomere length, declined from approximately 9.4 kilobases at early passage to 5.8 kilobases at the point of growth arrest, a loss of roughly 70 base pairs per population doubling that was consistent across three separate donor cell lines despite differing overall replicative lifespans. Expression of p16INK4a, quantified by quantitative PCR, remained low and stable through the majority of the culture's lifespan before rising sharply over the final five to eight doublings preceding arrest, a pattern distinct from p21, which rose more gradually and earlier, suggesting the two cyclin-dependent kinase inhibitors are engaged at different stages of the approach to senescence. Ectopic expression of the catalytic subunit of telomerase, introduced by retroviral transduction at population doubling 30, stabilized telomere length and permitted continued proliferation well beyond the doubling count at which matched untransduced cultures arrested, with transduced cultures still dividing at doubling 110 when the experiment was terminated. Cultures maintained at atmospheric, rather than physiological, oxygen tension senesced after significantly fewer doublings, around 34 on average, and showed faster telomere attrition per doubling, indicating that oxidative stress accelerates telomere-dependent senescence. Reactive oxygen species scavenger supplementation partially, but not fully, restored replicative lifespan under atmospheric oxygen conditions.",
"nearMiss": "Keratinocytes isolated from the epidermis of surgical skin discards from six donors were expanded in a low-calcium medium and then switched to media containing 0.1, 0.5 or 1.2 millimolar calcium for six days to characterize the dose dependence of the differentiation program in a defined system. At the lowest calcium concentration the cells remained a proliferating monolayer with a doubling time near 28 hours, whereas at 1.2 millimolar they stratified into three to four layers within four days and expressed involucrin and loricrin, the late differentiation markers, in more than 70 percent of cells in the upper layers. Desmosome formation, scored by electron microscopy, was detectable within two hours of the calcium switch and complete by twelve hours, well before any change in marker expression. Colony-forming efficiency of cells replated from the 1.2 millimolar flasks fell to under 2 percent, compared with 18 percent from the low-calcium flasks, indicating that the stem-like fraction is lost rapidly once stratification begins. Blocking the calcium-sensing receptor with a selective antagonist prevented stratification and marker expression at every concentration tested, and the authors propose the antagonist as a tool for maintaining undifferentiated keratinocyte stocks."
},
{
"query": "CONSTANS protein stability",
"passage": "Wild-type Arabidopsis plants grown under long-day conditions of sixteen hours light and eight hours dark flowered at a median of 21 days after germination, whereas plants grown under short days of eight hours light flowered at a median of 58 days, establishing the baseline photoperiodic response used to interpret subsequent mutant comparisons. CONSTANS transcript abundance, measured by quantitative PCR at two-hour intervals across a full daily cycle, peaked during the light period only under long-day conditions, coinciding with a window in which the CONSTANS protein, otherwise targeted for degradation in darkness, remained stable long enough to activate downstream transcription. Plants carrying a loss-of-function constans allele failed to show accelerated flowering under long days, flowering instead at a rate statistically indistinguishable from short-day-grown wild-type plants, confirming that this stability window rather than transcript abundance alone is the operative signal. FT transcript, induced directly downstream of functional CONSTANS protein in leaf phloem companion cells, was undetectable in the constans mutant background under any photoperiod tested. Grafting experiments joining a wild-type rootstock capable of long-day FT induction to a shoot carrying a nonfunctional FT allele restored early flowering in the mutant shoot, and FT protein was subsequently detected in shoot apex tissue by tagged-protein imaging, supporting long-distance movement of FT protein from leaf to shoot apex as the mechanism linking peripheral day-length perception to the central flowering decision. A gain-of-function allele producing constitutively stable CONSTANS protein induced early flowering even under short days, though not as early as long-day wild-type plants, indicating additional long-day-specific inputs contribute to the full photoperiodic response.",
"nearMiss": "Freshly harvested Arabidopsis seeds from the Cvi accession, which shows deep primary dormancy, were stored dry at 20 degrees Celsius and sampled every two weeks over sixteen weeks to follow how after-ripening releases dormancy and how the hormone balance in the seed changes over that period. Germination on water at harvest was below 5 percent and rose to 92 percent by week fourteen, with the steepest increase between weeks six and ten. Abscisic acid content in the dry seed declined only modestly over storage, but the amount produced on imbibition fell sharply, from a peak of 480 picograms per seed in freshly harvested lots to 90 in fully after-ripened ones, while imbibition-induced gibberellin synthesis rose over the same interval. Seeds stored at 4 degrees Celsius after-ripened far more slowly, reaching only 30 percent germination at sixteen weeks, and seeds held at 30 degrees completed the process in under six weeks but showed reduced viability thereafter. Applying a gibberellin biosynthesis inhibitor to after-ripened seeds restored dormancy-like behavior, and exogenous gibberellin broke dormancy in fresh seed, placing the hormonal switch downstream of whatever change storage produces in the dry seed."
},
{
"query": "keyhole porosity",
"passage": "The test coupons were built from a nitrogen-atomized Inconel 718 powder with a D50 particle size of 34 micrometers, processed on a 400 W fiber laser system at scan speeds ranging from 700 to 1400 mm/s and hatch spacings of 90 to 130 micrometers. At the lowest volumetric energy densities, below roughly 45 J/mm3, X-ray computed tomography of thirty vertically oriented cylinders revealed irregular, lack-of-fusion voids concentrated along layer boundaries, with an average porosity fraction of 1.8 percent. Raising the energy density above 90 J/mm3 shifted the dominant defect population toward small, spherical cavities consistent with vapor-cavity collapse during unstable melt pool behavior; these pores averaged 40 micrometers in diameter and were distributed more uniformly through the bulk rather than clustered at boundaries. A high-speed coaxial camera operating at 20,000 frames per second captured intermittent plume ejection and melt pool depression events coinciding with the onset of this second defect mode, consistent with a transition into keyhole-mode melting once the local energy density exceeded a threshold near 80 J/mm3 for this powder and layer thickness combination. Tensile specimens machined from the two defect regimes showed a 22 percent reduction in elongation at failure for the keyhole-affected condition despite only a modest change in ultimate tensile strength, suggesting that pore morphology and connectivity matter more than total void fraction for ductility. A processing window bounded by these two thresholds, roughly 55 to 75 J/mm3, produced parts with porosity below 0.2 percent in all thirty specimens. The result should be treated cautiously outside this specific powder lot and layer thickness of 40 micrometers, since prior recoater blade wear and powder oxygen content are known to shift both boundaries independently.",
"nearMiss": "Gas-atomized titanium alloy powder was tracked through twelve reuse cycles in a laser powder bed system, with samples taken after each build and sieving step to follow how repeated exposure to the build chamber alters the feedstock. Oxygen content measured by inert gas fusion rose from 0.09 to 0.16 weight percent over the twelve cycles, most of the increase occurring in the first four, and nitrogen rose more slowly from 0.01 to 0.03 percent. Particle size distribution narrowed slightly as fines were lost to sieving and to the filter, with the tenth-percentile diameter shifting from 19 to 24 micrometers, and flowability measured by Hall funnel improved from 31 to 27 seconds per 50 grams, an effect the authors attribute both to the loss of fines and to the smoothing of satellite particles. Scanning electron microscopy showed a growing fraction of irregular spatter particles with oxide-rich surfaces, reaching about 3 percent by number at cycle twelve. Tensile specimens built from cycle-twelve powder met the alloy's minimum strength requirement but showed elongation at failure reduced from 12 to 9 percent, which the authors link to the interstitial pickup, and they recommend blending in at least 30 percent virgin powder every fourth cycle to hold oxygen below 0.13 percent."
},
{
"query": "carbonation front depth measurement technique",
"passage": "Mortar prisms with water-to-cement ratios of 0.45, 0.55, and 0.65 were cured for 28 days under standard moist conditions before being transferred to a chamber holding a constant 4 percent CO2 concentration, 65 percent relative humidity, and 20 degrees Celsius, conditions chosen to accelerate carbonation relative to ambient exposure by a factor estimated at roughly fifteen. Carbonation depth was tracked at 7, 14, 28, 56, and 90 days by splitting companion specimens and spraying freshly fractured surfaces with a 1 percent phenolphthalein solution in ethanol, then measuring the uncolored zone at eight points per face with digital calipers. Depth progressed approximately with the square root of exposure time for all three mixes, but the proportionality constant varied from 3.1 mm per root-week for the 0.45 mix to 6.4 mm per root-week for the 0.65 mix, a difference attributed to the higher capillary porosity and lower calcium hydroxide reserve in the higher water-to-cement ratio paste. Thermogravimetric analysis of powder samples taken at 2 mm intervals from the exposed face confirmed that portlandite content dropped to near zero within the phenolphthalein-defined carbonated zone, while calcite content rose correspondingly, supporting the indicator method as a reasonable proxy for the reaction front in this mix design. A parallel set of prisms containing 30 percent fly ash replacement carbonated roughly 40 percent faster than the plain cement mix at equal water-to-cement ratio, consistent with the lower portlandite buffering capacity of blended cements reported in an earlier survey of supplementary cementitious materials. These accelerated-condition constants should not be extrapolated directly to natural exposure without a diffusion-based correction, since the elevated CO2 concentration alters the moisture profile within the pore network relative to field conditions.",
"nearMiss": "A ready-mix plant supplying slab pours in a hot inland climate ran a season-long trial to find the retarding admixture dose that keeps its standard structural mix workable through deliveries of up to ninety minutes, since drivers had been adding water at the site to recover slump, with predictable damage to strength. Trucks carrying the plain mix and the same mix dosed with a lignosulfonate retarder at three levels were held under agitation in the yard with concrete temperatures between 32 and 35 degrees Celsius, and slump was taken every fifteen minutes for two hours. The undosed mix lost half of its initial 180 millimeter slump within 45 minutes, while the middle retarder dose held slump above 120 millimeters for 90 minutes and the highest dose for the full two hours at the cost of a setting delay of nearly three hours by penetration resistance, which the plant judged unacceptable for the late-afternoon pours. Cylinders cast from each truck at the end of its holding period showed no strength penalty at 28 days for any dosed mix, whereas cylinders from trucks that had been re-tempered with water at the site, sampled in parallel from real deliveries, averaged 9 percent below the plant's design strength. The plant adopted the middle dose as standard for any delivery scheduled beyond forty minutes and issued drivers with written instructions forbidding site water additions."
},
{
"query": "How does multipath reflection affect GPS positioning accuracy in urban canyons?",
"passage": "Signal traces were logged from a dual-frequency receiver mounted on a survey cart pushed along a fixed 400 meter route between buildings averaging eleven stories, with a static rooftop reference station providing carrier-phase corrections for a real-time kinematic baseline. Under open-sky conditions along a control segment, horizontal positioning error stayed below 3 centimeters for 95 percent of epochs, but within the canyon segment the code-minus-carrier statistic showed multipath-induced noise spikes exceeding 1.5 meters on the L1 C/A signal whenever a satellite's elevation angle dropped below 25 degrees on the side facing the taller building row. Reflected signals arriving with path-length excess greater than the receiver's correlator spacing produced a distinctive bimodal error distribution rather than the single Gaussian lobe seen in open sky, with roughly 18 percent of epochs falling in a secondary error mode centered near 0.9 meters. Excluding satellites below a 30 degree elevation mask reduced the fix availability from continuous to about 71 percent of epochs but cut the 95th percentile horizontal error to 0.4 meters, illustrating the standard tradeoff between availability and accuracy in this geometry. A ray-tracing model built from a simplified extruded building footprint reproduced the timing and rough magnitude of the worst multipath excursions to within about 30 percent, though it consistently underpredicted diffraction-related errors near building corners, where measured errors reached nearly double the modeled values. Combining the code-based solution with a tightly coupled inertial measurement unit bridged most of the gaps during complete signal blockage, though heading drift after blockages longer than 12 seconds occasionally required a subsequent open-sky segment to fully reconverge.",
"nearMiss": "An airborne laser scanning campaign over a 640 square kilometer forested catchment produced 38 overlapping flight strips, and the residual misalignment between strips was quantified before adjustment by comparing elevations on 212 planar surfaces such as road pavements and flat roofs falling in the overlap zones. Strip-to-strip elevation differences had a standard deviation of 9 centimeters and a maximum of 31, with a systematic pattern that grew toward the outer edges of each strip, consistent with a small residual roll error in the scanner's mounting. A least-squares strip adjustment estimating three angular boresight corrections and a per-strip vertical offset reduced the standard deviation on the check surfaces to 3 centimeters, and the roll correction converged to 0.011 degrees, matching the value obtained from a dedicated calibration flight over the airport a week earlier. Terrain models generated before and after adjustment differed by more than 15 centimeters over about 8 percent of the catchment, mostly on steep slopes at strip edges, where the uncorrected model produced artificial steps that a hydrological flow-routing algorithm interpreted as small dams. Canopy height estimates were less sensitive, changing by more than half a meter over only 1 percent of the area."
},
{
"query": "adjacent channel interference",
"passage": "Two small cell units operating in the n77 band were installed on lamp posts 45 meters apart along a pedestrian corridor, one configured with a 20 MHz channel centered at 3550 MHz and the other at 3570 MHz, separated by only a 10 MHz guard region rather than the 20 MHz typically recommended for this equipment class. Spectrum analyzer sweeps at the boundary between the two cells' coverage areas showed the unwanted emissions from the first unit's power amplifier extending roughly 8 dB above the expected spectral mask at the edge of the second unit's channel, consistent with intermodulation products generated under full-load traffic rather than idle conditions. Throughput measurements taken with a commercial test phone dropped from an unloaded 780 Mbps to 210 Mbps in the overlap region once both cells were simultaneously serving four active users each, and the physical downlink control channel error rate rose above 4 percent, well past the 1 percent threshold typically tolerated before retransmission overhead becomes noticeable to users. Increasing the guard band to the full 20 MHz by retuning the second unit to 3580 MHz recovered throughput to 690 Mbps under the same loading, though at the cost of losing one 5 MHz slice of otherwise usable spectrum for the operator. A software update that enabled dynamic power back-off on the amplifier when adjacent-channel occupancy was detected achieved a similar 640 Mbps recovery without sacrificing spectrum, at the expense of roughly 6 percent lower peak cell throughput during isolated single-cell operation. Interference of this kind is easy to miss in single-site lab testing, since it only appears once two units share overlapping coverage under simultaneous heavy load, a condition the isolated bench characterization in the equipment's own datasheet did not reproduce.",
"nearMiss": "Handover parameters on a university campus network of 46 cells were tuned over a semester to reduce the ping-pong handovers observed when users walked along corridors where two cells overlapped with nearly equal signal strength. Baseline traces from 1,200 test walks showed that 23 percent of handovers were followed by a reverse handover within five seconds, and that these events accounted for most of the brief throughput dips users reported in the library building. Raising the hysteresis margin from 2 to 4 decibels cut the ping-pong rate to 9 percent but increased the number of dropped sessions on the stairwells, where signal from the serving cell fell too quickly for the delayed handover to complete, while extending the time-to-trigger from 160 to 320 milliseconds produced a smaller improvement without the stairwell penalty. The combination finally adopted, a 3 decibel margin with a 256 millisecond trigger and cell-specific offsets on the four stairwell cells, brought ping-pong down to 6 percent with no increase in drops. Mean session throughput along the corridor route rose by 11 percent, almost entirely from the elimination of the repeated handover interruptions rather than from any change in radio conditions."
},
{
"query": "resonance based fatigue test rig design",
"passage": "The 62 meter blade was mounted horizontally on a rigid steel test stand and excited near its first flapwise natural frequency of 0.62 Hz using a pair of eccentric mass actuators bolted at the 70 percent span station, a resonance-based approach chosen to reach the target 25 million cycle count within an eleven month test window rather than the several years a forced-displacement rig would require at equivalent load levels. Strain gauges bonded at twelve spanwise stations on both the pressure and suction sides tracked root bending moment throughout the test, with the actuator amplitude adjusted every few hours by a closed-loop controller to hold the root moment range within 2 percent of the target derived from the blade's design load envelope. A trailing edge adhesive disbond initiated near the 35 meter station after approximately 9.4 million cycles, first detected through a subtle but consistent shift in the local strain gauge phase relative to the root sensor rather than through visual inspection, which did not reveal any external sign of the disbond until nearly 1.2 million cycles later. Thermographic imaging performed at scheduled inspection intervals confirmed the disbond had grown to roughly 40 centimeters in length by the time it became visually apparent, growing at an accelerating rate consistent with a fracture-mechanics-based prediction using measured local strain energy release rates. The test was halted at 14.6 million cycles once disbond growth reached a length considered unsafe for continued resonant operation, short of the 25 million cycle target, prompting a design revision to the adhesive bond line thickness in that region for the next blade iteration. Because resonance testing constrains the applied load spectrum to a narrow frequency band, the rig cannot fully replicate the broader load spectrum a blade experiences in variable wind conditions, a limitation acknowledged when translating these results into a revised design life estimate.",
"nearMiss": "Icing on the blades of 22 turbines in an upland wind farm was studied over two winters by comparing each turbine's ten-minute power output with the value expected from its own summer power curve at the same nacelle wind speed, treating a shortfall of more than 15 percent lasting at least an hour at temperatures below 2 degrees Celsius as an icing event. The method identified 61 events over the two winters, with durations from three hours to four days, and cross-checking against the six turbines fitted with a commercial ice sensor showed agreement on onset within one hour in 48 of the 52 events those turbines experienced. Production lost to icing amounted to 3.8 percent of winter output in the first season and 6.1 percent in the second, which had more freezing-fog days, and losses were concentrated on the seven turbines on the exposed western ridge. Events detected by the power-curve method tended to persist for several hours after the sensor reported the blades clear, which the operator attributes to residual ice on the outer third of the blade where the sensor is not located. The operator now uses the power-curve indicator to schedule inspections and to decide when to run the blade heating on the four turbines that have it."
},
{
"query": "What causes phosphorus poisoning of three-way catalytic converters?",
"passage": "Engine dynamometer testing accumulated 80,000 kilometers of equivalent mileage on matched catalyst bricks using two lubricant formulations differing primarily in zinc dialkyldithiophosphate content, 0.11 percent phosphorus by mass in the standard oil and 0.04 percent in a low-ash variant. Post-test cross-sections of the standard-oil catalysts, examined by scanning electron microscopy with energy-dispersive X-ray spectroscopy, showed a phosphorus-rich glassy layer coating the washcoat surface at the inlet face, with phosphorus concentration falling from roughly 6 weight percent at the first 5 millimeters of brick length to under 1 percent beyond 30 millimeters, indicating that the inlet region bears the majority of the poisoning burden as engine oil combustion byproducts pass through the exhaust stream. Conversion efficiency for carbon monoxide measured on a subsequent light-off test dropped from an as-new 98 percent to 84 percent at the standard operating temperature of 450 degrees Celsius for the high-phosphorus bricks, while the low-ash oil bricks retained 95 percent conversion efficiency after the same mileage. Nitrogen oxide conversion proved more sensitive to the poisoning than carbon monoxide, falling to 71 percent in the high-phosphorus case, attributed to preferential blocking of the rhodium sites responsible for the reduction reaction by the phosphate glass layer. Light-off temperature, defined as the temperature at which 50 percent conversion is reached, shifted upward by 38 degrees Celsius for the poisoned bricks relative to fresh catalyst, a change with direct consequences for cold-start emissions given that a large share of a typical drive cycle's total tailpipe emissions occur before the catalyst reaches steady operating temperature. These findings reinforce oil formulation as a meaningful lever for catalyst durability independent of any change to the catalyst washcoat chemistry itself.",
"nearMiss": "Non-exhaust particulate emissions from disc brakes were measured on a brake dynamometer enclosed in a filtered airflow chamber, using four commercial pad formulations run against grey cast iron discs over a 300-stop cycle representative of urban driving, with particle number and mass sampled from the chamber outlet by a condensation particle counter and gravimetric filters. Mass emissions ranged from 4.1 to 9.7 milligrams per kilometer per brake across the four formulations, with the low-metallic pads at the top of the range and the ceramic-fiber pads at the bottom, and more than 70 percent of the emitted mass in every case fell in the size fraction above 2.5 micrometers. Particle number emissions were dominated by the sub-100 nanometer fraction and rose sharply, by roughly an order of magnitude, whenever disc temperature exceeded 170 degrees Celsius during the repeated hard stops at the end of the cycle, a threshold the authors relate to the onset of resin decomposition in the pad matrix. Wear of the disc contributed about a third of the emitted mass by elemental analysis of iron in the filter deposits. A prototype shroud with a suction port near the caliper captured 62 percent of the emitted mass in a follow-up run, and the authors suggest that a regulatory test for brake emissions would need to fix disc temperature limits to be reproducible across laboratories."
},
{
"query": "finger table",
"passage": "Each node in the 4096-node overlay maintained a finger table of up to 160 entries spanning the full identifier space, with the ith entry pointing to the first live node at or after a distance of 2^i from the local identifier, following the reference implementation's original routing scheme. Lookup latency was measured across 50,000 randomly generated key queries injected into a simulated network with churn rates ranging from one node join or departure per second up to twenty per second, representative of a moderately volatile peer population. At the lowest churn rate, median lookup latency held steady at 3.2 routing hops, close to the theoretical expectation of half the base-2 logarithm of network size, but at twenty churn events per second median hop count rose to 5.8 as stale finger table entries increasingly pointed to departed nodes, forcing fallback to successor-list traversal for a growing fraction of lookups. Periodic finger table stabilization every 30 seconds reduced staleness but consumed bandwidth roughly proportional to network size; halving the stabilization interval to 15 seconds cut median hop count under high churn to 4.1 at the cost of doubling background maintenance traffic per node. A modified scheme that triggered finger table repair reactively upon detecting a failed hop during an actual lookup, rather than relying solely on periodic stabilization, recovered median hop count to 3.9 under the same high-churn condition while adding only about 20 percent more maintenance traffic than the slow periodic baseline, suggesting reactive repair captures most of the benefit of aggressive periodic stabilization at a fraction of the bandwidth cost. These results depend on the assumed uniform random churn model, and the reactive scheme's advantage narrowed considerably when churn was instead concentrated in bursts correlated across many nodes simultaneously, a pattern more representative of real deployments experiencing shared network outages.",
"nearMiss": "A five-node consensus cluster running a Raft implementation was subjected to 400 injected network partitions of varying duration and shape on a testbed with configurable link latency, to measure how election timeout settings trade recovery speed against spurious leader changes. With the timeout drawn uniformly from 150 to 300 milliseconds, the cluster elected a new leader within a median of 410 milliseconds after a partition isolated the old one, but also suffered 37 unnecessary elections during the test in which a slow but healthy leader was deposed by a follower whose timeout fired first. Widening the timeout range to 300 to 600 milliseconds eliminated all but four of the spurious elections at the cost of doubling median recovery time, and introducing a pre-vote phase, in which a candidate first confirms that a majority also considers the leader dead, removed the spurious elections entirely while keeping recovery near the original figure. Asymmetric partitions, in which the leader could send but not receive, produced the longest outages under every setting, since the leader continued to issue heartbeats that prevented followers from timing out. Client-visible write latency during stable operation was unaffected by any of the timeout settings."
},
{
"query": "adaptive mesh refinement stress concentration",
"passage": "The bracket geometry, a right-angle steel fitting with a 6 millimeter fillet at the load-bearing corner, was first analyzed on a uniform tetrahedral mesh with an average element edge length of 2 millimeters, yielding a peak von Mises stress of 210 MPa at the fillet under the specified 4 kN transverse load. An h-adaptive refinement loop was then applied, using an element-wise error estimator based on the recovered stress gradient discontinuity between adjacent elements, subdividing any element whose estimated error exceeded 5 percent of the current maximum stress in the model. After three refinement passes, local element size at the fillet had dropped to approximately 0.15 millimeters while the bulk of the bracket remained at the original coarse resolution, and the peak stress prediction rose to 268 MPa, a 28 percent increase from the initial coarse-mesh estimate that had not yet converged. A fourth refinement pass changed the peak value by less than 1.5 percent, taken as evidence of practical convergence for this feature. Without adaptive refinement, reaching comparable local resolution everywhere in the model would have required roughly 40 times more elements than the adaptively refined mesh's final count of 310,000, illustrating the computational advantage of concentrating resolution where the error estimator indicates it is needed rather than refining uniformly. A separate check using a submodeling approach, in which a small cutout around the fillet was re-analyzed at very fine resolution with boundary conditions imported from the coarse global solution, produced a peak stress within 2 percent of the adaptive result, cross-validating the adaptive scheme's converged answer. The error estimator did underperform near a secondary stress concentration at a bolt hole 40 millimeters from the fillet, where its refinement decisions lagged the true error by one additional pass, a known weakness of gradient-recovery estimators when two concentration features lie within each other's influence zone.",
"nearMiss": "A welded steel frame of the kind used to support pump skids was instrumented with 24 accelerometers and excited with an impact hammer at three points to identify its first eight vibration modes, which were then compared with a finite element beam model to establish how well the model's joint stiffness assumptions represent the fabricated structure. Measured natural frequencies ranged from 14.2 to 96.5 hertz, and the model built with rigid joints overpredicted the first three by between 9 and 14 percent while matching the mode shapes closely, with modal assurance criterion values above 0.9. Replacing the rigid joints with rotational springs and fitting a single stiffness value to the first three measured frequencies brought all eight within 4 percent, and the fitted stiffness corresponded to about 60 percent of the value expected for a fully welded connection, which inspection attributed to partial-penetration welds on the inner faces of the columns. Damping ratios extracted from the measurements were between 0.8 and 1.6 percent, higher than the 0.5 percent the designers had assumed, and the updated model predicted a resonant response under the pump's operating speed about a third lower than the original design calculation."
},
{
"query": "How does reservoir permeability affect long-term heat extraction rates in enhanced geothermal systems?",
"passage": "A doublet configuration consisting of one injection and one production well, spaced 500 meters apart within a fractured granite formation at 4.2 kilometers depth and an initial rock temperature of 195 degrees Celsius, was modeled using a coupled thermal-hydraulic-mechanical simulator calibrated against a hydraulic stimulation test that had raised the effective fracture network permeability from an initial 0.5 millidarcy to approximately 8 millidarcies. At this stimulated permeability, sustained circulation of 80 liters per second maintained a production temperature above 160 degrees Celsius for a simulated 25 years before thermal breakthrough began driving output temperature down more steeply, consistent with cold injected fluid reaching the production well along the shortest, most permeable fracture pathways. Reducing the assumed permeability to 3 millidarcies in a sensitivity run, representing incomplete stimulation, cut the sustainable flow rate to 45 liters per second for the same thermal drawdown criterion, and thermal breakthrough arrived roughly six years earlier despite the lower flow rate, because flow became concentrated into fewer dominant fracture channels rather than spreading across the network. Conversely, a higher permeability case of 15 millidarcies allowed 110 liters per second but shortened the time to breakthrough to only 14 years, illustrating a tradeoff in which higher permeability improves near-term output at the cost of faster long-term thermal decline. Electrical power output estimates using a representative binary cycle conversion efficiency of 12 percent put the stimulated 8 millidarcy case at approximately 1.8 MWe sustained over the 25 year window, a figure sensitive enough to the assumed fracture network geometry that the modelers flagged tracer-test-derived flow partitioning, rather than permeability alone, as the more reliable predictor of breakthrough timing in this formation.",
"nearMiss": "Heat losses along the 11 kilometer buried transmission line of a district heating scheme supplied from two deep hot-water wells were measured over one heating season by comparing supply and return temperatures at six chambers along the route with the flow rates logged at the pumping station, to check the design assumption of 2 percent loss and to locate sections where the pre-insulated pipe had deteriorated. Total loss across the season was 3.4 percent of delivered heat, with more than half occurring in a 1.8 kilometer section laid in 1994 where the polyurethane insulation had absorbed groundwater after its outer casing was damaged during later cable works. Thermal imaging of the ground surface along that section in February showed a continuous warm strip up to 2 degrees above the surroundings, and excavation at two points confirmed waterlogged insulation with a measured conductivity three times the specification. The remaining sections performed within 0.3 percentage points of design. Replacing the damaged section was estimated to pay back in about seven years at current heat prices, and the operator has also added moisture-sensing wires to the replacement pipe so that future casing damage is detected before the insulation saturates."
},
{
"query": "spill cost heuristic",
"passage": "The allocator under evaluation builds an interference graph from live ranges computed over SSA-form intermediate code, then applies Chaitin-style simplification, removing nodes with degree below the available register count k before resorting to a spill cost heuristic when no such node remains. The heuristic ranked candidate spill nodes by the ratio of estimated dynamic use-and-definition frequency, weighted by an estimated loop nesting depth multiplier of 10 per level, to the node's interference graph degree, spilling the node with the lowest ratio first on the theory that it is used rarely relative to how much it constrains the rest of the graph. Across a suite of 42 benchmark programs compiled for a 16 general-purpose register target, this heuristic produced 9 percent fewer spill loads and stores than a simpler heuristic that ignored loop depth and ranked purely by static use count, with the improvement concentrated almost entirely in benchmarks containing nested loops three levels deep or more. A small number of benchmarks, mostly straight-line numerical kernels with little control flow, showed no measurable difference between the two heuristics since their interference graphs rarely forced any spilling regardless of the ranking function used. Compile time overhead from computing the loop-depth-weighted frequency estimate added approximately 3 percent to total allocation time, a cost the authors considered acceptable given the runtime benefit. One benchmark exhibited a 4 percent runtime regression under the improved heuristic, traced to a case where a variable with high loop-weighted priority was retained in a register at the cost of spilling several lower-priority but still frequently accessed variables in an unrelated basic block outside the loop, a pathology suggesting the heuristic's frequency estimate would benefit from being scoped more locally rather than applied uniformly across a variable's entire live range.",
"nearMiss": "The automatic vectorizer of a production compiler was evaluated on 1,400 hot loops extracted from a numerical benchmark suite to determine how often it succeeds and, where it fails, which analysis limitation is responsible. The vectorizer transformed 61 percent of the loops at the default optimization level, and profiling attributed 83 percent of the suite's total runtime improvement to just 140 of them. Among the 546 loops left scalar, the compiler's own diagnostics cited a possible aliasing between pointer arguments in 41 percent of cases, an unknown trip count or a non-unit stride in 27 percent, and a loop-carried dependence through a reduction variable of a type the pass did not recognize in 14 percent. Adding restrict qualifiers to the pointer arguments of the 224 alias-blocked loops allowed 189 of them to vectorize and improved suite runtime by a further 7 percent, while the dependence-blocked loops mostly involved floating-point reductions that require a reassociation flag the benchmark rules forbid. Comparing the same loops under a second compiler showed a similar overall rate but a different failure profile, with far fewer alias refusals and more trip-count refusals, which the authors attribute to that compiler's more aggressive interprocedural alias analysis."
},
{
"query": "radar absorbing material layer thickness",
"passage": "Flat panel samples of a carbon-loaded polyurethane foam absorber were tested in an anechoic chamber against an X-band radar operating from 8 to 12 GHz, with panel thickness varied from 3 to 12 millimeters in 1 millimeter increments backed by a grounded aluminum plate. Reflectivity, expressed in decibels relative to a bare metal reference, showed a single absorption minimum near negative 22 dB at 10.5 GHz for the 6 millimeter panel, consistent with a quarter-wavelength resonance condition given the material's measured complex permittivity at that frequency. Thinner panels shifted this minimum to higher frequency and reduced its depth, with the 3 millimeter sample bottoming out at only negative 11 dB near 11.8 GHz, while thicker panels below 12 millimeters broadened the absorption band but at the cost of a shallower minimum, since the resonance condition was satisfied at a frequency where the material's loss tangent was lower. A two-layer configuration pairing a 4 millimeter high-loss front layer with an 8 millimeter lower-loss backing layer achieved better than negative 15 dB reflectivity across the full 8 to 12 GHz band, broader bandwidth than any single-layer thickness tested, at the expense of 50 percent more total mass per unit area than the best single-layer result. Oblique incidence measurements at 45 degrees showed the absorption minimum shifting down by roughly 0.6 GHz relative to normal incidence for the two-layer panel, attributed to the effective path length increase through the absorber at the steeper angle. Mass loading remains the primary practical constraint on absorber selection for airborne applications, and the study's authors noted that the two-layer design's bandwidth advantage would need to be weighed against this weight penalty for any specific aircraft structural application rather than treated as a universal improvement.",
"nearMiss": "Conducted emissions from a 150 watt switch-mode power supply intended for laboratory instruments were measured on a line impedance stabilization network across the 150 kilohertz to 30 megahertz range to identify which parts of the input filter mattered for meeting the class B limit that the product's market requires. In its initial form the supply exceeded the limit by up to 9 decibels between 300 and 800 kilohertz, and separating the common-mode and differential-mode components with a current probe showed that the excess was almost entirely common-mode, coupled from the switching transistor's heat sink to the chassis. Adding a second common-mode choke reduced the excess by 5 decibels but the remaining margin was recovered only by bonding the heat sink to the primary-side return through a 1 nanofarad capacitor, which removed the coupling path rather than filtering its result. Differential-mode emissions, already 6 decibels under the limit, rose slightly when the choke was added because of a resonance between its leakage inductance and the existing X capacitor, and moving the capacitor to the line side of the choke removed it. The final configuration passed with a 4 decibel margin at every frequency, and the authors note that the chassis bonding change cost less than the choke it made redundant."
},
{
"query": "What mechanisms drive long-term subsidence in sedimentary basins undergoing groundwater extraction?",
"passage": "Extensometer and GPS benchmark records spanning 34 years across the basin's central well field showed cumulative land subsidence reaching 2.1 meters at the location of heaviest groundwater withdrawal, with subsidence rate closely tracking the seasonal drawdown and recovery cycle of the underlying aquifer's piezometric surface during the first two decades of the record. Core samples from the compacting silty clay interbeds revealed that roughly 70 percent of the total measured compaction had become inelastic and irrecoverable by the later years of the record, meaning the aquifer system's storage capacity had been permanently reduced even during periods when groundwater levels partially recovered following managed recharge efforts. A one-dimensional consolidation model calibrated against the extensometer data attributed the transition from elastic to predominantly inelastic compaction to the piezometric surface dropping below its historical minimum, or preconsolidation stress equivalent, sometime around the eleventh year of the record, after which further drawdown produced compaction at roughly four times the rate per unit head decline seen in the earlier elastic-dominated period. Managed aquifer recharge introduced in year 22, delivering treated surface water through injection wells at the basin margins, slowed the regional subsidence rate by about 35 percent within five years but did not reverse it, consistent with the earlier finding that most compaction in the affected interbeds could not be recovered by raising water levels alone. Comparison with an independent basin lacking significant clay interbeds, where withdrawal-induced head decline of similar magnitude produced almost no measurable subsidence, reinforced the interpretation that interbedded fine-grained sediment compressibility, rather than head decline alone, is the primary control on whether groundwater extraction translates into significant permanent land subsidence in a given basin.",
"nearMiss": "Repeat terrestrial laser scanning of a 1.8 kilometer stretch of sandstone sea cliff, carried out from fixed stations four times a year for twelve years, produced a record of volumetric loss detailed enough to separate gradual surface lowering from discrete block falls. Mean retreat averaged 0.11 meters per year but was highly episodic, with 62 percent of the total volume lost in fourteen individual failures each exceeding 400 cubic meters, and the largest, in the ninth winter, removing a 3,100 cubic meter slab along a pre-existing joint after three weeks of onshore storms. Small-scale loss, made up of grain-by-grain and flake detachment across the whole face, proceeded at a near-constant rate of about 2 centimeters per year and showed no dependence on season or wave climate, whereas the large failures were concentrated in winters with the highest recorded wave energy and clustered in the six weeks after prolonged rainfall. Notch formation at the cliff toe preceded the major falls by one to three years in every case, and the notch depth at which failure occurred scaled with the spacing of the vertical joints. The record suggests that cliff-top setback distances based on average retreat rates underestimate the hazard at joint-bounded sections by a factor of two or more."
},
{
"query": "Poincare section",
"passage": "A double pendulum apparatus with arm lengths of 0.25 and 0.20 meters and lumped end masses of 0.15 and 0.10 kilograms was released from a range of initial angles and tracked using a 240 frames per second overhead camera with sub-millimeter marker resolution on each arm. For small release angles below about 15 degrees from vertical, the motion remained quasi-periodic over the full 90 second recording window, and a Poincare section constructed by sampling the second arm's angle and angular velocity each time the first arm crossed zero angle with positive velocity produced a closed, smooth curve typical of regular motion confined to a low-dimensional torus in phase space. Above a release angle of roughly 55 degrees, the same Poincare section construction instead produced a diffuse scatter of points filling a bounded region of the phase plane, the signature of chaotic trajectories exploring a larger accessible volume of phase space rather than remaining confined to a simple orbit. The largest Lyapunov exponent, estimated from the divergence rate of two trajectories released from nearly identical initial conditions differing by 0.001 radians, was measured at approximately 2.8 per second in this chaotic regime, meaning an initial angular difference too small to detect by eye grew to order-one size within roughly 0.4 seconds. Energy dissipation from joint friction, quantified separately by tracking the amplitude decay of small-angle oscillations, was small enough over the 90 second window to be treated as a minor correction rather than the dominant factor shaping the qualitative transition between regular and chaotic behavior. The measured transition angle of 55 degrees was somewhat lower than the value predicted by the idealized frictionless double pendulum equations, a discrepancy the authors attributed to the finite size and moment of inertia of the physical end masses relative to the point-mass idealization used in the reference equations of motion.",
"nearMiss": "A batch Belousov-Zhabotinsky reaction in a stirred 50 milliliter cell was driven by a periodic light pulse of variable frequency to map the range of driving periods over which the chemical oscillator locks to the external forcing. Unforced, the reaction oscillated with a natural period of 42 seconds that lengthened slowly as reagents were consumed, and the color change of the ferroin indicator was tracked by a photodiode to give a continuous record of phase. When the light pulse period was within about 8 percent of the natural period the oscillation locked one-to-one within three cycles, and two-to-one and three-to-two locking were observed over narrower ranges near twice and one and a half times the natural period, producing the expected tongue structure when locking range was plotted against pulse intensity. Outside the tongues the phase drifted quasi-periodically, and near the tongue boundaries the record showed intermittent slips in which the reaction skipped a pulse after a run of locked cycles. Increasing the pulse intensity widened the primary tongue linearly up to a threshold beyond which the light suppressed the oscillation altogether for the duration of each pulse, and this suppression threshold fell as the reagents aged."
},
{
"query": "ion implantation dose uniformity control",
"passage": "Boron implantation into 200 millimeter silicon wafers was performed at 40 keV to a nominal dose of 2 times 10^15 ions per square centimeter using a batch implanter with wafers mounted on a rotating disk scanned mechanically beneath a ribbon-shaped ion beam. Dose uniformity across each wafer was measured by four-point probe sheet resistance mapping at 49 points following a rapid thermal anneal at 1000 degrees Celsius for 10 seconds, with the as-received beam scan producing a within-wafer nonuniformity of 3.2 percent, one standard deviation, dominated by a radial gradient attributed to a slight mismatch between the disk rotation speed and the beam's mechanical scan rate across the batch. Adjusting the scan overlap pattern to compensate for the measured beam current profile, which had been characterized separately with a Faraday cup array positioned at the wafer plane, reduced within-wafer nonuniformity to 1.1 percent without any hardware modification to the beamline itself. Wafer-to-wafer dose repeatability across a lot of 25 wafers processed consecutively stayed within 0.8 percent of the target dose as tracked by an integrated beam current dosimetry system, though a gradual drift of about 0.4 percent per hour in apparent dose was traced to slow warming of the Faraday cup electronics over a multi-hour processing run, correctable by periodic recalibration every two hours. Secondary ion mass spectrometry depth profiles confirmed that the improved scan pattern did not alter the implant's depth distribution, only its lateral uniformity, with peak concentration remaining at the expected depth of approximately 130 nanometers for this energy and species. The authors noted that nonuniformity contributions from beam profile mismatch and from channeling variation across slightly miscut wafers were difficult to separate using sheet resistance mapping alone, since both produce similar radially symmetric signatures on this wafer geometry.",
"nearMiss": "Line-edge roughness in a chemically amplified photoresist was studied as a function of post-exposure bake temperature and duration on 300 millimeter wafers patterned with 45 nanometer dense lines, using critical-dimension scanning electron microscopy to measure edge position along 2 micrometer lengths at 40 sites per wafer. Raising the bake from 100 to 120 degrees Celsius at a fixed 60 seconds reduced the three-sigma roughness from 4.8 to 3.6 nanometers, with most of the improvement occurring between 110 and 115 degrees, but also increased the critical dimension by 3.1 nanometers through additional acid diffusion, requiring an exposure adjustment of about 6 percent to hold the target width. Extending the bake at 115 degrees from 60 to 120 seconds gave a further 0.3 nanometer reduction that the authors judged not worth the throughput penalty. Power spectral analysis of the edge profiles showed the improvement concentrated at spatial frequencies above 1 per 100 nanometers, consistent with smoothing by acid diffusion, while the low-frequency component associated with the mask and aerial image was unchanged. Across-wafer variation in roughness correlated with the bake plate's temperature map, and replacing the plate reduced the site-to-site spread by half."
},
{
"query": "How does multipath propagation limit data rates in underwater acoustic communication?",
"passage": "Field trials in a 45 meter deep coastal channel used a vertical transmitter array and a single hydrophone receiver moored 800 meters away, transmitting a phase-shift-keyed signal centered at 12 kHz with a symbol rate of 2000 symbols per second. Channel impulse response measurements, obtained by correlating received signals against a known probe sequence transmitted every 30 seconds, showed a dominant direct path arrival followed by surface- and bottom-reflected arrivals spread over a delay spread of approximately 8 milliseconds, corresponding to roughly 16 symbol periods at the tested symbol rate, a spread large enough to cause severe intersymbol interference without equalization. An adaptive decision-feedback equalizer with 32 feedforward and 16 feedback taps, updated using a recursive least squares algorithm, reduced the measured bit error rate from an unequalized 1.8 times 10^-1 to 3 times 10^-3 under calm sea-surface conditions, but performance degraded to a bit error rate of 4 times 10^-2 once wind speed exceeded 8 meters per second and surface wave motion introduced time-varying reflection delays that outpaced the equalizer's adaptation rate. Doppler spreading from platform drift, measured at up to 0.8 knots during the trials, further degraded performance when combined with high sea state, since the resulting frequency spreading interacted with the already time-varying multipath structure in a way the equalizer's fixed update rate could not fully track. Reducing the symbol rate to 1000 symbols per second, effectively halving the delay spread's impact in symbol periods, restored bit error rate to below 5 times 10^-3 even under the rougher conditions, at the direct cost of halving the achievable data throughput, illustrating the fundamental tradeoff between data rate and multipath robustness that constrains underwater acoustic link design in shallow, reflective channels.",
"nearMiss": "A bottom-mounted hydrophone recording continuously at a depth of 1,100 meters on a continental slope over 26 months was used to describe the seasonal occurrence of fin whale 20 hertz calls, which were detected automatically by matching the spectrogram against a template and then verified by an analyst on a subsample. Calls were present on 71 percent of days overall, with a pronounced peak from October to February when the characteristic song pattern of regularly repeated pulses was recorded on nearly every day, and a minimum in June and July when only sporadic irregular calls were detected. Received levels of the loudest calls exceeded 130 decibels relative to one micropascal, and comparison with a calibrated source towed past the site at known ranges placed the closest singers within 15 kilometers of the instrument. Song inter-pulse interval, measured on 4,200 song bouts, lengthened from a mean of 13.1 seconds in the first winter to 13.9 in the second, continuing a slow drift reported from other basins. Detection rates dropped by about half during periods of high wind when ambient noise in the band rose by more than 6 decibels, and the authors caution that seasonal patterns in such records partly reflect the noise climate rather than the animals."
},
{
"query": "birthday bound",
"passage": "The compression function under study produces a 160-bit digest from 512-bit message blocks using a Davies-Meyer construction built on a dedicated block cipher, and the analysis focused on how closely practical collision-finding attempts approached the generic birthday bound of roughly 2^80 operations expected for an ideal random function of this output size. A parallel collision search implemented across a cluster of 64 GPUs, using distinguished-point tracking to reduce memory requirements for cycle detection in the underlying random walk, found collisions in a deliberately reduced 48-bit-output variant of the function after approximately 2^25 evaluations, closely matching the birthday-bound prediction for that truncated size and validating the implementation's fidelity to the theoretical model before scaling conclusions up. Differential cryptanalysis of the full 160-bit function identified a characteristic through 44 of the compression function's 80 rounds with a predicted probability of 2^-61, meaning that if such a characteristic could be exploited end to end it would represent a shortcut well below the generic birthday bound, though extending the characteristic to the full round count reduced its probability below the level needed to offer any practical advantage over generic search. No full-round collision attack faster than generic birthday search was demonstrated, and the authors' best reduced-round results reached 52 of 80 rounds before the attack's complexity exceeded 2^80, at which point it no longer improved on brute-force birthday search and was reported mainly to characterize the security margin rather than as a practical threat. The reduced-round results nonetheless prompted the recommendation of an increased round count for any new deployment of the construction, since the margin between the best attack at 52 rounds and the full 80 rounds, while currently comfortable, had eroded compared to the margin reported for this function family in an earlier survey of hash function cryptanalysis.",
"nearMiss": "An authentication server for a web application with 1.9 million registered accounts was profiled under production load to choose a work factor for its password hashing scheme, balancing resistance to offline guessing against the latency users see at login. At the existing setting the hash took a mean of 38 milliseconds on the server's cores and login requests averaged 61 milliseconds end to end, of which the hash accounted for roughly two thirds. Doubling the work factor raised the hash to 77 milliseconds and total login latency to 102, and during the evening peak of 210 logins per second the hashing alone consumed 16 of the server's 24 cores, leaving too little headroom for the rest of the request path and producing a queue that pushed 95th-percentile latency to 340 milliseconds. Moving the hash computation to a dedicated pool of four machines removed the contention and allowed the doubled work factor to run with a 95th-percentile login latency of 118 milliseconds, a change user testing rated as unnoticeable. The team estimated that the doubled setting raises the cost of an offline guessing attack against a stolen database by the same factor, and adopted a policy of doubling again whenever the median hash time on new hardware falls below 40 milliseconds."
},
{
"query": "Bruges fullers' ordinance",
"passage": "The 1361 ordinance preserved in the Bruges stadsarchief, register series 114, folio 22, laid out in unusual detail the obligations imposed on fullers working along the Speyestraat canal, where the mechanical fulling mills drew their power from a controlled sluice gate maintained jointly with the dyers' guild. Fullers who allowed cloth to sit unattended in the trough for more than a full tide were fined four groten, a sum equivalent to roughly two days' wages for an unskilled laborer, and repeat offenders faced confiscation of their tools for a season. What makes the register valuable is not the fine schedule itself, which resembles similar provisions in Ghent and Ypres, but the marginal annotations added by a later clerk, apparently around 1390, correcting the currency conversion after the debasement of the Flemish groot. These corrections suggest the ordinance remained in active use for at least three decades, contradicting an older assumption that such municipal regulations were largely aspirational and rarely enforced past their first decade. Complaints filed by individual fullers, six of which survive in loose sheets tucked into the same register, describe disputes over water rights during dry summers, when the sluice keeper allegedly favored dyers' vats over the fulling troughs downstream. One complainant, a fuller named Coppin van Male, petitioned the aldermen directly in 1372, arguing that the dyers' guild had bribed the sluice keeper, though the outcome of his petition is not recorded. Taken together, the ordinance and its accompanying complaints offer a rare glimpse of how craft regulation functioned not as a fixed code but as a contested, continually renegotiated arrangement between guilds sharing scarce urban infrastructure, one in which formal fines coexisted uneasily with informal favoritism that the written record only partially conceals.",
"nearMiss": "Excise accounts kept by the aldermen of Leuven for the years 1372 to 1379 record the duty paid on each brew by the town's licensed brewers, and because the duty was levied per vat rather than per barrel sold, the accounts allow the output of individual brewhouses to be reconstructed week by week. Sixty-one brewers appear in the earliest year, falling to forty-eight by the last, while total taxed output rose by roughly a fifth, so that the average brewhouse grew and the four largest came to account for a third of the town's production. Brewing was strongly seasonal, with output in the weeks before Lent and around the autumn fairs running at nearly twice the summer level, when the accounts also record several brewers paying the reduced duty for small beer only. The clerks noted fines for brewing with untaxed grain on nineteen occasions, all but two in years when grain prices, recorded separately in the town's purchase accounts, stood above their decade average. Marginal notes identify seven brewhouses as operated by widows continuing their husbands' licenses, and two of these are among the ten largest producers, which qualifies the assumption that female-run enterprises in the trade were marginal."
},
{
"query": "how neume spacing indicated rhythmic duration",
"passage": "The gradual copied at the abbey of Saint-Ouarn sometime before 1140, now held in fragmentary form as four bifolia bound into a later breviary, uses a form of Aquitanian neume in which the horizontal spacing between individual signs varies according to the syllable's presumed duration, a practice that diverges from the more rigidly proportioned notation found in contemporary manuscripts from Cluny. Where a syllable carried a single note, the copyist left barely enough space for the neume itself before beginning the next syllable; where a syllable carried an extended melisma, the neumes are spread across nearly twice the horizontal distance, with faint dry-point ruling visible under raking light that appears to have guided this expansion. An earlier survey of the archive catalogued the manuscript simply as a damaged gradual of local Aquitanian type without commenting on this feature, but a closer comparison with the abbey's customary, which specifies that certain feast-day chants be sung 'more slowly, in the manner of the elders,' suggests the spacing may encode an oral performance tradition rather than a purely notational convention. Three of the four surviving bifolia show this expanded spacing exclusively on chants assigned to major feasts, while ordinary weekday chants are notated with uniform, compressed spacing throughout. If the correlation holds across the lost portions of the manuscript, it would indicate that the scribe, or whoever supervised the copying, treated notational space itself as a marker of liturgical solemnity, independent of any explicit rhythmic sign. This reading remains provisional, since only a single scribal hand is represented and no comparable gradual from Saint-Ouarn survives to confirm whether the practice was idiosyncratic to this copyist or reflected a broader house convention.",
"nearMiss": "The paper stock of a polyphonic choirbook long assigned to the 1480s on stylistic grounds was examined leaf by leaf under transmitted light, recording 214 watermarks belonging to eleven distinct molds, and each mold was compared against dated watermark repertories and against archival paper of known date in the same region. Eight of the eleven molds could be matched to twin pairs used in documents dated between 1496 and 1503, and the three remaining molds, all bearing a variant of a bull's head with a serpent, appeared in notarial registers of 1501 and 1502 only. The distribution of the molds through the gatherings was not random: the first four gatherings use a single pair of molds throughout, while the later gatherings mix five or more, suggesting that the scribe began with a purchased ream and later drew on smaller lots as work proceeded. Chain-line interval measurements confirmed the mold identifications where the watermark itself was damaged by trimming. Taken together the evidence moves the copying of the manuscript forward by at least a decade, to about 1500 to 1504, which brings it into the period when the chapel that owned it is documented as employing the singers whose names appear in its margins."
},
{
"query": "What did an overturned glass symbolize in Dutch still life paintings?",
"passage": "An inventory drawn up in 1671 following the death of a Delft wine merchant lists among his household goods 'a small piece with a fallen roemer and half-peeled lemon,' almost certainly a banquet still life of the type produced in quantity by painters working in the wake of the Haarlem tradition. The overturned roemer, tipped so that a thin stream of wine spills across the tablecloth without yet reaching the table's edge, recurs across dozens of surviving banquet pieces from the 1650s and 1660s, and the consistency of the motif has led students of the genre to treat it as a stock reminder of life's precariousness, on a par with the guttering candle or the half-eaten pie left to attract flies. Yet the merchant's inventory pairs this description with an appraisal value nearly double that of a comparable flower piece listed two lines below, suggesting that contemporary buyers did not necessarily receive such paintings primarily as moral instruction; the higher valuation more plausibly reflects the technical difficulty of rendering glass and reflected light convincingly, a skill for which certain workshops charged a premium regardless of subject. Notarial records from the same decade occasionally specify glassware still lifes as suitable gifts for wedding settlements, which sits awkwardly with a purely admonitory reading of the overturned glass as a warning against excess. It seems more likely that the motif carried a double valence, legible simultaneously as a conventional memento of transience and as a display piece prized for its illusionistic virtuosity, with individual buyers and rooms determining which register predominated. Later collectors, writing in the following century, tended to flatten this ambiguity, describing such works uniformly as moralizing vanitas, a simplification that later cataloguers have been slow to revise.",
"nearMiss": "Registers of the painters' guild in Leiden surviving for 1648 to 1669 record the entrance fees paid by masters and the enrollment of apprentices, and cross-referencing them with the town's tax assessments makes it possible to follow how the trade was recruited and where its practitioners stood economically. Forty-three masters were admitted over the period, of whom twenty-nine had served apprenticeships in the town and the remainder arrived as journeymen from elsewhere, paying a fee half again as large. Apprentice enrollments peaked in the mid-1650s at eleven in a single year and fell to two or three annually after 1662, a decline the registers do not explain but which coincides with the collapse of prices recorded in the town's auction house after 1660. Tax assessments place most masters in the middle of the town's property distribution, with only four assessed among the top tenth of households, and the wealthiest of these owned a house and shop on the market square, dealt in pictures by others as well as his own, and enrolled six apprentices over his career, more than any other master. Five masters appear in the poor-relief accounts in their final years."
},
{
"query": "avunculocal residence",
"passage": "Fieldnotes compiled during an eighteen-month stay on the island of Tarawau describe a residence pattern in which a young married man customarily relocates to the household of his mother's brother rather than remaining with his own father, an arrangement classified in the kinship literature as avunculocal residence and documented in relatively few Pacific societies compared with the far more common patrilocal or matrilocal patterns. Genealogies collected from four hamlets show that of sixty-one married men surveyed, forty-four had taken up residence with a maternal uncle within two years of marriage, typically the uncle who held the largest yam terraces, while the remainder had either remained patrilocal or, in a handful of cases, moved to a wife's village where no eligible maternal uncle survived. Informants explained the practice not in terms of descent ideology but in practical terms of land access, since terrace rights on Tarawau pass preferentially through the mother's line even though formal title and ceremonial authority remain vested in the father's lineage, producing a split between economic and ceremonial affiliation that the residence pattern appears designed to reconcile. This arrangement complicates any simple equation between matrilineal descent and matrilocal residence, since Tarawau kinship reckoning is itself only partially matrilineal, tracing ceremonial rank patrilineally while treating land and, notably, fishing rights as inherited through women. Disputes over terrace boundaries recorded during the fieldwork period frequently invoked a nephew's residence with his uncle as evidence of legitimate claim, suggesting that avunculocal residence functions locally as a kind of continuous, embodied title registration in a society without written land records. Whether this pattern predates European contact or represents an adaptation to population pressure on limited terrace land could not be determined from oral testimony alone.",
"nearMiss": "Marriage payments recorded among a cattle-keeping population in the eastern savanna were compiled from 312 unions contracted between 1962 and 1995, drawing on household histories collected in three villages and on the ledgers kept by two local courts that registered disputed payments. The number of cattle transferred at marriage rose from a median of eleven in the 1960s to nineteen in the 1990s, while the proportion of the payment made in cash rather than in animals grew from almost nothing to about a third, most often as a substitute for the smaller stock traditionally included. Payments were spread over longer periods in the later decades, with the final installment typically delivered after the birth of the couple's second child rather than before the wedding, and court records show disputes over unpaid installments becoming the most common category of case by the 1980s. Informants attributed the rise to competition among families whose sons had earned wages in the mining towns, and the household histories show that men with wage income married at a median age two years younger than those without. The authors argue that the shift toward deferred payment allowed the institution to persist under inflation that would otherwise have priced many young men out of marriage."
},
{
"query": "reasons small parties favored proportional representation",
"passage": "Parliamentary debates recorded in the Belgian lower house between 1894 and 1899 reveal that support for the d'Hondt method of seat allocation came disproportionately from deputies representing the Catholic party's rural wing and from the small but growing socialist bloc, an alliance that appears counterintuitive only if one assumes proportional systems chiefly serve numerically weak parties. The Catholic rural deputies calculated, correctly as it turned out, that first-past-the-post arrangements in mixed urban-rural districts tended to hand disproportionate advantage to Liberal candidates who could concentrate their vote in provincial towns, while a proportional formula would preserve Catholic representation drawn from dispersed rural support that could not easily carry a plurality in any single district. Socialist deputies, for their part, held almost no seats under the existing arrangement despite polling respectably in industrial constituencies, since their vote was spread too thinly to win outright in most districts, and party leaders argued openly in committee that proportional allocation was the only realistic path to parliamentary representation matching their share of the electorate. Liberal deputies, who benefited most from the status quo, resisted the change through 1898, but their position weakened after the 1898 elections produced a legislature widely regarded as unrepresentative of the actual vote distribution, prompting even some previously skeptical Liberal deputies to reconsider. The eventual 1899 reform adopted the d'Hondt divisor method largely intact from the proposal drafted by a parliamentary commission two years earlier, with only minor adjustments to district magnitude. What the debates make clear is that support for proportional representation in this instance emerged from a tactical convergence between ideologically opposed parties rather than from any shared commitment to proportionality as a democratic principle in the abstract.",
"nearMiss": "The introduction of verbatim stenographic recording of debates in the Danish Folketing in 1850 offers a chance to test whether members changed how they spoke once every word would be printed, since the preceding decade of the advisory assemblies had been recorded only in summary by a secretary. Comparing the summaries of 1846 to 1848 with the transcripts of 1850 to 1853 for the same forty-two members who served in both bodies, speeches became longer, from an estimated median of eight minutes to about fourteen, and members increasingly read prepared texts, a practice the presidents complained about in six recorded rulings. Interjections and exchanges between speakers, frequent in the secretaries' accounts, all but vanish from the transcripts, and the number of members speaking in a typical sitting fell by a quarter. Newspapers in Copenhagen and the provinces reprinted transcript extracts within two days of a sitting, and letters to the editor citing a member's exact words appear for the first time in 1851. The members most active in the earlier assemblies were not those who spoke most under the new system, and several rural members who had been prominent before 1850 rarely appear in the transcripts at all."
},
{
"query": "Why did 19th century industrial cities retain more heat at night?",
"passage": "Meteorological logs kept by the observatory at Kesterbridge between 1861 and 1889 record nighttime temperatures in the town center that averaged nearly four degrees Fahrenheit higher than simultaneous readings taken at a rural station eleven miles distant, a gap that widened noticeably during the winter months when coal consumption for heating and manufacturing peaked. The observatory's keeper attributed the difference at the time to the shelter provided by dense building, but a later reexamination of the logs alongside brick and paving records suggests a more specific mechanism: the extensive replacement of packed earth streets with fired clay brick paving between 1855 and 1870 coincided almost exactly with the widening of the temperature gap, since brick and stone absorb solar heat during the day and release it slowly overnight, unlike bare soil, which releases stored heat far more quickly after sunset. Coal smoke, thick enough by the 1870s to reduce recorded sunshine hours by roughly a third compared with the rural station, likely compounded the effect by trapping outgoing radiation close to the ground rather than allowing it to escape freely into the night sky, though the observatory's instruments were not designed to isolate this contribution separately from the paving effect. Industrial furnaces themselves, concentrated along the canal district, added a further direct source of heat that operated continuously regardless of the diurnal cycle. Comparable gaps recorded in nearby Ashcombe, a town of similar population that retained unpaved streets through the 1880s, were consistently smaller, roughly half the magnitude observed at Kesterbridge, lending some support to the paving hypothesis over smoke alone, though the two towns differed enough in industrial composition that the comparison cannot be treated as a controlled test.",
"nearMiss": "Toll registers kept at a stone bridge over a river in the eastern lowlands between 1802 and 1878 record the day each spring on which the ferry service that operated alongside the bridge resumed after the winter ice went out, providing a continuous record of ice break-up date for a period before instrumental observations began in the district. The break-up date ranged from the last week of February to the third week of April, with a mean around 24 March, and the ten earliest years all fell after 1850, while the latest cluster in the 1810s and 1830s. Comparison with the temperature series from an observatory 140 kilometers away, which begins in 1841, shows the break-up date advancing by about six days for each degree Celsius of warming in the mean March temperature over the overlapping years, and applying that relation to the toll record suggests that the decade of the 1810s was colder than any decade in the instrumental period. Entries also record eleven springs in which the ice went out in a single day with flooding of the approach road, and these coincide with years the parish accounts list payments for repairs to the bridge abutments."
},
{
"query": "hospital of Roncesvalles",
"passage": "Account books kept by the hospital of Roncesvalles for the years 1275 to 1281, among the more complete surviving records for any pilgrim hostel along the Iberian route to Santiago, list nightly admissions that rose sharply each year in the weeks following Easter and again, though less steeply, around the feast of Saint James in late July, patterns consistent with the two principal pilgrimage seasons noted in guidebooks of the period. The hospital's provisioning entries record bread, wine, and salted fish purchased in bulk from merchants in nearby Pamplona, with quantities scaled closely to the previous year's admission figures, suggesting the hospital's stewards kept and consulted earlier account books rather than provisioning reactively. Notable among the entries are periodic references to pilgrims arriving 'sine denariis,' without money, who were nonetheless admitted and fed at the hospital's expense, a practice the accounts treat as routine rather than exceptional, appearing in roughly one entry in six across the surviving years. This pattern complicates a reading of medieval hospital charity as primarily symbolic or occasional, since the scale of uncompensated provisioning implied by these entries would have required a stable and apparently anticipated subsidy, likely drawn from the substantial land grants the hospital held across Navarre. Separate entries record disputes with a neighboring monastery over grazing rights on land the hospital claimed had been granted alongside its founding charter, disputes that recur in the accounts for at least four consecutive years without clear resolution noted. Taken as a whole, the account books suggest an institution operating at a scale and with a degree of logistical foresight considerably beyond what its modest surviving physical remains, largely a single fortified building, would suggest to a visitor unfamiliar with the documentary record.",
"nearMiss": "The wax accounts of a lay confraternity in a Tuscan hill town, kept in a single ledger covering 1381 to 1394, record every purchase of candles and every allocation of them to the brotherhood's own services, to funerals of members, and to the feasts of the parish church, and so allow the rhythm of the confraternity's devotional year to be reconstructed in some detail. Annual expenditure on wax ranged between 18 and 31 lire, roughly a fifth of the brotherhood's total spending, and the largest single allocation each year, between 40 and 55 pounds of wax, went to the Corpus Christi procession, followed by the vigil of the patron saint and the feast of the Assumption. Funeral candles account for a rising share over the period, from 9 to 22 percent, which the entries themselves explain by noting the deaths of members in the plague years of 1383 and 1390. Prices per pound recorded in the ledger rose by a third across the period, and from 1388 the brotherhood began buying tallow for its ordinary weekly services, reserving beeswax for the major feasts, a substitution the ledger records without comment but which a later chapter statute forbids."
},
{
"query": "disputes over amateur radio wavelength allocation",
"passage": "Correspondence preserved in the files of a regional broadcasting board established in 1923 documents a running dispute between commercial station operators and amateur wireless operators over the allocation of wavelengths in the lower band, a conflict that intensified after the board granted priority clearance to a commercial station whose transmissions had previously overlapped with an amateur club operating from a rented room above a hardware store. The amateur operators, organized loosely under a regional wireless society with roughly ninety members, petitioned the board directly, arguing that amateur experimentation had itself established many of the technical practices commercial broadcasters now relied upon and that reassignment to a narrower band would render much of their existing equipment obsolete. Board minutes record a compromise proposal, floated by a junior engineer on the regulatory staff, to shift amateur transmissions to a shorter wavelength band still considered experimental at the time, a proposal the amateur society initially rejected on the grounds that receivers capable of tuning the shorter band were not yet widely available to hobbyists. The commercial station's owner, for his part, argued in submitted correspondence that interference from amateur transmissions during evening hours, when both amateur activity and commercial listening peaked simultaneously, represented a direct threat to advertising revenue, since sponsors had begun to condition their contracts on reliable evening reception. The board ultimately adopted a modified version of the shorter-wavelength proposal in 1925, phased in over eighteen months to allow amateur operators time to acquire suitable equipment, though correspondence from the following year indicates continued informal interference complaints from listeners in outlying districts where reception had never been reliable under either arrangement.",
"nearMiss": "Records of the telegraph department's tariff committee for 1905 to 1908 document the campaign by provincial newspapers for a reduced press rate for night transmission of news copy, and the department's calculation of what such a rate would cost it. The newspapers argued that the existing rate of one penny per word made it uneconomic for any but the metropolitan dailies to take a full evening service from the capital, and that the wires stood idle for much of the night in any case. The department's traffic returns supported the second claim, showing average night loading on the main trunk circuits at under a fifth of capacity, and its engineers estimated that a press rate of one third of a penny per word would fill about half the idle capacity without requiring new lines. The committee nonetheless resisted for two years, chiefly on the argument advanced by the accountant that the metropolitan papers, which already paid the full rate, would simply transfer their traffic to the cheaper night hours and reduce revenue overall. A one-year trial on two trunk routes in 1907 showed night press traffic rising sevenfold while daytime press traffic fell by less than a tenth, and the rate was extended to the whole network the following year."
},
{
"query": "How did lenition affect intervocalic consonants in early Romance dialects?",
"passage": "Charters copied at the scriptorium of a monastery in the Pyrenean foothills between roughly 920 and 980 show increasingly frequent spelling variants in which the Latin intervocalic voiceless stops -p-, -t-, and -k- are rendered with letters conventionally associated with their voiced counterparts, a pattern scribes seem to have introduced gradually rather than adopting wholesale, since earlier charters in the same collection preserve the classical spelling far more consistently. A charter dated 942 spells the place name later attested as 'Lagata' with a medial -t-, while a charter from the same scriptorium dated 971 renders what appears to be the identical place name with -d-, and by the 990s the voiced spelling predominates in the collection's later documents, suggesting the underlying pronunciation had shifted to a lenited, voiced articulation well before scribal orthography caught up with it. This lag between phonetic change and its written representation is unsurprising given that Latin spelling conventions remained institutionally prestigious long after everyday speech had diverged from them, but the collection's density of dated charters allows the transition to be tracked with more precision than is usually possible for this period and region. Notably, the shift appears earlier and more consistently in common nouns than in personal names, where scribes seem to have preserved traditional spellings out of respect for established naming conventions even after abandoning classical orthography elsewhere, a distinction that complicates any simple model of orthographic change as a uniform process across the lexicon. Place names occupy an intermediate position, shifting later than common nouns but earlier than personal names, plausibly because they lacked the same weight of dynastic tradition while still functioning as fixed referential labels.",
"nearMiss": "A corpus of 214 Old English homiletic and historical prose texts, tagged for clause type and for the position of the finite verb, allows the distribution of verb-final order in subordinate clauses to be compared across three roughly dated bands of composition spanning the late ninth to the late eleventh century. In the earliest band, subordinate clauses placed the finite verb in final position in 74 percent of instances, falling to 61 percent in the middle band and 43 percent in the latest, while main clauses showed no comparable trend. The decline was steepest in clauses containing a heavy object phrase of three or more words, where verb-final order fell from 68 to 29 percent, and slowest in short clauses with a pronominal object, which retained the older order in more than half of instances even in the latest texts. Texts translated closely from Latin sources showed higher rates of verb-final order than original compositions of the same date, by about ten percentage points, which the authors take as a caution against reading translated prose as direct evidence for the spoken language. Two texts assigned to the latest band on paleographic grounds pattern with the middle band, and the authors suggest they may be copies of earlier compositions."
},
{
"query": "rural literacy brigades",
"passage": "Government reports filed between 1944 and 1947 describe the deployment of rural literacy brigades, teams typically consisting of a trained teacher and two or three secondary students recruited from provincial normal schools, to villages across the northern highlands where school attendance had historically been lowest. Each brigade carried a standardized primer built around vocabulary drawn from agricultural life, a deliberate departure from the urban-oriented primers used in city schools, on the reasoning that adult learners would engage more readily with reading material describing tasks and objects already familiar to them. Enrollment figures collected by district supervisors show substantial variation in completion rates across villages, ranging from as low as twelve percent in communities where the agricultural calendar left little time for evening classes during harvest season to over sixty percent in villages closer to market towns where brigades could offer classes during a wider range of hours. Supervisors' correspondence attributes much of this variation to whether brigades successfully recruited a respected local figure, often a landowner or parish official, to publicly endorse the program, since villages lacking such endorsement showed markedly lower initial enrollment regardless of brigade effort. An earlier survey of the archive found relatively little attention paid to gender differences in enrollment, but a closer reading of the district reports shows that women's enrollment lagged men's by roughly half in most villages, a gap supervisors attributed variously to domestic labor demands and to reluctance among some husbands to permit evening attendance, though the reports rarely probe this reluctance further. The campaign's own internal evaluations, prepared for the ministry in 1948, judged the brigades a qualified success.",
"nearMiss": "Provincial ledgers kept between 1951 and 1954 record a government loan program that paid independent contractors a fixed rate per kilometer to grade dirt access roads connecting isolated farmsteads in the eastern lowlands to the existing market roads, with payment released only after a district engineer certified that the surface met a minimum width and drainage standard. In the program's first year 412 kilometers were certified across four districts, rising to 1,070 kilometers in the third, and the ledgers show the average cost per kilometer falling by a quarter over the period as contractors acquired their own graders in place of hired animal teams. Rejection rates at inspection ran near 30 percent in the first year and fell to 8 percent by the last, with the most common fault recorded as inadequate side ditches. Farmsteads gained access in an order that tracked the contractors' convenience rather than the program's published priority list, and the two districts with the highest completion also show the largest number of complaints preserved in the ledgers from farmers who had been passed over. The program was wound up in 1955 when responsibility for local roads passed to the newly created municipal councils."
},
{
"query": "chain migration shaped settlement patterns",
"passage": "Parish registers and factory employment cards from the industrial district of Marlowick show that migrants arriving between 1902 and 1914 clustered overwhelmingly by village of origin, with entire streets in the district's eastern quarter populated almost exclusively by families originating from a cluster of six villages some ninety miles to the south. Factory hiring records indicate that this clustering was not incidental but actively reproduced through workplace recommendation, since foremen at the district's largest textile mill routinely hired new workers on the recommendation of an already-employed relative or fellow villager, a practice mill management appears to have tolerated because it reduced the administrative burden of screening unfamiliar applicants. Lodging records kept by several boardinghouse keepers in the same quarter show a similar pattern, with newly arrived migrants typically lodging with an established household from their home village for their first several months before securing independent accommodation, often in the same street. This settlement pattern had consequences beyond housing, since mutual aid societies formed in the district during this period drew membership almost entirely along village-of-origin lines rather than uniting migrants by trade or parish more broadly, and disputes between rival mutual aid societies over access to a single burial ground in 1911 reveal how deeply these village-based divisions structured social life even after a generation of shared urban residence. An earlier survey of the archive treated the district's migrant population as a relatively undifferentiated mass defined chiefly by rural origin in general, but the finer-grained record of village-specific clustering suggests that the relevant social unit for understanding settlement in Marlowick was the specific sending village, whose internal networks migrants appear to have carried with them largely intact.",
"nearMiss": "Sickness benefit records of a miners' friendly society in a coalfield town survive for 1858 to 1877 and record for each claim the member's name, occupation, the number of days paid and, from 1863, the clerk's note of the cause, allowing the course of working-class morbidity to be followed across two decades in a single industrial community. Membership grew from 340 to 910 over the period, and the annual rate of claims fluctuated between 21 and 34 per hundred members with peaks in the winters of 1864 and 1871, both years in which the town's medical officer reported epidemic respiratory illness. Injury claims made up a third of the total and averaged 23 days paid, against 15 for illness, and hewers claimed for injury at nearly twice the rate of surface workers. The society's rules limited full benefit to 26 weeks, and 41 members exhausted it over the period, most of them older men with chest complaints who then appear on the reduced allowance the rules provided for chronic cases. The clerk's marginal notes record eleven members struck off for claiming while seen at work, and the society's accounts show that its reserve fund fell below the level its actuary had advised in six of the twenty years."
},
{
"query": "Did guild monopolies raise prices for consumers in early modern towns?",
"passage": "Price records maintained by the wardens of a cutlers' guild in a provincial English town between 1580 and 1640 allow a direct comparison between guild-regulated prices charged within the town and prices charged by cutlers working in nearby villages outside the guild's jurisdiction, a comparison complicated by the fact that surviving village price records are sparse and unevenly dated. Where comparable entries exist, guild-regulated knives of standard quality sold within the town at prices roughly fifteen to twenty percent above those recorded for comparable knives sold in villages beyond the guild's reach, a gap that held reasonably consistent across the six decades covered by the records despite considerable fluctuation in raw material costs over the same period. The guild's own justification, recorded in wardens' minutes, held that regulated pricing protected consumers from shoddy workmanship by excluding untrained producers, and the minutes do record a handful of prosecutions against town cutlers for selling defective blades, suggesting the guild's quality enforcement was not purely nominal. Whether this justification fully accounts for the price gap is less clear, since the same minutes also record repeated efforts to restrict the number of new masters admitted to the guild, a restriction with an obvious effect on supply independent of any quality rationale, and admission numbers did in fact decline noticeably after 1610 even as the town's population continued to grow. Complaints preserved in the town's court records, filed by residents accusing the guild of artificially restricting supply, appear with some regularity after 1615, though the court's rulings in these cases mostly upheld the guild's exclusive rights, leaving the two effects, quality assurance and monopoly rent, thoroughly entangled in the surviving evidence.",
"nearMiss": "In 1604 the corporation of a small English cathedral town resolved to bring water from a spring some two and a half miles distant by means of a bored elm conduit, and the chamberlain's accounts for the following fourteen years record how the work was financed and how badly the original estimate held. The undertaking was funded by a loan of 400 pounds from three aldermen at 8 percent, to be repaid from a rate levied on the households that took a private supply, and the accounts show that fewer than half the projected sixty subscribers had connected by 1610, leaving the corporation to meet the interest from its general revenue for most of the period. Pipe replacement proved the largest recurring cost, with sections of the conduit failing every winter where it crossed a stretch of waterlogged meadow, and in 1612 the corporation paid a plumber from a neighboring city to relay that stretch in lead, at a cost that exceeded the whole of the original estimate for the timber pipe. The public cistern in the market place, the ostensible purpose of the scheme, was not completed until 1618, and the loan was finally discharged in 1621 through a sale of corporation land."
},
{
"query": "multiple narrators device",
"passage": "The novel's structure relies on letters attributed to five separate correspondents, a device that allows the same disputed event, a broken engagement recounted differently by the jilted suitor, the bride's mother, and a servant who claims to have witnessed the final confrontation, to be presented three times with materially inconsistent details, and the novel never resolves which account, if any, is fully reliable. This refusal to adjudicate among competing narrators departs from the more common practice in contemporary epistolary fiction of using a single dominant correspondent whose letters the reader is implicitly encouraged to trust, and early reviewers found the technique disorienting, with one periodical complaining that the reader was left 'without any settled ground on which to stand.' The servant's letters are of particular interest, since they are the only correspondence in the novel written by a character of markedly lower social station, and their prose style, looser in syntax and more colloquial in diction than the letters of the gentry characters, has led some readers to treat them as a deliberate authorial experiment in representing class-inflected voice rather than a mere stylistic inconsistency. Manuscript drafts held among the author's surviving papers show that the servant's letters were added in a later revision, replacing a shorter framing device in which the same information had been conveyed through the mother's letters alone, a change that considerably increases the narrative's structural complexity and its ambiguity regarding truth. Whether this revision was motivated by aesthetic considerations or by a publisher's request for a longer manuscript cannot be determined from the surviving correspondence between author and publisher, which addresses length but not content.",
"nearMiss": "The serialization of a three-volume novel in twenty monthly parts between 1864 and 1866 can be followed in the publisher's surviving ledger, which records the sales of each part alongside the author's letters negotiating where each installment should end, and the two sources together show how the demands of the format shaped the finished book. Sales of the first three parts held near 14,000 copies, fell to 9,800 by the sixth, and recovered to over 12,000 after the eighth, the installment in which the author, at the publisher's urging, moved the discovery of the forged will forward from the position it occupied in the manuscript plan. Of the twenty installment endings, fourteen fall on an unresolved event, and the author's letters show him revising six of these after the publisher objected that an installment closed too quietly. The chapter divisions in the volume edition preserve all twenty installment breaks, but the author removed the recapitulating paragraphs that had opened eleven of the parts, and two chapters written to fill installments that had run short were cut entirely. Reviewers of the volume edition complained of an uneven pace in the middle third, which corresponds exactly to the parts published during the sales decline."
},
{
"query": "ceramic cargo revealed trade routes",
"passage": "Excavation of a wreck site off the coast near Cape Ferrando, conducted over three seasons beginning in 1988, recovered an estimated fourteen hundred ceramic vessels from the ship's hold, the majority coarse storage amphorae but including a smaller quantity of finer tableware whose distinctive glaze and rim profile match production centers documented along a stretch of coastline several hundred miles to the east. This combination of cargo, bulk storage vessels alongside a modest consignment of luxury tableware, suggests a vessel engaged in mixed trade rather than the specialized bulk transport more commonly associated with amphora-laden wrecks from the same broad period. Stamped handles on a subset of the amphorae, roughly one in eight, carry a workshop mark previously attested at only two other sites, both considerably closer to the presumed production region, extending the known distribution of that workshop's output substantially further along the trade route than earlier evidence had indicated. Sediment analysis of residue preserved inside several intact amphorae identified traces consistent with preserved fish products rather than wine or oil, the two contents most frequently assumed for amphora cargoes of this general type, a finding that has prompted a reassessment of how readily amphora contents can be inferred from vessel shape alone without direct residue analysis. The ship's timber, sampled from surviving hull fragments, was identified as a pine species native to a region distinct from both the cargo's presumed origin and its likely destination, indicating that the vessel itself may have been built in a third location entirely, a detail that complicates any straightforward narrative of a single bilateral trade route.",
"nearMiss": "Sediment cores and repeated photomosaics taken across a late Roman wreck lying at 38 meters were used over five years to understand why the forward third of the hull survives in good condition while the stern has all but vanished, a question with direct bearing on how the site should be managed. The forward section lies beneath 40 to 70 centimeters of fine, poorly oxygenated silt whose pore water was found to be anoxic below the top 5 centimeters, and wood samples recovered from it showed no borer galleries and only shallow bacterial degradation of the outer millimeter. The stern lies on a coarser, mobile sand sheet that the photomosaics show shifting by up to 30 centimeters between winter and summer surveys, exposing timbers seasonally, and every exposed timber sampled was riddled with shipworm tunnels to a depth of 4 centimeters. Current meters deployed for one year recorded near-bottom velocities exceeding 25 centimeters per second on eleven occasions, all during winter storms, coinciding with the largest changes in sand cover. Trial reburial of one exposed stern timber under sandbags and a geotextile mat halted borer activity over the following two years, and the site plan now recommends covering the whole stern rather than excavating it."
},
{
"query": "How did prize courts determine ownership of captured cargo during wartime?",
"passage": "The proceedings recorded before the vice-admiralty prize court sitting at a colonial port in 1747 concerned a merchant vessel captured while sailing under neutral flag but carrying cargo manifests that the capturing privateer's counsel argued were fraudulently altered to disguise enemy ownership of the goods. The court's task, as the presiding judge framed it in his opening remarks, was not to determine the vessel's own national character, which the neutral flag and ship's papers established without serious dispute, but to determine whether the cargo itself, considered separately from the ship carrying it, belonged beneficially to a subject of the enemy power, since prize law of the period treated ship and cargo as potentially subject to different rulings even when travelling together. Depositions taken from the ship's supercargo and from a merchant correspondent in the port of origin diverged sharply on the question of who held ultimate title to the goods at the moment of capture, with the supercargo insisting the cargo had been purchased outright by a neutral merchant house while the correspondent's deposition suggested the neutral house acted merely as a forwarding agent for an enemy principal, receiving a fixed commission rather than bearing ownership risk. The court gave considerable weight to a set of insurance papers entered into evidence, reasoning that the identity of the party who had insured the cargo, and who therefore bore its risk of loss, offered stronger evidence of beneficial ownership than the bill of sale alone, which the court noted could be, and in other cases had been shown to be, drawn up after the fact to disguise true ownership. On this basis the court condemned the cargo as enemy property subject to confiscation while releasing the vessel itself.",
"nearMiss": "Accounts kept by the storekeeper of a naval dockyard on the Hampshire coast between 1738 and 1752 record in detail the arrival, measurement and seasoning of oak and fir timber drawn from three supply routes: coppice woodland within thirty miles of the yard, rafted deliveries down a navigable river from inland estates, and imports of fir masts from the Baltic. Oak arriving overland cost the yard about a third more per load than river-borne timber but was consistently rated higher for quality, with only 6 percent rejected at measurement against 19 percent for the river deliveries, which the storekeeper repeatedly attributed to damage from prolonged immersion. Seasoning times in the yard's timber ponds and open stacks averaged four years for compass timber, and the accounts show the yard's stock falling below the two-year reserve the Navy Board required in 1744 and again in 1748, both years in which building programs were accelerated. Baltic fir arrived in only seven of the fifteen years, always in the autumn, and the accounts record that in the years without deliveries the yard resorted to jointing shorter native fir for smaller masts. The storekeeper's tallies of loads received differ from the contractors' invoices in most years by a small margin that the auditors accepted without comment."
},
{
"query": "reductionism about testimonial justification",
"passage": "A recurring disagreement in the epistemology of testimony concerns whether a hearer's justification for believing what a speaker tells her can be reduced, without remainder, to justification the hearer already possesses independently for trusting that speaker or that type of report, a position generally labeled reductionism about testimonial justification, or whether testimony instead confers justification in a manner not fully explicable by appeal to independent, non-testimonial evidence. The reductionist position, in its more demanding form, requires that a hearer possess some positive inductive basis, however minimal, for regarding a given speaker or type of testimony as reliable before that testimony can justify belief, a requirement critics have argued sets an implausibly high bar given how little independent evidence most people actually possess about the reliability of the vast majority of speakers whose testimony they nonetheless accept without hesitation. A weaker reductionist position allows the required inductive basis to be extremely general, amounting to little more than a background assumption that speakers tend to be more reliable than chance, but this weakening has struck some critics as draining the reductionist thesis of much of its original content, since almost any anti-reductionist could accept so minimal a background condition without abandoning the claim that testimony confers justification in its own right. Anti-reductionists typically appeal to the situation of young children, who plainly acquire justified beliefs from testimony well before they could possess any inductive track record concerning speaker reliability, as evidence that testimonial justification cannot uniformly depend on prior, independently acquired grounds for trust. Reductionists have replied that the child's case may be disanalogous to adult testimonial exchange in ways that limit its evidential force.",
"nearMiss": "Two hundred and eighty undergraduates were presented with paired vignettes in which two drivers behave identically, both briefly checking a phone at the wheel, but only one strikes a child who steps into the road, and were asked to rate how much blame, punishment and moral badness each driver deserves on nine-point scales. Ratings of blame and punishment were sharply higher for the driver who caused harm, by more than three scale points on average, while ratings of how bad a person each driver was differed by less than half a point, a split that held across three further vignette pairs involving a surgeon, a hunter and a babysitter. When participants were asked to explain their ratings, most who assigned more blame to the unlucky driver nonetheless agreed, when asked directly, that the two had done the same thing and had been equally careless, and a quarter revised their blame ratings downward on being shown their own earlier answer. Participants who scored high on a measure of belief in a just world showed the largest gap between the two drivers and were least likely to revise. The authors read the pattern as evidence that ordinary judgment separates assessment of the agent from assessment of the outcome, and that the influence of outcomes on blame is partly a response to the harm rather than a considered verdict on the person."
}
],
"zh": [
{
"query": "诱饵摆动频率",
"passage": "在2019年至2021年间,对采自马里亚纳海沟西侧约850米至1200米水深处的47尾角鮟鱇雌性个体进行了解剖学与微生物学联合分析,其中28尾采自850至1000米水层,19尾采自1000至1200米水层,以比较不同深度种群诱饵器官结构的差异。诱饵器官内部的腺体组织切片经革兰氏染色后,在光学显微镜下可见密集排列的弧菌属共生发光菌,菌落密度平均达到每立方毫米2.3乘以10的8次方个细胞,深水层个体的菌落密度略低,约为1.8乘以10的8次方。通过16S rRNA基因测序比对,确认其中约六成个体携带的优势菌株与费氏弧菌关系密切,但存在若干此前未见的碱基替换位点,测序覆盖深度平均为42倍。进一步的行为观察显示,诱饵摆动频率与光强度呈现正相关,平均每分钟摆动12至18次的个体捕获糠虾类猎物的成功率比摆动频率低于8次的个体高出约37%,该效应在两个深度组之间无显著差异。解剖还发现,发光器官周围分布有密集的血管网络与色素细胞层,推测其兼具遮光与调控光色的功能,光谱仪测量显示发光峰值波长集中在478至483纳米之间。研究人员认为,这种共生关系并非单纯的营养交换,雌鱼可能通过分泌特定黏液成分对菌群密度进行季节性调节,从而在食物匮乏期降低代谢消耗,但受限于样本量,这一推论仍需更大规模的跨季节采样加以验证。",
"nearMiss": "研究人员于2021年夏季在南海北部冷泉区约1100米水深处,利用遥控潜水器采集了三十九只深海海参,并同步采集其所在位置的表层沉积物,以分析这类动物对沉积物粒径的摄食选择。肠道内容物经筛分后与环境沉积物比较,结果显示海参肠道中粒径小于63微米的细颗粒占比达到71%,明显高于环境沉积物中的48%,而粗砂组分几乎不见于肠道。肠道内容物的有机碳含量平均为环境沉积物的2.3倍,提示其摄食时对富含有机质的细颗粒有主动选择。不同体长个体之间的选择强度差异不显著。研究人员还比较了两个采样季节的数据,发现秋季个体肠道中细颗粒占比略高于夏季,推测与该季节沉积物有机质输入增加有关。基于摄食速率的估算表明,这一海参种群每年可翻动其分布区表层约一点五厘米厚的沉积物。"
},
{
"query": "夯土墙体保温隔热性能",
"passage": "选取福建省南靖县田螺坑周边五座建于清代中后期的圆形土楼作为测量对象,分别在2022年7月与2023年1月两个季度各布设三十六个温湿度记录仪,采样间隔设定为十分钟,另在其中两座土楼的一层与三层加装了八个热流传感器以获取墙体热流密度数据。数据显示,厚度普遍在1.2米至1.6米之间的夯土外墙使楼内二层房间的日温差被压缩至4.3摄氏度以内,而同期室外日温差可达11.8摄氏度,一层房间由于地面蓄热效应,日温差进一步压缩至3.1摄氏度。红外热像仪扫描结果表明,墙体夯土层中掺入的糯米浆与竹筋网络在墙体内部形成了多处热阻不连续区域,这些区域的表面温度比周边墙面低约1.5至2摄氏度,热流传感器数据显示该区域热流密度波动幅度达到周边墙面的1.6倍。冬季测量还发现,中央天井的烟囱效应使得二层通风廊道内的空气流速平均达到每秒0.35米,显著高于底层的0.12米,且该流速在正午前后达到峰值。研究人员据此推算,若将现代保温材料简单叠加于夯土墙外侧,反而可能破坏原有的湿度调节机制,导致墙体内部结露风险上升,团队建议后续研究应延长监测周期至完整年度,以排除单一季节测量带来的偏差。",
"nearMiss": "2022年秋季,研究组对安徽黟县两处村落中十四座清代民居的马头墙进行了测绘与病害记录,重点考察叠落式山墙的砌筑构造与屋面交接处的防水做法。测绘显示,马头墙顶部普遍采用三层小青瓦覆盖,瓦下以石灰砂浆找坡,墀头部位挑出的青砖以两层或三层为主,挑出距离多在十八至二十四厘米之间。病害调查记录到的主要问题是瓦面松动导致墙顶渗水,进而引起下方石灰粉刷层剥落,十四座民居中有九座存在此类情况,且集中于屋面坡度较缓的建筑。对三处剥落部位取样分析表明,原粉刷层以石灰为主并掺有少量麻刀,后期修补则多用水泥砂浆,二者收缩性能差异是修补处再次开裂的主要原因。研究组据此建议在修缮中恢复原有材料与瓦下找坡层,而非简单更换瓦片。"
},
{
"query": "学区房溢价为什么下降?",
"passage": "以成都市三环至绕城高速之间的十二个住宅片区为样本,收集了2020年1月至2023年6月期间共计四万一千余套二手住宅的成交记录,并按房龄、朝向与楼层进行了分层控制,以降低样本结构差异对回归结果的干扰。回归分析显示,地铁站步行距离每增加一百米,单价平均下降约186元每平方米,而这一效应在建成年代超过二十年的老旧小区中被削弱近四成,研究者推测老旧小区购房者对通勤便利性的敏感度相对较低。挂牌到成交的平均周期从2021年的58天延长至2023年上半年的94天,同期议价空间的中位数由3.2%扩大到6.7%,其中改善型住房的议价空间扩大幅度明显高于刚需型住房。分片区来看,天府新区周边次新房的价格波动幅度明显小于二环内老城区,后者的价格标准差是前者的近两倍,研究者认为这与老城区业主结构中高龄自住群体占比较高、议价耐性更强有关。此外,学区资格与房源价格的关联性在样本期内呈现逐年减弱趋势,某重点小学对应片区的房价溢价率从18%降至11%,与近年学区房政策调整的时间点大致吻合,但研究者同时指出,样本未覆盖2023年下半年后续政策变化,结论的时效性有一定局限。",
"nearMiss": "课题组于2022年3月至2023年5月对西安市两个中心城区六十四个建于1990年代的多层住宅小区开展了加装电梯意愿调查,共回收有效问卷两千三百一十份,并对已完成加装的十九个单元核查了造价与工期。调查显示,三层及以上住户的支持率为81%,一层与二层住户分别为12%与29%,反对理由中采光受遮挡与通行噪声最为集中。已完成项目的平均造价为六十八万元,政府补贴约占三成,其余按楼层系数分摊,平均工期为四个半月。课题组发现,由业主委员会而非物业公司牵头的项目从意向到开工的周期短约两个月。"
},
{
"query": "手势透明度评分",
"passage": "研究团队邀请了三十二名来自不同省份的聋人手语使用者,对中国手语词典中收录的两百个具象名词手势进行了透明度评分实验,受试者平均手语使用年限为十九年,均以手语作为日常主要交流方式。评分采用七级量表,由未接触过手语的听人被试根据手势形态猜测词义,再与实际词义比对计算命中率,听人被试共计六十名,平均年龄二十六岁。结果显示,表示动物类词汇的手势平均命中率达到54%,显著高于表示抽象情感类词汇的19%,而表示日常工具类词汇的命中率居中,约为38%。进一步的手部运动捕捉数据分析发现,高透明度手势往往伴随更大幅度的手部空间位移,平均位移距离为28厘米,而低透明度手势的平均位移仅为9厘米,两组之间的手部运动速度差异同样显著。研究者还比较了北方与南方手语社群中同一词汇的手势变体,发现约四分之一的具象词汇存在地域性形态差异,但这些差异并未显著影响听人被试的猜测准确率。研究认为,象似性程度与手势的社会规约化程度之间可能存在此消彼长的关系,越常用的词汇其手势越趋于简化和抽象化,团队计划在后续研究中扩大方言手语社群的取样范围以验证这一假设的普遍性。",
"nearMiss": "研究组于2022年在北京、成都与广州三地各招募二十名以手语为第一语言的聋人,利用同一套一百二十幅图片进行词汇诱导,以考察三地手语在日常词汇上的地域差异。结果显示,三地在食物、交通与亲属称谓三类词汇上的差异最为明显,其中食物类词汇有近四成在三地使用完全不同的手形,而数字与颜色类词汇的一致率超过九成。年龄在五十岁以上的受访者保留地方变体的比例明显高于年轻受访者,后者更多使用电视新闻手语中的通用形式。研究组认为,聋校集中办学与网络视频的普及是词汇趋同的主要推动因素。此外,三地受访者对同一图片给出两种以上手形的比例平均为一成二,以广州最高。"
},
{
"query": "氧化石墨烯膜层间距调控",
"passage": "实验采用改进的哈默斯法制备氧化石墨烯片层,通过真空抽滤在聚醚砜基底上组装出厚度约180纳米的层状膜,层间距经X射线衍射测定为0.83纳米,制膜过程中真空度控制在0.08兆帕以确保片层排列均匀。将该膜应用于模拟苦咸水淡化测试,进水氯化钠浓度为2000毫克每升,在1.5兆帕操作压力下,渗透通量稳定在每平方米每小时18.6升,截盐率达到96.2%,测试期间进水温度维持在25摄氏度左右。连续运行720小时后,通量下降至初始值的78%,断面扫描电镜显示层间出现轻微塌陷,间距缩小至0.71纳米,能谱分析未见明显的元素污染沉积。研究人员尝试用戊二醛对层间进行化学交联处理,交联后的膜在相同测试条件下运行720小时后通量保持率提升至91%,但初始通量略微下降约6%,交联时间延长至十二小时以上时通量下降幅度进一步增大。此外,膜表面接枝聚乙二醇链段后,对腐殖酸的抗污染能力明显增强,污染实验中通量恢复率从交联前的65%提升至89%,研究团队指出该膜在实际苦咸水体系中面对复杂离子组成时的长期稳定性仍需进一步验证。",
"nearMiss": "某市政污水处理厂的膜生物反应器在运行第三年出现跨膜压差快速上升的问题,运行方与研究组合作,在2022年3月至11月间对三组并联的聚偏氟乙烯中空纤维膜组件分别施加每平方米每小时5、8与12立方米的曝气强度,考察曝气对膜面生物污染的抑制效果。跨膜压差记录显示,低曝气组在四十天内由12千帕升至35千帕并触发化学清洗,中曝气组维持了七十二天,高曝气组在整个试验期内未超过28千帕。取膜丝表面污染层进行测序发现,低曝气组以丝状菌为主,而高曝气组的污染层更薄且以絮体状细菌为主。能耗核算表明,曝气强度由8提高到12时,单位处理水量的电耗增加约百分之十九,而清洗频率仅由每年五次降至四次,运行方最终选择了中曝气方案。"
},
{
"query": "漕运总督衙门批文",
"passage": "依据清代漕运总督衙门存档的批文与《漕运则例》相关条目统计,乾隆二十五年至三十年间,经由京杭大运河北上的漕粮总量年均约为三百二十万石,其中江苏、浙江两省承担的份额合计超过六成,安徽与江西两省合计承担约两成五。档案记载,每年漕船自扬州仪征一带集结北上,途经淮安、临清等六处漕运重镇,全程设有四十余处闸坝用于调节水位,漕船编队通常按十艘一组分批过闸。因黄河夺淮导致河道淤积,乾隆二十八年的漕运记录显示,山东段运河水位一度低于正常通航标准三尺有余,当年有近两百艘漕船滞留待闸,延误时间最长者达四十五天,部分船只因此错过预定抵通州的验收期限。为应对淤积问题,官府在当年冬季征调民夫逾三万人次疏浚河道,耗银约十八万两,疏浚工程集中在临清至东昌一段约六十里河道。此外,漕粮运输过程中的损耗率也被详细记录在案,平均每万石漕粮途中损耗约二百三十石,损耗原因多归于船只渗漏与仓储受潮,地方官员因此被要求按月呈报仓储湿度情况,现存档案中此类月报仅完整保存至乾隆二十九年。",
"nearMiss": "嘉庆朝二十五年间顺天乡试共中式举人两千七百余人,其籍贯构成可据现存题名录逐科统计。直隶本省中式者约占五成八,其中顺天府所属州县又占本省的四成以上;其余为寄籍顺天应试的各省士子,以山西与山东两省居多,合计超过两成。就出身而言,监生中式的比例由嘉庆初年的两成二升至嘉庆末年的三成四,与同期国子监捐监人数的增长趋势一致。题名录所载年龄显示,中式者平均约三十岁,四十岁以上者占一成三,最年长者六十一岁。将各科数据与《清实录》所载对冒籍的查处记载对照可见,凡上一科有查处记录,下一科寄籍中式比例即有所回落,但两科之后又恢复原有水平。"
},
{
"query": "购物决策疲劳如何影响选择行为?",
"passage": "实验招募了一百六十名在线购物平台的活跃用户,要求其在受控环境下完成连续四十五分钟的模拟选购任务,任务中需要在每组八到十二个选项中做出购买决策,共完成三十组决策,被试年龄分布在二十二岁至四十五岁之间,男女比例大致相当。研究通过瞳孔追踪与反应时记录评估决策疲劳程度,结果显示,任务进行到第二十组前后,被试的平均决策反应时间从最初的6.8秒延长至11.4秒,选择默认推荐选项的比例也从17%上升至43%,瞳孔直径的基线波动幅度同期缩小约22%。进一步分析发现,在决策后期,被试对商品评价数量的关注权重明显下降,而对价格排序位置的依赖程度显著上升,倾向于直接选择价格排序中靠前的选项,这一趋势在女性被试中表现得更为明显。研究还设置了中途插入五分钟休息的对照组,结果该组在后续任务中的反应时间回落至8.1秒,默认选项选择比例也降至28%,表明短暂休息能够部分缓解决策疲劳带来的认知资源损耗。研究者指出,由于实验任务为模拟情境,被试未使用真实资金支付,实际网购场景中的决策疲劳程度可能与本实验结果存在一定偏差。",
"nearMiss": "研究选取某综合电商平台2023年第二季度的两千四百场服装类直播录像,结合平台提供的匿名观看日志,分析主播语速与观众停留时长之间的关联,日志覆盖观众一百九十万人次。语音转写后的统计显示,主播平均语速为每分钟二百四十八字,语速位于二百二十至二百六十字区间的场次,观众停留时长中位数为一百零六秒,明显高于语速超过三百字的场次的七十一秒。进一步分析发现,语速过快的场次中观众在主播报价环节的流失最为集中,而语速适中的场次在该环节反而出现停留时长的小幅上升。研究建议平台在主播培训中加入语速反馈。"
},
{
"query": "第三染色体主效位点",
"passage": "选育团队自2016年起,以来自黄淮海地区的两百一十四份玉米地方品种为材料,在河南新乡与内蒙古通辽两个干旱胁迫试验点连续种植六个生长季,通过控水处理模拟拔节期至抽雄期的阶段性干旱,两试验点的年均降水量分别约为580毫米与360毫米。田间考种数据显示,入选的十七份抗旱材料在干旱处理下的根系深度平均达到112厘米,比对照品种深出约34厘米,且次生根数量增加近五成,根系性状的测定采用挖掘法结合数字图像分析完成。进一步的基因型分析定位到位于第三染色体上的一个主效数量性状位点,该位点附近标记与干旱条件下的产量保持率呈显著相关,携带优异等位基因的材料在干旱年份的减产幅度控制在12%以内,而不携带该等位基因的材料减产普遍超过30%,该位点在两个试验点间的效应方向保持一致。2022年在通辽点的区域试验中,以该位点为选择标记聚合选育出的新品系较对照品种增产9.6%,且千粒重未见明显下降,目前已提交省级区域试验申请,后续还需完成至少两年多点试验以满足审定要求。",
"nearMiss": "试验于2019年至2022年在河北藁城与山东德州两地进行,设置三个播期与四个种植密度的裂区处理,供试品种为当地主推的中熟夏玉米,小区面积三十平方米,重复三次。结果显示,6月12日播种处理的籽粒灌浆持续期平均为四十一天,较6月26日播种处理延长六天,灌浆速率峰值出现在授粉后十八至二十二天。种植密度由每公顷六万株增至九万株时,单株粒重下降约两成,但群体产量提高11%,再增至十万五千株时产量不再增加而倒伏率明显上升。两地三年的数据均支持早播配合每公顷九万株的组合。"
},
{
"query": "热木星大气水汽吸收特征",
"passage": "利用架设于智利阿塔卡马高原的地面望远镜,研究团队于2023年三月至六月间对编号为TOI-2109b的热木星型系外行星进行了七次凌星光谱观测,每次观测覆盖波长范围为0.6至1.7微米,单次观测时长平均为四小时二十分钟。透射光谱分析显示,在1.4微米附近存在明显的水汽吸收特征,吸收深度约为320个百万分之一,据此推算该行星大气中水汽的体积混合比约为百万分之五十,该数值与此前基于空间望远镜的初步估计基本吻合。此外,光谱在0.76微米处观测到钠元素的双峰吸收线,谱线宽度暗示行星大气层顶部风速可能达到每秒四公里量级,谱线不对称性进一步提示昼夜半球间存在明显的大气环流。团队还结合此前的次食观测数据,估算该行星白昼半球平均温度约为2200开尔文,昼夜温差可能超过800开尔文,这一结果支持该行星大气环流受潮汐锁定影响较强的假设。研究人员指出,由于该行星母恒星活动性较高,七次观测中有两次因耀斑事件被剔除,剩余五次观测的信噪比也存在一定波动,后续仍需更多凌星数据以降低测量不确定性。",
"nearMiss": "一颗V星等8.6的K型矮星在视向速度数据中呈现出周期分别为16.3天与41.7天的两个显著信号,虚警概率均低于千分之一,对应的半振幅分别为每秒4.2米与3.1米。这些数据来自青海冷湖台址2.5米望远镜所配高分辨率光谱仪自2021年9月至2023年2月的一百七十二次有效测量,曝光时间依天气条件在一刻钟到半小时之间调整,仪器零点的长期漂移优于每秒一米。为排除恒星活动的干扰,团队同时分析了钙线活动指数与谱线不对称度指标,二者均未在上述两个周期附近出现信号,而恒星自转周期由光变数据确定为约二十九天,与两个信号均不重合。若两个信号均由行星引起,其最小质量分别约为地球的十一倍与十四倍,轨道均位于宜居带之内。"
},
{
"query": "积分入学资格",
"passage": "课题组于2021年至2022年对东莞市六个镇街的随迁子女入学情况进行了追踪调查,共回收有效问卷一千三百四十份,并对四十二个家庭进行了深度访谈,访谈对象涵盖制造业、建筑业与家政服务业三类主要务工群体。数据显示,持有积分入学资格的随迁子女中,约有61%最终被分配至公办学校,而未达到积分门槛的家庭中,超过七成子女进入民办及非正规教育机构就读,部分非正规机构未取得正式办学资质。访谈发现,积分不足的家庭平均需要额外承担每年约六千元的择校及交通支出,部分家庭因此选择将子女送回原籍由祖辈照料,样本中此类留守安排占比达到23%,留守时长平均为一年零四个月。进一步分析显示,父母社保连续缴纳年限是影响积分高低的关键变量,而这一变量与家庭所从事行业的稳定性密切相关,从事建筑与家政行业的家长因工作流动性较大,社保断缴现象更为普遍,断缴次数在三次以上的家庭积分达标率不足两成。研究认为,现行积分入学制度在客观上加剧了不同职业背景随迁家庭子女教育机会的分化,课题组建议进一步追踪这批儿童升入初中阶段后的教育路径分化情况。",
"nearMiss": "上海市中心城区六十五岁以上老年人中,过去一个月内使用过社区助餐点的比例为34%,独居老年人的使用比例达到52%,而与子女同住者仅为19%。这一结果来自2023年4月至9月在三个区二十六个街道进行的问卷调查,有效样本一千八百六十份。未使用助餐点的老年人中,最常提到的原因是助餐点距离住所超过十分钟步行路程,占四成一;其次是对菜品口味不满,占两成六;认为价格偏高者不足一成。十二个助餐点运营方的访谈则集中反映补贴结算周期过长的问题,平均结算周期为四个月,导致日均供餐量超过两百份的助餐点普遍出现流动资金紧张。研究者由此提出应按老年人口密度而非行政区划来布局助餐点,并缩短补贴结算周期。"
}
],
"other": [
{
"query": "Quelle résine a été identifiée dans la couche de vernis intermédiaire ?",
"passage": "La restauration du retable de l'ancienne chapelle des Ursulines de Nancy, entreprise entre septembre 2021 et mars 2022, a mobilisé une équipe de quatre restauratrices spécialisées dans la polychromie sur bois. L'examen préalable sous lumière ultraviolette a révélé la présence d'au moins trois couches de vernis successives, la plus ancienne datant probablement de la fin du dix-septième siècle et présentant un jaunissement prononcé attribuable à l'oxydation des résines naturelles. Les prélèvements microscopiques effectués sur douze points distincts du panneau central ont montré une épaisseur moyenne de vernis de 45 microns, avec des variations locales atteignant 80 microns dans les zones d'ombre où les restaurateurs du dix-neuvième siècle avaient appliqué des repeints. L'analyse par spectroscopie infrarouge à transformée de Fourier a permis d'identifier une résine dammar mélangée à de l'huile de lin dans la couche intermédiaire, ce qui complique singulièrement le décapage sans endommager la couche picturale originale sous-jacente. Le protocole retenu a consisté en un allègement progressif au moyen de gels solvants à base d'éthanol et d'acétone en proportions variables, appliqués par petites zones de trois centimètres carrés sous contrôle binoculaire. Ce travail, réparti sur environ six cents heures cumulées, a permis de dégager la carnation originale des figures du panneau représentant sainte Ursule, dont les teintes rosées s'étaient auparavant confondues avec le vernis oxydé. Les radiographies prises avant intervention ont également révélé un repentir important au niveau de la main gauche de la sainte, initialement positionnée plus bas de sept centimètres. Une fois le nettoyage achevé, les lacunes de la couche picturale, représentant environ 4% de la surface totale du panneau central, ont été comblées avec un mastic à base de craie et de colle de peau, puis retouchées en glacis réversibles selon la méthode dite du tratteggio. Le rapport final souligne que l'humidité relative de la chapelle, mesurée à 68% en moyenne durant les mois d'hiver, reste préoccupante pour la conservation à long terme et recommande l'installation d'un système de régulation climatique avant toute nouvelle exposition permanente de l'œuvre.",
"nearMiss": "L'analyse des colorants d'une tapisserie d'Aubusson datée par son cartouche de 1712 et conservée dans un château du Limousin a été menée sur quarante-trois micro-prélèvements de fils de laine et de soie, traités par extraction acide puis analysés par chromatographie liquide couplée à un détecteur à barrette de diodes. Les rouges se sont révélés obtenus à partir de garance pour les laines et de cochenille pour les rares fils de soie, tandis que les jaunes, très altérés, provenaient de la gaude, dont le flavonoïde principal n'a été détecté qu'à l'état de traces dans les zones exposées à la lumière. Les bleus et les verts, mieux conservés, contenaient de l'indigotine, les verts résultant d'une double teinture par la gaude. Trois fils de laine rouge orangé situés dans la bordure ont livré un profil différent, dominé par un colorant synthétique, ce qui permet d'identifier une restauration postérieure à 1870 que les archives du château ne mentionnent pas. La comparaison avec les échantillons de référence teints à l'atelier selon les recettes du dix-huitième siècle montre que la perte des jaunes explique à elle seule le virage général des feuillages vers le bleu observé aujourd'hui."
},
{
"query": "Delamination Holmgurt",
"passage": "Am Prüfstand für Rotorblätter der Technischen Hochschule in Bremerhaven wurden zwischen Februar und November 2022 insgesamt sechs Rotorblätter des Typs mit 63 Metern Länge einer beschleunigten Ermüdungsprüfung unterzogen. Die Blätter stammten aus einer Serienproduktion, die ursprünglich für Offshore-Anlagen in der Nordsee vorgesehen war, jedoch aufgrund geringfügiger Abweichungen in der Harzaushärtung nicht zur Auslieferung freigegeben wurde. Die Prüfeinrichtung brachte über hydraulische Erreger eine wechselnde Biegelast auf, wobei die Lastspielzahl pro Blatt auf zehn Millionen Zyklen festgelegt wurde, was rechnerisch etwa fünfundzwanzig Betriebsjahren entspricht. Dehnungsmessstreifen an insgesamt vierzig Positionen entlang der Blattschale zeichneten die Verformung mit einer Abtastrate von zweihundert Hertz auf. Bei vier der sechs Blätter traten erste Delaminationen im Bereich des Holmgurts bereits nach etwa 6,2 Millionen Lastspielen auf, deutlich früher als die in der Auslegung angenommenen 8,5 Millionen Zyklen. Die anschließende Computertomographie der betroffenen Abschnitte zeigte Hohlräume von bis zu drei Millimetern Durchmesser zwischen den Kohlefaserlagen, die auf lokale Harzarmut während der Fertigung zurückgeführt wurden. Interessanterweise blieben die beiden Blätter, die aus einer späteren Fertigungscharge mit angepasstem Vakuuminfusionsdruck stammten, bis zum Erreichen der vollen zehn Millionen Zyklen nahezu schadensfrei, mit einer maximal gemessenen Steifigkeitsabnahme von lediglich 3 Prozent gegenüber dem Ausgangszustand. Die Ingenieure schlossen daraus, dass der Infusionsdruck während der Fertigung einen stärkeren Einfluss auf die Ermüdungslebensdauer hat als die zuvor als kritisch angesehene Harzrezeptur selbst. Für künftige Serien wurde empfohlen, den Infusionsdruck kontinuierlich mit Drucksensoren im Formwerkzeug zu überwachen und Chargen unterhalb eines Schwellenwerts von 2,3 bar automatisch zur Nachprüfung freizugeben, bevor sie in Serienblätter verbaut werden.",
"nearMiss": "An acht Onshore-Windenergieanlagen der Drei-Megawatt-Klasse in einem Windpark bei Husum wurden zwischen 2020 und 2023 halbjährlich Präzisionsnivellements der Fundamentoberkanten durchgeführt, um die Setzung der Flachgründungen auf den dort anstehenden Kleiböden zu verfolgen. Als Bezug diente ein tiefgegründeter Festpunkt außerhalb des Einflussbereichs der Anlagen. Die gemessenen Gesamtsetzungen lagen nach drei Betriebsjahren zwischen 6 und 19 Millimetern, wobei die drei Anlagen auf der niedrigsten Geländestufe die größten Werte zeigten. Die Schiefstellung der Fundamente blieb bei allen Anlagen unter 1,5 Millimetern pro Meter und damit deutlich innerhalb des zulässigen Bereichs. Die Setzungsverläufe flachten nach dem zweiten Jahr erkennbar ab, was auf ein weitgehendes Abklingen der Konsolidation hindeutet. Ein Vergleich mit den in der Gründungsstatik prognostizierten Werten ergab, dass die Berechnung die Setzungen auf den weichen Böden um etwa ein Drittel unterschätzt hatte. Der Betreiber hat die Messungen in den regulären Wartungsplan aufgenommen."
},
{
"query": "adherencia real a inhaladores en niños",
"passage": "En una clínica pediátrica de referencia ubicada en la localidad de Suba, en Bogotá, se realizó un seguimiento durante catorce meses a ciento dos niños de entre cuatro y once años diagnosticados con asma persistente moderada, con el propósito de evaluar el uso real de inhaladores de corticoide inhalado frente a lo prescrito. Cada dispositivo fue equipado con un sensor electrónico que registraba la fecha y hora de cada activación, sin que las familias conocieran el propósito exacto del registro durante los primeros seis meses, a fin de minimizar el sesgo de comportamiento. Los datos mostraron que la adherencia promedio real fue del 54%, considerablemente inferior al 89% que las mismas familias habían reportado en los cuestionarios trimestrales de autorreporte. El análisis por franjas horarias reveló que las dosis de la noche se omitían con mayor frecuencia, alcanzando una tasa de omisión del 41%, en comparación con el 22% registrado para las dosis matutinas. Los niños cuyos cuidadores principales trabajaban en turnos rotativos presentaron una adherencia significativamente menor, de apenas 43%, frente al 61% observado en hogares con horarios laborales estables. Durante las visitas de seguimiento, se identificó que catorce familias desconocían la técnica correcta de inhalación con cámara espaciadora, lo que se tradujo en una entrega efectiva del fármaco notablemente reducida según las mediciones con dispositivo trazador. Tras una intervención educativa breve, consistente en dos sesiones prácticas de quince minutos con enfermeras capacitadas, la adherencia electrónica promedio del grupo intervenido aumentó a 68% en el mes posterior, mientras que el grupo control sin intervención adicional se mantuvo estable en 52%. Los investigadores concluyeron que las discrepancias entre adherencia autorreportada y adherencia electrónica objetiva son sustanciales en este contexto y recomendaron incorporar sensores de monitoreo como herramienta rutinaria en el seguimiento de asma pediátrica en centros de atención primaria de la ciudad.",
"nearMiss": "Ciento ochenta y seis niños de veinticuatro meses atendidos en el programa de crecimiento y desarrollo de cuatro centros de salud del municipio de Bello, en el área metropolitana de Medellín, fueron evaluados entre agosto de 2021 y julio de 2022 con un cuestionario de vocabulario expresivo respondido por los padres, con el fin de estimar la frecuencia de retraso del lenguaje en la población atendida y la capacidad del instrumento para identificarlo en la consulta habitual. El 14 por ciento de los niños se situó por debajo del percentil diez de vocabulario para su edad, con una proporción mayor entre los varones y entre los hijos de madres con menos de nueve años de escolaridad. Cuarenta y dos de los niños con puntuaciones bajas fueron remitidos a valoración por fonoaudiología, y en treinta y uno de ellos se confirmó un retraso que ameritaba intervención, lo que arroja un valor predictivo positivo del 74 por ciento. El tiempo medio de aplicación del cuestionario fue de siete minutos, y las enfermeras que lo administraron señalaron que la principal dificultad fue la baja escolaridad de algunos cuidadores para completarlo sin ayuda. Los autores proponen incorporarlo al control de los dos años en toda la red municipal."
},
{
"query": "Por que a recolonização de caranguejos foi mais lenta no manguezal restaurado?",
"passage": "No estuário do rio Caeté, no litoral nordeste do Pará, uma iniciativa de reflorestamento de manguezais iniciada em 2019 acompanhou a recuperação de oito hectares degradados por antigas atividades de carcinicultura abandonadas na década de 1990. As mudas de Rhizophora mangle e Avicennia germinans foram plantadas em espaçamento de um metro e meio, totalizando aproximadamente vinte e seis mil indivíduos, com monitoramento trimestral da taxa de sobrevivência ao longo de quatro anos. Os resultados indicam uma sobrevivência média de 71% para Rhizophora mangle, consideravelmente superior aos 48% observados para Avicennia germinans nas parcelas mais próximas ao canal de maré, onde a salinidade intersticial do solo chegou a atingir 38 partes por mil durante os meses de estiagem. Medições de sedimentação instaladas em oito estações mostraram acúmulo médio de 2,3 centímetros de sedimento por ano nas áreas replantadas, taxa próxima à registrada em manguezais maduros de referência situados a cerca de dois quilômetros de distância. A recolonização por caranguejos do gênero Ucides, indicadores frequentemente usados para avaliar a funcionalidade ecológica do manguezal restaurado, começou a ser detectada apenas no terceiro ano, com densidade de tocas ainda equivalente a menos da metade da observada em áreas não degradadas. Entrevistas com trinta e duas famílias de pescadores artesanais da região revelaram que a percepção de recuperação dos estoques de caranguejo-uçá é mais lenta do que sugerem os dados biofísicos, gerando certa desconfiança quanto à efetividade do projeto entre os moradores mais antigos. Os pesquisadores responsáveis pelo monitoramento defendem que a ausência de canais de maré secundários nas parcelas mais internas, removidos durante o período de uso para carcinicultura, pode estar limitando a troca hídrica necessária para acelerar tanto o crescimento das mudas quanto o retorno da fauna associada, e recomendam a escavação de canais artificiais como próxima etapa da intervenção.",
"nearMiss": "Em um trecho de floresta de restinga do litoral norte do Espírito Santo, próximo à localidade de Barra Nova, uma equipe de ecologia vegetal instalou em 2019 doze parcelas permanentes de vinte por cinquenta metros para acompanhar a regeneração natural após um incêndio que atingiu a área em setembro de 2018. Todos os indivíduos lenhosos com altura superior a trinta centímetros foram marcados e medidos anualmente até 2023. A densidade de indivíduos passou de 1.240 por hectare no primeiro ano para 3.860 no quinto, com predominância de espécies que rebrotam a partir de estruturas subterrâneas, sobretudo dos gêneros Byrsonima e Protium, enquanto o recrutamento a partir de sementes permaneceu baixo e concentrado nas parcelas mais próximas da mata não queimada. A cobertura de gramíneas exóticas, ausente antes do fogo, atingiu 22 por cento nas parcelas mais afastadas da borda no terceiro ano e declinou em seguida à medida que o dossel se fechava. Os autores estimam que a riqueza de espécies lenhosas retornou a cerca de 70 por cento do valor registrado em um levantamento anterior ao incêndio."
},
{
"query": "raschietti a lama piatta",
"passage": "Nel tratto sotterraneo dell'acquedotto del Serino che attraversa il territorio dell'attuale comune di Sirignano, in provincia di Avellino, le indagini condotte tra il 2020 e il 2022 hanno permesso di documentare un sistema di pozzi di ispezione posti a intervalli medi di trentacinque metri, utilizzati con ogni probabilità per le operazioni periodiche di rimozione dei depositi calcarei. I carotaggi effettuati all'interno dello specus hanno rivelato uno strato di concrezioni calcaree dello spessore medio di dodici centimetri sul fondo del canale, con punte di diciotto centimetri in corrispondenza dei tratti a pendenza più ridotta, dove la velocità di scorrimento dell'acqua doveva risultare inferiore a venti centimetri al secondo. L'analisi stratigrafica delle concrezioni, condotta mediante sezioni sottili, ha permesso di distinguere almeno sette fasi di accumulo separate da sottili strati di materiale organico, interpretate come tracce di altrettanti interventi di pulizia che scandiscono un intervallo di manutenzione stimato in circa sei o sette anni. In tre dei pozzi esaminati sono stati rinvenuti utensili in ferro, tra cui raschietti a lama piatta con lunghezza compresa tra i quaranta e i sessanta centimetri, verosimilmente impiegati per il distacco meccanico delle incrostazioni dalle pareti del canale. Il confronto con le fonti epigrafiche locali, in particolare un'iscrizione frammentaria rinvenuta nei pressi di uno dei pozzi, menziona un collegio di operai addetti alla cura dell'acquedotto, corroborando l'ipotesi di una manutenzione organizzata e non occasionale. Gli archeologi coinvolti nello scavo ritengono che la frequenza di pulizia rilevata fosse dettata non tanto da un calo drastico della portata, quanto dalla necessità di preservare la pendenza costante dello specus, elemento critico per il funzionamento a gravità dell'intero sistema idraulico che riforniva le città della Campania settentrionale.",
"nearMiss": "Lo scavo di una villa rustica di età imperiale nel territorio di Nola, condotto tra il 2018 e il 2021, ha restituito circa 9.400 resti faunistici provenienti da tre contesti chiusi: una fossa di scarico presso la cucina, il riempimento di una cisterna dismessa e lo strato di abbandono del quartiere servile. L'analisi archeozoologica ha mostrato una netta prevalenza dei suini, che rappresentano il 58 per cento dei resti determinati nella fossa di scarico, seguiti da ovicaprini e bovini, mentre nel riempimento della cisterna la proporzione si inverte a favore dei bovini adulti, con tracce di macellazione sistematica sulle vertebre e sulle scapole. L'età di abbattimento dei suini, stimata dall'eruzione dentaria, si concentra tra i dodici e i ventiquattro mesi, un profilo compatibile con un allevamento orientato alla produzione di carne piuttosto che alla riproduzione. Sono stati inoltre riconosciuti resti di pollame, di lepre e di almeno quattro specie di pesci marini, questi ultimi quasi esclusivamente nel quartiere servile. Gli autori collegano la differenza tra i contesti a una distinzione tra il consumo della famiglia proprietaria e quello della manodopera."
},
{
"query": "桃収穫用ソフトグリッパーの把持成功率",
"passage": "2022年から2023年にかけて、愛知県内の農業機械メーカーと共同で、桃の収穫作業を想定した空気圧駆動式ソフトグリッパーの試作機が開発された。グリッパーは三本のシリコーン製指部から構成され、内部に埋め込まれた光ファイバーセンサーにより把持圧力をリアルタイムで推定する仕組みを採用しており、センサーの応答遅延は平均八ミリ秒に抑えられている。実験では、直径七十から八十五ミリメートルの模擬果実百二十個を用い、把持成功率と表面損傷の有無を評価した。その結果、内部圧力を十二キロパスカルに設定した条件で把持成功率は96%に達し、表面に目視可能な圧痕が生じた個体は全体の3%にとどまった。一方、圧力を十八キロパスカルまで高めた条件では成功率が98%まで向上したものの、圧痕発生率も11%に上昇し、収穫作業における圧力設定の難しさが浮き彫りになった。さらに、実際の果樹園における屋外試験では、枝の間隔が狭い箇所での接近動作に平均二・四秒を要し、屋内実験時の一・一秒と比べて大幅に時間を要することが判明した。屋外試験は九月上旬の収穫期に三日間にわたり実施され、延べ二百四十個の果実を対象とした。研究チームは、この遅延の主な原因が枝葉による視覚センサーの遮蔽であると分析し、今後は接触型センサーを補助的に併用する設計への改良を検討しているほか、次期試作機では指部の材質を変更して耐久性を高める計画も進めている。",
"nearMiss": "水稲圃場を対象としたマルチスペクトルカメラ搭載ドローンによる生育診断について、新潟県内の三つの生産法人と共同で実証試験を行った結果を報告する。対象は合計四十二筆、約十八ヘクタールで、出穂前の三時期に高度五十メートルから撮影し、正規化植生指数を筆ごとに算出した。指数と同時期に採取した葉身の窒素含有率との相関係数は0.78であり、指数に基づいて追肥量を筆ごとに調整した区では、一律に追肥した区と比べて玄米タンパク質含有率のばらつきが約三割縮小した。収量に有意な差は認められなかった。撮影時の風速が毎秒五メートルを超えた日には指数の筆間ばらつきが大きくなり、診断精度が低下することも確認された。生産法人からは、筆ごとの追肥調整に要する作業時間の増加が課題として指摘され、可変施肥機との連携が次年度の検討事項となった。"
},
{
"query": "간호사 소진에 근무 형태가 어떤 영향을 미치는가?",
"passage": "2021년부터 2022년까지 서울 소재 3차 병원 두 곳의 중환자실 간호사 214명을 대상으로 근무 형태와 소진 정도의 관계를 분석한 조사가 진행되었다. 대상자는 3교대 근무자와 2교대 근무자로 구분되었으며, 평균 임상 경력은 6.8년이었다. 마슬라크 소진 척도를 이용해 정서적 고갈, 비인격화, 개인성취감 저하의 세 하위 영역을 3개월 간격으로 총 네 차례 측정하였다. 분석 결과, 2교대 근무자의 정서적 고갈 점수는 평균 34.2점으로 3교대 근무자의 29.6점보다 유의하게 높게 나타났으며, 이러한 차이는 야간 근무 이후 최소 휴식 시간이 11시간에 못 미치는 경우에 특히 두드러졌다. 근무 기록을 세부적으로 분석한 결과, 2교대 근무자 가운데 연속 근무일이 5일을 초과한 사례가 전체의 27%에 달했고, 이 집단의 비인격화 점수는 연속 근무일이 4일 이하인 집단에 비해 약 1.6배 높았다. 한편 개인성취감 저하는 근무 형태보다는 환자 대 간호사 비율과 더 밀접한 관련을 보였는데, 비율이 1대 3을 초과하는 병동에서 근무하는 간호사의 경우 형태와 무관하게 저하 점수가 유의하게 높았다. 조사 기간 중 이직을 고려한다고 응답한 비율은 2교대 근무자 집단에서 41%로, 3교대 근무자 집단의 26%보다 높게 나타났다. 연구진은 근무 형태 자체보다 연속 근무일수와 휴식 간격을 조정하는 스케줄링 방식이 소진 완화에 더 실질적인 개입 지점이 될 수 있다고 제언하였다.",
"nearMiss": "국외에서 개발된 낙상 위험 사정 도구의 한국어판이 노인요양시설 입소자에게도 유효한지 확인하기 위해, 2022년 4월부터 12월까지 충청북도 소재 요양시설 다섯 곳의 65세 이상 입소자 312명을 대상으로 타당도 검증이 이루어졌다. 도구는 입소 시점과 3개월 후 두 차례 적용되었고, 이후 6개월 동안 발생한 낙상 사건을 시설 기록으로 확인하였다. 추적 기간 중 낙상은 87명에게서 총 121건 발생하였으며, 도구 총점이 높은 집단에서 낙상 발생률이 2.4배 높았다. 민감도는 0.79, 특이도는 0.62로 나타났고, 보행 보조기구 사용과 야간 배뇨 횟수 항목이 예측력에 가장 크게 기여하였다. 연구진은 절단점을 원 도구보다 한 점 낮게 조정할 것과, 인지 기능 저하가 있는 입소자에게는 보호자 보고 항목을 추가할 것을 제안하였다."
},
{
"query": "карбонитридное торможение границ зёрен",
"passage": "На металлургическом комбинате в Нижнем Тагиле в течение 2022 года были проведены испытания партии холоднокатаной низкоуглеродистой стали толщиной 0,7 миллиметра с целью изучения влияния режима рекристаллизационного отжига на конечную микроструктуру полосы. Образцы, прокатанные с суммарным обжатием 68%, подвергались отжигу в колпаковых печах при трёх температурных режимах: 620, 680 и 720 градусов Цельсия, с выдержкой от четырёх до восьми часов в защитной азотно-водородной атмосфере. Металлографический анализ показал, что при температуре 620 градусов рекристаллизация происходила не полностью, доля остаточных деформированных зёрен составляла около 23%, что впоследствии приводило к анизотропии механических свойств вдоль и поперёк направления прокатки. При повышении температуры до 680 градусов средний размер зерна феррита стабилизировался на уровне 11,4 микрометра, а показатель нормальной анизотропии достиг значения 1,8, что соответствует требованиям для последующей глубокой вытяжки автомобильных кузовных деталей. Дальнейшее повышение температуры до 720 градусов приводило к росту зерна свыше 16 микрометров и, как следствие, к снижению предела текучести до 178 мегапаскалей, что для ряда заказчиков оказалось ниже допустимого технического порога. Отдельно исследовалась партия с добавлением 0,03% ниобия, где даже при 720 градусах размер зерна не превышал 12 микрометров благодаря карбонитридному торможению границ зёрен. Специалисты лаборатории пришли к выводу, что для базовой марки стали оптимальным следует считать режим отжига при 680 градусах с выдержкой шесть часов, обеспечивающий баланс между штампуемостью и прочностными характеристиками готового проката.",
"nearMiss": "На линии горячего цинкования одного из прокатных заводов Липецкой области в период с апреля по октябрь 2022 года проводились наблюдения за равномерностью толщины покрытия на полосе шириной 1250 миллиметров при скорости движения от 90 до 140 метров в минуту. Толщина цинкового слоя измерялась бесконтактным рентгенофлуоресцентным датчиком в семи точках по ширине полосы с шагом записи в две секунды. Установлено, что при скоростях выше 120 метров в минуту разброс толщины между кромками и центром полосы возрастал с 4 до 11 микрометров, а при снижении давления воздушных ножей ниже 30 килопаскалей на кромках появлялись локальные утолщения. Корректировка зазора между ножами и полосой с 8 до 6 миллиметров на верхних скоростях позволила снизить разброс до 6 микрометров без потери производительности. Дополнительно было отмечено, что колебания температуры ванны в пределах 8 градусов не оказывали заметного влияния на равномерность покрытия."
},
{
"query": "verhoogd fietspad en bijna-ongelukken",
"passage": "Bij de herinrichting van het kruispunt Vleutenseweg-Marnixlaan in Utrecht, uitgevoerd tussen april en oktober 2021, is onderzocht welk effect een verhoogd en in rood asfalt uitgevoerd fietspad heeft op het aantal bijna-ongelukken tussen fietsers en afslaand gemotoriseerd verkeer. Voorafgaand aan de herinrichting werden gedurende zes weken camerabeelden geanalyseerd, waarbij 412 bijna-aanrijdingen werden geregistreerd, gedefinieerd als situaties waarin de tijd tot een mogelijke botsing minder dan anderhalve seconde bedroeg. Na de aanpassing, waarbij het fietspad met tien centimeter werd verhoogd en voorzien van een duidelijk contrasterende belijning, daalde dit aantal in een vergelijkbare meetperiode naar 187 gevallen, een afname van ruim 54%. Met name bij rechtsafslaand verkeer vanuit de Marnixlaan nam het aantal kritieke situaties af, van 261 naar 96 gevallen, terwijl de afname bij linksafslaand verkeer vanaf de Vleutenseweg met 38% aanzienlijk beperkter was. Snelheidsmetingen wezen uit dat de gemiddelde fietssnelheid ter hoogte van het kruispunt licht daalde, van 19,4 naar 17,8 kilometer per uur, wat door de onderzoekers deels wordt toegeschreven aan de verhoogde opstelling die fietsers dwingt iets voorzichtiger te remmen bij nadering. Enquêtes onder 340 dagelijkse fietsers, afgenomen zowel voor als na de ingreep, toonden een toename van het ervaren veiligheidsgevoel van gemiddeld 6,1 naar 7,8 op een schaal van tien. Automobilisten meldden daarentegen vaker verminderd zicht op overstekende voetgangers door de verhoogde drempel, een punt dat de gemeente meeneemt bij de evaluatie van vergelijkbare kruispunten elders in de stad.",
"nearMiss": "In een woonwijk in Haarlem met ongeveer 2.100 woningen is tussen mei 2022 en juni 2023 de parkeerdruk gemeten voor en na de invoering van vergunningparkeren, om na te gaan of de maatregel het gebruik van de openbare parkeerplaatsen door bezoekers van het nabijgelegen station terugdringt. Op zes werkdagen en twee zaterdagen per meetronde werden alle 1.640 parkeerplaatsen in de wijk om tien uur, veertien uur en twintig uur geteld, waarbij kentekens werden vastgelegd om de verblijfsduur van geparkeerde voertuigen te bepalen. Vóór de invoering bedroeg de gemiddelde bezetting op werkdagen om tien uur 94 procent, waarvan naar schatting een kwart voor rekening kwam van voertuigen die de hele werkdag bleven staan en niet aan een adres in de wijk konden worden gekoppeld. Na de invoering daalde de bezetting op dat tijdstip tot 71 procent en verdween het langparkeren door niet-bewoners nagenoeg geheel, terwijl de avondbezetting met 88 procent vrijwel gelijk bleef. In de aangrenzende wijk zonder vergunningstelsel steeg de bezetting op werkdagen met acht procentpunten."
},
{
"query": "pieczenie taśm szpulowych",
"passage": "W latach 2020-2023 zespół etnomuzykologów z ośrodka badawczego w Lublinie przeprowadził cyfryzację zbioru taśm szpulowych z nagraniami pieśni ludowych, zarejestrowanych podczas ekspedycji terenowych na Lubelszczyźnie i Podlasiu w latach sześćdziesiątych i siedemdziesiątych. Zbiór obejmował 1840 nagrań o łącznym czasie trwania niemal 620 godzin, przechowywanych na taśmach o szerokości 6,25 milimetra, z których część wykazywała już wyraźne oznaki degradacji spoiwa magnetycznego. Przed digitalizacją każdą taśmę poddawano procesowi tak zwanego pieczenia w temperaturze 54 stopni Celsjusza przez osiem godzin, co pozwoliło odzyskać sygnał z 96% materiału uznanego wcześniej za zagrożony całkowitą utratą. Nagrania przekonwertowano do formatu cyfrowego o rozdzielczości 24 bitów i częstotliwości próbkowania 96 kilogerców, a następnie poddano ręcznej korekcie szumów przez zespół złożony z czterech specjalistów, przy czym każda godzina nagrania wymagała średnio dwóch i pół godziny pracy redakcyjnej. Analiza treści wykazała, że 62% zarejestrowanych pieśni stanowiły utwory obrzędowe związane z porą żniw i dożynkami, podczas gdy pieśni weselne stanowiły jedynie 14% zbioru, co badacze tłumaczą specyfiką pory roku, w której prowadzono większość nagrań terenowych. Porównanie wariantów melodycznych tej samej pieśni żniwnej zarejestrowanej w trzech różnych wsiach oddalonych od siebie o mniej niż dwadzieścia kilometrów ujawniło zaskakująco duże różnice w ornamentyce wykonawczej, co skłoniło zespół do postawienia hipotezy o silnym zróżnicowaniu lokalnych szkół śpiewaczych nawet na stosunkowo niewielkim obszarze geograficznym.",
"nearMiss": "Kapele weselne w dziewięciu wsiach Podhala grają dziś w tradycyjnym składzie skrzypiec prymu, sekundu i basów w niespełna połowie przypadków, a w pozostałych pojawił się akordeon lub instrumenty elektroniczne. Taki obraz wyłania się z badań terenowych prowadzonych w latach 2021 i 2022 przez zespół ośrodka etnograficznego w Nowym Targu, obejmujących sześćdziesiąt cztery wywiady z muzykantami oraz rejestrację przebiegu czternastu wesel. Repertuar taneczny liczył średnio dwadzieścia dwa utwory tradycyjne na wesele, przy czym nuty ozwodne i krzesane wykonywano niemal wyłącznie na życzenie starszych gości, a po północy dominowały utwory estradowe. Muzykanci poniżej trzydziestego roku życia stanowili jedną trzecią rozmówców i w większości zdobywali umiejętności w szkołach muzycznych, a nie w rodzinie. Autorzy zwracają uwagę, że honorarium kapeli wzrosło w ciągu dekady blisko dwukrotnie, co zdaniem starszych muzykantów ogranicza liczbę wesel z muzyką na żywo na rzecz oprawy odtwarzanej z nagrań."
}
]
}

View file

@ -1,4 +1,4 @@
-- 129
-- 130
-- Copyright (c) 2009 Center for History and New Media
-- George Mason University, Fairfax, Virginia, USA

View file

@ -106,6 +106,8 @@
@import "elements/attachmentRow";
@import "elements/attachmentAnnotationsBox";
@import "elements/annotationRow";
@import "elements/searchResultsBox";
@import "elements/itemPaneSearchResults";
@import "elements/noteRow";
@import "elements/librariesCollectionsBox";
@import "elements/duplicatesMergePane";

View file

@ -95,6 +95,7 @@ $item-pane-sections: (
"libraries-collections": var(--accent-teal),
"tags": var(--accent-orange),
"related": var(--accent-wood),
"search-results": var(--accent-gold),
);
$tagColorsLookup: (

View file

@ -357,12 +357,70 @@
flex-grow: 2;
flex-basis: 0;
}
// The text cells size to their content, so push the relevance
// bar to the row's end, where the Relevance column sits while
// a best-match search is active (the column's width rules
// still apply -- they're injected with !important)
&.relevance {
margin-inline-start: auto;
}
}
&.tight .cell {
// The tight annotation layout trades cell padding for text room, but
// the relevance bar fills its cell's content box, so dropping the
// padding would render a wider bar than every other row's. Keep the
// standard padding on that one cell so the bars match.
&.tight .cell:not(.relevance) {
padding: 0 2px 0 0;
}
}
.search-match-row {
// One cell across the row: a match shows a passage, which the
// columns say nothing about
.cell {
font-size: $font-size-small;
flex: 1 1 auto;
}
// Location above quote, beside the indent and twisty the tree adds
// to the first cell. The row's height is set from JS for exactly
// these two lines (ItemTree#_getSearchMatchRowHeight()), so
// neither may wrap.
.search-match-lines {
display: flex;
flex-direction: column;
justify-content: center;
flex: 1 1 auto;
// Without this a flex item refuses to shrink below its text,
// and nothing is ever clipped to an ellipsis
min-width: 0;
}
.search-match-location,
.cell-text {
display: block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.search-match-location {
color: var(--fill-secondary);
}
&.selected {
.search-match-location {
color: var(--color-accent-text);
opacity: 0.7;
// An unfocused selection is a grey fill with ordinary text
@include state(".virtualized-table:not(:focus-within)") {
color: var(--fill-secondary);
opacity: 1;
}
}
}
.search-match-highlight {
font-weight: 600;
}
}
.cell:not(.hasAttachment) .item-icon {
margin-inline-end: 4px;
}

View file

@ -1,4 +1,8 @@
annotation-row {
// search-result-row (the search-results section's chunk cards) shares this
// structure and look, minus the parts it doesn't have (icon, action, tags):
// the section already carries the magnifier in its head, so repeating it on
// every card says nothing
annotation-row, search-result-row {
display: flex;
flex-direction: column;

View file

@ -0,0 +1,48 @@
search-results-pane {
display: flex;
flex-direction: column;
overflow-y: auto;
@include elements-custom-head;
// Keep the attachment's name legible however long it is, cutting it by
// letter rather than dropping the whole line
collapsible-section > .head .title-box .summary {
opacity: 1 !important;
width: 0;
white-space: wrap;
word-break: break-all;
display: inline;
}
// The section's icon is the quick-search magnifier, the same one the
// search bar uses
collapsible-section > .head .title::before {
content: '';
width: 16px;
height: 16px;
background: icon-url("16/universal/magnifier.svg") no-repeat center;
-moz-context-properties: fill, fill-opacity, stroke, stroke-opacity;
fill: var(--accent-gold);
stroke: var(--accent-gold);
}
collapsible-section > .body {
display: flex;
flex-direction: column;
gap: 4px;
@include comfortable {
gap: 8px;
}
}
collapsible-section:not(:last-child) {
border-bottom: 1px solid var(--fill-quinary);
}
// A passage shown on its own is there to be read, so it isn't clamped
search-result-row .body .quote {
-webkit-line-clamp: inherit !important;
}
}

View file

@ -0,0 +1,85 @@
// The section's icon is the quick-search magnifier, the same one the search
// bar uses, rather than a per-size copy of its own. Both the sidenav button
// and the section head otherwise derive their icon from the pane name (see
// _itemPaneSidenav.scss and _collapsibleSection.scss), so each needs pointing
// at the shared file.
item-pane-sidenav .btn[data-pane="search-results"] {
background-image: url("chrome://zotero/skin/20/universal/magnifier.svg");
}
collapsible-section[data-pane="search-results"] > .head .title::before {
background-image: icon-url("16/universal/magnifier.svg");
}
search-results-box {
display: flex;
flex-direction: column;
&[hidden] {
display: none;
}
& > collapsible-section {
& > .body {
display: flex;
flex-direction: column;
gap: 4px;
@include comfortable {
gap: 8px;
}
}
}
}
// The parts search-result-row adds on top of the shared annotation-row look
// (see _annotationRow.scss): the page label in the head, and the clamp
// toggle under the quote
search-result-row {
.head {
.location {
margin-inline-start: auto;
color: var(--fill-secondary);
white-space: nowrap;
}
}
// A passage's paragraph breaks are kept
.body .quote {
white-space: pre-line;
}
&.expanded .body .quote {
-webkit-line-clamp: none;
}
// Where the query's words are in the passage
.body .quote .match {
font-weight: 600;
}
// The line the tree row quotes, in the primary color against the
// quote's secondary
.body .quote .snippet {
color: var(--fill-primary);
}
.show-more {
align-self: flex-start;
margin: 0 8px 4px 16px;
padding: 0;
border: none;
background: transparent;
color: var(--fill-secondary);
font: inherit;
cursor: pointer;
&:hover {
text-decoration: underline;
}
&[hidden] {
display: none;
}
}
}

View file

@ -98,6 +98,76 @@
margin-inline-start: 6px;
}
#semantic-search {
// One gap between rows; the controls' own block margins would make
// the gap depend on which control a row holds
row-gap: 6px;
> hbox > :is(menulist, button, .html-input),
> vbox > hbox > :is(menulist, button) {
margin-block: 0;
}
#best-match-margin {
width: 5em;
}
.semantic-search-hint {
opacity: 0.7;
font-size: 0.9em;
}
#semantic-search-status {
row-gap: 6px;
}
#semantic-search-diagnostics-toggle {
align-self: flex-start;
margin-inline: 0;
}
// Bars and figures share one label column
#semantic-search-progress,
#semantic-search-diagnostics {
display: grid;
align-items: center;
column-gap: 12px;
row-gap: 4px;
}
#semantic-search-progress {
grid-template-columns: max-content 1fr max-content;
.semantic-search-progress-row {
display: contents;
&[hidden] {
display: none;
}
}
progress {
width: 100%;
}
}
#semantic-search-diagnostics {
grid-template-columns: max-content 1fr;
&[hidden] {
display: none;
}
label:nth-child(odd) {
color: var(--fill-secondary);
}
label:nth-child(even) {
white-space: normal;
}
}
}
#db-maintenance-options {
display: flex;
gap: 6px;
@ -111,10 +181,27 @@
}
}
#semantic-search-model-description {
margin-bottom: 10px;
}
#zotero-embeddings-endpoint-container {
.endpoint-note {
opacity: 0.7;
font-size: 0.9em;
}
#semantic-search-libraries {
padding-top: 10px;
}
.endpoint-facts {
display: grid;
grid-template-columns: max-content 1fr;
column-gap: 8px;
margin: 4px 0 4px 12px;
}
.endpoint-command {
flex: 1;
font-family: monospace;
font-size: 0.9em;
resize: none;
}
.endpoint-error {
color: var(--accent-red);
}
}

View file

@ -174,6 +174,29 @@ describe("Advanced Search", function () {
await childOnly.eraseTx();
});
it("should prefill a Best Match quick search as the ranking field, keeping its conditions", async function () {
var item = await createDataObject('item', { title: "alpha beta" });
item.addTag('zztag');
await item.saveTx();
await zp.openAdvancedSearchFromQuickSearch('tag:zztag alpha beta', 'bestMatch');
var iv = zp.itemsView;
await iv.waitForLoad();
// The clause stays a condition; the free text ranks rather than
// filtering, so no word becomes a condition of its own
var conds = Object.values(deck.pane.search.getConditions());
assert.isTrue(conds.some(c => c.condition === 'tag' && c.value === 'zztag'));
assert.isFalse(conds.some(c => c.condition === 'anyField'));
assert.isTrue(conds.some(c => c.condition === 'resultLevel' && c.operator === 'item'));
assert.equal(deck.pane.search.getBestMatchQuery(), 'alpha beta');
assert.equal(deck.pane.querySelector('.best-match-input').value, 'alpha beta');
await zp.setAdvancedSearchState('closed');
await iv.waitForLoad();
await item.eraseTx();
});
it("should run a cross-level search across a multi-collection selection", async function () {
var word = 'zmc' + Zotero.Utilities.randomString();
var makeMatch = async function (collection) {

1121
test/tests/bestMatchTest.js Normal file

File diff suppressed because it is too large Load diff

View file

@ -202,18 +202,91 @@ describe("CollectionViewItemTree", function () {
assert.equal(quicksearch.value, "test");
});
describe("in best-match mode without embeddings", function () {
var stubs = [];
beforeEach(function () {
stubs.push(sinon.stub(Zotero.Embeddings, 'isEnabled').returns(false));
Zotero.Prefs.set('search.quicksearch-mode', 'bestMatch');
});
afterEach(async function () {
Zotero.Prefs.set('search.quicksearch-mode', 'fields');
await zp.itemsView.setFilter('search', '');
await selectLibrary(win);
stubs.forEach(stub => stub.restore());
stubs = [];
});
it("should rank items lexically", async function () {
// Ranking is under test, not membership: keep the two-word
// match that the result margin would otherwise cut
Zotero.Prefs.set('search.bestMatchMargin', 100);
try {
let col = await createDataObject('collection');
// Unrelated items, so the query words the matches share still
// separate documents in this corpus -- in a corpus of nothing
// but matches, FTS5 floors their idf as separating nothing
// and no match earns a score
for (let i = 0; i < 6; i++) {
await createDataObject('item', { title: `unrelated filler number ${i}` });
}
let full = await createDataObject('item',
{ title: 'Lexint owl migration patterns', collections: [col.id] });
// Three of the query's four terms, ranked below the full match
let partial = await createDataObject('item',
{ title: 'Lexint owl migration handbook', collections: [col.id] });
// Two of four, ranked below both
let sparse = await createDataObject('item',
{ title: 'Lexint owl guidebook', collections: [col.id] });
await select(win, col);
let itemsView = zp.itemsView;
await itemsView.setFilter('search', 'lexint owl migration patterns');
// Scored items only, ranked by coverage, most relevant first
assert.deepEqual(itemsView._rows.map(row => row.id),
[full.id, partial.id, sparse.id]);
assert.equal(itemsView.getSortField(), 'relevance');
// The bars carry the lexical scores directly
let fractions = itemsView.rowProvider.getBestMatchBarFractions();
assert.isAbove(fractions.get(full.id), fractions.get(partial.id));
assert.isAbove(fractions.get(partial.id), fractions.get(sparse.id));
assert.isAtMost(fractions.get(full.id), 1);
assert.isAbove(fractions.get(sparse.id), 0);
}
finally {
Zotero.Prefs.clear('search.bestMatchMargin');
}
});
});
describe("in best-match mode", function () {
var stubs = [];
// Embeddings scoreItemIDs fakes below supply bare score Maps (or a
// function returning one); wrap them in the engine's real
// { scores, previewableIDs } envelope
function scoreEnvelope(fake) {
return async (...args) => ({
scores: await (typeof fake == 'function' ? fake(...args) : fake),
previewableIDs: new Set()
});
}
beforeEach(function () {
stubs.push(sinon.stub(Zotero.Embeddings, 'isEnabled').returns(true));
stubs.push(sinon.stub(Zotero.Embeddings, 'getScoreFraction').callsFake(score => score));
// No lexical matches, so the fused ranking and the bars carry
// the semantic scores these tests control
stubs.push(sinon.stub(Zotero.Lexical, 'scoreItemIDs').resolves(new Map()));
// A fully built index by default, so no indexing banner appears
stubs.push(sinon.stub(Zotero.Embeddings.Indexing, 'getStatus').returns({
enabled: true,
indexing: false,
paused: false,
libraries: [{ libraryID: Zotero.Libraries.userLibraryID, indexed: 0, eligible: 0 }]
items: { done: 1, total: 1 },
attachments: { done: 0, total: 0, awaiting: 0 }
}));
Zotero.Prefs.set('search.quicksearch-mode', 'bestMatch');
});
@ -233,7 +306,7 @@ describe("CollectionViewItemTree", function () {
let itemA = await createDataObject('item', { title: "A", collections: [col.id] });
let itemB = await createDataObject('item', { title: "B", collections: [col.id] });
let itemC = await createDataObject('item', { title: "C", collections: [col.id] });
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(async (query, itemIDs) => {
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(async (query, itemIDs) => {
let scores = new Map();
if (itemIDs.includes(itemA.id)) {
scores.set(itemA.id, 0.5);
@ -242,7 +315,7 @@ describe("CollectionViewItemTree", function () {
scores.set(itemB.id, 0.9);
}
return scores;
}));
})));
await select(win, col);
itemsView = zp.itemsView;
@ -258,9 +331,12 @@ describe("CollectionViewItemTree", function () {
// The Relevance cells show the ranks
assert.equal(itemsView.getCellText(0, 'relevance'), 1);
assert.equal(itemsView.getCellText(1, 'relevance'), 2);
// Score fractions for the bars
assert.equal(itemsView.rowProvider.getBestMatchBarFractions().get(itemB.id), 0.9);
assert.equal(itemsView.rowProvider.getBestMatchBarFractions().get(itemA.id), 0.5);
// The bars carry the fused scores, so they agree with the
// ranking: with no lexical matches, a semantic match at rank
// 1 fuses to half its fraction, rank 2 to fraction * 61/124
let barFractions = itemsView.rowProvider.getBestMatchBarFractions();
assert.closeTo(barFractions.get(itemB.id), 0.9 / 2, 1e-12);
assert.closeTo(barFractions.get(itemA.id), 0.5 * 61 / 124, 1e-12);
assert.isFalse(itemsView._getColumns().find(c => c.dataKey == 'relevance').hidden);
// The rendered header shows the column
assert.ok(win.document.querySelector('.virtualized-table-header .cell.relevance'));
@ -295,13 +371,115 @@ describe("CollectionViewItemTree", function () {
);
});
it("should rank a row by the best match beneath it, with the bar reporting only its own score", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "liftrank paper", collections: [col.id] });
let attachment = await importPDFAttachment(item);
let annotation = await createAnnotation('highlight', attachment,
{ comment: 'lift comment' });
let other = await createDataObject('item', { title: "liftrank other", collections: [col.id] });
// Only the annotation and the unrelated peer match on their own
// text -- the paper's own abstract says nothing about the query
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(async (query, itemIDs) => {
let scores = new Map();
if (itemIDs.includes(annotation.id)) {
scores.set(annotation.id, 0.9);
}
if (itemIDs.includes(other.id)) {
scores.set(other.id, 0.5);
}
return scores;
})));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
// The annotation lifts its paper above the peer that matched
// on its own text
assert.deepEqual(
itemsView._rows.filter(row => row.level == 0).map(row => row.id),
[item.id, other.id]
);
// The whole chain under the annotation carries its rank
let ranks = itemsView.rowProvider.getBestMatchRanks();
assert.equal(ranks.get(annotation.id), 1);
assert.equal(ranks.get(attachment.id), 1);
assert.equal(ranks.get(item.id), 1);
assert.equal(ranks.get(other.id), 2);
// The bar reports only the row's own score: the annotation gets
// its match (fused: a rank-1 semantic match at half its
// fraction), the rows ranked by it get an empty bar
let fractions = itemsView.rowProvider.getBestMatchBarFractions();
assert.closeTo(fractions.get(annotation.id), 0.45, 1e-12);
assert.equal(fractions.get(attachment.id), 0);
assert.equal(fractions.get(item.id), 0);
assert.closeTo(fractions.get(other.id), 0.5 * 61 / 124, 1e-12);
// The matched annotation's ancestors auto-expand, and its row
// renders a relevance bar of its own. Row painting is async, so
// poll (the test times out on failure).
let annotationRow = itemsView.getRowIndexByID(annotation.id);
assert.notEqual(annotationRow, false);
let bar;
while (!bar) {
bar = itemsView.tree._jsWindow.getElementByIndex(annotationRow)
?.querySelector('.cell.relevance .relevance-bar');
if (!bar) {
await Zotero.Promise.delay(10);
}
}
assert.equal(bar.firstChild.style.width, '45%');
// An annotation row's bar has to be the same width as every
// other row's. The tight annotation layout drops cell padding,
// and the bar fills its cell's content box, so without an
// exception for this cell the bar would render wider.
let peerRow = itemsView.getRowIndexByID(other.id);
let peerBar = itemsView.tree._jsWindow.getElementByIndex(peerRow)
.querySelector('.cell.relevance .relevance-bar');
assert.isTrue(bar.closest('.row').classList.contains('tight'),
"annotation row should use the tight layout for this to be meaningful");
assert.equal(
bar.getBoundingClientRect().width,
peerBar.getBoundingClientRect().width
);
});
it("should move the Relevance column to the far right and restore it when cleared", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "farright A", collections: [col.id] });
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, 0.5]))
)));
await select(win, col);
itemsView = zp.itemsView;
let visibleBefore = itemsView._getColumns()
.filter(column => !column.hidden).map(column => column.dataKey);
await itemsView.setFilter('search', 'some query');
let visible = itemsView._getColumns().filter(column => !column.hidden);
assert.equal(visible[visible.length - 1].dataKey, 'relevance');
// The rendered header agrees
let headerCells = [...win.document.querySelectorAll('.virtualized-table-header .cell')];
assert.isTrue(headerCells[headerCells.length - 1].classList.contains('relevance'));
// Clearing the search puts the columns back
await itemsView.setFilter('search', '');
assert.deepEqual(
itemsView._getColumns().filter(column => !column.hidden).map(column => column.dataKey),
visibleBefore
);
});
it("should override a persisted column sort while a best-match search is active", async function () {
let col = await createDataObject('collection');
let itemA = await createDataObject('item', { title: "persistsort A", collections: [col.id] });
let itemB = await createDataObject('item', { title: "persistsort B", collections: [col.id] });
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == itemB.id ? 0.9 : 0.5]))
));
)));
await select(win, col);
itemsView = zp.itemsView;
@ -337,16 +515,442 @@ describe("CollectionViewItemTree", function () {
}
});
it("should show match rows under matched attachments, derived before the rows appear", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "matchrow A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ source: 'title', text: 'matchrow owls', ranges: [[9, 13]], strength: 1 },
{ source: 'abstract', text: 'about owls', ranges: [[6, 10]], strength: 0.5 }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
// One row per derived entry, already in place when the search
// resolves
let matchRow = itemsView.getRowIndexByID('SM' + attachment.id + '-0');
assert.notStrictEqual(matchRow, false);
assert.notStrictEqual(itemsView.getRowIndexByID('SM' + attachment.id + '-1'), false);
assert.equal(itemsView.getRow(matchRow).ref.entry.text, 'matchrow owls');
// The parent and the matched attachment both auto-expanded
assert.isTrue(itemsView.isContainerOpen(itemsView.getRowIndexByID(item.id)));
assert.isTrue(itemsView.isContainerOpen(itemsView.getRowIndexByID(attachment.id)));
assert.equal(itemsView.getLevel(matchRow), 2);
// Clearing the search removes the match rows
await itemsView.setFilter('search', '');
assert.isFalse(itemsView.getRowIndexByID('SM' + attachment.id + '-0'));
});
it("should place a semantic match's rows under its attachment", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "chunkcount A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
async (query, itemIDs) => ({
scores: new Map(itemIDs.includes(attachment.id)
? [[attachment.id, 0.9]] : []),
previewableIDs: new Set(itemIDs.includes(attachment.id)
? [attachment.id] : [])
})
));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ source: 'content', text: 'chunkcount owls', ranges: [], strength: 1 }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let matchRow = itemsView.getRowIndexByID('SM' + attachment.id + '-0');
assert.notStrictEqual(matchRow, false);
// Under the attachment, which auto-expanded to show it
let attachmentRow = itemsView.getRowIndexByID(attachment.id);
assert.equal(itemsView.getParentIndex(matchRow), attachmentRow);
});
it("should drop a non-matching annotation row when a search starts", async function () {
Zotero.Prefs.set("hideContextAnnotationRows", true);
try {
let col = await createDataObject('collection');
let item = await createDataObject('item',
{ title: "annstale A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
// One annotation the search matches and one it doesn't --
// with a match among them, the non-matching one is context
let matching = await createAnnotation('highlight', attachment);
let context = await createAnnotation('highlight', attachment);
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
[attachment.id, matching.id]
.filter(id => itemIDs.includes(id))
.map(id => [id, 0.8])));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype,
'getMatchingExcerpts').resolves([]));
await select(win, col);
itemsView = zp.itemsView;
// Open the container before searching, so its children are
// built while no search is filtering them
itemsView.expandAllRows(true);
assert.notStrictEqual(itemsView.getRowIndexByID(context.id), false,
'the annotation is a row to begin with');
await itemsView.setFilter('search', 'some query');
// The search excludes it and the pref hides non-matching
// annotations, so its row shouldn't have survived
assert.isFalse(itemsView.getRowIndexByID(context.id));
// ...while the one it matched stays
assert.notStrictEqual(itemsView.getRowIndexByID(matching.id), false);
}
finally {
Zotero.Prefs.set("hideContextAnnotationRows", false);
}
});
it("should show no match rows for a matched note", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "notematch A", collections: [col.id] });
let note = new Zotero.Item('note');
note.parentID = item.id;
note.setNote('<p>notematch text</p>');
await note.saveTx();
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(note.id) ? [[note.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
// Only file attachments show match rows, so a matched note
// stays a plain, childless row
let noteRow = itemsView.getRowIndexByID(note.id);
assert.notStrictEqual(noteRow, false);
assert.isFalse(itemsView.isContainer(noteRow));
assert.isFalse(itemsView.getRowIndexByID('SM' + note.id + '-0'));
});
it("should show selected match rows as passages in the item pane", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "matchselect A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ key: 0, text: 'matchselect owls', ranges: [], strength: 1,
snippet: { start: 0, end: 16 } },
{ key: 1, text: 'more about owls', ranges: [], strength: 0.5,
snippet: { start: 0, end: 15 } }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let first = itemsView.getRowIndexByID('SM' + attachment.id + '-0');
let second = itemsView.getRowIndexByID('SM' + attachment.id + '-1');
itemsView.selection.select(second);
// A passage isn't an item, so no item is selected
assert.lengthOf(itemsView.getSelectedItems(), 0);
let matches = itemsView.getSelectedSearchMatches();
assert.lengthOf(matches, 1);
assert.equal(matches[0].itemID, attachment.id);
assert.equal(matches[0].entry.key, 1);
await zp.itemSelected();
// The passage is shown on its own, not the attachment's fields
assert.equal(zp.itemPane.mode, 'search-results');
let pane = zp.itemPane.querySelector('#zotero-search-results-pane');
assert.lengthOf(pane.querySelectorAll('search-result-row'), 1);
// Every selected passage gets a card, under its attachment
itemsView.selection.clearSelection();
itemsView.selection.rangedSelect(first, second, true);
await zp.itemSelected();
assert.lengthOf(itemsView.getSelectedSearchMatches(), 2);
assert.equal(zp.itemPane.mode, 'search-results');
assert.lengthOf(pane.querySelectorAll('search-result-row'), 2);
assert.lengthOf(pane.querySelectorAll('collapsible-section'), 1);
// A selection holding anything that isn't a passage names none
itemsView.selection.clearSelection();
itemsView.selection.rangedSelect(
itemsView.getRowIndexByID(attachment.id), second, true);
assert.isEmpty(itemsView.getSelectedSearchMatches());
});
it("should mark a cut in a match row's line, but not a line opening on a sentence", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "matchcut A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
let text = 'First part. Second sentence here, and more text follows.';
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ key: 0, text, ranges: [], strength: 1,
snippet: { start: 12, end: 32, startsSentence: true } },
{ key: 1, text, ranges: [], strength: 0.5,
snippet: { start: 34, end: 56 } }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let line = id => itemsView.getRow(itemsView.getRowIndexByID(id)).getQuotedLine().text;
assert.equal(line('SM' + attachment.id + '-0'), 'Second sentence here…');
assert.equal(line('SM' + attachment.id + '-1'), '…and more text follows.');
});
it("should mark the quoted line in a selected match row's passage", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "matchmark A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
let text = 'An opening sentence. The owl sentence is quoted. A closing one.';
let start = text.indexOf('The owl');
let end = text.indexOf(' A closing');
let owl = text.indexOf('owl');
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ key: 0, text, ranges: [[owl, owl + 3]], strength: 1,
snippet: { start, end, startsSentence: true } }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'owl');
itemsView.selection.select(itemsView.getRowIndexByID('SM' + attachment.id + '-0'));
await zp.itemSelected();
let quote = zp.itemPane.querySelector('#zotero-search-results-pane search-result-row .quote');
// The whole passage, with the tree's line marked and the match
// bold inside it
assert.equal(quote.textContent, text);
let marked = quote.querySelector('.snippet');
assert.equal(marked.textContent, 'The owl sentence is quoted.');
assert.equal(marked.querySelector('.match').textContent, 'owl');
});
it("should hide non-matching annotations when the attachment has match rows", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "annhide A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
// Random text, so it never matches the query
await createAnnotation('highlight', attachment);
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ key: 0, text: 'annhide owls', ranges: [], strength: 1,
snippet: { start: 0, end: 12 } }
]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let attachmentRow = itemsView.getRowIndexByID(attachment.id);
let childrenFor = () => itemsView.getRow(attachmentRow).getChildItems({
searchMode: true,
searchItemIDs: new Set([attachment.id]),
getMatchPreviews: itemsView.bestMatchSession.getPreviews
});
Zotero.Prefs.set("hideContextAnnotationRows", true);
try {
// The annotation didn't match, and the match rows are what
// the attachment expands to instead
let children = childrenFor();
assert.isFalse(children.some(ref => ref.isAnnotation?.()));
assert.isAbove(children.length, 0);
}
finally {
Zotero.Prefs.set("hideContextAnnotationRows", false);
}
// With the pref off it stays alongside the match rows
assert.isTrue(childrenFor().some(ref => ref.isAnnotation?.()));
});
it("should open a match at its passage from the tree and the item pane", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "openmatch A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
let position = { pageIndex: 3, rects: [[1, 2, 3, 4]] };
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([
{ key: 0, text: 'openmatch owls', ranges: [], strength: 1,
snippet: { start: 0, end: 14 }, position },
// A passage with no geometry to navigate to
{ key: 1, text: 'openmatch more owls', ranges: [], strength: 0.5,
snippet: { start: 0, end: 19 } }
]));
let viewAttachment = sinon.stub(zp, 'viewAttachment').resolves();
stubs.push(viewAttachment);
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let withPosition = itemsView.getRowIndexByID('SM' + attachment.id + '-0');
// From the tree: the attachment, at the passage's geometry
await itemsView.handleActivate({}, [withPosition]);
assert.isTrue(viewAttachment.calledOnce);
assert.equal(viewAttachment.firstCall.args[0], attachment.id);
assert.deepEqual(viewAttachment.firstCall.args[3], { location: { position } });
// A passage with no geometry opens the attachment as it stands
viewAttachment.resetHistory();
await itemsView.handleActivate(
{}, [itemsView.getRowIndexByID('SM' + attachment.id + '-1')]);
assert.isTrue(viewAttachment.calledOnce);
assert.isUndefined(viewAttachment.firstCall.args[3]);
// From the item pane: selecting the row shows its card, and
// double-clicking the card opens the same place
viewAttachment.resetHistory();
itemsView.selection.select(withPosition);
await zp.itemSelected();
let card = zp.itemPane
.querySelector('#zotero-search-results-pane search-result-row');
assert.ok(card);
card.dispatchEvent(new win.MouseEvent('dblclick', { bubbles: true }));
await Zotero.Promise.delay(50);
assert.isTrue(viewAttachment.calledOnce);
assert.equal(viewAttachment.firstCall.args[0], attachment.id);
assert.deepEqual(viewAttachment.firstCall.args[3], { location: { position } });
});
it("should show only the quoted matches as rows", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "quotedrows A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
let entries = [0, 1, 2, 3, 4].map(i => ({
key: i,
text: `quotedrows passage ${i}`,
ranges: [],
strength: 1 - i / 10,
snippet: i < 3 ? { start: 0, end: 10 } : undefined
}));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts')
.resolves(entries));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
// The tree shows the quoted passages; the rest are read in
// the item pane
for (let i = 0; i < Zotero.BestMatch.MAX_QUOTED_PASSAGES; i++) {
assert.notStrictEqual(
itemsView.getRowIndexByID('SM' + attachment.id + '-' + i), false,
`passage ${i} has a row`);
}
assert.isFalse(itemsView.getRowIndexByID(
'SM' + attachment.id + '-' + Zotero.BestMatch.MAX_QUOTED_PASSAGES));
// The preview still holds them all
assert.lengthOf(
itemsView.bestMatchSession.getPreviews(attachment.id).entries, 5);
});
it("should show match rows for previews derived after the search resolves", async function () {
this.timeout(60000);
// One more attachment than score() derives before resolving,
// so the last preview arrives after the results are on screen
let preloaded = Zotero.BestMatch.PRELOADED_MATCH_PREVIEWS;
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "backfill A", collections: [col.id] });
let attachments = [];
for (let i = 0; i <= preloaded; i++) {
attachments.push(await importFileAttachment('test.pdf', { parentID: item.id }));
}
// Descending, so the extra attachment is the one left over
let ids = attachments.map(att => att.id);
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
ids.filter(id => itemIDs.includes(id)).map((id, i) => [id, 0.9 - i / 100])));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
// Held open, so the search has to resolve without it
let release;
let held = new Promise(resolve => release = resolve);
let derived = 0;
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts')
.callsFake(async () => {
if (++derived > preloaded) {
await held;
}
return [{ source: 'content', text: 'backfill owls', ranges: [], strength: 1 }];
}));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
let last = attachments[attachments.length - 1];
// The preloaded ones are on screen with the results...
assert.notStrictEqual(itemsView.getRowIndexByID('SM' + attachments[0].id + '-0'), false);
// ...while the leftover is still deriving, so it has no rows
assert.isFalse(itemsView.getRowIndexByID('SM' + last.id + '-0'));
release();
await itemsView.bestMatchSession.previewsSettled;
// Filling expanded the attachment and added the match row
let matchRow = itemsView.getRowIndexByID('SM' + last.id + '-0');
assert.notStrictEqual(matchRow, false);
assert.equal(itemsView.getParentIndex(matchRow),
itemsView.getRowIndexByID(last.id));
});
it("should show no match rows when derivation finds nothing to show", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "emptyfill A", collections: [col.id] });
let attachment = await importFileAttachment('test.pdf', { parentID: item.id });
Zotero.Lexical.scoreItemIDs.callsFake(async (query, itemIDs) => new Map(
itemIDs.includes(attachment.id) ? [[attachment.id, 0.8]] : []));
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.callsFake(scoreEnvelope(new Map())));
stubs.push(sinon.stub(Zotero.BestMatch.Session.prototype, 'getMatchingExcerpts').resolves([]));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
assert.isFalse(itemsView.getRowIndexByID('SM' + attachment.id + '-0'));
// The item stays -- it's still a scored result
assert.notStrictEqual(itemsView.getRowIndexByID(item.id), false);
});
it("should show an indexing-progress banner while the index is incomplete", async function () {
let col = await createDataObject('collection');
let item = await createDataObject('item', { title: "A", collections: [col.id] });
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs')
.resolves(new Map([[item.id, 0.7]])));
.callsFake(scoreEnvelope(new Map([[item.id, 0.7]]))));
// The counts are split between the item and attachment pairs, so
// the banner's totals prove the two are summed
Zotero.Embeddings.Indexing.getStatus.returns({
enabled: true,
indexing: true,
paused: false,
libraries: [{ libraryID: Zotero.Libraries.userLibraryID, indexed: 752, eligible: 9553 }]
items: { done: 700, total: 9000 },
attachments: { done: 52, total: 553, awaiting: 0 }
});
await select(win, col);
@ -368,7 +972,8 @@ describe("CollectionViewItemTree", function () {
enabled: true,
indexing: false,
paused: false,
libraries: [{ libraryID: Zotero.Libraries.userLibraryID, indexed: 752, eligible: 9553 }]
items: { done: 700, total: 9000 },
attachments: { done: 52, total: 553, awaiting: 0 }
});
await itemsView.setFilter('search', 'between runs query');
banner = win.document.querySelector('.best-match-index-banner');
@ -377,7 +982,8 @@ describe("CollectionViewItemTree", function () {
enabled: true,
indexing: false,
paused: true,
libraries: [{ libraryID: Zotero.Libraries.userLibraryID, indexed: 752, eligible: 9553 }]
items: { done: 700, total: 9000 },
attachments: { done: 52, total: 553, awaiting: 0 }
});
await itemsView.setFilter('search', 'paused query');
banner = win.document.querySelector('.best-match-index-banner');
@ -388,7 +994,8 @@ describe("CollectionViewItemTree", function () {
enabled: true,
indexing: false,
paused: false,
libraries: [{ libraryID: Zotero.Libraries.userLibraryID, indexed: 9553, eligible: 9553 }]
items: { done: 9000, total: 9000 },
attachments: { done: 553, total: 553, awaiting: 0 }
});
await itemsView.setFilter('search', 'another query');
assert.notOk(win.document.querySelector('.best-match-index-banner'));
@ -399,9 +1006,9 @@ describe("CollectionViewItemTree", function () {
let col2 = await createDataObject('collection');
let shared = await createDataObject('item', { collections: [col1.id, col2.id] });
let other = await createDataObject('item', { collections: [col2.id] });
let scoreStub = sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
let scoreStub = sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == shared.id ? 0.9 : 0.5]))
);
));
stubs.push(scoreStub);
await cv.selectByID("C" + col1.id);
@ -427,11 +1034,11 @@ describe("CollectionViewItemTree", function () {
let itemB = await createDataObject('item', { title: "savedsimtest B" });
// Install the stub first: creating the saved search auto-selects
// it, which already runs a best-match refresh
let stub = sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
let stub = sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(itemIDs.map((id) => {
let best = query == 'saved query' ? itemA.id : itemB.id;
return [id, id == best ? 0.9 : 0.5];
}))
})))
);
stubs.push(stub);
let search = new Zotero.Search();
@ -463,11 +1070,11 @@ describe("CollectionViewItemTree", function () {
// Install the stub first: creating the saved search auto-selects
// it, which already runs its top-K search
let scores = new Map([[kItem1.id, 0.9], [kItem2.id, 0.5], [colItem2.id, 0.7]]);
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(
itemIDs.filter(id => scores.has(id)).map(id => [id, scores.get(id)])
)
));
)));
let search = new Zotero.Search();
search.name = "Top-K best-match test";
search.libraryID = col.libraryID;
@ -514,73 +1121,27 @@ describe("CollectionViewItemTree", function () {
await itemsView.setFilter('advanced-search', null);
});
it("should rerank when the indexer announces changed embeddings", async function () {
let col = await createDataObject('collection');
let itemA = await createDataObject('item', { title: "rerank A", collections: [col.id] });
let itemB = await createDataObject('item', { title: "rerank B", collections: [col.id] });
let best = itemA.id;
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == best ? 0.9 : 0.5]))
));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
assert.deepEqual(itemsView._rows.map(row => row.id), [itemA.id, itemB.id]);
// The indexer's coalesced notification after new/removed vectors
best = itemB.id;
await Zotero.Notifier.trigger('refresh', 'item', [itemA.id, itemB.id], { embeddingsUpdate: true });
await itemsView._refreshPromise;
assert.deepEqual(itemsView._rows.map(row => row.id), [itemB.id, itemA.id]);
});
it("shouldn't rerank on a refresh that isn't an embeddings update", async function () {
it("shouldn't rerank on a refresh", async function () {
let col = await createDataObject('collection');
let itemA = await createDataObject('item', { title: "norerank A", collections: [col.id] });
let itemB = await createDataObject('item', { title: "norerank B", collections: [col.id] });
let best = itemA.id;
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(scoreEnvelope(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == best ? 0.9 : 0.5]))
));
)));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
assert.deepEqual(itemsView._rows.map(row => row.id), [itemA.id, itemB.id]);
// An unrelated refresh (e.g. a field change) leaves the ranking alone
// A refresh (e.g. a field change) leaves the ranking alone
best = itemB.id;
await Zotero.Notifier.trigger('refresh', 'item', [itemA.id, itemB.id]);
await itemsView._refreshPromise;
assert.deepEqual(itemsView._rows.map(row => row.id), [itemA.id, itemB.id]);
});
it("should rerank without re-running the search when embeddings change", async function () {
let col = await createDataObject('collection');
let itemA = await createDataObject('item', { title: "reuse A", collections: [col.id] });
let itemB = await createDataObject('item', { title: "reuse B", collections: [col.id] });
let best = itemA.id;
stubs.push(sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == best ? 0.9 : 0.5]))
));
await select(win, col);
itemsView = zp.itemsView;
await itemsView.setFilter('search', 'some query');
assert.deepEqual(itemsView._rows.map(row => row.id), [itemA.id, itemB.id]);
// The underlying search is dropped and re-run by clearing the row cache
let clearCacheSpy = sinon.spy(Zotero.CollectionTreeRow.prototype, 'clearCache');
stubs.push(clearCacheSpy);
best = itemB.id;
await Zotero.Notifier.trigger('refresh', 'item', [itemA.id, itemB.id], { embeddingsUpdate: true });
await itemsView._refreshPromise;
assert.deepEqual(itemsView._rows.map(row => row.id), [itemB.id, itemA.id]);
assert.isFalse(clearCacheSpy.called);
});
});
it("should expand parent item and attachment for an annotation match", async function () {

File diff suppressed because it is too large Load diff

View file

@ -922,4 +922,91 @@ describe("Zotero.FullText", function () {
);
});
});
describe("Item-text indexing", function () {
function itemTextMatch(match, cjk = false) {
let table = cjk ? 'fulltextItemTextCJK' : 'fulltextItemText';
return Zotero.DB.columnQueryAsync(
`SELECT rowid FROM ftindex.${table} WHERE ${table} MATCH ?`,
[match]
);
}
it("should index a regular item's title and abstract on save", async function () {
let item = await createDataObject('item', { title: 'Owl Migrátion Atlas' });
item.setField('abstractNote', 'Ztracking methods overview');
await item.saveTx();
// Words match diacritic-insensitively, in their own column
assert.include(await itemTextMatch('title:"migration"'), item.id);
assert.include(await itemTextMatch('abstract:"ztracking"'), item.id);
assert.notInclude(await itemTextMatch('title:"ztracking"'), item.id);
assert.ok(await Zotero.DB.valueQueryAsync(
"SELECT COUNT(*) FROM ftindex.fulltextItemTextState WHERE itemID=?",
item.id
));
});
it("should replace the indexed text when the item changes", async function () {
let item = await createDataObject('item', { title: 'Zfirsttitle here' });
item.setField('title', 'Zsecondtitle now');
await item.saveTx();
assert.include(await itemTextMatch('title:"zsecondtitle"'), item.id);
assert.notInclude(await itemTextMatch('title:"zfirsttitle"'), item.id);
});
it("should index an annotation's passage and comment on save", async function () {
let attachment = await importFileAttachment('test.pdf');
let annotation = await createAnnotation('highlight', attachment,
{ comment: 'zmethodology concern here' });
assert.include(await itemTextMatch('annotation:"zmethodology"'), annotation.id);
// Editing re-indexes
annotation.annotationComment = 'zrevised remark';
await annotation.saveTx();
assert.include(await itemTextMatch('annotation:"zrevised"'), annotation.id);
assert.notInclude(await itemTextMatch('annotation:"zmethodology"'), annotation.id);
});
it("should index a note's text into its note column when the note queue runs", async function () {
let note = new Zotero.Item('note');
note.setNote('<p>Znotable owl observations</p>');
await note.saveTx();
await Zotero.FullText.processNoteIndexQueue();
assert.include(await itemTextMatch('note:"znotable"'), note.id);
});
it("should index CJK item text as 2-grams", async function () {
let item = await createDataObject('item', { title: '疫情控制研究' });
assert.include(await itemTextMatch('title:"疫情"', true), item.id);
});
it("should clear an erased item's entries", async function () {
let item = await createDataObject('item', { title: 'Zephemeral title' });
assert.include(await itemTextMatch('title:"zephemeral"'), item.id);
await item.eraseTx();
assert.notInclude(await itemTextMatch('title:"zephemeral"'), item.id);
assert.equal(await Zotero.DB.valueQueryAsync(
"SELECT COUNT(*) FROM ftindex.fulltextItemTextState WHERE itemID=?",
item.id
), 0);
});
it("should backfill items missing from the index", async function () {
let item = await createDataObject('item', { title: 'Zbackfill target item' });
// Simulate an item that predates the index (e.g., after a rebuild)
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemText WHERE rowid=?", item.id);
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextItemTextState WHERE itemID=?", item.id);
assert.notInclude(await itemTextMatch('title:"zbackfill"'), item.id);
assert.isAbove(await Zotero.FullText.getItemTextIndexQueueCount(), 0);
await Zotero.FullText.processItemTextIndexQueue();
assert.include(await itemTextMatch('title:"zbackfill"'), item.id);
// The queue converges: everything eligible is indexed
assert.equal(await Zotero.FullText.getItemTextIndexQueueCount(), 0);
// The index stats report the same state
let stats = await Zotero.FullText.getIndexStats();
assert.equal(stats.itemTextQueue, 0);
assert.isAbove(stats.itemTextIndexed, 0);
});
});
})

671
test/tests/lexicalTest.js Normal file
View file

@ -0,0 +1,671 @@
"use strict";
describe("Zotero.Lexical", function () {
describe("#parseQuery()", function () {
it("should split a query into normalized word terms", function () {
let terms = Zotero.Lexical.parseQuery("OWL Migrátion");
assert.deepEqual(terms, [
{ type: 'word', text: 'owl', prefix: false },
{ type: 'word', text: 'migration', prefix: false }
]);
});
it("should flag a word the query put an asterisk after as a prefix", function () {
let terms = Zotero.Lexical.parseQuery("owl migr*");
assert.deepEqual(terms, [
{ type: 'word', text: 'owl', prefix: false },
{ type: 'word', text: 'migr', prefix: true }
]);
// Any word can carry one, not just the last
terms = Zotero.Lexical.parseQuery("migr* owl");
assert.isTrue(terms[0].prefix);
assert.isFalse(terms[1].prefix);
// An unfinished word is exact unless the query says otherwise
assert.isFalse(Zotero.Lexical.parseQuery("owl migr")[1].prefix);
});
it("should treat an asterisk inside quotes as literal text", function () {
assert.deepEqual(Zotero.Lexical.parseQuery('"migr*"'), [
{ type: 'word', text: 'migr', prefix: false }
]);
});
it("should keep a quoted multi-word part as one exact phrase", function () {
let terms = Zotero.Lexical.parseQuery('"united states" owl');
assert.deepEqual(terms, [
{ type: 'phrase', text: 'united states', tokens: ['united', 'states'] },
{ type: 'word', text: 'owl', prefix: false }
]);
});
it("should treat a quoted single word as an exact word", function () {
let terms = Zotero.Lexical.parseQuery('"owl"');
assert.deepEqual(terms, [
{ type: 'word', text: 'owl', prefix: false }
]);
});
it("should split a mixed-script part into word and CJK terms", function () {
let terms = Zotero.Lexical.parseQuery("covid疫情");
assert.deepEqual(terms, [
{ type: 'word', text: 'covid', prefix: false },
{ type: 'cjk', text: '疫情', bigrams: '疫情' }
]);
});
it("should carry a CJK run as its overlapping bigrams", function () {
let terms = Zotero.Lexical.parseQuery("疫情控制 ");
assert.deepEqual(terms, [
{ type: 'cjk', text: '疫情控制', bigrams: '疫情 情控 控制' }
]);
});
it("should carry a single CJK character with no bigrams", function () {
let terms = Zotero.Lexical.parseQuery("疫 ");
assert.deepEqual(terms, [
{ type: 'cjk', text: '疫', bigrams: null }
]);
});
it("should count a repeated term once, preferring its exact form", function () {
// One 'owl' carries an asterisk, but the query also asked for it
// exactly
let terms = Zotero.Lexical.parseQuery("owl migration owl*");
assert.deepEqual(terms, [
{ type: 'word', text: 'owl', prefix: false },
{ type: 'word', text: 'migration', prefix: false }
]);
});
it("should return no terms for text with nothing to match", function () {
assert.deepEqual(Zotero.Lexical.parseQuery(""), []);
assert.deepEqual(Zotero.Lexical.parseQuery(" "), []);
assert.deepEqual(Zotero.Lexical.parseQuery("!!! ..."), []);
});
});
describe("#buildExpression()", function () {
function pieces(queryText, family) {
let built = Zotero.Lexical.buildExpression(
Zotero.Lexical.parseQuery(queryText), family);
return built && built.pieces.map(
piece => piece.match + ' x' + piece.repetitions);
}
it("should OR the query's words together, repeated against its phrases", function () {
assert.deepEqual(pieces("special education ", 'word'), [
'"special" x3',
'"education" x3',
'"special education" x1'
]);
});
it("should add a phrase term for each consecutive pair of words", function () {
assert.deepEqual(pieces("fall of communism ", 'word'), [
'"fall" x3',
'"of" x3',
'"communism" x3',
'"fall of" x1',
'"of communism" x1'
]);
});
it("should expand the expression from the pieces and their repetitions", function () {
let built = Zotero.Lexical.buildExpression(
Zotero.Lexical.parseQuery("owl migration "), 'word');
assert.equal(built.expression,
'"owl" OR "owl" OR "owl" OR "migration" OR "migration" OR "migration" '
+ 'OR "owl migration"');
});
it("should carry a prefix into the word and its phrase", function () {
assert.deepEqual(pieces("special educat*", 'word'), [
'"special" x3',
'"educat"* x3',
'"special educat"* x1'
]);
});
it("should keep a quoted phrase as one term, unrepeated", function () {
assert.deepEqual(pieces('"united states" owl ', 'word'), [
'"united states" x1',
'"owl" x3'
]);
});
it("should build CJK runs from their bigrams, and nothing else", function () {
assert.deepEqual(pieces("疫情控制 ", 'cjk'), ['"疫情 情控 控制" x1']);
// A single character has no bigram, so it matches every bigram
// starting with it
assert.deepEqual(pieces("疫 ", 'cjk'), ['"疫"* x1']);
// Words never go to the 2-gram index
assert.isNull(pieces("owl ", 'cjk'));
});
it("should route a mixed-script query to both families", function () {
assert.deepEqual(pieces("covid疫情 ", 'word'), ['"covid" x3']);
assert.deepEqual(pieces("covid疫情 ", 'cjk'), ['"疫情" x1']);
});
it("should return null when the query has nothing for a family", function () {
assert.isNull(Zotero.Lexical.buildExpression([], 'word'));
assert.isNull(Zotero.Lexical.buildExpression([], 'cjk'));
});
});
describe("#buildProximityExpression()", function () {
function expression(queryText, family = 'word') {
return Zotero.Lexical.buildProximityExpression(
Zotero.Lexical.parseQuery(queryText), family);
}
it("should gather every word of a short query in one window", function () {
assert.equal(expression("fall communism "),
'NEAR("fall" "communism", 200)');
assert.equal(expression("owl migration patterns "),
'NEAR("owl" "migration" "patterns", 200)');
});
it("should let a longer query miss a word, in any position", function () {
assert.equal(expression("a1 b2 c3 d4 "), [
'NEAR("a1" "b2" "c3", 200)',
'NEAR("a1" "b2" "d4", 200)',
'NEAR("a1" "c3" "d4", 200)',
'NEAR("b2" "c3" "d4", 200)'
].join(' OR '));
});
it("should carry prefixes and phrases as terms", function () {
assert.equal(expression('owl* "barn owl" '),
'NEAR("owl"* "barn owl", 200)');
});
it("should gate each family on its own terms", function () {
assert.equal(expression("owl 猫头鹰 ", 'word'), null);
assert.equal(expression("猫头鹰 迁徙 ", 'cjk'),
'NEAR("猫头 头鹰" "迁徙", 200)');
});
it("should return null for a single term, which gathers anywhere", function () {
assert.isNull(expression("owl "));
assert.isNull(expression(""));
});
});
describe("#scoreItemIDs()", function () {
const BASE_ROWID = 940000000;
var inserted = [];
async function addContentDoc(id, text) {
let rowid = BASE_ROWID + id;
await Zotero.DB.queryAsync(
"INSERT INTO ftindex.fulltextContent (rowid, text) VALUES (?, ?)",
[rowid, Zotero.Utilities.Internal.normalizeForSearch(text) || '']
);
await Zotero.DB.queryAsync(
"REPLACE INTO ftindex.fulltextIndexState (itemID, version) VALUES (?, 1)",
[rowid]
);
inserted.push(rowid);
return rowid;
}
// A corpus of unrelated documents and items, so a test's words are
// rare in it: in a corpus of two, every query word is in more than
// half of it, and FTS5 weighs such a word as nothing
before(async function () {
let subjects = ['tides', 'granite', 'sonnets', 'bridges', 'yeast',
'glaciers', 'violins', 'ledgers', 'orchards', 'compasses'];
for (let i = 0; i < subjects.length; i++) {
await addContentDoc(900 + i,
`notes on ${subjects[i]} and their study through the years`);
await createDataObject('item', { title: `A survey of ${subjects[i]}` });
}
});
after(async function () {
for (let rowid of inserted) {
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextContent WHERE rowid=?", rowid);
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextIndexState WHERE itemID=?", rowid);
}
});
it("should match a document only where the query's words share a passage", async function () {
let near = await addContentDoc(60,
'the lexproxfall of lexproxcommunism in eastern europe');
let far = await addContentDoc(61,
'lexproxcommunism spread widely ' + 'filler '.repeat(300)
+ 'in the lexproxfall season the leaves drop');
// An item's own text is a passage already, so one word matches it
let item = await createDataObject('item',
{ title: 'Lexproxcommunism in the twentieth century' });
let scores = await Zotero.Lexical.scoreItemIDs(
'lexproxfall lexproxcommunism', [near, far, item.id]);
assert.isTrue(scores.has(near));
assert.isFalse(scores.has(far));
assert.isTrue(scores.has(item.id));
});
it("should admit a document saying most of a longer query in one passage", async function () {
let most = await addContentDoc(62,
'lexpcova lexpcovb and lexpcovc discussed here');
let half = await addContentDoc(63,
'lexpcova lexpcovb discussed here');
let scattered = await addContentDoc(64,
'lexpcova lexpcovb ' + 'filler '.repeat(300) + 'lexpcovc lexpcovd');
let scores = await Zotero.Lexical.scoreItemIDs(
'lexpcova lexpcovb lexpcovc lexpcovd', [most, half, scattered]);
// Three of four words in one passage is saying the query...
assert.isTrue(scores.has(most));
// ...two of four isn't, and neither is all four spread apart
assert.isFalse(scores.has(half));
assert.isFalse(scores.has(scattered));
});
it("should return scores as 0-1 fractions", async function () {
let item = await createDataObject('item',
{ title: 'Lexfrac owls of the northern coast' });
let scores = await Zotero.Lexical.scoreItemIDs('lexfrac ', [item.id]);
let score = scores.get(item.id);
assert.isAbove(score, 0);
assert.isAtMost(score, 1);
});
it("should rank the query's words adjacent above the words scattered", async function () {
let adjacent = await createDataObject('item',
{ title: 'Lexadjacent special education handbook' });
let scattered = await createDataObject('item',
{ title: 'Lexadjacent education for the special handbook' });
let scores = await Zotero.Lexical.scoreItemIDs(
'lexadjacent special education ', [adjacent.id, scattered.id]);
// Not clearing the floor is the weakest ranking there is, so a
// missing score reads as 0 either way
assert.isAbove(scores.get(adjacent.id), scores.get(scattered.id) ?? 0);
});
it("should rank coverage above partial matches, wherever they land", async function () {
// Full coverage in a title, partial coverage in a title, partial
// coverage in a document, and noise containing none of the query
let full = await createDataObject('item',
{ title: 'Lexsowl lexsmigration in the lexsunited lexsstates' });
let census = await createDataObject('item',
{ title: 'Lexsunited lexsstates census records' });
let norway = await addContentDoc(1,
'lexsowl lexsmigration routes across norway seasons');
let noise = await addContentDoc(2, 'lexsother archive entry');
let scores = await Zotero.Lexical.scoreItemIDs(
'lexsowl lexsmigration in the lexsunited lexsstates ',
[full.id, census.id, norway, noise]
);
// Full coverage outranks either partial match...
assert.isAbove(scores.get(full.id), scores.get(census.id) ?? 0);
assert.isAbove(scores.get(full.id), scores.get(norway) ?? 0);
// ...and matching nothing the query asked for is no match at all
assert.isFalse(scores.has(noise));
});
it("should rank a title match above the same words in a document", async function () {
// The item-text index is scored with per-column weights, so the
// column a match lands in moves the score
let titled = await createDataObject('item',
{ title: 'Lexcolumn owls of the coast' });
let document = await addContentDoc(3, 'lexcolumn owls of the coast');
let scores = await Zotero.Lexical.scoreItemIDs(
'lexcolumn owls ', [titled.id, document]);
assert.isAbove(scores.get(titled.id), scores.get(document));
});
it("should rank a document returning to a term above one passing mention", async function () {
let filler = Array.from({ length: 300 }, (x, i) => `lexsfill${i}`).join(' ');
let buried = await addContentDoc(4, `lexsdeep ${filler}`);
let focused = await addContentDoc(5, 'lexsdeep lexsdeep lexsdeep summary');
let scores = await Zotero.Lexical.scoreItemIDs('lexsdeep ', [buried, focused]);
assert.isAbove(scores.get(focused), scores.get(buried));
});
it("should only score the requested candidates", async function () {
let wanted = await addContentDoc(6, 'lexcandidate archive one');
await addContentDoc(7, 'lexcandidate archive two');
let scores = await Zotero.Lexical.scoreItemIDs('lexcandidate ', [wanted]);
assert.deepEqual([...scores.keys()], [wanted]);
});
it("should match a note the index doesn't have yet", async function () {
let note = new Zotero.Item('note');
note.setNote('<p>Lexunindexed owls in a note never indexed</p>');
await note.saveTx();
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextNoteIndexState WHERE itemID=?", note.id);
let scores = await Zotero.Lexical.scoreItemIDs('lexunindexed ', [note.id]);
assert.isAbove(scores.get(note.id), 0);
});
it("should match a just-edited note by its current text", async function () {
let note = new Zotero.Item('note');
note.setNote('<p>Lexoldword only</p>');
await note.saveTx();
await Zotero.FullText.indexItems([note.id]);
note.setNote('<p>Lexnewword only</p>');
await note.saveTx();
let scores = await Zotero.Lexical.scoreItemIDs('lexnewword ', [note.id]);
assert.isAbove(scores.get(note.id), 0);
});
it("should weigh terms only as far as the corpus can tell them apart", async function () {
// FTS5's BM25 floors the inverse document frequency of a term in
// more than about half a corpus, so in a corpus too small (or too
// uniform) for any term to be rare, every term weighs the same and
// ranking falls back to term frequency alone. Ordering survives;
// term importance doesn't.
let corpusSize = await Zotero.DB.valueQueryAsync(
"SELECT COUNT(*) FROM ftindex.fulltextItemText");
let common = await createDataObject('item',
{ title: 'Lexweigh lexweigh lexweigh repeated' });
let single = await createDataObject('item', { title: 'Lexweigh once only' });
let scores = await Zotero.Lexical.scoreItemIDs('lexweigh ', [common.id, single.id]);
// Whatever the corpus can say about the term, repetition still ranks
assert.isAbove(scores.get(common.id), scores.get(single.id) ?? 0);
assert.isAtLeast(corpusSize, 1);
});
it("should return nothing for a query with no terms", async function () {
assert.equal((await Zotero.Lexical.scoreItemIDs('', [1])).size, 0);
assert.equal((await Zotero.Lexical.scoreItemIDs('!!! ...', [1])).size, 0);
});
it("should return nothing for no candidates", async function () {
assert.equal((await Zotero.Lexical.scoreItemIDs('owl ', [])).size, 0);
});
it("should abandon scoring when cancelled", async function () {
let item = await createDataObject('item', { title: 'Lexscancel target' });
let e = await getPromiseError(Zotero.Lexical.scoreItemIDs(
'lexscancel ', [item.id], { shouldCancel: () => true }));
assert.instanceOf(e, Zotero.Lexical.ScoringCancelledError);
});
});
describe("#getMatchingExcerpts()", function () {
function rangeTexts(excerpt) {
return excerpt.ranges.map(([start, end]) => excerpt.text.slice(start, end));
}
it("should excerpt a regular item's title and abstract, marking matches as written", async function () {
let item = await createDataObject('item', { title: 'Lexkwic Owl Migrátion Atlas' });
// Long enough that the match sits past any excerpt window
item.setField('abstractNote',
'Filler sentences pad this abstract out well past the window. '.repeat(6)
+ 'Here the owl population finally appears.');
await item.saveTx();
let excerpts = await Zotero.Lexical.getMatchingExcerpts('owl migration ', item.id);
assert.lengthOf(excerpts, 2);
// The title shows both of the query's words, so it leads, whole
// and unelided, in the original casing and diacritics
let title = excerpts[0];
assert.equal(title.source, 'title');
assert.equal(title.text, 'Lexkwic Owl Migrátion Atlas');
assert.deepEqual(rangeTexts(title), ['Owl', 'Migrátion']);
// The abstract excerpt is a window around its match, elided at
// the start but running to the text's end
let abstract = excerpts[1];
assert.equal(abstract.source, 'abstract');
assert.isTrue(abstract.text.startsWith('…'));
assert.isFalse(abstract.text.endsWith('…'));
assert.deepEqual(rangeTexts(abstract), ['owl']);
// Strength is the share of the query's terms an excerpt shows
assert.equal(title.strength, 1);
assert.equal(abstract.strength, 0.5);
});
it("should highlight the whole word a prefix term matches", async function () {
let item = await createDataObject('item', { title: 'Lexkwic Migration Study' });
let excerpts = await Zotero.Lexical.getMatchingExcerpts('migr*', item.id);
assert.lengthOf(excerpts, 1);
assert.deepEqual(rangeTexts(excerpts[0]), ['Migration']);
});
it("should excerpt attachment content around a phrase", async function () {
this.timeout(60000);
let item = await createDataObject('item');
let attachment = await importPDFAttachment(item);
await Zotero.Fulltext.indexItems([attachment.id]);
// "easy-to-use" in the document: the phrase's spaces match its
// hyphens, and the range marks the hyphenated original
let excerpts = await Zotero.Lexical.getMatchingExcerpts('"easy to use"', attachment.id);
assert.isNotEmpty(excerpts);
assert.equal(excerpts[0].source, 'content');
assert.match(rangeTexts(excerpts[0])[0], /^easy[\s-]+to[\s-]+use$/i);
});
it("should excerpt a note's text without its markup", async function () {
let note = new Zotero.Item('note');
note.setNote('<p>Lexkwic owls fly at <b>night</b> over water</p>');
await note.saveTx();
let excerpts = await Zotero.Lexical.getMatchingExcerpts('lexkwic night ', note.id);
assert.lengthOf(excerpts, 1);
assert.equal(excerpts[0].source, 'note');
assert.notInclude(excerpts[0].text, '<');
assert.deepEqual(rangeTexts(excerpts[0]), ['Lexkwic', 'night']);
});
it("should excerpt an annotation's comment", async function () {
this.timeout(60000);
let item = await createDataObject('item');
let attachment = await importPDFAttachment(item);
let annotation = await createAnnotation('highlight', attachment,
{ comment: 'Lexkwic owls in the comment' });
let excerpts = await Zotero.Lexical.getMatchingExcerpts('lexkwic ', annotation.id);
assert.isNotEmpty(excerpts);
let comment = excerpts.find(excerpt => rangeTexts(excerpt)[0] == 'Lexkwic');
assert.equal(comment.source, 'annotation');
});
it("should cap the excerpts at the limit, strongest first", async function () {
// Every fixture word is lex-prefixed so it can't be turned up by
// another suite's quick search
let item = await createDataObject('item', { title: 'Lexcap lexheading' });
// Five matches, each in its own window's worth of filler
item.setField('abstractNote',
Array.from({ length: 5 },
(_, i) => `lexkwic sighting ${i} ` + 'filler '.repeat(40)).join(''));
await item.saveTx();
let excerpts = await Zotero.Lexical.getMatchingExcerpts('lexkwic ', item.id,
{ limit: 2 });
assert.lengthOf(excerpts, 2);
for (let excerpt of excerpts) {
assert.equal(excerpt.source, 'abstract');
assert.deepEqual(rangeTexts(excerpt), ['lexkwic']);
}
});
it("should return nothing for an empty query or an unmatched item", async function () {
let item = await createDataObject('item', { title: 'Lexkwic quiet title' });
assert.isEmpty(await Zotero.Lexical.getMatchingExcerpts('', item.id));
assert.isEmpty(await Zotero.Lexical.getMatchingExcerpts('absentword ', item.id));
});
});
describe("#findMatchRanges()", function () {
it("should mark where the query matches in given texts", async function () {
let [first, second, third] = await Zotero.Lexical.findMatchRanges(
'owl migration ',
['The Owl Migrátion Atlas', 'nothing relevant here', 'an owl alone']
);
assert.deepEqual(first, [[4, 7], [8, 17]]);
assert.deepEqual(second, []);
assert.deepEqual(third, [[3, 6]]);
});
it("should mark nothing for an empty query", async function () {
assert.deepEqual(
await Zotero.Lexical.findMatchRanges('', ['some text']),
[[]]
);
});
});
describe("#scoreTexts()", function () {
it("should score a text carrying the query above one that doesn't", async function () {
let [carrying, unrelated] = await Zotero.Lexical.scoreTexts(
'owl migration',
['owl migration patterns across the tundra', 'a note about something else']
);
assert.isAbove(carrying, unrelated);
assert.equal(unrelated, 0);
assert.isAtMost(carrying, 1);
});
it("should keep every score within 0-1", async function () {
let scores = await Zotero.Lexical.scoreTexts(
'owl',
['owl owl owl owl owl owl owl owl owl owl', 'owl']
);
for (let score of scores) {
assert.isAtLeast(score, 0);
assert.isAtMost(score, 1);
}
});
it("should weigh terms by their rarity within the given texts", async function () {
// Every passage of a dinosaur book says 'dinosaur', so among its
// passages the word separates nothing -- 'weight' is what picks
// out the passage answering the query, however loudly the others
// repeat the word the whole document is about
let spam = 'dinosaur dinosaur dinosaur dinosaur dinosaur everywhere';
let answer = 'the weight of a grown dinosaur';
let scores = await Zotero.Lexical.scoreTexts(
'dinosaur weight',
[spam, spam, spam, answer]
);
let best = Math.max(...scores);
assert.equal(scores.indexOf(best), 3);
});
it("should score nothing for a query with no terms", async function () {
assert.deepEqual(await Zotero.Lexical.scoreTexts('', ['owl']), [0]);
});
});
describe("#pickQuote()", function () {
const BASE_ROWID = 950000000;
const OPTIONS = { longSentence: 200, width: 150 };
var inserted = [];
var pick = text => Zotero.Lexical.pickQuote(
'lexweighrare lexweighcommon', text, Zotero.BestMatch.splitSentences(text), OPTIONS);
var said = (text, extent) => text.slice(extent.start, extent.end);
// A long sentence of filler with the given words spliced in where asked
var longSentence = (words) => {
let filler = 'the survey covered several plots over many seasons and noted everything seen there '
.repeat(4).trim().split(' ');
filler[0] = 'The';
for (let [index, word] of Object.entries(words)) {
filler[index] = word;
}
return filler.join(' ') + '.';
};
// Ten documents: one says the rare word, three the common one
before(async function () {
for (let i = 0; i < 10; i++) {
let words = ['filler text about unrelated matters'];
if (i < 3) {
words.push('lexweighcommon');
}
if (i == 0) {
words.push('lexweighrare');
}
let rowid = BASE_ROWID + i;
await Zotero.DB.queryAsync(
"INSERT INTO ftindex.fulltextContent (rowid, text) VALUES (?, ?)",
[rowid, words.join(' ')]
);
inserted.push(rowid);
}
});
after(async function () {
for (let rowid of inserted) {
await Zotero.DB.queryAsync(
"DELETE FROM ftindex.fulltextContent WHERE rowid=?", rowid);
}
});
it("should pick the sentence where the query weighs the most", async function () {
// The common word three times, then the rare one once: a repeat
// adds nothing, and the rarer word weighs more
let first = 'The lexweighcommon, the lexweighcommon and the lexweighcommon again.';
let second = 'Only the lexweighrare turns up in this one.';
let text = `${first} ${second}`;
let { sentence, line } = await pick(text);
assert.equal(said(text, sentence), second);
// A sentence that fits the line is quoted whole
assert.isUndefined(line);
});
it("should add up distinct terms, and take the first sentence on a tie", async function () {
let sentences = [
'An opening that says lexweighrare alone.',
'A middle that says lexweighrare and lexweighcommon both.',
'A close that says lexweighcommon and lexweighrare too.'
];
let text = sentences.join(' ');
assert.equal(said(text, (await pick(text)).sentence), sentences[1]);
});
it("should quote a long sentence as the line with its match toward the middle", async function () {
let sentence = longSentence({ 30: 'lexweighrare' });
let text = `A short sentence opens the passage. ${sentence}`;
let { line } = await pick(text);
let quoted = said(text, line);
// Whole words, at least the width of them
assert.isAtLeast(quoted.length, OPTIONS.width);
assert.match(text[line.start - 1], /\s/);
assert.isTrue(line.end == text.length || /\s/.test(text[line.end]));
let offset = quoted.indexOf('lexweighrare');
assert.isAbove(offset, 40);
assert.isBelow(offset, quoted.length - 40);
assert.isFalse(line.startsSentence);
});
it("should keep a long sentence's start when its match comes early", async function () {
let sentence = longSentence({ 2: 'lexweighrare' });
let text = `A short sentence opens the passage. ${sentence}`;
let { line } = await pick(text);
assert.equal(line.start, text.indexOf(sentence));
assert.isTrue(line.startsSentence);
assert.isBelow(line.end, text.length);
});
it("should end a long sentence's line on a match at its end, with what else fits", async function () {
let words = longSentence({}).split(' ');
let sentence = longSentence({ [words.length - 12]: 'lexweighcommon', [words.length - 1]: 'lexweighrare' });
let { line } = await pick(sentence);
let quoted = said(sentence, line);
assert.isTrue(quoted.endsWith('lexweighrare.'));
assert.include(quoted, 'lexweighcommon');
});
it("should pick nothing where no sentence says the query", async function () {
assert.isNull(await pick('Nothing here says it. Nor here.'));
assert.isNull(await pick(''));
});
});
});

View file

@ -156,41 +156,36 @@ describe("Advanced Preferences", function () {
})
describe("Best-Match Search", function () {
it("should confirm mode changes only when leaving an enabled mode", async function () {
var stubs = [
sinon.stub(Zotero.Embeddings.Indexing, 'startIndexing').resolves(),
sinon.stub(Zotero.Embeddings, 'pruneModels').resolves()
];
it("should show the indexing status only while semantic search is on", async function () {
var stub = sinon.stub(Zotero.Embeddings.Indexing, 'startIndexing').resolves();
var win = await loadPrefPane('advanced');
var menu = win.document.getElementById('semantic-search-model');
try {
// Enabling from Disabled prompts nothing (the test would hang on
// an unexpected modal prompt)
menu.value = 'bge-small-en-v1.5';
await win.Zotero_Preferences.Advanced.handleSemanticSearchModeChange();
assert.equal(Zotero.Prefs.get('embeddings.model'), 'bge-small-en-v1.5');
var menu = win.document.getElementById('semantic-search-enable');
var statusBox = win.document.getElementById('semantic-search-status');
assert.equal(menu.value, 'false');
assert.isTrue(statusBox.hidden);
// Cancelling a switch restores the menu and leaves the pref alone
var promise = waitForDialog(null, 'cancel');
menu.value = 'multilingual-e5-small';
await win.Zotero_Preferences.Advanced.handleSemanticSearchModeChange();
await promise;
assert.equal(Zotero.Prefs.get('embeddings.model'), 'bge-small-en-v1.5');
assert.equal(menu.value, 'bge-small-en-v1.5');
// Choosing Enabled writes a boolean, not the menu's string
menu.value = 'true';
menu.dispatchEvent(new win.Event('command'));
assert.isTrue(Zotero.Prefs.get('search.bestMatch.enableSemantic'));
// The pref observer re-renders the status block
while (statusBox.hidden) {
await Zotero.Promise.delay(10);
}
// Confirming applies the change
promise = waitForDialog();
menu.value = '';
await win.Zotero_Preferences.Advanced.handleSemanticSearchModeChange();
await promise;
assert.equal(Zotero.Prefs.get('embeddings.model'), '');
// And the menu follows the pref when something else writes it
Zotero.Prefs.set('search.bestMatch.enableSemantic', false);
while (!statusBox.hidden) {
await Zotero.Promise.delay(10);
}
assert.equal(menu.value, 'false');
}
finally {
win.close();
Zotero.Prefs.set('embeddings.model', '');
await Zotero.Embeddings.Indexing.waitForPendingModelSwitch();
Zotero.Prefs.clear('search.bestMatch.enableSemantic');
Zotero.Prefs.clear('embeddings.indexingPaused');
stubs.forEach(stub => stub.restore());
stub.restore();
}
});
})

View file

@ -31,6 +31,128 @@ describe("Zotero.SDT", function () {
assert.deepEqual(progress, []);
});
it("should cut a cached pack into chunks", async function () {
// Cold worker startup
this.timeout(60000);
let item = await importFileAttachment('test.pdf');
let pako = getTestRequire()('pako');
let bytes = makeTestSDTPackV1WithContent(documentWorkerMetadata, pako, {
outline: [
{ title: 'Introduction', ref: [1] },
],
pages: [
{ label: '1', contentRange: [[0], [2]] },
],
blocks: [
testBlock('Front matter on the title page', 0, [10, 700, 300, 720]),
testBlock('Introduction', 0, [10, 650, 300, 670], 'heading'),
testBlock('Owls are nocturnal birds of prey.', 0, [10, 600, 300, 620]),
],
});
await writeTestSDTCache(item, bytes);
let result = await Zotero.SDT.getItemChunks(item.id);
assert.isTrue(result.ok);
assert.equal(result.contentHash, TEST_PDF_HASH);
let { chunks } = result;
assert.lengthOf(chunks, 1);
assert.include(chunks[0].text, 'Front matter on the title page');
assert.include(chunks[0].text, 'Owls are nocturnal birds of prey.');
assert.isTrue(chunks[0].anchor.pageRects.every(rect => rect[0] === 0));
// A pack that isn't cached isn't generated when the caller says so
let other = await importFileAttachment('test.pdf');
assert.deepEqual(await Zotero.SDT.getItemChunks(other.id, { cachedOnly: true }),
{ ok: false, reason: 'not-cached' });
});
it("should read chunk anchors back as their text", async function () {
// Cold worker startup fetches the wasm runtime
this.timeout(60000);
let item = await importFileAttachment('test.pdf');
let pako = getTestRequire()('pako');
let bytes = makeTestSDTPackV1WithContent(documentWorkerMetadata, pako, {
outline: [
{ title: 'Introduction', ref: [0] },
{ title: 'Methods', ref: [2] },
],
pages: [
{ label: 'ix', contentRange: [[0], [2]] },
{ label: '10', contentRange: [[2], [4]] },
],
blocks: [
testBlock('Introduction', 0, [10, 700, 300, 720], 'heading'),
testBlock('Owls are nocturnal birds of prey. '.repeat(40), 0, [10, 600, 300, 680]),
testBlock('Methods', 1, [10, 700, 300, 720], 'heading'),
testBlock('We tracked forty owls with GPS loggers. '.repeat(40), 1, [10, 600, 300, 680]),
],
});
await writeTestSDTCache(item, bytes);
let { chunks } = await Zotero.SDT.getItemChunks(item.id, { positions: true });
assert.lengthOf(chunks, 2);
// Each chunk is anchored on its own page
assert.deepEqual(chunks.map(chunk => chunk.anchor.pageRects.map(rect => rect[0])), [[0], [1]]);
// Each chunk's anchor gives back its text, with where it sits -- the
// section it starts in and its page -- and the reader position to
// open it at, the one the chunk was cut with. Anchors on nothing the
// document has read back as nothing.
let anchors = [
chunks[0].anchor,
chunks[1].anchor,
null,
{ pageRects: [[5, 0, 0, 1, 1]] },
{ pageRects: [[0, 0, 0, 1, 1]] },
];
let spy = sinon.spy(Zotero.PDFWorker, 'readStructuredDocumentTextAnchors');
try {
var read = await Zotero.SDT.readAnchors(item.id, anchors);
assert.isTrue(spy.calledOnce);
}
finally {
spy.restore();
}
assert.isTrue(read.ok);
assert.lengthOf(read.chunks, 5);
assert.equal(read.chunks[0].text, chunks[0].text);
assert.equal(read.chunks[0].outlinePath, 'Introduction');
assert.equal(read.chunks[0].pageLabel, 'ix');
assert.equal(read.chunks[0].position.pageIndex, 0);
assert.equal(read.chunks[1].text, chunks[1].text);
assert.equal(read.chunks[1].outlinePath, 'Methods');
assert.equal(read.chunks[1].pageLabel, '10');
assert.equal(read.chunks[1].position.pageIndex, 1);
assert.isNull(read.chunks[2]);
assert.isNull(read.chunks[3]);
assert.isNull(read.chunks[4]);
assert.deepEqual(read.chunks[1].position, chunks[1].position);
// A worker failure is reported, not worked around
let stub = sinon.stub(Zotero.PDFWorker, 'readStructuredDocumentTextAnchors')
.rejects(new Error('Worker down'));
try {
assert.deepEqual(await Zotero.SDT.readAnchors(item.id, anchors),
{ ok: false, reason: 'failed' });
}
finally {
stub.restore();
}
});
it("should report the expected extraction identity from getProcessorVersion()", async function () {
let item = await importFileAttachment('test.pdf');
let version = await Zotero.SDT.getProcessorVersion(item);
let { CHUNKER_VERSION } = getTestRequire()(
'resource://zotero/document-worker/structured-document-text-chunker.js');
assert.equal(version, 'pdf/' + documentWorkerMetadata.SDT_PROCESSOR_VERSIONS.pdf
+ '/' + parseInt(documentWorkerMetadata.SDT_SCHEMA_VERSION)
+ '/' + CHUNKER_VERSION);
// Unsupported attachment types have no extraction identity
let unsupported = await importFileAttachment('test.png');
assert.isNull(await Zotero.SDT.getProcessorVersion(unsupported));
});
it("should generate the pack when missing", async function () {
let item = await importFileAttachment('test.pdf');
let cachePath = getSDTCachePath(item);
@ -180,6 +302,25 @@ describe("Zotero.SDT", function () {
}
});
it("should regenerate a stale-processor pack before returning it with allowStale: false", async function () {
let item = await importFileAttachment('test.pdf');
await writeTestSDTCache(item, getStaleProcessorVersionSDTPackBytes());
let workerStub = sinon.stub(Zotero.PDFWorker, 'getStructuredDocumentText')
.resolves({ buf: getTestSDTPackBuffer() });
try {
// A consumer that stores references into the pack's content gets
// the current extraction, never one about to be replaced
let result = await Zotero.SDT.getPack(item.id, { allowStale: false });
assert.isTrue(result.ok);
assert.deepEqual(new Uint8Array(result.bytes), getTestSDTPackBytes());
assert.isTrue(workerStub.calledOnce);
}
finally {
workerStub.restore();
}
});
it("should regenerate a pack with the wrong processor type", async function () {
let item = await importFileAttachment('test.pdf');
await writeTestSDTCache(item, getWrongProcessorTypeSDTPackBytes());
@ -294,6 +435,95 @@ describe("Zotero.SDT", function () {
assert.equal(result.reason, 'unavailable');
});
it("should cut chunks in the document worker", async function () {
// Cold worker startup
this.timeout(60000);
let item = await importFileAttachment('test.pdf');
let pako = getTestRequire()('pako');
let bytes = makeTestSDTPackV1WithContent(documentWorkerMetadata, pako, {
outline: [
{ title: 'Introduction', ref: [0] },
{ title: 'Methods', ref: [2] },
],
pages: [
{ label: 'ix', contentRange: [[0], [2]] },
{ label: '10', contentRange: [[2], [4]] },
],
blocks: [
testBlock('Introduction', 0, [10, 700, 300, 720], 'heading'),
testBlock('Owls are nocturnal birds of prey. '.repeat(40), 0, [10, 600, 300, 680]),
testBlock('Methods', 1, [10, 700, 300, 720], 'heading'),
testBlock('We tracked forty owls with GPS loggers. '.repeat(40), 1, [10, 600, 300, 680]),
],
});
await writeTestSDTCache(item, bytes);
let spy = sinon.spy(Zotero.PDFWorker, 'getStructuredDocumentTextChunks');
try {
let result = await Zotero.SDT.getItemChunks(item.id);
assert.isTrue(spy.calledOnce);
assert.isTrue(result.ok);
assert.equal(result.contentHash, TEST_PDF_HASH);
let { chunks } = result;
assert.lengthOf(chunks, 2);
assert.include(chunks[0].text, 'Owls are nocturnal');
assert.include(chunks[1].text, 'We tracked forty owls');
assert.deepEqual(chunks.map(chunk => chunk.outlinePath), ['Introduction', 'Methods']);
assert.deepEqual(chunks.map(chunk => chunk.anchor.pageRects.map(rect => rect[0])), [[0], [1]]);
assert.isFalse('position' in chunks[0]);
// Asked for, each chunk also carries the position it opens at
result = await Zotero.SDT.getItemChunks(item.id, { positions: true });
assert.isTrue(result.ok);
assert.deepEqual(result.chunks.map(chunk => chunk.position.pageIndex), [0, 1]);
assert.isFalse('positions' in result.chunks[0]);
}
finally {
spy.restore();
}
// A worker failure, or a chunker the metadata doesn't describe, is
// reported as a failed cut rather than worked around
let stub = sinon.stub(Zotero.PDFWorker, 'getStructuredDocumentTextChunks')
.rejects(new Error('Worker down'));
try {
assert.deepEqual(await Zotero.SDT.getItemChunks(item.id),
{ ok: false, reason: 'cut-failed' });
stub.resolves({ chunks: [], sourceHash: TEST_PDF_HASH, chunkerVersion: 999 });
assert.deepEqual(await Zotero.SDT.getItemChunks(item.id),
{ ok: false, reason: 'cut-failed' });
}
finally {
stub.restore();
}
});
it("should fail pending worker requests on a worker error and start afresh", async function () {
this.timeout(60000);
let item = await importFileAttachment('test.pdf');
await writeTestSDTCache(item);
// A request in flight when the worker errors
Zotero.PDFWorker._init();
let pending = Zotero.PDFWorker._query('sdt.readAnchors', { buf: new ArrayBuffer(0), anchors: [] }, []);
Zotero.PDFWorker._worker.dispatchEvent(new ErrorEvent('error', { message: 'Worker down' }));
let error = null;
try {
await pending;
}
catch (e) {
error = e;
}
assert.include(error?.message, 'Worker down');
assert.isNull(Zotero.PDFWorker._worker);
// The next request gets a new worker
let result = await Zotero.SDT.getItemChunks(item.id);
assert.isTrue(result.ok);
assert.isNotNull(Zotero.PDFWorker._worker);
});
it("should generate and open a pack with the real document worker", async function () {
// Cold worker startup fetches the wasm runtime and segmentation models
this.timeout(120000);
@ -317,12 +547,11 @@ describe("Zotero.SDT", function () {
assert.equal(progress.at(-1), 100);
assert.isTrue(progress.some(value => value > 0 && value < 100));
// getReader() should return a parsed pack from the cache without
// re-extracting
let reader = await Zotero.SDT.getReader(item.id);
assert.isOk(reader);
let metadata = await reader.getMetadata();
assert.equal(metadata.source.hash, TEST_PDF_HASH);
// The cached pack cuts in the worker without re-extracting
let cut = await Zotero.SDT.getItemChunks(item.id);
assert.isTrue(cut.ok);
assert.equal(cut.contentHash, TEST_PDF_HASH);
assert.isNotEmpty(cut.chunks);
});
function getSDTCachePath(item) {
@ -450,6 +679,80 @@ describe("Zotero.SDT", function () {
return bytes;
}
// A block of one text node laid out on a page, with the character
// geometry the chunker requires of a PDF: one run along the rect, each
// non-whitespace character an equal share of it
function testBlock(text, pageIndex, [x1, y1, x2, y2], type = 'paragraph') {
let count = text.replace(/\s/g, '').length;
let width = (x2 - x1) / Math.max(1, count);
return {
type,
anchor: { pageRects: [[pageIndex, x1, y1, x2, y2]] },
content: [{
text,
anchor: { textMap: JSON.stringify([[0, pageIndex, x1, y1, x2, y2, ...new Array(count).fill(width)]]) }
}]
};
}
// A v1 pack with real content blocks and catalog, for section/outline
// consumers (see makeEmptyTestSDTPackV1() for the layout)
function makeTestSDTPackV1WithContent(metadata, pako, { outline = [], pages = [], blocks = [] } = {}) {
if (metadata.SDT_PACK_VERSION !== 1) {
throw new Error('Unsupported test SDT pack version');
}
const HEADER_LENGTH = 16;
// Two entries each of chunk byte offsets and chunk block starts, after
// the metadata and catalog lengths
const INDEX_LENGTH = 8 + 2 * 4 + 2 * 4;
let encoder = new TextEncoder();
let schemaVersion = metadata.SDT_SCHEMA_VERSION.split('.').map(Number);
let metadataBytes = pako.deflateRaw(JSON.stringify({
processor: {
type: 'pdf',
version: metadata.SDT_PROCESSOR_VERSIONS.pdf,
},
dateCreated: '2026-01-01T00:00:00.000Z',
source: { hash: TEST_PDF_HASH },
}));
let catalogBytes = pako.deflateRaw(JSON.stringify({ pages, outline }));
// One content chunk: an offset table, then the block JSON back to back
let blockByteArrays = blocks.map(block => encoder.encode(JSON.stringify(block)));
let chunkBytes = new Uint8Array(
blocks.length * 4 + blockByteArrays.reduce((sum, b) => sum + b.byteLength, 0)
);
let chunkView = new DataView(chunkBytes.buffer);
let blockOffset = 0;
let writeOffset = blocks.length * 4;
for (let i = 0; i < blockByteArrays.length; i++) {
chunkView.setUint32(i * 4, blockOffset, true);
blockOffset += blockByteArrays[i].byteLength;
chunkBytes.set(blockByteArrays[i], writeOffset);
writeOffset += blockByteArrays[i].byteLength;
}
let compressedChunk = pako.deflateRaw(chunkBytes);
let payloadOffset = HEADER_LENGTH + INDEX_LENGTH;
let bytes = new Uint8Array(
payloadOffset + metadataBytes.byteLength + catalogBytes.byteLength
+ compressedChunk.byteLength
);
bytes.set(SDT_PACK_MAGIC, 0);
bytes.set([metadata.SDT_PACK_VERSION, ...schemaVersion], 8);
let view = new DataView(bytes.buffer);
view.setUint32(12, INDEX_LENGTH, true);
view.setUint32(HEADER_LENGTH, metadataBytes.byteLength, true);
view.setUint32(HEADER_LENGTH + 4, catalogBytes.byteLength, true);
// chunkByteOffsets [0, byteLength], chunkBlockStarts [0, blockCount]
view.setUint32(HEADER_LENGTH + 12, compressedChunk.byteLength, true);
view.setUint32(HEADER_LENGTH + 20, blocks.length, true);
bytes.set(metadataBytes, payloadOffset);
bytes.set(catalogBytes, payloadOffset + metadataBytes.byteLength);
bytes.set(compressedChunk,
payloadOffset + metadataBytes.byteLength + catalogBytes.byteLength);
return bytes;
}
function decodeBase64Bytes(base64) {
let binary = atob(base64);
let bytes = new Uint8Array(binary.length);

View file

@ -750,8 +750,9 @@ describe("Zotero.SearchQuery", function () {
assert.notInclude(conditions('fields'), 'fulltextContent');
assert.include(conditions('fields'), 'tag');
assert.include(conditions('everything'), 'fulltextContent');
// Best Match with no index falls back to matching text
assert.notInclude(conditions('bestMatch'), 'bestMatch');
// Best Match ranks by the text rather than filtering by it
assert.include(conditions('bestMatch'), 'bestMatch');
assert.notInclude(conditions('bestMatch'), 'tag');
});
it("should find items matching both the clauses and the text", async function () {

View file

@ -242,9 +242,12 @@ describe("Zotero.Search", function () {
sEmpty.addCondition('bestMatch', 'contains', '""');
assert.isFalse(sEmpty.getBestMatchQuery());
// With a cutoff, membership is the K most similar
let stub = sinon.stub(Zotero.Embeddings, 'scoreItemIDs').callsFake(
async (query, itemIDs) => new Map(itemIDs.map(id => [id, id == itemB.id ? 0.9 : 0.5]))
// With a cutoff, membership is the K most relevant
let stub = sinon.stub(Zotero.BestMatch, 'scoreItemIDs').callsFake(
async (query, itemIDs) => ({
scores: new Map(itemIDs.map(id => [id, id == itemB.id ? 0.9 : 0.5])),
matches: { lexical: new Set(itemIDs), semantic: new Set() }
})
);
try {
let s2 = new Zotero.Search();

View file

@ -513,7 +513,22 @@ describe("Zotero.Utilities.Internal", function () {
});
});
describe("#isLikelyReference()", function () {
it("should match a reference list's title in any case and punctuation", function () {
assert.isTrue(Zotero.Utilities.Internal.isLikelyReference('References'));
assert.isTrue(Zotero.Utilities.Internal.isLikelyReference('WORKS CITED:'));
assert.isTrue(Zotero.Utilities.Internal.isLikelyReference('Literaturverzeichnis'));
assert.isTrue(Zotero.Utilities.Internal.isLikelyReference('参考文献'));
});
it("should not match a string that only mentions references", function () {
assert.isFalse(Zotero.Utilities.Internal.isLikelyReference('Sources of law'));
assert.isFalse(Zotero.Utilities.Internal.isLikelyReference('References to prior work'));
assert.isFalse(Zotero.Utilities.Internal.isLikelyReference(''));
});
});
describe("#getNextName()", function () {
it("should get the next available numbered name", function () {
var existing = ['Name', 'Name 1', 'Name 3'];