mirror of
https://github.com/zotero/zotero.git
synced 2026-10-07 02:58:09 +00:00
Mirror of https://github.com/zotero/zotero.git
- added environment to run embedding models locally (transformers.js, ONNX Runtime WASM binary, etc.). The actual inference execution happens in a separate worker environment (worker.js) - added local itemEmbeddings table to store embeddings locally - in advanced preferences, one can select two options for semantic search model: english and multilingual. English model (bge-small-en-v1.5) is better for english-only corpus but multilingual (multilingual-e5-small) is necessary to handle abstracts with any other language than english. We can add more language-specific models as needed. - when the model is selected, Zotero.Embeddings.download will download the model (quantized ~100mb) and store it locally. - Zotero.Embeddings.Indexing will start a process to index all regular items with title+abstract. It happens in batches and takes some time. The progress will appear in the advanced preferences pane. Embeddings are inserted into itemEmbeddings SQL table. For now, the table is local only, no syncing is involved. - when embedding model pref is set to "Disabled", the model is deleted and embeddings table is cleared. - when an embedding model is selected, quick search dropdown has a new "Similarity" mode, which will run semantic search on the current scope of items. - semantic search does not clearly define what counts as "relevant" and what is "not relevant". In addition, it will change depending on the library and query. So we cannot semantically filter out items the way it is done via SQL. Semantic search returns the ranking but items cannot be sorted because it is done by the itemTree based on column selection. So in "similarity" quicksearch mode, there is also a dropdown to select how many top relevant items to keep (top 5 - top 100). It allows the user to keep the most relevant items depending on the context, without conflicting with itemTree sorting. - semantic search happens in-memory. On a large 5K library it's fast, but we could consider sqlite-vec extension if needed. |
||
|---|---|---|
| .github/workflows | ||
| app | ||
| chrome | ||
| defaults/preferences | ||
| document-worker@6d0c0ce45d | ||
| js-build | ||
| note-editor@acec74d09b | ||
| reader@c6edbfff75 | ||
| resource | ||
| scripts | ||
| scss | ||
| styles@dff7452b24 | ||
| test | ||
| translators@60f2d542ff | ||
| types/gecko | ||
| .babelrc | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| chrome.manifest | ||
| CLAUDE.md | ||
| CONTRIBUTING.md | ||
| COPYING | ||
| eslint.config.mjs | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| update.rdf | ||
| version | ||
Zotero
Zotero is a free, easy-to-use tool to help you collect, organize, cite, and share your research sources.
Please post feature requests or bug reports to the Zotero Forums. If you're having trouble with Zotero, see Getting Help.
For more information on how to use this source code, see the Zotero documentation.