mirror of
https://github.com/zotero/zotero.git
synced 2026-10-07 02:58:09 +00:00
Mirror of https://github.com/zotero/zotero.git
Move the generation of mean vectors into its own Zotero.Embeddings.Calibration object from tests. Now a test run is not required to get mean vector for a newly added model. Instead, the corpus on which we generate embeddings to calculate mean vector, as well as the logic for actual mean vector calculation, lives in Zotero.Embeddings.Calibration so it can be done on demand - at the start of the first indexing pass for a model that hasn't been measured yet. Each passage from Zotero.Embeddings.Calibration corpus also has a matching query, to which the passage is an expected search result. Those queries are also embedded during calibration, which is used to calculate the appropriate min score cutoff and max display score for relevance bar. Score floor is computed by embedding all queries and passages, finding the similarity between all pairs and locating the cutoff that only 1% of non-matching query/passage pairs exceed. Max display score is the median of the matching pairs, so half of genuine matches fill the bar completely. Calculated score cutoff, max display value and mean vector are stored in the embeddings database to be reused later. During calibration of language-specific models, passages/queries for languages that the model cannot handle are excluded (e.g drop russian/chinese/spanish that english models are not meant to handle). For this, model config expects a new optional language field. A model that omits it is measured against the whole corpus. Chunk size is now bounded by CHUNK_MAX_TOKENS rather than by the model's context window, which still caps it. A model with a large window would otherwise stop notes being chunked at all, putting a whole note into a single vector and defeating scoring an item by its best chunk. With the mean vector and both score bounds measured rather than configured, each model's hardcoded meanVector, minScore and maxDisplayScore are gone from MODELS, along with an unused files array. A model entry is now only facts that can be read off its model card. Added bge-small-zh-v1.5 as a model specifically for chinese, as well as a number of multilingual and english models for testing. |
||
|---|---|---|
| .github/workflows | ||
| app | ||
| chrome | ||
| defaults/preferences | ||
| document-worker@6d0c0ce45d | ||
| js-build | ||
| note-editor@acec74d09b | ||
| reader@c6edbfff75 | ||
| resource | ||
| scripts | ||
| scss | ||
| styles@dff7452b24 | ||
| test | ||
| translators@60f2d542ff | ||
| types/gecko | ||
| .babelrc | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| chrome.manifest | ||
| CLAUDE.md | ||
| CONTRIBUTING.md | ||
| COPYING | ||
| eslint.config.mjs | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| update.rdf | ||
| version | ||
Zotero
Zotero is a free, easy-to-use tool to help you collect, organize, cite, and share your research sources.
Please post feature requests or bug reports to the Zotero Forums. If you're having trouble with Zotero, see Getting Help.
For more information on how to use this source code, see the Zotero documentation.