A condition that matched a child item in the trash rolled up to its
parent when the search had a top-level (or other ancestor) result level,
or when "Include parent and child items of matching items" was checked.
https://forums.zotero.org/discussion/133836/
(cherry picked from commit d7751b93d8)
Any Field expands to a generic 'field' condition, which the cross-level
code treated as matching only on top-level items. At an attachment
result level it therefore matched attachments whose parent item had the
value, instead of attachments with the value themselves. Since the
condition stands in for every field, it's now treated as matching at any
level those fields live at.
(cherry picked from commit ad98e84d24)
These conditions expanded into an OR-group across their underlying
fields, so a "does not contain"/"is not" operator matched almost every
item: any item missing one of the fields satisfied that field's negated
condition. Use an AND-group for negative operators, so the value must be
absent from every field.
https://forums.zotero.org/discussion/132835/
Replace the trigram FTS5 index for attachment content with a unicode61
word index, so terms match whole words with the final token as a prefix
("archive" matches "archives", but "ion" doesn't match "condition"), as
in the pre-FTS5 word index. A multi-word phrase gets adjacent-token
candidates from the index and is then verified against the cached text
of just those items, since FTS5 ignores what separates adjacent tokens;
the verification treats whitespace and hyphen runs as equivalent
(they're frequently extraction layout or styling) but requires other
punctuation to match literally. Notes keep the trigram index and CJK
matching is unchanged; the index database version is bumped so the
index is rebuilt.
Follow-up to #5979
Note content is indexed into fulltext.sqlite, making note searches
accent- and case-insensitive and matching the note's plain text rather
than its HTML markup. To avoid re-indexing on every auto-save, a save
flags the note for background indexing, and searches match a flagged
note from its normalized text in memory until it's indexed.
Closes#378
Index attachment content into a contentless trigram FTS5 table in a
separate, attached fulltext.sqlite, normalized so matching is accent-
and case-insensitive. For content containing CJK characters, a companion
'ascii'-tokenized table holds bigrams so 1-2 character CJK queries, which
the trigram tokenizer can't match, still work. The extracted text still
lives in the .zotero-ft-cache files, so the index is fully derived and
rebuildable.
Use the FTS index for the fulltextContent condition, falling back to the
cached-text scan for queries too short to index, and point quick
search's content matching at the FTS index in place of the now-removed
word index. (One side effect: quick search now matches attachment
content by substring rather than by word.)
Already-extracted content is migrated into the index at startup, slowing
down on active usage. A background queue then extracts not-yet-indexed
attachments gradually when Zotero is idle. Attachments with no local
file or full-text content are recorded as missing. Content downloaded
via sync is processed into the index immediately when the sync finishes,
rather than waiting for idle like it did before, so it's searchable
immediately in on-demand file-download mode.
The index DB is tied to the main DB via the local user key and rebuilt
if they don't match (e.g., after a delete-and-resync). We compact it by
running FTS5's 'optimize' command once the indexing queue drains, and we
vacuum the attached database when necessary to reclaim disk space.
Closes#2038, #2044
Addresses #1595
Search now ignores accents, so "seance" matches "séance" and vice versa.
Text is normalized with Unicode NFKD compatibility decomposition (which
also handles typographic ligatures, superscripts, full-width forms,
etc.) plus a small map for letters NFKD leaves alone (ø, œ, æ, ß, ...)
and the fraction slash, via Z.Utilities.Internal.normalizeForSearch().
The HTML tags we support in item fields are stripped, so markup isn't
matched (#81). Typographic quotes (#29, #1876) and dashes are folded to
ASCII.
Each searchable column gets a normalized shadow column --
itemDataValues.valueNormalized, tags.nameNormalized,
creators.firstNameNormalized/lastNameNormalized, and
itemAnnotations.textNormalized/commentNormalized -- populated at write
time and matched via COALESCE(normalized, raw) LIKE. NULL is stored when
normalizing only changes case, so plain-ASCII values are only stored
once. This covers the contains/doesNotContain/beginsWith operators in
both quick search and Advanced Search.
The new columns are local-only derived data and aren't synced. Older
clients will ignore them, so this doesn't break DB compatibility.
Existing rows are backfilled after the startup sync by
Zotero.Schema.populateNormalizedSearchColumns(), which should only take
a few seconds on most databases.
Closes#29, #81, #1300, #1876
Notes and attachments are counted on regular items, and annotations on an
attachment or across a regular item's attachments; other rows are excluded
rather than always matching with a count of 0. Trashed children aren't
counted.
_rollUpAnyToLevel() only followed an annotation's parent when the
attachment itself had a parent, so in a search for top-level items, a
tag on a standalone attachment's annotation matched nothing.
The condition seeded by _loadConditions() used `mode: undefined`, but
toJSON() only omits the "/mode" suffix when the mode is exactly false
(what parseCondition() returns), so a saved search migrated from the
obsolete childNote condition serialized -- and synced -- the condition
as "resultLevel/undefined".
A field that exists at more than one level (e.g., Title, on both
top-level items and attachments) targeting a descendant result level
only matched the closest ancestor, so searching annotations by Title
found nothing, since it checked the parent attachment's title rather
than the top-level item's. Map down from each ancestor level and union
them, testing the predicate once so its bound parameters aren't
duplicated.
Addresses #5978
Add annotationType, annotationColor, and annotationAuthor conditions,
each tagged `level: 'annotation'` so the cross-level search logic maps
and negates them correctly. The value fields are drop-down menus: the
annotation types, the reader's color palette, and the library's
existing annotation authors.
Ported from #5839. The PR kept a negated annotation condition (e.g.,
"Annotation Color" "is not" "yellow") from matching every non-annotation
item by checking whether the condition name contained "annotation". This
does the same using the condition's level, which the cross-level logic
already handles, so the existing annotationText and annotationComment
conditions are covered too.
Fixes#5837
The flag forced a condition to be ANDed even in "any" mode, but
condition groups now express that directly. It was never exposed in the
search UI and nothing seems to have been using it.
addCondition()/updateCondition() now throw if passed a truthy
`required`. Not dropping the column now to preserve DB compatibility.
Opening Advanced Search from a non-empty quick search reproduces the
current quick search mode as editable conditions, one per word (or quoted
phrase) joined with "all":
- Title/Creator/Year: a single Title, Creator, Year condition per word,
with the result level set to item (the mode matches only top-level items)
- All Fields & Tags: a single Any Field condition per word
- Everything: Any Field + Full Text Content as an "any" group per word
(relies on grouped full-text composing correctly in SQL)
The Title/Creator/Year mode needs a condition to map to, so add a "Title,
Creator, Year" search condition that expands to the same field set as the
quick search mode (title, publication title, short title, court, year,
citation key, creator), mirroring how Any Field matches All Fields & Tags.
Like Any Field, it expands at query-build time, so the saved search stores
a single condition and its sub-fields don't need their own entries in the
condition menu.
An Any Field condition expands into field/tag/note/creator conditions for
the term. Splice the expansion in right after the condition so it stays at
the same nesting depth, rather than appending it to the end of the
processing queue, which would emit it at the top level instead of within
its group. (Top-level Any Field is unaffected.)
Add annotation text and comments to the conditions the Any Field
condition expands to, matching the fields covered by the All Fields &
Tags quick search mode, as the comment already says is intended. Key
detection and quoted-phrase splitting still differ, since those depend
on the quick search string parsing that Any Field doesn't do.
Give a search a result level -- top-level item, attachment, note, or
annotation -- and map every condition to that level, so one search can
mix conditions that match at different levels of the item hierarchy
(e.g., a top-level item with a given author and a red annotation on one
of its PDFs). Each condition carries the level(s) it matches at: a match
is mapped up to an ancestor or down to a descendant, level-agnostic
conditions (tags) roll up to the result level, and fields that exist on
both items and attachments (title, url, accessDate) match natively at
either. The result level is stored as a `resultLevel` marker condition
alongside the join mode.
This also removes the temporary annotation-parent hacks, which the
general cross-level mapping replaces.
A fulltextContent condition was evaluated as a global post-filter keyed
on the search's top-level join mode -- correct for a top-level
condition, but not for one inside a group, which must combine with its
siblings under the group's own join mode. The new condition grouping UI
allows fulltextContent to be placed within groups, and we need to do so
to prefill the advanced-search pane from an "Everything" quicksearch.
Materialize a grouped fulltextContent into an itemID set and emit it as an
ordinary itemID IN/NOT IN predicate, so combineConditions composes it
under the group's join mode. Top-level fulltextContent keeps the existing
post-filter unchanged.
This also removes the quicksearch full-text post-filter special case. A
quick search puts its full-text in per-word "any" groups, so the
_hasQuicksearch flag was needed to make the post-filter union those
matches rather than intersect them under the top-level "all" join. Now
that the grouped full-text is composed in SQL it never reaches the
post-filter, so the flag is gone.
Replace the flat anySQL/quicksearch-block assembly in _buildQuery with a
recursive tree of AND/OR groups, built and reduced by a new
Zotero.Search.combineConditions() helper. groupStart/groupEnd markers
delimit nested groups and a joinMode marker sets each group's mode, so a
saved search can combine conditions with arbitrary nesting and per-group
join modes. The per-condition SQL generation is unchanged.
For example, a search built as
search.addCondition('joinMode', 'all');
search.addCondition('title', 'contains', 'foo');
search.addCondition('groupStart', 'true', '');
search.addCondition('joinMode', 'any');
search.addCondition('tag', 'is', 'x');
search.addCondition('tag', 'is', 'y');
search.addCondition('groupEnd', 'true', '');
means "title contains 'foo' AND (tag is 'x' OR tag is 'y')". The 'true'
operator on the group markers is an unused placeholder -- they carry no
value, but a condition's operator can't be empty.
The quick search (matching multiple words) and the Any Field condition
previously had their own special handling in the query builder; they now
use the same grouping as everything else, so that code is gone. Behavior
for existing non-grouped searches is unchanged; new tests cover nested
groups and combineConditions() directly.
When conditions are removed, shift conditionIDs so that
conditionIDs always go in increments of 1 (0, 1, 2 ...).
It prevents conditionIDs from conflicting with each other
when conditions are rearranged.
Fixes: #3434
- Annotations are displayed in itemTree under their file attachments
on a third level. The annotation spans the entire row.
- The title is constructed on the go. When possible, it
includes annotation quote and comment as pseudo-columns
of the row. The comment occupies about twice as much space as
the quote. Otherwise, (if there is no quote) only the annotation
comment
is included as the "title" part of the row
- Non-CJK segments of the quote part of the annotation row are
italicized. CJK segments are left as is. If CJK segments are present,
there is more padding between quote and comment parts.
- Search matches the actual attachment instead of its parent file.
- Can create child notes from annotations of the same item or
standalone notes from annotations across different items from
the context menu or the header button.
- When an annotation (or multiple annotations) are selected, the
annotationItemPane component is displayed where annotations are
grouped by their top-level item. Annotations are displayed fully,
without having their content cut off.
- Special treatment for annotations to always prompt
to erase the item regardless of what collectionTree row
is selected (e.g., if a collection is selected, we
still want one to be able to delete the annotation).
This only applies if all selected items are annotations.
If multiple items are selected, some annotations and
some not, do nothing. This is until the trash is
ready. In the future, we may send annotations to trash
- strip all HTML tags from annotation for now, until the logic to
properly render annotation markup is copied over from the reader
(applies to both annotation-row component and the annotation
item rendered in the itemTree)
- Added a generalized "expandToItem" function to itemTree to
expand all ancestors of a given item, similar to "expandToCollection"
from collectionTree
- add annotation conditions to advanced search
- show [Image not available] if no annotation file for ink or image
annotations
- only keep annotation-specific context menu options when some
annotations are selected in itemTree
- enable Quick Copy of annotations from itemTree via drag-drop,
shortcut key, or Edit → Copy Annotation
- Minor refactoring of Zotero.Annotation.toJSON() to pull out async code
that handles ink and image annotations, so that
Zotero.Annotation.toJSONsync() for highlight, underline, and note
annotations does not have to be awaited. Since ink and image
annotation don't seem to work for drag-drop Quick Copy, they are just
skipped for now.
A browser window was no longer actually needed for charset detection
(on macOS, at least, and hopefully elsewhere) because we switched to a
HiddenFrame-based hidden browser. Remaining uses now call
`loadZoteroWindow()`.
Fix the glitch where having anyField search condition
along with any other condition would return an empty
result set if joinMode="any".
This would happen due to a conflict between conditions
caused by wrapping "anyField" condition set into a quickSearch
block, which enforces "AND" operator between conditions.
Fixes: #4830
- revert change from 2401a34031
that only loads un-trashed collections in _loadCollections.
If an item only belongs to deleted collections, item._loaded.collections = true
from _loadCollections will never run, so an exception
will be thrown in item.toJSON() when syncing happens.
Instead, to address the problem of item.getCollections()
having stale data #4307, add 'includeTrashed' parameter to
item.getCollections() based on which item._collections
will be filtered. Fixes: #4346
- revert earlier, no more necessary, changes from a532cfb475
to not alter item._collections cache when collections are being trashed or restored.
Collection is removed from item._collections only when it is permanently
erased.
- removed unnecessary test checking for consistent item._collections
value before and after reload, since item._collections is no longer
modified
- fix encountered bug where a trashed child collection is not
unloaded if a parent collection is erased without being trashed first.
- tweaked Zotero.Search sql construction to count items
that only belong to trashed collections into 'unfiled'. Fixes: #4347
---------
Co-authored-by: Dan Stillman
When a collection or a saved search is deleted, it appears in
trash among other trashed items. From there, it can be restored
or permanently deleted.
Items of trashed collections are not affected my the trashing/permanent
deletion of a collection and need to be deleted separately like before.
Subcollections of a trashed collection do not appear in the trash and
are restored or permanently deleted with the top-most trashed parent.
Prevents bug in zotero-citation plugin (at least on macOS) from creating
a search that breaks syncing
We were already checking for a missing name in `saveTx()`, but the
plugin is saving the same search twice in rapid succession, the second
time without a name, and the second attempt clears the search object's
name value after the first save's `_initSave()` check and before its SQL
write. The second save fails, but the first save goes through without a
name, resulting in a sync error.
https://forums.zotero.org/discussion/104274/id-1702002152-cannot-synchttps://github.com/MuiseDestiny/zotero-citation/issues/31
Expose annotation tags in tag selector and match parent attachments when
filtering/searching
This also fixes searching for annotation text or comments when using
Everything quick search.
This is temporary until we display annotations in the items list
directly.
This lays the groundwork for moving collections and searches to the
trash instead of deleting them outright. We're not doing that yet, so
the `deleted` property will never be set (except for items), but this
will allow clients from this point forward to sync collections and
searches with that property for when it's used in the future. For now,
such objects will just be hidden from the collections pane as if they
had been deleted.
This changes the way item types, item fields, creator types, and CSL
mappings are defined and handled, in preparation for updated types and
fields.
Instead of being predefined in SQL files or code, type/field info is
read from a bundled JSON file shared with other parts of the Zotero
ecosystem [1], referred to as the "global schema". Updates to the
bundled schema file are automatically applied to the database at first
run, allowing changes to be made consistently across apps.
When syncing, invalid JSON properties are now rejected instead of being
ignored and processed later, which will allow for schema changes to be
made without causing problems in existing clients. We considered many
alternative approaches, but this approach is by far the simplest,
safest, and most transparent to the user.
For now, there are no actual changes to types and fields, since we'll
first need to do a sync cut-off for earlier versions that don't reject
invalid properties.
For third-party code, the main change is that type and field IDs should
no longer be hard-coded, since they may not be consistent in new
installs. For example, code should use `Zotero.ItemTypes.getID('note')`
instead of hard-coding `1`.
[1] https://github.com/zotero/zotero-schema
- Fix incorrect results for ANY search with multiple "Attachment
Content" conditions and no other conditions
- Dramatically speed up single-word searches by avoiding unnecessary
text scans (which probably addresses #1595)
- Clean up code