Commit graph

65 commits

Author SHA1 Message Date
Dan Stillman
d31051f85b Don't match an item by a trashed child item in a search
A condition that matched a child item in the trash rolled up to its
parent when the search had a top-level (or other ancestor) result level,
or when "Include parent and child items of matching items" was checked.

https://forums.zotero.org/discussion/133836/
(cherry picked from commit d7751b93d8)
2026-09-29 11:40:57 -04:00
Dan Stillman
95a6d8821e Fix Any Field searches at a non-item result level
Any Field expands to a generic 'field' condition, which the cross-level
code treated as matching only on top-level items. At an attachment
result level it therefore matched attachments whose parent item had the
value, instead of attachments with the value themselves. Since the
condition stands in for every field, it's now treated as matching at any
level those fields live at.

(cherry picked from commit ad98e84d24)
2026-08-21 12:56:53 -04:00
Dan Stillman
e81e72dccf Fix negated Title/Creator/Year and Any Field advanced searches
These conditions expanded into an OR-group across their underlying
fields, so a "does not contain"/"is not" operator matched almost every
item: any item missing one of the fields satisfied that field's negated
condition. Use an AND-group for negative operators, so the value must be
absent from every field.

https://forums.zotero.org/discussion/132835/
2026-07-21 14:15:45 -04:00
Dan Stillman
0ce289a7fe Use a word index for full-text content search
Replace the trigram FTS5 index for attachment content with a unicode61
word index, so terms match whole words with the final token as a prefix
("archive" matches "archives", but "ion" doesn't match "condition"), as
in the pre-FTS5 word index. A multi-word phrase gets adjacent-token
candidates from the index and is then verified against the cached text
of just those items, since FTS5 ignores what separates adjacent tokens;
the verification treats whitespace and hyphen runs as equivalent
(they're frequently extraction layout or styling) but requires other
punctuation to match literally. Notes keep the trigram index and CJK
matching is unchanged; the index database version is bumped so the
index is rebuilt.

Follow-up to #5979
2026-07-17 12:41:19 -04:00
Dan Stillman
3ec407d861 Add full-text searching of note content
Note content is indexed into fulltext.sqlite, making note searches
accent- and case-insensitive and matching the note's plain text rather
than its HTML markup. To avoid re-indexing on every auto-save, a save
flags the note for background indexing, and searches match a flagged
note from its normalized text in memory until it's indexed.

Closes #378
2026-07-15 15:36:40 -04:00
Dan Stillman
7c2a1d127d Add full-text content search via FTS5
Index attachment content into a contentless trigram FTS5 table in a
separate, attached fulltext.sqlite, normalized so matching is accent-
and case-insensitive. For content containing CJK characters, a companion
'ascii'-tokenized table holds bigrams so 1-2 character CJK queries, which
the trigram tokenizer can't match, still work. The extracted text still
lives in the .zotero-ft-cache files, so the index is fully derived and
rebuildable.

Use the FTS index for the fulltextContent condition, falling back to the
cached-text scan for queries too short to index, and point quick
search's content matching at the FTS index in place of the now-removed
word index. (One side effect: quick search now matches attachment
content by substring rather than by word.)

Already-extracted content is migrated into the index at startup, slowing
down on active usage. A background queue then extracts not-yet-indexed
attachments gradually when Zotero is idle. Attachments with no local
file or full-text content are recorded as missing. Content downloaded
via sync is processed into the index immediately when the sync finishes,
rather than waiting for idle like it did before, so it's searchable
immediately in on-demand file-download mode.

The index DB is tied to the main DB via the local user key and rebuilt
if they don't match (e.g., after a delete-and-resync). We compact it by
running FTS5's 'optimize' command once the indexing queue drains, and we
vacuum the attached database when necessary to reclaim disk space.

Closes #2038, #2044
Addresses #1595
2026-07-15 15:36:40 -04:00
Dan Stillman
3bd8d641b8 Make item search accent-insensitive
Search now ignores accents, so "seance" matches "séance" and vice versa.

Text is normalized with Unicode NFKD compatibility decomposition (which
also handles typographic ligatures, superscripts, full-width forms,
etc.) plus a small map for letters NFKD leaves alone (ø, œ, æ, ß, ...)
and the fraction slash, via Z.Utilities.Internal.normalizeForSearch().
The HTML tags we support in item fields are stripped, so markup isn't
matched (#81). Typographic quotes (#29, #1876) and dashes are folded to
ASCII.

Each searchable column gets a normalized shadow column --
itemDataValues.valueNormalized, tags.nameNormalized,
creators.firstNameNormalized/lastNameNormalized, and
itemAnnotations.textNormalized/commentNormalized -- populated at write
time and matched via COALESCE(normalized, raw) LIKE. NULL is stored when
normalizing only changes case, so plain-ASCII values are only stored
once. This covers the contains/doesNotContain/beginsWith operators in
both quick search and Advanced Search.

The new columns are local-only derived data and aren't synced. Older
clients will ignore them, so this doesn't break DB compatibility.
Existing rows are backfilled after the startup sync by
Zotero.Schema.populateNormalizedSearchColumns(), which should only take
a few seconds on most databases.

Closes #29, #81, #1300, #1876
2026-07-15 15:36:40 -04:00
nexdep
2c6185d8f1
Add Attachment Storage Type search condition (#5875)
---------

Co-authored-by: Dan Stillman <dstillman@zotero.org>
2026-07-13 15:57:37 -04:00
Dan Stillman
cdb9134537 Add "# of Notes", "# of Attachments", and "# of Annotations" search conditions
Notes and attachments are counted on regular items, and annotations on an
attachment or across a regular item's attachments; other rows are excluded
rather than always matching with a count of 0. Trashed children aren't
counted.
2026-07-13 14:18:59 -04:00
Dan Stillman
7cdd74bd2b Add "# of Tags" advanced search condition
Closes #158
2026-07-13 13:38:22 -04:00
Dan Stillman
24fe50cfd5 Add "is empty"/"is not empty" advanced search operators
Available for text, date, and number fields and creators. Previously only
possible via a doesNotContain hack with an empty value.
2026-07-13 13:38:22 -04:00
Dan Stillman
cc28fad6da Roll a standalone attachment's annotations up to the attachment itself
_rollUpAnyToLevel() only followed an annotation's parent when the
attachment itself had a parent, so in a search for top-level items, a
tag on a standalone attachment's annotation matched nothing.
2026-07-10 11:45:28 -04:00
Dan Stillman
d95d8f294d Fix serialization of the result level seeded for a migrated childNote
The condition seeded by _loadConditions() used `mode: undefined`, but
toJSON() only omits the "/mode" suffix when the mode is exactly false
(what parseCondition() returns), so a saved search migrated from the
obsolete childNote condition serialized -- and synced -- the condition
as "resultLevel/undefined".
2026-07-10 11:45:28 -04:00
Dan Stillman
02b19d0cde Match multi-level search condition against any ancestor level for descendant results
Some checks failed
CI / Build, Upload, Test (push) Has been cancelled
A field that exists at more than one level (e.g., Title, on both
top-level items and attachments) targeting a descendant result level
only matched the closest ancestor, so searching annotations by Title
found nothing, since it checked the parent attachment's title rather
than the top-level item's. Map down from each ancestor level and union
them, testing the predicate once so its bound parameters aren't
duplicated.

Addresses #5978
2026-06-26 16:53:52 -04:00
Dan Stillman
cf7ee984ea Advanced search: Hide obsolete Child Note condition and migrate to Note
Addresses #5978
2026-06-26 15:53:07 -04:00
Dan Stillman
b56c153125 Fix Attachment Last Read search condition by matching at attachment level
Addresses #5978
2026-06-26 15:50:17 -04:00
Dan Stillman
2717aafedd Add annotation type, color, and author search conditions
Add annotationType, annotationColor, and annotationAuthor conditions,
each tagged `level: 'annotation'` so the cross-level search logic maps
and negates them correctly. The value fields are drop-down menus: the
annotation types, the reader's color palette, and the library's
existing annotation authors.

Ported from #5839. The PR kept a negated annotation condition (e.g.,
"Annotation Color" "is not" "yellow") from matching every non-annotation
item by checking whether the condition name contained "annotation". This
does the same using the condition's level, which the cross-level logic
already handles, so the existing annotationText and annotationComment
conditions are covered too.

Fixes #5837
2026-06-26 13:08:15 -04:00
Dan Stillman
60807d552c Remove the search condition required flag (#5962)
The flag forced a condition to be ANDed even in "any" mode, but
condition groups now express that directly. It was never exposed in the
search UI and nothing seems to have been using it.

addCondition()/updateCondition() now throw if passed a truthy
`required`. Not dropping the column now to preserve DB compatibility.
2026-06-25 16:14:38 -04:00
Dan Stillman
ccf6f18643 Prefill Advanced Search from the quick search (#5962)
Opening Advanced Search from a non-empty quick search reproduces the
current quick search mode as editable conditions, one per word (or quoted
phrase) joined with "all":

- Title/Creator/Year: a single Title, Creator, Year condition per word,
  with the result level set to item (the mode matches only top-level items)
- All Fields & Tags: a single Any Field condition per word
- Everything: Any Field + Full Text Content as an "any" group per word
  (relies on grouped full-text composing correctly in SQL)

The Title/Creator/Year mode needs a condition to map to, so add a "Title,
Creator, Year" search condition that expands to the same field set as the
quick search mode (title, publication title, short title, court, year,
citation key, creator), mirroring how Any Field matches All Fields & Tags.
Like Any Field, it expands at query-build time, so the saved search stores
a single condition and its sub-fields don't need their own entries in the
condition menu.
2026-06-25 16:14:38 -04:00
Dan Stillman
786b85bf2d Expand Any Field in place so it nests correctly (#5962)
An Any Field condition expands into field/tag/note/creator conditions for
the term. Splice the expansion in right after the condition so it stays at
the same nesting depth, rather than appending it to the end of the
processing queue, which would emit it at the top level instead of within
its group. (Top-level Any Field is unaffected.)
2026-06-25 16:14:38 -04:00
Dan Stillman
0529c9a579 Match Any Field search condition to the All Fields & Tags quick search mode (#5962)
Add annotation text and comments to the conditions the Any Field
condition expands to, matching the fields covered by the All Fields &
Tags quick search mode, as the comment already says is intended. Key
detection and quoted-phrase splitting still differ, since those depend
on the quick search string parsing that Any Field doesn't do.
2026-06-25 16:14:38 -04:00
Dan Stillman
8b5a77a75e Support cross-level conditions and a result level in search (#5962)
Give a search a result level -- top-level item, attachment, note, or
annotation -- and map every condition to that level, so one search can
mix conditions that match at different levels of the item hierarchy
(e.g., a top-level item with a given author and a red annotation on one
of its PDFs). Each condition carries the level(s) it matches at: a match
is mapped up to an ancestor or down to a descendant, level-agnostic
conditions (tags) roll up to the result level, and fields that exist on
both items and attachments (title, url, accessDate) match natively at
either. The result level is stored as a `resultLevel` marker condition
alongside the join mode.

This also removes the temporary annotation-parent hacks, which the
general cross-level mapping replaces.
2026-06-25 16:14:38 -04:00
Dan Stillman
cdc70d1280 Compose grouped full-text conditions in SQL (#5962)
A fulltextContent condition was evaluated as a global post-filter keyed
on the search's top-level join mode -- correct for a top-level
condition, but not for one inside a group, which must combine with its
siblings under the group's own join mode. The new condition grouping UI
allows fulltextContent to be placed within groups, and we need to do so
to prefill the advanced-search pane from an "Everything" quicksearch.

Materialize a grouped fulltextContent into an itemID set and emit it as an
ordinary itemID IN/NOT IN predicate, so combineConditions composes it
under the group's join mode. Top-level fulltextContent keeps the existing
post-filter unchanged.

This also removes the quicksearch full-text post-filter special case. A
quick search puts its full-text in per-word "any" groups, so the
_hasQuicksearch flag was needed to make the post-filter union those
matches rather than intersect them under the top-level "all" join. Now
that the grouped full-text is composed in SQL it never reaches the
post-filter, so the flag is gone.
2026-06-25 16:14:38 -04:00
Dan Stillman
c01def2a85 Support nested condition groups in saved searches (#5962)
Replace the flat anySQL/quicksearch-block assembly in _buildQuery with a
recursive tree of AND/OR groups, built and reduced by a new
Zotero.Search.combineConditions() helper. groupStart/groupEnd markers
delimit nested groups and a joinMode marker sets each group's mode, so a
saved search can combine conditions with arbitrary nesting and per-group
join modes. The per-condition SQL generation is unchanged.

For example, a search built as

    search.addCondition('joinMode', 'all');
    search.addCondition('title', 'contains', 'foo');
    search.addCondition('groupStart', 'true', '');
    search.addCondition('joinMode', 'any');
    search.addCondition('tag', 'is', 'x');
    search.addCondition('tag', 'is', 'y');
    search.addCondition('groupEnd', 'true', '');

means "title contains 'foo' AND (tag is 'x' OR tag is 'y')". The 'true'
operator on the group markers is an unused placeholder -- they carry no
value, but a condition's operator can't be empty.

The quick search (matching multiple words) and the Any Field condition
previously had their own special handling in the query builder; they now
use the same grouping as everything else, so that code is gone. Behavior
for existing non-grouped searches is unchanged; new tests cover nested
groups and combineConditions() directly.
2026-06-25 16:14:38 -04:00
Abe Jellinek
00527332c4 Move Advanced Search and saved search editing to the main window (#5658)
---------

Co-authored-by: Dan Stillman <dstillman@zotero.org>
2026-06-12 15:21:59 -04:00
Bogdan Abaev
d2f1c56250 keep search conditionIDs in arithmetic sequence
When conditions are removed, shift conditionIDs so that
conditionIDs always go in increments of 1 (0, 1, 2 ...).
It prevents conditionIDs from conflicting with each other
when conditions are rearranged.

Fixes: #3434
2026-06-12 15:16:46 -04:00
abaevbog
22f62138a3
Fix quicksearch not finding annotations of standalone attachments (#5756)
Fix search not finding annotations of standalone attachments
when searching within a scope.

Fixes: #5751
2026-01-28 13:30:36 -05:00
Abe Jellinek
9509c62918 Fix search test descriptions, remove outdated tests
Some descriptions weren't changed to match test behavior changes in
48fd23c, and the doesn't-match tests are no longer relevant after those
changes.
2025-10-10 12:39:26 -04:00
Abe Jellinek
67d2e1cead fx140: Asyncify/ESMify tests
They seem to be succeeding when run individually, but failing when
run as a whole. Not sure why yet.
2025-07-30 22:30:53 -04:00
abaevbog
48fd23ccec
annotations showing in itemTree (#3416)
- Annotations are displayed in itemTree under their file attachments
  on a third level. The annotation spans the entire row.
- The title is constructed on the go. When possible, it
  includes annotation quote and comment as pseudo-columns
  of the row. The comment occupies about twice as much space as
  the quote. Otherwise, (if there is no quote) only the annotation
  comment
  is included as the "title" part of the row
- Non-CJK segments of the quote part of the annotation row are
  italicized. CJK segments are left as is. If CJK segments are present,
  there is more padding between quote and comment parts.
- Search matches the actual attachment instead of its parent file.
- Can create child notes from annotations of the same item or
  standalone notes from annotations across different items from
  the context menu or the header button.
- When an annotation (or multiple annotations) are selected, the
  annotationItemPane component is displayed where annotations are
  grouped by their top-level item. Annotations are displayed fully,
  without having their content cut off.
- Special treatment for annotations to always prompt
  to erase the item regardless of what collectionTree row
  is selected (e.g., if a collection is selected, we
  still want one to be able to delete the annotation).
  This only applies if all selected items are annotations.
  If multiple items are selected, some annotations and
  some not, do nothing. This is until the trash is
  ready. In the future, we may send annotations to trash
- strip all HTML tags from annotation for now, until the logic to
  properly render annotation markup is copied over from the reader
  (applies to both annotation-row component and the annotation
  item rendered in the itemTree)
- Added a generalized "expandToItem" function to itemTree to
  expand all ancestors of a given item, similar to "expandToCollection"
  from collectionTree
- add annotation conditions to advanced search
- show [Image not available] if no annotation file for ink or image
  annotations
- only keep annotation-specific context menu options when some
  annotations are selected in itemTree
- enable Quick Copy of annotations from itemTree via drag-drop,
  shortcut key, or Edit → Copy Annotation
- Minor refactoring of Zotero.Annotation.toJSON() to pull out async code
  that handles ink and image annotations, so that
  Zotero.Annotation.toJSONsync() for highlight, underline, and note
  annotations does not have to be awaited. Since ink and image
  annotation don't seem to work for drag-drop Quick Copy, they are just
  skipped for now.
2025-04-28 04:10:28 -04:00
Abe Jellinek
3c8d50dd47 Remove loadBrowserWindow() test support function (#5050)
A browser window was no longer actually needed for charset detection
(on macOS, at least, and hopefully elsewhere) because we switched to a
HiddenFrame-based hidden browser. Remaining uses now call
`loadZoteroWindow()`.
2025-02-26 03:00:04 -05:00
abaevbog
247826194a
fix advanced search anyField condition breakage (#4873)
Process "joinMode" before other conditions that may rely on it

Fixes: #4871

Co-authored-by: Dan Stillman <dstillman@zotero.org>
2024-11-27 23:43:09 -05:00
abaevbog
e3c18ee1c7
fix "anyField" conflicting with other conditions (#4845)
Fix the glitch where having anyField search condition
along with any other condition would return an empty
result set if joinMode="any".
This would happen due to a conflict between conditions
caused by wrapping "anyField" condition set into a quickSearch
block, which enforces "AND" operator between conditions.

Fixes: #4830
2024-11-15 02:07:11 -05:00
abaevbog
82d50676d3
fix quicksearch phrase search not respecting scope (#4608)
Fixes: #4381
2024-09-12 06:27:29 -04:00
abaevbog
2d3375e9f6
Revised logic of trashed collections in item cache (#4358)
- revert change from 2401a34031
that only loads un-trashed collections in _loadCollections.
If an item only belongs to deleted collections, item._loaded.collections = true
from _loadCollections will never run, so an exception
will be thrown in item.toJSON() when syncing happens.
Instead, to address the problem of item.getCollections()
having stale data #4307, add 'includeTrashed' parameter to
item.getCollections() based on which item._collections
will be filtered. Fixes: #4346
- revert earlier, no more necessary, changes from a532cfb475
to not alter item._collections cache when collections are being trashed or restored.
Collection is removed from item._collections only when it is permanently
erased.
- removed unnecessary test checking for consistent item._collections
value before and after reload, since item._collections is no longer
modified
- fix encountered bug where a trashed child collection is not
unloaded if a parent collection is erased without being trashed first.
- tweaked Zotero.Search sql construction to count items
that only belong to trashed collections into 'unfiled'. Fixes: #4347

---------

Co-authored-by: Dan Stillman
2024-07-09 03:34:12 -04:00
Bogdan Abaev
a532cfb475 trash functionality for collections and searches (#3307)
When a collection or a saved search is deleted, it appears in
trash among other trashed items. From there, it can be restored
or permanently deleted.

Items of trashed collections are not affected my the trashing/permanent
deletion of a collection and need to be deleted separately like before.

Subcollections of a trashed collection do not appear in the trash and
are restored or permanently deleted with the top-most trashed parent.
2024-06-17 23:14:21 -04:00
Dan Stillman
33ef7b1641 Prevent setting search .name to empty value
Prevents bug in zotero-citation plugin (at least on macOS) from creating
a search that breaks syncing

We were already checking for a missing name in `saveTx()`, but the
plugin is saving the same search twice in rapid succession, the second
time without a name, and the second attempt clears the search object's
name value after the first save's `_initSave()` check and before its SQL
write. The second save fails, but the first save goes through without a
name, resulting in a sync error.

https://forums.zotero.org/discussion/104274/id-1702002152-cannot-sync
https://github.com/MuiseDestiny/zotero-citation/issues/31
2023-04-12 22:23:13 -04:00
Dan Stillman
76f2f0c783 Don't show items with annotated attachments after moving to trash
https://forums.zotero.org/discussion/100775/deleted-items-keep-reappearing-in-my-library

Regression from c3ee588bf
2022-11-28 04:34:49 -05:00
Dan Stillman
3a77eb85ed Don't match all attachments with annotations for "not" search conditions
Fixes #2867
2022-10-27 03:46:18 -04:00
Dan Stillman
5103d904f3 Include proper test for b373291c02 for #2771 2022-08-19 12:05:30 -04:00
Dan Stillman
e9e1add9b8 Fixed filed items with annotations appearing in Unfiled Items
Fixes #2771

Regression from 20c6fe67
2022-08-19 12:05:30 -04:00
Dan Stillman
c3ee588bfe Match parent attachments for annotation tags
Expose annotation tags in tag selector and match parent attachments when
filtering/searching

This also fixes searching for annotation text or comments when using
Everything quick search.

This is temporary until we display annotations in the items list
directly.
2022-08-17 03:35:28 -04:00
Dan Stillman
ead8c6bb45 Fix Everything search after annotations
And replace ancient 'annotation' search condition with
'annotationText'/'annotationComment'
2021-03-02 17:39:39 -05:00
Dan Stillman
e45ca4edad Support deleted property for collections and searches
This lays the groundwork for moving collections and searches to the
trash instead of deleting them outright. We're not doing that yet, so
the `deleted` property will never be set (except for items), but this
will allow clients from this point forward to sync collections and
searches with that property for when it's used in the future. For now,
such objects will just be hidden from the collections pane as if they
had been deleted.
2021-01-13 00:49:12 -05:00
Dan Stillman
a53f363b8d Additional fix for search crash with includeParentsAndChildren
Follow-up to 76081ab05
2020-02-11 13:09:54 -05:00
Dan Stillman
76081ab05f Fix crash when search uses no-op condition and includeParentsAndChildren
E.g., a nonexistent saved search
2020-02-11 00:23:45 -05:00
Dan Stillman
bbbd02444b Restore 'yesterday'/'today'/'tomorrow' parsing for dates in searches
Follow-up to a549a64de9, which removed it from strToDate()
2019-12-06 03:12:48 -07:00
Dan Stillman
4b60c6ca27 Type/field handling overhaul
This changes the way item types, item fields, creator types, and CSL
mappings are defined and handled, in preparation for updated types and
fields.

Instead of being predefined in SQL files or code, type/field info is
read from a bundled JSON file shared with other parts of the Zotero
ecosystem [1], referred to as the "global schema". Updates to the
bundled schema file are automatically applied to the database at first
run, allowing changes to be made consistently across apps.

When syncing, invalid JSON properties are now rejected instead of being
ignored and processed later, which will allow for schema changes to be
made without causing problems in existing clients. We considered many
alternative approaches, but this approach is by far the simplest,
safest, and most transparent to the user.

For now, there are no actual changes to types and fields, since we'll
first need to do a sync cut-off for earlier versions that don't reject
invalid properties.

For third-party code, the main change is that type and field IDs should
no longer be hard-coded, since they may not be consistent in new
installs. For example, code should use `Zotero.ItemTypes.getID('note')`
instead of hard-coding `1`.

[1] https://github.com/zotero/zotero-schema
2019-09-16 02:27:22 -04:00
Dan Stillman
1061893998 "Attachment Content" search improvements
- Fix incorrect results for ANY search with multiple "Attachment
  Content" conditions and no other conditions
- Dramatically speed up single-word searches by avoiding unnecessary
  text scans (which probably addresses #1595)
- Clean up code
2019-02-19 04:10:25 -05:00
Dan Stillman
223f582aa7 Fix search error on nonexistent collection in recursive mode
And don't return results for a nonexistent parent search
2018-11-28 15:31:57 -07:00