ReMe/reme4/utils/wikilink_handler.py
Sen Huang bb354cc580
refactor(steps): reorganize step modules and remove demo steps (#255)
* feat(config): add comprehensive job definitions for vault operations

- Add utility jobs like version, search, traverse, list, read, stat
- Include file operations like move, delete, upload, download
- Add daily workspace management jobs: daily_list, daily_resolve, daily_reindex
- Update descriptions to reflect vault-based operations instead of working_dir
- Add proper section headers and documentation for each job category

refactor(steps): reorganize step modules and remove demo steps

- Move steps into categorized packages: common, crud, frontmatter, daily, jobs
- Remove demo steps (DemoEchoStep1, DemoEchoStep2, StreamDemoStep1, StreamDemoStep2)
- Add new steps: InitStep for vault initialization, TraverseStep for graph traversal
- Update __init__.py to auto-import all step modules
- Organize imports by functionality (common, CRUD operations, frontmatter, daily)

feat(vault): implement vault-centric file operations and configuration

- Change default config to use vault_dir instead of working_dir
- Add environment variable support for embedding configuration
- Implement file watcher with lite backend for daily/digest directories
- Update search step to use 'name' instead of 'title' from frontmatter
- Create ResourceEntry schema for tracking uploaded assets

docs(steps): add comprehensive documentation for all step categories

- Document file-I/O split by blast radius (crud vs frontmatter packages)
- Add detailed descriptions for each step category and functionality
- Explain the purpose and usage patterns for different types of file operations
- Provide clear parameter documentation for all new job configurations

* fix(config): correct vault directory path and remove unused job configurations

- Fix vault_dir from 'vaultd' to 'vault' in default configuration
- Remove deprecated traverse and list job configurations
- Remove unused tag tooling configurations
- Remove background watch_file job configuration

refactor(steps): remove unused jobs module import

- Comment out jobs module import in steps/__init__.py
- This removes unused synchronizer and digester step registrations

refactor(tests): update import path and add pylint directive

- Update ResourceEntry import from reme4.schema to reme4.schema.resource_meta
- Add pylint disable directive for unused argument in test datetime mocks

* efactor(steps): remove unused modules from __all__

- Remove "background" module from __all__ list
- Remove "jobs" module from __all__ list
- These modules were no longer being used in the steps package

* feat(config): update vault directory structure and remove file watcher

- Change vault_dir reference from ./vault to ./vault in CLI example
- Add daily_dir, digest_dir, and resource_dir configuration options
- Remove file_watcher component configuration as it's no longer needed
- Update comment to reflect correct module name (reme4vault)

refactor(steps): add background step and remove deprecated init step

- Import and register background step module
- Remove deprecated InitStep from common steps
- Update __all__ export list to include background step

refactor(reindex): improve reindex step to scan vault directly

- Update docstring to reflect vault scanning instead of watcher sync
- Replace file watcher stop/start logic with direct vault path walking
- Add support for suffix filtering during reindex operation
- Use index_changes job to process found files

refactor(wikilink_utils): enhance inbound source lookup with link scope

- Import LinkScopeEnum for proper type handling
- Update get_inlinks call to use ALL scope for virtual targets
- Improve documentation for reverse-index lookup behavior

test(refactor): clean up test suite removing deprecated functionality

- Remove test_init_job and test_demo_job unit tests
- Update help job assertion to check for literal command format
- Change test directory from .reme to vault in CRUD tests
- Remove init and demo job calls from integration test

BREAKING CHANGE: Removes file_watcher component and init step

* style(steps): fix import formatting in __init__.py

Add proper spacing in the background module import statement
to maintain consistent code style and readability.

* refactor(config): change default vault directory from vault to .reme

Default dev config now points vault_dir at ./.reme so `python -m
reme4 start` can be run from the repo root and exercise the full
atomic-tool surface against the seeded test data.

BREAKING CHANGE: The default vault directory has been changed from
'vault' to '.reme' in the configuration.

* docs(reme4_report): fix markdown formatting and remove extra content

* refactor(file_parser): delegate wikilink extraction to WikilinkHandler

* fix(search): handle empty query case gracefully

- Replace assertion with conditional check for empty query
- Set response success to false when query is empty
- Return error message instead of throwing assertion error
- Maintain existing validation for other parameters
2026-05-25 17:52:51 +08:00

387 lines
14 KiB
Python

"""Wikilink handler — single source of truth for ``[[...]]`` syntax.
One class, :class:`WikilinkHandler`, owning every wikilink concern:
* **Pure text** — regex, Dataview predicate inference, validation:
:meth:`~WikilinkHandler.extract_links` (used by
:mod:`reme4.components.file_parser.linked_file_parser`),
:meth:`~WikilinkHandler.scan_and_rewrite`,
:meth:`~WikilinkHandler.validate_src_dst` /
:meth:`~WikilinkHandler.validate_scope` /
:meth:`~WikilinkHandler.within_scope`.
* **Async, file_graph-aware** —
:meth:`~WikilinkHandler.find_inbound` (called by ``file_delete``
to surface references the caller might want to clean up) and
:meth:`~WikilinkHandler.retarget_links` (called by ``file_move``
post-rename to point inbound ``[[src]]`` at the new path). Source
candidates come from the file_graph's reverse index — no fs scan.
Wikilink convention. Targets are taken **literally** — ``[[X]]`` →
``target="X"``, no implicit ``.md``, no short-form basename search,
no folder-note expansion. Anchor and alias survive a rewrite
verbatim. Image marker (``!``) and Dataview predicate (``pred::``
outside the brackets) sit outside ``[[...]]`` and are not touched by
a rewrite. Recommended form: full path relative to the vault with
extension (``[[topics/x.md]]``).
Stale graph entries are harmless (``scan_and_rewrite`` returns
count=0 and the file is skipped), but a graph missing recent writes
will miss those sources — keep the watcher in sync.
"""
import re
from dataclasses import dataclass
from pathlib import Path
from ..enumeration import LinkScopeEnum
from ..schema import FileLink
@dataclass(frozen=True)
class WikilinkMatch:
"""One ``[[...]]`` occurrence with parts surfaced.
``anchor`` / ``alias`` are stored **without** the leading ``#`` /
``|`` so they map cleanly to :class:`FileLink.target_anchor`; the
rewrite path reads the raw regex groups (with delimiters) directly
and doesn't go through this dataclass.
"""
target: str
anchor: str | None
alias: str | None
bang: bool
start: int
end: int
class WikilinkHandler:
"""Pure-text wikilink operations: parse, extract, rewrite, validate."""
# Captures: optional image marker (``!``), the bare target, an
# optional ``#anchor`` slice (with ``#``), and an optional ``|alias``
# slice (with ``|``). The anchor / alias inner classes exclude ``[``
# defensively so a runaway match on malformed input can't swallow
# following links.
WIKILINK_RE = re.compile(
r"""
(?P<bang>!?)
\[\[
(?P<target>[^\[\]\|\#\n]+?)
(?P<anchor>\#[^\[\]\|\n]+)?
(?P<alias>\|[^\[\]\n]+)?
\]\]
""",
re.VERBOSE,
)
FORBIDDEN_IN_NEW = ("[", "]", "#", "|", "\n", "\r")
_DATAVIEW_LINE_RE = re.compile(
r"^[ \t]*(?:[-*+][ \t]+)?(?P<predicate>[A-Za-z][A-Za-z0-9_]*)\s*::\s*(?P<value>.+?)\s*$",
re.MULTILINE,
)
_INLINE_FIELD_OPEN_RE = re.compile(r"\[(?P<predicate>[A-Za-z][A-Za-z0-9_]*)\s*::\s*")
# -- Low-level scan ------------------------------------------------
@classmethod
def iter_matches(cls, text: str):
"""Yield :class:`WikilinkMatch` for every ``[[...]]`` in ``text``.
Skips matches whose target is empty after strip (defensive).
"""
for m in cls.WIKILINK_RE.finditer(text):
target = m.group("target").strip()
if not target:
continue
anchor_raw = m.group("anchor")
alias_raw = m.group("alias")
yield WikilinkMatch(
target=target,
anchor=anchor_raw[1:].strip() if anchor_raw else None,
alias=alias_raw[1:].strip() if alias_raw else None,
bang=bool(m.group("bang")),
start=m.start(),
end=m.end(),
)
# -- FileLink extraction (with predicate inference) ---------------
@classmethod
def extract_links(cls, text: str, source_path: str) -> list[FileLink]:
"""Emit :class:`FileLink` edges for every wikilink in ``text``.
No resolution: ``target_path`` is the bracket contents verbatim.
Results are deduped by ``(target_path, predicate, target_anchor)``
preserving order.
"""
if not text:
return []
inline_spans = cls._iter_inline_fields(text)
out: list[FileLink] = []
seen: set[tuple] = set()
for wm in cls.iter_matches(text):
predicate = cls._predicate_for(text, wm.start, inline_spans)
key = (wm.target, predicate, wm.anchor)
if key in seen:
continue
seen.add(key)
out.append(
FileLink(
source_path=source_path,
target_path=wm.target,
target_anchor=wm.anchor,
predicate=predicate,
),
)
return out
# -- Find / rewrite by literal target match ------------------------
@classmethod
def scan_and_rewrite(
cls,
text: str,
old: str,
new: str | None,
) -> tuple[str, int]:
"""Find (and optionally rewrite) wikilinks whose target equals ``old``.
Returns ``(new_text, count)``. When ``new`` is ``None`` no rewrite
happens (the original text is returned), but the count is still
populated — used by ``find_inbound``. Matching is literal:
``target == old``. No short-link, no implicit ``.md``, no
folder-note expansion.
"""
count = 0
def sub(match: re.Match) -> str:
nonlocal count
target = match.group("target").strip()
if target != old:
return match.group(0)
count += 1
if new is None:
return match.group(0)
anchor = match.group("anchor") or ""
alias = match.group("alias") or ""
bang = match.group("bang") or ""
return f"{bang}[[{new}{anchor}{alias}]]"
new_text = cls.WIKILINK_RE.sub(sub, text)
return new_text, count
# -- Validation ----------------------------------------------------
@classmethod
def validate_src_dst(cls, src: str, dst: str) -> str | None:
"""Return an error message for bad rewrite inputs, or None when OK."""
if not src or not dst:
return "src and dst are required"
if any(ch in dst for ch in cls.FORBIDDEN_IN_NEW):
return "dst must not contain [ ] # | newline"
if Path(src).is_absolute() or Path(dst).is_absolute():
return "src and dst must be relative to the vault"
return None
@staticmethod
def validate_scope(scope: str) -> str | None:
"""Return an error message for a bad scope, or None when OK."""
if scope and Path(scope).is_absolute():
return "scope must be relative to the vault"
return None
@staticmethod
def within_scope(rel: str, scope: str) -> bool:
"""``rel`` (relative to the vault) is inside ``scope`` (empty = anywhere)."""
if not scope:
return True
prefix = scope.rstrip("/") + "/"
return rel == scope or rel.startswith(prefix)
# -- Predicate helpers (internal) ---------------------------------
@classmethod
def _iter_inline_fields(cls, text: str) -> list[tuple[int, int, str]]:
"""Find inline-bracketed ``[predicate:: …]`` field spans by depth scan."""
out: list[tuple[int, int, str]] = []
for m in cls._INLINE_FIELD_OPEN_RE.finditer(text):
depth = 1
i = m.end()
n = len(text)
while i < n:
c = text[i]
if c == "\n":
break
if c == "[":
depth += 1
elif c == "]":
depth -= 1
if depth == 0:
out.append((m.start(), i + 1, m.group("predicate")))
break
i += 1
return out
@classmethod
def _predicate_for(
cls,
text: str,
pos: int,
inline_spans: list[tuple[int, int, str]],
) -> str | None:
"""Resolve the predicate governing a wikilink at offset ``pos``.
Precedence: inline-bracketed > line-level Dataview > none.
"""
for field_start, field_end, predicate in inline_spans:
if field_start <= pos < field_end:
return predicate
line_start = text.rfind("\n", 0, pos) + 1
line_end = text.find("\n", pos)
if line_end == -1:
line_end = len(text)
m = cls._DATAVIEW_LINE_RE.match(text[line_start:line_end])
if m and line_start + m.start("value") <= pos:
return m.group("predicate")
return None
# -- Async file_graph-aware operations -----------------------------
@classmethod
async def _inbound_sources(cls, file_store, target: str) -> list[str]:
"""Source paths the file_graph reports as referencing ``target``.
Reverse-index lookup via ``file_graph.get_inlinks(target, scope=ALL)`` —
``target`` is typically virtual here (the move/delete callers query for
references to a path that has just been removed), so ``scope=ALL`` is
required to surface sources whose edges sit in the pending bucket.
Each returned ``FileLink`` carries the linking node's ``source_path``;
we dedupe to a sorted list since one source can host multiple edges
(different anchor/predicate) to the same target. Returns ``[]`` when
there is no file_graph attached or no source references the target.
"""
if not file_store.file_graph:
return []
inlinks = await file_store.file_graph.get_inlinks(target, scope=LinkScopeEnum.ALL)
return sorted({link.source_path for link in inlinks if link.source_path})
@classmethod
async def find_inbound(cls, file_store, target: str, scope: str = "") -> dict:
"""Count wikilinks across the vault that point at ``target``.
Literal matching: ``[[target]]`` only. The target file itself is
excluded — self-references don't survive a delete and aren't
actionable for the caller. Sources come from the file_graph's
reverse index; per-file counts come from reading each candidate
source (the graph dedupes by ``(target, predicate, anchor)`` so
it can't count repeated bare-wikilink occurrences directly).
Result shape::
{
"target": str,
"scope": str | None,
"files_touched": int, # number of OTHER files containing >=1 ref
"links_total": int, # total ref count across those files
"by_file": [{"path": str, "count": int}, ...],
}
On bad inputs returns ``{"target": ..., "error": str}``.
"""
if not target:
return {"target": target, "error": "target is required"}
if Path(target).is_absolute():
return {"target": target, "error": "target must be relative to the vault"}
err = cls.validate_scope(scope)
if err is not None:
return {"target": target, "error": err}
vault_dir = Path(file_store.vault_path or ".").resolve()
by_file: list[dict] = []
total = 0
for rel in await cls._inbound_sources(file_store, target):
if rel == target:
continue # self-references not actionable for delete cleanup
if not cls.within_scope(rel, scope):
continue
try:
text = (vault_dir / rel).read_text(encoding="utf-8")
except Exception:
continue
_, count = cls.scan_and_rewrite(text, old=target, new=None)
if count > 0:
by_file.append({"path": rel, "count": count})
total += count
return {
"target": target,
"scope": scope or None,
"files_touched": len(by_file),
"links_total": total,
"by_file": by_file,
}
@classmethod
async def retarget_links(
cls,
file_store,
src: str,
dst: str,
scope: str = "",
dry_run: bool = False,
) -> dict:
"""Rewrite every wikilink pointing at ``src`` to point at ``dst``.
Pure helper — called directly by ``file_move`` post-rename. Literal
matching only; candidate sources come from the file_graph's reverse
index.
"""
err = cls.validate_src_dst(src, dst)
if err is not None:
return {"src": src, "dst": dst, "error": err}
if src == dst:
return {
"src": src,
"dst": dst,
"scope": scope or None,
"dry_run": dry_run,
"files_touched": 0,
"links_changed": 0,
"by_file": [],
}
err = cls.validate_scope(scope)
if err is not None:
return {"src": src, "dst": dst, "error": err}
vault_dir = Path(file_store.vault_path or ".").resolve()
by_file: list[dict] = []
total_changes = 0
for rel in await cls._inbound_sources(file_store, src):
if not cls.within_scope(rel, scope):
continue
abs_path = vault_dir / rel
try:
text = abs_path.read_text(encoding="utf-8")
except Exception:
continue
new_text, count = cls.scan_and_rewrite(text, old=src, new=dst)
if count > 0:
by_file.append({"path": rel, "count": count})
total_changes += count
if not dry_run:
abs_path.write_text(new_text, encoding="utf-8")
return {
"src": src,
"dst": dst,
"scope": scope or None,
"dry_run": dry_run,
"files_touched": len(by_file),
"links_changed": total_changes,
"by_file": by_file,
}