mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-09-29 01:41:38 +00:00
* feat(config): add comprehensive job definitions for vault operations - Add utility jobs like version, search, traverse, list, read, stat - Include file operations like move, delete, upload, download - Add daily workspace management jobs: daily_list, daily_resolve, daily_reindex - Update descriptions to reflect vault-based operations instead of working_dir - Add proper section headers and documentation for each job category refactor(steps): reorganize step modules and remove demo steps - Move steps into categorized packages: common, crud, frontmatter, daily, jobs - Remove demo steps (DemoEchoStep1, DemoEchoStep2, StreamDemoStep1, StreamDemoStep2) - Add new steps: InitStep for vault initialization, TraverseStep for graph traversal - Update __init__.py to auto-import all step modules - Organize imports by functionality (common, CRUD operations, frontmatter, daily) feat(vault): implement vault-centric file operations and configuration - Change default config to use vault_dir instead of working_dir - Add environment variable support for embedding configuration - Implement file watcher with lite backend for daily/digest directories - Update search step to use 'name' instead of 'title' from frontmatter - Create ResourceEntry schema for tracking uploaded assets docs(steps): add comprehensive documentation for all step categories - Document file-I/O split by blast radius (crud vs frontmatter packages) - Add detailed descriptions for each step category and functionality - Explain the purpose and usage patterns for different types of file operations - Provide clear parameter documentation for all new job configurations * fix(config): correct vault directory path and remove unused job configurations - Fix vault_dir from 'vaultd' to 'vault' in default configuration - Remove deprecated traverse and list job configurations - Remove unused tag tooling configurations - Remove background watch_file job configuration refactor(steps): remove unused jobs module import - Comment out jobs module import in steps/__init__.py - This removes unused synchronizer and digester step registrations refactor(tests): update import path and add pylint directive - Update ResourceEntry import from reme4.schema to reme4.schema.resource_meta - Add pylint disable directive for unused argument in test datetime mocks * efactor(steps): remove unused modules from __all__ - Remove "background" module from __all__ list - Remove "jobs" module from __all__ list - These modules were no longer being used in the steps package * feat(config): update vault directory structure and remove file watcher - Change vault_dir reference from ./vault to ./vault in CLI example - Add daily_dir, digest_dir, and resource_dir configuration options - Remove file_watcher component configuration as it's no longer needed - Update comment to reflect correct module name (reme4vault) refactor(steps): add background step and remove deprecated init step - Import and register background step module - Remove deprecated InitStep from common steps - Update __all__ export list to include background step refactor(reindex): improve reindex step to scan vault directly - Update docstring to reflect vault scanning instead of watcher sync - Replace file watcher stop/start logic with direct vault path walking - Add support for suffix filtering during reindex operation - Use index_changes job to process found files refactor(wikilink_utils): enhance inbound source lookup with link scope - Import LinkScopeEnum for proper type handling - Update get_inlinks call to use ALL scope for virtual targets - Improve documentation for reverse-index lookup behavior test(refactor): clean up test suite removing deprecated functionality - Remove test_init_job and test_demo_job unit tests - Update help job assertion to check for literal command format - Change test directory from .reme to vault in CRUD tests - Remove init and demo job calls from integration test BREAKING CHANGE: Removes file_watcher component and init step * style(steps): fix import formatting in __init__.py Add proper spacing in the background module import statement to maintain consistent code style and readability. * refactor(config): change default vault directory from vault to .reme Default dev config now points vault_dir at ./.reme so `python -m reme4 start` can be run from the repo root and exercise the full atomic-tool surface against the seeded test data. BREAKING CHANGE: The default vault directory has been changed from 'vault' to '.reme' in the configuration. * docs(reme4_report): fix markdown formatting and remove extra content * refactor(file_parser): delegate wikilink extraction to WikilinkHandler * fix(search): handle empty query case gracefully - Replace assertion with conditional check for empty query - Set response success to false when query is empty - Return error message instead of throwing assertion error - Maintain existing validation for other parameters
119 lines
4.9 KiB
Python
119 lines
4.9 KiB
Python
"""``graph_traverse_step`` — BFS over wikilink edges from a seed file.
|
|
|
|
Single tool for relationship browsing. ``depth=1`` covers the trivial
|
|
"what does this link to / what links here" lookups (set ``direction``
|
|
accordingly); higher depth opens up multi-hop exploration.
|
|
|
|
Output is one record per edge traversed (not per node), so the same
|
|
target can appear multiple times if reached via different predicates
|
|
or paths — agents dedupe at the call site if they want a flat node
|
|
set. Each record carries ``via`` (the predecessor) and the link's
|
|
``predicate`` / ``anchor`` so the agent can reconstruct the path.
|
|
|
|
Adjacency is loaded once via ``file_graph.get_nodes(None)`` — every
|
|
real node arrives with its full ``links`` payload, and we build both
|
|
the outbound and the inbound index in a single pass. The BFS then
|
|
runs purely in memory: no per-frontier-node graph round-trips, no
|
|
filesystem walk. The ``get_inlinks`` / ``get_outlinks`` contract
|
|
methods stay unused here because they'd add network round-trips for
|
|
data we already have.
|
|
|
|
Direction vocabulary accepts both the standard convention
|
|
(``forward`` / ``backward`` / ``both``) and the engine convention
|
|
(``out`` / ``in`` / ``both``).
|
|
|
|
The seed ``path`` is taken as-is (vault-relative). A seed that doesn't
|
|
match any graph node yields an empty result (no error).
|
|
"""
|
|
|
|
from collections import deque
|
|
|
|
from ..base_step import BaseStep
|
|
from ...components import R
|
|
from ...schema import FileLink
|
|
|
|
|
|
_FORWARD = {"out", "forward"}
|
|
_BACKWARD = {"in", "backward"}
|
|
_BOTH = {"both"}
|
|
_VALID_DIRECTIONS = _FORWARD | _BACKWARD | _BOTH
|
|
|
|
|
|
@R.register("graph_traverse_step")
|
|
class GraphTraverseStep(BaseStep):
|
|
"""BFS from a seed file to explore wikilink relationships.
|
|
|
|
Parameters:
|
|
path — seed path (vault-relative).
|
|
direction — ``forward`` / ``backward`` / ``both`` (or ``out`` / ``in`` / ``both``).
|
|
depth — hop limit (default 1 = immediate neighbors).
|
|
predicate — optional edge-type filter; ``None`` = no filter.
|
|
"""
|
|
|
|
async def execute(self):
|
|
"""BFS from ``path`` and emit one record per traversed edge."""
|
|
assert self.context is not None
|
|
seed = str(self.context.get("path") or "").strip()
|
|
assert seed, "path is required"
|
|
max_depth = int(self.context.get("depth") or 1)
|
|
direction = (self.context.get("direction") or "both").lower()
|
|
predicate = self.context.get("predicate")
|
|
assert (
|
|
direction in _VALID_DIRECTIONS
|
|
), f"direction must be one of {sorted(_VALID_DIRECTIONS)}, got {direction!r}"
|
|
|
|
# Build outbound / inbound adjacency in one pass over all nodes.
|
|
outbound: dict[str, list[tuple[str, FileLink]]] = {}
|
|
inbound: dict[str, list[tuple[str, FileLink]]] = {}
|
|
if self.file_store.file_graph:
|
|
for node in await self.file_store.file_graph.get_nodes():
|
|
for link in node.links:
|
|
if not link.target_path:
|
|
continue
|
|
outbound.setdefault(node.path, []).append((link.target_path, link))
|
|
inbound.setdefault(link.target_path, []).append((node.path, link))
|
|
|
|
walk_out = direction in _FORWARD or direction in _BOTH
|
|
walk_in = direction in _BACKWARD or direction in _BOTH
|
|
|
|
visited_edges: set[tuple[str, str, str | None]] = set()
|
|
results: list[dict] = []
|
|
queue: deque[tuple[str, int]] = deque([(seed, 0)])
|
|
|
|
while queue:
|
|
current, depth = queue.popleft()
|
|
if depth >= max_depth:
|
|
continue
|
|
|
|
edges: list[tuple[str, str | None, str | None]] = []
|
|
if walk_out:
|
|
for tgt, link in outbound.get(current, ()):
|
|
if predicate is not None and link.predicate != predicate:
|
|
continue
|
|
edges.append((tgt, link.predicate, link.target_anchor))
|
|
if walk_in:
|
|
for src, link in inbound.get(current, ()):
|
|
if predicate is not None and link.predicate != predicate:
|
|
continue
|
|
edges.append((src, link.predicate, link.target_anchor))
|
|
|
|
for next_path, pred, anchor in edges:
|
|
edge_key = (current, next_path, pred)
|
|
if edge_key in visited_edges:
|
|
continue
|
|
visited_edges.add(edge_key)
|
|
results.append(
|
|
{
|
|
"path": next_path,
|
|
"depth": depth + 1,
|
|
"via": current,
|
|
"predicate": pred,
|
|
"anchor": anchor,
|
|
},
|
|
)
|
|
if depth + 1 < max_depth:
|
|
queue.append((next_path, depth + 1))
|
|
|
|
self.context.response.success = True
|
|
self.context.response.answer = f"Traversed {len(results)} edge(s) from {seed}"
|
|
self.context.response.metadata.update({"edges": results, "count": len(results)})
|