mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-11 03:40:03 +00:00
* feat(search): add tool_context-scoped chunk dedup with TTL
Introduce _ToolContextDedupMixin shared by search/vector_search/bm25_search
to skip already-seen chunks within one agent tool_context. Per-context state
lives in app_context.metadata with configurable TTL (default 24h).
* feat(search): unify chunk answer rendering with merge and explicit empty messages
- Refactor SearchStep/VectorSearchStep/Bm25SearchStep to share format_chunks_answer for consistent source rendering and adjacent session-chunk merging.
- Distinguish empty results: ALL_RETURNED_MESSAGE when dedup removes everything vs NO_RESULTS_MESSAGE when nothing matched.
- Bump JsonlFileChunker default max_chars to 4000.
- Add unit tests for source-format merge and empty-result messages.
* refactor(config): reorganize file_chunker components and move jsonl max_chars into config
- Register explicit markdown/json/jsonl chunkers in beam.yaml and lme.yaml with markdown options (embed_toc, max_ast_sections, frontmatter handling) and jsonl max_chars=4000.
- Restrict default chunker to txt/log extensions.
- Revert JsonlFileChunker code default max_chars back to 2000; the 4000 value now lives in config.
* chore(benchmark): increase longmemeval num_items from 64 to 500
* refactor(search): split SearchStep into simplified and v2 variants, extract counter utility
- Extract global_counter_next from ApplicationContext into reme/utils/counter.py
as a standalone function operating on metadata dict with lazy initialization.
- Split SearchStep into two variants:
- SearchStep (simplified): inline chunk.id dedup, single-branch vector/keyword
optimization based on vector_weight, inline answer formatting.
- SearchV2Step (full): preserves _ToolContextDedupMixin with interval-subset-aware
dedup and format_chunks_answer with session-aware chunk merging.
- Update beam.yaml and lme.yaml to use search_v2_step for benchmark jobs.
- Rename existing search tests to test_search_v2_step_* and add new
test_search_step_* tests covering the simplified variant.
* fix: normalise missing trailing newline in _build_union_chunk to prevent line collision
* refactor: lazy-init counter tree in ApplicationContext metadata
- Remove hardcoded _counter_tree and _counter_tree_lock initialization
from ApplicationContext.metadata; rely on lazy initialization in
reme.utils.counter.global_counter_next on first call
- Set longmemeval num_items back to 500
- Remove obsolete trailing-newline collision tests
---------
Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
32 lines
1.5 KiB
YAML
32 lines
1.5 KiB
YAML
# LongMemEval evaluation configuration
|
|
# This file controls what/how to evaluate.
|
|
|
|
dataset:
|
|
path: "benchmark/datasets/longmemeval/longmemeval_s_reme_cleaned.json"
|
|
start_index: 0 # first item index
|
|
num_items: 500 # how many items to evaluate (starting from start_index)
|
|
max_sessions: 0 # 0 = all sessions; >0 = limit sessions per item for testing
|
|
question_types: [] # filter by question_type; empty list = no filtering (all types)
|
|
workspace_root: "benchmark/memory_workspaces/longmemeval-s" # workspace root for item workspaces
|
|
|
|
evaluation:
|
|
# LLM-as-judge uses the 'judge' as_llm component defined in lme.yaml
|
|
# Model and credentials are configured there (reading from .env)
|
|
# Judgment is always binary (yes/no) — defined in lme/llm_judge.yaml
|
|
num_workers: 32 # 0 = auto (cpu_count - 2, min 1); 1 = sequential; >1 = parallel
|
|
filter_future_sessions: true # true = only ingest sessions with timestamp <= question_date
|
|
|
|
reme:
|
|
config: "lme.yaml" # reme config to use (in reme/config/)
|
|
# Dream trigger: when gap between consecutive sessions crosses this hour (23:00)
|
|
dream_trigger_hour: 23
|
|
# Dream scan_days for each trigger
|
|
dream_scan_days: 2
|
|
dream_max_units: 5
|
|
|
|
output:
|
|
dir: "benchmark/results/longmemeval"
|
|
log_dir: "logs" # log directory (relative to project root)
|
|
log_prefix: "longmemeval" # benchmark name used in log filenames
|
|
log_to_console: true
|
|
log_to_file: true
|