* feat(search): add tool_context-scoped chunk dedup with TTL
Introduce _ToolContextDedupMixin shared by search/vector_search/bm25_search
to skip already-seen chunks within one agent tool_context. Per-context state
lives in app_context.metadata with configurable TTL (default 24h).
* feat(search): unify chunk answer rendering with merge and explicit empty messages
- Refactor SearchStep/VectorSearchStep/Bm25SearchStep to share format_chunks_answer for consistent source rendering and adjacent session-chunk merging.
- Distinguish empty results: ALL_RETURNED_MESSAGE when dedup removes everything vs NO_RESULTS_MESSAGE when nothing matched.
- Bump JsonlFileChunker default max_chars to 4000.
- Add unit tests for source-format merge and empty-result messages.
* refactor(config): reorganize file_chunker components and move jsonl max_chars into config
- Register explicit markdown/json/jsonl chunkers in beam.yaml and lme.yaml with markdown options (embed_toc, max_ast_sections, frontmatter handling) and jsonl max_chars=4000.
- Restrict default chunker to txt/log extensions.
- Revert JsonlFileChunker code default max_chars back to 2000; the 4000 value now lives in config.
* chore(benchmark): increase longmemeval num_items from 64 to 500
* refactor(search): split SearchStep into simplified and v2 variants, extract counter utility
- Extract global_counter_next from ApplicationContext into reme/utils/counter.py
as a standalone function operating on metadata dict with lazy initialization.
- Split SearchStep into two variants:
- SearchStep (simplified): inline chunk.id dedup, single-branch vector/keyword
optimization based on vector_weight, inline answer formatting.
- SearchV2Step (full): preserves _ToolContextDedupMixin with interval-subset-aware
dedup and format_chunks_answer with session-aware chunk merging.
- Update beam.yaml and lme.yaml to use search_v2_step for benchmark jobs.
- Rename existing search tests to test_search_v2_step_* and add new
test_search_step_* tests covering the simplified variant.
* fix: normalise missing trailing newline in _build_union_chunk to prevent line collision
* refactor: lazy-init counter tree in ApplicationContext metadata
- Remove hardcoded _counter_tree and _counter_tree_lock initialization
from ApplicationContext.metadata; rely on lazy initialization in
reme.utils.counter.global_counter_next on first call
- Set longmemeval num_items back to 500
- Remove obsolete trailing-newline collision tests
---------
Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
|
||
|---|---|---|
| .. | ||
| beam | ||
| datasets | ||
| longmemeval | ||
| kill.sh | ||
| README.md | ||
| README_ZH.md | ||
| result-beam.md | ||
| result-longmemeval.md | ||
ReMe Benchmarks
Reproduction guide for the two memory benchmarks shipped with ReMe:
- LongMemEval — long-term memory over multi-session chat histories.
- BEAM — memory capability over long-context chat cases with rubric-based judging.
Each benchmark runs its own end-to-end pipeline: ingest sessions into an isolated per-item workspace, answer probing questions via an agentic (ReAct) mode, then score answers with an LLM-as-judge.
1. Prerequisites
Install ReMe with dev + core extras (Python 3.11+):
pip install -e ".[dev,core]"
Configure model credentials in a project-root .env file (copied from example.env).
The runners auto-load .env from the repository root. Required variables typically include:
LLM_API_KEY=...
LLM_BASE_URL=...
EMBEDDING_API_KEY=...
EMBEDDING_BASE_URL=...
Model names and component wiring live in the ReMe configs referenced by each benchmark
(reme/config/lme.yaml and reme/config/beam.yaml).
2. Download Datasets
See datasets/README_EN.md for full details.
LongMemEval (downloaded from a HuggingFace mirror):
cd benchmark/datasets/longmemeval
python download.py # downloads the cleaned-S dataset; skips if already present
BEAM (public repository, cloned into benchmark/datasets/):
cd benchmark/datasets
git clone https://github.com/mohammadtavakoli78/BEAM.git
3. Run LongMemEval
From the repository root:
python benchmark/longmemeval/run.py
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
python benchmark/longmemeval/run.py --eval_only # reuse existing workspaces, query + judge only
Pipeline
- Load the dataset (ground truth is embedded in the data file).
- For each item, create an isolated workspace and ingest sessions in chronological order.
- Trigger
auto_dreamwhen consecutive sessions cross the configured hour (default 23:00). - Answer each question via agentic (ReAct) mode.
- Judge the answer (binary yes/no) with the
answer_judgejob and print per-type accuracy.
Key config — benchmark/longmemeval/config.yaml
| Key | Meaning |
|---|---|
dataset.path |
Dataset file to evaluate (e.g. longmemeval_s_reme_cleaned.json); ground truth is included. |
dataset.start_index / num_items |
Slice of items to evaluate. |
dataset.question_types |
Filter by question type; empty = all. |
dataset.workspace_root |
Per-item workspace root (benchmark/memory_workspaces/longmemeval-s). |
evaluation.num_workers |
0 = auto (cpu-2), 1 = sequential, >1 = parallel. |
evaluation.filter_future_sessions |
Only ingest sessions with timestamp ≤ question_date. |
reme.config |
ReMe config used (lme.yaml). |
reme.dream_trigger_hour / dream_scan_days / dream_max_units |
Dream triggering behavior. |
output.dir |
Results directory (benchmark/results/longmemeval). |
4. Run BEAM
From the repository root:
python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q # quiet
python benchmark/beam/run.py --eval_only # reuse existing workspaces, query + judge only
Pipeline
- For each case, load
chat.jsonand convert each batch into a ReMe session. - Ingest sessions in chronological order into an isolated workspace, then
digest_update. - Answer each probing question via agentic (ReAct) mode.
- Score answers with BEAM's rubric-based
answer_judgejob and print per-type averages.
Key config — benchmark/beam/config.yaml
| Key | Meaning |
|---|---|
dataset.beam_root |
BEAM dataset root (benchmark/datasets/BEAM). |
dataset.chat_size |
Variant to run: 100K / 500K / 1M / 10M. |
dataset.case_ids |
Specific cases (e.g. ["1","2"]); empty = all cases. |
dataset.start_index / num_items |
Case pagination (num_items 0 = all). |
dataset.workspace_root |
Per-case workspace root (benchmark/memory_workspaces/beam). |
evaluation.num_workers |
0 = auto, 1 = sequential, >1 = parallel. |
reme.config |
ReMe config used (beam.yaml). |
output.dir |
Results directory (benchmark/results/beam). |
5. Outputs & Logs
- Results: JSON files written to
output.dir(results_<timestamp>.jsonfor LongMemEval,results_<chat_size>_<timestamp>.jsonfor BEAM). A summary with per-type accuracy/score is also printed to the console. - Logs: when
output.log_to_fileis enabled, per-run logs are written tologs/<log_prefix>_<timestamp>/(arunner.logplus oneworker-<pid>.logper worker process).
6. Stopping a Run
Parallel runs spawn a process tree. To terminate a run and all its workers cleanly:
bash benchmark/kill.sh <PID>
The script gracefully sends SIGTERM to the whole process tree, then escalates to
SIGKILL for any process that does not exit within 5 seconds.
7. Reference Results
Recorded evaluation results are available in: