* chore(benchmark): isolate dataset/workspaces/results per benchmark
- Move shared benchmark/{datasets,memory_workspaces,results} into per-benchmark subdirs benchmark/<name>/{dataset,workspaces,results}
- Update beam/longmemeval config.yaml and run.py path defaults
- Relocate longmemeval download.py to benchmark/longmemeval/ (downloads into dataset/ subdir); inline dataset download docs into README
- Update .gitignore: benchmark/*/{dataset,workspaces,results}/
- Move result-{beam,longmemeval}.md to benchmark/results_md/ and drop result- prefix; update README links
- Fix stale path refs in llm_judge.py and logs/demo_search_format.py
* feat(benchmark): add read tool to agentic answer and update BEAM results
- Add 'read' to job_tools in BaseAgenticAnswerStep for file reading capability
- Document read tool usage in lme/agentic_answer.yaml system prompt
- Update result-beam.md with latest evaluation scores (OVERALL: 0.623/0.580)
* feat(auto_memory): add source line-number markers for note traceability
- Add _format_history hook in AutoMemoryStep with line-number annotation
- Override in BeamAutoMemoryStep to prefix each turn with [Ln] for citation
- Add session_file variable to prompt templates for source marker paths
- Simplify repeated extraction rules by referencing system prompt
- Enhance agentic_answer search strategy (multi-search, read tool hint)
- Add warning log on ReadStep failure
* feat(beam): enhance auto_memory with source markers and pilot ingest tooling
* refactor(beam): rename max_chunk_words to max_segment_words, drop one-off pilot scripts
* feat: add CompressorStep and search_v2 dual-mode session compression
- Add CompressorStep (reme/steps/evolve/compressor.py) for direct LLM
text compression with optional query-guided relevance filtering
- Extend search_v2_step to support query-aware and query-independent
session transcript compression via _compress injected kwargs
- Refactor _source_format.py: split into render_chunk_entries +
join_chunk_entries; session chunks now render line-aligned with
L<n>: prefixes for verbatim/compressed parity
- Add JOB_TOOLS and INJECTED_JOB_KWARGS to BaseAgenticAnswerStep for
per-subclass tool and parameter injection
- LmeAgenticAnswerStep injects _search._compress payload to enable
query-aware compression during benchmark evaluation
- Record compression ablation results in result-longmemeval.md
- Add unit tests for CompressorStep and search compression paths
* refactor(compress): relax session compression to lenient format-preserving strategy and update LME results
* refactor(benchmark): make session compression config-driven via compress_session flag
Move session-transcript compression from LME hard-coded injection to a
runtime context flag set by evaluation.compress_session in each
benchmark config. Compression is off by default for both BEAM and LME,
and BaseAgenticAnswerStep now conditionally injects the _search compress
payload only when the flag is truthy.
* feat(lme/auto_memory): add source attribution markers with line numbers
Add _format_history to annotate each turn with [Ln] line numbers and
expose {session_file} in prompts so the agent can emit bare wikilink-style
source markers like [[session/dialog/s1.jsonl#L1-L2,L5-L6]] at the end
of factual entries. Consolidate the per-prompt body/format rules into
references to the system prompt to avoid drift, and add frontmatter-
protection guidance for the edit tool.
* feat: improve agentic answer prompt and update beam 100K results
- Strengthen abstention rule: prohibit extrapolation from related but
non-direct evidence
- Add multi-angle search after preliminary answer to check for
conflicting/supplementary/updated information
- Add max-iteration fallback to 'Information not found'
- Update beam.md with 100K results (agentscope 2.0.4.post1, from scratch)
including per-type token consumption and memory construction stats
- config.yaml: 100K dataset, 20 workers for BEAM evaluation
- run.py: add memory construction token usage tracking (default agent)
- Overall: 0.635 → 0.654 (+0.019), contradiction_resolution: 0.338 → 0.478
(+0.140), abstention: 0.500 → 0.525 (+0.025)
* feat(read): add session-aware formatting for read tool and update BEAM eval
- Add truncate_session_output in _file_io.py to render jsonl session
lines as [speaker @ time] content before byte-budget truncation
- Add read_step_format_session flag to ReadStep, honoring injected
job kwargs (precedence) and YAML fallback
- Inject read_step_format_session=True into BaseAgenticAnswerStep
so agentic answer reads render session transcripts human-readably
- Refine BEAM agentic_answer prompt: continue multi-angle search
after preliminary answer, forbid fabrication/extrapolation
- Update BEAM config to 1M variant and add sequential 100K-eval /
1M-build shell script
- Refresh benchmark/results_md/beam.md with latest results
* chore(config): disable expand_links in beam and lme search_v2 configs
* refactor(beam): drop one-off sequential 100K-eval-then-1M-build script
* fix(benchmark): add compressor job to beam config and fix BEAM clone instructions
- Add compressor job and compressor as_llm component to reme/config/beam.yaml
(aligned with lme.yaml) so that compress_session: true works for BEAM
- Add graceful degradation guard in search_v2._compress_session_entries:
when the compressor job is missing from the active config, log a warning
and skip compression instead of raising 'Job compressor not found'.
Skipped when there is no app_context so unit tests mocking run_job still
drive compression behavior.
- Fix BEAM download instructions in README.md/README_ZH.md: add mkdir -p
before cd benchmark/beam/dataset (the directory is gitignored and absent
in a fresh clone)
* fix(steps): guard compressor exceptions and fix ReadStep boolean override
1. search_v2: catch per-entry exceptions from run_job('compressor') inside
compress() so asyncio.gather never propagates a compressor failure (e.g.
temporary LLM outage). The failing entry keeps its original body while
remaining entries are still compressed, preserving already-retrieved
search results.
2. read: replace 'context_value or yaml_value' with an existence check so
that a runtime-injected False can explicitly disable a YAML-true
read_step_format_session flag.
Add focused unit tests for both paths.
* fix(search_v2): use existence check for strict_date_filter boolean override
Replace 'context_value or yaml_value' with an existence-based check so
that a runtime-injected False can explicitly disable a YAML-true
strict_date_filter flag, consistent with the read_step_format_session fix.
* refactor(search): simplify strict_date_filter fallback to truthiness-or
* style(test): rename unused param to satisfy pylint W0613
---------
Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
|
||
|---|---|---|
| .. | ||
| beam | ||
| longmemeval | ||
| results_md | ||
| toolmemory | ||
| kill.sh | ||
| README.md | ||
| README_ZH.md | ||
ReMe Benchmarks
Reproduction guide for the two memory benchmarks shipped with ReMe:
- LongMemEval — long-term memory over multi-session chat histories.
- BEAM — memory capability over long-context chat cases with rubric-based judging.
Each benchmark runs its own end-to-end pipeline: ingest sessions into an isolated per-item workspace, answer probing questions via an agentic (ReAct) mode, then score answers with an LLM-as-judge.
1. Prerequisites
Install ReMe with dev + core extras (Python 3.11+):
pip install -e ".[dev,core]"
Configure model credentials in a project-root .env file (copied from example.env).
The runners auto-load .env from the repository root. Required variables typically include:
LLM_API_KEY=...
LLM_BASE_URL=...
EMBEDDING_API_KEY=...
EMBEDDING_BASE_URL=...
Model names and component wiring live in the ReMe configs referenced by each benchmark
(reme/config/lme.yaml and reme/config/beam.yaml).
2. Download Datasets
Each benchmark keeps its own data under its directory:
benchmark/<name>/dataset (input data), benchmark/<name>/workspaces
(per-item memory workspaces), and benchmark/<name>/results (evaluation outputs).
All three are excluded from Git.
LongMemEval — ReMe uses only the cleaned-S split, hosted on HuggingFace:
agentscope-ai/ReMe_longmemeval_clean_s_v2
(the script downloads via the hf-mirror.com mirror; to use a different mirror,
modify BASE_URL in download.py):
cd benchmark/longmemeval
python download.py # saves dataset/longmemeval_s_reme_cleaned.json; skips if already present
BEAM (public repository, cloned into benchmark/beam/dataset/):
mkdir -p benchmark/beam/dataset
cd benchmark/beam/dataset
git clone https://github.com/mohammadtavakoli78/BEAM.git
After cloning, benchmark/beam/dataset/BEAM/ should contain chats/, src/,
topics/ and other subdirectories.
3. Run LongMemEval
From the repository root:
python benchmark/longmemeval/run.py
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
python benchmark/longmemeval/run.py --eval_only # reuse existing workspaces, query + judge only
Pipeline
- Load the dataset (ground truth is embedded in the data file).
- For each item, create an isolated workspace and ingest sessions in chronological order.
- Trigger
auto_dreamwhen consecutive sessions cross the configured hour (default 23:00). - Answer each question via agentic (ReAct) mode.
- Judge the answer (binary yes/no) with the
answer_judgejob and print per-type accuracy.
Key config — benchmark/longmemeval/config.yaml
| Key | Meaning |
|---|---|
dataset.path |
Dataset file to evaluate (e.g. longmemeval_s_reme_cleaned.json); ground truth is included. |
dataset.start_index / num_items |
Slice of items to evaluate. |
dataset.question_types |
Filter by question type; empty = all. |
dataset.workspace_root |
Per-item workspace root (benchmark/longmemeval/workspaces/longmemeval-s). |
evaluation.num_workers |
0 = auto (cpu-2), 1 = sequential, >1 = parallel. |
evaluation.filter_future_sessions |
Only ingest sessions with timestamp ≤ question_date. |
reme.config |
ReMe config used (lme.yaml). |
reme.dream_trigger_hour / dream_scan_days / dream_max_units |
Dream triggering behavior. |
output.dir |
Results directory (benchmark/longmemeval/results). |
4. Run BEAM
From the repository root:
python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q # quiet
python benchmark/beam/run.py --eval_only # reuse existing workspaces, query + judge only
Pipeline
- For each case, load
chat.jsonand convert each batch into a ReMe session. - Ingest sessions in chronological order into an isolated workspace, then
digest_update. - Answer each probing question via agentic (ReAct) mode.
- Score answers with BEAM's rubric-based
answer_judgejob and print per-type averages.
Key config — benchmark/beam/config.yaml
| Key | Meaning |
|---|---|
dataset.beam_root |
BEAM dataset root (benchmark/beam/dataset/BEAM). |
dataset.chat_size |
Variant to run: 100K / 500K / 1M / 10M. |
dataset.case_ids |
Specific cases (e.g. ["1","2"]); empty = all cases. |
dataset.start_index / num_items |
Case pagination (num_items 0 = all). |
dataset.workspace_root |
Per-case workspace root (benchmark/beam/workspaces/beam). |
evaluation.num_workers |
0 = auto, 1 = sequential, >1 = parallel. |
reme.config |
ReMe config used (beam.yaml). |
output.dir |
Results directory (benchmark/beam/results). |
5. Outputs & Logs
- Results: JSON files written to
output.dir(results_<timestamp>.jsonfor LongMemEval,results_<chat_size>_<timestamp>.jsonfor BEAM). A summary with per-type accuracy/score is also printed to the console. - Logs: when
output.log_to_fileis enabled, per-run logs are written tologs/<log_prefix>_<timestamp>/(arunner.logplus oneworker-<pid>.logper worker process).
6. Stopping a Run
Parallel runs spawn a process tree. To terminate a run and all its workers cleanly:
bash benchmark/kill.sh <PID>
The script gracefully sends SIGTERM to the whole process tree, then escalates to
SIGKILL for any process that does not exit within 5 seconds.
7. Reference Results
Recorded evaluation results are available in: