ReMe/benchmark/beam
xyf2020 23d4c96c15
refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422)
* chore(benchmark): isolate dataset/workspaces/results per benchmark

- Move shared benchmark/{datasets,memory_workspaces,results} into per-benchmark subdirs benchmark/<name>/{dataset,workspaces,results}
- Update beam/longmemeval config.yaml and run.py path defaults
- Relocate longmemeval download.py to benchmark/longmemeval/ (downloads into dataset/ subdir); inline dataset download docs into README
- Update .gitignore: benchmark/*/{dataset,workspaces,results}/
- Move result-{beam,longmemeval}.md to benchmark/results_md/ and drop result- prefix; update README links
- Fix stale path refs in llm_judge.py and logs/demo_search_format.py

* feat(benchmark): add read tool to agentic answer and update BEAM results

- Add 'read' to job_tools in BaseAgenticAnswerStep for file reading capability
- Document read tool usage in lme/agentic_answer.yaml system prompt
- Update result-beam.md with latest evaluation scores (OVERALL: 0.623/0.580)

* feat(auto_memory): add source line-number markers for note traceability

- Add _format_history hook in AutoMemoryStep with line-number annotation
- Override in BeamAutoMemoryStep to prefix each turn with [Ln] for citation
- Add session_file variable to prompt templates for source marker paths
- Simplify repeated extraction rules by referencing system prompt
- Enhance agentic_answer search strategy (multi-search, read tool hint)
- Add warning log on ReadStep failure

* feat(beam): enhance auto_memory with source markers and pilot ingest tooling

* refactor(beam): rename max_chunk_words to max_segment_words, drop one-off pilot scripts

* feat: add CompressorStep and search_v2 dual-mode session compression

- Add CompressorStep (reme/steps/evolve/compressor.py) for direct LLM
  text compression with optional query-guided relevance filtering
- Extend search_v2_step to support query-aware and query-independent
  session transcript compression via _compress injected kwargs
- Refactor _source_format.py: split into render_chunk_entries +
  join_chunk_entries; session chunks now render line-aligned with
  L<n>: prefixes for verbatim/compressed parity
- Add JOB_TOOLS and INJECTED_JOB_KWARGS to BaseAgenticAnswerStep for
  per-subclass tool and parameter injection
- LmeAgenticAnswerStep injects _search._compress payload to enable
  query-aware compression during benchmark evaluation
- Record compression ablation results in result-longmemeval.md
- Add unit tests for CompressorStep and search compression paths

* refactor(compress): relax session compression to lenient format-preserving strategy and update LME results

* refactor(benchmark): make session compression config-driven via compress_session flag

Move session-transcript compression from LME hard-coded injection to a
runtime context flag set by evaluation.compress_session in each
benchmark config. Compression is off by default for both BEAM and LME,
and BaseAgenticAnswerStep now conditionally injects the _search compress
payload only when the flag is truthy.

* feat(lme/auto_memory): add source attribution markers with line numbers

Add _format_history to annotate each turn with [Ln] line numbers and
expose {session_file} in prompts so the agent can emit bare wikilink-style
source markers like [[session/dialog/s1.jsonl#L1-L2,L5-L6]] at the end
of factual entries. Consolidate the per-prompt body/format rules into
references to the system prompt to avoid drift, and add frontmatter-
protection guidance for the edit tool.

* feat: improve agentic answer prompt and update beam 100K results

- Strengthen abstention rule: prohibit extrapolation from related but
  non-direct evidence
- Add multi-angle search after preliminary answer to check for
  conflicting/supplementary/updated information
- Add max-iteration fallback to 'Information not found'
- Update beam.md with 100K results (agentscope 2.0.4.post1, from scratch)
  including per-type token consumption and memory construction stats
- config.yaml: 100K dataset, 20 workers for BEAM evaluation
- run.py: add memory construction token usage tracking (default agent)
- Overall: 0.635 → 0.654 (+0.019), contradiction_resolution: 0.338 → 0.478
  (+0.140), abstention: 0.500 → 0.525 (+0.025)

* feat(read): add session-aware formatting for read tool and update BEAM eval

- Add truncate_session_output in _file_io.py to render jsonl session
  lines as [speaker @ time] content before byte-budget truncation
- Add read_step_format_session flag to ReadStep, honoring injected
  job kwargs (precedence) and YAML fallback
- Inject read_step_format_session=True into BaseAgenticAnswerStep
  so agentic answer reads render session transcripts human-readably
- Refine BEAM agentic_answer prompt: continue multi-angle search
  after preliminary answer, forbid fabrication/extrapolation
- Update BEAM config to 1M variant and add sequential 100K-eval /
  1M-build shell script
- Refresh benchmark/results_md/beam.md with latest results

* chore(config): disable expand_links in beam and lme search_v2 configs

* refactor(beam): drop one-off sequential 100K-eval-then-1M-build script

* fix(benchmark): add compressor job to beam config and fix BEAM clone instructions

- Add compressor job and compressor as_llm component to reme/config/beam.yaml
  (aligned with lme.yaml) so that compress_session: true works for BEAM
- Add graceful degradation guard in search_v2._compress_session_entries:
  when the compressor job is missing from the active config, log a warning
  and skip compression instead of raising 'Job compressor not found'.
  Skipped when there is no app_context so unit tests mocking run_job still
  drive compression behavior.
- Fix BEAM download instructions in README.md/README_ZH.md: add mkdir -p
  before cd benchmark/beam/dataset (the directory is gitignored and absent
  in a fresh clone)

* fix(steps): guard compressor exceptions and fix ReadStep boolean override

1. search_v2: catch per-entry exceptions from run_job('compressor') inside
   compress() so asyncio.gather never propagates a compressor failure (e.g.
   temporary LLM outage). The failing entry keeps its original body while
   remaining entries are still compressed, preserving already-retrieved
   search results.

2. read: replace 'context_value or yaml_value' with an existence check so
   that a runtime-injected False can explicitly disable a YAML-true
   read_step_format_session flag.

Add focused unit tests for both paths.

* fix(search_v2): use existence check for strict_date_filter boolean override

Replace 'context_value or yaml_value' with an existence-based check so
that a runtime-injected False can explicitly disable a YAML-true
strict_date_filter flag, consistent with the read_step_format_session fix.

* refactor(search): simplify strict_date_filter fallback to truthiness-or

* style(test): rename unused param to satisfy pylint W0613

* refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt

- Move shared benchmark/README, README_ZH, kill.sh, and results_md/*.md into
  per-benchmark subdirs (benchmark/beam/, benchmark/longmemeval/) so each
  benchmark owns its own docs, scripts, and result snapshots.
- Simplify lme/agentic_answer.yaml system prompt: drop verbose memory-system
  description, keep search strategy, draft tool, and answer rules concise.

* docs(benchmark): update LME README_ZH results to latest eval run

---------

Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
2026-08-06 15:13:52 +08:00
..
config.yaml feat(benchmark): enhance session memory retrieval and isolate benchmark assets (#409) 2026-08-05 19:23:42 +08:00
kill.sh refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422) 2026-08-06 15:13:52 +08:00
README.md refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422) 2026-08-06 15:13:52 +08:00
README_ZH.md refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422) 2026-08-06 15:13:52 +08:00
run.py feat(benchmark): enhance session memory retrieval and isolate benchmark assets (#409) 2026-08-05 19:23:42 +08:00

中文版 / Chinese version

BEAM Benchmark

BEAM is a benchmark for memory capability over long-context chat cases. Each case contains a very long chat history split into batches; ReMe converts each batch into a session, ingests them in chronological order, then answers probing questions via an agentic (ReAct) mode. Answers are scored with BEAM's rubric-based answer_judge job, which produces both a graded score and a binary verdict, and per-type averages are reported.

BEAM ships dataset variants by chat size — 100K / 500K / 1M / 10M — so memory systems can be stressed at different context lengths. Question types include abstention, contradiction resolution, event ordering, information extraction, instruction following, knowledge update, multi-session reasoning, preference following, summarization, and temporal reasoning.

For the shared setup (dependencies, credentials, log conventions) see the top-level benchmark README.

1. Get the Dataset

BEAM is a public repository, cloned into benchmark/beam/dataset/:

mkdir -p benchmark/beam/dataset
cd benchmark/beam/dataset
git clone https://github.com/mohammadtavakoli78/BEAM.git

After cloning, benchmark/beam/dataset/BEAM/ should contain chats/, src/, topics/ and other subdirectories.

2. Run

From the repository root:

python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q                        # quiet
python benchmark/beam/run.py --eval_only               # reuse existing workspaces, query + judge only

3. Pipeline

  1. For each case, load chat.json and convert each batch into a ReMe session.
  2. Ingest sessions in chronological order into an isolated workspace, then digest_update.
  3. Answer each probing question via agentic (ReAct) mode.
  4. Score answers with BEAM's rubric-based answer_judge job and print per-type averages.

4. Key config — benchmark/beam/config.yaml

Key Meaning
dataset.beam_root BEAM dataset root (benchmark/beam/dataset/BEAM).
dataset.chat_size Variant to run: 100K / 500K / 1M / 10M.
dataset.case_ids Specific cases (e.g. ["1","2"]); empty = all cases.
dataset.start_index / num_items Case pagination (num_items 0 = all).
dataset.workspace_root Per-case workspace root (benchmark/beam/workspaces/beam).
evaluation.num_workers 0 = auto, 1 = sequential, >1 = parallel.
reme.config ReMe config used (beam.yaml).
output.dir Results directory (benchmark/beam/results).

5. Outputs

Results are JSON files written to output.dir as results_<chat_size>_<timestamp>.json, with a per-type score summary also printed to the console. Logging conventions are shared across benchmarks — see the top-level README.

6. Reference Results

The results below use the longmemeval-version prompt.

100K

agentscope==2.0.4.post1, conda reme env, 20 workers, eval-only (reusing prebuilt memory) (2026-08-05, 20 cases / 400 Qs, total 46.0 min)

Type Agentic Binary input tok/q output tok/q total tok/q tool calls/q
abstention 0.550 0.550 96,031 1,070 97,101 4.58
contradiction_resolution 0.438 0.412 32,263 872 33,135 2.48
event_ordering 0.501 0.423 140,195 5,163 145,358 4.70
information_extraction 0.873 0.832 50,245 883 51,128 3.15
instruction_following 0.750 0.725 37,986 848 38,834 2.67
knowledge_update 0.688 0.675 31,198 651 31,849 2.27
multi_session_reasoning 0.626 0.584 85,038 4,563 89,601 4.28
preference_following 0.925 0.912 34,281 989 35,270 2.50
summarization 0.623 0.461 89,657 2,056 91,713 4.12
temporal_reasoning 0.637 0.625 34,563 1,049 35,612 2.52
OVERALL 0.661 0.620 63,146 1,814 64,960 3.33

Memory Construction average token consumption (default agent, full build over 20 cases):

Agent input tok/case output tok/case total tok/case
default 2,172,316 136,697 2,309,013

1M

agentscope==2.0.4.post1, conda reme env, 20 workers, full memory build (2026-08-05, 35 cases / 700 Qs, total 459.2 min)

Type Agentic Binary input tok/q output tok/q total tok/q tool calls/q
abstention 0.429 0.429 118,707 1,178 119,886 4.20
contradiction_resolution 0.391 0.364 49,787 810 50,597 2.50
event_ordering 0.558 0.456 201,514 3,889 205,403 4.79
information_extraction 0.809 0.772 78,950 894 79,844 3.00
instruction_following 0.852 0.832 55,757 924 56,681 2.81
knowledge_update 0.779 0.771 45,981 665 46,646 2.37
multi_session_reasoning 0.658 0.612 138,133 2,873 141,006 4.40
preference_following 0.798 0.777 51,796 920 52,716 2.53
summarization 0.693 0.537 158,794 2,905 161,700 4.44
temporal_reasoning 0.536 0.536 100,176 3,148 103,324 3.90
OVERALL 0.650 0.609 99,959 1,821 101,780 3.49

Memory Construction average token consumption (default agent, full build over 35 cases):

Agent input tok/case output tok/case total tok/case
default 31,943,817 1,417,061 33,360,878