ReMe/plugins/lme/README.md
xyf2020 67936d5a43
refactor(plugins): move BaseAgenticAnswerStep out of the core benchmark steps package into the lme/beam plugins (#576)
* refactor(plugins): move BaseAgenticAnswerStep into lme/beam plugins

Move the shared benchmark answer base class from the core
reme.steps.benchmark package into each plugin's own src tree
(reme_lme.base_agentic_answer / reme_beam.base_agentic_answer) with
absolute imports, drop the core benchmark steps package, and update
plugin READMEs accordingly.

* refactor(plugins): simplify benchmark answer steps

* docs(plugins): clarify benchmark answer step migration

---------

Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com>
2026-09-30 12:19:32 +08:00

42 lines
2.3 KiB
Markdown

# LongMemEval plugin
[中文说明](./README_ZH.md)
This plugin owns the LongMemEval memory and agentic-answer Steps and their Job defaults.
The trusted judge Step, prompts and `answer_judge` Job live in `plugins/lme-judge`. ReMe's built-in `benchmark.yaml` owns the
shared evaluation Jobs and components. Dataset handling, the runner and results remain
in [`benchmark/longmemeval`](../../benchmark/longmemeval/README.md).
From the repository root, install ReMe and this plugin in editable mode before running the benchmark:
```bash
python -m pip install -e ".[as]"
reme plugins install ./plugins/lme --editable
reme plugins install ./plugins/lme-judge --editable
reme plugins validate lme
reme plugins validate lme-judge
python benchmark/longmemeval/run.py
```
Editable installation registers the `lme` entry point while keeping source changes immediately
visible. The runner selects the built-in `benchmark` preset and explicitly enables `lme` and `lme-judge` for
each Application. Installing the plugin makes it discoverable but does not enable it globally.
`plugin.yaml` registers backends and contributes the plugin-owned `auto_memory`,
`agentic_answer` Job defaults. Start a full benchmark application with
`reme start config=benchmark plugins='["lme", "lme-judge"]'`. The shared preset does not inherit
`default`: only declared Jobs run, indexing is manual, and neither scheduled dream
nor the optional `auto_dream` Job is enabled.
The existing `auto_memory`, `agentic_answer`, `answer_judge`, `bench` and `judge`
names and model environment variables are unchanged. Explicit application/CLI overrides
still take precedence. Installing this plugin does not start an evaluation.
Custom Python callers should import memory, search and answer Steps from `reme_lme`, and install
`lme-judge` before importing the judge Step from `judge_lme`. After uninstalling,
Applications and CLI services must omit the plugin until it is installed again.
`LmeAgenticAnswerStep` now implements the answer behavior directly. The former
`reme.steps.benchmark.BaseAgenticAnswerStep` import is gone; custom subclasses can
extend `reme_lme.LmeAgenticAnswerStep` for LongMemEval behavior, or implement their
own Step using `reme.steps.base_step.BaseStep`.
Uninstallation never removes datasets, workspaces or results.
Restart an existing service after changing plugins.