mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-10 03:30:56 +00:00
* refactor(plugins): move BaseAgenticAnswerStep into lme/beam plugins Move the shared benchmark answer base class from the core reme.steps.benchmark package into each plugin's own src tree (reme_lme.base_agentic_answer / reme_beam.base_agentic_answer) with absolute imports, drop the core benchmark steps package, and update plugin READMEs accordingly. * refactor(plugins): simplify benchmark answer steps * docs(plugins): clarify benchmark answer step migration --------- Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com>
42 lines
2.3 KiB
Markdown
42 lines
2.3 KiB
Markdown
# LongMemEval plugin
|
|
|
|
[中文说明](./README_ZH.md)
|
|
|
|
This plugin owns the LongMemEval memory and agentic-answer Steps and their Job defaults.
|
|
The trusted judge Step, prompts and `answer_judge` Job live in `plugins/lme-judge`. ReMe's built-in `benchmark.yaml` owns the
|
|
shared evaluation Jobs and components. Dataset handling, the runner and results remain
|
|
in [`benchmark/longmemeval`](../../benchmark/longmemeval/README.md).
|
|
|
|
From the repository root, install ReMe and this plugin in editable mode before running the benchmark:
|
|
|
|
```bash
|
|
python -m pip install -e ".[as]"
|
|
reme plugins install ./plugins/lme --editable
|
|
reme plugins install ./plugins/lme-judge --editable
|
|
reme plugins validate lme
|
|
reme plugins validate lme-judge
|
|
python benchmark/longmemeval/run.py
|
|
```
|
|
|
|
Editable installation registers the `lme` entry point while keeping source changes immediately
|
|
visible. The runner selects the built-in `benchmark` preset and explicitly enables `lme` and `lme-judge` for
|
|
each Application. Installing the plugin makes it discoverable but does not enable it globally.
|
|
|
|
`plugin.yaml` registers backends and contributes the plugin-owned `auto_memory`,
|
|
`agentic_answer` Job defaults. Start a full benchmark application with
|
|
`reme start config=benchmark plugins='["lme", "lme-judge"]'`. The shared preset does not inherit
|
|
`default`: only declared Jobs run, indexing is manual, and neither scheduled dream
|
|
nor the optional `auto_dream` Job is enabled.
|
|
The existing `auto_memory`, `agentic_answer`, `answer_judge`, `bench` and `judge`
|
|
names and model environment variables are unchanged. Explicit application/CLI overrides
|
|
still take precedence. Installing this plugin does not start an evaluation.
|
|
|
|
Custom Python callers should import memory, search and answer Steps from `reme_lme`, and install
|
|
`lme-judge` before importing the judge Step from `judge_lme`. After uninstalling,
|
|
Applications and CLI services must omit the plugin until it is installed again.
|
|
`LmeAgenticAnswerStep` now implements the answer behavior directly. The former
|
|
`reme.steps.benchmark.BaseAgenticAnswerStep` import is gone; custom subclasses can
|
|
extend `reme_lme.LmeAgenticAnswerStep` for LongMemEval behavior, or implement their
|
|
own Step using `reme.steps.base_step.BaseStep`.
|
|
Uninstallation never removes datasets, workspaces or results.
|
|
Restart an existing service after changing plugins.
|