* refactor(plugins): move BaseAgenticAnswerStep into lme/beam plugins Move the shared benchmark answer base class from the core reme.steps.benchmark package into each plugin's own src tree (reme_lme.base_agentic_answer / reme_beam.base_agentic_answer) with absolute imports, drop the core benchmark steps package, and update plugin READMEs accordingly. * refactor(plugins): simplify benchmark answer steps * docs(plugins): clarify benchmark answer step migration --------- Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com> |
||
|---|---|---|
| .. | ||
| src/reme_lme | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
| README_ZH.md | ||
LongMemEval plugin
This plugin owns the LongMemEval memory and agentic-answer Steps and their Job defaults.
The trusted judge Step, prompts and answer_judge Job live in plugins/lme-judge. ReMe's built-in benchmark.yaml owns the
shared evaluation Jobs and components. Dataset handling, the runner and results remain
in benchmark/longmemeval.
From the repository root, install ReMe and this plugin in editable mode before running the benchmark:
python -m pip install -e ".[as]"
reme plugins install ./plugins/lme --editable
reme plugins install ./plugins/lme-judge --editable
reme plugins validate lme
reme plugins validate lme-judge
python benchmark/longmemeval/run.py
Editable installation registers the lme entry point while keeping source changes immediately
visible. The runner selects the built-in benchmark preset and explicitly enables lme and lme-judge for
each Application. Installing the plugin makes it discoverable but does not enable it globally.
plugin.yaml registers backends and contributes the plugin-owned auto_memory,
agentic_answer Job defaults. Start a full benchmark application with
reme start config=benchmark plugins='["lme", "lme-judge"]'. The shared preset does not inherit
default: only declared Jobs run, indexing is manual, and neither scheduled dream
nor the optional auto_dream Job is enabled.
The existing auto_memory, agentic_answer, answer_judge, bench and judge
names and model environment variables are unchanged. Explicit application/CLI overrides
still take precedence. Installing this plugin does not start an evaluation.
Custom Python callers should import memory, search and answer Steps from reme_lme, and install
lme-judge before importing the judge Step from judge_lme. After uninstalling,
Applications and CLI services must omit the plugin until it is installed again.
LmeAgenticAnswerStep now implements the answer behavior directly. The former
reme.steps.benchmark.BaseAgenticAnswerStep import is gone; custom subclasses can
extend reme_lme.LmeAgenticAnswerStep for LongMemEval behavior, or implement their
own Step using reme.steps.base_step.BaseStep.
Uninstallation never removes datasets, workspaces or results.
Restart an existing service after changing plugins.