ReMe/plugins/lme/README.md
xyf2020 9975bb37b9
Some checks failed
CI / Python tests / Unit Tests - py3.12 (push) Waiting to run
CI / Python tests / Unit Tests - py3.13 (push) Waiting to run
CI / Windows / CLI smoke - py3.11 (push) Waiting to run
Deploy / Documentation / Build documentation (push) Waiting to run
Deploy / Documentation / deploy (push) Blocked by required conditions
CI / Python packages / Build and verify distributions (push) Waiting to run
CI / Python quality / Pre-commit (push) Waiting to run
CI / Python tests / Unit Tests - py3.11 (push) Waiting to run
Security / CodeQL / Analyze javascript-typescript (push) Waiting to run
Security / CodeQL / Analyze python (push) Waiting to run
CI / Documentation / Test and build documentation (push) Has been cancelled
Separate benchmark judge plugins (#535)
2026-09-10 20:16:46 +08:00

2.2 KiB

LongMemEval plugin

中文说明

This plugin owns the LongMemEval memory and agentic-answer Steps and their Job defaults. The trusted judge Step, prompts and answer_judge Job live in plugins/lme-judge. ReMe's built-in benchmark.yaml owns the shared evaluation Jobs and components. Dataset handling, the runner and results remain in benchmark/longmemeval.

From the repository root, install ReMe and this plugin in editable mode before running the benchmark:

python -m pip install -e ".[as]"
reme plugins install ./plugins/lme --editable
reme plugins install ./plugins/lme-judge --editable
reme plugins validate lme
reme plugins validate lme-judge
python benchmark/longmemeval/run.py

Editable installation registers the lme entry point while keeping source changes immediately visible. The runner selects the built-in benchmark preset and explicitly enables lme and lme-judge for each Application. Installing the plugin makes it discoverable but does not enable it globally.

plugin.yaml registers backends and contributes the plugin-owned auto_memory, agentic_answer Job defaults. Start a full benchmark application with reme start config=benchmark plugins='["lme", "lme-judge"]'. The shared preset does not inherit default: only declared Jobs run, indexing is manual, and neither scheduled dream nor the optional auto_dream Job is enabled. The existing auto_memory, agentic_answer, answer_judge, bench and judge names and model environment variables are unchanged. Explicit application/CLI overrides still take precedence. Installing this plugin does not start an evaluation.

The shared answer base class lives in reme.steps.benchmark.base_agentic_answer. The old core-owned reme.steps.benchmark.lme Python import path is removed. Custom Python callers should import memory, search and answer Steps from reme_lme, and install lme-judge before importing the judge Step from judge_lme. After uninstalling, Applications and CLI services must omit the plugin until it is installed again. Uninstallation never removes datasets, workspaces or results. Restart an existing service after changing plugins.