mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-09 03:20:54 +00:00
Compare commits
151 commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
084c02e43a | ||
|
|
dc12798526 | ||
|
|
090c8c24ab | ||
|
|
eb2f466a30 | ||
|
|
4c54c2b650 | ||
|
|
49bbbc93ff | ||
|
|
d529ec5256 | ||
|
|
1648b7ce93 | ||
|
|
65b48ec5fc | ||
|
|
81c16d3e2c | ||
|
|
67135cfc57 | ||
|
|
67936d5a43 | ||
|
|
30227e509f | ||
|
|
bebad36745 | ||
|
|
8b5456641f | ||
|
|
c9a9728164 | ||
|
|
ce386f5391 | ||
|
|
873bcee220 | ||
|
|
19472233be | ||
|
|
59e4ed7d3b | ||
|
|
07d4d6e838 | ||
|
|
6125fc197d | ||
|
|
cae613d1c4 | ||
|
|
5231f3970c | ||
|
|
9ebe17a89d | ||
|
|
16269c9a76 | ||
|
|
6c5d2194c0 | ||
|
|
5f2c693ddb | ||
|
|
dab56fc794 | ||
|
|
6f4bdfd416 | ||
|
|
b9caae1e50 | ||
|
|
d67f1490f5 | ||
|
|
fe336da566 | ||
|
|
fd6337fe42 | ||
|
|
6ed97f033b | ||
|
|
46eca95bb9 | ||
|
|
4f7c8786e3 | ||
|
|
a518dd168b | ||
|
|
9ad3dafce5 | ||
|
|
05958d4d8b | ||
|
|
dff2d33cec | ||
|
|
6cb81e7921 | ||
|
|
1be61b1e4c | ||
|
|
7e25d4679b | ||
|
|
9975bb37b9 | ||
|
|
06fb46fa48 | ||
|
|
1f67a6ce29 | ||
|
|
354837f9af | ||
|
|
f04eedb3ab | ||
|
|
36e3a87c75 | ||
|
|
5c17874f73 | ||
|
|
0eba6ea831 | ||
|
|
193fd418fb | ||
|
|
8c3d3016e3 | ||
|
|
dc28e62526 | ||
|
|
88ed21165b | ||
|
|
3f2eb6235f | ||
|
|
8c4898999d | ||
|
|
c1b85f8241 | ||
|
|
f9a45a319a | ||
|
|
65cb4ebdd6 | ||
|
|
bec7e48772 | ||
|
|
fc4a5398a8 | ||
|
|
21f7757c80 | ||
|
|
c85917a812 | ||
|
|
157b096448 | ||
|
|
99afc2604f | ||
|
|
2dd2255760 | ||
|
|
d8d667c6ac | ||
|
|
940a923f06 | ||
|
|
3d2ecc60d2 | ||
|
|
6f38d201b6 | ||
|
|
ef3f99f019 | ||
|
|
b78e32ef03 | ||
|
|
6a6e0b3c29 | ||
|
|
a457bf7542 | ||
|
|
513fb5b7f4 | ||
|
|
1a6b584274 | ||
|
|
15d12be6b6 | ||
|
|
626c850ccb | ||
|
|
01ef1a6efb | ||
|
|
efcc2b34d1 | ||
|
|
c8e1248769 | ||
|
|
8416fd3ac9 | ||
|
|
f44f52d919 | ||
|
|
ebcb154e37 | ||
|
|
39233f4e62 | ||
|
|
94b7dedc26 | ||
|
|
87187c1d25 | ||
|
|
f5ec230fef | ||
|
|
2f5fd46b44 | ||
|
|
618e8cec66 | ||
|
|
d3aee1adf5 | ||
|
|
6b9a75267b | ||
|
|
c792fd197c | ||
|
|
fd2894f939 | ||
|
|
2a05914150 | ||
|
|
29eb51d7ba | ||
|
|
da9a8b7810 | ||
|
|
dbf2a17da6 | ||
|
|
64249873ce | ||
|
|
28fa636506 | ||
|
|
52fdd446fb | ||
|
|
ab66f2bb56 | ||
|
|
215c1f72f2 | ||
|
|
9533c17d51 | ||
|
|
b8f48c8004 | ||
|
|
3924f89bb4 | ||
|
|
c7dbf31c3f | ||
|
|
58276f740b | ||
|
|
3095564313 | ||
|
|
21057931a9 | ||
|
|
5a5855f5ff | ||
|
|
072cb6a55b | ||
|
|
fca42f4e6c | ||
|
|
d5e0d2837b | ||
|
|
e05b201da9 | ||
|
|
765103a597 | ||
|
|
168b7194ab | ||
|
|
e7b9274190 | ||
|
|
c5d92a24ab | ||
|
|
9218a2d0e3 | ||
|
|
6503e1271c | ||
|
|
23d4c96c15 | ||
|
|
f31daf1949 | ||
|
|
b00eb0a9ea | ||
|
|
ad4f23e4dc | ||
|
|
5bc46c88b6 | ||
|
|
e256c556ca | ||
|
|
d2b8872f2e | ||
|
|
eac8223387 | ||
|
|
dc7df26e95 | ||
|
|
a9ec334adc | ||
|
|
6b035c6553 | ||
|
|
3d487d8d45 | ||
|
|
f3d32e203d | ||
|
|
550317c3bf | ||
|
|
a367c2ce13 | ||
|
|
c937be9d94 | ||
|
|
4eb2adf961 | ||
|
|
f34dcdb09b | ||
|
|
2f79977df0 | ||
|
|
0522135791 | ||
|
|
11fe50d89c | ||
|
|
1687179f84 | ||
|
|
46adb5ae1e | ||
|
|
630f26b119 | ||
|
|
7b1da5a9ee | ||
|
|
e7d44f6f3b | ||
|
|
b4333fbef8 | ||
|
|
55ef4bd6ad |
794 changed files with 115179 additions and 12883 deletions
40
.dockerignore
Normal file
40
.dockerignore
Normal file
|
|
@ -0,0 +1,40 @@
|
||||||
|
.git
|
||||||
|
**/.DS_Store
|
||||||
|
**/.ssh
|
||||||
|
**/.codex
|
||||||
|
**/.claude
|
||||||
|
**/private*
|
||||||
|
**/.env
|
||||||
|
**/.env.*
|
||||||
|
**/.reme
|
||||||
|
**/.venv
|
||||||
|
**/venv
|
||||||
|
**/__pycache__
|
||||||
|
**/*.py[cod]
|
||||||
|
**/.pytest_cache
|
||||||
|
**/.mypy_cache
|
||||||
|
**/.ruff_cache
|
||||||
|
**/.coverage*
|
||||||
|
**/htmlcov
|
||||||
|
**/node_modules
|
||||||
|
**/dist
|
||||||
|
**/dist-static
|
||||||
|
**/.generated
|
||||||
|
**/.next
|
||||||
|
**/.vite
|
||||||
|
**/.wrangler
|
||||||
|
**/.cache
|
||||||
|
**/*.egg-info
|
||||||
|
**/build
|
||||||
|
**/logs
|
||||||
|
**/*.log
|
||||||
|
**/*.tmp
|
||||||
|
reme_studio/src/reme_studio/static
|
||||||
|
benchmark
|
||||||
|
cookbook
|
||||||
|
docs
|
||||||
|
github-pages
|
||||||
|
integrations
|
||||||
|
plugins
|
||||||
|
skills
|
||||||
|
tests
|
||||||
97
.github/ISSUE_TEMPLATE/bug_report.yml
vendored
Normal file
97
.github/ISSUE_TEMPLATE/bug_report.yml
vendored
Normal file
|
|
@ -0,0 +1,97 @@
|
||||||
|
name: Bug report
|
||||||
|
description: Report reproducible incorrect or unexpected ReMe behavior
|
||||||
|
title: "[Bug]: "
|
||||||
|
labels: [bug]
|
||||||
|
body:
|
||||||
|
- type: markdown
|
||||||
|
attributes:
|
||||||
|
value: |
|
||||||
|
Thanks for helping improve ReMe. Please remove secrets, API keys, and private memory content before submitting.
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: description
|
||||||
|
attributes:
|
||||||
|
label: Description
|
||||||
|
description: What happened, and what did you expect instead?
|
||||||
|
placeholder: Describe the observed and expected behavior.
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: reproduce
|
||||||
|
attributes:
|
||||||
|
label: Steps to reproduce
|
||||||
|
description: Provide the smallest configuration and command sequence that reproduces the problem.
|
||||||
|
placeholder: |
|
||||||
|
1. Configure ...
|
||||||
|
2. Run ...
|
||||||
|
3. Observe ...
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: config
|
||||||
|
attributes:
|
||||||
|
label: Relevant configuration
|
||||||
|
description: Include only relevant values and redact credentials, tokens, endpoints, and private paths.
|
||||||
|
render: yaml
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: logs
|
||||||
|
attributes:
|
||||||
|
label: Logs or traceback
|
||||||
|
description: Paste relevant output after removing secrets and private workspace content.
|
||||||
|
render: shell
|
||||||
|
|
||||||
|
- type: input
|
||||||
|
id: reme-version
|
||||||
|
attributes:
|
||||||
|
label: ReMe version
|
||||||
|
placeholder: e.g. 0.4.1.8 or a commit SHA
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: input
|
||||||
|
id: python-version
|
||||||
|
attributes:
|
||||||
|
label: Python version
|
||||||
|
placeholder: e.g. 3.11.9
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: dropdown
|
||||||
|
id: os
|
||||||
|
attributes:
|
||||||
|
label: Operating system
|
||||||
|
options:
|
||||||
|
- Linux
|
||||||
|
- macOS
|
||||||
|
- Windows
|
||||||
|
- Other
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: dropdown
|
||||||
|
id: area
|
||||||
|
attributes:
|
||||||
|
label: Affected area
|
||||||
|
options:
|
||||||
|
- CLI or configuration
|
||||||
|
- HTTP, MCP, or local service
|
||||||
|
- Memory or workspace files
|
||||||
|
- Search, catalog, graph, or index
|
||||||
|
- Model or agent integration
|
||||||
|
- ReMe Studio
|
||||||
|
- Plugin or external integration
|
||||||
|
- Packaging or installation
|
||||||
|
- Other
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: checkboxes
|
||||||
|
id: safety
|
||||||
|
attributes:
|
||||||
|
label: Data safety
|
||||||
|
options:
|
||||||
|
- label: I removed credentials and private memory content from this report.
|
||||||
|
required: true
|
||||||
8
.github/ISSUE_TEMPLATE/config.yml
vendored
Normal file
8
.github/ISSUE_TEMPLATE/config.yml
vendored
Normal file
|
|
@ -0,0 +1,8 @@
|
||||||
|
blank_issues_enabled: false
|
||||||
|
contact_links:
|
||||||
|
- name: ReMe documentation
|
||||||
|
url: https://reme.agentscope.io
|
||||||
|
about: Read the installation, configuration, and usage guides.
|
||||||
|
- name: Existing issues
|
||||||
|
url: https://github.com/agentscope-ai/ReMe/issues
|
||||||
|
about: Search for existing reports and discussions before opening a new issue.
|
||||||
64
.github/ISSUE_TEMPLATE/feature_request.yml
vendored
Normal file
64
.github/ISSUE_TEMPLATE/feature_request.yml
vendored
Normal file
|
|
@ -0,0 +1,64 @@
|
||||||
|
name: Feature request
|
||||||
|
description: Propose a focused enhancement to ReMe
|
||||||
|
title: "[Feature]: "
|
||||||
|
labels: [enhancement]
|
||||||
|
body:
|
||||||
|
- type: textarea
|
||||||
|
id: problem
|
||||||
|
attributes:
|
||||||
|
label: Problem
|
||||||
|
description: What user problem or limitation should this change address?
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: proposal
|
||||||
|
attributes:
|
||||||
|
label: Proposed behavior
|
||||||
|
description: Describe the desired behavior and its user-visible contract.
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: dropdown
|
||||||
|
id: area
|
||||||
|
attributes:
|
||||||
|
label: Area
|
||||||
|
options:
|
||||||
|
- CLI or configuration
|
||||||
|
- Jobs or steps
|
||||||
|
- Memory or workspace files
|
||||||
|
- Search, catalog, graph, or index
|
||||||
|
- Service or client
|
||||||
|
- Model or agent integration
|
||||||
|
- ReMe Studio
|
||||||
|
- Plugin or external integration
|
||||||
|
- Documentation
|
||||||
|
- Other
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: ownership
|
||||||
|
attributes:
|
||||||
|
label: Local-first and compatibility considerations
|
||||||
|
description: Explain any effect on user-owned files, rebuildable state, configuration, schemas, or service interfaces.
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: alternatives
|
||||||
|
attributes:
|
||||||
|
label: Alternatives considered
|
||||||
|
description: Describe workarounds or alternative designs you considered.
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: examples
|
||||||
|
attributes:
|
||||||
|
label: Example usage
|
||||||
|
description: Show the proposed CLI, configuration, API, or UI behavior when useful.
|
||||||
|
render: shell
|
||||||
|
|
||||||
|
- type: checkboxes
|
||||||
|
id: contribution
|
||||||
|
attributes:
|
||||||
|
label: Contribution
|
||||||
|
options:
|
||||||
|
- label: I am willing to help implement or test this feature.
|
||||||
53
.github/ISSUE_TEMPLATE/question.yml
vendored
Normal file
53
.github/ISSUE_TEMPLATE/question.yml
vendored
Normal file
|
|
@ -0,0 +1,53 @@
|
||||||
|
name: Usage question
|
||||||
|
description: Ask for help using or configuring ReMe
|
||||||
|
title: "[Question]: "
|
||||||
|
labels: [question]
|
||||||
|
body:
|
||||||
|
- type: markdown
|
||||||
|
attributes:
|
||||||
|
value: Please check the documentation and existing issues before asking a new question.
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: goal
|
||||||
|
attributes:
|
||||||
|
label: What are you trying to achieve?
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: textarea
|
||||||
|
id: attempted
|
||||||
|
attributes:
|
||||||
|
label: What have you tried?
|
||||||
|
description: Include relevant commands or configuration, with secrets and private memory content removed.
|
||||||
|
validations:
|
||||||
|
required: true
|
||||||
|
|
||||||
|
- type: input
|
||||||
|
id: reme-version
|
||||||
|
attributes:
|
||||||
|
label: ReMe version
|
||||||
|
placeholder: e.g. 0.4.1.8 or a commit SHA
|
||||||
|
|
||||||
|
- type: dropdown
|
||||||
|
id: area
|
||||||
|
attributes:
|
||||||
|
label: Area
|
||||||
|
options:
|
||||||
|
- Installation
|
||||||
|
- Configuration
|
||||||
|
- CLI or service usage
|
||||||
|
- Memory and workspace management
|
||||||
|
- Search and retrieval
|
||||||
|
- ReMe Studio
|
||||||
|
- Plugin or integration
|
||||||
|
- Other
|
||||||
|
|
||||||
|
- type: checkboxes
|
||||||
|
id: checked
|
||||||
|
attributes:
|
||||||
|
label: Before submitting
|
||||||
|
options:
|
||||||
|
- label: I checked the [ReMe documentation](https://reme.agentscope.io) and searched existing issues.
|
||||||
|
required: true
|
||||||
|
- label: I removed credentials and private memory content.
|
||||||
|
required: true
|
||||||
35
.github/PULL_REQUEST_TEMPLATE.md
vendored
Normal file
35
.github/PULL_REQUEST_TEMPLATE.md
vendored
Normal file
|
|
@ -0,0 +1,35 @@
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
<!-- Explain the problem and the smallest coherent change that addresses it. -->
|
||||||
|
|
||||||
|
## Related issue
|
||||||
|
|
||||||
|
<!-- Use "Fixes #123" when applicable. -->
|
||||||
|
|
||||||
|
## Contract and data impact
|
||||||
|
|
||||||
|
- [ ] No public configuration, schema, CLI, endpoint, streaming, or workspace-layout contract changes
|
||||||
|
- [ ] No user-owned memory files are deleted or rewritten
|
||||||
|
- [ ] Derived indexes, catalogs, graphs, caches, and metadata remain rebuildable
|
||||||
|
|
||||||
|
<!-- If any item is unchecked, describe the impact and migration or recovery path. -->
|
||||||
|
|
||||||
|
## Validation
|
||||||
|
|
||||||
|
<!-- List the exact checks run and their results. Explain relevant checks that were not run. -->
|
||||||
|
|
||||||
|
- [ ] Focused tests pass
|
||||||
|
- [ ] Unit tests pass, or omitted tests are explained below
|
||||||
|
- [ ] `pre-commit run --all-files` passes, or omitted checks are explained below
|
||||||
|
- [ ] Frontend checks were run when `reme_studio/` changed
|
||||||
|
|
||||||
|
## Checklist
|
||||||
|
|
||||||
|
- [ ] I reviewed the diff for unrelated changes and sensitive data
|
||||||
|
- [ ] Tests cover intentional behavior changes
|
||||||
|
- [ ] Defaults, schemas, and concise documentation were updated together when required
|
||||||
|
- [ ] Long-lived clients, tasks, services, and executors follow the application lifecycle
|
||||||
|
|
||||||
|
## Screenshots or additional notes
|
||||||
|
|
||||||
|
<!-- Include UI screenshots, compatibility notes, or follow-up work when relevant. -->
|
||||||
35
.github/dependabot.yml
vendored
Normal file
35
.github/dependabot.yml
vendored
Normal file
|
|
@ -0,0 +1,35 @@
|
||||||
|
version: 2
|
||||||
|
updates:
|
||||||
|
- package-ecosystem: "github-actions"
|
||||||
|
directory: "/"
|
||||||
|
target-branch: "main"
|
||||||
|
schedule:
|
||||||
|
interval: "weekly"
|
||||||
|
day: "monday"
|
||||||
|
time: "09:30"
|
||||||
|
timezone: "Asia/Shanghai"
|
||||||
|
groups:
|
||||||
|
codeql:
|
||||||
|
patterns:
|
||||||
|
- "github/codeql-action/*"
|
||||||
|
open-pull-requests-limit: 5
|
||||||
|
commit-message:
|
||||||
|
prefix: "chore"
|
||||||
|
include: "scope"
|
||||||
|
|
||||||
|
- package-ecosystem: "pip"
|
||||||
|
directory: "/"
|
||||||
|
target-branch: "main"
|
||||||
|
schedule:
|
||||||
|
interval: "cron"
|
||||||
|
cronjob: "30 9 * * *"
|
||||||
|
timezone: "Asia/Shanghai"
|
||||||
|
allow:
|
||||||
|
- dependency-name: "agentscope"
|
||||||
|
cooldown:
|
||||||
|
exclude:
|
||||||
|
- "agentscope"
|
||||||
|
open-pull-requests-limit: 1
|
||||||
|
commit-message:
|
||||||
|
prefix: "chore"
|
||||||
|
include: "scope"
|
||||||
80
.github/workflows/_build-docs.yml
vendored
Normal file
80
.github/workflows/_build-docs.yml
vendored
Normal file
|
|
@ -0,0 +1,80 @@
|
||||||
|
name: _Build documentation
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_call:
|
||||||
|
inputs:
|
||||||
|
run_tests:
|
||||||
|
description: Run the documentation test suite before building
|
||||||
|
required: false
|
||||||
|
default: true
|
||||||
|
type: boolean
|
||||||
|
upload_pages_artifact:
|
||||||
|
description: Upload the build for a later GitHub Pages deployment job
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
name: Build documentation
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 20
|
||||||
|
defaults:
|
||||||
|
run:
|
||||||
|
working-directory: github-pages
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
fetch-depth: 0
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Set up Node
|
||||||
|
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: '22.22.3'
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: |
|
||||||
|
github-pages/package-lock.json
|
||||||
|
reme_studio/package-lock.json
|
||||||
|
|
||||||
|
- name: Install dependencies
|
||||||
|
run: npm ci
|
||||||
|
|
||||||
|
- name: Install Studio dependencies
|
||||||
|
run: npm ci --prefix ../reme_studio
|
||||||
|
|
||||||
|
- name: Check Studio types and lint
|
||||||
|
if: inputs.run_tests
|
||||||
|
working-directory: reme_studio
|
||||||
|
run: npm run format:check && npm run lint
|
||||||
|
|
||||||
|
- name: Test browser demo behavior
|
||||||
|
if: inputs.run_tests
|
||||||
|
working-directory: reme_studio
|
||||||
|
run: node --test tests/demo-workspace.test.mjs tests/wikilinks.test.mjs
|
||||||
|
|
||||||
|
- name: Run tests
|
||||||
|
if: inputs.run_tests
|
||||||
|
run: npm test
|
||||||
|
|
||||||
|
- name: Build documentation
|
||||||
|
run: npm run build
|
||||||
|
|
||||||
|
- name: Verify browser demo bundle
|
||||||
|
if: inputs.run_tests
|
||||||
|
working-directory: reme_studio
|
||||||
|
run: node --test tests/demo-build.test.mjs
|
||||||
|
|
||||||
|
- name: Configure Pages
|
||||||
|
if: inputs.upload_pages_artifact
|
||||||
|
uses: actions/configure-pages@45bfe0192ca1faeb007ade9deae92b16b8254a0d # v6
|
||||||
|
|
||||||
|
- name: Upload Pages artifact
|
||||||
|
if: inputs.upload_pages_artifact
|
||||||
|
uses: actions/upload-pages-artifact@fc324d3547104276b827a68afc52ff2a11cc49c9 # v5.0.0
|
||||||
|
with:
|
||||||
|
path: github-pages/dist
|
||||||
89
.github/workflows/_build-python-packages.yml
vendored
Normal file
89
.github/workflows/_build-python-packages.yml
vendored
Normal file
|
|
@ -0,0 +1,89 @@
|
||||||
|
name: _Build Python packages
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_call:
|
||||||
|
inputs:
|
||||||
|
expected_version:
|
||||||
|
description: Expected release version; omit for a consistency-only check
|
||||||
|
required: false
|
||||||
|
default: ''
|
||||||
|
type: string
|
||||||
|
upload_artifacts:
|
||||||
|
description: Upload distributions for later publish jobs
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
distributions:
|
||||||
|
name: Build Python distributions
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 45
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Set up Python
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
|
||||||
|
- name: Install build dependencies
|
||||||
|
run: |
|
||||||
|
python -m pip install --upgrade pip
|
||||||
|
python -m pip install build packaging pytest twine
|
||||||
|
|
||||||
|
- name: Validate package versions
|
||||||
|
if: inputs.expected_version == ''
|
||||||
|
run: python scripts/bump_version.py --check
|
||||||
|
|
||||||
|
- name: Validate release version
|
||||||
|
if: inputs.expected_version != ''
|
||||||
|
env:
|
||||||
|
EXPECTED_VERSION: ${{ inputs.expected_version }}
|
||||||
|
run: python scripts/bump_version.py --check --expected-version "${EXPECTED_VERSION}"
|
||||||
|
|
||||||
|
- name: Run package tests
|
||||||
|
run: PYTHONPATH=. python -m pytest tests/unit/test_package_versions.py -q
|
||||||
|
|
||||||
|
- name: Build and check distributions
|
||||||
|
run: |
|
||||||
|
mkdir -p dist/reme
|
||||||
|
python -m build --outdir dist/reme
|
||||||
|
python -m twine check dist/reme/*
|
||||||
|
|
||||||
|
- name: Verify distributions and isolated installation
|
||||||
|
run: |
|
||||||
|
REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
|
||||||
|
python -m zipfile -l "${REME_WHEEL}" | (! grep 'reme/web/')
|
||||||
|
python -m zipfile -l "${REME_WHEEL}" | (! grep 'reme_studio/')
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-package-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-package-smoke/bin/python" -m pip install "${REME_WHEEL}[as]"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-package-smoke/bin/python" -c "import reme"
|
||||||
|
|
||||||
|
- name: Verify released core dependencies
|
||||||
|
if: inputs.expected_version != ''
|
||||||
|
run: |
|
||||||
|
REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-core-package-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-core-package-smoke/bin/python" -m pip install "${REME_WHEEL}[core]"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-core-package-smoke/bin/python" - <<'PY'
|
||||||
|
from reme_studio import static_dir
|
||||||
|
|
||||||
|
assert (static_dir() / "index.html").is_file()
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Upload ReMe distributions
|
||||||
|
if: inputs.upload_artifacts
|
||||||
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
with:
|
||||||
|
name: reme-distributions
|
||||||
|
path: dist/reme/
|
||||||
|
if-no-files-found: error
|
||||||
155
.github/workflows/_release-npm-plugin.yml
vendored
Normal file
155
.github/workflows/_release-npm-plugin.yml
vendored
Normal file
|
|
@ -0,0 +1,155 @@
|
||||||
|
name: _Release npm plugin
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_call:
|
||||||
|
inputs:
|
||||||
|
directory:
|
||||||
|
description: Repository-relative package directory
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
package_name:
|
||||||
|
description: Exact public npm package name
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
artifact_name:
|
||||||
|
description: Prefix for the packed package artifact
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
version:
|
||||||
|
description: Exact package.json version; an optional v prefix is accepted
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
npm_tag:
|
||||||
|
description: npm distribution tag
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
use_npm_token:
|
||||||
|
description: Use the npm environment NPM_TOKEN instead of Trusted Publishing
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
validate_clawhub:
|
||||||
|
description: Validate the package against the ClawHub contract
|
||||||
|
required: false
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
outputs:
|
||||||
|
version:
|
||||||
|
description: Normalized package version
|
||||||
|
value: ${{ jobs.build.outputs.version }}
|
||||||
|
secrets:
|
||||||
|
NPM_TOKEN:
|
||||||
|
description: Optional bootstrap or recovery token for npm publishing
|
||||||
|
required: false
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 30
|
||||||
|
outputs:
|
||||||
|
version: ${{ steps.validate.outputs.version }}
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "24.16.0"
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: ${{ inputs.directory }}/package-lock.json
|
||||||
|
|
||||||
|
- name: Validate package identity and version
|
||||||
|
id: validate
|
||||||
|
working-directory: ${{ inputs.directory }}
|
||||||
|
env:
|
||||||
|
EXPECTED_NAME: ${{ inputs.package_name }}
|
||||||
|
RELEASE_VERSION: ${{ inputs.version }}
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
run: |
|
||||||
|
node --input-type=module <<'JS'
|
||||||
|
import { appendFileSync, readFileSync } from 'node:fs';
|
||||||
|
const manifest = JSON.parse(readFileSync('package.json', 'utf8'));
|
||||||
|
const expected = process.env.RELEASE_VERSION.replace(/^v/, '');
|
||||||
|
if (manifest.name !== process.env.EXPECTED_NAME) {
|
||||||
|
throw new Error(`Expected ${process.env.EXPECTED_NAME}, found ${manifest.name}`);
|
||||||
|
}
|
||||||
|
if (manifest.version !== expected) throw new Error(`package.json is ${manifest.version}, workflow input is ${expected}`);
|
||||||
|
if (manifest.version.includes('-') !== (process.env.NPM_TAG === 'next')) {
|
||||||
|
throw new Error('Prereleases must use next; stable releases must use latest');
|
||||||
|
}
|
||||||
|
appendFileSync(process.env.GITHUB_OUTPUT, `version=${manifest.version}\n`);
|
||||||
|
JS
|
||||||
|
|
||||||
|
- run: npm ci
|
||||||
|
working-directory: ${{ inputs.directory }}
|
||||||
|
- name: Validate package
|
||||||
|
working-directory: ${{ inputs.directory }}
|
||||||
|
run: |
|
||||||
|
npm run format:check
|
||||||
|
npm run lint
|
||||||
|
npm run typecheck
|
||||||
|
npm test
|
||||||
|
npm run test:package
|
||||||
|
- name: Validate ClawHub contract
|
||||||
|
if: inputs.validate_clawhub
|
||||||
|
working-directory: ${{ inputs.directory }}
|
||||||
|
run: npx --yes clawhub@0.23.3 package validate . --json
|
||||||
|
- name: Pack
|
||||||
|
working-directory: ${{ inputs.directory }}
|
||||||
|
run: |
|
||||||
|
mkdir -p "$RUNNER_TEMP/plugin-package"
|
||||||
|
npm pack --pack-destination "$RUNNER_TEMP/plugin-package"
|
||||||
|
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
with:
|
||||||
|
name: ${{ inputs.artifact_name }}-${{ steps.validate.outputs.version }}
|
||||||
|
path: ${{ runner.temp }}/plugin-package/*.tgz
|
||||||
|
if-no-files-found: error
|
||||||
|
|
||||||
|
publish:
|
||||||
|
if: github.ref == 'refs/heads/main'
|
||||||
|
needs: build
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: npm
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
steps:
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "24"
|
||||||
|
registry-url: https://registry.npmjs.org
|
||||||
|
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: ${{ inputs.artifact_name }}-${{ needs.build.outputs.version }}
|
||||||
|
path: dist/plugin
|
||||||
|
- name: Reject an existing package version
|
||||||
|
env:
|
||||||
|
PACKAGE_NAME: ${{ inputs.package_name }}
|
||||||
|
PACKAGE_VERSION: ${{ needs.build.outputs.version }}
|
||||||
|
run: |
|
||||||
|
if npm view "${PACKAGE_NAME}@${PACKAGE_VERSION}" version >/dev/null 2>&1; then
|
||||||
|
echo "${PACKAGE_NAME}@${PACKAGE_VERSION} already exists" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
- name: Publish to npm with Trusted Publishing
|
||||||
|
if: ${{ !inputs.use_npm_token }}
|
||||||
|
env:
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
run: npm publish dist/plugin/*.tgz --access public --tag "$NPM_TAG" --provenance
|
||||||
|
- name: Publish to npm with NPM_TOKEN
|
||||||
|
if: inputs.use_npm_token
|
||||||
|
env:
|
||||||
|
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
run: |
|
||||||
|
if [[ -z "${NODE_AUTH_TOKEN}" ]]; then
|
||||||
|
echo "NPM_TOKEN is required when use_npm_token is enabled" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
npm publish dist/plugin/*.tgz --access public --tag "$NPM_TAG" --provenance
|
||||||
40
.github/workflows/ci-docs.yml
vendored
Normal file
40
.github/workflows/ci-docs.yml
vendored
Normal file
|
|
@ -0,0 +1,40 @@
|
||||||
|
name: CI / Documentation
|
||||||
|
|
||||||
|
on:
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- '.github/workflows/ci-docs.yml'
|
||||||
|
- '.github/workflows/_build-docs.yml'
|
||||||
|
- 'AGENTS.md'
|
||||||
|
- 'README.md'
|
||||||
|
- 'README_ZH.md'
|
||||||
|
- 'docs/**'
|
||||||
|
- 'github-pages/**'
|
||||||
|
- 'reme/config/default.yaml'
|
||||||
|
- 'integrations/claude_code/README.md'
|
||||||
|
- 'integrations/hermes_agent/README.md'
|
||||||
|
- 'integrations/hermes_agent/figures/**'
|
||||||
|
- 'reme_studio/**'
|
||||||
|
- 'integrations/dsh/README*.md'
|
||||||
|
- 'integrations/dsh/figures/**'
|
||||||
|
- 'integrations/openclaw/README*.md'
|
||||||
|
- 'integrations/openclaw/figures/**'
|
||||||
|
- 'plugins/*/README*.md'
|
||||||
|
- 'benchmark/*/README*.md'
|
||||||
|
- 'benchmark/toolmemory/gitcha.png'
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
documentation:
|
||||||
|
name: Test and build documentation
|
||||||
|
uses: ./.github/workflows/_build-docs.yml
|
||||||
|
with:
|
||||||
|
run_tests: true
|
||||||
52
.github/workflows/ci-dsh-plugin.yml
vendored
Normal file
52
.github/workflows/ci-dsh-plugin.yml
vendored
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
name: CI / DSH plugin
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- ".github/workflows/ci-dsh-plugin.yml"
|
||||||
|
- ".github/workflows/_release-npm-plugin.yml"
|
||||||
|
- ".github/workflows/release-dsh-plugin.yml"
|
||||||
|
- "integrations/dsh/**"
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- ".github/workflows/ci-dsh-plugin.yml"
|
||||||
|
- ".github/workflows/_release-npm-plugin.yml"
|
||||||
|
- ".github/workflows/release-dsh-plugin.yml"
|
||||||
|
- "integrations/dsh/**"
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
package:
|
||||||
|
name: Validate DSH plugin
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 20
|
||||||
|
defaults:
|
||||||
|
run:
|
||||||
|
working-directory: integrations/dsh
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "24.16.0"
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: integrations/dsh/package-lock.json
|
||||||
|
|
||||||
|
- run: npm ci
|
||||||
|
- run: npm run format:check
|
||||||
|
- run: npm run lint
|
||||||
|
- run: npm run typecheck
|
||||||
|
- run: npm test
|
||||||
|
- run: npm run test:package
|
||||||
54
.github/workflows/ci-openclaw-plugin.yml
vendored
Normal file
54
.github/workflows/ci-openclaw-plugin.yml
vendored
Normal file
|
|
@ -0,0 +1,54 @@
|
||||||
|
name: CI / OpenClaw plugin
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- ".github/workflows/ci-openclaw-plugin.yml"
|
||||||
|
- ".github/workflows/_release-npm-plugin.yml"
|
||||||
|
- ".github/workflows/release-openclaw-plugin.yml"
|
||||||
|
- "integrations/openclaw/**"
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- ".github/workflows/ci-openclaw-plugin.yml"
|
||||||
|
- ".github/workflows/_release-npm-plugin.yml"
|
||||||
|
- ".github/workflows/release-openclaw-plugin.yml"
|
||||||
|
- "integrations/openclaw/**"
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
package:
|
||||||
|
name: Validate OpenClaw plugin
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 20
|
||||||
|
defaults:
|
||||||
|
run:
|
||||||
|
working-directory: integrations/openclaw
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "24.16.0"
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: integrations/openclaw/package-lock.json
|
||||||
|
|
||||||
|
- run: npm ci
|
||||||
|
- run: npm run format:check
|
||||||
|
- run: npm run lint
|
||||||
|
- run: npm run typecheck
|
||||||
|
- run: npm test
|
||||||
|
- run: npm run test:package
|
||||||
|
- name: Validate ClawHub contract
|
||||||
|
run: npx --yes clawhub@0.23.3 package validate . --json
|
||||||
40
.github/workflows/ci-packages.yml
vendored
Normal file
40
.github/workflows/ci-packages.yml
vendored
Normal file
|
|
@ -0,0 +1,40 @@
|
||||||
|
name: CI / Python packages
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- '.github/workflows/ci-packages.yml'
|
||||||
|
- '.github/workflows/_build-python-packages.yml'
|
||||||
|
- '.github/workflows/release-python.yml'
|
||||||
|
- 'pyproject.toml'
|
||||||
|
- 'README.md'
|
||||||
|
- 'reme/**'
|
||||||
|
- 'scripts/bump_version.py'
|
||||||
|
- 'tests/unit/test_package_versions.py'
|
||||||
|
- 'LICENSE'
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- '.github/workflows/ci-packages.yml'
|
||||||
|
- '.github/workflows/_build-python-packages.yml'
|
||||||
|
- '.github/workflows/release-python.yml'
|
||||||
|
- 'pyproject.toml'
|
||||||
|
- 'README.md'
|
||||||
|
- 'reme/**'
|
||||||
|
- 'scripts/bump_version.py'
|
||||||
|
- 'tests/unit/test_package_versions.py'
|
||||||
|
- 'LICENSE'
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
distributions:
|
||||||
|
name: Build and verify distributions
|
||||||
|
uses: ./.github/workflows/_build-python-packages.yml
|
||||||
65
.github/workflows/ci-python-quality.yml
vendored
Normal file
65
.github/workflows/ci-python-quality.yml
vendored
Normal file
|
|
@ -0,0 +1,65 @@
|
||||||
|
name: CI / Python quality
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
actionlint:
|
||||||
|
name: GitHub Actions
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Validate workflows with actionlint
|
||||||
|
env:
|
||||||
|
ACTIONLINT_VERSION: 1.7.12
|
||||||
|
ACTIONLINT_SHA256: 8aca8db96f1b94770f1b0d72b6dddcb1ebb8123cb3712530b08cc387b349a3d8
|
||||||
|
run: |
|
||||||
|
archive="actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz"
|
||||||
|
curl --fail --location --proto '=https' --retry 3 --silent --show-error \
|
||||||
|
--output "${RUNNER_TEMP}/${archive}" \
|
||||||
|
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/${archive}"
|
||||||
|
echo "${ACTIONLINT_SHA256} ${RUNNER_TEMP}/${archive}" | sha256sum --check
|
||||||
|
tar -xzf "${RUNNER_TEMP}/${archive}" -C "${RUNNER_TEMP}" actionlint
|
||||||
|
"${RUNNER_TEMP}/actionlint" .github/workflows/*.yml
|
||||||
|
|
||||||
|
pre-commit:
|
||||||
|
name: Pre-commit
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 30
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Setup Python
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
cache: pip
|
||||||
|
|
||||||
|
- name: Update setuptools
|
||||||
|
run: |
|
||||||
|
pip install -U setuptools wheel
|
||||||
|
|
||||||
|
- name: Install
|
||||||
|
run: |
|
||||||
|
pip install -q -e reme_studio -e ".[dev,core]"
|
||||||
|
pip install -q --no-deps -e plugins/auto-fin -e plugins/daily_paper -e plugins/lme -e plugins/beam
|
||||||
|
|
||||||
|
- name: Pre-commit starts
|
||||||
|
run: pre-commit run --all-files
|
||||||
|
|
@ -1,30 +1,36 @@
|
||||||
name: Tests ReMe
|
name: CI / Python tests
|
||||||
|
|
||||||
on:
|
on:
|
||||||
push:
|
push:
|
||||||
branches: [main, master, dev, develop]
|
branches: [main]
|
||||||
pull_request:
|
pull_request:
|
||||||
branches: [main, master, dev, develop]
|
branches: [main]
|
||||||
workflow_dispatch:
|
workflow_dispatch:
|
||||||
|
|
||||||
concurrency:
|
concurrency:
|
||||||
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
cancel-in-progress: true
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
unit-tests:
|
unit-tests:
|
||||||
name: Unit Tests - py${{ matrix.python-version }}
|
name: Unit Tests - py${{ matrix.python-version }}
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 90
|
||||||
strategy:
|
strategy:
|
||||||
fail-fast: false
|
fail-fast: false
|
||||||
matrix:
|
matrix:
|
||||||
python-version: ["3.11", "3.12", "3.13"]
|
python-version: ["3.11", "3.12", "3.13", "3.14"]
|
||||||
|
|
||||||
steps:
|
steps:
|
||||||
- uses: actions/checkout@v4
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
- name: Set up Python ${{ matrix.python-version }}
|
- name: Set up Python ${{ matrix.python-version }}
|
||||||
uses: actions/setup-python@v5
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
with:
|
with:
|
||||||
python-version: ${{ matrix.python-version }}
|
python-version: ${{ matrix.python-version }}
|
||||||
cache: 'pip'
|
cache: 'pip'
|
||||||
|
|
@ -32,12 +38,15 @@ jobs:
|
||||||
- name: Install dependencies
|
- name: Install dependencies
|
||||||
run: |
|
run: |
|
||||||
python -m pip install --upgrade pip setuptools wheel
|
python -m pip install --upgrade pip setuptools wheel
|
||||||
pip install -e ".[dev,core]"
|
pip install -e reme_studio -e ".[dev,core]"
|
||||||
|
pip install --no-deps -e plugins/auto-fin
|
||||||
|
pip install -e plugins/daily_paper
|
||||||
|
pip install -e plugins/lme -e plugins/beam
|
||||||
pip install coverage
|
pip install coverage
|
||||||
|
|
||||||
- name: Run unit tests
|
- name: Run unit tests
|
||||||
run: |
|
run: |
|
||||||
coverage run -m pytest tests/unit \
|
coverage run -m pytest tests/unit plugins/auto-fin plugins/daily_paper plugins/lme plugins/beam \
|
||||||
-v \
|
-v \
|
||||||
--tb=long \
|
--tb=long \
|
||||||
-s \
|
-s \
|
||||||
93
.github/workflows/ci-reme-studio.yml
vendored
Normal file
93
.github/workflows/ci-reme-studio.yml
vendored
Normal file
|
|
@ -0,0 +1,93 @@
|
||||||
|
name: CI / ReMe Studio
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- "reme_studio/**"
|
||||||
|
- ".github/workflows/ci-reme-studio.yml"
|
||||||
|
- ".github/workflows/release-reme-studio.yml"
|
||||||
|
- "scripts/package_studio.py"
|
||||||
|
- "tests/unit/test_package_versions.py"
|
||||||
|
- "pyproject.toml"
|
||||||
|
- "LICENSE"
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- "reme_studio/**"
|
||||||
|
- ".github/workflows/ci-reme-studio.yml"
|
||||||
|
- ".github/workflows/release-reme-studio.yml"
|
||||||
|
- "scripts/package_studio.py"
|
||||||
|
- "tests/unit/test_package_versions.py"
|
||||||
|
- "pyproject.toml"
|
||||||
|
- "LICENSE"
|
||||||
|
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
studio:
|
||||||
|
name: Studio checks
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 30
|
||||||
|
defaults:
|
||||||
|
run:
|
||||||
|
working-directory: reme_studio
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Setup Node
|
||||||
|
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "22.22.3"
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: reme_studio/package-lock.json
|
||||||
|
|
||||||
|
- name: Install dependencies
|
||||||
|
run: npm ci
|
||||||
|
|
||||||
|
- name: Run format check
|
||||||
|
run: npm run format:check
|
||||||
|
|
||||||
|
- name: Run lint
|
||||||
|
run: npm run lint
|
||||||
|
|
||||||
|
- name: Run tests
|
||||||
|
run: npm test
|
||||||
|
|
||||||
|
- name: Verify npm package
|
||||||
|
run: |
|
||||||
|
npm pack --pack-destination "${RUNNER_TEMP}"
|
||||||
|
tar -tzf "${RUNNER_TEMP}"/agentscope-ai-reme_studio-*.tgz | grep '^package/dist-static/index.html$'
|
||||||
|
|
||||||
|
- name: Set up Python
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: "3.11"
|
||||||
|
|
||||||
|
- name: Build and verify Python package
|
||||||
|
working-directory: .
|
||||||
|
run: |
|
||||||
|
python -m pip install build packaging pytest twine
|
||||||
|
PYTHONPATH=. python -m pytest tests/unit/test_package_versions.py -q
|
||||||
|
python scripts/package_studio.py
|
||||||
|
python -m build reme_studio --outdir dist/studio
|
||||||
|
python -m twine check dist/studio/*
|
||||||
|
STUDIO_WHEEL="$(pwd)/$(ls dist/studio/reme_studio-*.whl)"
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-studio-package-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" -m pip install "${STUDIO_WHEEL}"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" - <<'PY'
|
||||||
|
from reme_studio import static_dir
|
||||||
|
|
||||||
|
assert (static_dir() / "index.html").is_file()
|
||||||
|
PY
|
||||||
52
.github/workflows/ci-windows.yml
vendored
Normal file
52
.github/workflows/ci-windows.yml
vendored
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
name: CI / Windows
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
cli-smoke:
|
||||||
|
name: CLI smoke - py${{ matrix.python-version }}
|
||||||
|
runs-on: windows-latest
|
||||||
|
timeout-minutes: 30
|
||||||
|
strategy:
|
||||||
|
fail-fast: false
|
||||||
|
matrix:
|
||||||
|
python-version: ["3.11"]
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Set up Python ${{ matrix.python-version }}
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: ${{ matrix.python-version }}
|
||||||
|
cache: 'pip'
|
||||||
|
|
||||||
|
- name: Install package
|
||||||
|
run: |
|
||||||
|
python -m pip install --upgrade pip setuptools wheel
|
||||||
|
pip install -e ".[dev,as]"
|
||||||
|
|
||||||
|
- name: Run version job
|
||||||
|
run: reme start config=tests/fixtures/config/version-smoke.yaml job=version
|
||||||
|
|
||||||
|
- name: Run Windows path tests
|
||||||
|
run: |
|
||||||
|
python -m pytest `
|
||||||
|
tests/unit/test_auto_dream.py::test_scan_day_files_includes_only_markdown_day_files `
|
||||||
|
tests/unit/test_auto_dream.py::test_dream_extract_matches_posix_catalog_paths `
|
||||||
|
tests/unit/test_read_with_neighbors.py::test_read_with_neighbors_uses_posix_nested_path `
|
||||||
|
-v
|
||||||
60
.github/workflows/deploy-docs.yml
vendored
Normal file
60
.github/workflows/deploy-docs.yml
vendored
Normal file
|
|
@ -0,0 +1,60 @@
|
||||||
|
name: Deploy / Documentation
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- "github-pages/**"
|
||||||
|
- "docs/**"
|
||||||
|
- "reme/config/default.yaml"
|
||||||
|
- "integrations/claude_code/README.md"
|
||||||
|
- "integrations/hermes_agent/README.md"
|
||||||
|
- "integrations/hermes_agent/figures/**"
|
||||||
|
- "README.md"
|
||||||
|
- "README_ZH.md"
|
||||||
|
- "reme_studio/**"
|
||||||
|
- "integrations/dsh/README*.md"
|
||||||
|
- "integrations/dsh/figures/**"
|
||||||
|
- "integrations/openclaw/README*.md"
|
||||||
|
- "integrations/openclaw/figures/**"
|
||||||
|
- "plugins/*/README*.md"
|
||||||
|
- "benchmark/*/README*.md"
|
||||||
|
- "benchmark/toolmemory/gitcha.png"
|
||||||
|
- "AGENTS.md"
|
||||||
|
- ".github/workflows/deploy-docs.yml"
|
||||||
|
- ".github/workflows/_build-docs.yml"
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: pages
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
name: Build documentation
|
||||||
|
uses: ./.github/workflows/_build-docs.yml
|
||||||
|
with:
|
||||||
|
run_tests: true
|
||||||
|
upload_pages_artifact: true
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
pages: write
|
||||||
|
id-token: write
|
||||||
|
|
||||||
|
deploy:
|
||||||
|
environment:
|
||||||
|
name: github-pages
|
||||||
|
url: ${{ steps.deployment.outputs.page_url }}
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
needs: build
|
||||||
|
timeout-minutes: 10
|
||||||
|
permissions:
|
||||||
|
pages: write
|
||||||
|
id-token: write
|
||||||
|
steps:
|
||||||
|
- name: Deploy
|
||||||
|
id: deployment
|
||||||
|
uses: actions/deploy-pages@368f82528645a54fb793d4d04e342629a3f51346 # v5.0.1
|
||||||
221
.github/workflows/docker.yml
vendored
Normal file
221
.github/workflows/docker.yml
vendored
Normal file
|
|
@ -0,0 +1,221 @@
|
||||||
|
name: CI and Release / Docker
|
||||||
|
|
||||||
|
on:
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- '.github/workflows/docker.yml'
|
||||||
|
- 'Dockerfile'
|
||||||
|
- '.dockerignore'
|
||||||
|
- 'docker-compose.yml'
|
||||||
|
- 'deploy/docker/**'
|
||||||
|
- 'pyproject.toml'
|
||||||
|
- 'README.md'
|
||||||
|
- 'LICENSE'
|
||||||
|
- 'reme/**'
|
||||||
|
- 'reme_studio/**'
|
||||||
|
- 'scripts/package_studio.py'
|
||||||
|
- 'scripts/test_docker_image.py'
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
paths:
|
||||||
|
- '.github/workflows/docker.yml'
|
||||||
|
- 'Dockerfile'
|
||||||
|
- '.dockerignore'
|
||||||
|
- 'docker-compose.yml'
|
||||||
|
- 'deploy/docker/**'
|
||||||
|
- 'pyproject.toml'
|
||||||
|
- 'README.md'
|
||||||
|
- 'LICENSE'
|
||||||
|
- 'reme/**'
|
||||||
|
- 'reme_studio/**'
|
||||||
|
- 'scripts/package_studio.py'
|
||||||
|
- 'scripts/test_docker_image.py'
|
||||||
|
release:
|
||||||
|
types: [published]
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
||||||
|
|
||||||
|
env:
|
||||||
|
PUBLISH_IMAGE: ${{ github.event_name == 'release' || (github.event_name != 'pull_request' && github.ref == 'refs/heads/main') }}
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
name: Build and test / ${{ matrix.arch }}
|
||||||
|
runs-on: ${{ matrix.runner }}
|
||||||
|
timeout-minutes: 60
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
packages: write
|
||||||
|
strategy:
|
||||||
|
fail-fast: false
|
||||||
|
matrix:
|
||||||
|
include:
|
||||||
|
- arch: amd64
|
||||||
|
runner: ubuntu-24.04
|
||||||
|
- arch: arm64
|
||||||
|
runner: ubuntu-24.04-arm
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
|
||||||
|
- name: Validate package version
|
||||||
|
env:
|
||||||
|
RELEASE_VERSION: ${{ github.event.release.tag_name }}
|
||||||
|
run: |
|
||||||
|
python -m pip install packaging
|
||||||
|
if [ -n "$RELEASE_VERSION" ]; then
|
||||||
|
python scripts/bump_version.py --check --expected-version "$RELEASE_VERSION"
|
||||||
|
else
|
||||||
|
python scripts/bump_version.py --check
|
||||||
|
fi
|
||||||
|
|
||||||
|
- name: Normalize image name
|
||||||
|
id: image
|
||||||
|
env:
|
||||||
|
REPOSITORY: ${{ github.repository }}
|
||||||
|
run: echo "name=ghcr.io/${REPOSITORY,,}" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
|
- name: Validate Compose
|
||||||
|
run: docker compose config --quiet
|
||||||
|
|
||||||
|
- uses: docker/setup-buildx-action@f87e5991a6d7451dcb8d9637bfbc97413f497069 # v4
|
||||||
|
|
||||||
|
- uses: docker/metadata-action@dc802804100637a589fabce1cb79ff13a1411302 # v6
|
||||||
|
id: metadata
|
||||||
|
with:
|
||||||
|
images: ${{ steps.image.outputs.name }}
|
||||||
|
flavor: latest=${{ github.event_name == 'release' && !github.event.release.prerelease && 'auto' || 'false' }}
|
||||||
|
tags: |
|
||||||
|
type=raw,value=main,enable=${{ github.ref == 'refs/heads/main' }}
|
||||||
|
type=pep440,pattern={{version}},value=${{ github.event.release.tag_name }},enable=${{ github.event_name == 'release' }}
|
||||||
|
type=sha
|
||||||
|
|
||||||
|
- name: Build local image
|
||||||
|
uses: docker/build-push-action@c3c9e263c25d99ce0380d002d59b67737d91b0dc # v7
|
||||||
|
with:
|
||||||
|
context: .
|
||||||
|
platforms: linux/${{ matrix.arch }}
|
||||||
|
load: true
|
||||||
|
tags: reme:smoke
|
||||||
|
labels: ${{ steps.metadata.outputs.labels }}
|
||||||
|
cache-from: type=gha,scope=reme-${{ matrix.arch }}
|
||||||
|
cache-to: type=gha,mode=max,scope=reme-${{ matrix.arch }}
|
||||||
|
|
||||||
|
- name: Test installed image and persistent workspace
|
||||||
|
run: python scripts/test_docker_image.py --image reme:smoke
|
||||||
|
|
||||||
|
- name: Log in to GHCR
|
||||||
|
if: env.PUBLISH_IMAGE == 'true'
|
||||||
|
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
|
||||||
|
with:
|
||||||
|
registry: ghcr.io
|
||||||
|
username: ${{ github.actor }}
|
||||||
|
password: ${{ secrets.GITHUB_TOKEN }}
|
||||||
|
|
||||||
|
- name: Publish tested architecture by digest
|
||||||
|
id: publish
|
||||||
|
if: env.PUBLISH_IMAGE == 'true'
|
||||||
|
uses: docker/build-push-action@c3c9e263c25d99ce0380d002d59b67737d91b0dc # v7
|
||||||
|
with:
|
||||||
|
context: .
|
||||||
|
platforms: linux/${{ matrix.arch }}
|
||||||
|
outputs: type=image,name=${{ steps.image.outputs.name }},push-by-digest=true,name-canonical=true,push=true
|
||||||
|
labels: ${{ steps.metadata.outputs.labels }}
|
||||||
|
cache-from: type=gha,scope=reme-${{ matrix.arch }}
|
||||||
|
provenance: mode=max
|
||||||
|
sbom: true
|
||||||
|
|
||||||
|
- name: Record image digest
|
||||||
|
if: env.PUBLISH_IMAGE == 'true'
|
||||||
|
env:
|
||||||
|
IMAGE_DIGEST: ${{ steps.publish.outputs.digest }}
|
||||||
|
run: |
|
||||||
|
mkdir -p "$RUNNER_TEMP/digests"
|
||||||
|
touch "$RUNNER_TEMP/digests/${IMAGE_DIGEST#sha256:}"
|
||||||
|
|
||||||
|
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
if: env.PUBLISH_IMAGE == 'true'
|
||||||
|
with:
|
||||||
|
name: docker-digest-${{ matrix.arch }}
|
||||||
|
path: ${{ runner.temp }}/digests/*
|
||||||
|
if-no-files-found: error
|
||||||
|
retention-days: 1
|
||||||
|
|
||||||
|
manifest:
|
||||||
|
name: Publish multi-platform tags
|
||||||
|
needs: build
|
||||||
|
if: github.event_name == 'release' || (github.event_name != 'pull_request' && github.ref == 'refs/heads/main')
|
||||||
|
runs-on: ubuntu-24.04
|
||||||
|
timeout-minutes: 15
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
packages: write
|
||||||
|
steps:
|
||||||
|
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
pattern: docker-digest-*
|
||||||
|
merge-multiple: true
|
||||||
|
path: ${{ runner.temp }}/digests
|
||||||
|
|
||||||
|
- name: Normalize image name
|
||||||
|
id: image
|
||||||
|
env:
|
||||||
|
REPOSITORY: ${{ github.repository }}
|
||||||
|
run: echo "name=ghcr.io/${REPOSITORY,,}" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
|
- uses: docker/setup-buildx-action@f87e5991a6d7451dcb8d9637bfbc97413f497069 # v4
|
||||||
|
|
||||||
|
- uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
|
||||||
|
with:
|
||||||
|
registry: ghcr.io
|
||||||
|
username: ${{ github.actor }}
|
||||||
|
password: ${{ secrets.GITHUB_TOKEN }}
|
||||||
|
|
||||||
|
- uses: docker/metadata-action@dc802804100637a589fabce1cb79ff13a1411302 # v6
|
||||||
|
id: metadata
|
||||||
|
with:
|
||||||
|
images: ${{ steps.image.outputs.name }}
|
||||||
|
flavor: latest=${{ github.event_name == 'release' && !github.event.release.prerelease && 'auto' || 'false' }}
|
||||||
|
tags: |
|
||||||
|
type=raw,value=main,enable=${{ github.ref == 'refs/heads/main' }}
|
||||||
|
type=pep440,pattern={{version}},value=${{ github.event.release.tag_name }},enable=${{ github.event_name == 'release' }}
|
||||||
|
type=sha
|
||||||
|
|
||||||
|
- name: Assemble and verify tags
|
||||||
|
env:
|
||||||
|
IMAGE_NAME: ${{ steps.image.outputs.name }}
|
||||||
|
IMAGE_TAGS: ${{ steps.metadata.outputs.tags }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
image_refs=()
|
||||||
|
for digest_file in "$RUNNER_TEMP"/digests/*; do
|
||||||
|
image_refs+=("${IMAGE_NAME}@sha256:$(basename "$digest_file")")
|
||||||
|
done
|
||||||
|
[ "${#image_refs[@]}" -eq 2 ]
|
||||||
|
tag_args=()
|
||||||
|
while IFS= read -r tag; do
|
||||||
|
tag_args+=(--tag "$tag")
|
||||||
|
done <<< "$IMAGE_TAGS"
|
||||||
|
docker buildx imagetools create "${tag_args[@]}" "${image_refs[@]}"
|
||||||
|
while IFS= read -r tag; do
|
||||||
|
manifest=$(docker buildx imagetools inspect "$tag" --raw)
|
||||||
|
IMAGE_MANIFEST="$manifest" python - <<'PY'
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
manifest = json.loads(os.environ["IMAGE_MANIFEST"])
|
||||||
|
platforms = {(item["platform"]["os"], item["platform"]["architecture"]) for item in manifest["manifests"]}
|
||||||
|
assert {("linux", "amd64"), ("linux", "arm64")} <= platforms, platforms
|
||||||
|
PY
|
||||||
|
done <<< "$IMAGE_TAGS"
|
||||||
|
|
@ -1,16 +1,21 @@
|
||||||
name: PR Title Check
|
name: Policy / PR title
|
||||||
|
|
||||||
on:
|
on:
|
||||||
pull_request:
|
pull_request:
|
||||||
branches: [main, master, dev, develop]
|
branches: [main]
|
||||||
types: [opened, edited, synchronize, reopened]
|
types: [opened, edited, reopened]
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
pull-requests: read
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
check-pr-title:
|
check-pr-title:
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
steps:
|
steps:
|
||||||
- name: Check PR title format
|
- name: Check PR title format
|
||||||
uses: amannn/action-semantic-pull-request@v6.1.1
|
uses: amannn/action-semantic-pull-request@48f256284bd46cdaab1048c3721360e808335d50 # v6.1.1
|
||||||
env:
|
env:
|
||||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||||
with:
|
with:
|
||||||
38
.github/workflows/pre-commit.yml
vendored
38
.github/workflows/pre-commit.yml
vendored
|
|
@ -1,38 +0,0 @@
|
||||||
name: Pre-commit
|
|
||||||
|
|
||||||
on: [ push, pull_request ]
|
|
||||||
|
|
||||||
jobs:
|
|
||||||
run:
|
|
||||||
runs-on: ${{ matrix.os }}
|
|
||||||
strategy:
|
|
||||||
fail-fast: True
|
|
||||||
matrix:
|
|
||||||
os: [ ubuntu-latest ]
|
|
||||||
env:
|
|
||||||
OS: ${{ matrix.os }}
|
|
||||||
PYTHON: '3.11'
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@v4
|
|
||||||
- name: Setup Python
|
|
||||||
uses: actions/setup-python@v5
|
|
||||||
with:
|
|
||||||
python-version: '3.11'
|
|
||||||
- name: Update setuptools
|
|
||||||
run: |
|
|
||||||
pip install -U setuptools wheel
|
|
||||||
- name: Install
|
|
||||||
run: |
|
|
||||||
pip install -q -e ".[dev,core]"
|
|
||||||
- name: Install pre-commit
|
|
||||||
run: |
|
|
||||||
pre-commit install
|
|
||||||
- name: Pre-commit starts
|
|
||||||
run: |
|
|
||||||
pre-commit run --all-files > pre-commit.log 2>&1 || true
|
|
||||||
cat pre-commit.log
|
|
||||||
if grep -q Failed pre-commit.log; then
|
|
||||||
echo -e "\e[41m [**FAIL**] Please install pre-commit and format your code first. \e[0m"
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
echo -e "\e[46m ********************************Passed******************************** \e[0m"
|
|
||||||
45
.github/workflows/python-publish.yml
vendored
45
.github/workflows/python-publish.yml
vendored
|
|
@ -1,45 +0,0 @@
|
||||||
# This workflow will upload a Python Package using Twine when a release is created
|
|
||||||
# For more information see: https://docs.github.com/en/actions/automating-builds-and-tests/building-and-testing-python#publishing-to-package-registries
|
|
||||||
|
|
||||||
# This workflow uses actions that are not certified by GitHub.
|
|
||||||
# They are provided by a third-party and are governed by
|
|
||||||
# separate terms of service, privacy policy, and support
|
|
||||||
# documentation.
|
|
||||||
|
|
||||||
name: Publish Python Package to Pypi
|
|
||||||
|
|
||||||
on:
|
|
||||||
workflow_dispatch:
|
|
||||||
release:
|
|
||||||
types: [published]
|
|
||||||
|
|
||||||
permissions:
|
|
||||||
contents: read
|
|
||||||
|
|
||||||
jobs:
|
|
||||||
deploy:
|
|
||||||
|
|
||||||
runs-on: ubuntu-latest
|
|
||||||
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@v6
|
|
||||||
- name: Set up Python
|
|
||||||
uses: actions/setup-python@v6
|
|
||||||
with:
|
|
||||||
python-version: '3.11'
|
|
||||||
- name: Install dependencies
|
|
||||||
run: |
|
|
||||||
python -m pip install --upgrade pip
|
|
||||||
pip install setuptools wheel build
|
|
||||||
- name: Build package
|
|
||||||
run: python -m build
|
|
||||||
- name: Test installation
|
|
||||||
run: |
|
|
||||||
WHEEL="$(ls dist/*.whl)"
|
|
||||||
pip install "${WHEEL}[core]"
|
|
||||||
python -c "import reme; print(reme.__version__)"
|
|
||||||
- name: Publish package to PyPI
|
|
||||||
uses: pypa/gh-action-pypi-publish@release/v1
|
|
||||||
with:
|
|
||||||
user: __token__
|
|
||||||
password: ${{ secrets.PYPI_API_TOKEN }}
|
|
||||||
178
.github/workflows/release-auto-fin.yml
vendored
Normal file
178
.github/workflows/release-auto-fin.yml
vendored
Normal file
|
|
@ -0,0 +1,178 @@
|
||||||
|
# 发布操作手册:
|
||||||
|
# 1. 先将 plugins/auto-fin/pyproject.toml 中的 project.version 更新为待发布版本并合入目标分支。
|
||||||
|
# 2. 确认插件依赖的 reme-ai 版本已经发布到 PyPI;本工作流会在构建阶段验证该依赖可下载。
|
||||||
|
# 3. 确认 PyPI Trusted Publisher 已绑定本仓库、此工作流和 pypi environment,且 PyPI 上不存在相同版本。
|
||||||
|
# 4. 在 GitHub 仓库的 Actions 页面选择“Release / Auto Fin plugin”,点击“Run workflow”。
|
||||||
|
# 5. 输入与 project.version 完全一致的版本号(例如 0.1.0)后运行;版本也可以带 v 前缀。
|
||||||
|
#
|
||||||
|
# 推荐发布顺序:reme-ai -> reme-auto-fin -> QwenPaw 更新依赖并通过 plugins: [auto-fin] 启用。
|
||||||
|
# 当前仅支持 workflow_dispatch 手动触发,不会因 push、tag 或 release 自动发布。
|
||||||
|
|
||||||
|
name: Release / Auto Fin plugin
|
||||||
|
|
||||||
|
run-name: Publish reme-auto-fin ${{ inputs.version }}
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Version from plugins/auto-fin/pyproject.toml (for example, 0.1.0)
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-auto-fin
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 45
|
||||||
|
env:
|
||||||
|
RELEASE_VERSION: ${{ inputs.version }}
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Set up Python
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
|
||||||
|
- name: Install test and build dependencies
|
||||||
|
run: |
|
||||||
|
python -m pip install --upgrade pip
|
||||||
|
python -m pip install build packaging pytest pytest-asyncio twine
|
||||||
|
python -m pip install -e ".[core]"
|
||||||
|
python -m pip install --no-deps -e plugins/auto-fin
|
||||||
|
|
||||||
|
- name: Validate package name and release version
|
||||||
|
id: package
|
||||||
|
run: |
|
||||||
|
python - "${RELEASE_VERSION}" <<'PY'
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import tomllib
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from packaging.requirements import Requirement
|
||||||
|
from packaging.version import Version
|
||||||
|
|
||||||
|
project = tomllib.loads(Path("plugins/auto-fin/pyproject.toml").read_text(encoding="utf-8"))["project"]
|
||||||
|
expected = Version(sys.argv[1].removeprefix("v"))
|
||||||
|
actual = Version(project["version"])
|
||||||
|
if project["name"] != "reme-auto-fin":
|
||||||
|
raise SystemExit(f"Expected project name 'reme-auto-fin', found {project['name']!r}")
|
||||||
|
if actual != expected:
|
||||||
|
raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
|
||||||
|
requirements = [requirement for requirement in project["dependencies"] if requirement.startswith("reme-ai")]
|
||||||
|
if len(requirements) != 1:
|
||||||
|
raise SystemExit(f"Expected one reme-ai dependency, found {requirements!r}")
|
||||||
|
reme_requirement = Requirement(requirements[0])
|
||||||
|
if reme_requirement.name != "reme-ai" or reme_requirement.extras:
|
||||||
|
raise SystemExit(f"Expected a base reme-ai dependency, found {requirements[0]!r}")
|
||||||
|
if Version("0.4.1.11") in reme_requirement.specifier or Version("0.4.1.12") not in reme_requirement.specifier:
|
||||||
|
raise SystemExit(f"Expected reme-ai>=0.4.1.12, found {requirements[0]!r}")
|
||||||
|
root_project = tomllib.loads(Path("pyproject.toml").read_text(encoding="utf-8"))["project"]
|
||||||
|
agentscope_requirements = [
|
||||||
|
Requirement(value) for value in root_project["optional-dependencies"]["as"]
|
||||||
|
]
|
||||||
|
if len(agentscope_requirements) != 1 or agentscope_requirements[0].name != "agentscope":
|
||||||
|
raise SystemExit(f"Expected one agentscope dependency, found {agentscope_requirements!r}")
|
||||||
|
agentscope_requirement = agentscope_requirements[0]
|
||||||
|
agentscope_specifiers = list(agentscope_requirement.specifier)
|
||||||
|
if (
|
||||||
|
agentscope_requirement.extras != {"model-ollama"}
|
||||||
|
or agentscope_requirement.marker is not None
|
||||||
|
or len(agentscope_specifiers) != 1
|
||||||
|
or agentscope_specifiers[0].operator != "=="
|
||||||
|
or Version(agentscope_specifiers[0].version).is_prerelease
|
||||||
|
):
|
||||||
|
raise SystemExit(f"Expected a stable exact agentscope[model-ollama] pin, found {agentscope_requirement}")
|
||||||
|
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
|
||||||
|
print(f"reme_requirement={reme_requirement}", file=output)
|
||||||
|
print(f"agentscope_requirement={agentscope_requirement}", file=output)
|
||||||
|
print(f"Publishing {project['name']} {actual}")
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Run Auto Fin tests
|
||||||
|
run: python -m pytest plugins/auto-fin -q
|
||||||
|
|
||||||
|
- name: Require the plugin-enabled ReMe release on PyPI
|
||||||
|
env:
|
||||||
|
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
|
||||||
|
run: |
|
||||||
|
python -m pip download --no-deps \
|
||||||
|
--dest "${RUNNER_TEMP}/reme-auto-fin-base" \
|
||||||
|
"${REME_REQUIREMENT}"
|
||||||
|
|
||||||
|
- name: Build and check distributions
|
||||||
|
run: |
|
||||||
|
mkdir -p dist/auto-fin
|
||||||
|
python -m build plugins/auto-fin --outdir dist/auto-fin
|
||||||
|
python -m twine check dist/auto-fin/*
|
||||||
|
|
||||||
|
- name: Verify distributions and isolated installation
|
||||||
|
env:
|
||||||
|
AGENTSCOPE_REQUIREMENT: ${{ steps.package.outputs.agentscope_requirement }}
|
||||||
|
run: |
|
||||||
|
AUTO_FIN_WHEEL="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.whl)"
|
||||||
|
AUTO_FIN_SDIST="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.tar.gz)"
|
||||||
|
python -m zipfile -l "${AUTO_FIN_WHEEL}" | grep 'dist-info/licenses/LICENSE'
|
||||||
|
python -m tarfile -l "${AUTO_FIN_SDIST}" | grep '/LICENSE'
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-auto-fin-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" -m pip install \
|
||||||
|
"${AGENTSCOPE_REQUIREMENT}" "${AUTO_FIN_WHEEL}"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" - <<'PY'
|
||||||
|
from importlib.metadata import distribution
|
||||||
|
|
||||||
|
from reme.plugin_manifest import load_package_manifest
|
||||||
|
|
||||||
|
package = distribution("reme-auto-fin")
|
||||||
|
plugins = {entry.name: entry for entry in package.entry_points if entry.group == "reme.plugins"}
|
||||||
|
assert plugins["auto-fin"].value == "reme_auto_fin"
|
||||||
|
manifest = load_package_manifest("reme_auto_fin", plugin_name="auto-fin")
|
||||||
|
assert set(manifest.backends) == {
|
||||||
|
"auto_fin_data_step",
|
||||||
|
"auto_fin_topic_step",
|
||||||
|
"auto_fin_merge_step",
|
||||||
|
}
|
||||||
|
assert set(manifest.application_defaults["jobs"]) == {
|
||||||
|
"auto_fin",
|
||||||
|
"auto_fin_cron",
|
||||||
|
}
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Upload distributions
|
||||||
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
with:
|
||||||
|
name: reme-auto-fin-${{ inputs.version }}
|
||||||
|
path: dist/auto-fin/
|
||||||
|
if-no-files-found: error
|
||||||
|
|
||||||
|
publish:
|
||||||
|
needs: build
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: pypi
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- name: Download distributions
|
||||||
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: reme-auto-fin-${{ inputs.version }}
|
||||||
|
path: dist/auto-fin
|
||||||
|
|
||||||
|
- name: Publish reme-auto-fin
|
||||||
|
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
|
||||||
|
with:
|
||||||
|
packages-dir: dist/auto-fin
|
||||||
178
.github/workflows/release-daily-paper.yml
vendored
Normal file
178
.github/workflows/release-daily-paper.yml
vendored
Normal file
|
|
@ -0,0 +1,178 @@
|
||||||
|
# Release checklist:
|
||||||
|
# 1. Update project.version in plugins/daily_paper/pyproject.toml and merge it into the target branch.
|
||||||
|
# 2. Publish the required reme-ai version before this plugin; the build verifies that dependency on PyPI.
|
||||||
|
# 3. Configure PyPI Trusted Publishing for this repository/workflow and its pypi environment.
|
||||||
|
# 4. Run "Release / Daily Paper plugin" from GitHub Actions with the exact project version (a v prefix is accepted).
|
||||||
|
#
|
||||||
|
# Recommended order: reme-ai -> reme-daily-paper -> downstream applications enabling plugins: [daily-paper].
|
||||||
|
# This workflow is intentionally manual and never publishes from a push, tag, or GitHub release event.
|
||||||
|
|
||||||
|
name: Release / Daily Paper plugin
|
||||||
|
|
||||||
|
run-name: Publish reme-daily-paper ${{ inputs.version }}
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Version from plugins/daily_paper/pyproject.toml (for example, 0.1.0)
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-daily-paper
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 45
|
||||||
|
env:
|
||||||
|
RELEASE_VERSION: ${{ inputs.version }}
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Set up Python
|
||||||
|
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: '3.11'
|
||||||
|
|
||||||
|
- name: Install test and build dependencies
|
||||||
|
run: |
|
||||||
|
python -m pip install --upgrade pip
|
||||||
|
python -m pip install build packaging pytest pytest-asyncio twine
|
||||||
|
python -m pip install -e ".[core]"
|
||||||
|
python -m pip install -e plugins/daily_paper
|
||||||
|
|
||||||
|
- name: Validate package name, dependencies, and release version
|
||||||
|
id: package
|
||||||
|
run: |
|
||||||
|
python - "${RELEASE_VERSION}" <<'PY'
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import tomllib
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from packaging.requirements import Requirement
|
||||||
|
from packaging.version import Version
|
||||||
|
|
||||||
|
project = tomllib.loads(Path("plugins/daily_paper/pyproject.toml").read_text(encoding="utf-8"))["project"]
|
||||||
|
expected = Version(sys.argv[1].removeprefix("v"))
|
||||||
|
actual = Version(project["version"])
|
||||||
|
if project["name"] != "reme-daily-paper":
|
||||||
|
raise SystemExit(f"Expected project name 'reme-daily-paper', found {project['name']!r}")
|
||||||
|
if actual != expected:
|
||||||
|
raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
|
||||||
|
requirements = [Requirement(value) for value in project["dependencies"]]
|
||||||
|
reme_requirements = [requirement for requirement in requirements if requirement.name == "reme-ai"]
|
||||||
|
if len(reme_requirements) != 1 or reme_requirements[0].extras:
|
||||||
|
raise SystemExit(f"Expected one base reme-ai dependency, found {reme_requirements!r}")
|
||||||
|
if Version("0.4.1.11") in reme_requirements[0].specifier or Version("0.4.1.12") not in reme_requirements[0].specifier:
|
||||||
|
raise SystemExit(f"Expected reme-ai>=0.4.1.12, found {reme_requirements!r}")
|
||||||
|
if sum(requirement.name == "pypdf" for requirement in requirements) != 1:
|
||||||
|
raise SystemExit("Expected exactly one pypdf dependency")
|
||||||
|
root_project = tomllib.loads(Path("pyproject.toml").read_text(encoding="utf-8"))["project"]
|
||||||
|
agentscope_requirements = [
|
||||||
|
Requirement(value) for value in root_project["optional-dependencies"]["as"]
|
||||||
|
]
|
||||||
|
if len(agentscope_requirements) != 1 or agentscope_requirements[0].name != "agentscope":
|
||||||
|
raise SystemExit(f"Expected one agentscope dependency, found {agentscope_requirements!r}")
|
||||||
|
agentscope_requirement = agentscope_requirements[0]
|
||||||
|
agentscope_specifiers = list(agentscope_requirement.specifier)
|
||||||
|
if (
|
||||||
|
agentscope_requirement.extras != {"model-ollama"}
|
||||||
|
or agentscope_requirement.marker is not None
|
||||||
|
or len(agentscope_specifiers) != 1
|
||||||
|
or agentscope_specifiers[0].operator != "=="
|
||||||
|
or Version(agentscope_specifiers[0].version).is_prerelease
|
||||||
|
):
|
||||||
|
raise SystemExit(f"Expected a stable exact agentscope[model-ollama] pin, found {agentscope_requirement}")
|
||||||
|
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
|
||||||
|
print(f"reme_requirement={reme_requirements[0]}", file=output)
|
||||||
|
print(f"agentscope_requirement={agentscope_requirement}", file=output)
|
||||||
|
print(f"Publishing {project['name']} {actual}")
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Run Daily Paper tests
|
||||||
|
run: python -m pytest plugins/daily_paper -q
|
||||||
|
|
||||||
|
- name: Require the plugin-enabled ReMe release on PyPI
|
||||||
|
env:
|
||||||
|
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
|
||||||
|
run: |
|
||||||
|
python -m pip download --no-deps \
|
||||||
|
--dest "${RUNNER_TEMP}/reme-daily-paper-base" \
|
||||||
|
"${REME_REQUIREMENT}"
|
||||||
|
|
||||||
|
- name: Build and check distributions
|
||||||
|
run: |
|
||||||
|
mkdir -p dist/daily-paper
|
||||||
|
python -m build plugins/daily_paper --outdir dist/daily-paper
|
||||||
|
python -m twine check dist/daily-paper/*
|
||||||
|
|
||||||
|
- name: Verify distributions and isolated installation
|
||||||
|
env:
|
||||||
|
AGENTSCOPE_REQUIREMENT: ${{ steps.package.outputs.agentscope_requirement }}
|
||||||
|
run: |
|
||||||
|
DAILY_PAPER_WHEEL="$(pwd)/$(ls dist/daily-paper/reme_daily_paper-*.whl)"
|
||||||
|
DAILY_PAPER_SDIST="$(pwd)/$(ls dist/daily-paper/reme_daily_paper-*.tar.gz)"
|
||||||
|
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'reme_daily_paper/plugin.yaml'
|
||||||
|
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'reme_daily_paper/analyze.yaml'
|
||||||
|
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'dist-info/licenses/LICENSE'
|
||||||
|
python -m tarfile -l "${DAILY_PAPER_SDIST}" | grep '/LICENSE'
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-daily-paper-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" -m pip install \
|
||||||
|
"${AGENTSCOPE_REQUIREMENT}" "${DAILY_PAPER_WHEEL}"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" - <<'PY'
|
||||||
|
from importlib.metadata import distribution
|
||||||
|
|
||||||
|
from reme.plugin_manifest import load_package_manifest
|
||||||
|
|
||||||
|
package = distribution("reme-daily-paper")
|
||||||
|
plugins = {entry.name: entry for entry in package.entry_points if entry.group == "reme.plugins"}
|
||||||
|
assert plugins["daily-paper"].value == "reme_daily_paper"
|
||||||
|
manifest = load_package_manifest("reme_daily_paper", plugin_name="daily-paper")
|
||||||
|
assert set(manifest.backends) == {
|
||||||
|
"daily_paper_collect_step",
|
||||||
|
"daily_paper_rank_step",
|
||||||
|
"daily_paper_select_step",
|
||||||
|
"daily_paper_analyze_step",
|
||||||
|
"daily_paper_digest_step",
|
||||||
|
}
|
||||||
|
assert set(manifest.application_defaults["jobs"]) == {"daily_paper", "daily_paper_cron"}
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Upload distributions
|
||||||
|
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
with:
|
||||||
|
name: reme-daily-paper-${{ inputs.version }}
|
||||||
|
path: dist/daily-paper/
|
||||||
|
if-no-files-found: error
|
||||||
|
|
||||||
|
publish:
|
||||||
|
needs: build
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: pypi
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- name: Download distributions
|
||||||
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: reme-daily-paper-${{ inputs.version }}
|
||||||
|
path: dist/daily-paper
|
||||||
|
|
||||||
|
- name: Publish reme-daily-paper
|
||||||
|
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
|
||||||
|
with:
|
||||||
|
packages-dir: dist/daily-paper
|
||||||
49
.github/workflows/release-dsh-plugin.yml
vendored
Normal file
49
.github/workflows/release-dsh-plugin.yml
vendored
Normal file
|
|
@ -0,0 +1,49 @@
|
||||||
|
# Configure npm Trusted Publishing for this caller filename and the npm environment.
|
||||||
|
name: Release / DSH plugin
|
||||||
|
|
||||||
|
run-name: Publish ReMe DSH plugin ${{ inputs.version }} (${{ inputs.npm_tag }})
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Exact package.json version; an optional v prefix is accepted
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
npm_tag:
|
||||||
|
description: npm distribution tag
|
||||||
|
required: true
|
||||||
|
default: latest
|
||||||
|
type: choice
|
||||||
|
options:
|
||||||
|
- next
|
||||||
|
- latest
|
||||||
|
use_npm_token:
|
||||||
|
description: Use the npm environment NPM_TOKEN instead of Trusted Publishing
|
||||||
|
required: true
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-dsh-plugin
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
release:
|
||||||
|
if: github.ref == 'refs/heads/main'
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
uses: ./.github/workflows/_release-npm-plugin.yml
|
||||||
|
secrets:
|
||||||
|
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||||
|
with:
|
||||||
|
directory: integrations/dsh
|
||||||
|
package_name: '@agentscope-ai/reme-dsh-plugin'
|
||||||
|
artifact_name: agentscope-ai-reme-dsh-plugin
|
||||||
|
version: ${{ inputs.version }}
|
||||||
|
npm_tag: ${{ inputs.npm_tag }}
|
||||||
|
use_npm_token: ${{ inputs.use_npm_token }}
|
||||||
75
.github/workflows/release-openclaw-plugin.yml
vendored
Normal file
75
.github/workflows/release-openclaw-plugin.yml
vendored
Normal file
|
|
@ -0,0 +1,75 @@
|
||||||
|
# Configure npm Trusted Publishing for this caller filename and the npm environment.
|
||||||
|
name: Release / OpenClaw plugin
|
||||||
|
|
||||||
|
run-name: Publish ReMe OpenClaw plugin ${{ inputs.version }} (${{ inputs.npm_tag }})
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Exact package.json version; an optional v prefix is accepted
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
npm_tag:
|
||||||
|
description: npm distribution tag
|
||||||
|
required: true
|
||||||
|
default: latest
|
||||||
|
type: choice
|
||||||
|
options:
|
||||||
|
- next
|
||||||
|
- latest
|
||||||
|
use_npm_token:
|
||||||
|
description: Use the npm environment NPM_TOKEN instead of Trusted Publishing
|
||||||
|
required: true
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
publish_clawhub:
|
||||||
|
description: Also publish the package to ClawHub
|
||||||
|
required: true
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-openclaw-plugin
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
release:
|
||||||
|
if: github.ref == 'refs/heads/main'
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
uses: ./.github/workflows/_release-npm-plugin.yml
|
||||||
|
secrets:
|
||||||
|
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||||
|
with:
|
||||||
|
directory: integrations/openclaw
|
||||||
|
package_name: '@agentscope-ai/reme-openclaw-plugin'
|
||||||
|
artifact_name: agentscope-ai-reme-openclaw-plugin
|
||||||
|
version: ${{ inputs.version }}
|
||||||
|
npm_tag: ${{ inputs.npm_tag }}
|
||||||
|
use_npm_token: ${{ inputs.use_npm_token }}
|
||||||
|
validate_clawhub: true
|
||||||
|
|
||||||
|
publish-clawhub:
|
||||||
|
if: ${{ inputs.publish_clawhub && github.ref == 'refs/heads/main' }}
|
||||||
|
needs: release
|
||||||
|
permissions:
|
||||||
|
actions: read
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
uses: openclaw/clawhub/.github/workflows/package-publish.yml@cacf5ec1b0ee3cb532ab4555a68c1db9e0c1aac7 # v0.24.0
|
||||||
|
with:
|
||||||
|
family: code-plugin
|
||||||
|
version: ${{ needs.release.outputs.version }}
|
||||||
|
tags: ${{ inputs.npm_tag }}
|
||||||
|
source_repo: ${{ github.repository }}
|
||||||
|
source_commit: ${{ github.sha }}
|
||||||
|
source_ref: ${{ github.sha }}
|
||||||
|
source_path: integrations/openclaw
|
||||||
|
package_artifact_name: agentscope-ai-reme-openclaw-plugin-${{ needs.release.outputs.version }}
|
||||||
|
dry_run: false
|
||||||
|
wait_for_publication: true
|
||||||
52
.github/workflows/release-python.yml
vendored
Normal file
52
.github/workflows/release-python.yml
vendored
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
name: Release / ReMe Python package
|
||||||
|
|
||||||
|
# Publishing a GitHub Release is the primary release trigger for reme-ai; workflow_dispatch is the recovery path.
|
||||||
|
# Configure a PyPI Trusted Publisher for this repository, workflow, and its pypi environment first.
|
||||||
|
|
||||||
|
run-name: Publish reme-ai ${{ github.event.release.tag_name || inputs.version }}
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Exact reme-ai version; an optional v prefix is accepted
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
release:
|
||||||
|
types: [published]
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-ai
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
name: Build and verify distributions
|
||||||
|
uses: ./.github/workflows/_build-python-packages.yml
|
||||||
|
with:
|
||||||
|
expected_version: ${{ github.event.release.tag_name || inputs.version }}
|
||||||
|
upload_artifacts: true
|
||||||
|
|
||||||
|
publish-reme:
|
||||||
|
needs: build
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: pypi
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
steps:
|
||||||
|
- name: Download ReMe distributions
|
||||||
|
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: reme-distributions
|
||||||
|
path: dist/reme
|
||||||
|
|
||||||
|
- name: Publish ReMe
|
||||||
|
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
|
||||||
|
with:
|
||||||
|
packages-dir: dist/reme
|
||||||
|
skip-existing: true
|
||||||
243
.github/workflows/release-reme-studio.yml
vendored
Normal file
243
.github/workflows/release-reme-studio.yml
vendored
Normal file
|
|
@ -0,0 +1,243 @@
|
||||||
|
# Release checklist:
|
||||||
|
# 1. Update reme_studio/pyproject.toml, package.json, and package-lock.json to the same Studio version.
|
||||||
|
# 2. Configure npm Trusted Publishing with the npm environment and PyPI Trusted Publishing with the pypi environment.
|
||||||
|
# 3. Run this workflow manually with the exact Studio version.
|
||||||
|
|
||||||
|
name: Release / ReMe Studio
|
||||||
|
|
||||||
|
run-name: Publish ReMe Studio ${{ inputs.version }} (${{ inputs.npm_tag }})
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
version:
|
||||||
|
description: Version from the Studio Python and npm manifests
|
||||||
|
required: true
|
||||||
|
type: string
|
||||||
|
npm_tag:
|
||||||
|
description: npm distribution tag
|
||||||
|
required: true
|
||||||
|
default: latest
|
||||||
|
type: choice
|
||||||
|
options:
|
||||||
|
- next
|
||||||
|
- latest
|
||||||
|
use_npm_token:
|
||||||
|
description: Use the npm environment NPM_TOKEN instead of Trusted Publishing
|
||||||
|
required: true
|
||||||
|
default: false
|
||||||
|
type: boolean
|
||||||
|
publish_target:
|
||||||
|
description: Packages to publish; single-package modes are for release recovery
|
||||||
|
required: true
|
||||||
|
default: both
|
||||||
|
type: choice
|
||||||
|
options:
|
||||||
|
- both
|
||||||
|
- pypi
|
||||||
|
- npm
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: publish-reme-studio
|
||||||
|
cancel-in-progress: false
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
build:
|
||||||
|
if: github.ref == 'refs/heads/main'
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 45
|
||||||
|
env:
|
||||||
|
RELEASE_VERSION: ${{ inputs.version }}
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "22.22.3"
|
||||||
|
cache: npm
|
||||||
|
cache-dependency-path: reme_studio/package-lock.json
|
||||||
|
|
||||||
|
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
|
||||||
|
with:
|
||||||
|
python-version: "3.11"
|
||||||
|
|
||||||
|
- name: Validate Studio package names and version
|
||||||
|
run: |
|
||||||
|
python - <<'PY'
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import tomllib
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
studio = Path("reme_studio")
|
||||||
|
python_manifest = tomllib.loads((studio / "pyproject.toml").read_text(encoding="utf-8"))["project"]
|
||||||
|
npm_manifest = json.loads((studio / "package.json").read_text(encoding="utf-8"))
|
||||||
|
expected = os.environ["RELEASE_VERSION"].removeprefix("v")
|
||||||
|
if python_manifest["name"] != "reme_studio":
|
||||||
|
raise SystemExit(f"Unexpected Python package name: {python_manifest['name']}")
|
||||||
|
if npm_manifest["name"] != "@agentscope-ai/reme_studio":
|
||||||
|
raise SystemExit(f"Unexpected npm package name: {npm_manifest['name']}")
|
||||||
|
if python_manifest["version"] != expected or npm_manifest["version"] != expected:
|
||||||
|
raise SystemExit(
|
||||||
|
f"Studio manifests are {python_manifest['version']} and {npm_manifest['version']}; "
|
||||||
|
f"workflow input is {expected}",
|
||||||
|
)
|
||||||
|
prerelease = "-" in expected
|
||||||
|
if prerelease != (os.environ["NPM_TAG"] == "next"):
|
||||||
|
raise SystemExit("Prereleases must use next; stable releases must use latest")
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Install dependencies and run checks
|
||||||
|
working-directory: reme_studio
|
||||||
|
run: |
|
||||||
|
npm ci
|
||||||
|
npm run format:check
|
||||||
|
npm run lint
|
||||||
|
npm test
|
||||||
|
|
||||||
|
- name: Build Studio distributions
|
||||||
|
run: |
|
||||||
|
python -m pip install build twine
|
||||||
|
mkdir -p dist/studio-python dist/studio-npm
|
||||||
|
npm pack ./reme_studio --pack-destination dist/studio-npm
|
||||||
|
python scripts/package_studio.py
|
||||||
|
python -m build reme_studio --outdir dist/studio-python
|
||||||
|
python -m twine check dist/studio-python/*
|
||||||
|
|
||||||
|
- name: Verify Studio distributions and isolated installation
|
||||||
|
run: |
|
||||||
|
STUDIO_WHEEL="$(pwd)/$(ls dist/studio-python/reme_studio-*.whl)"
|
||||||
|
tar -tzf dist/studio-npm/*.tgz | grep '^package/dist-static/index.html$'
|
||||||
|
python -m venv "${RUNNER_TEMP}/reme-studio-package-smoke"
|
||||||
|
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" -m pip install "${STUDIO_WHEEL}"
|
||||||
|
cd "${RUNNER_TEMP}"
|
||||||
|
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" - <<'PY'
|
||||||
|
from reme_studio import static_dir
|
||||||
|
|
||||||
|
assert (static_dir() / "index.html").is_file()
|
||||||
|
PY
|
||||||
|
|
||||||
|
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
||||||
|
with:
|
||||||
|
name: reme-studio-${{ inputs.version }}
|
||||||
|
path: |
|
||||||
|
dist/studio-python/*
|
||||||
|
dist/studio-npm/*
|
||||||
|
if-no-files-found: error
|
||||||
|
|
||||||
|
publish-python:
|
||||||
|
if: inputs.publish_target != 'npm'
|
||||||
|
needs: build
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: pypi
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
steps:
|
||||||
|
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: reme-studio-${{ inputs.version }}
|
||||||
|
path: dist
|
||||||
|
|
||||||
|
- name: Publish ReMe Studio to PyPI
|
||||||
|
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
|
||||||
|
with:
|
||||||
|
packages-dir: dist/studio-python
|
||||||
|
skip-existing: true
|
||||||
|
|
||||||
|
publish-npm:
|
||||||
|
if: >-
|
||||||
|
!cancelled() &&
|
||||||
|
needs.build.result == 'success' &&
|
||||||
|
(inputs.publish_target == 'npm' ||
|
||||||
|
(inputs.publish_target == 'both' && needs.publish-python.result == 'success'))
|
||||||
|
needs: [build, publish-python]
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 10
|
||||||
|
environment: npm
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
id-token: write
|
||||||
|
steps:
|
||||||
|
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
|
||||||
|
with:
|
||||||
|
node-version: "24"
|
||||||
|
registry-url: https://registry.npmjs.org
|
||||||
|
|
||||||
|
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
||||||
|
with:
|
||||||
|
name: reme-studio-${{ inputs.version }}
|
||||||
|
path: dist
|
||||||
|
|
||||||
|
- name: Verify the matching PyPI release for npm-only recovery
|
||||||
|
if: inputs.publish_target == 'npm'
|
||||||
|
env:
|
||||||
|
PACKAGE_VERSION: ${{ inputs.version }}
|
||||||
|
run: |
|
||||||
|
python - <<'PY'
|
||||||
|
import os
|
||||||
|
import urllib.error
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
version = os.environ["PACKAGE_VERSION"].removeprefix("v")
|
||||||
|
url = f"https://pypi.org/pypi/reme-studio/{version}/json"
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(url, timeout=30) as response:
|
||||||
|
if response.status != 200:
|
||||||
|
raise SystemExit(f"Unexpected PyPI response for reme-studio {version}: {response.status}")
|
||||||
|
except urllib.error.HTTPError as exc:
|
||||||
|
raise SystemExit(f"reme-studio {version} must exist on PyPI before npm-only recovery") from exc
|
||||||
|
PY
|
||||||
|
|
||||||
|
- name: Check for an identical existing npm package
|
||||||
|
id: npm-version
|
||||||
|
env:
|
||||||
|
PACKAGE_VERSION: ${{ inputs.version }}
|
||||||
|
run: |
|
||||||
|
package_file=$(find dist/studio-npm -maxdepth 1 -name '*.tgz' -print -quit)
|
||||||
|
if [[ -z "${package_file}" ]]; then
|
||||||
|
echo "Studio npm artifact is missing" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
local_integrity=$(node --input-type=module - "${package_file}" <<'JS'
|
||||||
|
import { createHash } from 'node:crypto';
|
||||||
|
import { readFileSync } from 'node:fs';
|
||||||
|
const digest = createHash('sha512').update(readFileSync(process.argv[2])).digest('base64');
|
||||||
|
console.log(`sha512-${digest}`);
|
||||||
|
JS
|
||||||
|
)
|
||||||
|
if remote_integrity=$(npm view "@agentscope-ai/reme_studio@${PACKAGE_VERSION#v}" dist.integrity 2>/dev/null); then
|
||||||
|
if [[ "${remote_integrity}" != "${local_integrity}" ]]; then
|
||||||
|
echo "Existing npm package has different contents" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "exists=true" >> "${GITHUB_OUTPUT}"
|
||||||
|
echo "The identical npm package already exists; nothing to publish"
|
||||||
|
else
|
||||||
|
echo "exists=false" >> "${GITHUB_OUTPUT}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
- name: Publish ReMe Studio to npm with Trusted Publishing
|
||||||
|
if: ${{ steps.npm-version.outputs.exists != 'true' && !inputs.use_npm_token }}
|
||||||
|
env:
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
run: npm publish dist/studio-npm/*.tgz --access public --tag "${NPM_TAG}" --provenance
|
||||||
|
- name: Publish ReMe Studio to npm with NPM_TOKEN
|
||||||
|
if: ${{ steps.npm-version.outputs.exists != 'true' && inputs.use_npm_token }}
|
||||||
|
env:
|
||||||
|
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||||
|
NPM_TAG: ${{ inputs.npm_tag }}
|
||||||
|
run: |
|
||||||
|
if [[ -z "${NODE_AUTH_TOKEN}" ]]; then
|
||||||
|
echo "NPM_TOKEN is required when use_npm_token is enabled" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
npm publish dist/studio-npm/*.tgz --access public --tag "${NPM_TAG}" --provenance
|
||||||
45
.github/workflows/security-codeql.yml
vendored
Normal file
45
.github/workflows/security-codeql.yml
vendored
Normal file
|
|
@ -0,0 +1,45 @@
|
||||||
|
name: Security / CodeQL
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches: [main]
|
||||||
|
pull_request:
|
||||||
|
branches: [main]
|
||||||
|
schedule:
|
||||||
|
- cron: '0 1 * * 1'
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: read
|
||||||
|
security-events: write
|
||||||
|
|
||||||
|
concurrency:
|
||||||
|
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
||||||
|
cancel-in-progress: true
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
analyze:
|
||||||
|
name: Analyze ${{ matrix.language }}
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
timeout-minutes: 60
|
||||||
|
strategy:
|
||||||
|
fail-fast: false
|
||||||
|
matrix:
|
||||||
|
language: [python, javascript-typescript]
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- name: Checkout repository
|
||||||
|
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
||||||
|
with:
|
||||||
|
persist-credentials: false
|
||||||
|
|
||||||
|
- name: Initialize CodeQL
|
||||||
|
uses: github/codeql-action/init@1c5b675653bb5c22dbe9b12b556ec555138e09fd # v4
|
||||||
|
with:
|
||||||
|
languages: ${{ matrix.language }}
|
||||||
|
build-mode: none
|
||||||
|
|
||||||
|
- name: Perform CodeQL analysis
|
||||||
|
uses: github/codeql-action/analyze@1c5b675653bb5c22dbe9b12b556ec555138e09fd # v4
|
||||||
|
with:
|
||||||
|
category: /language:${{ matrix.language }}
|
||||||
38
.github/workflows/windows-smoke.yml
vendored
38
.github/workflows/windows-smoke.yml
vendored
|
|
@ -1,38 +0,0 @@
|
||||||
name: Windows Smoke
|
|
||||||
|
|
||||||
on:
|
|
||||||
push:
|
|
||||||
branches: [main, master, dev, develop]
|
|
||||||
pull_request:
|
|
||||||
branches: [main, master, dev, develop]
|
|
||||||
workflow_dispatch:
|
|
||||||
|
|
||||||
concurrency:
|
|
||||||
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
|
|
||||||
cancel-in-progress: true
|
|
||||||
|
|
||||||
jobs:
|
|
||||||
cli-smoke:
|
|
||||||
name: CLI smoke - py${{ matrix.python-version }}
|
|
||||||
runs-on: windows-latest
|
|
||||||
strategy:
|
|
||||||
fail-fast: false
|
|
||||||
matrix:
|
|
||||||
python-version: ["3.11"]
|
|
||||||
|
|
||||||
steps:
|
|
||||||
- uses: actions/checkout@v4
|
|
||||||
|
|
||||||
- name: Set up Python ${{ matrix.python-version }}
|
|
||||||
uses: actions/setup-python@v5
|
|
||||||
with:
|
|
||||||
python-version: ${{ matrix.python-version }}
|
|
||||||
cache: 'pip'
|
|
||||||
|
|
||||||
- name: Install package
|
|
||||||
run: |
|
|
||||||
python -m pip install --upgrade pip setuptools wheel
|
|
||||||
pip install -e ".[core,benchmark]"
|
|
||||||
|
|
||||||
- name: Run version job
|
|
||||||
run: reme start service.backend=cli job=version
|
|
||||||
26
.gitignore
vendored
26
.gitignore
vendored
|
|
@ -2,8 +2,8 @@
|
||||||
.DS_Store
|
.DS_Store
|
||||||
.idea/
|
.idea/
|
||||||
.vscode/
|
.vscode/
|
||||||
*.code-workspace
|
|
||||||
.qoder/
|
.qoder/
|
||||||
|
*.code-workspace
|
||||||
|
|
||||||
# Local environment
|
# Local environment
|
||||||
.env
|
.env
|
||||||
|
|
@ -30,7 +30,9 @@ htmlcov/
|
||||||
# Packaging / build outputs
|
# Packaging / build outputs
|
||||||
build/
|
build/
|
||||||
dist/
|
dist/
|
||||||
|
node_modules/
|
||||||
*.egg-info/
|
*.egg-info/
|
||||||
|
integrations/*/reports/
|
||||||
|
|
||||||
# Logs / temporary files
|
# Logs / temporary files
|
||||||
*.log
|
*.log
|
||||||
|
|
@ -45,6 +47,8 @@ temp*/
|
||||||
|
|
||||||
# ReMe runtime data
|
# ReMe runtime data
|
||||||
.reme/
|
.reme/
|
||||||
|
reme_workspace/
|
||||||
|
reme_workspace_auto_fin_real_test*/
|
||||||
vault/
|
vault/
|
||||||
*.db
|
*.db
|
||||||
*.sqlite
|
*.sqlite
|
||||||
|
|
@ -55,4 +59,24 @@ docs/_build/
|
||||||
site/
|
site/
|
||||||
|
|
||||||
evaluation/
|
evaluation/
|
||||||
|
# The pi-Bench suite ships its own trace-history render config, which must
|
||||||
|
# stay in git even though it lives under an evaluation/ directory.
|
||||||
|
!benchmark/pibench/config/bench/evaluation/
|
||||||
|
!benchmark/pibench/config/bench/evaluation/**
|
||||||
datasets/
|
datasets/
|
||||||
|
|
||||||
|
# Claude Code skills (local only)
|
||||||
|
.claude/skills/
|
||||||
|
|
||||||
|
# Benchmark memory workspaces (created on demand by run.py via mkdir)
|
||||||
|
benchmark/*/workspaces/
|
||||||
|
|
||||||
|
# Benchmark datasets (LongMemEval via download.py, BEAM via git clone)
|
||||||
|
benchmark/*/dataset/
|
||||||
|
|
||||||
|
# Benchmark outputs (created on demand by run.py via mkdir)
|
||||||
|
benchmark/*/results/
|
||||||
|
|
||||||
|
# integration tests outputs
|
||||||
|
tests/integration/logs/
|
||||||
|
daily/
|
||||||
|
|
|
||||||
|
|
@ -1,3 +1,5 @@
|
||||||
|
exclude: ^skills/
|
||||||
|
|
||||||
repos:
|
repos:
|
||||||
- repo: https://github.com/pre-commit/pre-commit-hooks
|
- repo: https://github.com/pre-commit/pre-commit-hooks
|
||||||
rev: v6.0.0
|
rev: v6.0.0
|
||||||
|
|
|
||||||
275
AGENTS.md
275
AGENTS.md
|
|
@ -1,20 +1,19 @@
|
||||||
# AGENTS.md
|
# AGENTS.md
|
||||||
|
|
||||||
This file guides coding agents working in the ReMe repository. Keep changes small,
|
This file guides coding agents working in the ReMe repository. Keep changes small, testable, and consistent with the
|
||||||
testable, and consistent with the contracts already expressed by the code.
|
contracts expressed by the current code.
|
||||||
|
|
||||||
## Project Principles
|
## Project Principles
|
||||||
|
|
||||||
ReMe is a local-first, file-native memory system for agents.
|
ReMe is a local-first, file-native memory system for agents.
|
||||||
|
|
||||||
- User-owned memory files are the source of truth.
|
- User-owned workspace files are the durable source of truth.
|
||||||
- Indexes, caches, metadata, and generated state must be rebuildable.
|
- Indexes, catalogs, graphs, caches, and generated metadata must remain rebuildable.
|
||||||
- Prefer transparent formats and behavior over hidden state.
|
- Prefer transparent formats and predictable behavior over hidden state.
|
||||||
- Preserve user control over storage, configuration, and service boundaries.
|
- Preserve user control over workspace paths, configuration, and service boundaries.
|
||||||
- Keep concepts focused on project intent; let code and schemas describe implementation.
|
- Keep concepts focused on project intent; let code and schemas describe implementation.
|
||||||
|
|
||||||
When a proposed convenience conflicts with these principles, favor data ownership,
|
When convenience conflicts with these principles, favor data ownership, recoverability, and explicit behavior.
|
||||||
recoverability, and predictable behavior.
|
|
||||||
|
|
||||||
## Sources of Truth
|
## Sources of Truth
|
||||||
|
|
||||||
|
|
@ -22,166 +21,198 @@ Use this order when documentation and implementation disagree:
|
||||||
|
|
||||||
1. Current code and public Pydantic schemas.
|
1. Current code and public Pydantic schemas.
|
||||||
2. Tests that describe supported behavior.
|
2. Tests that describe supported behavior.
|
||||||
3. CLI help and the built-in configuration.
|
3. CLI behavior and the built-in configuration.
|
||||||
4. Development documentation and historical notes.
|
4. README files and other development documentation.
|
||||||
|
|
||||||
Do not copy large implementation descriptions into documentation. Link to the relevant
|
Do not duplicate large implementation descriptions in documentation. Express the stable contract and link to the
|
||||||
module or express the stable contract instead. If behavior changes intentionally, update
|
relevant module where useful. When behavior changes intentionally, update the implementation, schemas, tests, defaults,
|
||||||
the code, schema, tests, configuration, and concise documentation together as needed.
|
and concise documentation together.
|
||||||
|
|
||||||
## Repository Map
|
## Repository Map
|
||||||
|
|
||||||
- `reme/reme.py`: CLI entry point and client/server dispatch.
|
- `reme/reme.py`: CLI entry point; dispatches `start`, `find_reme`, and client calls.
|
||||||
- `reme/application.py`: application assembly, dependency ordering, and lifecycle.
|
- `reme/application.py`: application assembly, dependency ordering, job execution, and lifecycle.
|
||||||
- `reme/components/application_context.py`: application-wide wiring and shared in-memory metadata.
|
- `reme/config/config_parser.py`: YAML/JSON loading, environment expansion, dot-notation parsing, and deep config
|
||||||
- `reme/components/runtime_context.py`: scratch state shared by steps within one execution.
|
merging.
|
||||||
- `reme/config/default.yaml`: built-in jobs, components, and defaults.
|
- `reme/config/default.yaml`: default service, jobs, steps, and components. Other files in
|
||||||
- `reme/schema/`: public and runtime Pydantic contracts.
|
`reme/config/` are named configuration variants.
|
||||||
- `reme/components/`: services, stores, clients, jobs, and component registration.
|
- `reme/schema/application_config.py`: typed application, component, and job configuration.
|
||||||
- `reme/steps/`: executable job steps.
|
- `reme/schema/`: request, response, streaming, memory, graph, and file contracts.
|
||||||
- `tests/unit/`: primary fast validation suite.
|
- `reme/components/application_context.py`: application-wide wiring and in-memory shared state.
|
||||||
- `tests/integration/`: tests that may require real credentials or services.
|
- `reme/components/runtime_context.py`: request-scoped data, response, streaming queue, and stop event.
|
||||||
- `tests/vector/` and `tests/light/`: specialized suites.
|
- `reme/components/base_component.py`: component lifecycle, dependency binding, and workspace helpers.
|
||||||
- `plugins/reme/`: Claude Code integration.
|
- `reme/components/component_registry.py`: the frozen built-in registry template and application-local registry factory.
|
||||||
- `skills/reme_memory/`: skill that communicates with the ReMe service.
|
- `reme/components/job/`: base, stream, background, and cron job implementations.
|
||||||
- `skills/qwenpaw_memory/`: separate direct-file memory convention; it does not call ReMe.
|
- `reme/components/service/`: local CLI, HTTP, and MCP service backends.
|
||||||
- `docs/`: pages and assets that support the repository README; not the deployed docs site.
|
- `reme/components/`: agent wrappers, model adapters, stores, catalogs, graphs, indexes, clients, tokenizers, and
|
||||||
|
outbound proxies.
|
||||||
|
- `reme/steps/`: registered job steps grouped by common, file I/O, index, evolve, cookbook, and transfer
|
||||||
|
concerns.
|
||||||
|
- `reme/utils/`: shared utilities, including service discovery, logging, web-static resolution, session I/O, token
|
||||||
|
accounting, and wikilink handling.
|
||||||
|
- `tests/unit/`: primary fast, isolated validation suite.
|
||||||
|
- `tests/integration/`: service/model tests that may need credentials or external processes.
|
||||||
|
- `reme_studio/`: ReMe Studio frontend source plus the independently published `reme_studio` Python package and
|
||||||
|
`@agentscope-ai/reme_studio` npm static distribution.
|
||||||
|
- `plugins/`: installable ReMe extensions, including Auto Fin and LME/BEAM plugins.
|
||||||
|
- `integrations/`: adapters that connect ReMe to external agent hosts, including the independent, self-contained DSH
|
||||||
|
and OpenClaw TypeScript plugins plus the Claude Code and Hermes Agent integrations.
|
||||||
|
- `skills/`: standalone skills; `reme_memory` calls ReMe, while other skills may use separate tools or direct-file
|
||||||
|
conventions.
|
||||||
|
- `benchmark/` and `cookbook/`: runnable evaluations and example workflows.
|
||||||
|
- `docs/`: README-linked supporting pages and figures.
|
||||||
|
- `github-pages/`: VitePress build shell, generated-content assembly, documentation checks, and GitHub Pages output. The
|
||||||
|
canonical theme and guides remain under `docs/`; `.generated/` and `dist/` are disposable.
|
||||||
|
|
||||||
## Development Setup
|
## Development Setup
|
||||||
|
|
||||||
ReMe requires Python 3.11 or newer.
|
ReMe requires Python 3.11 or newer. Install the editable development environment with:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install -e ".[dev,core]"
|
pip install -e reme_studio -e ".[dev,core]"
|
||||||
```
|
```
|
||||||
|
|
||||||
Before changing behavior, inspect the adjacent implementation, schemas, configuration,
|
Before changing behavior, inspect the adjacent implementation, schema, built-in config, and focused tests. Follow
|
||||||
and focused tests. Follow existing patterns unless the task explicitly calls for a new
|
existing async and typing patterns unless the task explicitly requires a new contract.
|
||||||
contract or architecture.
|
|
||||||
|
|
||||||
## Change Workflow
|
## Configuration and CLI Contracts
|
||||||
|
|
||||||
1. Identify the narrowest supported contract affected by the request.
|
- CLI syntax is `reme ACTION key=value ...`; leading `-` or `--` on arguments is accepted.
|
||||||
2. Read the relevant implementation and tests before editing.
|
- Nested overrides use dot notation. Values support null, booleans, numbers, JSON collections, and quoted JSON strings;
|
||||||
3. Make the smallest coherent change; avoid unrelated cleanup.
|
leading-zero numeric-looking values remain strings.
|
||||||
4. Update related schemas, defaults, registrations, and imports when required.
|
- `config=<name-or-path>` loads a discovered config name or a `.yaml`, `.yml`, or `.json` file. With no explicit config
|
||||||
5. Add or adjust focused tests for observable behavior.
|
path, `default` is loaded when available.
|
||||||
6. Run proportionate validation and report anything not run.
|
- Config files expand `${VAR}` and `${VAR:-default}` recursively. An undefined variable without a default is an error.
|
||||||
|
- CLI/config overrides are deep-merged over the loaded file. Do not silently change this merge behavior or stable
|
||||||
|
configuration keys.
|
||||||
|
- `ApplicationConfig` normalizes `workspace_dir` to an expanded absolute path. `session_dir`
|
||||||
|
must remain workspace-relative; standard transcripts live under `{session_dir}/dialog`.
|
||||||
|
- `reme start` runs the configured service. `reme start job=<name> ...` switches to the one-shot CLI service and runs
|
||||||
|
the job through the normal application lifecycle.
|
||||||
|
- Other actions use a client selected from the running service configuration when discoverable, otherwise from local
|
||||||
|
config. Client-selection arguments must not leak into the job payload.
|
||||||
|
|
||||||
Component and step discovery depends on registration imports:
|
## Registration and Application Lifecycle
|
||||||
|
|
||||||
- Components use `R.register(...)` in `reme/components/component_registry.py`.
|
Component and Step discovery is import-driven:
|
||||||
- Component packages must be reachable through `reme/components/__init__.py`.
|
|
||||||
- Step modules must be reachable through `reme/steps/__init__.py`.
|
|
||||||
|
|
||||||
Adding an implementation without its registration import can leave it undiscoverable at
|
- Implementations declare a non-`BASE` `component_type` and register with `@R.register("backend")`
|
||||||
runtime. Treat the implementation, registry entry, and import side effect as one change.
|
or `R.register(Class, "backend")`.
|
||||||
|
- Component packages must be imported through `reme/components/__init__.py`.
|
||||||
|
- Step packages/modules must be reachable through their package `__init__.py` chain and ultimately
|
||||||
|
`reme/steps/__init__.py`.
|
||||||
|
- Adding an implementation without its registration import leaves it undiscoverable at runtime. Treat implementation,
|
||||||
|
registration, import side effect, defaults, and tests as one change.
|
||||||
|
|
||||||
Do not silently change stable CLI flags, configuration keys, workspace layouts, serialized
|
`Application` validates config through `ApplicationContext`, creates workspace directories, instantiates the service,
|
||||||
schemas, or service interfaces. When such a change is required, preserve compatibility
|
configured components, and jobs, and then manages lifecycle as follows:
|
||||||
where practical and make the migration explicit.
|
|
||||||
|
|
||||||
## Step State Model
|
- Components start in topological dependency order. Missing required dependencies and cycles fail explicitly; optional
|
||||||
|
dependencies may resolve to `None`.
|
||||||
|
- Jobs start after components in this order: base jobs, stream jobs, background jobs, then cron jobs.
|
||||||
|
- Shutdown closes everything in reverse start order and then shuts down the optional thread pool.
|
||||||
|
- If startup fails, already-started resources are closed.
|
||||||
|
- `BaseComponent.start()` and `close()` are lock-protected and idempotent. Dependencies created by a standalone
|
||||||
|
`default_factory` are owned and closed by the parent component.
|
||||||
|
|
||||||
Treat every Step as stateless. `BaseJob` stores Step specifications and builds fresh Step
|
Keep async clients, tasks, executors, and services under this lifecycle. Do not introduce an untracked long-lived
|
||||||
instances for each Job invocation. A Step instance must not use `self` or class variables to
|
resource.
|
||||||
retain mutable runtime state between calls.
|
|
||||||
|
|
||||||
Place state according to its lifetime:
|
## Jobs, Steps, and State
|
||||||
|
|
||||||
- Constructor fields on `self`: immutable Step configuration and resolved dependencies only.
|
`BaseJob` resolves configured Step classes during job startup and constructs fresh Step instances for every invocation.
|
||||||
- `self.context` (`RuntimeContext`): request data and intermediate results for one Job
|
Job-level kwargs are merged into each `RuntimeContext`, with call-time kwargs taking precedence. Sequential Steps in one
|
||||||
execution; sequential Steps share this context.
|
invocation share the same `RuntimeContext` and `Response`.
|
||||||
- `self.app_context.metadata`: in-memory state that must be shared across Step or Job
|
|
||||||
invocations for the lifetime of the Application.
|
|
||||||
- Workspace files or a dedicated Component/store: durable state that must survive an
|
|
||||||
Application restart.
|
|
||||||
|
|
||||||
Use narrow, namespaced keys in `app_context.metadata`, following existing patterns such as
|
Treat Step instances as invocation-scoped:
|
||||||
`tool_contexts` and `channel_sink`. The ApplicationContext is shared, so account for
|
|
||||||
concurrent access when values are mutable. New Step code must not fall back to `self.kwargs`
|
|
||||||
or another Step field to emulate shared state when `app_context` is absent; tests of shared
|
|
||||||
state should construct an `ApplicationContext`. If shared state grows into a stable
|
|
||||||
service-level contract or needs its own lifecycle, locking, or persistence, promote it to a
|
|
||||||
typed ApplicationContext field or a dedicated Component instead of expanding an ad hoc
|
|
||||||
metadata bucket.
|
|
||||||
|
|
||||||
Do not use `Response.metadata` as a state store. It is request-scoped output for callers and
|
- Constructor fields and `self.kwargs` hold Step configuration and resolved dependencies. They may be cached or adjusted
|
||||||
diagnostics, distinct from `ApplicationContext.metadata`.
|
during that one invocation, but must not be relied on across Job calls.
|
||||||
|
- `self.context.data` holds request inputs and intermediate values shared by sequential Steps.
|
||||||
|
- `self.context.response.answer`, `success`, and `metadata` are request-scoped output. Because the same response travels
|
||||||
|
through the Step chain, later Steps may consume metadata produced earlier, but it is not application-lifetime or
|
||||||
|
durable storage.
|
||||||
|
- `self.app_context.metadata` holds in-memory state shared across Job/Step invocations for the life of one
|
||||||
|
`Application`, such as counters, tool-context state, session maps, or locks.
|
||||||
|
- Workspace files or a dedicated Component/store hold durable state that must survive restart.
|
||||||
|
|
||||||
|
Use narrow, namespaced keys in `app_context.metadata` and protect shared mutable values against concurrent access. The
|
||||||
|
search/draft helpers intentionally mirror tool-context state into
|
||||||
|
`self.kwargs` only when no `ApplicationContext` exists for standalone use and unit tests; do not generalize that
|
||||||
|
compatibility fallback into persistent runtime state. If shared state becomes a stable service contract or needs
|
||||||
|
dedicated lifecycle, locking, or persistence, promote it to a typed context field or Component.
|
||||||
|
|
||||||
|
Additional Step contracts:
|
||||||
|
|
||||||
|
- `Ref` dependencies resolve in this order: Step kwargs, current `RuntimeContext`, then the named application component.
|
||||||
|
The value is cached only on the current Step instance and cleared before each call.
|
||||||
|
- `input_mapping` and `output_mapping` copy keys within `RuntimeContext.data`; missing sources are ignored.
|
||||||
|
- Dispatched Steps receive the current `RuntimeContext`, so their data and response are shared.
|
||||||
|
- Base jobs convert uncaught Step errors into `Response(success=False)`; stream jobs emit an error chunk and always a
|
||||||
|
terminal `DONE`; background jobs let errors reach their supervisor.
|
||||||
|
- Background jobs are never service-exposed. MCP also skips stream jobs. Respect `enable_serve`
|
||||||
|
and any configured service job allowlist.
|
||||||
|
|
||||||
|
## Workspace and File Safety
|
||||||
|
|
||||||
|
- Application startup creates the workspace plus configured metadata, session, memory-session, resource, daily, and
|
||||||
|
digest directories.
|
||||||
|
- File-operation paths are resolved against the workspace and must stay inside it. Home-relative paths are unsupported,
|
||||||
|
traversal escapes are rejected, and `_allowed_paths` restrictions fail closed when invalid.
|
||||||
|
- Preserve per-path locking, encoding detection, byte limits, truncation behavior, and optimistic
|
||||||
|
`expected_mtime` checks when modifying file operations.
|
||||||
|
- Do not bypass the existing file steps or stores in a way that weakens workspace containment.
|
||||||
|
- Never write test state into the repository's `.reme/`; use `tmp_path` or another isolated workspace.
|
||||||
|
- Do not delete or rewrite user memory to repair an index or make a test pass. Rebuild derived state from source files
|
||||||
|
instead.
|
||||||
|
|
||||||
## Validation
|
## Validation
|
||||||
|
|
||||||
Use the narrowest useful check while iterating, then broaden it according to risk.
|
Use the narrowest useful check while iterating, then broaden it according to risk.
|
||||||
|
|
||||||
Run a focused test:
|
Focused test:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pytest tests/unit/path/to/test_file.py -v
|
pytest tests/unit/path/to/test_file.py -v
|
||||||
```
|
```
|
||||||
|
|
||||||
Run the main unit suite:
|
Main unit suite:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pytest tests/unit -v --tb=long -s --log-cli-level=WARNING
|
pytest tests/unit -v --tb=long -s --log-cli-level=WARNING
|
||||||
```
|
```
|
||||||
|
|
||||||
Run repository formatting and lint checks when the change warrants it:
|
Repository formatting and lint checks:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pre-commit run --all-files
|
pre-commit run --all-files
|
||||||
```
|
```
|
||||||
|
|
||||||
Formatting and lint configuration is authoritative. Python code currently uses a maximum
|
Black and Flake8 use a 120-character line limit and Python 3.11 formatting; Pylint is also run by pre-commit. If
|
||||||
line length of 120 for Black and Flake8, with Pylint also run by pre-commit.
|
`reme_studio/` changes, use its Node 22.13+ scripts and run the proportionate checks from that directory, such as
|
||||||
|
`npm run format:check`, `npm run lint`, or `npm test`.
|
||||||
|
|
||||||
Integration tests may contact real services and require credentials such as
|
Integration tests may contact real model providers, services, or agent subprocesses and can require credentials. Do not
|
||||||
`LLM_API_KEY` or `EMBEDDING_API_KEY`. Do not run credentialed or externally mutating tests
|
run credentialed or externally mutating tests automatically; run them only when the task requires them and the necessary
|
||||||
automatically. Run them only when the task requires them and the user has supplied or
|
environment has been supplied or authorized. Mock network, model, and subprocess boundaries in unit tests.
|
||||||
authorized the necessary environment.
|
|
||||||
|
|
||||||
## Coding and Test Conventions
|
If documentation or the documentation theme changes, run `npm test` and `npm run build` from `github-pages/`. The Job
|
||||||
|
reference is generated from `reme/config/default.yaml`; do not edit generated pages directly.
|
||||||
|
|
||||||
- Target Python 3.11+ and follow the surrounding typing and async style.
|
## Change Guardrails
|
||||||
- Steps are stateless. If a step needs to persist state, store it in
|
|
||||||
`self.app_context.metadata` rather than on the step instance.
|
|
||||||
- Keep public schemas explicit and backward-compatible where practical.
|
|
||||||
- Close async clients, services, tasks, and other lifecycle resources deterministically.
|
|
||||||
- Prefer clear failures over silently falling back to corrupt or ambiguous state.
|
|
||||||
- Keep indexes and caches derivable from user-owned source files.
|
|
||||||
- Use `tmp_path` or another isolated temporary workspace in tests.
|
|
||||||
- Never write test state into the repository's `.reme/` directory.
|
|
||||||
- Mock network or model boundaries in unit tests.
|
|
||||||
- Do not commit `.env` files, credentials, runtime memory, logs, indexes, or caches.
|
|
||||||
|
|
||||||
## Documentation Boundaries
|
|
||||||
|
|
||||||
ReMe's local docs and the deployed documentation site have separate responsibilities.
|
|
||||||
|
|
||||||
- Keep `docs/` focused on content and assets used by `README.md` and `README_ZH.md`.
|
|
||||||
- Preserve README-linked pages under `docs/en/` and `docs/zh/`, including their relative
|
|
||||||
paths, unless the README is updated in the same change.
|
|
||||||
- Keep README-required images under `docs/figure/`.
|
|
||||||
- Keep the README's main documentation index pointed at `docs.agentscope.io` or the
|
|
||||||
`agentscope-ai/docs` repository, following the existing link style.
|
|
||||||
- Do not treat local README-supporting pages as the source for the deployed website.
|
|
||||||
|
|
||||||
The separate `agentscope-ai/docs` repository owns website content, navigation, versioning,
|
|
||||||
and deployment. Public ReMe pages live there under `reme/<version>/`. Make website changes
|
|
||||||
in that repository and follow its existing version-management conventions.
|
|
||||||
|
|
||||||
Do not add website build configuration or deployment workflows to ReMe unless the task
|
|
||||||
explicitly changes this repository boundary.
|
|
||||||
|
|
||||||
## Agent Guardrails
|
|
||||||
|
|
||||||
- Preserve unrelated user changes in a dirty working tree.
|
- Preserve unrelated user changes in a dirty working tree.
|
||||||
- Do not edit generated output when the source can be changed instead.
|
- Make the smallest coherent change and avoid unrelated cleanup or broad refactors.
|
||||||
- Do not delete or rewrite user data to make a test pass.
|
- Do not edit generated output when the source can be changed instead. The publish workflow builds
|
||||||
- Avoid broad refactors unless they are necessary for the requested outcome.
|
`reme_studio/dist-static` and stages it under `reme_studio/src/reme_studio/static`; change `reme_studio/` source for
|
||||||
- Do not introduce dependencies without a concrete need and repository-level justification.
|
frontend work.
|
||||||
- Treat network access, real credentials, and external service mutations as opt-in.
|
- Do not silently change CLI flags, configuration keys, workspace layouts, serialized schemas, endpoint shapes,
|
||||||
- State which validations passed and which were not run in the final handoff.
|
streaming termination, or service interfaces. Preserve compatibility where practical and document intentional
|
||||||
|
migrations.
|
||||||
|
- Do not introduce dependencies without a concrete repository-level need.
|
||||||
|
- Do not commit `.env` files, credentials, runtime memory, logs, indexes, caches, benchmark outputs, or generated
|
||||||
|
Studio distributions.
|
||||||
|
- State which validations passed and which relevant checks were not run in the final handoff.
|
||||||
|
|
||||||
If a requirement is ambiguous, first infer intent from nearby code, tests, and schemas. Ask
|
If a requirement is ambiguous, infer intent from nearby code, schemas, defaults, and tests. Ask the user only when the
|
||||||
the user only when the remaining choice would materially alter a public contract, user data,
|
remaining choice would materially alter a public contract, user data, or an external system.
|
||||||
or external system.
|
|
||||||
|
|
|
||||||
54
Dockerfile
Normal file
54
Dockerfile
Normal file
|
|
@ -0,0 +1,54 @@
|
||||||
|
# syntax=docker/dockerfile:1
|
||||||
|
|
||||||
|
FROM node:22-bookworm-slim AS studio-builder
|
||||||
|
WORKDIR /build/reme_studio
|
||||||
|
COPY reme_studio/package.json reme_studio/package-lock.json ./
|
||||||
|
RUN --mount=type=cache,target=/root/.npm npm ci
|
||||||
|
COPY reme_studio/ ./
|
||||||
|
RUN npm run build:static && test -f dist-static/index.html
|
||||||
|
|
||||||
|
FROM python:3.11-slim-bookworm AS python-builder
|
||||||
|
RUN apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends build-essential \
|
||||||
|
&& rm -rf /var/lib/apt/lists/*
|
||||||
|
WORKDIR /build
|
||||||
|
COPY pyproject.toml README.md LICENSE ./
|
||||||
|
COPY reme/ reme/
|
||||||
|
COPY reme_studio/pyproject.toml reme_studio/README.md reme_studio/LICENSE reme_studio/
|
||||||
|
COPY reme_studio/src/ reme_studio/src/
|
||||||
|
COPY --from=studio-builder /build/reme_studio/dist-static/ reme_studio/dist-static/
|
||||||
|
COPY scripts/package_studio.py scripts/package_studio.py
|
||||||
|
RUN python scripts/package_studio.py && python -m venv /opt/venv
|
||||||
|
ENV PATH="/opt/venv/bin:${PATH}"
|
||||||
|
RUN --mount=type=cache,target=/root/.cache/pip \
|
||||||
|
python -m pip install --upgrade pip \
|
||||||
|
&& python -m pip install ./reme_studio ".[core,image-heif]" \
|
||||||
|
&& python -m pip check
|
||||||
|
# Check installed resources away from the checkout, so source files cannot mask
|
||||||
|
# incomplete wheels or a Studio package fetched accidentally from PyPI.
|
||||||
|
WORKDIR /tmp
|
||||||
|
RUN python -I -c "import reme; from reme_studio import static_dir; from reme.config import resolve_app_config; assert (static_dir() / 'index.html').is_file(); assert resolve_app_config(log_config=False)['service']['backend'] == 'http'"
|
||||||
|
|
||||||
|
FROM python:3.11-slim-bookworm AS runtime
|
||||||
|
RUN apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends ca-certificates git libgomp1 libstdc++6 tini tzdata \
|
||||||
|
&& rm -rf /var/lib/apt/lists/* \
|
||||||
|
&& groupadd --gid 1000 reme \
|
||||||
|
&& useradd --uid 1000 --gid reme --no-create-home reme \
|
||||||
|
&& mkdir -p /app /data \
|
||||||
|
&& chown reme:reme /app /data
|
||||||
|
COPY --from=python-builder /opt/venv /opt/venv
|
||||||
|
COPY deploy/docker/reme_container.py /usr/local/lib/reme_container.py
|
||||||
|
ENV PATH="/opt/venv/bin:${PATH}" \
|
||||||
|
PYTHONUNBUFFERED=1 \
|
||||||
|
PYTHONDONTWRITEBYTECODE=1 \
|
||||||
|
HOME=/tmp/reme-home \
|
||||||
|
REME_WORKSPACE_DIR=/data \
|
||||||
|
REME_HOST=0.0.0.0
|
||||||
|
WORKDIR /app
|
||||||
|
USER reme
|
||||||
|
EXPOSE 2333
|
||||||
|
HEALTHCHECK --interval=30s --timeout=5s --start-period=120s --retries=3 \
|
||||||
|
CMD ["python", "/usr/local/lib/reme_container.py", "--healthcheck"]
|
||||||
|
ENTRYPOINT ["/usr/bin/tini", "--", "python", "/usr/local/lib/reme_container.py"]
|
||||||
|
CMD ["start"]
|
||||||
384
README.md
384
README.md
|
|
@ -1,5 +1,5 @@
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
|
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
|
|
@ -8,6 +8,7 @@
|
||||||
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
|
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
|
||||||
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
|
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
|
||||||
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
|
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
|
||||||
|
<a href="https://reme.agentscope.io"><img src="https://img.shields.io/badge/docs-ReMe-blue" alt="Documentation"></a>
|
||||||
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
|
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
|
||||||
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
|
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
|
||||||
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
|
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
|
||||||
|
|
@ -19,49 +20,79 @@
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<strong>An agent memory layer that turns conversations and resources into readable, editable, searchable Markdown memory.</strong><br>
|
<strong>A local-first, self-evolving personal knowledge base for AI agents.</strong><br>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
> Previous versions: [0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
|
> Previous versions: [0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
|
||||||
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
|
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
|
||||||
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
|
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
|
||||||
|
|
||||||
🧠 ReMe is a local-first memory layer for **AI agents**. It turns conversations and resources into file-based long-term
|
## ✨ Why ReMe?
|
||||||
memory, then continuously indexes, links, and consolidates that memory for future recall.
|
|
||||||
|
|
||||||
## ✨ Core Ideas
|
🧠 ReMe turns conversations and resources into readable, editable, searchable, and interconnected Markdown memory. Agents
|
||||||
|
such as QwenPaw and DeepSeek Harness can share the same workspace to retrieve, maintain, and evolve knowledge, while
|
||||||
|
users retain control of the durable files.
|
||||||
|
|
||||||
- **Memory as File**: Markdown files with frontmatter and wikilinks serve as memory nodes that both users and agents can
|
- **Memory as File, File as Memory**: ReMe stores durable memory as ordinary Markdown with frontmatter and wikilinks.
|
||||||
read and write directly.
|
Users and agents can inspect, edit, move, sync, and back it up with familiar tools, while indexes and generated
|
||||||
- **Self-evolving knowledge base**: Auto Memory, Auto Resource, and Auto Dream progressively transform conversations and
|
metadata remain rebuildable.
|
||||||
resources into long-term memories, while automatically building wikilink relationships.
|
- **Self-evolving knowledge base**: ReMe progressively turns conversations and resources into daily notes and long-term
|
||||||
- **Progressive hybrid search**: ReMe combines wikilinks, BM25, and embeddings for hybrid retrieval across keyword
|
knowledge, preserving sources while refining facts, preferences, procedures, and relationships over time.
|
||||||
matching, semantic recall, and relationship expansion.
|
- **Recall is precise and context-aware.** BM25, optional embeddings, and wikilink expansion retrieve relevant
|
||||||
- **Agent-friendly integration**: SKILL.md + CLI integration makes it easy for different agents to read, write,
|
line-level passages and their relationships without loading the entire knowledge base into the agent context.
|
||||||
maintain, and reuse memory.
|
- **One memory workspace works across agents.** Personal assistants, coding agents, and other agent runtimes can share
|
||||||
|
the same local workspace through native integrations, SKILL.md, CLI, HTTP, MCP, or Python APIs.
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="docs/figure/design-philosophy.svg" alt="ReMe Design Philosophy" width="92%">
|
<img src="docs/figure/design-philosophy.svg" alt="ReMe Design Philosophy" width="92%">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
## 🔭 Use Cases
|
## 📰 Latest Updates
|
||||||
|
|
||||||
- **Personal assistants**: Give personal assistants such as
|
- [2026.10] - **[ReMe Studio Playground](https://reme.agentscope.io/studio/?lang=en) is live**: explore example memory
|
||||||
[QwenPaw](https://github.com/agentscope-ai/QwenPaw), [OpenClaw](https://github.com/openclaw/openclaw), and
|
files, edit Markdown, and browse linked memory graphs right in your browser—no installation or backend required.
|
||||||
[Hermes](https://github.com/nousresearch/hermes-agent) a user-editable long-term memory layer.
|
Everyone is welcome to [try it out](https://reme.agentscope.io/studio/?lang=en)!
|
||||||
- **Coding agents**: Preserve coding style, project background, repository decisions, and workflow
|
|
||||||
experience across sessions when integrating with coding agents such as [Claude Code](plugins/reme).
|
|
||||||
- **LLM Wiki**: Turn conversations, notes, and resources into a searchable, traceable, and linked Markdown
|
|
||||||
knowledge base that both users and agents can maintain.
|
|
||||||
- **Self-evolving agents**: Support agents that learn from experience by saving successful paths, failed attempts,
|
|
||||||
reusable procedures, and periodic reflections as memory.
|
|
||||||
|
|
||||||
## 📰 News
|
<p align="center">
|
||||||
|
<a href="https://reme.agentscope.io/studio/?lang=en">
|
||||||
|
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/reme_studio/figures/studio-overview.png" alt="ReMe Studio workspace preview — click to try the Playground" width="480" style="margin: 0 auto;">
|
||||||
|
</a>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
- [2026.09] - **[ReMe Memory Tags](https://reme.agentscope.io/en/blog_20260920) published**: an introduction
|
||||||
|
to file-native entity tags, rebuildable tag indexes, and tag-filtered memory search.
|
||||||
|
- [2026.09] - **[Hermes Agent memory provider](https://reme.agentscope.io/en/integrations/hermes) available**: choose HTTP or embedded
|
||||||
|
mode for automatic recall before model calls and asynchronous `auto_memory` after completed turns. The integration
|
||||||
|
supports Hermes Agent 0.21+ and includes profile-aware background work.
|
||||||
|
- [2026.09] - **[OpenClaw plugin](https://reme.agentscope.io/en/integrations/openclaw) released**: install it from
|
||||||
|
[ClawHub](https://clawhub.ai/agentscope-ai/plugins/reme-openclaw-plugin) or
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-openclaw-plugin) to add native memory recall, automatic
|
||||||
|
conversation capture, and scheduled consolidation to OpenClaw.
|
||||||
|
- [2026.09] - **[DeepSeek Harness plugin](https://reme.agentscope.io/en/integrations/dsh) released**: install it from
|
||||||
|
[Awesome DSH Plugin](https://awesome-dsh-plugin.com/p/agentscope-ai/ReMe--integrations-dsh/) or
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-dsh-plugin) for long-term-memory guidance, `reme_search`,
|
||||||
|
automatic memory, Auto Dream, and ReMe Status.
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary>More updates</summary>
|
||||||
|
|
||||||
|
- [2026.08] - **ReMe blog published**: the [ReMe blog](https://reme.agentscope.io/en/reme-blog) introduces the
|
||||||
|
local-first memory architecture, self-evolving workflows, hybrid search, proactive discovery, and benchmark results.
|
||||||
|
- [2026.08] - **New ReMe ecosystem plugins**: [Daily Paper](https://reme.agentscope.io/en/plugins/daily-paper)
|
||||||
|
discovers and analyzes papers and generates file-native briefs, while
|
||||||
|
[Auto Fin](https://reme.agentscope.io/en/plugins/auto-fin) researches the latest 24 hours of topic-related CLS news
|
||||||
|
and builds traceable reports with local memory. Try them out.
|
||||||
|
- [2026.08] - **Plugin development support released**: use [Plugin Development](https://reme.agentscope.io/en/plugin_development) and
|
||||||
|
[Plugin Management](https://reme.agentscope.io/en/plugin_management) to extend ReMe with Components, Steps, and Jobs. Contributions and
|
||||||
|
new community plugins are welcome.
|
||||||
|
- [2026.08] - ReMe's [experience-driven enhancement method](https://reme.agentscope.io/en/benchmarks/toolmemory) for
|
||||||
|
agent tool use is available on [arXiv:2608.03403](https://arxiv.org/abs/2608.03403).
|
||||||
- [2026.07] - Our
|
- [2026.07] - Our
|
||||||
paper [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
|
paper [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
|
||||||
has been accepted to Findings of ACL 2026.
|
has been accepted to Findings of ACL 2026.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
## 🚀 Quick Start
|
## 🚀 Quick Start
|
||||||
|
|
||||||
### Installation
|
### Installation
|
||||||
|
|
@ -79,9 +110,28 @@ Install from source:
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/agentscope-ai/ReMe.git
|
git clone https://github.com/agentscope-ai/ReMe.git
|
||||||
cd ReMe
|
cd ReMe
|
||||||
pip install -e ".[core]"
|
pip install -e reme_studio -e ".[core]"
|
||||||
|
cd reme_studio
|
||||||
|
npm ci
|
||||||
|
npm run build:static
|
||||||
|
cd ..
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The static build requires Node.js 22.13 or newer and makes Studio available from the source tree.
|
||||||
|
|
||||||
|
### Docker
|
||||||
|
|
||||||
|
With Docker and Compose 2.24.0+, build and start ReMe with the bundled Studio:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p .reme
|
||||||
|
docker compose up --build -d
|
||||||
|
```
|
||||||
|
|
||||||
|
Open <http://127.0.0.1:2333>. The complete workspace persists in `./.reme`. On Linux, set `REME_UID` and `REME_GID` to your
|
||||||
|
user's IDs when they differ from 1000. See [Docker deployment](https://reme.agentscope.io/en/docker) for model credentials,
|
||||||
|
custom paths, published images, and upgrades.
|
||||||
|
|
||||||
### Environment Variables
|
### Environment Variables
|
||||||
|
|
||||||
Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are
|
Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are
|
||||||
|
|
@ -93,7 +143,7 @@ cat > .env <<'EOF'
|
||||||
# EMBEDDING_API_KEY=sk-xxx
|
# EMBEDDING_API_KEY=sk-xxx
|
||||||
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||||
|
|
||||||
# Required for auto_memory, auto_resource, and auto_dream.
|
# Required for auto_memory, auto_resource, auto_dream, and proactive refresh.
|
||||||
LLM_API_KEY=sk-xxx
|
LLM_API_KEY=sk-xxx
|
||||||
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||||
EOF
|
EOF
|
||||||
|
|
@ -105,7 +155,7 @@ Basic file operations, BM25 search, wikilink traversal, and reading proactive to
|
||||||
> To enable embedding-based semantic retrieval, uncomment `components.as_embedding` and
|
> To enable embedding-based semantic retrieval, uncomment `components.as_embedding` and
|
||||||
> `components.embedding_store` in [`reme/config/default.yaml`](reme/config/default.yaml), then change
|
> `components.embedding_store` in [`reme/config/default.yaml`](reme/config/default.yaml), then change
|
||||||
> `components.file_store.default.embedding_store` from `""` to `default`. See the
|
> `components.file_store.default.embedding_store` from `""` to `default`. See the
|
||||||
> [memory search guide](docs/en/memory_search.md) for details.
|
> [memory search guide](https://reme.agentscope.io/en/memory_search) for details.
|
||||||
|
|
||||||
### Start the Service
|
### Start the Service
|
||||||
|
|
||||||
|
|
@ -120,10 +170,10 @@ reme start service.port=8181
|
||||||
# reme start workspace_dir=/tmp/reme-demo service.port=8181
|
# reme start workspace_dir=/tmp/reme-demo service.port=8181
|
||||||
```
|
```
|
||||||
|
|
||||||
After startup, check the service status. If you use a custom port, replace `2333` in the URL below with that port.
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
reme version
|
reme version
|
||||||
|
reme health_check
|
||||||
|
reme help
|
||||||
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
|
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -161,94 +211,26 @@ ReMe stores agent memory as readable Markdown.
|
||||||
Related: [[digest/wiki/memory-as-file.md]]
|
Related: [[digest/wiki/memory-as-file.md]]
|
||||||
```
|
```
|
||||||
|
|
||||||
## 📁 Memory System
|
### ReMe Studio (Optional)
|
||||||
|
|
||||||
> Memory as File, File as Memory.
|
The `core` installation includes Studio. After starting ReMe, open <http://127.0.0.1:2333/> to browse, edit, and search
|
||||||
|
the workspace. To add Studio to a base installation, use `pip install "reme-ai[web]"`. See the
|
||||||
|
[ReMe Studio guide](https://reme.agentscope.io/en/workspace/studio) for source builds, configuration, and development.
|
||||||
|
|
||||||
ReMe treats **memory as files**, progressively processing raw conversations and external resources from `session/` and
|
## 🤝 Use ReMe with Your Agent
|
||||||
`resource/` into `daily/`, then consolidating them into reusable long-term memory nodes under `digest/`.
|
|
||||||
|
|
||||||
### Directory Structure
|
ReMe can run as a local memory service accessed through the CLI, HTTP API, or MCP server, or it can be embedded in the
|
||||||
|
host process through its Python API. Host integrations can add memory guidance, recall, and capture to the agent
|
||||||
|
lifecycle according to the capabilities of each runtime.
|
||||||
|
|
||||||
```text
|
| Agent | Recommended path | Available after integration |
|
||||||
<workspace_dir>/
|
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
|
||||||
├── metadata/ # Persistent system state such as indexes, graphs, and catalogs
|
| **DeepSeek Harness** | Install [`@agentscope-ai/reme-dsh-plugin`](https://reme.agentscope.io/en/integrations/dsh) with `dsh plugin --profile web add @agentscope-ai/reme-dsh-plugin`. | Configurable memory guidance, `reme_search`, automatic turn capture, scheduled Auto Dream, and ReMe Status. |
|
||||||
├── session/ # Raw conversations and agent sessions
|
| **OpenClaw** | Install [`@agentscope-ai/reme-openclaw-plugin`](https://reme.agentscope.io/en/integrations/openclaw) with `openclaw plugins install clawhub:@agentscope-ai/reme-openclaw-plugin`. | Native memory tools, recall before user-triggered runs, and automatic turn capture. |
|
||||||
│ ├── dialog/
|
| **QwenPaw** | Embed ReMe in-process through its Python API. | Reuse the host lifecycle and model config while keeping memory local and file-based. |
|
||||||
│ │ └── <session_id>.jsonl
|
| **Claude Code** | Start the shared streamable HTTP MCP service and install [the ReMe plugin](https://reme.agentscope.io/en/integrations/claude-code). | Semantic, graph, and state recall through MCP, plus asynchronous session capture through a Stop hook. |
|
||||||
│ ├── agentscope/
|
| **Hermes** | Install [the ReMe provider](https://reme.agentscope.io/en/integrations/hermes) and choose HTTP or embedded mode. | Recall before model calls and asynchronous `auto_memory` after each completed turn. |
|
||||||
│ └── claude_code/
|
| **Codex and other CLI agents** | Install or copy the [ReMe Memory skill](skills/reme_memory/SKILL.md). | Search, read, and write memory through the CLI; automatic capture requires host lifecycle integration. |
|
||||||
├── resource/ # External raw materials
|
|
||||||
│ └── YYYY-MM-DD/
|
|
||||||
│ └── <resource>.<ext>
|
|
||||||
├── daily/ # Lightly processed memory: daily facts, conversation summaries, resource readings
|
|
||||||
│ ├── YYYY-MM-DD.md
|
|
||||||
│ └── YYYY-MM-DD/
|
|
||||||
│ ├── <session_event>.md
|
|
||||||
│ ├── <resource_stem>.md
|
|
||||||
│ └── interests.yaml
|
|
||||||
└── digest/ # Long-term memory: personal facts, procedural experience, knowledge nodes
|
|
||||||
├── personal/
|
|
||||||
│ └── {topic/event}.md
|
|
||||||
├── procedure/
|
|
||||||
│ └── {topic/event}.md
|
|
||||||
└── wiki/
|
|
||||||
└── {topic/event}.md
|
|
||||||
```
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<img src="docs/figure/reme-overview.svg" alt="ReMe file-based memory system overview" width="92%">
|
|
||||||
</p>
|
|
||||||
|
|
||||||
## 🧭 Memory Design Philosophy
|
|
||||||
|
|
||||||
> Capture raw dialogs and resources, refine them into long-term preferences, reusable experience, and valuable
|
|
||||||
> knowledge,
|
|
||||||
> while keeping the result editable by humans and agents.
|
|
||||||
|
|
||||||
### Automatic Memory Flow
|
|
||||||
|
|
||||||
ReMe follows a capture → index → consolidate → recall loop. Conversations and resources first become daily memory cards;
|
|
||||||
background jobs keep files searchable; `auto_dream` distills stable knowledge into `digest/`; agents recall memory
|
|
||||||
through search, wikilinks, or proactive topics.
|
|
||||||
|
|
||||||
| Capability | Entry point | What it does | Output |
|
|
||||||
|---------------------------------------------|-------------------------------------------------|-------------------------------------------------------------------------------------------------|---------------------------------------------------------|
|
|
||||||
| [`auto_memory`](docs/en/auto_memory.md) | Agent hook or `reme auto_memory` | Distills useful conversation facts while preserving the raw session. | `session/dialog/*.jsonl`, `daily/<date>/<session>.md` |
|
|
||||||
| [`auto_resource`](docs/en/auto_resource.md) | Resource watcher or `reme auto_resource` | Turns files under `resource/<date>/` into source-linked daily cards. | `daily/<date>/<resource-card>.md` |
|
|
||||||
| [`auto_index`](docs/en/memory_search.md) | Background watcher or `reme reindex` | Maintains chunks, the BM25 index, the wikilink graph, and the optional embedding index. | Searchable `daily/`, `digest/`, and `resource/` content |
|
|
||||||
| [`auto_dream`](docs/en/auto_dream.md) | `dream_cron` or `reme auto_dream` | Consolidates changed daily cards into long-term personal, procedure, and wiki memory. | `digest/**`, `daily/<date>/interests.yaml` |
|
|
||||||
| [`proactive`](docs/en/proactive.md) | `reme proactive` before an agent decides to act | Reads topics generated by `auto_dream`; the host agent decides whether and how to mention them. | Structured topics from `daily/<date>/interests.yaml` |
|
|
||||||
|
|
||||||
<table>
|
|
||||||
<tr>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
|
|
||||||
</td>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
<tr>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
|
|
||||||
</td>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
</table>
|
|
||||||
|
|
||||||
## 🤝 Agent-friendly Integration
|
|
||||||
|
|
||||||
ReMe runs as a local memory service and offers multiple integration paths: CLI, HTTP API, MCP server, and SDK. Different
|
|
||||||
agents can choose the path that fits their runtime while sharing the same local memory workspace.
|
|
||||||
|
|
||||||
| Agents | Recommended path | What works out of the box |
|
|
||||||
|------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------|
|
|
||||||
| **QwenPaw** | Embed ReMe via the Python SDK. | Reuse the app's own lifecycle and model config while keeping memory local and file-based. |
|
|
||||||
| **Claude Code** | Start ReMe as an MCP service and install [plugins/reme](plugins/reme). | MCP recall tools, a `reme-memory` skill, and a Stop hook that records sessions automatically. |
|
|
||||||
| **Other CLI-capable agents (OpenClaw/Hermes/Codex)** | Copy or install [skills/reme_memory/SKILL.md](skills/reme_memory/SKILL.md). | Search/read/write memory and call `auto_memory`, `auto_dream`, and `proactive` via the CLI. |
|
|
||||||
|
|
||||||
<p align="center"><b>Integration demos</b></p>
|
<p align="center"><b>Integration demos</b></p>
|
||||||
|
|
||||||
|
|
@ -278,39 +260,171 @@ agents can choose the path that fits their runtime while sharing the same local
|
||||||
</tr>
|
</tr>
|
||||||
</table>
|
</table>
|
||||||
|
|
||||||
## 🛠️ ReMe Operations
|
## 🧠 How ReMe Works
|
||||||
|
|
||||||
ReMe operates the workspace through a unified job interface exposed by the CLI. Agents usually only need retrieval,
|
> Memory as File, File as Memory.
|
||||||
reading, writing, editing, and automatic memory commands. Lower-level indexing, frontmatter, and file operation commands
|
|
||||||
are mainly for maintenance, debugging, or advanced integration. Run `reme help` for the full job list.
|
ReMe treats **memory as files**, progressively processing filtered conversation source records and external resources
|
||||||
|
from `session/` and `resource/` into `daily/`, then `digest/`. The default workspace is `.reme/` under the current
|
||||||
|
directory; `workspace_dir=...` selects a different user-owned location.
|
||||||
|
|
||||||
|
### Workspace Layout
|
||||||
|
|
||||||
|
```text
|
||||||
|
<workspace_dir>/
|
||||||
|
├── metadata/ # Rebuildable indexes, graphs, catalogs, and caches
|
||||||
|
├── session/ # Conversation source records and agent sessions
|
||||||
|
│ ├── dialog/
|
||||||
|
│ │ └── <session_id>.jsonl # Source messages saved by auto_memory
|
||||||
|
│ └── claude_code/
|
||||||
|
│ └── <session_id>.jsonl # ReMe copy used by auto_memory_cc
|
||||||
|
├── mem_session/ # Generated agent-wrapper sessions/config, not user memory
|
||||||
|
│ ├── agentscope/
|
||||||
|
│ ├── claude_config/
|
||||||
|
│ └── codex/
|
||||||
|
├── resource/ # External raw materials
|
||||||
|
│ ├── <resource>.<ext> # Root-level files enter today's daily layer
|
||||||
|
│ └── YYYY-MM-DD/
|
||||||
|
│ └── <resource>.<ext>
|
||||||
|
├── daily/ # Lightly processed memory: daily facts, conversation summaries, resource readings
|
||||||
|
│ ├── YYYY-MM-DD.md
|
||||||
|
│ └── YYYY-MM-DD/
|
||||||
|
│ ├── <generated_name>.md # Topic-named conversation or resource card
|
||||||
|
│ └── interests.yaml
|
||||||
|
└── digest/ # Long-term memory: personal facts, procedural experience, knowledge nodes
|
||||||
|
├── personal/
|
||||||
|
│ └── {topic/event}.md
|
||||||
|
├── procedure/
|
||||||
|
│ └── {topic/event}.md
|
||||||
|
└── wiki/
|
||||||
|
└── {topic/event}.md
|
||||||
|
```
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img src="docs/figure/reme-overview.svg" alt="ReMe file-based memory system overview" width="92%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
### Memory Lifecycle
|
||||||
|
|
||||||
|
ReMe follows a capture → index → consolidate → recall loop. Workspace files remain the durable source of truth;
|
||||||
|
everything under `metadata/` is rebuildable.
|
||||||
|
|
||||||
|
| Capability | Entry point | What it does | Output |
|
||||||
|
| ------------------------------------------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
|
||||||
|
| [`auto_memory`](https://reme.agentscope.io/en/auto_memory) | Agent hook or `reme auto_memory` | Distills useful conversation facts while preserving a filtered conversation source record. | `session/dialog/*.jsonl`, `daily/<date>/<generated-name>.md` |
|
||||||
|
| [`auto_resource`](https://reme.agentscope.io/en/auto_resource) | Resource watcher or `reme auto_resource` | Turns files under `resource/` into source-linked, content-named daily cards. | `daily/<date>/<resource-card>.md` |
|
||||||
|
| [`auto_index`](https://reme.agentscope.io/en/memory_search) | Background watcher or `reme reindex` | The watcher ingests Markdown from `daily/` and `digest/`; `reindex` only rebuilds BM25 and embeddings from already-ingested chunks. | Searchable chunks, BM25, wikilink graph, and optional vectors |
|
||||||
|
| [`auto_dream`](https://reme.agentscope.io/en/auto_dream) | `dream_cron` or `reme auto_dream` | By default, extracts up to five reusable units from changed files in the latest two-day window, then creates, corroborates, refines, or corrects digest nodes. | `digest/**` |
|
||||||
|
| [`proactive_read`](https://reme.agentscope.io/en/proactive) | `reme proactive_read` before an agent decides to act | Reads topics generated by the independent proactive refresh flow; the host agent decides whether and how to mention them. | Structured topics from `daily/<date>/interests.yaml` |
|
||||||
|
|
||||||
|
<table>
|
||||||
|
<tr>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
|
||||||
|
</td>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
|
||||||
|
</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
|
||||||
|
</td>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
|
||||||
|
</td>
|
||||||
|
</tr>
|
||||||
|
</table>
|
||||||
|
|
||||||
|
Search returns matching chunks with line ranges and bounded wikilink neighbors. Optional vector results are fused with
|
||||||
|
BM25 through reciprocal rank fusion (RRF).
|
||||||
|
|
||||||
|
> [!IMPORTANT]
|
||||||
|
>
|
||||||
|
> `proactive_read` only reads and exposes interest topics produced by proactive refresh. It does not independently browse the web,
|
||||||
|
> send notifications, or rewrite the knowledge base; the host agent decides whether and how to act on a topic.
|
||||||
|
|
||||||
|
## 📊 Benchmarks
|
||||||
|
|
||||||
|
ReMe evaluates multi-session and long-context memory with agentic search-and-read workflows. The figures below are the
|
||||||
|
published reference runs in this repository; model, prompt, dataset, and judging details are documented with each
|
||||||
|
benchmark.
|
||||||
|
|
||||||
|
| Benchmark | Setting | Sample size | Agentic score | Focus |
|
||||||
|
| --------------------------------------------------------------------------- | ------------ | -----------------------: | ------------: | ------------------------------------------------------------------ |
|
||||||
|
| **[LongMemEval cleaned-s](https://reme.agentscope.io/en/benchmarks/longmemeval)** | **Overall** | **500 questions** | **89.4%** | Cross-session retrieval, knowledge updates, and temporal reasoning |
|
||||||
|
| [BEAM](https://reme.agentscope.io/en/benchmarks/beam) | 100K context | 20 cases / 400 questions | 66.1% | Ten types of long-context memory tasks |
|
||||||
|
| [BEAM](https://reme.agentscope.io/en/benchmarks/beam) | 1M context | 35 cases / 700 questions | 65.0% | Ultra-long conversation settings |
|
||||||
|
|
||||||
|
ReMe also achieved a **0.580 PROC score across five user personas** in the repository's
|
||||||
|
[π-Bench evaluation](https://reme.agentscope.io/en/benchmarks/pibench), 2.4% above NanoBot under the same test-model configuration. PROC
|
||||||
|
measures proactive handling of hidden intent, clarification, cross-session preferences and conventions, task
|
||||||
|
dependencies, and underspecified requests.
|
||||||
|
|
||||||
|
## 🧩 Extensions and Plugins
|
||||||
|
|
||||||
|
Plugins are optional Python distributions that contribute Component, Step, or Job backends and configuration. They are
|
||||||
|
installed separately and enabled explicitly by configuration. Daily Paper and Auto Fin are independently packaged
|
||||||
|
plugins; see their documentation for [Daily Paper](https://reme.agentscope.io/en/plugins/daily-paper) and
|
||||||
|
[Auto Fin](https://reme.agentscope.io/en/plugins/auto-fin).
|
||||||
|
|
||||||
|
| Plugin | Capability |
|
||||||
|
| ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
|
||||||
|
| [Daily Paper](https://reme.agentscope.io/en/plugins/daily-paper) | Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief. |
|
||||||
|
| [Auto Fin](https://reme.agentscope.io/en/plugins/auto-fin) | Fetch topic-related CLS news, search ReMe history, and generate wikilink-backed Markdown reports. |
|
||||||
|
|
||||||
|
See [Plugin Management](https://reme.agentscope.io/en/plugin_management) to install, inspect, validate, enable, and uninstall ReMe plugins.
|
||||||
|
|
||||||
|
## 📚 Documentation
|
||||||
|
|
||||||
|
These guides cover the main user workflows and the runtime contracts implemented by the current code.
|
||||||
|
|
||||||
|
| Guide | What you will learn |
|
||||||
|
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
|
||||||
|
| [Quick Start](https://reme.agentscope.io/en/quick_start) | Install ReMe, start the service, and run the first file and memory operations. |
|
||||||
|
| [Configuration](https://reme.agentscope.io/en/configuration) | Configure the workspace, models, Service, Jobs, Components, plugins, and CLI overrides. |
|
||||||
|
| [Services and Deployment](https://reme.agentscope.io/en/services) | Use HTTP, SSE, MCP, and Studio while respecting the default security boundary. |
|
||||||
|
| [Memory as File](https://reme.agentscope.io/en/memory_as_file) | Understand workspace layers, frontmatter, wikilinks, chunks, and the file-as-source-of-truth model. |
|
||||||
|
| [Auto Memory](https://reme.agentscope.io/en/auto_memory) | Preserve source conversations and distill reusable daily memory cards. |
|
||||||
|
| [Auto Resource](https://reme.agentscope.io/en/auto_resource) | Import supported text and image resources as source-linked daily cards. |
|
||||||
|
| [Auto Dream](https://reme.agentscope.io/en/auto_dream) and [Auto Link](https://reme.agentscope.io/en/auto_link) | Consolidate daily notes into evolving digest nodes and readable wikilink relationships. |
|
||||||
|
| [Memory Search](https://reme.agentscope.io/en/memory_search) | Use BM25, optional vectors, RRF fusion, line-range recall, and progressive link expansion. |
|
||||||
|
| [Proactive](https://reme.agentscope.io/en/proactive) | Read interest topics safely and integrate them into a host agent's decision flow. |
|
||||||
|
| [Application Scenarios](https://reme.agentscope.io/en/reme_scene) | Follow concrete financial research, coding-memory, and personal knowledge-base examples. |
|
||||||
|
| [Framework](https://reme.agentscope.io/en/framework) | Understand Application, Job, Step, Component, service, configuration, and lifecycle boundaries. |
|
||||||
|
| [Agent Integrations](https://reme.agentscope.io/en/integrations) | Choose an interface and connect DSH, Claude Code, OpenClaw, Hermes, Codex, or another agent. |
|
||||||
|
| [DSH plugin](https://reme.agentscope.io/en/integrations/dsh) and [Claude Code plugin](https://reme.agentscope.io/en/integrations/claude-code) | Configure host-native recall, automatic capture, consolidation, and diagnostics. |
|
||||||
|
| [CLI and Job API](https://reme.agentscope.io/en/reference/cli) | Learn command syntax and use the generated default Job parameter reference. |
|
||||||
|
| [Operations and Recovery](https://reme.agentscope.io/en/operations) | Diagnose services, maintain indexes, and back up, migrate, or recover a workspace. |
|
||||||
|
| [ReMe Blog](https://reme.agentscope.io/en/reme-blog) | Read the product story, design rationale, examples, and benchmark summary. |
|
||||||
|
|
||||||
|
## 🛠️ Common Commands
|
||||||
|
|
||||||
|
Run `reme help` for the full job list. Common workspace and maintenance commands are:
|
||||||
|
|
||||||
| Command | Purpose |
|
| Command | Purpose |
|
||||||
|-------------------------------------------|----------------------------------------------------------------------------------------|
|
| ----------------------------------------- | --------------------------------------------------------------------------------- |
|
||||||
| `reme start` | Start the local ReMe service. |
|
|
||||||
| `reme version` / `reme health_check` | Check package and component status. |
|
|
||||||
| `reme status` | Show stateful data-component memory estimates and process RSS. |
|
| `reme status` | Show stateful data-component memory estimates and process RSS. |
|
||||||
| [`reme search`](docs/en/memory_search.md) | Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled. |
|
| [`reme search`](https://reme.agentscope.io/en/memory_search) | Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled. |
|
||||||
| `reme read` / `reme write` / `reme edit` | Inspect and maintain Markdown memory files. |
|
| `reme read` / `reme write` / `reme edit` | Inspect and maintain Markdown memory files. |
|
||||||
| `reme auto_memory` | Turn conversation messages into daily memory cards. Requires LLM credentials. |
|
| `reme traverse` / `reme graph_snapshot` | Explore wikilink neighborhoods or the category-rooted digest graph. |
|
||||||
| `reme auto_resource` | Interpret files under `resource/` into daily resource cards. Requires LLM credentials. |
|
| `reme chat` | Stream a read-only, workspace-aware agent conversation. Requires LLM credentials. |
|
||||||
| `reme auto_dream` / `reme proactive` | Consolidate daily memory into long-term digest and surface topics worth attention. |
|
| `reme reindex` | Rebuild BM25 and embedding indexes from already-ingested chunks. |
|
||||||
| `reme reindex` | Rebuild search and wikilink indexes from existing files. |
|
|
||||||
|
|
||||||
## 🤝 Community and Support
|
## 🤝 Community and Contributing
|
||||||
|
|
||||||
- **Issues and requests**: Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) first. If there is no
|
- **Issues, requests, and help**: Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) first. If there is no
|
||||||
related discussion, open a new issue with background, expected behavior, and impact scope.
|
related discussion, open one with the background, expected behavior, and impact scope.
|
||||||
- **Code contributions**: Before making changes, read
|
- **Code contributions**: Before making changes, read the repository's
|
||||||
the [contribution guide](https://docs.agentscope.io/reme/stable/en/contributing). Source,
|
[contribution guide](https://reme.agentscope.io/en/contributing). Source, schemas, and tests are the authoritative architecture and
|
||||||
schemas, and tests are the authoritative architecture and extension guide.
|
extension guide.
|
||||||
- **Documentation contributions**: Submit user-facing documentation changes to the
|
- **Documentation contributions**: Update the canonical files under `docs/en/`, `docs/zh/`, or the relevant package
|
||||||
[unified documentation repository](https://github.com/agentscope-ai/docs) under `reme/<version>/{en,zh}/`.
|
directory in this repository. The documentation site is generated from these files.
|
||||||
- **Commit convention**: Conventional Commits are recommended, for example `feat(search): add link expansion option` or
|
- **Commit convention**: Conventional Commits are recommended, for example `feat(search): add link expansion option` or
|
||||||
`docs(zh): update quick start`.
|
`docs(zh): update quick start`.
|
||||||
- **Pre-submit checks**: Before submitting a PR, try to run `pre-commit run --all-files` and `pytest`. If tests that
|
- **Pre-submit checks**: Before submitting a PR, try to run `pre-commit run --all-files` and `pytest`. If tests that
|
||||||
depend on LLMs, embeddings, or external services cannot run, explain that in the PR.
|
depend on LLMs, embeddings, or external services cannot run, explain that in the PR.
|
||||||
- **Get help**: Use [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues) for bugs and feature requests. Project
|
- **Documentation**: Visit [reme.agentscope.io](https://reme.agentscope.io).
|
||||||
documentation is available at [https://docs.agentscope.io/](https://docs.agentscope.io/reme/stable/en/).
|
|
||||||
|
|
||||||
### Contributors
|
### Contributors
|
||||||
|
|
||||||
|
|
|
||||||
354
README_ZH.md
354
README_ZH.md
|
|
@ -1,5 +1,5 @@
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
|
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
|
|
@ -8,6 +8,7 @@
|
||||||
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
|
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
|
||||||
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
|
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
|
||||||
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
|
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
|
||||||
|
<a href="https://reme.agentscope.io"><img src="https://img.shields.io/badge/docs-ReMe-blue" alt="文档"></a>
|
||||||
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
|
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
|
||||||
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
|
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
|
||||||
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
|
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
|
||||||
|
|
@ -19,42 +20,71 @@
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<strong>一个将对话和资料转化为可读、可编辑、可检索 Markdown 记忆的 Agent 记忆层。</strong><br>
|
<strong>面向 AI Agent 的 local-first 自进化个人知识库。</strong><br>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
> 历史版本:[0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
|
> 历史版本:[0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
|
||||||
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
|
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
|
||||||
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
|
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
|
||||||
|
|
||||||
🧠 ReMe 是一个面向 **AI 智能体** 的 local-first 记忆层。它把对话和资料沉淀为文件化长期记忆,并持续完成索引、链接和整理,让后续
|
## ✨ 为什么选择 ReMe?
|
||||||
Agent 能够可靠召回。
|
|
||||||
|
|
||||||
## ✨ 核心创新
|
🧠 ReMe 将对话和资料持续沉淀为可读、可编辑、可检索、相互链接的 Markdown 记忆。QwenPaw、DeepSeek Harness 等 Agent
|
||||||
|
可以共享同一个 workspace,共同检索、维护和演化知识,而持久文件始终由用户掌控。
|
||||||
|
|
||||||
- **Memory as File**:以带 frontmatter 和 wikilink 的 Markdown 作为记忆节点,让用户和 Agent 都能直接读写。
|
- **Memory as File, File as Memory**:ReMe 使用带 frontmatter 和 wikilink 的普通 Markdown 保存持久记忆。用户和 Agent
|
||||||
- **自进化知识库**:通过 Auto Memory、Auto Resource 和 Auto Dream,把对话与资料逐步加工为长期记忆,并自动建立 wikilink 关系。
|
都可以使用熟悉的工具查看、编辑、移动、同步和备份;索引及生成的元数据均可重建。
|
||||||
- **渐进式混合搜索**:融合 wikilink、BM25 和 embedding,支持从关键词匹配到语义召回、关系扩展的混合检索。
|
- **自进化知识库**:ReMe 将对话和资料逐步加工为 daily note 与长期知识,在保留来源的同时,持续提炼事实、偏好、
|
||||||
- **Agent 友好集成**:通过 SKILL.md + CLI 接入,方便不同 Agent 读写、维护与复用记忆。
|
流程经验及其关系。
|
||||||
|
- **精准召回所需上下文。** ReMe 结合 BM25、可选 embedding 和 wikilink 展开,召回带行号的相关片段及其关系,无需把整个知识库塞入
|
||||||
|
Agent 上下文。
|
||||||
|
- **一个 workspace,可供不同 Agent 共同使用。** 个人助理、coding agent 和其他 Agent runtime 可以通过原生集成、SKILL.md、CLI、
|
||||||
|
HTTP、MCP 或 Python API 共享同一个本地记忆空间。
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="docs/figure/design-philosophy.svg" alt="ReMe 设计理念" width="92%">
|
<img src="docs/figure/design-philosophy.svg" alt="ReMe 设计理念" width="92%">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
## 🔭 适用场景
|
## 📰 最新动态
|
||||||
|
|
||||||
- **Personal assistants**:为 [QwenPaw](https://github.com/agentscope-ai/QwenPaw)、
|
- [2026.10] - **[ReMe Studio Playground](https://reme.agentscope.io/studio/?lang=zh) 上线**:无需安装或启动后端,
|
||||||
[OpenClaw](https://github.com/openclaw/openclaw)、[Hermes](https://github.com/nousresearch/hermes-agent)
|
即可在浏览器中浏览示例记忆文件、编辑 Markdown、探索记忆关联图谱。欢迎大家[来体验](https://reme.agentscope.io/studio/?lang=zh)!
|
||||||
等个人助理提供用户可编辑的长期记忆层。
|
|
||||||
- **Coding agents**:在接入 [Claude Code](plugins/reme) 等 coding agent 时,跨会话保留代码风格、项目背景、仓库决策和流程经验。
|
|
||||||
- **LLM Wiki**:把对话、笔记和资料转化为可检索、可追溯、可链接的 Markdown 知识库,由用户和 Agent 共同维护。
|
|
||||||
- **Self-evolving agents**:帮助 Agent 从经验中学习,把成功路径、失败尝试、可复用流程和阶段性反思沉淀为记忆。
|
|
||||||
|
|
||||||
## 📰 新闻
|
<p align="center">
|
||||||
|
<a href="https://reme.agentscope.io/studio/?lang=zh">
|
||||||
|
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/reme_studio/figures/studio-overview.png" alt="ReMe Studio 工作区预览,点击体验 Playground" width="480" style="margin: 0 auto;">
|
||||||
|
</a>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
- [2026.09] - **[给记忆加上“标签”](https://reme.agentscope.io/zh/blog_20260920)发布**:介绍基于 Markdown 的实体标签、
|
||||||
|
可重建 Tag Index 与标签过滤检索。
|
||||||
|
- [2026.09] - **[Hermes Agent 记忆 Provider](https://reme.agentscope.io/zh/integrations/hermes) 已可使用**:支持 HTTP 和 Embedded
|
||||||
|
两种模式,在模型调用前自动召回、每轮对话结束后异步执行 `auto_memory`。集成支持 Hermes Agent 0.21 及以上版本,后台任务也会继承当前 profile 上下文。
|
||||||
|
- [2026.09] - **[OpenClaw 插件](https://reme.agentscope.io/zh/integrations/openclaw) 发布**:可通过
|
||||||
|
[ClawHub](https://clawhub.ai/agentscope-ai/plugins/reme-openclaw-plugin) 或
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-openclaw-plugin) 安装,为 OpenClaw 提供原生记忆召回、自动对话捕获和定时整理能力。
|
||||||
|
- [2026.09] - **[DeepSeek Harness 插件](https://reme.agentscope.io/zh/integrations/dsh) 发布**:可通过
|
||||||
|
[Awesome DSH Plugin](https://awesome-dsh-plugin.com/p/agentscope-ai/ReMe--integrations-dsh/) 或
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-dsh-plugin) 安装,提供长期记忆指引、`reme_search`、自动记忆、Auto Dream 和 ReMe Status。
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary>更多更新</summary>
|
||||||
|
|
||||||
|
- [2026.08] - **ReMe 博客发布**:[ReMe 博客](https://reme.agentscope.io/zh/reme-blog) 系统介绍了本地优先的记忆架构、
|
||||||
|
自进化工作流、混合检索、主动发现与评测结果。
|
||||||
|
- [2026.08] - **新增 ReMe 生态插件**:[每日论文](https://reme.agentscope.io/zh/plugins/daily-paper) 可自动发现、解析论文并生成文件化简报;
|
||||||
|
[Auto Fin](https://reme.agentscope.io/zh/plugins/auto-fin) 可研究最近 24 小时的主题相关财联社新闻,并结合本地记忆构建可追溯报告。欢迎体验。
|
||||||
|
- [2026.08] - **插件开发能力上线**:参考 [插件开发](https://reme.agentscope.io/zh/plugin_development) 与 [插件管理](https://reme.agentscope.io/zh/plugin_management),
|
||||||
|
为 ReMe 扩展 Component、Step 和 Job;欢迎开发并分享你的插件。
|
||||||
|
- [2026.08] - 基于 ReMe 的智能体工具使用
|
||||||
|
[经验驱动增强方法](https://reme.agentscope.io/zh/benchmarks/toolmemory) 已发布,见
|
||||||
|
[arXiv:2608.03403](https://arxiv.org/abs/2608.03403)。
|
||||||
- [2026.07] -
|
- [2026.07] -
|
||||||
我们的论文 [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
|
我们的论文 [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
|
||||||
已被 Findings of ACL 2026 接收。
|
已被 Findings of ACL 2026 接收。
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
## 🚀 快速开始
|
## 🚀 快速开始
|
||||||
|
|
||||||
### 安装
|
### 安装
|
||||||
|
|
@ -72,12 +102,30 @@ pip install "reme-ai[core]"
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/agentscope-ai/ReMe.git
|
git clone https://github.com/agentscope-ai/ReMe.git
|
||||||
cd ReMe
|
cd ReMe
|
||||||
pip install -e ".[core]"
|
pip install -e reme_studio -e ".[core]"
|
||||||
|
cd reme_studio
|
||||||
|
npm ci
|
||||||
|
npm run build:static
|
||||||
|
cd ..
|
||||||
```
|
```
|
||||||
|
|
||||||
### 环境变量
|
静态构建要求 Node.js 22.13 或更高版本,并让源码安装可以直接使用 Studio。
|
||||||
|
|
||||||
如果需要 LLM 驱动的记忆演化或 embedding 检索,可以配置环境变量。embedding 默认关闭,因此默认配置不会启动
|
### Docker
|
||||||
|
|
||||||
|
使用 Docker 和 Compose 2.24.0+,构建并启动包含 Studio 的 ReMe:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p .reme
|
||||||
|
docker compose up --build -d
|
||||||
|
```
|
||||||
|
|
||||||
|
打开 <http://127.0.0.1:2333>,完整工作区保存在宿主机的 `./.reme`。Linux 用户的 UID/GID 不是 1000 时,请设置对应的
|
||||||
|
`REME_UID` 和 `REME_GID`。模型凭证、自定义路径、发布镜像和升级方式见 [Docker 部署](https://reme.agentscope.io/zh/docker)。
|
||||||
|
|
||||||
|
### 环境变量配置
|
||||||
|
|
||||||
|
如果需要 LLM 驱动的记忆演化或 embedding 检索,请在启动服务前配置环境变量。embedding 默认关闭,因此默认配置不会启动
|
||||||
embedding 模型,也不需要 embedding API key。
|
embedding 模型,也不需要 embedding API key。
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
|
@ -86,19 +134,19 @@ cat > .env <<'EOF'
|
||||||
# EMBEDDING_API_KEY=sk-xxx
|
# EMBEDDING_API_KEY=sk-xxx
|
||||||
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||||
|
|
||||||
# 必须:auto_memory、auto_resource 和 auto_dream 需要 LLM。
|
# 必须:auto_memory、auto_resource、auto_dream 和 proactive refresh 需要 LLM。
|
||||||
LLM_API_KEY=sk-xxx
|
LLM_API_KEY=sk-xxx
|
||||||
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||||
EOF
|
EOF
|
||||||
```
|
```
|
||||||
|
|
||||||
基础文件读写、BM25 检索、wikilink 遍历和 proactive topics 读取可以先不配置 LLM 凭证。
|
基础文件读写、BM25 检索、wikilink 遍历和 proactive topics 读取不需要 LLM 凭证,可以跳过此步骤并直接启动服务。
|
||||||
|
|
||||||
> [!NOTE]
|
> [!NOTE]
|
||||||
> 如需启用基于 embedding 的语义检索,请取消 [`reme/config/default.yaml`](reme/config/default.yaml) 中
|
> 如需启用基于 embedding 的语义检索,请取消 [`reme/config/default.yaml`](reme/config/default.yaml) 中
|
||||||
> `components.as_embedding` 和 `components.embedding_store` 的注释,并将
|
> `components.as_embedding` 和 `components.embedding_store` 的注释,并将
|
||||||
> `components.file_store.default.embedding_store` 从 `""` 改为 `default`。完整说明见
|
> `components.file_store.default.embedding_store` 从 `""` 改为 `default`。完整说明见
|
||||||
> [记忆检索文档](docs/zh/memory_search.md)。
|
> [记忆检索文档](https://reme.agentscope.io/zh/memory_search)。
|
||||||
|
|
||||||
### 启动服务
|
### 启动服务
|
||||||
|
|
||||||
|
|
@ -113,10 +161,10 @@ reme start service.port=8181
|
||||||
# reme start workspace_dir=/tmp/reme-demo service.port=8181
|
# reme start workspace_dir=/tmp/reme-demo service.port=8181
|
||||||
```
|
```
|
||||||
|
|
||||||
启动后可以检查服务状态;如果使用了自定义端口,请将下面 URL 中的 `2333` 替换为对应端口。
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
reme version
|
reme version
|
||||||
|
reme health_check
|
||||||
|
reme help
|
||||||
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
|
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -154,91 +202,25 @@ ReMe 会把 Agent 记忆保存为可读的 Markdown。
|
||||||
相关链接:[[digest/wiki/memory-as-file.md]]
|
相关链接:[[digest/wiki/memory-as-file.md]]
|
||||||
```
|
```
|
||||||
|
|
||||||
## 📁 记忆系统
|
### ReMe Studio(可选)
|
||||||
|
|
||||||
> Memory as File, File as Memory.
|
上面的 `core` 安装已包含 Studio。启动 ReMe 后,打开 <http://127.0.0.1:2333/> 即可浏览、编辑和搜索 workspace。
|
||||||
|
如需为基础安装单独添加 Studio,可使用 `pip install "reme-ai[web]"`。源码构建、配置和开发说明见
|
||||||
|
[ReMe Studio 指南](https://reme.agentscope.io/zh/workspace/studio)。
|
||||||
|
|
||||||
ReMe 将**记忆视为文件**,让原始对话和外部资料从 `session/`、`resource/` 渐进加工到 `daily/`,再沉淀为 `digest/`
|
## 🤝 将 ReMe 接入你的 Agent
|
||||||
中可长期复用的知识节点。
|
|
||||||
|
|
||||||
### 目录结构
|
ReMe 既可以作为本地记忆服务,通过 CLI、HTTP API 或 MCP server 接入,也可以通过 Python API 嵌入宿主进程。宿主集成可根据不同
|
||||||
|
runtime 的能力,将记忆指引、召回和捕获接入 Agent 生命周期。
|
||||||
|
|
||||||
```text
|
| Agent | 推荐接入方式 | 接入后能力 |
|
||||||
<workspace_dir>/
|
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
|
||||||
├── metadata/ # 系统索引、图谱、catalog 等持久状态
|
| **DeepSeek Harness** | 使用 `dsh plugin --profile web add @agentscope-ai/reme-dsh-plugin` 安装 [`@agentscope-ai/reme-dsh-plugin`](https://reme.agentscope.io/zh/integrations/dsh)。 | 可配置记忆指引、`reme_search`、自动对话捕获、定时 Auto Dream 和 ReMe Status。 |
|
||||||
├── session/ # 原始对话和 Agent session
|
| **OpenClaw** | 使用 `openclaw plugins install clawhub:@agentscope-ai/reme-openclaw-plugin` 安装 [`@agentscope-ai/reme-openclaw-plugin`](https://reme.agentscope.io/zh/integrations/openclaw)。 | 原生记忆工具、用户触发运行前召回和自动对话捕获。 |
|
||||||
│ ├── dialog/
|
| **QwenPaw** | 通过 Python API 在进程内嵌入 ReMe。 | 复用宿主生命周期和模型配置,同时保持记忆本地、文件化。 |
|
||||||
│ │ └── <session_id>.jsonl
|
| **Claude Code** | 启动共享的 streamable HTTP MCP service,并安装 [ReMe 插件](https://reme.agentscope.io/zh/integrations/claude-code)。 | 通过 MCP 进行语义、图关系和状态召回,并由 Stop Hook 异步捕获会话。 |
|
||||||
│ ├── agentscope/
|
| **Hermes** | 安装 [ReMe provider](https://reme.agentscope.io/zh/integrations/hermes),并选择 HTTP 或 Embedded 模式。 | 模型调用前召回,每轮对话完成后异步执行 `auto_memory`。 |
|
||||||
│ └── claude_code/
|
| **Codex 及其他 CLI Agent** | 安装或复制 [ReMe Memory skill](skills/reme_memory/SKILL.md)。 | 通过 CLI 搜索、读取和写入记忆;自动捕获需要显式接入宿主生命周期。 |
|
||||||
├── resource/ # 外部原始材料
|
|
||||||
│ └── YYYY-MM-DD/
|
|
||||||
│ └── <resource>.<ext>
|
|
||||||
├── daily/ # 浅加工记忆:当天事实、对话摘要、资源解读
|
|
||||||
│ ├── YYYY-MM-DD.md
|
|
||||||
│ └── YYYY-MM-DD/
|
|
||||||
│ ├── <session_event>.md
|
|
||||||
│ ├── <resource_stem>.md
|
|
||||||
│ └── interests.yaml
|
|
||||||
└── digest/ # 长期记忆:个人事实、流程经验、知识节点
|
|
||||||
├── personal/
|
|
||||||
│ └── {topic/event}.md
|
|
||||||
├── procedure/
|
|
||||||
│ └── {topic/event}.md
|
|
||||||
└── wiki/
|
|
||||||
└── {topic/event}.md
|
|
||||||
```
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<img src="docs/figure/reme-overview.svg" alt="ReMe 文件化记忆系统总览" width="92%">
|
|
||||||
</p>
|
|
||||||
|
|
||||||
## 🧭 记忆设计理念
|
|
||||||
|
|
||||||
> 捕获原始对话和资料,将其整理为长期偏好、可复用经验和有价值的知识,并让结果始终能被用户和 Agent 直接编辑。
|
|
||||||
|
|
||||||
### 自动记忆流程
|
|
||||||
|
|
||||||
ReMe 遵循 capture → index → consolidate → recall 的循环。对话和资料先变成 daily 记忆卡片;后台任务保持文件可检索;
|
|
||||||
`auto_dream` 将稳定知识沉淀到 `digest/`;Agent 再通过搜索、wikilink 或 proactive topics 召回记忆。
|
|
||||||
|
|
||||||
| 能力 | 入口 | 作用 | 输出 |
|
|
||||||
|---------------------------------------------|----------------------------------|----------------------------------------------------|------------------------------------------------------|
|
|
||||||
| [`auto_memory`](docs/zh/auto_memory.md) | Agent hook 或 `reme auto_memory` | 提炼有长期价值的对话事实,同时保留原始 session。 | `session/dialog/*.jsonl`、`daily/<date>/<session>.md` |
|
|
||||||
| [`auto_resource`](docs/zh/auto_resource.md) | 资源监听或 `reme auto_resource` | 将 `resource/<date>/` 下的文件转为带来源链接的 daily 卡片。 | `daily/<date>/<resource-card>.md` |
|
|
||||||
| [`auto_index`](docs/zh/memory_search.md) | 后台监听或 `reme reindex` | 维护 chunks、BM25 索引、wikilink 图谱及可选的 embedding 索引。 | 可检索的 `daily/`、`digest/`、`resource/` 内容 |
|
|
||||||
| [`auto_dream`](docs/zh/auto_dream.md) | `dream_cron` 或 `reme auto_dream` | 将变化的 daily 卡片整理为长期 personal、procedure 和 wiki 记忆。 | `digest/**`、`daily/<date>/interests.yaml` |
|
|
||||||
| [`proactive`](docs/zh/proactive.md) | Agent 决定主动行动前调用 `reme proactive` | 读取 `auto_dream` 生成的 topics;是否以及如何提醒用户由宿主 Agent 决定。 | 来自 `daily/<date>/interests.yaml` 的结构化 topics |
|
|
||||||
|
|
||||||
<table>
|
|
||||||
<tr>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
|
|
||||||
</td>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
<tr>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
|
|
||||||
</td>
|
|
||||||
<td align="center" width="50%">
|
|
||||||
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
</table>
|
|
||||||
|
|
||||||
## 🤝 Agent-friendly Integration
|
|
||||||
|
|
||||||
ReMe 作为本地记忆服务运行,并提供 CLI、HTTP API、MCP server 和 SDK 等多种接入方式。不同 Agent 可以选择适合自身 runtime
|
|
||||||
的路径,同时共享同一个本地 memory workspace。
|
|
||||||
|
|
||||||
| Agent | 推荐接入方式 | 开箱可用能力 |
|
|
||||||
|------------------------------------------------------|-------------------------------------------------------------------|-----------------------------------------------------------------|
|
|
||||||
| **QwenPaw** | 通过 Python SDK 嵌入 ReMe。 | 复用应用自身生命周期和模型配置,同时保持 memory 本地、文件化。 |
|
|
||||||
| **Claude Code** | 以 MCP service 启动 ReMe,并安装 [plugins/reme](plugins/reme)。 | MCP recall tools、`reme-memory` skill,以及自动记录会话的 Stop hook。 |
|
|
||||||
| **Other CLI-capable agents (OpenClaw/Hermes/Codex)** | 复制或安装 [skills/reme_memory/SKILL.md](skills/reme_memory/SKILL.md)。 | 通过 CLI 搜索/读取/写入记忆,并调用 `auto_memory`、`auto_dream` 和 `proactive`。 |
|
|
||||||
|
|
||||||
<p align="center"><b>集成演示</b></p>
|
<p align="center"><b>集成演示</b></p>
|
||||||
|
|
||||||
|
|
@ -268,36 +250,160 @@ ReMe 作为本地记忆服务运行,并提供 CLI、HTTP API、MCP server 和
|
||||||
</tr>
|
</tr>
|
||||||
</table>
|
</table>
|
||||||
|
|
||||||
## 🛠️ ReMe Operations
|
## 🧠 ReMe 如何工作
|
||||||
|
|
||||||
ReMe 通过 CLI 暴露的统一 job interface 操作 workspace。Agent 通常只需要使用检索、读取、写入、编辑和自动记忆相关命令;更底层的索引、
|
> Memory as File, File as Memory.
|
||||||
frontmatter 和文件操作接口主要用于维护、调试或高级集成。完整 job 列表可以运行 `reme help` 查看。
|
|
||||||
|
ReMe 将 **记忆视为文件**,让过滤后的对话来源记录和外部资料从 `session/`、`resource/` 渐进加工到 `daily/`,再沉淀为
|
||||||
|
`digest/`。默认 workspace 是当前目录下的 `.reme/`;可通过 `workspace_dir=...` 选择其他由用户控制的位置。
|
||||||
|
|
||||||
|
### Workspace 结构
|
||||||
|
|
||||||
|
```text
|
||||||
|
<workspace_dir>/
|
||||||
|
├── metadata/ # 可重建的索引、图谱、catalog 和缓存
|
||||||
|
├── session/ # 对话来源记录和 Agent session
|
||||||
|
│ ├── dialog/
|
||||||
|
│ │ └── <session_id>.jsonl # auto_memory 保存的来源消息
|
||||||
|
│ └── claude_code/
|
||||||
|
│ └── <session_id>.jsonl # auto_memory_cc 使用的 ReMe 副本
|
||||||
|
├── mem_session/ # Agent wrapper 生成的 session/配置,不是用户记忆
|
||||||
|
│ ├── agentscope/
|
||||||
|
│ ├── claude_config/
|
||||||
|
│ └── codex/
|
||||||
|
├── resource/ # 外部原始材料
|
||||||
|
│ ├── <resource>.<ext> # 根目录文件进入当天 daily 层
|
||||||
|
│ └── YYYY-MM-DD/
|
||||||
|
│ └── <resource>.<ext>
|
||||||
|
├── daily/ # 浅加工记忆:当天事实、对话摘要、资源解读
|
||||||
|
│ ├── YYYY-MM-DD.md
|
||||||
|
│ └── YYYY-MM-DD/
|
||||||
|
│ ├── <generated_name>.md # 按主题命名的对话或资源卡片
|
||||||
|
│ └── interests.yaml
|
||||||
|
└── digest/ # 长期记忆:个人事实、流程经验、知识节点
|
||||||
|
├── personal/
|
||||||
|
│ └── {topic/event}.md
|
||||||
|
├── procedure/
|
||||||
|
│ └── {topic/event}.md
|
||||||
|
└── wiki/
|
||||||
|
└── {topic/event}.md
|
||||||
|
```
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img src="docs/figure/reme-overview.svg" alt="ReMe 文件化记忆系统总览" width="92%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
### 记忆生命周期
|
||||||
|
|
||||||
|
ReMe 遵循 capture → index → consolidate → recall 的循环。workspace 文件是持久化的事实来源,`metadata/` 中的内容均可重建。
|
||||||
|
|
||||||
|
| 能力 | 入口 | 作用 | 输出 |
|
||||||
|
| ------------------------------------------- | ----------------------------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
|
||||||
|
| [`auto_memory`](https://reme.agentscope.io/zh/auto_memory) | Agent hook 或 `reme auto_memory` | 提炼有长期价值的对话事实,同时保留过滤后的对话来源记录。 | `session/dialog/*.jsonl`、`daily/<date>/<generated-name>.md` |
|
||||||
|
| [`auto_resource`](https://reme.agentscope.io/zh/auto_resource) | 资源监听或 `reme auto_resource` | 将 `resource/` 下的文件转为带来源链接、按内容命名的 daily 卡片。 | `daily/<date>/<resource-card>.md` |
|
||||||
|
| [`auto_index`](https://reme.agentscope.io/zh/memory_search) | 后台监听或 `reme reindex` | watcher 摄取 `daily/` 和 `digest/` 中的 Markdown;`reindex` 只基于已摄取的 chunks 重建 BM25 和 Embedding。 | 可检索的 chunks、BM25、wikilink 图谱和可选向量 |
|
||||||
|
| [`auto_dream`](https://reme.agentscope.io/zh/auto_dream) | `dream_cron` 或 `reme auto_dream` | 默认从最近两天内变化的文件中最多提取 5 个可复用 unit,再创建、印证、补充或修正 digest 节点。 | `digest/**` |
|
||||||
|
| [`proactive_read`](https://reme.agentscope.io/zh/proactive) | Agent 决定主动行动前调用 `reme proactive_read` | 读取独立 proactive refresh 流程生成的 topics;是否以及如何提醒用户由宿主 Agent 决定。 | 来自 `daily/<date>/interests.yaml` 的结构化 topics |
|
||||||
|
|
||||||
|
<table>
|
||||||
|
<tr>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
|
||||||
|
</td>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
|
||||||
|
</td>
|
||||||
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
|
||||||
|
</td>
|
||||||
|
<td align="center" width="50%">
|
||||||
|
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
|
||||||
|
</td>
|
||||||
|
</tr>
|
||||||
|
</table>
|
||||||
|
|
||||||
|
搜索返回带行号范围的相关 chunks 和数量受限的 wikilink 邻居;可选向量结果通过 RRF 与 BM25 融合。
|
||||||
|
|
||||||
|
> [!IMPORTANT]
|
||||||
|
>
|
||||||
|
> `proactive_read` 只读取并暴露 proactive refresh 生成的兴趣主题,不会自行联网、发送通知或改写知识库;是否以及如何使用主题,由宿主 Agent
|
||||||
|
> 决定。
|
||||||
|
|
||||||
|
## 📊 评测结果
|
||||||
|
|
||||||
|
ReMe 通过 Agent 多轮搜索与读取的方式,评测多会话和超长上下文中的记忆能力。下表为仓库中已公开的参考实验结果;模型、prompt、数据集和评判细节见各评测文档。
|
||||||
|
|
||||||
|
| 基准 | 设置 | 样本量 | Agentic 得分 | 主要检验内容 |
|
||||||
|
| --------------------------------------------------------------------------- | ----------- | ----------------: | -----------: | ------------------------------ |
|
||||||
|
| **[LongMemEval cleaned-s](https://reme.agentscope.io/zh/benchmarks/longmemeval)** | **整体** | **500 题** | **89.4%** | 跨会话检索、知识更新与时间推理 |
|
||||||
|
| [BEAM](https://reme.agentscope.io/zh/benchmarks/beam) | 100K 上下文 | 20 cases / 400 题 | 66.1% | 十类长上下文记忆任务 |
|
||||||
|
| [BEAM](https://reme.agentscope.io/zh/benchmarks/beam) | 1M 上下文 | 35 cases / 700 题 | 65.0% | 超长对话设置 |
|
||||||
|
|
||||||
|
在仓库的 [π-Bench 评测](https://reme.agentscope.io/zh/benchmarks/pibench)中,ReMe Agent 在 5 种用户角色上的平均 **PROC 得分为 0.580**
|
||||||
|
,比相同测试模型配置的 NanoBot 高 2.4%。PROC 用于评估隐藏意图完成、针对性澄清、跨会话偏好和规范复用、跨任务依赖推断以及欠规格请求推进等主动性能力。
|
||||||
|
|
||||||
|
## 🧩 扩展与插件
|
||||||
|
|
||||||
|
插件是可选的独立 Python distribution,可以贡献 Component、Step、Job backend 和配置,并通过配置显式启用。每日论文与 Auto Fin
|
||||||
|
均已独立打包,使用说明分别见[每日论文](https://reme.agentscope.io/zh/plugins/daily-paper)和
|
||||||
|
[Auto Fin](https://reme.agentscope.io/zh/plugins/auto-fin)。
|
||||||
|
|
||||||
|
| 插件 | 能力 |
|
||||||
|
| ---------------------------------------------------------- | ------------------------------------------------------------------------------ |
|
||||||
|
| [每日论文](https://reme.agentscope.io/zh/plugins/daily-paper) | 发现并排序论文,使用 Agent 解读 PDF,生成文件化论文笔记和五分钟简报。 |
|
||||||
|
| [Auto Fin](https://reme.agentscope.io/zh/plugins/auto-fin) | 拉取主题相关财联社新闻,搜索 ReMe 历史材料并生成带 wikilink 的 Markdown 报告。 |
|
||||||
|
|
||||||
|
安装、查看、校验、启用和卸载 ReMe 插件的方法见[插件管理](https://reme.agentscope.io/zh/plugin_management)。
|
||||||
|
|
||||||
|
## 📚 文档
|
||||||
|
|
||||||
|
下列文档覆盖主要使用流程,并以当前代码的运行时契约为准。
|
||||||
|
|
||||||
|
| 文档 | 主要内容 |
|
||||||
|
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------- |
|
||||||
|
| [快速开始](https://reme.agentscope.io/zh/quick_start) | 安装 ReMe、启动服务,并执行首次文件和记忆操作。 |
|
||||||
|
| [基础配置](https://reme.agentscope.io/zh/configuration) | 配置 workspace、模型、Service、Job、Component、插件和命令行覆盖。 |
|
||||||
|
| [服务与部署](https://reme.agentscope.io/zh/services) | 使用 HTTP、SSE、MCP 和 Studio,并理解默认安全边界。 |
|
||||||
|
| [Memory as File](https://reme.agentscope.io/zh/memory_as_file) | 理解 workspace 分层、frontmatter、wikilink、chunk 和文件事实来源模型。 |
|
||||||
|
| [Auto Memory](https://reme.agentscope.io/zh/auto_memory) | 保留过滤后的对话来源记录,并提炼可复用的 daily 记忆卡片。 |
|
||||||
|
| [Auto Resource](https://reme.agentscope.io/zh/auto_resource) | 导入支持的文本与图像资料,转换为可追溯来源的 daily 卡片。 |
|
||||||
|
| [Auto Dream](https://reme.agentscope.io/zh/auto_dream) 与 [Auto Link](https://reme.agentscope.io/zh/auto_link) | 将 daily 记忆整理为持续演化的 digest 节点和可读 wikilink 关系。 |
|
||||||
|
| [记忆检索](https://reme.agentscope.io/zh/memory_search) | 使用 BM25、可选向量、RRF 融合、行号范围召回和渐进式链接扩展。 |
|
||||||
|
| [Proactive](https://reme.agentscope.io/zh/proactive) | 安全读取兴趣主题,并将其接入宿主 Agent 的决策流程。 |
|
||||||
|
| [应用场景](https://reme.agentscope.io/zh/reme_scene) | 查看金融研究、研发记忆和个人知识库的完整使用示例。 |
|
||||||
|
| [框架说明](https://reme.agentscope.io/zh/framework) | 理解 Application、Job、Step、Component、service、配置和生命周期边界。 |
|
||||||
|
| [Agent 集成](https://reme.agentscope.io/zh/integrations) | 选择接口,并将 DSH、Claude Code、OpenClaw、Hermes、Codex 或其他 Agent 接入 ReMe。 |
|
||||||
|
| [DSH 插件](https://reme.agentscope.io/zh/integrations/dsh) 与 [Claude Code 插件](https://reme.agentscope.io/zh/integrations/claude-code) | 配置宿主原生召回、自动捕获、记忆整理与诊断。 |
|
||||||
|
| [CLI 与 Job API](https://reme.agentscope.io/zh/reference/cli) | 查询命令语法,以及由默认配置自动生成的 Job 参数参考。 |
|
||||||
|
| [运维与恢复](https://reme.agentscope.io/zh/operations) | 诊断服务、维护索引,并备份、迁移和恢复 workspace。 |
|
||||||
|
| [ReMe 博客](https://reme.agentscope.io/zh/reme-blog) | 了解完整产品故事、设计动机、使用示例和评测摘要。 |
|
||||||
|
|
||||||
|
## 🛠️ 常用命令
|
||||||
|
|
||||||
|
运行 `reme help` 可查看完整 job 列表。常用 workspace 与维护命令如下:
|
||||||
|
|
||||||
| 命令 | 作用 |
|
| 命令 | 作用 |
|
||||||
|-------------------------------------------|---------------------------------------------|
|
| ----------------------------------------- | ------------------------------------------------------------- |
|
||||||
| `reme start` | 启动本地 ReMe 服务。 |
|
|
||||||
| `reme version` / `reme health_check` | 检查包版本和组件状态。 |
|
|
||||||
| `reme status` | 查看有状态数据组件的内存估算及进程 RSS。 |
|
| `reme status` | 查看有状态数据组件的内存估算及进程 RSS。 |
|
||||||
| [`reme search`](docs/zh/memory_search.md) | 默认使用 BM25 和 wikilink 检索,启用后增加向量检索。 |
|
| [`reme search`](https://reme.agentscope.io/zh/memory_search) | 默认使用 BM25 和 wikilink 检索,启用后增加向量检索。 |
|
||||||
| `reme read` / `reme write` / `reme edit` | 检查和维护 Markdown 记忆文件。 |
|
| `reme read` / `reme write` / `reme edit` | 检查和维护 Markdown 记忆文件。 |
|
||||||
| `reme auto_memory` | 将对话 messages 转为 daily 记忆卡片;需要 LLM 凭证。 |
|
| `reme traverse` / `reme graph_snapshot` | 浏览 wikilink 邻域或按类别组织的 digest 图。 |
|
||||||
| `reme auto_resource` | 将 `resource/` 下的文件解读为 daily 资料卡片;需要 LLM 凭证。 |
|
| `reme chat` | 与可感知 workspace 的只读 Agent 进行流式对话;需要 LLM 凭证。 |
|
||||||
| `reme auto_dream` / `reme proactive` | 将 daily 记忆整理为长期 digest,并暴露值得关注的主题。 |
|
| `reme reindex` | 基于已摄取的 chunks 重建 BM25 和 Embedding 索引。 |
|
||||||
| `reme reindex` | 基于已有文件重建检索和 wikilink 索引。 |
|
|
||||||
|
|
||||||
## 🤝 社区与支持
|
## 🤝 社区与贡献
|
||||||
|
|
||||||
- **问题反馈与需求**:请先查看 [Open Issues](https://github.com/agentscope-ai/ReMe/issues);如无相关讨论,可新建 Issue
|
- **问题反馈、需求与帮助**:请先查看 [Open Issues](https://github.com/agentscope-ai/ReMe/issues);如无相关讨论,可新建 Issue
|
||||||
说明背景、目标行为和影响范围。
|
说明背景、目标行为和影响范围。
|
||||||
- **代码贡献**:改动前建议阅读 [贡献指南](https://docs.agentscope.io/reme/stable/zh/contributing)。架构与扩展方式以源码、schema
|
- **代码贡献**:改动前建议阅读[贡献指南](https://reme.agentscope.io/zh/contributing)。架构与扩展方式以源码、schema 和测试为准。
|
||||||
和测试为准。
|
- **文档贡献**:请直接更新本仓库 `docs/en/`、`docs/zh/` 或对应 package 目录中的规范源文件;文档站点会从这些文件生成。
|
||||||
- **文档贡献**:用户可见文档请提交到[统一文档仓库](https://github.com/agentscope-ai/docs)的 `reme/<version>/{en,zh}/` 目录。
|
|
||||||
- **提交规范**:建议使用 Conventional Commits,例如 `feat(search): add link expansion option`、
|
- **提交规范**:建议使用 Conventional Commits,例如 `feat(search): add link expansion option`、
|
||||||
`docs(zh): update quick start`。
|
`docs(zh): update quick start`。
|
||||||
- **提交前检查**:提交 PR 前请尽量运行 `pre-commit run --all-files` 和 `pytest`;如有依赖 LLM、embedding 或外部服务的测试无法运行,请在
|
- **提交前检查**:提交 PR 前请尽量运行 `pre-commit run --all-files` 和 `pytest`;如有依赖 LLM、embedding 或外部服务的测试无法运行,请在
|
||||||
PR 中说明。
|
PR 中说明。
|
||||||
- **获取帮助**:如需反馈 Bug 或功能请求,请使用 [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues);项目文档见
|
- **项目文档**:访问 [reme.agentscope.io](https://reme.agentscope.io)。
|
||||||
[https://docs.agentscope.io/](https://docs.agentscope.io/reme/stable/zh/)。
|
|
||||||
|
|
||||||
### 贡献者
|
### 贡献者
|
||||||
|
|
||||||
|
|
|
||||||
137
benchmark/beam/README.md
Normal file
137
benchmark/beam/README.md
Normal file
|
|
@ -0,0 +1,137 @@
|
||||||
|
[中文版 / Chinese version](./README_ZH.md)
|
||||||
|
|
||||||
|
# BEAM Benchmark
|
||||||
|
|
||||||
|
BEAM is a benchmark for **memory capability over long-context chat cases**. Each
|
||||||
|
case contains a very long chat history split into batches; ReMe converts each
|
||||||
|
batch into a session, ingests them in chronological order, then answers probing
|
||||||
|
questions via an agentic (ReAct) mode. Answers are scored with BEAM's
|
||||||
|
rubric-based `answer_judge` job, which produces both a graded score and a binary
|
||||||
|
verdict, and per-type averages are reported.
|
||||||
|
|
||||||
|
BEAM ships dataset variants by chat size — `100K` / `500K` / `1M` / `10M` — so
|
||||||
|
memory systems can be stressed at different context lengths. Question types
|
||||||
|
include abstention, contradiction resolution, event ordering, information
|
||||||
|
extraction, instruction following, knowledge update, multi-session reasoning,
|
||||||
|
preference following, summarization, and temporal reasoning.
|
||||||
|
|
||||||
|
Install ReMe and the BEAM plugin in editable mode from the repository root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m pip install -e ".[as]"
|
||||||
|
reme plugins install ./plugins/beam --editable
|
||||||
|
reme plugins install ./plugins/beam-judge --editable
|
||||||
|
reme plugins validate beam
|
||||||
|
```
|
||||||
|
|
||||||
|
The runner explicitly enables the installed `beam` plugin and combines its defaults with
|
||||||
|
ReMe's built-in `benchmark` preset. Editable installation keeps changes under
|
||||||
|
[`plugins/beam`](../../plugins/beam/README.md) visible without reinstalling the plugin.
|
||||||
|
Custom application config paths still work through `reme.config` and can use `extends: benchmark`.
|
||||||
|
This directory continues to own the runner, evaluation settings, dataset and outputs.
|
||||||
|
Model credentials use the environment variables declared by the shared benchmark configuration.
|
||||||
|
|
||||||
|
## 1. Get the Dataset
|
||||||
|
|
||||||
|
BEAM is a public repository, cloned into `benchmark/beam/dataset/`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p benchmark/beam/dataset
|
||||||
|
cd benchmark/beam/dataset
|
||||||
|
git clone https://github.com/mohammadtavakoli78/BEAM.git
|
||||||
|
```
|
||||||
|
|
||||||
|
After cloning, `benchmark/beam/dataset/BEAM/` should contain `chats/`, `src/`,
|
||||||
|
`topics/` and other subdirectories.
|
||||||
|
|
||||||
|
## 2. Run
|
||||||
|
|
||||||
|
From the repository root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python benchmark/beam/run.py
|
||||||
|
python benchmark/beam/run.py --config benchmark/beam/config.yaml
|
||||||
|
python benchmark/beam/run.py -q # quiet
|
||||||
|
python benchmark/beam/run.py --eval_only # reuse existing workspaces, query + judge only
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. Pipeline
|
||||||
|
|
||||||
|
1. For each case, load `chat.json` and convert each batch into a ReMe session.
|
||||||
|
2. Ingest sessions in chronological order into an isolated workspace, then `digest_update`.
|
||||||
|
3. Answer each probing question via agentic (ReAct) mode.
|
||||||
|
4. Score answers with BEAM's rubric-based `answer_judge` job and print per-type averages.
|
||||||
|
|
||||||
|
## 4. Key config — `benchmark/beam/config.yaml`
|
||||||
|
|
||||||
|
| Key | Meaning |
|
||||||
|
| --- | --- |
|
||||||
|
| `dataset.beam_root` | BEAM dataset root (`benchmark/beam/dataset/BEAM`). |
|
||||||
|
| `dataset.chat_size` | Variant to run: `100K` / `500K` / `1M` / `10M`. |
|
||||||
|
| `dataset.case_ids` | Specific cases (e.g. `["1","2"]`); empty = all cases. |
|
||||||
|
| `dataset.start_index` / `num_items` | Case pagination (`num_items` `0` = all). |
|
||||||
|
| `dataset.workspace_root` | Per-case workspace root (`benchmark/beam/workspaces/beam`). |
|
||||||
|
| `evaluation.num_workers` | `0` = auto, `1` = sequential, `>1` = parallel. |
|
||||||
|
| `reme.config` | ReMe config used (`benchmark`). |
|
||||||
|
| `output.dir` | Results directory (`benchmark/beam/results`). |
|
||||||
|
|
||||||
|
## 5. Outputs
|
||||||
|
|
||||||
|
Results are JSON files written to `output.dir` as
|
||||||
|
`results_<chat_size>_<timestamp>.json`, with a per-type score summary also
|
||||||
|
printed to the console. Logging conventions are shared across benchmarks — see
|
||||||
|
the [top-level README](../README.md#outputs--logs).
|
||||||
|
|
||||||
|
## 6. Reference Results
|
||||||
|
|
||||||
|
> The results below use the longmemeval-version prompt.
|
||||||
|
|
||||||
|
### 100K
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1, conda reme env, 20 workers, eval-only (reusing prebuilt memory)
|
||||||
|
(2026-08-05, 20 cases / 400 Qs, total 46.0 min)
|
||||||
|
|
||||||
|
| Type | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|---|
|
||||||
|
| abstention | 0.550 | 0.550 | 96,031 | 1,070 | 97,101 | 4.58 |
|
||||||
|
| contradiction_resolution | 0.438 | 0.412 | 32,263 | 872 | 33,135 | 2.48 |
|
||||||
|
| event_ordering | 0.501 | 0.423 | 140,195 | 5,163 | 145,358 | 4.70 |
|
||||||
|
| information_extraction | 0.873 | 0.832 | 50,245 | 883 | 51,128 | 3.15 |
|
||||||
|
| instruction_following | 0.750 | 0.725 | 37,986 | 848 | 38,834 | 2.67 |
|
||||||
|
| knowledge_update | 0.688 | 0.675 | 31,198 | 651 | 31,849 | 2.27 |
|
||||||
|
| multi_session_reasoning | 0.626 | 0.584 | 85,038 | 4,563 | 89,601 | 4.28 |
|
||||||
|
| preference_following | 0.925 | 0.912 | 34,281 | 989 | 35,270 | 2.50 |
|
||||||
|
| summarization | 0.623 | 0.461 | 89,657 | 2,056 | 91,713 | 4.12 |
|
||||||
|
| temporal_reasoning | 0.637 | 0.625 | 34,563 | 1,049 | 35,612 | 2.52 |
|
||||||
|
| **OVERALL** | **0.661** | **0.620** | **63,146** | **1,814** | **64,960** | **3.33** |
|
||||||
|
|
||||||
|
Memory Construction average token consumption (default agent, full build over 20 cases):
|
||||||
|
|
||||||
|
| Agent | input tok/case | output tok/case | total tok/case |
|
||||||
|
|---|---|---|---|
|
||||||
|
| default | 2,172,316 | 136,697 | 2,309,013 |
|
||||||
|
|
||||||
|
### 1M
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1, conda reme env, 20 workers, full memory build
|
||||||
|
(2026-08-05, 35 cases / 700 Qs, total 459.2 min)
|
||||||
|
|
||||||
|
| Type | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|---|
|
||||||
|
| abstention | 0.429 | 0.429 | 118,707 | 1,178 | 119,886 | 4.20 |
|
||||||
|
| contradiction_resolution | 0.391 | 0.364 | 49,787 | 810 | 50,597 | 2.50 |
|
||||||
|
| event_ordering | 0.558 | 0.456 | 201,514 | 3,889 | 205,403 | 4.79 |
|
||||||
|
| information_extraction | 0.809 | 0.772 | 78,950 | 894 | 79,844 | 3.00 |
|
||||||
|
| instruction_following | 0.852 | 0.832 | 55,757 | 924 | 56,681 | 2.81 |
|
||||||
|
| knowledge_update | 0.779 | 0.771 | 45,981 | 665 | 46,646 | 2.37 |
|
||||||
|
| multi_session_reasoning | 0.658 | 0.612 | 138,133 | 2,873 | 141,006 | 4.40 |
|
||||||
|
| preference_following | 0.798 | 0.777 | 51,796 | 920 | 52,716 | 2.53 |
|
||||||
|
| summarization | 0.693 | 0.537 | 158,794 | 2,905 | 161,700 | 4.44 |
|
||||||
|
| temporal_reasoning | 0.536 | 0.536 | 100,176 | 3,148 | 103,324 | 3.90 |
|
||||||
|
| **OVERALL** | **0.650** | **0.609** | **99,959** | **1,821** | **101,780** | **3.49** |
|
||||||
|
|
||||||
|
Memory Construction average token consumption (default agent, full build over 35 cases):
|
||||||
|
|
||||||
|
| Agent | input tok/case | output tok/case | total tok/case |
|
||||||
|
|---|---|---|---|
|
||||||
|
| default | 31,943,817 | 1,417,061 | 33,360,878 |
|
||||||
132
benchmark/beam/README_ZH.md
Normal file
132
benchmark/beam/README_ZH.md
Normal file
|
|
@ -0,0 +1,132 @@
|
||||||
|
# BEAM 评测
|
||||||
|
|
||||||
|
[English version](./README.md)
|
||||||
|
|
||||||
|
BEAM 是一个面向**长上下文对话场景**的记忆能力评测基准。每个 case 包含一段被切分为多个
|
||||||
|
batch 的超长对话;ReMe 将每个 batch 转换为一个会话,按时间顺序摄入后,以 agentic(ReAct)
|
||||||
|
模式回答探测问题。答案由 BEAM 基于 rubric 的 `answer_judge` 任务打分,同时给出分级分数与二元
|
||||||
|
判定,并输出各类型平均分。
|
||||||
|
|
||||||
|
BEAM 按对话规模提供多种数据变体 —— `100K` / `500K` / `1M` / `10M`,可在不同上下文长度下
|
||||||
|
压测记忆系统。题型包括 abstention(拒答)、contradiction resolution(矛盾消解)、event
|
||||||
|
ordering(事件排序)、information extraction(信息抽取)、instruction following(指令遵循)、
|
||||||
|
knowledge update(知识更新)、multi-session reasoning(多会话推理)、preference following
|
||||||
|
(偏好遵循)、summarization(摘要)与 temporal reasoning(时间推理)。
|
||||||
|
|
||||||
|
在仓库根目录以 editable 模式安装 ReMe 和 BEAM 插件:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m pip install -e ".[as]"
|
||||||
|
reme plugins install ./plugins/beam --editable
|
||||||
|
reme plugins install ./plugins/beam-judge --editable
|
||||||
|
reme plugins validate beam
|
||||||
|
```
|
||||||
|
|
||||||
|
runner 显式启用已安装的 `beam` 插件,并将插件默认配置与 ReMe 内置的 `benchmark` 配置组合。
|
||||||
|
editable 安装会让 [`plugins/beam`](../../plugins/beam/README_ZH.md) 下的源码修改直接生效,无需重复安装。
|
||||||
|
本目录继续保留评测参数、数据集及输出。自定义完整应用配置路径仍可通过 `reme.config` 指定,
|
||||||
|
并可使用 `extends: benchmark`。
|
||||||
|
模型凭据通过公共 benchmark 配置中声明的环境变量设置。
|
||||||
|
|
||||||
|
## 1. 获取数据集
|
||||||
|
|
||||||
|
BEAM 是公开仓库,clone 到 `benchmark/beam/dataset/` 下:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p benchmark/beam/dataset
|
||||||
|
cd benchmark/beam/dataset
|
||||||
|
git clone https://github.com/mohammadtavakoli78/BEAM.git
|
||||||
|
```
|
||||||
|
|
||||||
|
clone 完成后,`benchmark/beam/dataset/BEAM/` 目录下应包含 `chats/`、`src/`、`topics/` 等子目录。
|
||||||
|
|
||||||
|
## 2. 运行
|
||||||
|
|
||||||
|
在仓库根目录执行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python benchmark/beam/run.py
|
||||||
|
python benchmark/beam/run.py --config benchmark/beam/config.yaml
|
||||||
|
python benchmark/beam/run.py -q # 安静模式
|
||||||
|
python benchmark/beam/run.py --eval_only # 复用已有工作区,仅执行查询 + 评判
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 流程
|
||||||
|
|
||||||
|
1. 为每个 case 加载 `chat.json`,将每个 batch 转换为一个 ReMe 会话。
|
||||||
|
2. 按时间顺序将会话摄入独立工作区,随后执行 `digest_update`。
|
||||||
|
3. 以 agentic(ReAct)模式回答每个探测问题。
|
||||||
|
4. 通过 BEAM 基于 rubric 的 `answer_judge` 任务打分,并输出各类型平均分。
|
||||||
|
|
||||||
|
## 4. 关键配置 —— `benchmark/beam/config.yaml`
|
||||||
|
|
||||||
|
| 配置项 | 含义 |
|
||||||
|
| --- | --- |
|
||||||
|
| `dataset.beam_root` | BEAM 数据集根目录(`benchmark/beam/dataset/BEAM`)。 |
|
||||||
|
| `dataset.chat_size` | 运行的变体:`100K` / `500K` / `1M` / `10M`。 |
|
||||||
|
| `dataset.case_ids` | 指定 case(如 `["1","2"]`),空表示全部。 |
|
||||||
|
| `dataset.start_index` / `num_items` | case 分页(`num_items` 为 `0` 表示全部)。 |
|
||||||
|
| `dataset.workspace_root` | case 工作区根目录(`benchmark/beam/workspaces/beam`)。 |
|
||||||
|
| `evaluation.num_workers` | `0` = 自动,`1` = 串行,`>1` = 并行。 |
|
||||||
|
| `reme.config` | 使用的 ReMe 配置(`benchmark`)。 |
|
||||||
|
| `output.dir` | 结果目录(`benchmark/beam/results`)。 |
|
||||||
|
|
||||||
|
## 5. 输出
|
||||||
|
|
||||||
|
结果以 JSON 文件写入 `output.dir`,文件名为 `results_<chat_size>_<timestamp>.json`,
|
||||||
|
同时控制台会打印含各类型分数的汇总。日志约定在各基准间通用,见
|
||||||
|
[总说明](../README_ZH.md#输出与日志)。
|
||||||
|
|
||||||
|
## 6. 参考结果
|
||||||
|
|
||||||
|
> 以下结果使用 longmemeval 版本的 prompt。
|
||||||
|
|
||||||
|
### 100K
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1,conda reme 环境,20 并发,eval-only(复用已构建 memory)
|
||||||
|
(2026-08-05,20 cases / 400 Qs,总耗时 46.0 min)
|
||||||
|
|
||||||
|
| 题型 | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|---|
|
||||||
|
| abstention | 0.550 | 0.550 | 96,031 | 1,070 | 97,101 | 4.58 |
|
||||||
|
| contradiction_resolution | 0.438 | 0.412 | 32,263 | 872 | 33,135 | 2.48 |
|
||||||
|
| event_ordering | 0.501 | 0.423 | 140,195 | 5,163 | 145,358 | 4.70 |
|
||||||
|
| information_extraction | 0.873 | 0.832 | 50,245 | 883 | 51,128 | 3.15 |
|
||||||
|
| instruction_following | 0.750 | 0.725 | 37,986 | 848 | 38,834 | 2.67 |
|
||||||
|
| knowledge_update | 0.688 | 0.675 | 31,198 | 651 | 31,849 | 2.27 |
|
||||||
|
| multi_session_reasoning | 0.626 | 0.584 | 85,038 | 4,563 | 89,601 | 4.28 |
|
||||||
|
| preference_following | 0.925 | 0.912 | 34,281 | 989 | 35,270 | 2.50 |
|
||||||
|
| summarization | 0.623 | 0.461 | 89,657 | 2,056 | 91,713 | 4.12 |
|
||||||
|
| temporal_reasoning | 0.637 | 0.625 | 34,563 | 1,049 | 35,612 | 2.52 |
|
||||||
|
| **OVERALL** | **0.661** | **0.620** | **63,146** | **1,814** | **64,960** | **3.33** |
|
||||||
|
|
||||||
|
Memory Construction 平均 token 消耗(default agent,20 cases 全量构建):
|
||||||
|
|
||||||
|
| Agent | input tok/case | output tok/case | total tok/case |
|
||||||
|
|---|---|---|---|
|
||||||
|
| default | 2,172,316 | 136,697 | 2,309,013 |
|
||||||
|
|
||||||
|
### 1M
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1,conda reme 环境,20 并发,全量构建 memory
|
||||||
|
(2026-08-05,35 cases / 700 Qs,总耗时 459.2 min)
|
||||||
|
|
||||||
|
| 题型 | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|---|
|
||||||
|
| abstention | 0.429 | 0.429 | 118,707 | 1,178 | 119,886 | 4.20 |
|
||||||
|
| contradiction_resolution | 0.391 | 0.364 | 49,787 | 810 | 50,597 | 2.50 |
|
||||||
|
| event_ordering | 0.558 | 0.456 | 201,514 | 3,889 | 205,403 | 4.79 |
|
||||||
|
| information_extraction | 0.809 | 0.772 | 78,950 | 894 | 79,844 | 3.00 |
|
||||||
|
| instruction_following | 0.852 | 0.832 | 55,757 | 924 | 56,681 | 2.81 |
|
||||||
|
| knowledge_update | 0.779 | 0.771 | 45,981 | 665 | 46,646 | 2.37 |
|
||||||
|
| multi_session_reasoning | 0.658 | 0.612 | 138,133 | 2,873 | 141,006 | 4.40 |
|
||||||
|
| preference_following | 0.798 | 0.777 | 51,796 | 920 | 52,716 | 2.53 |
|
||||||
|
| summarization | 0.693 | 0.537 | 158,794 | 2,905 | 161,700 | 4.44 |
|
||||||
|
| temporal_reasoning | 0.536 | 0.536 | 100,176 | 3,148 | 103,324 | 3.90 |
|
||||||
|
| **OVERALL** | **0.650** | **0.609** | **99,959** | **1,821** | **101,780** | **3.49** |
|
||||||
|
|
||||||
|
Memory Construction 平均 token 消耗(default agent,35 cases 全量构建):
|
||||||
|
|
||||||
|
| Agent | input tok/case | output tok/case | total tok/case |
|
||||||
|
|---|---|---|---|
|
||||||
|
| default | 31,943,817 | 1,417,061 | 33,360,878 |
|
||||||
25
benchmark/beam/config.yaml
Normal file
25
benchmark/beam/config.yaml
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
# BEAM evaluation configuration
|
||||||
|
# This file controls what/how to evaluate.
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
beam_root: "benchmark/beam/dataset/BEAM" # BEAM dataset root
|
||||||
|
chat_size: "1M" # 100K | 500K | 1M | 10M (dataset variant)
|
||||||
|
case_ids: [] # empty = all cases; or ["1", "2", "3"]
|
||||||
|
start_index: 0 # first case index (for pagination)
|
||||||
|
num_items: 0 # 0 = all cases; >0 = limit
|
||||||
|
workspace_root: "benchmark/beam/workspaces/beam" # workspace root for case workspaces
|
||||||
|
|
||||||
|
evaluation:
|
||||||
|
num_workers: 20 # 0 = auto; 1 = sequential; >1 = parallel (per-case)
|
||||||
|
compress_session: false # true = compress session chunks in search_v2 (query-aware); false = no compression
|
||||||
|
|
||||||
|
reme:
|
||||||
|
config: "benchmark" # shared ReMe benchmark preset
|
||||||
|
plugins: [beam, beam-judge]
|
||||||
|
|
||||||
|
output:
|
||||||
|
dir: "benchmark/beam/results"
|
||||||
|
log_dir: "logs" # log directory (relative to project root)
|
||||||
|
log_prefix: "beam" # benchmark name used in log filenames
|
||||||
|
log_to_console: true
|
||||||
|
log_to_file: true
|
||||||
76
benchmark/beam/kill.sh
Normal file
76
benchmark/beam/kill.sh
Normal file
|
|
@ -0,0 +1,76 @@
|
||||||
|
#!/bin/bash
|
||||||
|
# 杀死指定进程及其所有子进程
|
||||||
|
# Usage: bash kill.sh <PID>
|
||||||
|
|
||||||
|
if [ -z "$1" ]; then
|
||||||
|
echo "Usage: bash kill.sh <PID>"
|
||||||
|
echo " 杀死指定进程及其所有子进程"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
PID=$1
|
||||||
|
|
||||||
|
# 检查进程是否存在
|
||||||
|
if ! kill -0 "$PID" 2>/dev/null; then
|
||||||
|
echo "进程 $PID 不存在"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 递归收集所有子进程(包括子进程的子进程)
|
||||||
|
collect_children() {
|
||||||
|
local parent=$1
|
||||||
|
local children
|
||||||
|
children=$(ps -o pid= --ppid "$parent" 2>/dev/null | tr -d ' ')
|
||||||
|
for child in $children; do
|
||||||
|
collect_children "$child"
|
||||||
|
done
|
||||||
|
echo "$parent"
|
||||||
|
}
|
||||||
|
|
||||||
|
# 收集进程树(子进程在前,父进程在后,保证先杀子再杀父)
|
||||||
|
PROCESS_TREE=$(collect_children "$PID")
|
||||||
|
TOTAL=$(echo "$PROCESS_TREE" | wc -l | tr -d ' ')
|
||||||
|
|
||||||
|
echo "进程树(共 $TOTAL 个进程):"
|
||||||
|
while read -r p; do
|
||||||
|
cmd=$(ps -o args= -p "$p" 2>/dev/null | head -c 80)
|
||||||
|
printf " PID=%-8s %s\n" "$p" "$cmd"
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
# 先 SIGTERM 优雅终止
|
||||||
|
echo ""
|
||||||
|
echo "发送 SIGTERM..."
|
||||||
|
while read -r p; do
|
||||||
|
kill "$p" 2>/dev/null
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
# 等待最多 5 秒
|
||||||
|
for i in $(seq 1 5); do
|
||||||
|
alive=false
|
||||||
|
while read -r p; do
|
||||||
|
if kill -0 "$p" 2>/dev/null; then
|
||||||
|
alive=true
|
||||||
|
fi
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
if [ "$alive" = false ]; then
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
|
||||||
|
# 检查是否还有残留,强制 SIGKILL
|
||||||
|
remaining=false
|
||||||
|
while read -r p; do
|
||||||
|
if kill -0 "$p" 2>/dev/null; then
|
||||||
|
remaining=true
|
||||||
|
fi
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
if [ "$remaining" = true ]; then
|
||||||
|
echo "部分进程未响应,发送 SIGKILL..."
|
||||||
|
while read -r p; do
|
||||||
|
kill -9 "$p" 2>/dev/null
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "已终止进程树(根 PID=$PID,共 $TOTAL 个进程)"
|
||||||
906
benchmark/beam/run.py
Normal file
906
benchmark/beam/run.py
Normal file
|
|
@ -0,0 +1,906 @@
|
||||||
|
"""BEAM evaluation runner for ReMe.
|
||||||
|
|
||||||
|
Evaluates ReMe's memory capability using the BEAM dataset.
|
||||||
|
Each case gets an isolated workspace; chat.json batches are ingested as
|
||||||
|
sessions in chronological order; finally probing questions are answered
|
||||||
|
via an agentic (ReAct) approach, then
|
||||||
|
judged by BEAM's rubric-based LLM-as-judge.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python benchmark/beam/run.py
|
||||||
|
python benchmark/beam/run.py --config benchmark/beam/config.yaml
|
||||||
|
python benchmark/beam/run.py -q # quiet: only eval-level logs
|
||||||
|
python benchmark/beam/run.py --log-level WARNING # reduce eval runner logs
|
||||||
|
python benchmark/beam/run.py --reme-log-level WARNING # reduce reme internal logs
|
||||||
|
python benchmark/beam/run.py --eval_only # query+judge only, reuse existing workspace
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import time
|
||||||
|
import threading
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
from dotenv import load_dotenv
|
||||||
|
|
||||||
|
# Load .env from project root
|
||||||
|
_PROJECT_ROOT = Path(__file__).resolve().parent.parent.parent
|
||||||
|
load_dotenv(_PROJECT_ROOT / ".env")
|
||||||
|
|
||||||
|
# Workspace root — read from config.yaml (dataset.workspace_root)
|
||||||
|
_WORKSPACE_ROOT_DEFAULT = "benchmark/beam/workspaces/beam"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Logging
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
_DEFAULT_LOG_FORMAT = "%(asctime)s | %(levelname)s | %(message)s"
|
||||||
|
|
||||||
|
logging.basicConfig(level=logging.INFO, format=_DEFAULT_LOG_FORMAT)
|
||||||
|
logger = logging.getLogger("beam")
|
||||||
|
|
||||||
|
# Noisy library loggers silenced by default
|
||||||
|
_NOISY_LOGGERS = [
|
||||||
|
"httpx",
|
||||||
|
"httpcore",
|
||||||
|
"openai",
|
||||||
|
"uvicorn",
|
||||||
|
"multipart",
|
||||||
|
"asyncio",
|
||||||
|
"watchfiles",
|
||||||
|
"filelock",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def setup_logging(
|
||||||
|
log_level: str,
|
||||||
|
reme_log_level: str,
|
||||||
|
log_dir: str | None = None,
|
||||||
|
):
|
||||||
|
"""Configure logging for the eval runner and reme internals.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
log_level: Level for the eval runner logger (DEBUG/INFO/WARNING/ERROR).
|
||||||
|
reme_log_level: Level for reme's internal loguru logger.
|
||||||
|
log_dir: Per-run log directory (absolute path). None = no file logging.
|
||||||
|
"""
|
||||||
|
numeric = getattr(logging, log_level.upper(), logging.INFO)
|
||||||
|
# Eval runner logger
|
||||||
|
logging.getLogger().setLevel(numeric)
|
||||||
|
logger.setLevel(numeric)
|
||||||
|
|
||||||
|
# Suppress noisy library loggers when above DEBUG
|
||||||
|
if numeric > logging.DEBUG:
|
||||||
|
for name in _NOISY_LOGGERS:
|
||||||
|
lib_logger = logging.getLogger(name)
|
||||||
|
lib_logger.setLevel(max(numeric, logging.WARNING))
|
||||||
|
|
||||||
|
# Add file handler for eval runner if log_dir is specified
|
||||||
|
if log_dir:
|
||||||
|
os.makedirs(log_dir, exist_ok=True)
|
||||||
|
log_filepath = os.path.join(log_dir, "runner.log")
|
||||||
|
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
|
||||||
|
file_handler.setLevel(numeric)
|
||||||
|
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
|
||||||
|
logging.getLogger().addHandler(file_handler)
|
||||||
|
logger.info(f"Eval runner log file: {log_filepath}")
|
||||||
|
|
||||||
|
# Reme internal logger (loguru) — will be applied per-worker via _configure_worker
|
||||||
|
os.environ["REME_LOG_LEVEL"] = reme_log_level.upper()
|
||||||
|
if log_dir:
|
||||||
|
os.environ["REME_LOG_DIR"] = log_dir
|
||||||
|
|
||||||
|
|
||||||
|
def _configure_worker(
|
||||||
|
log_level: str,
|
||||||
|
reme_log_level: str,
|
||||||
|
log_dir: str | None = None,
|
||||||
|
):
|
||||||
|
"""Set up logging inside a multiprocessing worker process.
|
||||||
|
|
||||||
|
Must be called at the top of each worker because child processes inherit
|
||||||
|
parent state but loguru sinks are NOT shared across fork/spawn.
|
||||||
|
"""
|
||||||
|
numeric = getattr(logging, log_level.upper(), logging.INFO)
|
||||||
|
logging.basicConfig(level=numeric, format=_DEFAULT_LOG_FORMAT, force=True)
|
||||||
|
logging.getLogger("beam").setLevel(numeric)
|
||||||
|
if numeric > logging.DEBUG:
|
||||||
|
for name in _NOISY_LOGGERS:
|
||||||
|
logging.getLogger(name).setLevel(max(numeric, logging.WARNING))
|
||||||
|
|
||||||
|
# Add file handler for eval runner in worker process
|
||||||
|
if log_dir:
|
||||||
|
os.makedirs(log_dir, exist_ok=True)
|
||||||
|
pid = os.getpid()
|
||||||
|
log_filepath = os.path.join(log_dir, f"worker-{pid}.log")
|
||||||
|
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
|
||||||
|
file_handler.setLevel(numeric)
|
||||||
|
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
|
||||||
|
logging.getLogger().addHandler(file_handler)
|
||||||
|
|
||||||
|
# Re-initialize loguru for reme internals at the desired level
|
||||||
|
from reme.utils import get_logger
|
||||||
|
|
||||||
|
reme_log_dir = log_dir or "logs"
|
||||||
|
get_logger(log_dir=reme_log_dir, level=reme_log_level.upper(), force_init=True)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Config loading
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def load_eval_config(config_path: str | None = None) -> dict:
|
||||||
|
"""Load evaluation config yaml with env-var expansion."""
|
||||||
|
if config_path is None:
|
||||||
|
config_path = str(Path(__file__).parent / "config.yaml")
|
||||||
|
with open(config_path, encoding="utf-8") as f:
|
||||||
|
raw = f.read()
|
||||||
|
|
||||||
|
# Expand ${VAR} and ${VAR:-default}
|
||||||
|
def _expand(m):
|
||||||
|
expr = m.group(1)
|
||||||
|
if ":-" in expr:
|
||||||
|
key, default = expr.split(":-", 1)
|
||||||
|
return os.environ.get(key, default)
|
||||||
|
return os.environ.get(expr, "")
|
||||||
|
|
||||||
|
raw = re.sub(r"\$\{([^}]+)\}", _expand, raw)
|
||||||
|
return yaml.safe_load(raw)
|
||||||
|
|
||||||
|
|
||||||
|
def create_reme_app(config: str = "benchmark", **overrides):
|
||||||
|
"""Create an app with the BEAM candidate and judge plugins enabled.
|
||||||
|
|
||||||
|
Plugin discovery remains environment-based; editable installation keeps local
|
||||||
|
plugin source changes visible to every multiprocessing worker.
|
||||||
|
"""
|
||||||
|
from reme import Application
|
||||||
|
from reme.config import resolve_app_config
|
||||||
|
|
||||||
|
enabled_plugins = list(overrides.pop("plugins", ()) or ())
|
||||||
|
for plugin in ("beam", "beam-judge"):
|
||||||
|
if plugin not in enabled_plugins:
|
||||||
|
enabled_plugins.append(plugin)
|
||||||
|
app_config = resolve_app_config(config=config, plugins=enabled_plugins, **overrides)
|
||||||
|
return Application(**app_config)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# BEAM data loading
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def parse_beam_time_anchor(time_str: str) -> datetime:
|
||||||
|
"""Parse BEAM time_anchor format: 'March-15-2024' -> datetime."""
|
||||||
|
for fmt in ("%B-%d-%Y", "%b-%d-%Y"):
|
||||||
|
try:
|
||||||
|
return datetime.strptime(time_str, fmt)
|
||||||
|
except ValueError:
|
||||||
|
continue
|
||||||
|
raise ValueError(f"Cannot parse time_anchor: {time_str!r}")
|
||||||
|
|
||||||
|
|
||||||
|
def load_beam_chat(chat_path: Path, chat_size: str, case_id: str) -> list[dict]:
|
||||||
|
"""Load BEAM chat.json and convert to ReMe session format.
|
||||||
|
|
||||||
|
Each batch becomes one session with all its turns flattened.
|
||||||
|
Each turn resolves its own time_anchor independently; turns without
|
||||||
|
an explicit time_anchor inherit from the most recent preceding turn.
|
||||||
|
Returns list of sessions, each with:
|
||||||
|
- session_id: str
|
||||||
|
- date: str (YYYY-MM-DD) — derived from the *first* turn's time
|
||||||
|
- messages: list[dict] with name, role, content, created_at
|
||||||
|
"""
|
||||||
|
with open(chat_path, encoding="utf-8") as f:
|
||||||
|
batches = json.load(f)
|
||||||
|
|
||||||
|
sessions = []
|
||||||
|
for batch in batches:
|
||||||
|
batch_num = batch["batch_number"]
|
||||||
|
|
||||||
|
# Resolve batch-level fallback (used when no turn has a time_anchor)
|
||||||
|
batch_anchor = batch.get("time_anchor")
|
||||||
|
if not batch_anchor:
|
||||||
|
batch_anchor = "January-1-2024"
|
||||||
|
|
||||||
|
# Flatten all turns, resolving time_anchor per turn
|
||||||
|
messages = []
|
||||||
|
prev_dt = None # carries forward from previous turn
|
||||||
|
first_dt = None # for session-level date
|
||||||
|
|
||||||
|
for turn in batch["turns"]:
|
||||||
|
# Find this turn's own time_anchor from its messages
|
||||||
|
turn_anchor = None
|
||||||
|
for msg in turn:
|
||||||
|
if msg.get("time_anchor"):
|
||||||
|
turn_anchor = msg["time_anchor"]
|
||||||
|
break
|
||||||
|
|
||||||
|
if turn_anchor:
|
||||||
|
dt = parse_beam_time_anchor(turn_anchor)
|
||||||
|
elif prev_dt is not None:
|
||||||
|
dt = prev_dt # inherit from previous turn
|
||||||
|
else:
|
||||||
|
dt = parse_beam_time_anchor(batch_anchor)
|
||||||
|
|
||||||
|
if first_dt is None:
|
||||||
|
first_dt = dt
|
||||||
|
prev_dt = dt
|
||||||
|
|
||||||
|
for msg in turn:
|
||||||
|
role = msg["role"]
|
||||||
|
messages.append(
|
||||||
|
{
|
||||||
|
"name": role,
|
||||||
|
"role": role,
|
||||||
|
"content": msg["content"],
|
||||||
|
"created_at": dt.strftime("%Y-%m-%dT%H:%M:%S"),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
sessions.append(
|
||||||
|
{
|
||||||
|
"session_id": f"beam_{chat_size}_{case_id}_batch{batch_num}",
|
||||||
|
"date": first_dt.strftime("%Y-%m-%d"),
|
||||||
|
"messages": messages,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
return sessions
|
||||||
|
|
||||||
|
|
||||||
|
def get_available_cases(beam_root: Path, chat_size: str) -> list[str]:
|
||||||
|
"""Return sorted list of case IDs for a given chat size."""
|
||||||
|
chats_dir = beam_root / "chats" / chat_size
|
||||||
|
if not chats_dir.exists():
|
||||||
|
return []
|
||||||
|
return sorted(
|
||||||
|
[d.name for d in chats_dir.iterdir() if d.is_dir()],
|
||||||
|
key=int,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Answer generation
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
async def answer_question_agentic(app, question: str, compress_session: bool = False) -> tuple[str, dict]:
|
||||||
|
"""Answer a probing question using ReMe's agentic_answer job.
|
||||||
|
|
||||||
|
Returns (answer, metadata)
|
||||||
|
"""
|
||||||
|
from reme.utils.evaluation_interface import track_agent_token_usage, track_job_counts
|
||||||
|
|
||||||
|
with (
|
||||||
|
track_job_counts(["search"], app.context) as tool_counts,
|
||||||
|
track_agent_token_usage(
|
||||||
|
["bench"],
|
||||||
|
app.context,
|
||||||
|
) as token_usages,
|
||||||
|
):
|
||||||
|
query_resp = await app.run_job(
|
||||||
|
"agentic_answer",
|
||||||
|
query=question,
|
||||||
|
compress_session=compress_session,
|
||||||
|
)
|
||||||
|
answer = (query_resp.answer or "").strip()
|
||||||
|
|
||||||
|
return answer, {
|
||||||
|
"mode": "agentic",
|
||||||
|
"tool_counts": tool_counts,
|
||||||
|
"token_usage": token_usages["bench"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# BEAM rubric-based LLM-as-Judge
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
async def judge_answer(
|
||||||
|
app,
|
||||||
|
question: str,
|
||||||
|
llm_response: str,
|
||||||
|
rubric: list[str],
|
||||||
|
question_type: str = "",
|
||||||
|
) -> dict:
|
||||||
|
"""Judge an answer via the answer_judge job (beam_rubric_judge_step)."""
|
||||||
|
judge_resp = await app.run_job(
|
||||||
|
"answer_judge",
|
||||||
|
llm_response=llm_response,
|
||||||
|
rubric=rubric,
|
||||||
|
probing_question=question,
|
||||||
|
question_type=question_type,
|
||||||
|
)
|
||||||
|
result = {
|
||||||
|
"llm_judge_score": (judge_resp.metadata or {}).get("llm_judge_score", 0.0),
|
||||||
|
"llm_judge_responses": (judge_resp.metadata or {}).get("llm_judge_responses", []),
|
||||||
|
}
|
||||||
|
# Include event_ordering extra metrics if present
|
||||||
|
eo = (judge_resp.metadata or {}).get("event_ordering")
|
||||||
|
if eo:
|
||||||
|
result["event_ordering"] = eo
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Main evaluation pipeline
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
async def evaluate_case(eval_config: dict, case_id: str, eval_only: bool = False) -> dict:
|
||||||
|
"""Evaluate a single BEAM case end-to-end.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
eval_config: The evaluation configuration dict.
|
||||||
|
case_id: The case directory name (e.g. "1").
|
||||||
|
eval_only: If True, skip ingestion and only run query+judge
|
||||||
|
using the existing workspace.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
A results dict with all questions, answers, and judgments.
|
||||||
|
"""
|
||||||
|
|
||||||
|
dataset_cfg = eval_config["dataset"]
|
||||||
|
chat_size = dataset_cfg["chat_size"]
|
||||||
|
compress_session = bool(eval_config["evaluation"].get("compress_session", False))
|
||||||
|
beam_root = _PROJECT_ROOT / dataset_cfg.get("beam_root", "benchmark/beam/dataset/BEAM")
|
||||||
|
chat_path = beam_root / "chats" / chat_size / case_id / "chat.json"
|
||||||
|
probing_questions_path = beam_root / "chats" / chat_size / case_id / "probing_questions" / "probing_questions.json"
|
||||||
|
|
||||||
|
if not chat_path.exists():
|
||||||
|
raise FileNotFoundError(f"Chat file not found: {chat_path}")
|
||||||
|
if not probing_questions_path.exists():
|
||||||
|
raise FileNotFoundError(f"Probing questions not found: {probing_questions_path}")
|
||||||
|
|
||||||
|
logger.info(
|
||||||
|
"[Case %s] size=%s%s",
|
||||||
|
case_id,
|
||||||
|
chat_size,
|
||||||
|
" [eval_only]" if eval_only else "",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Workspace setup
|
||||||
|
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
|
||||||
|
case_dir = workspace_root / f"{chat_size}_{case_id}"
|
||||||
|
workspace_dir = str(case_dir / ".reme")
|
||||||
|
|
||||||
|
if eval_only:
|
||||||
|
if not case_dir.exists() or not Path(workspace_dir).exists():
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"[Case {case_id}] eval_only: workspace not found at {case_dir}. "
|
||||||
|
f"Run without --eval_only first to build the workspace.",
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
if case_dir.exists():
|
||||||
|
shutil.rmtree(case_dir)
|
||||||
|
logger.info(f"[Case {case_id}] Cleaned existing workspace: {case_dir}")
|
||||||
|
else:
|
||||||
|
logger.info(f"[Case {case_id}] Workspace not found, creating: {case_dir}")
|
||||||
|
case_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Pre-initialize ReMe's loguru logger with the correct log_dir
|
||||||
|
output_cfg = eval_config.get("output", {})
|
||||||
|
if output_cfg.get("log_to_file", False):
|
||||||
|
reme_log_dir = os.environ.get("REME_LOG_DIR")
|
||||||
|
if reme_log_dir:
|
||||||
|
from reme.utils import get_logger
|
||||||
|
|
||||||
|
get_logger(
|
||||||
|
log_dir=reme_log_dir,
|
||||||
|
level=os.environ.get("REME_LOG_LEVEL", "INFO"),
|
||||||
|
log_to_console=output_cfg.get("log_to_console", True),
|
||||||
|
log_to_file=True,
|
||||||
|
force_init=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
app = create_reme_app(
|
||||||
|
config=eval_config["reme"]["config"],
|
||||||
|
plugins=eval_config["reme"].get("plugins", ()),
|
||||||
|
workspace_dir=workspace_dir,
|
||||||
|
log_to_console=output_cfg.get("log_to_console", True),
|
||||||
|
log_to_file=output_cfg.get("log_to_file", False),
|
||||||
|
enable_logo=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
await app.start()
|
||||||
|
|
||||||
|
from reme.utils.evaluation_interface import check_agent_token_usage # noqa: E402
|
||||||
|
|
||||||
|
_MEM_AGENT_NAMES = ("default", "bench")
|
||||||
|
sessions_ingested = 0
|
||||||
|
memory_token_usage: dict[str, dict[str, int | None]] = {}
|
||||||
|
try:
|
||||||
|
if not eval_only:
|
||||||
|
# ── Phase 1: Ingest sessions (with token tracking) ─────────
|
||||||
|
sessions = load_beam_chat(chat_path, chat_size, case_id)
|
||||||
|
logger.info(f"[Case {case_id}] Loaded {len(sessions)} sessions from chat.json")
|
||||||
|
|
||||||
|
# Snapshot token counters before memory construction
|
||||||
|
mem_token_start = {name: check_agent_token_usage(name, app.context) for name in _MEM_AGENT_NAMES}
|
||||||
|
|
||||||
|
for i, session in enumerate(sessions):
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] Ingesting session {i+1}/{len(sessions)}: "
|
||||||
|
f"id={session['session_id']} date={session['date']} "
|
||||||
|
f"msgs={len(session['messages'])}",
|
||||||
|
)
|
||||||
|
resp = await app.run_job(
|
||||||
|
"auto_memory",
|
||||||
|
messages=session["messages"],
|
||||||
|
session_id=session["session_id"],
|
||||||
|
date=session["date"],
|
||||||
|
)
|
||||||
|
if not resp.success:
|
||||||
|
logger.warning(f"[Case {case_id}] auto_memory failed: {resp.answer}")
|
||||||
|
else:
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] auto_memory success: " f"{resp.answer[:100] if resp.answer else ''}",
|
||||||
|
)
|
||||||
|
await app.run_job("index_update")
|
||||||
|
sessions_ingested += 1
|
||||||
|
|
||||||
|
# Final digest update
|
||||||
|
logger.info(f"[Case {case_id}] Running digest_update...")
|
||||||
|
await app.run_job("digest_update")
|
||||||
|
logger.info(f"[Case {case_id}] Ingestion complete.")
|
||||||
|
|
||||||
|
# Compute memory construction token deltas
|
||||||
|
for name in _MEM_AGENT_NAMES:
|
||||||
|
end_usage = check_agent_token_usage(name, app.context)
|
||||||
|
delta: dict[str, int | None] = {}
|
||||||
|
for metric in _TOKEN_USAGE_METRICS:
|
||||||
|
current = end_usage[metric]
|
||||||
|
start = mem_token_start[name][metric]
|
||||||
|
delta[metric] = None if current is None else current - (start or 0)
|
||||||
|
memory_token_usage[name] = delta
|
||||||
|
logger.info(f"[Case {case_id}] Memory construction token usage: {memory_token_usage}")
|
||||||
|
|
||||||
|
# ── Phase 2: Answer + Judge probing questions ───────────────
|
||||||
|
with open(probing_questions_path, encoding="utf-8") as f:
|
||||||
|
probing_questions = json.load(f)
|
||||||
|
|
||||||
|
total_questions = sum(len(v) for v in probing_questions.values())
|
||||||
|
logger.info(f"[Case {case_id}] Total probing questions: {total_questions}")
|
||||||
|
|
||||||
|
all_question_results = []
|
||||||
|
q_idx = 0
|
||||||
|
|
||||||
|
for q_type in probing_questions:
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] Question type: {q_type} " f"({len(probing_questions[q_type])} questions)",
|
||||||
|
)
|
||||||
|
|
||||||
|
for i, q in enumerate(probing_questions[q_type]):
|
||||||
|
q_idx += 1
|
||||||
|
question = q["question"]
|
||||||
|
rubric = q.get("rubric", [])
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] [{q_idx}/{total_questions}] " f"{q_type} Q{i+1}: {question[:100]}...",
|
||||||
|
)
|
||||||
|
|
||||||
|
q_result = {
|
||||||
|
"question_type": q_type,
|
||||||
|
"question_index": i,
|
||||||
|
"question": question,
|
||||||
|
"rubric": rubric,
|
||||||
|
}
|
||||||
|
|
||||||
|
# Agentic answer
|
||||||
|
try:
|
||||||
|
agentic_answer, agentic_meta = await answer_question_agentic(
|
||||||
|
app,
|
||||||
|
question,
|
||||||
|
compress_session=compress_session,
|
||||||
|
)
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f"[Case {case_id}] Agentic answer failed: {e}")
|
||||||
|
agentic_answer = f"(error: {e})"
|
||||||
|
agentic_meta = {"error": str(e)}
|
||||||
|
|
||||||
|
if not agentic_answer:
|
||||||
|
agentic_answer = "(no answer generated)"
|
||||||
|
logger.info(f"[Case {case_id}] Agentic answer: {agentic_answer[:200]}...")
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] Agentic tool calls: {agentic_meta.get('tool_counts', {})}",
|
||||||
|
)
|
||||||
|
logger.info(f"[Case {case_id}] Bench token usage: {agentic_meta.get('token_usage', {})}")
|
||||||
|
|
||||||
|
# Judge agentic answer
|
||||||
|
logger.info(f"[Case {case_id}] Judging agentic ({q_type})...")
|
||||||
|
agentic_judgment = await judge_answer(
|
||||||
|
app,
|
||||||
|
question,
|
||||||
|
agentic_answer,
|
||||||
|
rubric,
|
||||||
|
question_type=q_type,
|
||||||
|
)
|
||||||
|
logger.info(
|
||||||
|
f"[Case {case_id}] Agentic score: " f"{agentic_judgment['llm_judge_score']:.3f}",
|
||||||
|
)
|
||||||
|
|
||||||
|
q_result["agentic_response"] = agentic_answer
|
||||||
|
q_result["agentic_judgment"] = agentic_judgment
|
||||||
|
q_result["agentic_metadata"] = agentic_meta
|
||||||
|
|
||||||
|
all_question_results.append(q_result)
|
||||||
|
|
||||||
|
finally:
|
||||||
|
await app.close()
|
||||||
|
|
||||||
|
return {
|
||||||
|
"case_id": case_id,
|
||||||
|
"chat_size": chat_size,
|
||||||
|
"sessions_ingested": sessions_ingested,
|
||||||
|
"total_questions": len(all_question_results),
|
||||||
|
"questions": all_question_results,
|
||||||
|
"memory_token_usage": memory_token_usage,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Worker: runs a single case in its own process with its own event loop
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def _evaluate_case_worker(task_input: tuple) -> dict:
|
||||||
|
"""Worker function for multiprocessing. Each process gets its own event loop."""
|
||||||
|
eval_config, case_id, log_level, reme_log_level, eval_only, log_dir = task_input
|
||||||
|
import asyncio # pylint: disable=import-outside-toplevel
|
||||||
|
|
||||||
|
_configure_worker(log_level, reme_log_level, log_dir=log_dir)
|
||||||
|
|
||||||
|
# Suppress httpx GC noise
|
||||||
|
logging.getLogger("asyncio").setLevel(logging.CRITICAL)
|
||||||
|
|
||||||
|
return asyncio.run(evaluate_case(eval_config, case_id, eval_only=eval_only))
|
||||||
|
|
||||||
|
|
||||||
|
def _indexed_worker(indexed_input: tuple) -> tuple:
|
||||||
|
"""Module-level wrapper for imap_unordered with index tracking."""
|
||||||
|
idx, task_input = indexed_input
|
||||||
|
return idx, _evaluate_case_worker(task_input)
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_num_workers(configured: int) -> int:
|
||||||
|
"""Resolve num_workers: 0=auto (cpu_count-2, min 1), 1=sequential, >1=parallel."""
|
||||||
|
if configured == 0:
|
||||||
|
return max(1, (os.cpu_count() or 4) - 2)
|
||||||
|
return max(1, configured)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Entry point
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def main( # pylint: disable=too-many-statements
|
||||||
|
config_path: str | None = None,
|
||||||
|
log_level: str = "INFO",
|
||||||
|
reme_log_level: str = "INFO",
|
||||||
|
eval_only: bool = False,
|
||||||
|
):
|
||||||
|
"""Run the BEAM evaluation pipeline.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
config_path: Path to the YAML config file.
|
||||||
|
log_level: Log level for the eval runner.
|
||||||
|
reme_log_level: Log level for reme internal logs.
|
||||||
|
eval_only: If True, skip ingestion and only run query+judge using
|
||||||
|
existing workspaces.
|
||||||
|
"""
|
||||||
|
from multiprocessing import Pool # pylint: disable=import-outside-toplevel
|
||||||
|
|
||||||
|
# Load config BEFORE logging setup so log_dir is available
|
||||||
|
eval_config = load_eval_config(config_path)
|
||||||
|
|
||||||
|
# Resolve per-run log directory from config
|
||||||
|
output_cfg = eval_config.get("output", {})
|
||||||
|
log_dir_abs = None
|
||||||
|
if output_cfg.get("log_to_file", False):
|
||||||
|
log_dir_raw = output_cfg.get("log_dir", "logs")
|
||||||
|
log_prefix = output_cfg.get("log_prefix", "beam")
|
||||||
|
run_ts = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
|
||||||
|
log_dir_abs = str(_PROJECT_ROOT / log_dir_raw / f"{log_prefix}_{run_ts}")
|
||||||
|
|
||||||
|
setup_logging(log_level, reme_log_level, log_dir=log_dir_abs)
|
||||||
|
dataset_cfg = eval_config["dataset"]
|
||||||
|
chat_size = dataset_cfg["chat_size"]
|
||||||
|
beam_root = _PROJECT_ROOT / dataset_cfg.get("beam_root", "benchmark/beam/dataset/BEAM")
|
||||||
|
|
||||||
|
# Determine which cases to run
|
||||||
|
case_ids = dataset_cfg.get("case_ids") or []
|
||||||
|
if not case_ids:
|
||||||
|
case_ids = get_available_cases(beam_root, chat_size)
|
||||||
|
|
||||||
|
# Pagination
|
||||||
|
start = dataset_cfg.get("start_index", 0)
|
||||||
|
num_items = dataset_cfg.get("num_items", 0)
|
||||||
|
if num_items > 0:
|
||||||
|
case_ids = case_ids[start : start + num_items]
|
||||||
|
elif start > 0:
|
||||||
|
case_ids = case_ids[start:]
|
||||||
|
|
||||||
|
if not case_ids:
|
||||||
|
logger.error(f"No cases found for chat_size={chat_size}")
|
||||||
|
return
|
||||||
|
|
||||||
|
logger.info(
|
||||||
|
"Evaluating %d case(s) for chat_size=%s: %s%s",
|
||||||
|
len(case_ids),
|
||||||
|
chat_size,
|
||||||
|
case_ids,
|
||||||
|
" [eval_only: query+judge only]" if eval_only else "",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Resolve parallelism
|
||||||
|
num_workers = _resolve_num_workers(eval_config["evaluation"].get("num_workers", 1))
|
||||||
|
logger.info(f"Using {num_workers} worker(s)")
|
||||||
|
|
||||||
|
# Create output directory
|
||||||
|
output_dir = _PROJECT_ROOT / output_cfg.get("dir", "benchmark/beam/results")
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Create workspace root directory
|
||||||
|
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
|
||||||
|
workspace_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Pre-check: verify all workspaces exist in eval_only mode
|
||||||
|
if eval_only:
|
||||||
|
missing_cases = []
|
||||||
|
for case_id in case_ids:
|
||||||
|
case_dir = workspace_root / f"{chat_size}_{case_id}"
|
||||||
|
if not case_dir.exists() or not (case_dir / ".reme").exists():
|
||||||
|
missing_cases.append(case_id)
|
||||||
|
if missing_cases:
|
||||||
|
preview = missing_cases[:10]
|
||||||
|
suffix = "..." if len(missing_cases) > 10 else ""
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"eval_only: {len(missing_cases)} workspace(s) not found under {workspace_root}. "
|
||||||
|
f"Missing cases: {preview}{suffix}. "
|
||||||
|
f"Run without --eval_only first to build the workspaces.",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Build task args
|
||||||
|
task_args = [(eval_config, case_id, log_level, reme_log_level, eval_only, log_dir_abs) for case_id in case_ids]
|
||||||
|
|
||||||
|
# Progress tracking
|
||||||
|
total_items = len(task_args)
|
||||||
|
completed_count = [0]
|
||||||
|
start_time = time.time()
|
||||||
|
progress_lock = threading.Lock()
|
||||||
|
|
||||||
|
def _print_progress(prefix: str = "PROGRESS"):
|
||||||
|
elapsed = time.time() - start_time
|
||||||
|
elapsed_min = elapsed / 60
|
||||||
|
done = completed_count[0]
|
||||||
|
pct = 100.0 * done / total_items if total_items else 0
|
||||||
|
eta_str = "N/A"
|
||||||
|
if done > 0:
|
||||||
|
eta_sec = elapsed / done * (total_items - done)
|
||||||
|
eta_str = f"{eta_sec/60:.1f}min"
|
||||||
|
print(
|
||||||
|
f"[{prefix}] {datetime.now().strftime('%Y-%m-%d %H:%M:%S')} | "
|
||||||
|
f"{done}/{total_items} ({pct:.1f}%) completed | "
|
||||||
|
f"elapsed={elapsed_min:.1f}min | ETA={eta_str}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _progress_timer():
|
||||||
|
"""Background thread: print progress every 10 minutes."""
|
||||||
|
while not _timer_stop.is_set():
|
||||||
|
_timer_stop.wait(600)
|
||||||
|
if not _timer_stop.is_set():
|
||||||
|
with progress_lock:
|
||||||
|
_print_progress()
|
||||||
|
|
||||||
|
_timer_stop = threading.Event()
|
||||||
|
timer_thread = threading.Thread(target=_progress_timer, daemon=True)
|
||||||
|
timer_thread.start()
|
||||||
|
|
||||||
|
# Run evaluation
|
||||||
|
if num_workers == 1:
|
||||||
|
results = []
|
||||||
|
for task_input in task_args:
|
||||||
|
result = _evaluate_case_worker(task_input)
|
||||||
|
results.append(result)
|
||||||
|
with progress_lock:
|
||||||
|
completed_count[0] += 1
|
||||||
|
else:
|
||||||
|
results = [None] * total_items
|
||||||
|
indexed_args = list(enumerate(task_args))
|
||||||
|
|
||||||
|
with Pool(processes=num_workers) as pool:
|
||||||
|
for idx, result in pool.imap_unordered(_indexed_worker, indexed_args):
|
||||||
|
results[idx] = result
|
||||||
|
with progress_lock:
|
||||||
|
completed_count[0] += 1
|
||||||
|
|
||||||
|
# Stop progress timer
|
||||||
|
_timer_stop.set()
|
||||||
|
timer_thread.join(timeout=2)
|
||||||
|
|
||||||
|
# Save results
|
||||||
|
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
|
||||||
|
output_file = output_dir / f"results_{chat_size}_{timestamp}.json"
|
||||||
|
with open(output_file, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(results, f, ensure_ascii=False, indent=2)
|
||||||
|
logger.info(f"Results saved to {output_file}")
|
||||||
|
|
||||||
|
# Final progress
|
||||||
|
_print_progress("FINAL")
|
||||||
|
|
||||||
|
# Print concise summary
|
||||||
|
print("\n" + "=" * 70)
|
||||||
|
print(f" BEAM EVALUATION RESULTS | size={chat_size} cases={len(results)}")
|
||||||
|
print("=" * 70)
|
||||||
|
|
||||||
|
# Per-type stats (agentic only)
|
||||||
|
type_scores: dict[str, list[float]] = {}
|
||||||
|
type_binary_scores: dict[str, list[float]] = {}
|
||||||
|
all_scores: list[float] = []
|
||||||
|
all_binary_scores: list[float] = []
|
||||||
|
all_tool_call_totals: list[int] = []
|
||||||
|
all_token_usages: list[dict[str, int | None]] = []
|
||||||
|
all_memory_token_usages: list[dict[str, dict[str, int | None]]] = []
|
||||||
|
|
||||||
|
for case_result in results:
|
||||||
|
if "error" in case_result:
|
||||||
|
continue
|
||||||
|
mem_usage = case_result.get("memory_token_usage", {})
|
||||||
|
if mem_usage:
|
||||||
|
all_memory_token_usages.append(mem_usage)
|
||||||
|
for q in case_result.get("questions", []):
|
||||||
|
judgment = q.get("agentic_judgment", {})
|
||||||
|
score = judgment.get("llm_judge_score", 0.0)
|
||||||
|
# Binary: convert each rubric item score to 0/1, then average
|
||||||
|
judge_responses = judgment.get("llm_judge_responses", [])
|
||||||
|
if judge_responses:
|
||||||
|
binary_scores_per_item = [1.0 if r.get("score", 0) >= 1.0 else 0.0 for r in judge_responses]
|
||||||
|
binary_score = sum(binary_scores_per_item) / len(binary_scores_per_item)
|
||||||
|
else:
|
||||||
|
binary_score = 1.0 if score > 0.99 else 0.0
|
||||||
|
qtype = q["question_type"]
|
||||||
|
if qtype not in type_scores:
|
||||||
|
type_scores[qtype] = []
|
||||||
|
type_binary_scores[qtype] = []
|
||||||
|
type_scores[qtype].append(score)
|
||||||
|
type_binary_scores[qtype].append(binary_score)
|
||||||
|
all_scores.append(score)
|
||||||
|
all_binary_scores.append(binary_score)
|
||||||
|
metadata = q.get("agentic_metadata", {})
|
||||||
|
all_tool_call_totals.append(sum(metadata.get("tool_counts", {}).values()))
|
||||||
|
all_token_usages.append(metadata.get("token_usage", {}))
|
||||||
|
|
||||||
|
# Memory construction token usage summary
|
||||||
|
if all_memory_token_usages:
|
||||||
|
print("\n ── Memory Construction Token Usage ──")
|
||||||
|
for agent_name in ("default", "bench"):
|
||||||
|
for metric in _TOKEN_USAGE_METRICS:
|
||||||
|
values = [
|
||||||
|
usage[agent_name][metric]
|
||||||
|
for usage in all_memory_token_usages
|
||||||
|
if usage.get(agent_name, {}).get(metric) is not None
|
||||||
|
]
|
||||||
|
if values:
|
||||||
|
total = sum(values)
|
||||||
|
mean, std = _mean_and_std(values)
|
||||||
|
print(
|
||||||
|
f" {agent_name}/{metric}: total={total} mean={mean:.2f} std={std:.2f} ({len(values)} cases)",
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
print(f" {agent_name}/{metric}: unavailable")
|
||||||
|
print()
|
||||||
|
|
||||||
|
print("\n ── AGENTIC ──")
|
||||||
|
if all_scores:
|
||||||
|
for qtype in sorted(type_scores.keys()):
|
||||||
|
scores = type_scores[qtype]
|
||||||
|
avg = sum(scores) / len(scores) if scores else 0
|
||||||
|
bin_scores = type_binary_scores[qtype]
|
||||||
|
bin_avg = sum(bin_scores) / len(bin_scores) if bin_scores else 0
|
||||||
|
print(f" {qtype:<40s}: {avg:.3f} binary={bin_avg:.3f} ({len(scores)} Qs)")
|
||||||
|
overall = sum(all_scores) / len(all_scores) if all_scores else 0
|
||||||
|
binary_overall = sum(all_binary_scores) / len(all_binary_scores) if all_binary_scores else 0
|
||||||
|
print(f" {'-'*38}")
|
||||||
|
print(f" {'OVERALL':<40s}: {overall:.3f} binary={binary_overall:.3f} ({len(all_scores)} Qs)")
|
||||||
|
tool_call_mean, tool_call_std = _mean_and_std(all_tool_call_totals)
|
||||||
|
print(f" Tool calls/query: mean={tool_call_mean:.2f} std={tool_call_std:.2f}")
|
||||||
|
print(" Bench reported tokens/query:")
|
||||||
|
for metric in _TOKEN_USAGE_METRICS:
|
||||||
|
values = [usage[metric] for usage in all_token_usages if usage.get(metric) is not None]
|
||||||
|
if values:
|
||||||
|
mean, std = _mean_and_std(values)
|
||||||
|
print(f" {metric}: mean={mean:.2f} std={std:.2f}")
|
||||||
|
else:
|
||||||
|
print(f" {metric}: unavailable")
|
||||||
|
else:
|
||||||
|
print(" (no results)")
|
||||||
|
|
||||||
|
# Per-case summary
|
||||||
|
print("\n ── Per-Case Summary ──")
|
||||||
|
for case_result in results:
|
||||||
|
case_id = case_result["case_id"]
|
||||||
|
if "error" in case_result:
|
||||||
|
print(f" Case {case_id}: ERROR — {case_result['error']}")
|
||||||
|
continue
|
||||||
|
n_qs = case_result.get("total_questions", 0)
|
||||||
|
n_sessions = case_result.get("sessions_ingested", 0)
|
||||||
|
mem_usage = case_result.get("memory_token_usage", {})
|
||||||
|
parts = [f"Case {case_id}: {n_sessions} sessions, {n_qs} questions"]
|
||||||
|
# Append memory construction total tokens if available
|
||||||
|
for agent_name in ("default", "bench"):
|
||||||
|
agent_usage = mem_usage.get(agent_name, {})
|
||||||
|
total = agent_usage.get("total_tokens")
|
||||||
|
if total is not None:
|
||||||
|
parts.append(f"mem_{agent_name}_tokens={total}")
|
||||||
|
questions = case_result.get("questions", [])
|
||||||
|
scores = [q.get("agentic_judgment", {}).get("llm_judge_score", 0.0) for q in questions]
|
||||||
|
if scores:
|
||||||
|
avg = sum(scores) / len(scores)
|
||||||
|
# Binary: 0/1 per rubric item, average per question, then across questions
|
||||||
|
bin_scores = []
|
||||||
|
for q in questions:
|
||||||
|
judge_responses = q.get("agentic_judgment", {}).get("llm_judge_responses", [])
|
||||||
|
if judge_responses:
|
||||||
|
item_bins = [1.0 if r.get("score", 0) >= 1.0 else 0.0 for r in judge_responses]
|
||||||
|
bin_scores.append(sum(item_bins) / len(item_bins))
|
||||||
|
else:
|
||||||
|
s = q.get("agentic_judgment", {}).get("llm_judge_score", 0.0)
|
||||||
|
bin_scores.append(1.0 if s > 0.99 else 0.0)
|
||||||
|
bin_avg = sum(bin_scores) / len(bin_scores)
|
||||||
|
parts.append(f"agentic={avg:.3f} binary={bin_avg:.3f}")
|
||||||
|
print(f" {' | '.join(parts)}")
|
||||||
|
|
||||||
|
print("=" * 70)
|
||||||
|
total_elapsed = time.time() - start_time
|
||||||
|
print(f"\n Total time: {total_elapsed/60:.1f} min")
|
||||||
|
print("\n" + "=" * 70)
|
||||||
|
print(" [DONE] BEAM EVALUATION COMPLETED SUCCESSFULLY")
|
||||||
|
print("=" * 70 + "\n")
|
||||||
|
|
||||||
|
|
||||||
|
_TOKEN_USAGE_METRICS = (
|
||||||
|
"input_tokens",
|
||||||
|
"output_tokens",
|
||||||
|
"total_tokens",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _mean_and_std(values: list[int]) -> tuple[float, float]:
|
||||||
|
"""Return population mean and standard deviation for one per-question metric."""
|
||||||
|
if not values:
|
||||||
|
return 0.0, 0.0
|
||||||
|
mean = sum(values) / len(values)
|
||||||
|
return mean, (sum((value - mean) ** 2 for value in values) / len(values)) ** 0.5
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
import argparse
|
||||||
|
|
||||||
|
parser = argparse.ArgumentParser(description="BEAM evaluation runner")
|
||||||
|
parser.add_argument("--config", type=str, default=None, help="Path to config.yaml")
|
||||||
|
parser.add_argument(
|
||||||
|
"--log-level",
|
||||||
|
type=str,
|
||||||
|
default="INFO",
|
||||||
|
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
|
||||||
|
help="Log level for the eval runner (default: INFO)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--reme-log-level",
|
||||||
|
type=str,
|
||||||
|
default="INFO",
|
||||||
|
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
|
||||||
|
help="Log level for reme internal logs — loguru (default: INFO)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-q",
|
||||||
|
"--quiet",
|
||||||
|
action="store_true",
|
||||||
|
help="Shortcut for --log-level WARNING --reme-log-level WARNING",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--eval_only",
|
||||||
|
action="store_true",
|
||||||
|
help="Skip ingestion. Reuse existing workspaces and only run query+judge.",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
if args.quiet:
|
||||||
|
args.log_level = "WARNING"
|
||||||
|
args.reme_log_level = "WARNING"
|
||||||
|
|
||||||
|
main(args.config, args.log_level, args.reme_log_level, eval_only=args.eval_only)
|
||||||
110
benchmark/longmemeval/README.md
Normal file
110
benchmark/longmemeval/README.md
Normal file
|
|
@ -0,0 +1,110 @@
|
||||||
|
[中文版 / Chinese version](./README_ZH.md)
|
||||||
|
|
||||||
|
# LongMemEval Benchmark
|
||||||
|
|
||||||
|
LongMemEval is a benchmark for **long-term memory over multi-session chat
|
||||||
|
histories**. Each item provides a chronologically ordered set of chat sessions
|
||||||
|
between a user and an assistant, followed by a probing question whose answer is
|
||||||
|
only recoverable by reasoning over the user-owned memory. ReMe ingests the
|
||||||
|
sessions into an isolated per-item workspace, answers the question via an
|
||||||
|
agentic (ReAct) mode, and scores the answer with an LLM-as-judge.
|
||||||
|
|
||||||
|
Question types include single-session (user / assistant / preference),
|
||||||
|
multi-session reasoning, knowledge update, and temporal reasoning.
|
||||||
|
|
||||||
|
Install ReMe and the LongMemEval plugin in editable mode from the repository root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m pip install -e ".[as]"
|
||||||
|
reme plugins install ./plugins/lme --editable
|
||||||
|
reme plugins install ./plugins/lme-judge --editable
|
||||||
|
reme plugins validate lme
|
||||||
|
```
|
||||||
|
|
||||||
|
The runner explicitly enables the installed `lme` plugin and combines its defaults with
|
||||||
|
ReMe's built-in `benchmark` preset. Editable installation keeps changes under
|
||||||
|
[`plugins/lme`](../../plugins/lme/README.md) visible without reinstalling the plugin.
|
||||||
|
Custom application config paths still work through `reme.config` and can use `extends: benchmark`.
|
||||||
|
This directory continues to own the runner, evaluation settings, dataset and outputs.
|
||||||
|
Model credentials use the environment variables declared by the shared benchmark configuration.
|
||||||
|
|
||||||
|
## 1. Get the Dataset
|
||||||
|
|
||||||
|
ReMe uses only the **cleaned-S** split, hosted on HuggingFace:
|
||||||
|
[agentscope-ai/ReMe_longmemeval_clean_s_v2](https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2).
|
||||||
|
The download script fetches it via the hf-mirror.com mirror; to use a different
|
||||||
|
mirror, modify `BASE_URL` in [`download.py`](./download.py).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd benchmark/longmemeval
|
||||||
|
python download.py # saves dataset/longmemeval_s_reme_cleaned.json; skips if already present
|
||||||
|
```
|
||||||
|
|
||||||
|
Ground truth is embedded in the data file.
|
||||||
|
|
||||||
|
## 2. Run
|
||||||
|
|
||||||
|
From the repository root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python benchmark/longmemeval/run.py
|
||||||
|
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
|
||||||
|
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
|
||||||
|
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
|
||||||
|
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
|
||||||
|
python benchmark/longmemeval/run.py --eval_only # reuse existing workspaces, query + judge only
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. Pipeline
|
||||||
|
|
||||||
|
1. Load the dataset (ground truth is embedded in the data file).
|
||||||
|
2. For each item, create an isolated workspace and ingest sessions in chronological order.
|
||||||
|
3. If a custom application configuration enables `auto_dream`, trigger it when sessions cross the configured hour
|
||||||
|
(default 23:00). The packaged preset leaves it disabled.
|
||||||
|
4. Answer each question via agentic (ReAct) mode.
|
||||||
|
5. Judge the answer (binary yes/no) with the `answer_judge` job and print per-type accuracy.
|
||||||
|
|
||||||
|
## 4. Key config — `benchmark/longmemeval/config.yaml`
|
||||||
|
|
||||||
|
| Key | Meaning |
|
||||||
|
| --- | --- |
|
||||||
|
| `dataset.path` | Dataset file to evaluate (e.g. `longmemeval_s_reme_cleaned.json`); ground truth is included. |
|
||||||
|
| `dataset.start_index` / `num_items` | Slice of items to evaluate. |
|
||||||
|
| `dataset.question_types` | Filter by question type; empty = all. |
|
||||||
|
| `dataset.workspace_root` | Per-item workspace root (`benchmark/longmemeval/workspaces/longmemeval-s`). |
|
||||||
|
| `evaluation.num_workers` | `0` = auto (cpu-2), `1` = sequential, `>1` = parallel. |
|
||||||
|
| `evaluation.filter_future_sessions` | Only ingest sessions with timestamp ≤ `question_date`. |
|
||||||
|
| `reme.config` | ReMe config used (`benchmark`). |
|
||||||
|
| `reme.dream_trigger_hour` / `dream_scan_days` / `dream_max_units` | Dream triggering behavior. |
|
||||||
|
| `output.dir` | Results directory (`benchmark/longmemeval/results`). |
|
||||||
|
|
||||||
|
## 5. Outputs
|
||||||
|
|
||||||
|
Results are JSON files written to `output.dir` as `results_<timestamp>.json`,
|
||||||
|
with a per-type accuracy summary also printed to the console. Logging
|
||||||
|
conventions are shared across benchmarks — see the
|
||||||
|
[top-level README](../README.md#outputs--logs).
|
||||||
|
|
||||||
|
## 6. Reference Results
|
||||||
|
|
||||||
|
### cleaned-s
|
||||||
|
|
||||||
|
**Basic settings**
|
||||||
|
|
||||||
|
1. Modified auto-memory prompt, auto-dream disabled.
|
||||||
|
2. All sessions in reme-memory are strictly earlier than the question time.
|
||||||
|
|
||||||
|
**Results**
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1, conda reme env, 32 workers, eval-only (reusing prebuilt memory)
|
||||||
|
(2026-08-06, 500 items, total 10.0 min)
|
||||||
|
|
||||||
|
| Type | Agentic | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| knowledge-update | 0.910 | 31,581 | 589 | 32,169 | 2.90 |
|
||||||
|
| multi-session | 0.842 | 52,837 | 1,474 | 54,311 | 4.21 |
|
||||||
|
| single-session-assistant | 1.000 | 15,596 | 279 | 15,875 | 1.89 |
|
||||||
|
| single-session-preference | 0.633 | 36,802 | 818 | 37,620 | 3.60 |
|
||||||
|
| single-session-user | 0.986 | 27,433 | 359 | 27,792 | 2.60 |
|
||||||
|
| temporal-reasoning | 0.902 | 62,674 | 985 | 63,659 | 4.97 |
|
||||||
|
| **OVERALL** | **0.894** | **43,448** | **876** | **44,324** | **3.69** |
|
||||||
103
benchmark/longmemeval/README_ZH.md
Normal file
103
benchmark/longmemeval/README_ZH.md
Normal file
|
|
@ -0,0 +1,103 @@
|
||||||
|
# LongMemEval 评测
|
||||||
|
|
||||||
|
[English version](./README.md)
|
||||||
|
|
||||||
|
LongMemEval 是一个面向**多轮多会话历史的长期记忆能力**的评测基准。每个条目提供一组按时间
|
||||||
|
顺序排列的用户与助手之间的会话,以及一个只能通过推理用户自有记忆才能回答的探测问题。ReMe
|
||||||
|
将会话摄入按条目隔离的工作区,以 agentic(ReAct)模式回答问题,最后由 LLM-as-judge 打分。
|
||||||
|
|
||||||
|
题型包括单会话(user / assistant / preference)、多会话推理、知识更新与时间推理等。
|
||||||
|
|
||||||
|
在仓库根目录以 editable 模式安装 ReMe 和 LongMemEval 插件:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m pip install -e ".[as]"
|
||||||
|
reme plugins install ./plugins/lme --editable
|
||||||
|
reme plugins install ./plugins/lme-judge --editable
|
||||||
|
reme plugins validate lme
|
||||||
|
```
|
||||||
|
|
||||||
|
runner 显式启用已安装的 `lme` 插件,并将插件默认配置与 ReMe 内置的 `benchmark` 配置组合。
|
||||||
|
editable 安装会让 [`plugins/lme`](../../plugins/lme/README_ZH.md) 下的源码修改直接生效,无需重复安装。
|
||||||
|
本目录继续保留评测参数、数据集及输出。自定义完整应用配置路径仍可通过 `reme.config` 指定,
|
||||||
|
并可使用 `extends: benchmark`。
|
||||||
|
模型凭据通过公共 benchmark 配置中声明的环境变量设置。
|
||||||
|
|
||||||
|
## 1. 获取数据集
|
||||||
|
|
||||||
|
ReMe 仅使用 **cleaned-S** 版本,数据托管在 HuggingFace:
|
||||||
|
[agentscope-ai/ReMe_longmemeval_clean_s_v2](https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2)。
|
||||||
|
下载脚本经 hf-mirror.com 镜像源获取,如需更换源请修改 [`download.py`](./download.py) 中的
|
||||||
|
`BASE_URL`。
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd benchmark/longmemeval
|
||||||
|
python download.py # 保存为 dataset/longmemeval_s_reme_cleaned.json,已存在则自动跳过
|
||||||
|
```
|
||||||
|
|
||||||
|
ground truth 已内嵌在数据文件中。
|
||||||
|
|
||||||
|
## 2. 运行
|
||||||
|
|
||||||
|
在仓库根目录执行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python benchmark/longmemeval/run.py
|
||||||
|
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
|
||||||
|
python benchmark/longmemeval/run.py -q # 安静模式:仅评测级日志
|
||||||
|
python benchmark/longmemeval/run.py --log-level WARNING # 降低评测 runner 日志
|
||||||
|
python benchmark/longmemeval/run.py --reme-log-level WARNING # 降低 reme 内部日志
|
||||||
|
python benchmark/longmemeval/run.py --eval_only # 复用已有工作区,仅执行查询 + 评判
|
||||||
|
```
|
||||||
|
|
||||||
|
## 3. 流程
|
||||||
|
|
||||||
|
1. 加载数据集(ground truth 已内嵌在数据文件中)。
|
||||||
|
2. 为每个条目创建独立工作区,按时间顺序摄入会话。
|
||||||
|
3. 若自定义应用配置启用了 `auto_dream`,在相邻会话跨越配置时刻(默认 23:00)时触发;插件预设保持关闭。
|
||||||
|
4. 以 agentic(ReAct)模式回答每个问题。
|
||||||
|
5. 通过 `answer_judge` 任务对答案做二元(yes/no)评判,并输出各类型准确率。
|
||||||
|
|
||||||
|
## 4. 关键配置 —— `benchmark/longmemeval/config.yaml`
|
||||||
|
|
||||||
|
| 配置项 | 含义 |
|
||||||
|
| --- | --- |
|
||||||
|
| `dataset.path` | 待评测的数据集文件(如 `longmemeval_s_reme_cleaned.json`),已包含 ground truth。 |
|
||||||
|
| `dataset.start_index` / `num_items` | 评测条目的切片范围。 |
|
||||||
|
| `dataset.question_types` | 按问题类型过滤,空表示全部。 |
|
||||||
|
| `dataset.workspace_root` | 条目工作区根目录(`benchmark/longmemeval/workspaces/longmemeval-s`)。 |
|
||||||
|
| `evaluation.num_workers` | `0` = 自动(cpu-2),`1` = 串行,`>1` = 并行。 |
|
||||||
|
| `evaluation.filter_future_sessions` | 仅摄入时间戳 ≤ `question_date` 的会话。 |
|
||||||
|
| `reme.config` | 使用的 ReMe 配置(`benchmark`)。 |
|
||||||
|
| `reme.dream_trigger_hour` / `dream_scan_days` / `dream_max_units` | dream 触发行为。 |
|
||||||
|
| `output.dir` | 结果目录(`benchmark/longmemeval/results`)。 |
|
||||||
|
|
||||||
|
## 5. 输出
|
||||||
|
|
||||||
|
结果以 JSON 文件写入 `output.dir`,文件名为 `results_<timestamp>.json`,
|
||||||
|
同时控制台会打印含各类型准确率的汇总。日志约定在各基准间通用,见
|
||||||
|
[总说明](../README_ZH.md#输出与日志)。
|
||||||
|
|
||||||
|
## 6. 参考结果
|
||||||
|
|
||||||
|
### cleaned-s
|
||||||
|
|
||||||
|
**基础设置**
|
||||||
|
|
||||||
|
1. 使用修改后的 auto-memory prompt,关闭 auto-dream 机制
|
||||||
|
2. reme-memory 中的全部 session 的时间一定早于 question 的时间
|
||||||
|
|
||||||
|
**结果**
|
||||||
|
|
||||||
|
agentscope==2.0.4.post1, conda reme env, 32 workers, eval-only(复用预构建记忆)
|
||||||
|
(2026-08-06,500 题,总计 10.0 min)
|
||||||
|
|
||||||
|
| 类型 | Agentic | input tok/q | output tok/q | total tok/q | tool calls/q |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| knowledge-update | 0.910 | 31,581 | 589 | 32,169 | 2.90 |
|
||||||
|
| multi-session | 0.842 | 52,837 | 1,474 | 54,311 | 4.21 |
|
||||||
|
| single-session-assistant | 1.000 | 15,596 | 279 | 15,875 | 1.89 |
|
||||||
|
| single-session-preference | 0.633 | 36,802 | 818 | 37,620 | 3.60 |
|
||||||
|
| single-session-user | 0.986 | 27,433 | 359 | 27,792 | 2.60 |
|
||||||
|
| temporal-reasoning | 0.902 | 62,674 | 985 | 63,659 | 4.97 |
|
||||||
|
| **OVERALL** | **0.894** | **43,448** | **876** | **44,324** | **3.69** |
|
||||||
|
|
@ -1,141 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Remove generated LongMemEval files while keeping source inputs.
|
|
||||||
|
|
||||||
For each ``datasets/longmemeval/<idx>`` workspace, this keeps only:
|
|
||||||
- query.json
|
|
||||||
- answer.json
|
|
||||||
- session/
|
|
||||||
|
|
||||||
All other files or directories in the sample root are considered generated
|
|
||||||
artifacts and can be removed. AppleDouble files whose names start with ``._``
|
|
||||||
are also removed recursively, including under ``session/``. The script is
|
|
||||||
dry-run by default; pass ``--apply`` to actually delete. To delete only specific
|
|
||||||
root-level generated files, pass one or more ``--filename`` values.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/clean_sample_outputs.py
|
|
||||||
python benchmark/longmemeval/clean_sample_outputs.py --apply
|
|
||||||
python benchmark/longmemeval/clean_sample_outputs.py --start 36 --end 79 --apply
|
|
||||||
python benchmark/longmemeval/clean_sample_outputs.py --filename check_golden.json --apply
|
|
||||||
python benchmark/longmemeval/clean_sample_outputs.py --filename session_review.json --apply
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import shutil
|
|
||||||
import time
|
|
||||||
from collections.abc import Iterator
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
KEEP = {"query.json", "answer.json", "session"}
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument("--start", type=int, default=0, help="first numeric sample id to clean, inclusive (default 0)")
|
|
||||||
p.add_argument("--end", type=int, default=499, help="last numeric sample id to clean, inclusive (default 499)")
|
|
||||||
p.add_argument("--limit", type=int, default=0, help="only clean the first N selected samples (0 = all)")
|
|
||||||
p.add_argument("--progress-every", type=int, default=25, help="print progress every N samples when applying")
|
|
||||||
p.add_argument(
|
|
||||||
"--filename",
|
|
||||||
action="append",
|
|
||||||
default=[],
|
|
||||||
help="delete only this root-level file or directory name; can be passed multiple times",
|
|
||||||
)
|
|
||||||
p.add_argument("--apply", action="store_true", help="actually delete files; default is dry-run")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all numeric sample IDs."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def delete_path(path: Path) -> None:
|
|
||||||
"""Delete a file, symlink, or directory."""
|
|
||||||
if path.is_dir() and not path.is_symlink():
|
|
||||||
shutil.rmtree(path)
|
|
||||||
else:
|
|
||||||
path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def iter_sample_targets(sample_dir: Path, filenames: set[str] | None = None) -> Iterator[Path]:
|
|
||||||
"""Yield generated artifacts for one sample.
|
|
||||||
|
|
||||||
Root-level generated directories are yielded as a whole, so there is no
|
|
||||||
need to recurse into them. AppleDouble files are only searched inside the
|
|
||||||
kept ``session/`` directory.
|
|
||||||
"""
|
|
||||||
if filenames:
|
|
||||||
for name in sorted(filenames):
|
|
||||||
path = sample_dir / name
|
|
||||||
if path.exists():
|
|
||||||
yield path
|
|
||||||
return
|
|
||||||
|
|
||||||
for path in sorted(sample_dir.iterdir(), key=lambda p: p.name):
|
|
||||||
if path.name not in KEEP:
|
|
||||||
yield path
|
|
||||||
|
|
||||||
session_dir = sample_dir / "session"
|
|
||||||
if session_dir.is_dir():
|
|
||||||
yield from session_dir.rglob("._*")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
"""Main entry point."""
|
|
||||||
args = parse_args()
|
|
||||||
if args.end < args.start:
|
|
||||||
raise ValueError(f"--end ({args.end}) must be >= --start ({args.start})")
|
|
||||||
filenames = {name.strip() for name in args.filename if name.strip()}
|
|
||||||
invalid_filenames = [name for name in filenames if Path(name).name != name]
|
|
||||||
if invalid_filenames:
|
|
||||||
raise ValueError(f"--filename only accepts root-level names, got: {invalid_filenames}")
|
|
||||||
|
|
||||||
ids = [idx for idx in sample_ids() if args.start <= int(idx) <= args.end]
|
|
||||||
if args.limit:
|
|
||||||
ids = ids[: args.limit]
|
|
||||||
|
|
||||||
total_targets = 0
|
|
||||||
deleted = 0
|
|
||||||
started_at = time.time()
|
|
||||||
for ordinal, idx in enumerate(ids, start=1):
|
|
||||||
sample_dir = DATA / idx
|
|
||||||
sample_started_at = time.time()
|
|
||||||
targets = list(iter_sample_targets(sample_dir, filenames=filenames))
|
|
||||||
total_targets += len(targets)
|
|
||||||
print(f"[sample {ordinal}/{len(ids)}] {idx} targets={len(targets)}", flush=True)
|
|
||||||
for path in targets:
|
|
||||||
if args.apply:
|
|
||||||
target_started_at = time.time()
|
|
||||||
print(f"[delete] {path}", flush=True)
|
|
||||||
delete_path(path)
|
|
||||||
deleted += 1
|
|
||||||
print(f"[deleted] {path} elapsed={time.time() - target_started_at:.1f}s", flush=True)
|
|
||||||
else:
|
|
||||||
print(f"[would-delete] {path}")
|
|
||||||
if args.apply and args.progress_every > 0 and (int(idx) + 1) % args.progress_every == 0:
|
|
||||||
elapsed = time.time() - started_at
|
|
||||||
print(
|
|
||||||
f"[progress] processed={ordinal}/{len(ids)} through={idx} " f"deleted={deleted} elapsed={elapsed:.1f}s",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
print(f"[sample-done] {idx} elapsed={time.time() - sample_started_at:.1f}s", flush=True)
|
|
||||||
|
|
||||||
mode = "DELETE" if args.apply else "DRY-RUN"
|
|
||||||
print(
|
|
||||||
f"{mode} LongMemEval generated artifacts: samples={len(ids)} "
|
|
||||||
f"targets={total_targets} deleted={deleted if args.apply else 0} range={args.start}..{args.end}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if not args.apply:
|
|
||||||
print("No files deleted. Re-run with --apply to delete these paths.", flush=True)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(main())
|
|
||||||
34
benchmark/longmemeval/config.yaml
Normal file
34
benchmark/longmemeval/config.yaml
Normal file
|
|
@ -0,0 +1,34 @@
|
||||||
|
# LongMemEval evaluation configuration
|
||||||
|
# This file controls what/how to evaluate.
|
||||||
|
|
||||||
|
dataset:
|
||||||
|
path: "benchmark/longmemeval/dataset/longmemeval_s_reme_cleaned.json"
|
||||||
|
start_index: 0 # first item index
|
||||||
|
num_items: 500 # how many items to evaluate (starting from start_index)
|
||||||
|
max_sessions: 0 # 0 = all sessions; >0 = limit sessions per item for testing
|
||||||
|
question_types: [] # filter by question_type; empty list = no filtering (all types)
|
||||||
|
workspace_root: "benchmark/longmemeval/workspaces/longmemeval-s" # workspace root for item workspaces
|
||||||
|
|
||||||
|
evaluation:
|
||||||
|
# LLM-as-judge uses the 'judge' as_llm component defined in benchmark.yaml
|
||||||
|
# Model and credentials are configured there (reading from .env)
|
||||||
|
# Judgment is always binary (yes/no) — defined in lme/llm_judge.yaml
|
||||||
|
num_workers: 32 # 0 = auto (cpu_count - 2, min 1); 1 = sequential; >1 = parallel
|
||||||
|
filter_future_sessions: true # true = only ingest sessions with timestamp <= question_date
|
||||||
|
compress_session: false # true = compress session chunks in search_v2 (query-aware); false = no compression
|
||||||
|
|
||||||
|
reme:
|
||||||
|
config: "benchmark" # runner enables both plugins below
|
||||||
|
plugins: [lme, lme-judge]
|
||||||
|
# Dream trigger: when gap between consecutive sessions crosses this hour (23:00)
|
||||||
|
dream_trigger_hour: 23
|
||||||
|
# Dream scan_days for each trigger
|
||||||
|
dream_scan_days: 2
|
||||||
|
dream_max_units: 5
|
||||||
|
|
||||||
|
output:
|
||||||
|
dir: "benchmark/longmemeval/results"
|
||||||
|
log_dir: "logs" # log directory (relative to project root)
|
||||||
|
log_prefix: "longmemeval" # benchmark name used in log filenames
|
||||||
|
log_to_console: true
|
||||||
|
log_to_file: true
|
||||||
67
benchmark/longmemeval/download.py
Normal file
67
benchmark/longmemeval/download.py
Normal file
|
|
@ -0,0 +1,67 @@
|
||||||
|
"""Download the LongMemEval cleaned-S dataset used by ReMe.
|
||||||
|
|
||||||
|
Source: https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2
|
||||||
|
(downloaded via the hf-mirror.com mirror for reliability).
|
||||||
|
|
||||||
|
The file ``longmemeval_s_reme_cleaned.json`` is saved under ``dataset/`` next to this
|
||||||
|
script using the same name as on the remote (``benchmark/longmemeval/config.yaml``
|
||||||
|
points to it).
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python download.py # download cleaned-S (skip if it already exists)
|
||||||
|
"""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
BASE_URL = "https://hf-mirror.com/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2/resolve/main"
|
||||||
|
TARGET_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "dataset")
|
||||||
|
|
||||||
|
# Files to download (saved with the same name as on the remote).
|
||||||
|
FILES = [
|
||||||
|
"longmemeval_s_reme_cleaned.json",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def download_file(filename: str):
|
||||||
|
"""Download a single file from the mirror to the target directory."""
|
||||||
|
url = f"{BASE_URL}/{filename}"
|
||||||
|
dest = os.path.join(TARGET_DIR, filename)
|
||||||
|
|
||||||
|
if os.path.exists(dest):
|
||||||
|
size = os.path.getsize(dest)
|
||||||
|
print(f" [skip] {filename} already exists ({size / 1024 / 1024:.1f} MB)")
|
||||||
|
return
|
||||||
|
|
||||||
|
print(f" [downloading] {filename} ...")
|
||||||
|
try:
|
||||||
|
urllib.request.urlretrieve(url, dest, reporthook=_progress)
|
||||||
|
size = os.path.getsize(dest)
|
||||||
|
print(f"\n [done] {filename} ({size / 1024 / 1024:.1f} MB)")
|
||||||
|
except Exception as e:
|
||||||
|
print(f"\n [error] {filename}: {e}")
|
||||||
|
if os.path.exists(dest):
|
||||||
|
os.remove(dest)
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
|
||||||
|
def _progress(block_num, block_size, total_size):
|
||||||
|
downloaded = block_num * block_size
|
||||||
|
if total_size > 0:
|
||||||
|
pct = min(100, downloaded * 100 / total_size)
|
||||||
|
mb = downloaded / 1024 / 1024
|
||||||
|
total_mb = total_size / 1024 / 1024
|
||||||
|
sys.stdout.write(f"\r {mb:.1f}/{total_mb:.1f} MB ({pct:.1f}%)")
|
||||||
|
else:
|
||||||
|
mb = downloaded / 1024 / 1024
|
||||||
|
sys.stdout.write(f"\r {mb:.1f} MB downloaded")
|
||||||
|
sys.stdout.flush()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
os.makedirs(TARGET_DIR, exist_ok=True)
|
||||||
|
print(f"Downloading LongMemEval cleaned-S dataset to: {TARGET_DIR}\n")
|
||||||
|
for fname in FILES:
|
||||||
|
download_file(fname)
|
||||||
|
print("\nAll files downloaded successfully!")
|
||||||
76
benchmark/longmemeval/kill.sh
Normal file
76
benchmark/longmemeval/kill.sh
Normal file
|
|
@ -0,0 +1,76 @@
|
||||||
|
#!/bin/bash
|
||||||
|
# 杀死指定进程及其所有子进程
|
||||||
|
# Usage: bash kill.sh <PID>
|
||||||
|
|
||||||
|
if [ -z "$1" ]; then
|
||||||
|
echo "Usage: bash kill.sh <PID>"
|
||||||
|
echo " 杀死指定进程及其所有子进程"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
PID=$1
|
||||||
|
|
||||||
|
# 检查进程是否存在
|
||||||
|
if ! kill -0 "$PID" 2>/dev/null; then
|
||||||
|
echo "进程 $PID 不存在"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# 递归收集所有子进程(包括子进程的子进程)
|
||||||
|
collect_children() {
|
||||||
|
local parent=$1
|
||||||
|
local children
|
||||||
|
children=$(ps -o pid= --ppid "$parent" 2>/dev/null | tr -d ' ')
|
||||||
|
for child in $children; do
|
||||||
|
collect_children "$child"
|
||||||
|
done
|
||||||
|
echo "$parent"
|
||||||
|
}
|
||||||
|
|
||||||
|
# 收集进程树(子进程在前,父进程在后,保证先杀子再杀父)
|
||||||
|
PROCESS_TREE=$(collect_children "$PID")
|
||||||
|
TOTAL=$(echo "$PROCESS_TREE" | wc -l | tr -d ' ')
|
||||||
|
|
||||||
|
echo "进程树(共 $TOTAL 个进程):"
|
||||||
|
while read -r p; do
|
||||||
|
cmd=$(ps -o args= -p "$p" 2>/dev/null | head -c 80)
|
||||||
|
printf " PID=%-8s %s\n" "$p" "$cmd"
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
# 先 SIGTERM 优雅终止
|
||||||
|
echo ""
|
||||||
|
echo "发送 SIGTERM..."
|
||||||
|
while read -r p; do
|
||||||
|
kill "$p" 2>/dev/null
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
# 等待最多 5 秒
|
||||||
|
for i in $(seq 1 5); do
|
||||||
|
alive=false
|
||||||
|
while read -r p; do
|
||||||
|
if kill -0 "$p" 2>/dev/null; then
|
||||||
|
alive=true
|
||||||
|
fi
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
if [ "$alive" = false ]; then
|
||||||
|
break
|
||||||
|
fi
|
||||||
|
sleep 1
|
||||||
|
done
|
||||||
|
|
||||||
|
# 检查是否还有残留,强制 SIGKILL
|
||||||
|
remaining=false
|
||||||
|
while read -r p; do
|
||||||
|
if kill -0 "$p" 2>/dev/null; then
|
||||||
|
remaining=true
|
||||||
|
fi
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
|
||||||
|
if [ "$remaining" = true ]; then
|
||||||
|
echo "部分进程未响应,发送 SIGKILL..."
|
||||||
|
while read -r p; do
|
||||||
|
kill -9 "$p" 2>/dev/null
|
||||||
|
done <<< "$PROCESS_TREE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
echo "已终止进程树(根 PID=$PID,共 $TOTAL 个进程)"
|
||||||
831
benchmark/longmemeval/run.py
Normal file
831
benchmark/longmemeval/run.py
Normal file
|
|
@ -0,0 +1,831 @@
|
||||||
|
"""LongMemEval evaluation runner for ReMe.
|
||||||
|
|
||||||
|
Evaluates ReMe's long-term memory capability using the LongMemEval dataset.
|
||||||
|
Each item gets an isolated workspace; sessions are ingested in chronological order;
|
||||||
|
dream is triggered when sessions cross midnight (23:00); finally questions are
|
||||||
|
answered via an agentic (ReAct) approach and judged by an LLM.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python benchmark/longmemeval/run.py
|
||||||
|
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
|
||||||
|
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
|
||||||
|
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
|
||||||
|
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
|
||||||
|
python benchmark/longmemeval/run.py --eval_only # query+judge only, reuse existing workspace
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import time
|
||||||
|
import threading
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
from dotenv import load_dotenv
|
||||||
|
|
||||||
|
# Load .env from project root
|
||||||
|
_PROJECT_ROOT = Path(__file__).resolve().parent.parent.parent
|
||||||
|
load_dotenv(_PROJECT_ROOT / ".env")
|
||||||
|
|
||||||
|
# Workspace root for evaluation items — read from config.yaml (dataset.workspace_root)
|
||||||
|
_WORKSPACE_ROOT_DEFAULT = "benchmark/longmemeval/workspaces/longmemeval-s"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Logging
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
_DEFAULT_LOG_FORMAT = "%(asctime)s | %(levelname)s | %(message)s"
|
||||||
|
|
||||||
|
logging.basicConfig(level=logging.INFO, format=_DEFAULT_LOG_FORMAT)
|
||||||
|
logger = logging.getLogger("longmemeval")
|
||||||
|
|
||||||
|
# Noisy library loggers silenced by default
|
||||||
|
_NOISY_LOGGERS = [
|
||||||
|
"httpx",
|
||||||
|
"httpcore",
|
||||||
|
"openai",
|
||||||
|
"uvicorn",
|
||||||
|
"multipart",
|
||||||
|
"asyncio",
|
||||||
|
"watchfiles",
|
||||||
|
"filelock",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def setup_logging(
|
||||||
|
log_level: str,
|
||||||
|
reme_log_level: str,
|
||||||
|
log_dir: str | None = None,
|
||||||
|
):
|
||||||
|
"""Configure logging for the eval runner and reme internals.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
log_level: Level for the eval runner logger (DEBUG/INFO/WARNING/ERROR).
|
||||||
|
reme_log_level: Level for reme's internal loguru logger.
|
||||||
|
log_dir: Per-run log directory (absolute path). None = no file logging.
|
||||||
|
"""
|
||||||
|
numeric = getattr(logging, log_level.upper(), logging.INFO)
|
||||||
|
# Eval runner logger
|
||||||
|
logging.getLogger().setLevel(numeric)
|
||||||
|
logger.setLevel(numeric)
|
||||||
|
|
||||||
|
# Suppress noisy library loggers when above DEBUG
|
||||||
|
if numeric > logging.DEBUG:
|
||||||
|
for name in _NOISY_LOGGERS:
|
||||||
|
lib_logger = logging.getLogger(name)
|
||||||
|
lib_logger.setLevel(max(numeric, logging.WARNING))
|
||||||
|
|
||||||
|
# Add file handler for eval runner if log_dir is specified
|
||||||
|
if log_dir:
|
||||||
|
os.makedirs(log_dir, exist_ok=True)
|
||||||
|
log_filepath = os.path.join(log_dir, "runner.log")
|
||||||
|
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
|
||||||
|
file_handler.setLevel(numeric)
|
||||||
|
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
|
||||||
|
logging.getLogger().addHandler(file_handler)
|
||||||
|
logger.info(f"Eval runner log file: {log_filepath}")
|
||||||
|
|
||||||
|
# Reme internal logger (loguru) — will be applied per-worker via _configure_worker
|
||||||
|
os.environ["REME_LOG_LEVEL"] = reme_log_level.upper()
|
||||||
|
if log_dir:
|
||||||
|
os.environ["REME_LOG_DIR"] = log_dir
|
||||||
|
|
||||||
|
|
||||||
|
def _configure_worker(
|
||||||
|
log_level: str,
|
||||||
|
reme_log_level: str,
|
||||||
|
log_dir: str | None = None,
|
||||||
|
):
|
||||||
|
"""Set up logging inside a multiprocessing worker process.
|
||||||
|
|
||||||
|
Must be called at the top of each worker because child processes inherit
|
||||||
|
parent state but loguru sinks are NOT shared across fork/spawn.
|
||||||
|
"""
|
||||||
|
numeric = getattr(logging, log_level.upper(), logging.INFO)
|
||||||
|
logging.basicConfig(level=numeric, format=_DEFAULT_LOG_FORMAT, force=True)
|
||||||
|
logging.getLogger("longmemeval").setLevel(numeric)
|
||||||
|
if numeric > logging.DEBUG:
|
||||||
|
for name in _NOISY_LOGGERS:
|
||||||
|
logging.getLogger(name).setLevel(max(numeric, logging.WARNING))
|
||||||
|
|
||||||
|
# Add file handler for eval runner in worker process
|
||||||
|
if log_dir:
|
||||||
|
os.makedirs(log_dir, exist_ok=True)
|
||||||
|
pid = os.getpid()
|
||||||
|
log_filepath = os.path.join(log_dir, f"worker-{pid}.log")
|
||||||
|
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
|
||||||
|
file_handler.setLevel(numeric)
|
||||||
|
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
|
||||||
|
logging.getLogger().addHandler(file_handler)
|
||||||
|
|
||||||
|
# Re-initialize loguru for reme internals at the desired level
|
||||||
|
from reme.utils import get_logger
|
||||||
|
|
||||||
|
reme_log_dir = log_dir or "logs"
|
||||||
|
get_logger(log_dir=reme_log_dir, level=reme_log_level.upper(), force_init=True)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Config loading
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def load_eval_config(config_path: str | None = None) -> dict:
|
||||||
|
"""Load evaluation config yaml with env-var expansion."""
|
||||||
|
if config_path is None:
|
||||||
|
config_path = str(Path(__file__).parent / "config.yaml")
|
||||||
|
with open(config_path, encoding="utf-8") as f:
|
||||||
|
raw = f.read()
|
||||||
|
|
||||||
|
# Expand ${VAR} and ${VAR:-default}
|
||||||
|
def _expand(m):
|
||||||
|
expr = m.group(1)
|
||||||
|
if ":-" in expr:
|
||||||
|
key, default = expr.split(":-", 1)
|
||||||
|
return os.environ.get(key, default)
|
||||||
|
return os.environ.get(expr, "")
|
||||||
|
|
||||||
|
raw = re.sub(r"\$\{([^}]+)\}", _expand, raw)
|
||||||
|
return yaml.safe_load(raw)
|
||||||
|
|
||||||
|
|
||||||
|
def create_reme_app(config: str = "benchmark", **overrides):
|
||||||
|
"""Create an app with the LongMemEval candidate and judge plugins enabled.
|
||||||
|
|
||||||
|
Plugin discovery remains environment-based; editable installation keeps local
|
||||||
|
plugin source changes visible to every multiprocessing worker.
|
||||||
|
"""
|
||||||
|
from reme import Application
|
||||||
|
from reme.config import resolve_app_config
|
||||||
|
|
||||||
|
enabled_plugins = list(overrides.pop("plugins", ()) or ())
|
||||||
|
for plugin in ("lme", "lme-judge"):
|
||||||
|
if plugin not in enabled_plugins:
|
||||||
|
enabled_plugins.append(plugin)
|
||||||
|
app_config = resolve_app_config(config=config, plugins=enabled_plugins, **overrides)
|
||||||
|
return Application(**app_config)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Date utilities
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def parse_haystack_date(date_str: str) -> datetime:
|
||||||
|
"""Parse LongMemEval date format: '2023/05/20 (Sat) 02:21' -> datetime."""
|
||||||
|
m = re.match(r"(\d{4}/\d{2}/\d{2})\s+\(\w+\)\s+(\d{2}:\d{2})", date_str)
|
||||||
|
if not m:
|
||||||
|
raise ValueError(f"Cannot parse haystack date: {date_str!r}")
|
||||||
|
return datetime.strptime(f"{m.group(1)} {m.group(2)}", "%Y/%m/%d %H:%M")
|
||||||
|
|
||||||
|
|
||||||
|
def to_iso(dt: datetime) -> str:
|
||||||
|
"""Convert datetime to ISO-8601 string precise to seconds."""
|
||||||
|
return dt.strftime("%Y-%m-%dT%H:%M:%S")
|
||||||
|
|
||||||
|
|
||||||
|
def should_trigger_dream(prev_dt: datetime, curr_dt: datetime, _trigger_hour: int = 23) -> bool:
|
||||||
|
"""Check if the time gap between two sessions crosses trigger_hour (e.g. 23:00)."""
|
||||||
|
if prev_dt.date() == curr_dt.date():
|
||||||
|
return False
|
||||||
|
# There's at least one midnight crossing; check if trigger_hour is between them
|
||||||
|
# Simple heuristic: if dates differ, dream should run for the previous day
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def sessions_sorted_by_time(item: dict) -> list[tuple[int, datetime, str, list[dict]]]:
|
||||||
|
"""Return (original_index, parsed_datetime, session_id, messages) sorted by time."""
|
||||||
|
entries = []
|
||||||
|
for i, (date_str, sid, msgs) in enumerate(
|
||||||
|
zip(item["haystack_dates"], item["haystack_session_ids"], item["haystack_sessions"]),
|
||||||
|
):
|
||||||
|
dt = parse_haystack_date(date_str)
|
||||||
|
entries.append((i, dt, sid, msgs))
|
||||||
|
# Sort by time (ascending)
|
||||||
|
entries.sort(key=lambda x: x[1])
|
||||||
|
return entries
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Message formatting
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def format_messages_for_reme(messages: list[dict], session_dt: datetime) -> list[dict]:
|
||||||
|
"""Convert LongMemEval messages to ReMe auto_memory format.
|
||||||
|
|
||||||
|
Adds: name, created_at (ISO seconds). All messages in a session share the
|
||||||
|
same created_at (the session timestamp).
|
||||||
|
"""
|
||||||
|
formatted = []
|
||||||
|
for msg in messages:
|
||||||
|
role = msg["role"]
|
||||||
|
formatted.append(
|
||||||
|
{
|
||||||
|
"name": role,
|
||||||
|
"role": role,
|
||||||
|
"content": msg["content"],
|
||||||
|
"created_at": to_iso(session_dt),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return formatted
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# LLM-as-Judge (delegated to answer_judge_step via app.run_job)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
async def judge_response_via_job(
|
||||||
|
app,
|
||||||
|
question: str,
|
||||||
|
ground_truth: str,
|
||||||
|
response: str,
|
||||||
|
question_type: str,
|
||||||
|
) -> dict:
|
||||||
|
"""Use the answer_judge_step to evaluate a response against the golden answer."""
|
||||||
|
judge_resp = await app.run_job(
|
||||||
|
"answer_judge",
|
||||||
|
query=question,
|
||||||
|
agent_answer=response,
|
||||||
|
golden_answer=ground_truth,
|
||||||
|
question_type=question_type,
|
||||||
|
)
|
||||||
|
|
||||||
|
verdict = (judge_resp.answer or "").strip().lower()
|
||||||
|
raw_answer = (judge_resp.metadata or {}).get("raw_answer_judgement", "")
|
||||||
|
|
||||||
|
return {
|
||||||
|
"verdict": verdict,
|
||||||
|
"reason": raw_answer if verdict not in ("yes", "no") else "",
|
||||||
|
"metric": "binary",
|
||||||
|
"question_type": question_type,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Main evaluation pipeline
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
async def evaluate_item(item: dict, eval_config: dict, item_index: int, eval_only: bool = False) -> dict:
|
||||||
|
"""Evaluate a single LongMemEval item end-to-end.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
item: The dataset item containing question, answer, sessions, etc.
|
||||||
|
eval_config: The evaluation configuration dict.
|
||||||
|
item_index: The index of this item in the dataset.
|
||||||
|
eval_only: If True, skip ingestion (phases 1-3) and only run query+judge
|
||||||
|
using the existing workspace. Useful for re-evaluating different query
|
||||||
|
configurations without re-ingesting sessions.
|
||||||
|
"""
|
||||||
|
from reme.utils.evaluation_interface import track_agent_token_usage, track_job_counts
|
||||||
|
|
||||||
|
reme_cfg = eval_config["reme"]
|
||||||
|
dream_trigger_hour = reme_cfg.get("dream_trigger_hour", 23)
|
||||||
|
dream_scan_days = reme_cfg.get("dream_scan_days", 2)
|
||||||
|
dream_max_units = reme_cfg.get("dream_max_units", 5)
|
||||||
|
|
||||||
|
# Sort sessions by time
|
||||||
|
sorted_sessions = sessions_sorted_by_time(item)
|
||||||
|
|
||||||
|
# Filter out sessions that occur after question_date (if enabled)
|
||||||
|
filter_future = eval_config["evaluation"].get("filter_future_sessions", True)
|
||||||
|
if filter_future and item.get("question_date"):
|
||||||
|
question_dt = parse_haystack_date(item["question_date"])
|
||||||
|
total_before_filter = len(sorted_sessions)
|
||||||
|
sorted_sessions = [(i, dt, sid, msgs) for i, dt, sid, msgs in sorted_sessions if dt <= question_dt]
|
||||||
|
if len(sorted_sessions) < total_before_filter:
|
||||||
|
logger.info(
|
||||||
|
f"[Item {item_index}] Filtered sessions: {total_before_filter} -> {len(sorted_sessions)} "
|
||||||
|
f"(removed {total_before_filter - len(sorted_sessions)} future sessions "
|
||||||
|
f"after question_date={item['question_date']})",
|
||||||
|
)
|
||||||
|
|
||||||
|
logger.info(
|
||||||
|
"[Item %s] question_id=%s type=%s sessions=%d%s",
|
||||||
|
item_index,
|
||||||
|
item["question_id"],
|
||||||
|
item["question_type"],
|
||||||
|
len(sorted_sessions),
|
||||||
|
" [eval_only]" if eval_only else "",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Use fixed workspace directory (clean it for fresh evaluation)
|
||||||
|
workspace_root = _PROJECT_ROOT / eval_config["dataset"].get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
|
||||||
|
item_dir = workspace_root / f"item_{item_index}"
|
||||||
|
workspace_dir = str(item_dir / ".reme")
|
||||||
|
if eval_only:
|
||||||
|
if not item_dir.exists() or not Path(workspace_dir).exists():
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"[Item {item_index}] eval_only: workspace not found at {item_dir}. "
|
||||||
|
f"Run without --eval_only first to build the workspace.",
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
if item_dir.exists():
|
||||||
|
shutil.rmtree(item_dir)
|
||||||
|
logger.info(f"[Item {item_index}] Cleaned existing workspace: {item_dir}")
|
||||||
|
else:
|
||||||
|
logger.info(f"[Item {item_index}] Workspace not found, creating: {item_dir}")
|
||||||
|
item_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Pre-initialize ReMe's loguru logger with the correct log_dir
|
||||||
|
# (singleton — Application.__init__ will reuse this instance)
|
||||||
|
output_cfg = eval_config.get("output", {})
|
||||||
|
if output_cfg.get("log_to_file", False):
|
||||||
|
reme_log_dir = os.environ.get("REME_LOG_DIR")
|
||||||
|
if reme_log_dir:
|
||||||
|
from reme.utils import get_logger
|
||||||
|
|
||||||
|
get_logger(
|
||||||
|
log_dir=reme_log_dir,
|
||||||
|
level=os.environ.get("REME_LOG_LEVEL", "INFO"),
|
||||||
|
log_to_console=output_cfg.get("log_to_console", True),
|
||||||
|
log_to_file=True,
|
||||||
|
force_init=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
app = create_reme_app(
|
||||||
|
config=reme_cfg["config"],
|
||||||
|
plugins=reme_cfg.get("plugins", ()),
|
||||||
|
workspace_dir=workspace_dir,
|
||||||
|
log_to_console=output_cfg.get("log_to_console", True),
|
||||||
|
log_to_file=output_cfg.get("log_to_file", False),
|
||||||
|
enable_logo=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
await app.start()
|
||||||
|
|
||||||
|
try:
|
||||||
|
dream_dates_triggered = set()
|
||||||
|
dream_available = True # Set to False if auto_dream job is not found
|
||||||
|
|
||||||
|
if not eval_only:
|
||||||
|
# ── Phase 1: Ingest sessions ──────────────────────────────
|
||||||
|
prev_dt = None
|
||||||
|
|
||||||
|
for idx, (_, session_dt, session_id, messages) in enumerate(sorted_sessions):
|
||||||
|
# Check if dream should be triggered before this session
|
||||||
|
if (
|
||||||
|
dream_available
|
||||||
|
and prev_dt is not None
|
||||||
|
and should_trigger_dream(prev_dt, session_dt, dream_trigger_hour)
|
||||||
|
):
|
||||||
|
dream_date = prev_dt.strftime("%Y-%m-%d")
|
||||||
|
if dream_date not in dream_dates_triggered:
|
||||||
|
logger.info(f"[Item {item_index}] Triggering dream for date={dream_date}")
|
||||||
|
try:
|
||||||
|
dream_resp = await app.run_job(
|
||||||
|
"auto_dream",
|
||||||
|
date=dream_date,
|
||||||
|
scan_days=dream_scan_days,
|
||||||
|
max_units=dream_max_units,
|
||||||
|
)
|
||||||
|
logger.info(
|
||||||
|
f"[Item {item_index}] Dream done: success={dream_resp.success} "
|
||||||
|
f"answer={dream_resp.answer[:100] if dream_resp.answer else ''}",
|
||||||
|
)
|
||||||
|
except Exception as e:
|
||||||
|
if "not found" in str(e).lower():
|
||||||
|
dream_available = False
|
||||||
|
logger.warning(f"[Item {item_index}] auto_dream job not found, skipping all dreams")
|
||||||
|
else:
|
||||||
|
logger.warning(f"[Item {item_index}] Dream failed for {dream_date}: {e}")
|
||||||
|
dream_dates_triggered.add(dream_date)
|
||||||
|
# Index update after dream to pick up new digest nodes
|
||||||
|
await app.run_job("index_update")
|
||||||
|
|
||||||
|
# Format and ingest the session
|
||||||
|
formatted_msgs = format_messages_for_reme(messages, session_dt)
|
||||||
|
date_str = session_dt.strftime("%Y-%m-%d")
|
||||||
|
|
||||||
|
logger.info(
|
||||||
|
f"[Item {item_index}] Ingesting session {idx+1}/{len(sorted_sessions)} "
|
||||||
|
f"id={session_id} date={date_str} msgs={len(formatted_msgs)}",
|
||||||
|
)
|
||||||
|
resp = await app.run_job(
|
||||||
|
"auto_memory",
|
||||||
|
messages=formatted_msgs,
|
||||||
|
session_id=session_id,
|
||||||
|
date=date_str,
|
||||||
|
)
|
||||||
|
if not resp.success:
|
||||||
|
logger.warning(
|
||||||
|
f"[Item {item_index}] auto_memory failed for session {session_id}: {resp.answer}",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Manual index update after each session
|
||||||
|
await app.run_job("index_update")
|
||||||
|
|
||||||
|
prev_dt = session_dt
|
||||||
|
|
||||||
|
# ── Phase 2: Final dream for the last day ─────────────────
|
||||||
|
if dream_available and prev_dt is not None:
|
||||||
|
last_dream_date = prev_dt.strftime("%Y-%m-%d")
|
||||||
|
if last_dream_date not in dream_dates_triggered:
|
||||||
|
logger.info(f"[Item {item_index}] Final dream for date={last_dream_date}")
|
||||||
|
try:
|
||||||
|
await app.run_job(
|
||||||
|
"auto_dream",
|
||||||
|
date=last_dream_date,
|
||||||
|
scan_days=dream_scan_days,
|
||||||
|
max_units=dream_max_units,
|
||||||
|
)
|
||||||
|
except Exception as e:
|
||||||
|
if "not found" in str(e).lower():
|
||||||
|
dream_available = False
|
||||||
|
logger.warning(f"[Item {item_index}] auto_dream job not found, skipping all dreams")
|
||||||
|
else:
|
||||||
|
logger.warning(f"[Item {item_index}] Final dream failed: {e}")
|
||||||
|
dream_dates_triggered.add(last_dream_date)
|
||||||
|
# Index update after final dream
|
||||||
|
await app.run_job("index_update")
|
||||||
|
|
||||||
|
# ── Phase 3: Digest update ────────────────────────────────
|
||||||
|
await app.run_job("digest_update")
|
||||||
|
|
||||||
|
# ── Phase 4: Ask question via agentic_answer job (ReAct agent) ──
|
||||||
|
question = item["question"]
|
||||||
|
compress_session = bool(eval_config["evaluation"].get("compress_session", False))
|
||||||
|
question_date_raw = item.get("question_date", "")
|
||||||
|
question_dt = parse_haystack_date(question_date_raw) if question_date_raw else None
|
||||||
|
query_time = to_iso(question_dt) if question_dt else ""
|
||||||
|
logger.info(
|
||||||
|
f"[Item {item_index}] Asking (agentic): {question[:80]}... query_time={query_time}",
|
||||||
|
)
|
||||||
|
|
||||||
|
with (
|
||||||
|
track_job_counts(["search"], app.context) as tool_counts,
|
||||||
|
track_agent_token_usage(
|
||||||
|
["bench"],
|
||||||
|
app.context,
|
||||||
|
) as token_usages,
|
||||||
|
):
|
||||||
|
query_resp = await app.run_job(
|
||||||
|
"agentic_answer",
|
||||||
|
query=question,
|
||||||
|
query_time=query_time,
|
||||||
|
compress_session=compress_session,
|
||||||
|
)
|
||||||
|
agentic_tool_counts = tool_counts
|
||||||
|
agentic_token_usage = token_usages["bench"]
|
||||||
|
agentic_response = (query_resp.answer or "").strip()
|
||||||
|
if not agentic_response:
|
||||||
|
agentic_response = "(no answer generated)"
|
||||||
|
|
||||||
|
logger.info(f"[Item {item_index}] Agentic response: {agentic_response[:200]}...")
|
||||||
|
logger.info(f"[Item {item_index}] Agentic tool calls: {agentic_tool_counts}")
|
||||||
|
logger.info(f"[Item {item_index}] Bench token usage: {agentic_token_usage}")
|
||||||
|
|
||||||
|
# ── Phase 5: Judge agentic response (via answer_judge_step) ──────────
|
||||||
|
logger.info(f"[Item {item_index}] Judging agentic (binary, type={item['question_type']})...")
|
||||||
|
agentic_judgment = await judge_response_via_job(
|
||||||
|
app=app,
|
||||||
|
question=question,
|
||||||
|
ground_truth=item["answer"],
|
||||||
|
response=agentic_response,
|
||||||
|
question_type=item["question_type"],
|
||||||
|
)
|
||||||
|
logger.info(f"[Item {item_index}] agentic binary result: {agentic_judgment}")
|
||||||
|
|
||||||
|
finally:
|
||||||
|
await app.close()
|
||||||
|
|
||||||
|
return {
|
||||||
|
"question_id": item["question_id"],
|
||||||
|
"question_type": item["question_type"],
|
||||||
|
"question": question,
|
||||||
|
"ground_truth": item["answer"],
|
||||||
|
"agentic_response": agentic_response,
|
||||||
|
"agentic_judgment": agentic_judgment,
|
||||||
|
"agentic_tool_counts": agentic_tool_counts,
|
||||||
|
"agentic_token_usage": agentic_token_usage,
|
||||||
|
"sessions_ingested": len(sorted_sessions),
|
||||||
|
"dreams_triggered": len(dream_dates_triggered),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Worker: runs a single item in its own process with its own event loop
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def _evaluate_item_worker(task_input: tuple) -> dict:
|
||||||
|
"""Worker function for multiprocessing. Each process gets its own event loop."""
|
||||||
|
item, eval_config, item_index, log_level, reme_log_level, eval_only, log_dir = task_input
|
||||||
|
import asyncio # pylint: disable=import-outside-toplevel
|
||||||
|
|
||||||
|
_configure_worker(log_level, reme_log_level, log_dir=log_dir)
|
||||||
|
|
||||||
|
# Permanently suppress "Task exception was never retrieved" /
|
||||||
|
# "Event loop is closed" noise from httpx AsyncClient GC cleanup.
|
||||||
|
# These fire AFTER asyncio.run() closes the loop, during Python's
|
||||||
|
# garbage collection of httpx connection-pool tasks — harmless.
|
||||||
|
logging.getLogger("asyncio").setLevel(logging.CRITICAL)
|
||||||
|
|
||||||
|
return asyncio.run(evaluate_item(item, eval_config, item_index, eval_only=eval_only))
|
||||||
|
|
||||||
|
|
||||||
|
def _indexed_worker(indexed_input: tuple) -> tuple:
|
||||||
|
"""Module-level wrapper for imap_unordered with index tracking."""
|
||||||
|
idx, task_input = indexed_input
|
||||||
|
return idx, _evaluate_item_worker(task_input)
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_num_workers(configured: int) -> int:
|
||||||
|
"""Resolve num_workers: 0=auto (cpu_count-2, min 1), 1=sequential, >1=parallel."""
|
||||||
|
if configured == 0:
|
||||||
|
return max(1, (os.cpu_count() or 4) - 2)
|
||||||
|
return max(1, configured)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Entry point
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def main(
|
||||||
|
config_path: str | None = None,
|
||||||
|
log_level: str = "INFO",
|
||||||
|
reme_log_level: str = "INFO",
|
||||||
|
eval_only: bool = False,
|
||||||
|
):
|
||||||
|
"""Run the LongMemEval evaluation pipeline.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
config_path: Path to the YAML config file.
|
||||||
|
log_level: Log level for the eval runner.
|
||||||
|
reme_log_level: Log level for reme internal logs.
|
||||||
|
eval_only: If True, skip ingestion and only run query+judge using
|
||||||
|
existing workspaces.
|
||||||
|
"""
|
||||||
|
from multiprocessing import Pool # pylint: disable=import-outside-toplevel
|
||||||
|
|
||||||
|
# Load config BEFORE logging setup so log_dir is available
|
||||||
|
eval_config = load_eval_config(config_path)
|
||||||
|
|
||||||
|
# Resolve per-run log directory from config
|
||||||
|
output_cfg = eval_config.get("output", {})
|
||||||
|
log_dir_abs = None
|
||||||
|
if output_cfg.get("log_to_file", False):
|
||||||
|
log_dir_raw = output_cfg.get("log_dir", "logs")
|
||||||
|
log_prefix = output_cfg.get("log_prefix", "longmemeval")
|
||||||
|
run_ts = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
|
||||||
|
log_dir_abs = str(_PROJECT_ROOT / log_dir_raw / f"{log_prefix}_{run_ts}")
|
||||||
|
|
||||||
|
setup_logging(log_level, reme_log_level, log_dir=log_dir_abs)
|
||||||
|
dataset_cfg = eval_config["dataset"]
|
||||||
|
|
||||||
|
# Load dataset
|
||||||
|
dataset_path = _PROJECT_ROOT / dataset_cfg["path"]
|
||||||
|
logger.info(f"Loading dataset from {dataset_path}")
|
||||||
|
with open(dataset_path, encoding="utf-8") as f:
|
||||||
|
data = json.load(f)
|
||||||
|
|
||||||
|
start = dataset_cfg.get("start_index", 0)
|
||||||
|
num_items = dataset_cfg.get("num_items", 0)
|
||||||
|
if num_items > 0:
|
||||||
|
raw_items = data[start : start + num_items]
|
||||||
|
else:
|
||||||
|
raw_items = data[start:]
|
||||||
|
|
||||||
|
# Build item list
|
||||||
|
items_with_idx = [(start + i, item) for i, item in enumerate(raw_items)]
|
||||||
|
|
||||||
|
# Filter by question_type if specified
|
||||||
|
question_types = dataset_cfg.get("question_types") or []
|
||||||
|
if question_types:
|
||||||
|
before_filter = len(items_with_idx)
|
||||||
|
items_with_idx = [(idx, item) for idx, item in items_with_idx if item.get("question_type") in question_types]
|
||||||
|
logger.info(
|
||||||
|
f"Filtered by question_types={question_types}: {before_filter} -> {len(items_with_idx)} items",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Filter by question_id if specified
|
||||||
|
question_ids = dataset_cfg.get("question_ids") or []
|
||||||
|
if question_ids:
|
||||||
|
qid_set = set(question_ids)
|
||||||
|
before_filter = len(items_with_idx)
|
||||||
|
items_with_idx = [(idx, item) for idx, item in items_with_idx if item.get("question_id") in qid_set]
|
||||||
|
logger.info(
|
||||||
|
f"Filtered by question_ids ({len(qid_set)} ids): {before_filter} -> {len(items_with_idx)} items",
|
||||||
|
)
|
||||||
|
|
||||||
|
logger.info(
|
||||||
|
"Evaluating %d item(s) starting from index %d%s",
|
||||||
|
len(items_with_idx),
|
||||||
|
start,
|
||||||
|
" [eval_only: query+judge only]" if eval_only else "",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Resolve parallelism
|
||||||
|
num_workers = _resolve_num_workers(eval_config["evaluation"].get("num_workers", 1))
|
||||||
|
logger.info(f"Using {num_workers} worker(s)")
|
||||||
|
|
||||||
|
# Create output directory
|
||||||
|
output_dir = _PROJECT_ROOT / output_cfg.get("dir", "benchmark/longmemeval/results")
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Create workspace root directory
|
||||||
|
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
|
||||||
|
workspace_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# Pre-check: verify all workspaces exist in eval_only mode
|
||||||
|
if eval_only:
|
||||||
|
missing_items = []
|
||||||
|
for orig_idx, _ in items_with_idx:
|
||||||
|
item_dir = workspace_root / f"item_{orig_idx}"
|
||||||
|
if not item_dir.exists() or not (item_dir / ".reme").exists():
|
||||||
|
missing_items.append(orig_idx)
|
||||||
|
if missing_items:
|
||||||
|
preview = missing_items[:10]
|
||||||
|
suffix = "..." if len(missing_items) > 10 else ""
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"eval_only: {len(missing_items)} workspace(s) not found under {workspace_root}. "
|
||||||
|
f"Missing item indices: {preview}{suffix}. "
|
||||||
|
f"Run without --eval_only first to build the workspaces.",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Build task args — include log levels, eval_only flag, and log paths (use original index for workspace lookup)
|
||||||
|
task_args = [
|
||||||
|
(item, eval_config, orig_idx, log_level, reme_log_level, eval_only, log_dir_abs)
|
||||||
|
for orig_idx, item in items_with_idx
|
||||||
|
]
|
||||||
|
|
||||||
|
# Progress tracking (force print regardless of log level, every 10 minutes)
|
||||||
|
total_items = len(task_args)
|
||||||
|
completed_count = [0] # use list for mutability in closure
|
||||||
|
start_time = time.time()
|
||||||
|
progress_lock = threading.Lock()
|
||||||
|
|
||||||
|
def _print_progress(prefix: str = "PROGRESS"):
|
||||||
|
elapsed = time.time() - start_time
|
||||||
|
elapsed_min = elapsed / 60
|
||||||
|
done = completed_count[0]
|
||||||
|
pct = 100.0 * done / total_items if total_items else 0
|
||||||
|
eta_str = "N/A"
|
||||||
|
if done > 0:
|
||||||
|
eta_sec = elapsed / done * (total_items - done)
|
||||||
|
eta_str = f"{eta_sec/60:.1f}min"
|
||||||
|
print(
|
||||||
|
f"[{prefix}] {datetime.now().strftime('%Y-%m-%d %H:%M:%S')} | "
|
||||||
|
f"{done}/{total_items} ({pct:.1f}%) completed | "
|
||||||
|
f"elapsed={elapsed_min:.1f}min | ETA={eta_str}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _progress_timer():
|
||||||
|
"""Background thread: print progress every 10 minutes."""
|
||||||
|
while not _timer_stop.is_set():
|
||||||
|
_timer_stop.wait(600) # 10 minutes
|
||||||
|
if not _timer_stop.is_set():
|
||||||
|
with progress_lock:
|
||||||
|
_print_progress()
|
||||||
|
|
||||||
|
_timer_stop = threading.Event()
|
||||||
|
timer_thread = threading.Thread(target=_progress_timer, daemon=True)
|
||||||
|
timer_thread.start()
|
||||||
|
|
||||||
|
# Run evaluation
|
||||||
|
if num_workers == 1:
|
||||||
|
# Sequential mode
|
||||||
|
results = []
|
||||||
|
for task_input in task_args:
|
||||||
|
result = _evaluate_item_worker(task_input)
|
||||||
|
results.append(result)
|
||||||
|
with progress_lock:
|
||||||
|
completed_count[0] += 1
|
||||||
|
else:
|
||||||
|
# Parallel mode — use imap_unordered for progress tracking
|
||||||
|
results = [None] * total_items
|
||||||
|
indexed_args = list(enumerate(task_args))
|
||||||
|
|
||||||
|
with Pool(processes=num_workers) as pool:
|
||||||
|
for idx, result in pool.imap_unordered(_indexed_worker, indexed_args):
|
||||||
|
results[idx] = result
|
||||||
|
with progress_lock:
|
||||||
|
completed_count[0] += 1
|
||||||
|
|
||||||
|
# Stop progress timer
|
||||||
|
_timer_stop.set()
|
||||||
|
timer_thread.join(timeout=2)
|
||||||
|
|
||||||
|
# Save results
|
||||||
|
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
|
||||||
|
output_file = output_dir / f"results_{timestamp}.json"
|
||||||
|
with open(output_file, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(results, f, ensure_ascii=False, indent=2)
|
||||||
|
logger.info(f"Results saved to {output_file}")
|
||||||
|
|
||||||
|
# Final progress
|
||||||
|
_print_progress("FINAL")
|
||||||
|
|
||||||
|
_print_summary(results, start_time)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Summary printing
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
def _print_summary(results: list[dict], start_time: float) -> None:
|
||||||
|
"""Print per-item verdicts and per-type accuracy."""
|
||||||
|
print("\n" + "=" * 60)
|
||||||
|
print("EVALUATION RESULTS")
|
||||||
|
print("=" * 60)
|
||||||
|
|
||||||
|
def _accumulate(judgment_key):
|
||||||
|
correct = 0
|
||||||
|
stats: dict = {} # {question_type: {correct: int, total: int}}
|
||||||
|
for r in results:
|
||||||
|
qtype = r["question_type"]
|
||||||
|
verdict = r.get(judgment_key, {}).get("verdict", "N/A")
|
||||||
|
if qtype not in stats:
|
||||||
|
stats[qtype] = {"correct": 0, "total": 0}
|
||||||
|
stats[qtype]["total"] += 1
|
||||||
|
if verdict == "yes":
|
||||||
|
correct += 1
|
||||||
|
stats[qtype]["correct"] += 1
|
||||||
|
return correct, stats
|
||||||
|
|
||||||
|
agentic_correct, agentic_type_stats = _accumulate("agentic_judgment")
|
||||||
|
|
||||||
|
total = len(results)
|
||||||
|
|
||||||
|
# Per-item verdict rows
|
||||||
|
for r in results:
|
||||||
|
a_verdict = r.get("agentic_judgment", {}).get("verdict", "N/A")
|
||||||
|
print(f" [{r['question_id']}] type={r['question_type']} agentic={a_verdict}")
|
||||||
|
|
||||||
|
print("\n" + "-" * 60)
|
||||||
|
print(f" Items: {total}")
|
||||||
|
|
||||||
|
# Agentic stats
|
||||||
|
print("\n ── Agentic (ReAct) ──")
|
||||||
|
print(f" Overall accuracy: {agentic_correct}/{total} ({100*agentic_correct/total:.1f}%)")
|
||||||
|
tool_call_totals = [sum(r.get("agentic_tool_counts", {}).values()) for r in results]
|
||||||
|
tool_call_mean, tool_call_std = _mean_and_std(tool_call_totals)
|
||||||
|
print(f" Tool calls/query: mean={tool_call_mean:.2f} std={tool_call_std:.2f}")
|
||||||
|
token_usages = [r.get("agentic_token_usage", {}) for r in results]
|
||||||
|
print(" Bench reported tokens/query:")
|
||||||
|
for metric in _TOKEN_USAGE_METRICS:
|
||||||
|
values = [usage[metric] for usage in token_usages if usage.get(metric) is not None]
|
||||||
|
if values:
|
||||||
|
mean, std = _mean_and_std(values)
|
||||||
|
print(f" {metric}: mean={mean:.2f} std={std:.2f}")
|
||||||
|
else:
|
||||||
|
print(f" {metric}: unavailable")
|
||||||
|
print(" Per-type accuracy:")
|
||||||
|
for qtype, stats in sorted(agentic_type_stats.items()):
|
||||||
|
acc = 100 * stats["correct"] / stats["total"] if stats["total"] else 0
|
||||||
|
print(f" {qtype}: {stats['correct']}/{stats['total']} ({acc:.1f}%)")
|
||||||
|
|
||||||
|
print("=" * 60)
|
||||||
|
total_elapsed = time.time() - start_time
|
||||||
|
print(f"\n Total time: {total_elapsed/60:.1f} min")
|
||||||
|
print("\n" + "=" * 60)
|
||||||
|
print(" [DONE] EVALUATION COMPLETED SUCCESSFULLY")
|
||||||
|
print("=" * 60 + "\n")
|
||||||
|
|
||||||
|
|
||||||
|
_TOKEN_USAGE_METRICS = (
|
||||||
|
"input_tokens",
|
||||||
|
"output_tokens",
|
||||||
|
"total_tokens",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _mean_and_std(values: list[int]) -> tuple[float, float]:
|
||||||
|
"""Return population mean and standard deviation for one per-query metric."""
|
||||||
|
if not values:
|
||||||
|
return 0.0, 0.0
|
||||||
|
mean = sum(values) / len(values)
|
||||||
|
return mean, (sum((value - mean) ** 2 for value in values) / len(values)) ** 0.5
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
import argparse
|
||||||
|
|
||||||
|
parser = argparse.ArgumentParser(description="LongMemEval evaluation runner")
|
||||||
|
parser.add_argument("--config", type=str, default=None, help="Path to config.yaml")
|
||||||
|
parser.add_argument(
|
||||||
|
"--log-level",
|
||||||
|
type=str,
|
||||||
|
default="INFO",
|
||||||
|
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
|
||||||
|
help="Log level for the eval runner (default: INFO)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--reme-log-level",
|
||||||
|
type=str,
|
||||||
|
default="INFO",
|
||||||
|
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
|
||||||
|
help="Log level for reme internal logs — loguru (default: INFO)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-q",
|
||||||
|
"--quiet",
|
||||||
|
action="store_true",
|
||||||
|
help="Shortcut for --log-level WARNING --reme-log-level WARNING",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--eval_only",
|
||||||
|
action="store_true",
|
||||||
|
help="Skip ingestion (phases 1-3). Reuse existing workspaces and only run query+judge.",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
if args.quiet:
|
||||||
|
args.log_level = "WARNING"
|
||||||
|
args.reme_log_level = "WARNING"
|
||||||
|
|
||||||
|
main(args.config, args.log_level, args.reme_log_level, eval_only=args.eval_only)
|
||||||
|
|
@ -1,343 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Drive the LongMemEval memory pipeline across all samples.
|
|
||||||
|
|
||||||
For every workspace under ``datasets/longmemeval/<idx>`` this launches one or more
|
|
||||||
``reme start config=jinli_lme job=<job>`` runs with ``LME_WORKSPACE_DIR`` pointed
|
|
||||||
at that sample. The pipeline jobs, in order, are:
|
|
||||||
|
|
||||||
1. auto_memory — distil every raw session into a daily note (``daily/*.md``)
|
|
||||||
2. update_index — clear the store and rebuild the index over ``daily/*.md``
|
|
||||||
3. agentic_answer — read ``query.json`` and answer it, writing ``mem_answer.json``
|
|
||||||
4. llm_judge — judge ``mem_answer.json`` against ``answer.json``
|
|
||||||
|
|
||||||
Pick one with ``--job``, or ``--job all`` to run the full pipeline *serially per sample*.
|
|
||||||
Runs are capped at ``--concurrency`` (default 1 for ``--job auto_memory``, otherwise
|
|
||||||
3) samples at once and each launch is staggered by ``--stagger`` seconds so they
|
|
||||||
do not all hit the LLM API at once.
|
|
||||||
|
|
||||||
By default every selected job is rerun for every sample — each job's own clear
|
|
||||||
step (configured in jinli_lme.yaml) wipes stale output first, so a run is always
|
|
||||||
a clean rebuild. Pass ``--resume`` to instead skip samples whose output already
|
|
||||||
exists (``daily/`` for auto_memory, ``metadata/embedding_store/`` for
|
|
||||||
update_index, ``mem_answer.json`` for agentic_answer, ``mem_answer.json`` with
|
|
||||||
``llm_judge.judgement`` for llm_judge) and continue an interrupted batch. Each
|
|
||||||
sample's stdout/stderr goes to ``logs/agentic_answer/<job>/<idx>.log``.
|
|
||||||
|
|
||||||
After an agentic_answer run finishes, the driver aggregates every sample's query,
|
|
||||||
golden answer, predicted answer, LLM judgement and a best-effort tool-call trail
|
|
||||||
into one big JSON at ``logs/agentic_answer/aggregate.json``.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py # agentic_answer, all 500, conc 3
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --job all # full pipeline serially per sample
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --job auto_memory # just step 1
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --job llm_judge # just judge existing answers
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --limit 5 --dry-run # list what would run
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --start 187 # samples 187..499
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --start 187 --end 499 # samples 187..499
|
|
||||||
python benchmark/longmemeval/run_agentic_answer.py --job all --resume # continue an interrupted batch
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import asyncio
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import time
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGDIR = REPO / "logs" / "agentic_answer"
|
|
||||||
AGGREGATE = LOGDIR / "aggregate.json"
|
|
||||||
|
|
||||||
# Pipeline jobs in execution order.
|
|
||||||
JOB_ORDER = ["auto_memory", "update_index", "agentic_answer", "llm_judge"]
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument(
|
|
||||||
"--job",
|
|
||||||
choices=[*JOB_ORDER, "all"],
|
|
||||||
default="agentic_answer",
|
|
||||||
help="which job to run per sample; 'all' runs the full pipeline serially (default: agentic_answer)",
|
|
||||||
)
|
|
||||||
p.add_argument("--concurrency", type=int, default=1, help="max samples running at once (default 3)")
|
|
||||||
p.add_argument("--stagger", type=float, default=1.0, help="seconds between consecutive launches (default 1)")
|
|
||||||
p.add_argument("--start", type=int, default=0, help="first numeric sample id to process, inclusive (default 0)")
|
|
||||||
p.add_argument(
|
|
||||||
"--end",
|
|
||||||
type=int,
|
|
||||||
default=0,
|
|
||||||
help="last numeric sample id to process, inclusive (0 = no upper bound)",
|
|
||||||
)
|
|
||||||
p.add_argument("--limit", type=int, default=0, help="only process the first N samples (0 = all)")
|
|
||||||
p.add_argument(
|
|
||||||
"--resume",
|
|
||||||
action="store_true",
|
|
||||||
help="skip a sample when the job's output already exists (resume an interrupted run); "
|
|
||||||
"by default every selected job is rerun so the config's clear step rebuilds cleanly",
|
|
||||||
)
|
|
||||||
p.add_argument("--dry-run", action="store_true", help="list what would run, launch nothing")
|
|
||||||
p.add_argument("--no-aggregate", action="store_true", help="skip writing aggregate.json after answer/judge jobs")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def selected_jobs(job: str) -> list[str]:
|
|
||||||
"""Expand the --job choice into an ordered list of jobs."""
|
|
||||||
return list(JOB_ORDER) if job == "all" else [job]
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs (numeric workspace dirs), numerically sorted."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def job_done(idx: str, job: str) -> bool:
|
|
||||||
"""Return True when ``job``'s expected output already exists for sample ``idx``."""
|
|
||||||
ws = DATA / idx
|
|
||||||
if job == "auto_memory":
|
|
||||||
daily = ws / "daily"
|
|
||||||
return daily.is_dir() and any(daily.rglob("*.md"))
|
|
||||||
if job == "update_index":
|
|
||||||
store = ws / "metadata" / "embedding_store"
|
|
||||||
return store.is_dir() and any(store.iterdir())
|
|
||||||
if job == "agentic_answer":
|
|
||||||
return (ws / "mem_answer.json").exists()
|
|
||||||
if job == "llm_judge":
|
|
||||||
judge = _load_json(ws / "mem_answer.json").get("llm_judge")
|
|
||||||
return isinstance(judge, dict) and bool(str(judge.get("judgement") or "").strip())
|
|
||||||
raise ValueError(f"unknown job: {job}")
|
|
||||||
|
|
||||||
|
|
||||||
async def run_job(idx: str, job: str, counters: dict) -> bool:
|
|
||||||
"""Run a single job for a single sample. Returns True on success."""
|
|
||||||
log = LOGDIR / job / f"{idx}.log"
|
|
||||||
log.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
env = dict(os.environ, LME_WORKSPACE_DIR=f"datasets/longmemeval/{idx}")
|
|
||||||
started = time.strftime("%H:%M:%S")
|
|
||||||
print(f"[start {started}] {idx}/{job}", flush=True)
|
|
||||||
with log.open("w", encoding="utf-8") as f:
|
|
||||||
proc = await asyncio.create_subprocess_exec(
|
|
||||||
"reme",
|
|
||||||
"start",
|
|
||||||
"config=jinli_lme",
|
|
||||||
f"job={job}",
|
|
||||||
cwd=str(REPO),
|
|
||||||
env=env,
|
|
||||||
stdout=f,
|
|
||||||
stderr=asyncio.subprocess.STDOUT,
|
|
||||||
)
|
|
||||||
rc = await proc.wait()
|
|
||||||
ok = rc == 0 and job_done(idx, job)
|
|
||||||
counters["done" if ok else "fail"] += 1
|
|
||||||
tag = "done" if ok else "fail"
|
|
||||||
print(f"[{tag}] {idx}/{job} rc={rc} ({counters['done']} done / {counters['fail']} fail)", flush=True)
|
|
||||||
return ok
|
|
||||||
|
|
||||||
|
|
||||||
async def run_one(idx: str, jobs: list[str], sem: asyncio.Semaphore, resume: bool, counters: dict) -> None:
|
|
||||||
"""Run the selected jobs for one sample, serially.
|
|
||||||
|
|
||||||
By default every selected job is rerun (the job's own clear step wipes stale
|
|
||||||
output first). With ``resume`` a job is skipped when its output already
|
|
||||||
exists, so an interrupted batch can continue without redoing finished work.
|
|
||||||
"""
|
|
||||||
async with sem:
|
|
||||||
for job in jobs:
|
|
||||||
if resume and job_done(idx, job):
|
|
||||||
counters["skip"] += 1
|
|
||||||
print(f"[skip] {idx}/{job} (output exists)", flush=True)
|
|
||||||
continue
|
|
||||||
ok = await run_job(idx, job, counters)
|
|
||||||
if not ok:
|
|
||||||
# Later jobs depend on earlier ones; don't waste a run on a broken workspace.
|
|
||||||
print(f"[abort] {idx}: {job} failed, skipping remaining jobs", flush=True)
|
|
||||||
break
|
|
||||||
|
|
||||||
|
|
||||||
# --------------------------------------------------------------------------- #
|
|
||||||
# Aggregation of agentic_answer results into one big JSON.
|
|
||||||
# --------------------------------------------------------------------------- #
|
|
||||||
|
|
||||||
# Match ``session_id=abc123`` headers and ``"...session_id": "abc123"`` fields in
|
|
||||||
# tool-result text, so we can list which sessions each search actually surfaced.
|
|
||||||
_SID_RE = re.compile(r'session_id["\s:=]+"?([A-Za-z0-9_\-]+)')
|
|
||||||
|
|
||||||
|
|
||||||
def _load_json(path: Path) -> dict:
|
|
||||||
"""Load a JSON object, returning {} on any error."""
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
return data if isinstance(data, dict) else {}
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return {}
|
|
||||||
|
|
||||||
|
|
||||||
def parse_tool_calls(idx: str, session_id: str) -> list[dict]:
|
|
||||||
"""Best-effort: parse the agent trajectory into an ordered tool-call summary.
|
|
||||||
|
|
||||||
Reads ``mem_session/agentscope/<session_id>.jsonl`` — the trajectory the
|
|
||||||
agentic_answer run dumped — and pairs every ``tool_call`` (name + parsed
|
|
||||||
args) with the ``session_id`` hits found in its ``tool_result``. Returns an
|
|
||||||
empty list if the file is missing or unreadable (never raises).
|
|
||||||
"""
|
|
||||||
if not session_id:
|
|
||||||
return []
|
|
||||||
path = DATA / idx / "mem_session" / "agentscope" / f"{session_id}.jsonl"
|
|
||||||
if not path.exists():
|
|
||||||
return []
|
|
||||||
|
|
||||||
calls: dict[str, dict] = {}
|
|
||||||
order: list[str] = []
|
|
||||||
try:
|
|
||||||
for line in path.read_text(encoding="utf-8").splitlines():
|
|
||||||
line = line.strip()
|
|
||||||
if not line:
|
|
||||||
continue
|
|
||||||
try:
|
|
||||||
msg = json.loads(line)
|
|
||||||
except json.JSONDecodeError:
|
|
||||||
continue
|
|
||||||
for c in msg.get("content") or []:
|
|
||||||
if not isinstance(c, dict):
|
|
||||||
continue
|
|
||||||
cid = c.get("id")
|
|
||||||
if c.get("type") == "tool_call" and cid:
|
|
||||||
try:
|
|
||||||
args = json.loads(c.get("input") or "{}")
|
|
||||||
except (json.JSONDecodeError, TypeError):
|
|
||||||
args = c.get("input")
|
|
||||||
calls[cid] = {"name": c.get("name"), "args": args, "hit_session_ids": []}
|
|
||||||
order.append(cid)
|
|
||||||
elif c.get("type") == "tool_result" and cid in calls:
|
|
||||||
text = ""
|
|
||||||
for o in c.get("output") or []:
|
|
||||||
if isinstance(o, dict) and isinstance(o.get("text"), str):
|
|
||||||
text += o["text"]
|
|
||||||
hits = list(dict.fromkeys(_SID_RE.findall(text)))
|
|
||||||
calls[cid]["hit_session_ids"] = hits
|
|
||||||
except OSError:
|
|
||||||
return []
|
|
||||||
|
|
||||||
return [{"iter": i + 1, **calls[cid]} for i, cid in enumerate(order)]
|
|
||||||
|
|
||||||
|
|
||||||
def build_record(idx: str) -> dict:
|
|
||||||
"""Assemble one sample's aggregate record from its on-disk artifacts."""
|
|
||||||
ws = DATA / idx
|
|
||||||
query = _load_json(ws / "query.json")
|
|
||||||
golden = _load_json(ws / "answer.json")
|
|
||||||
mem = _load_json(ws / "mem_answer.json")
|
|
||||||
|
|
||||||
pred = str(mem.get("answer") or "").strip()
|
|
||||||
session_id = str(mem.get("session_id") or "")
|
|
||||||
llm_judge = mem.get("llm_judge") if isinstance(mem.get("llm_judge"), dict) else {}
|
|
||||||
tool_calls = parse_tool_calls(idx, session_id) if mem else []
|
|
||||||
|
|
||||||
if not mem:
|
|
||||||
status = "missing"
|
|
||||||
elif not pred:
|
|
||||||
status = "empty"
|
|
||||||
elif "not provided" in pred.lower():
|
|
||||||
status = "not_provided"
|
|
||||||
else:
|
|
||||||
status = "answered"
|
|
||||||
|
|
||||||
return {
|
|
||||||
"idx": idx,
|
|
||||||
"question_id": query.get("question_id"),
|
|
||||||
"question_type": query.get("question_type"),
|
|
||||||
"question": query.get("question"),
|
|
||||||
"question_date": query.get("question_date"),
|
|
||||||
"golden_answer": golden.get("answer"),
|
|
||||||
"golden_answer_session_ids": golden.get("answer_session_ids"),
|
|
||||||
"pred_answer": pred,
|
|
||||||
"session_id": session_id,
|
|
||||||
"status": status,
|
|
||||||
"llm_judge": llm_judge.get("judgement"),
|
|
||||||
"llm_judge_raw": llm_judge.get("raw_judgement"),
|
|
||||||
"num_tool_calls": len(tool_calls),
|
|
||||||
"tool_calls": tool_calls,
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def write_aggregate(ids: list[str]) -> None:
|
|
||||||
"""Aggregate every sample's agentic_answer artifacts into one big JSON."""
|
|
||||||
records = [build_record(idx) for idx in ids]
|
|
||||||
finished = [r for r in records if r["status"] != "missing"]
|
|
||||||
by_status: dict[str, int] = {}
|
|
||||||
by_llm_judge: dict[str, int] = {}
|
|
||||||
for r in records:
|
|
||||||
by_status[r["status"]] = by_status.get(r["status"], 0) + 1
|
|
||||||
judgement = r.get("llm_judge") or "missing"
|
|
||||||
by_llm_judge[judgement] = by_llm_judge.get(judgement, 0) + 1
|
|
||||||
|
|
||||||
payload = {
|
|
||||||
"generated_at": time.strftime("%Y-%m-%d %H:%M:%S"),
|
|
||||||
"total": len(records),
|
|
||||||
"finished": len(finished),
|
|
||||||
"by_status": by_status,
|
|
||||||
"by_llm_judge": by_llm_judge,
|
|
||||||
"samples": records,
|
|
||||||
}
|
|
||||||
AGGREGATE.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
AGGREGATE.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
|
||||||
print(f"[aggregate] wrote {len(records)} samples ({len(finished)} finished) -> {AGGREGATE}", flush=True)
|
|
||||||
|
|
||||||
|
|
||||||
async def main() -> int:
|
|
||||||
"""Run the driver."""
|
|
||||||
args = parse_args()
|
|
||||||
LOGDIR.mkdir(parents=True, exist_ok=True)
|
|
||||||
jobs = selected_jobs(args.job)
|
|
||||||
|
|
||||||
ids = sample_ids()
|
|
||||||
if args.end and args.end < args.start:
|
|
||||||
raise ValueError(f"--end ({args.end}) must be >= --start ({args.start})")
|
|
||||||
ids = [i for i in ids if int(i) >= args.start and (not args.end or int(i) <= args.end)]
|
|
||||||
if args.limit:
|
|
||||||
ids = ids[: args.limit]
|
|
||||||
|
|
||||||
# Without --resume every job reruns; with --resume, jobs whose output exists are skipped.
|
|
||||||
def todo_jobs(i: str) -> list[str]:
|
|
||||||
return [j for j in jobs if not (args.resume and job_done(i, j))]
|
|
||||||
|
|
||||||
pending = [i for i in ids if todo_jobs(i)]
|
|
||||||
print(
|
|
||||||
f"jobs={jobs} resume={args.resume} samples total={len(ids)} pending={len(pending)} "
|
|
||||||
f"concurrency={args.concurrency} stagger={args.stagger}s",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if args.dry_run:
|
|
||||||
for i in pending:
|
|
||||||
print(f"[would-run] {i}: {todo_jobs(i)}")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
sem = asyncio.Semaphore(args.concurrency)
|
|
||||||
counters = {"done": 0, "fail": 0, "skip": 0}
|
|
||||||
tasks: list[asyncio.Task] = []
|
|
||||||
for n, idx in enumerate(ids):
|
|
||||||
if n and args.stagger > 0:
|
|
||||||
await asyncio.sleep(args.stagger) # stagger each launch relative to the previous
|
|
||||||
tasks.append(asyncio.create_task(run_one(idx, jobs, sem, args.resume, counters)))
|
|
||||||
|
|
||||||
await asyncio.gather(*tasks, return_exceptions=True)
|
|
||||||
print(
|
|
||||||
f"ALL FINISHED done={counters['done']} fail={counters['fail']} skip={counters['skip']}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if any(j in jobs for j in ("agentic_answer", "llm_judge")) and not args.no_aggregate:
|
|
||||||
write_aggregate(ids)
|
|
||||||
|
|
||||||
return 0 if counters["fail"] == 0 else 1
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(asyncio.run(main()))
|
|
||||||
|
|
@ -1,469 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Review every LongMemEval golden answer with the configured Claude Code job.
|
|
||||||
|
|
||||||
Every numeric ``datasets/longmemeval/<idx>`` workspace is processed sequentially.
|
|
||||||
The reference JSONL files are merged by ``question_id`` and supplied only when
|
|
||||||
they contain an alternative answer for that sample:
|
|
||||||
|
|
||||||
reme start config=jinli_lme job=final_answer_review
|
|
||||||
|
|
||||||
The job returns a plain four-field JSON object with ``reason``,
|
|
||||||
``golden_answer_correct``, ``answer``, and ``is_session_time_wrong``. After
|
|
||||||
every new success, this driver atomically rewrites the complete accumulated
|
|
||||||
output JSONL so an interrupted run can safely resume.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --exclude-reference-question-ids
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --only-reference-question-ids --rerun-selected
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --concurrency 2 --submit-interval-seconds 6
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --question-id e47becba
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --reference path/to/results.jsonl
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --limit 3
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --no-resume
|
|
||||||
python benchmark/longmemeval/run_final_answer_review.py --dry-run
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import concurrent.futures
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import subprocess
|
|
||||||
import sys
|
|
||||||
import tempfile
|
|
||||||
import time
|
|
||||||
from pathlib import Path
|
|
||||||
from typing import Any
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
DEFAULT_REFERENCES = (
|
|
||||||
REPO / "benchmark" / "longmemeval" / "golden_check_list_false.jsonl",
|
|
||||||
REPO / "benchmark" / "longmemeval" / "merge_confirm_jinli_false.jsonl",
|
|
||||||
)
|
|
||||||
DEFAULT_OUTPUT = REPO / "benchmark" / "longmemeval" / "final_answer_review.jsonl"
|
|
||||||
DEFAULT_LOG_DIR = REPO / "logs" / "final_answer_review"
|
|
||||||
REFERENCE_PATHS_ENV = "LME_FINAL_ANSWER_REFERENCE_PATHS"
|
|
||||||
MAX_CONCURRENCY = 3
|
|
||||||
MIN_SUBMIT_INTERVAL_SECONDS = 5.0
|
|
||||||
DEFAULT_SUBMIT_INTERVAL_SECONDS = 5.1
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
parser = argparse.ArgumentParser(
|
|
||||||
description=__doc__,
|
|
||||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--question-id",
|
|
||||||
dest="question_ids",
|
|
||||||
action="append",
|
|
||||||
help="process only this dataset question ID; repeat for multiple IDs (default: all)",
|
|
||||||
)
|
|
||||||
reference_selection = parser.add_mutually_exclusive_group()
|
|
||||||
reference_selection.add_argument(
|
|
||||||
"--exclude-reference-question-ids",
|
|
||||||
action="store_true",
|
|
||||||
help="skip question IDs found in the selected reference-answer JSONL files",
|
|
||||||
)
|
|
||||||
reference_selection.add_argument(
|
|
||||||
"--only-reference-question-ids",
|
|
||||||
action="store_true",
|
|
||||||
help="process only question IDs found in the selected reference-answer JSONL files",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--reference",
|
|
||||||
dest="references",
|
|
||||||
action="append",
|
|
||||||
type=Path,
|
|
||||||
help="reference-answer JSONL; repeat for multiple files (default: built-in disputed results)",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--output",
|
|
||||||
type=Path,
|
|
||||||
default=DEFAULT_OUTPUT,
|
|
||||||
help=f"output JSONL (default: {DEFAULT_OUTPUT})",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--log-dir",
|
|
||||||
type=Path,
|
|
||||||
default=DEFAULT_LOG_DIR,
|
|
||||||
help="directory for per-question logs",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--concurrency",
|
|
||||||
type=int,
|
|
||||||
default=MAX_CONCURRENCY,
|
|
||||||
help=f"maximum concurrent jobs, from 1 to {MAX_CONCURRENCY} (default: {MAX_CONCURRENCY})",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--submit-interval-seconds",
|
|
||||||
type=float,
|
|
||||||
default=DEFAULT_SUBMIT_INTERVAL_SECONDS,
|
|
||||||
help=f"minimum time between job submissions; must be > {MIN_SUBMIT_INTERVAL_SECONDS:g} "
|
|
||||||
f"(default: {DEFAULT_SUBMIT_INTERVAL_SECONDS:g})",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--limit",
|
|
||||||
type=int,
|
|
||||||
default=0,
|
|
||||||
help="process only the first N pending questions (0 = all)",
|
|
||||||
)
|
|
||||||
resume_mode = parser.add_mutually_exclusive_group()
|
|
||||||
resume_mode.add_argument(
|
|
||||||
"--no-resume",
|
|
||||||
action="store_true",
|
|
||||||
help="ignore existing output and rerun every selected question",
|
|
||||||
)
|
|
||||||
resume_mode.add_argument(
|
|
||||||
"--rerun-selected",
|
|
||||||
action="store_true",
|
|
||||||
help="rerun every selected question while preserving existing results until replacements finish",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--dry-run",
|
|
||||||
action="store_true",
|
|
||||||
help="show the selected cases without invoking ReMe",
|
|
||||||
)
|
|
||||||
return parser.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def _read_jsonl(path: Path) -> list[dict[str, Any]]:
|
|
||||||
"""Read a JSONL file and reject malformed or non-object rows."""
|
|
||||||
rows: list[dict[str, Any]] = []
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as file:
|
|
||||||
for line_number, line in enumerate(file, start=1):
|
|
||||||
if not line.strip():
|
|
||||||
continue
|
|
||||||
try:
|
|
||||||
row = json.loads(line)
|
|
||||||
except json.JSONDecodeError as exc:
|
|
||||||
raise ValueError(f"Invalid JSON at {path}:{line_number}") from exc
|
|
||||||
if not isinstance(row, dict):
|
|
||||||
raise ValueError(f"Expected a JSON object at {path}:{line_number}")
|
|
||||||
rows.append(row)
|
|
||||||
except OSError as exc:
|
|
||||||
raise FileNotFoundError(f"Cannot read JSONL file: {path}") from exc
|
|
||||||
return rows
|
|
||||||
|
|
||||||
|
|
||||||
def merge_references(paths: list[Path]) -> dict[str, list[dict[str, Any]]]:
|
|
||||||
"""Merge reference rows by question ID, preserving file and row order."""
|
|
||||||
merged: dict[str, list[dict[str, Any]]] = {}
|
|
||||||
seen_sources: set[tuple[str, str]] = set()
|
|
||||||
for path in paths:
|
|
||||||
for row in _read_jsonl(path):
|
|
||||||
question_id = str(row.get("question_id") or "").strip()
|
|
||||||
if not question_id:
|
|
||||||
raise ValueError(f"Reference row in {path} has no question_id")
|
|
||||||
source_key = (question_id, str(path.resolve()))
|
|
||||||
if source_key in seen_sources:
|
|
||||||
raise ValueError(f"Duplicate question_id={question_id!r} within {path}")
|
|
||||||
seen_sources.add(source_key)
|
|
||||||
merged.setdefault(question_id, []).append({"source": path.name, **row})
|
|
||||||
if not merged:
|
|
||||||
raise ValueError("No reference answers found")
|
|
||||||
return merged
|
|
||||||
|
|
||||||
|
|
||||||
def workspace_map() -> dict[str, Path]:
|
|
||||||
"""Map every dataset question ID to its numeric sample workspace."""
|
|
||||||
mapping: dict[str, Path] = {}
|
|
||||||
for workspace in sorted(
|
|
||||||
(path for path in DATA.iterdir() if path.is_dir() and path.name.isdigit()),
|
|
||||||
key=lambda p: int(p.name),
|
|
||||||
):
|
|
||||||
query_path = workspace / "query.json"
|
|
||||||
if not query_path.is_file():
|
|
||||||
continue
|
|
||||||
try:
|
|
||||||
with query_path.open(encoding="utf-8") as file:
|
|
||||||
query = json.load(file)
|
|
||||||
except (OSError, json.JSONDecodeError) as exc:
|
|
||||||
raise ValueError(f"Cannot parse {query_path}") from exc
|
|
||||||
if not isinstance(query, dict):
|
|
||||||
raise ValueError(f"Expected a JSON object in {query_path}")
|
|
||||||
question_id = str(query.get("question_id") or "").strip()
|
|
||||||
if not question_id:
|
|
||||||
raise ValueError(f"Missing question_id in {query_path}")
|
|
||||||
if question_id in mapping:
|
|
||||||
raise ValueError(
|
|
||||||
f"Duplicate dataset question_id={question_id!r}: {mapping[question_id]} and {workspace}",
|
|
||||||
)
|
|
||||||
mapping[question_id] = workspace
|
|
||||||
return mapping
|
|
||||||
|
|
||||||
|
|
||||||
def select_question_ids(
|
|
||||||
mapping: dict[str, Path],
|
|
||||||
requested: list[str] | None,
|
|
||||||
excluded: set[str] | None = None,
|
|
||||||
) -> list[str]:
|
|
||||||
"""Return all dataset IDs or validate an explicitly requested subset."""
|
|
||||||
excluded = excluded or set()
|
|
||||||
if not requested:
|
|
||||||
return [question_id for question_id in mapping if question_id not in excluded]
|
|
||||||
selected: list[str] = []
|
|
||||||
seen: set[str] = set()
|
|
||||||
for raw_question_id in requested:
|
|
||||||
question_id = raw_question_id.strip()
|
|
||||||
if not question_id:
|
|
||||||
raise ValueError("--question-id must not be empty")
|
|
||||||
if question_id in seen:
|
|
||||||
raise ValueError(f"Duplicate --question-id: {question_id}")
|
|
||||||
if question_id not in mapping:
|
|
||||||
raise ValueError(f"No dataset workspace for question ID: {question_id}")
|
|
||||||
if question_id not in excluded:
|
|
||||||
selected.append(question_id)
|
|
||||||
seen.add(question_id)
|
|
||||||
return selected
|
|
||||||
|
|
||||||
|
|
||||||
def _validate_result(value: Any, *, source: str) -> dict[str, Any]:
|
|
||||||
"""Validate the final four-field answer contract."""
|
|
||||||
expected_keys = {"reason", "golden_answer_correct", "answer", "is_session_time_wrong"}
|
|
||||||
if not isinstance(value, dict) or set(value) != expected_keys:
|
|
||||||
raise ValueError(
|
|
||||||
f"{source} must contain exactly 'reason', 'golden_answer_correct', 'answer', "
|
|
||||||
"and 'is_session_time_wrong'",
|
|
||||||
)
|
|
||||||
if not isinstance(value["reason"], str) or not value["reason"].strip():
|
|
||||||
raise ValueError(f"{source} has an invalid reason")
|
|
||||||
if not isinstance(value["golden_answer_correct"], bool):
|
|
||||||
raise ValueError(f"{source} has an invalid golden_answer_correct")
|
|
||||||
if not isinstance(value["answer"], str):
|
|
||||||
raise ValueError(f"{source} has an invalid answer")
|
|
||||||
answer = value["answer"].strip()
|
|
||||||
if value["golden_answer_correct"] and answer:
|
|
||||||
raise ValueError(f"{source} answer must be empty when golden_answer_correct is true")
|
|
||||||
if not value["golden_answer_correct"] and not answer:
|
|
||||||
raise ValueError(f"{source} answer must be non-empty when golden_answer_correct is false")
|
|
||||||
if not isinstance(value["is_session_time_wrong"], bool):
|
|
||||||
raise ValueError(f"{source} has an invalid is_session_time_wrong")
|
|
||||||
return {
|
|
||||||
"reason": value["reason"].strip(),
|
|
||||||
"golden_answer_correct": value["golden_answer_correct"],
|
|
||||||
"answer": answer,
|
|
||||||
"is_session_time_wrong": False,
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def load_existing(path: Path) -> dict[str, dict[str, Any]]:
|
|
||||||
"""Load resumable output, rejecting duplicate or malformed rows."""
|
|
||||||
if not path.exists():
|
|
||||||
return {}
|
|
||||||
results: dict[str, dict[str, Any]] = {}
|
|
||||||
for row in _read_jsonl(path):
|
|
||||||
question_id = str(row.get("question_id") or "").strip()
|
|
||||||
if not question_id:
|
|
||||||
raise ValueError(f"Existing output row in {path} has no question_id")
|
|
||||||
if question_id in results:
|
|
||||||
raise ValueError(
|
|
||||||
f"Duplicate question_id={question_id!r} in existing output {path}",
|
|
||||||
)
|
|
||||||
results[question_id] = _validate_result(
|
|
||||||
{key: value for key, value in row.items() if key != "question_id"},
|
|
||||||
source=f"existing result for {question_id}",
|
|
||||||
)
|
|
||||||
return results
|
|
||||||
|
|
||||||
|
|
||||||
def atomic_write_results(
|
|
||||||
path: Path,
|
|
||||||
order: list[str],
|
|
||||||
results: dict[str, dict[str, Any]],
|
|
||||||
) -> None:
|
|
||||||
"""Atomically rewrite all accumulated rows in stable merged-input order."""
|
|
||||||
path.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
temp_path: Path | None = None
|
|
||||||
try:
|
|
||||||
with tempfile.NamedTemporaryFile(
|
|
||||||
"w",
|
|
||||||
encoding="utf-8",
|
|
||||||
dir=path.parent,
|
|
||||||
prefix=f".{path.name}.",
|
|
||||||
delete=False,
|
|
||||||
) as file:
|
|
||||||
temp_path = Path(file.name)
|
|
||||||
for question_id in order:
|
|
||||||
if question_id not in results:
|
|
||||||
continue
|
|
||||||
row = {"question_id": question_id, **results[question_id]}
|
|
||||||
file.write(
|
|
||||||
json.dumps(row, ensure_ascii=False, separators=(",", ":")) + "\n",
|
|
||||||
)
|
|
||||||
file.flush()
|
|
||||||
os.fsync(file.fileno())
|
|
||||||
os.replace(temp_path, path)
|
|
||||||
finally:
|
|
||||||
if temp_path is not None and temp_path.exists():
|
|
||||||
temp_path.unlink()
|
|
||||||
|
|
||||||
|
|
||||||
def run_one(
|
|
||||||
question_id: str,
|
|
||||||
workspace: Path,
|
|
||||||
log_dir: Path,
|
|
||||||
reference_paths: list[Path],
|
|
||||||
) -> dict[str, Any]:
|
|
||||||
"""Run the configured one-shot job and validate its stdout JSON."""
|
|
||||||
env = dict(os.environ, LME_WORKSPACE_DIR=str(workspace.relative_to(REPO)))
|
|
||||||
env[REFERENCE_PATHS_ENV] = json.dumps(
|
|
||||||
[str(path.resolve()) for path in reference_paths],
|
|
||||||
ensure_ascii=False,
|
|
||||||
)
|
|
||||||
completed = subprocess.run(
|
|
||||||
[
|
|
||||||
sys.executable,
|
|
||||||
"-c",
|
|
||||||
"from reme.reme import main; main()",
|
|
||||||
"start",
|
|
||||||
"config=jinli_lme",
|
|
||||||
"job=final_answer_review",
|
|
||||||
],
|
|
||||||
cwd=REPO,
|
|
||||||
env=env,
|
|
||||||
text=True,
|
|
||||||
stdout=subprocess.PIPE,
|
|
||||||
stderr=subprocess.PIPE,
|
|
||||||
check=False,
|
|
||||||
)
|
|
||||||
log_dir.mkdir(parents=True, exist_ok=True)
|
|
||||||
log_path = log_dir / f"{question_id}.log"
|
|
||||||
log_text = (
|
|
||||||
f"workspace={workspace}\nreturncode={completed.returncode}\n\n"
|
|
||||||
f"[stdout]\n{completed.stdout}\n[stderr]\n{completed.stderr}"
|
|
||||||
)
|
|
||||||
log_path.write_text(
|
|
||||||
log_text,
|
|
||||||
encoding="utf-8",
|
|
||||||
)
|
|
||||||
if completed.returncode != 0:
|
|
||||||
raise RuntimeError(
|
|
||||||
f"Job failed for {question_id} with rc={completed.returncode}; see {log_path}",
|
|
||||||
)
|
|
||||||
try:
|
|
||||||
value = json.loads(completed.stdout.strip())
|
|
||||||
except json.JSONDecodeError as exc:
|
|
||||||
raise ValueError(
|
|
||||||
f"Job stdout is not JSON for {question_id}; see {log_path}",
|
|
||||||
) from exc
|
|
||||||
return _validate_result(value, source=f"job result for {question_id}")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
"""Review and checkpoint the selected dataset cases sequentially."""
|
|
||||||
args = parse_args()
|
|
||||||
if args.limit < 0:
|
|
||||||
raise ValueError("--limit must be >= 0")
|
|
||||||
if not 1 <= args.concurrency <= MAX_CONCURRENCY:
|
|
||||||
raise ValueError(f"--concurrency must be between 1 and {MAX_CONCURRENCY}")
|
|
||||||
if args.submit_interval_seconds <= MIN_SUBMIT_INTERVAL_SECONDS:
|
|
||||||
raise ValueError(
|
|
||||||
f"--submit-interval-seconds must be > {MIN_SUBMIT_INTERVAL_SECONDS:g}",
|
|
||||||
)
|
|
||||||
|
|
||||||
reference_paths = [path.resolve() for path in (args.references or DEFAULT_REFERENCES)]
|
|
||||||
mapping = workspace_map()
|
|
||||||
references = merge_references(reference_paths)
|
|
||||||
missing = [question_id for question_id in references if question_id not in mapping]
|
|
||||||
if missing:
|
|
||||||
raise ValueError(f"No dataset workspace for question IDs: {', '.join(missing)}")
|
|
||||||
|
|
||||||
full_order = list(mapping)
|
|
||||||
excluded = set(references) if args.exclude_reference_question_ids else set()
|
|
||||||
order = select_question_ids(mapping, args.question_ids, excluded)
|
|
||||||
if args.only_reference_question_ids:
|
|
||||||
order = [question_id for question_id in order if question_id in references]
|
|
||||||
results = {} if args.no_resume else load_existing(args.output.resolve())
|
|
||||||
pending = (
|
|
||||||
list(order) if args.rerun_selected else [question_id for question_id in order if question_id not in results]
|
|
||||||
)
|
|
||||||
if args.limit:
|
|
||||||
pending = pending[: args.limit]
|
|
||||||
|
|
||||||
no_reference = sum(question_id not in references for question_id in order)
|
|
||||||
one_reference = sum(len(references.get(question_id, [])) == 1 for question_id in order)
|
|
||||||
multiple_references = sum(len(references.get(question_id, [])) > 1 for question_id in order)
|
|
||||||
print(
|
|
||||||
f"total={len(order)} no_reference={no_reference} one_reference={one_reference} "
|
|
||||||
f"multiple_references={multiple_references} "
|
|
||||||
f"excluded={len(excluded)} "
|
|
||||||
f"only_reference_questions={args.only_reference_question_ids} "
|
|
||||||
f"concurrency={args.concurrency} submit_interval={args.submit_interval_seconds:g}s "
|
|
||||||
f"existing={len(results)} pending={len(pending)} output={args.output.resolve()}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if args.dry_run:
|
|
||||||
for question_id in pending:
|
|
||||||
print(
|
|
||||||
f"[would-run] question_id={question_id} workspace={mapping[question_id].name} "
|
|
||||||
f"references={len(references.get(question_id, []))}",
|
|
||||||
)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
executor = concurrent.futures.ThreadPoolExecutor(max_workers=args.concurrency)
|
|
||||||
active: dict[concurrent.futures.Future[dict[str, Any]], tuple[int, str]] = {}
|
|
||||||
next_position = 0
|
|
||||||
saved_count = 0
|
|
||||||
next_submit_at = 0.0
|
|
||||||
try:
|
|
||||||
while next_position < len(pending) or active:
|
|
||||||
can_submit = next_position < len(pending) and len(active) < args.concurrency
|
|
||||||
if can_submit and time.monotonic() >= next_submit_at:
|
|
||||||
question_id = pending[next_position]
|
|
||||||
position = next_position + 1
|
|
||||||
workspace = mapping[question_id]
|
|
||||||
print(
|
|
||||||
f"[submit {position}/{len(pending)}] question_id={question_id} "
|
|
||||||
f"workspace={workspace.name} references={len(references.get(question_id, []))}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
future = executor.submit(
|
|
||||||
run_one,
|
|
||||||
question_id,
|
|
||||||
workspace,
|
|
||||||
args.log_dir.resolve(),
|
|
||||||
reference_paths,
|
|
||||||
)
|
|
||||||
active[future] = (position, question_id)
|
|
||||||
next_position += 1
|
|
||||||
next_submit_at = time.monotonic() + args.submit_interval_seconds
|
|
||||||
continue
|
|
||||||
|
|
||||||
if not active:
|
|
||||||
time.sleep(max(0.0, next_submit_at - time.monotonic()))
|
|
||||||
continue
|
|
||||||
|
|
||||||
timeout = None
|
|
||||||
if can_submit:
|
|
||||||
timeout = max(0.0, next_submit_at - time.monotonic())
|
|
||||||
done, _ = concurrent.futures.wait(
|
|
||||||
active,
|
|
||||||
timeout=timeout,
|
|
||||||
return_when=concurrent.futures.FIRST_COMPLETED,
|
|
||||||
)
|
|
||||||
for future in done:
|
|
||||||
position, question_id = active.pop(future)
|
|
||||||
results[question_id] = future.result()
|
|
||||||
atomic_write_results(args.output.resolve(), full_order, results)
|
|
||||||
saved_count += 1
|
|
||||||
print(
|
|
||||||
f"[saved {saved_count}/{len(pending)}] submitted_position={position} " f"question_id={question_id}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
finally:
|
|
||||||
executor.shutdown(wait=True, cancel_futures=True)
|
|
||||||
|
|
||||||
print(
|
|
||||||
f"ALL FINISHED total_saved={sum(question_id in results for question_id in order)}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(main())
|
|
||||||
|
|
@ -1,216 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Run LongMemEval ``golden_check`` concurrently across samples.
|
|
||||||
|
|
||||||
For every workspace under ``datasets/longmemeval/<idx>`` in the selected numeric
|
|
||||||
range, this launches:
|
|
||||||
|
|
||||||
reme start config=jinli_lme job=golden_check
|
|
||||||
|
|
||||||
with ``LME_WORKSPACE_DIR`` pointed at that sample. Multiple samples can run at
|
|
||||||
once, capped by ``--concurrency``. The ``golden_check`` job itself waits for
|
|
||||||
``session_review.json`` when configured with ``wait_for_paths_step`` in
|
|
||||||
``jinli_lme.yaml``. Each sample's stdout/stderr goes to
|
|
||||||
``logs/golden_check/<idx>.log``.
|
|
||||||
|
|
||||||
By default the script processes samples 0..499 inclusive and reruns every sample
|
|
||||||
in that range. Pass ``--resume`` to skip samples whose ``check_golden.json``
|
|
||||||
already exists.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/run_golden_check.py
|
|
||||||
python benchmark/longmemeval/run_golden_check.py --start 187 --end 499
|
|
||||||
python benchmark/longmemeval/run_golden_check.py --concurrency 8 --stagger 1
|
|
||||||
python benchmark/longmemeval/run_golden_check.py --progress-interval 10
|
|
||||||
python benchmark/longmemeval/run_golden_check.py --resume
|
|
||||||
python benchmark/longmemeval/run_golden_check.py --limit 5 --dry-run
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import asyncio
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import time
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGDIR = REPO / "logs" / "golden_check"
|
|
||||||
OUTPUT_FILENAME = "check_golden.json"
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument("--start", type=int, default=0, help="first numeric sample id to process, inclusive (default 0)")
|
|
||||||
p.add_argument("--end", type=int, default=499, help="last numeric sample id to process, inclusive (default 499)")
|
|
||||||
p.add_argument("--limit", type=int, default=0, help="only process the first N selected samples (0 = all)")
|
|
||||||
p.add_argument("--concurrency", type=int, default=3, help="max samples running at once (default 3)")
|
|
||||||
p.add_argument("--stagger", type=float, default=1.0, help="seconds between consecutive launches (default 1)")
|
|
||||||
p.add_argument(
|
|
||||||
"--progress-interval",
|
|
||||||
type=float,
|
|
||||||
default=30.0,
|
|
||||||
help="seconds between progress reports while running (0 = disabled, default 30)",
|
|
||||||
)
|
|
||||||
p.add_argument(
|
|
||||||
"--resume",
|
|
||||||
action="store_true",
|
|
||||||
help=f"skip samples whose {OUTPUT_FILENAME} already exists",
|
|
||||||
)
|
|
||||||
p.add_argument("--dry-run", action="store_true", help="list what would run, launch nothing")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs (numeric workspace dirs), numerically sorted."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def output_is_current(idx: str) -> bool:
|
|
||||||
"""Return True when the sample already has a current-schema golden-check artifact."""
|
|
||||||
path = DATA / idx / OUTPUT_FILENAME
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return False
|
|
||||||
verdict = data.get("verdict") if isinstance(data, dict) else None
|
|
||||||
if not isinstance(verdict, dict):
|
|
||||||
return False
|
|
||||||
return isinstance(verdict.get("golden_answer_correct"), bool) and isinstance(
|
|
||||||
verdict.get("answer_session_ids_correct"),
|
|
||||||
bool,
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def print_progress(counters: dict, active: set[str], selected_total: int, started_at: float) -> None:
|
|
||||||
"""Print a one-line progress snapshot."""
|
|
||||||
finished = counters["done"] + counters["fail"] + counters["skip"]
|
|
||||||
running = len(active)
|
|
||||||
outstanding = max(selected_total - finished - running, 0)
|
|
||||||
elapsed = time.monotonic() - started_at
|
|
||||||
print(
|
|
||||||
f"[progress] selected={selected_total} done={counters['done']} fail={counters['fail']} "
|
|
||||||
f"skip={counters['skip']} running={running} outstanding={outstanding} "
|
|
||||||
f"elapsed={elapsed:.0f}s",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
async def progress_reporter(
|
|
||||||
counters: dict,
|
|
||||||
active: set[str],
|
|
||||||
selected_total: int,
|
|
||||||
started_at: float,
|
|
||||||
interval: float,
|
|
||||||
stop: asyncio.Event,
|
|
||||||
) -> None:
|
|
||||||
"""Periodically report progress until ``stop`` is set."""
|
|
||||||
if interval <= 0:
|
|
||||||
return
|
|
||||||
while not stop.is_set():
|
|
||||||
try:
|
|
||||||
await asyncio.wait_for(stop.wait(), timeout=interval)
|
|
||||||
except asyncio.TimeoutError:
|
|
||||||
print_progress(counters, active, selected_total, started_at)
|
|
||||||
|
|
||||||
|
|
||||||
async def run_one(idx: str, sem: asyncio.Semaphore, resume: bool, counters: dict, active: set[str]) -> None:
|
|
||||||
"""Run ``golden_check`` for one sample."""
|
|
||||||
if resume and output_is_current(idx):
|
|
||||||
counters["skip"] += 1
|
|
||||||
print(f"[skip] {idx} ({OUTPUT_FILENAME} exists)", flush=True)
|
|
||||||
return
|
|
||||||
|
|
||||||
async with sem:
|
|
||||||
active.add(idx)
|
|
||||||
log = LOGDIR / f"{idx}.log"
|
|
||||||
log.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
env = dict(os.environ, LME_WORKSPACE_DIR=f"datasets/longmemeval/{idx}")
|
|
||||||
|
|
||||||
started = time.strftime("%H:%M:%S")
|
|
||||||
print(f"[start {started}] {idx}", flush=True)
|
|
||||||
try:
|
|
||||||
with log.open("w", encoding="utf-8") as f:
|
|
||||||
proc = await asyncio.create_subprocess_exec(
|
|
||||||
"reme",
|
|
||||||
"start",
|
|
||||||
"config=jinli_lme",
|
|
||||||
"job=golden_check",
|
|
||||||
cwd=str(REPO),
|
|
||||||
env=env,
|
|
||||||
stdout=f,
|
|
||||||
stderr=asyncio.subprocess.STDOUT,
|
|
||||||
)
|
|
||||||
rc = await proc.wait()
|
|
||||||
|
|
||||||
ok = rc == 0 and output_is_current(idx)
|
|
||||||
counters["done" if ok else "fail"] += 1
|
|
||||||
tag = "done" if ok else "fail"
|
|
||||||
print(
|
|
||||||
f"[{tag}] {idx} rc={rc} log={log} ({counters['done']} done / {counters['fail']} fail)",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
finally:
|
|
||||||
active.discard(idx)
|
|
||||||
|
|
||||||
|
|
||||||
async def main() -> int:
|
|
||||||
"""Run the concurrent driver."""
|
|
||||||
args = parse_args()
|
|
||||||
if args.end < args.start:
|
|
||||||
raise ValueError(f"--end ({args.end}) must be >= --start ({args.start})")
|
|
||||||
if args.concurrency < 1:
|
|
||||||
raise ValueError("--concurrency must be >= 1")
|
|
||||||
if args.progress_interval < 0:
|
|
||||||
raise ValueError("--progress-interval must be >= 0")
|
|
||||||
|
|
||||||
LOGDIR.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
ids = [i for i in sample_ids() if args.start <= int(i) <= args.end]
|
|
||||||
if args.limit:
|
|
||||||
ids = ids[: args.limit]
|
|
||||||
|
|
||||||
pending = [i for i in ids if not (args.resume and output_is_current(i))]
|
|
||||||
print(
|
|
||||||
f"job=golden_check samples total={len(ids)} pending={len(pending)} "
|
|
||||||
f"range={args.start}..{args.end} resume={args.resume} "
|
|
||||||
f"concurrency={args.concurrency} stagger={args.stagger}s",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if args.dry_run:
|
|
||||||
for idx in pending:
|
|
||||||
print(f"[would-run] {idx}")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
sem = asyncio.Semaphore(args.concurrency)
|
|
||||||
counters = {"done": 0, "fail": 0, "skip": 0}
|
|
||||||
active: set[str] = set()
|
|
||||||
started_at = time.monotonic()
|
|
||||||
stop_progress = asyncio.Event()
|
|
||||||
progress_task = asyncio.create_task(
|
|
||||||
progress_reporter(counters, active, len(ids), started_at, args.progress_interval, stop_progress),
|
|
||||||
)
|
|
||||||
tasks: list[asyncio.Task] = []
|
|
||||||
try:
|
|
||||||
for n, idx in enumerate(ids):
|
|
||||||
if n and args.stagger > 0:
|
|
||||||
await asyncio.sleep(args.stagger)
|
|
||||||
tasks.append(asyncio.create_task(run_one(idx, sem, args.resume, counters, active)))
|
|
||||||
|
|
||||||
await asyncio.gather(*tasks)
|
|
||||||
finally:
|
|
||||||
stop_progress.set()
|
|
||||||
await progress_task
|
|
||||||
print_progress(counters, active, len(ids), started_at)
|
|
||||||
print(
|
|
||||||
f"ALL FINISHED done={counters['done']} fail={counters['fail']} skip={counters['skip']}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
return 0 if counters["fail"] == 0 else 1
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(asyncio.run(main()))
|
|
||||||
|
|
@ -1,203 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Run LongMemEval ``session_review`` concurrently across samples.
|
|
||||||
|
|
||||||
For every workspace under ``datasets/longmemeval/<idx>`` in the selected numeric
|
|
||||||
range, this launches:
|
|
||||||
|
|
||||||
reme start config=jinli_lme job=session_review
|
|
||||||
|
|
||||||
with ``LME_WORKSPACE_DIR`` pointed at that sample. Multiple samples can run at
|
|
||||||
once, capped by ``--concurrency``. By default this runner launches one sample at
|
|
||||||
a time; request submission is throttled inside each ``session_review`` process.
|
|
||||||
Each sample's stdout/stderr goes to ``logs/session_review/<idx>.log``.
|
|
||||||
|
|
||||||
By default the script processes samples 0..499 inclusive and reruns every sample
|
|
||||||
in that range. Pass ``--resume`` to skip samples whose ``session_review.json``
|
|
||||||
already exists.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/run_session_review.py
|
|
||||||
python benchmark/longmemeval/run_session_review.py --start 187 --end 499
|
|
||||||
python benchmark/longmemeval/run_session_review.py --concurrency 2
|
|
||||||
python benchmark/longmemeval/run_session_review.py --resume
|
|
||||||
python benchmark/longmemeval/run_session_review.py --limit 5 --dry-run
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import asyncio
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import time
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGDIR = REPO / "logs" / "session_review"
|
|
||||||
OUTPUT_FILENAME = "session_review.json"
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument("--start", type=int, default=0, help="first numeric sample id to process, inclusive (default 0)")
|
|
||||||
p.add_argument("--end", type=int, default=499, help="last numeric sample id to process, inclusive (default 499)")
|
|
||||||
p.add_argument("--limit", type=int, default=0, help="only process the first N selected samples (0 = all)")
|
|
||||||
p.add_argument("--concurrency", type=int, default=1, help="max samples running at once (default 1)")
|
|
||||||
p.add_argument("--stagger", type=float, default=1.0, help="seconds between worker launches (default 1)")
|
|
||||||
p.add_argument(
|
|
||||||
"--resume",
|
|
||||||
action="store_true",
|
|
||||||
help=f"skip samples whose {OUTPUT_FILENAME} already exists",
|
|
||||||
)
|
|
||||||
p.add_argument("--dry-run", action="store_true", help="list what would run, launch nothing")
|
|
||||||
p.add_argument("--stop-on-fail", action="store_true", help="stop immediately after the first failed sample")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs (numeric workspace dirs), numerically sorted."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def output_exists(idx: str) -> bool:
|
|
||||||
"""Return True when the sample already has a session review artifact."""
|
|
||||||
return (DATA / idx / OUTPUT_FILENAME).exists()
|
|
||||||
|
|
||||||
|
|
||||||
def output_is_healthy(idx: str) -> bool:
|
|
||||||
"""Return True when ``session_review.json`` exists and has no failed reviews."""
|
|
||||||
path = DATA / idx / OUTPUT_FILENAME
|
|
||||||
if not path.exists():
|
|
||||||
return False
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return False
|
|
||||||
review = data.get("review") if isinstance(data, dict) else None
|
|
||||||
if not isinstance(review, dict):
|
|
||||||
return False
|
|
||||||
raw = review.get("num_failed_reviews")
|
|
||||||
if isinstance(raw, int):
|
|
||||||
return raw == 0
|
|
||||||
failed_reviews = review.get("failed_reviews")
|
|
||||||
return not failed_reviews
|
|
||||||
|
|
||||||
|
|
||||||
async def run_one(idx: str, active: set[str]) -> bool:
|
|
||||||
"""Run ``session_review`` for one sample. Returns True on success."""
|
|
||||||
log = LOGDIR / f"{idx}.log"
|
|
||||||
log.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
env = dict(os.environ, LME_WORKSPACE_DIR=f"datasets/longmemeval/{idx}")
|
|
||||||
|
|
||||||
started = time.strftime("%H:%M:%S")
|
|
||||||
print(f"[start {started}] {idx}", flush=True)
|
|
||||||
active.add(idx)
|
|
||||||
try:
|
|
||||||
with log.open("w", encoding="utf-8") as f:
|
|
||||||
proc = await asyncio.create_subprocess_exec(
|
|
||||||
"reme",
|
|
||||||
"start",
|
|
||||||
"config=jinli_lme",
|
|
||||||
"job=session_review",
|
|
||||||
cwd=str(REPO),
|
|
||||||
env=env,
|
|
||||||
stdout=f,
|
|
||||||
stderr=asyncio.subprocess.STDOUT,
|
|
||||||
)
|
|
||||||
rc = await proc.wait()
|
|
||||||
finally:
|
|
||||||
active.discard(idx)
|
|
||||||
|
|
||||||
ok = rc == 0 and output_exists(idx)
|
|
||||||
tag = "done" if ok else "fail"
|
|
||||||
print(f"[{tag}] {idx} rc={rc} log={log}", flush=True)
|
|
||||||
return ok
|
|
||||||
|
|
||||||
|
|
||||||
async def worker(
|
|
||||||
name: int,
|
|
||||||
queue: asyncio.Queue[str],
|
|
||||||
args: argparse.Namespace,
|
|
||||||
counters: dict[str, int],
|
|
||||||
active: set[str],
|
|
||||||
stop: asyncio.Event,
|
|
||||||
) -> None:
|
|
||||||
"""Run samples from ``queue`` until exhausted or fail-fast is triggered."""
|
|
||||||
if name and args.stagger > 0:
|
|
||||||
await asyncio.sleep(args.stagger * name)
|
|
||||||
|
|
||||||
while not stop.is_set():
|
|
||||||
try:
|
|
||||||
idx = queue.get_nowait()
|
|
||||||
except asyncio.QueueEmpty:
|
|
||||||
return
|
|
||||||
|
|
||||||
try:
|
|
||||||
if args.resume and output_is_healthy(idx):
|
|
||||||
counters["skip"] += 1
|
|
||||||
print(f"[skip] {idx} (healthy {OUTPUT_FILENAME} exists)", flush=True)
|
|
||||||
continue
|
|
||||||
|
|
||||||
if await run_one(idx, active):
|
|
||||||
counters["done"] += 1
|
|
||||||
else:
|
|
||||||
counters["fail"] += 1
|
|
||||||
if args.stop_on_fail:
|
|
||||||
stop.set()
|
|
||||||
finally:
|
|
||||||
queue.task_done()
|
|
||||||
|
|
||||||
|
|
||||||
async def main() -> int:
|
|
||||||
"""Run the concurrent driver."""
|
|
||||||
args = parse_args()
|
|
||||||
if args.end < args.start:
|
|
||||||
raise ValueError(f"--end ({args.end}) must be >= --start ({args.start})")
|
|
||||||
if args.concurrency < 1:
|
|
||||||
raise ValueError("--concurrency must be >= 1")
|
|
||||||
if args.stagger < 0:
|
|
||||||
raise ValueError("--stagger must be >= 0")
|
|
||||||
|
|
||||||
LOGDIR.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
ids = [i for i in sample_ids() if args.start <= int(i) <= args.end]
|
|
||||||
if args.limit:
|
|
||||||
ids = ids[: args.limit]
|
|
||||||
|
|
||||||
pending = [i for i in ids if not (args.resume and output_exists(i))]
|
|
||||||
print(
|
|
||||||
f"job=session_review samples total={len(ids)} pending={len(pending)} "
|
|
||||||
f"range={args.start}..{args.end} resume={args.resume} "
|
|
||||||
f"concurrency={args.concurrency} stagger={args.stagger}s",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
|
|
||||||
if args.dry_run:
|
|
||||||
for idx in pending:
|
|
||||||
print(f"[would-run] {idx}")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
counters: dict[str, int] = {"done": 0, "fail": 0, "skip": 0}
|
|
||||||
active: set[str] = set()
|
|
||||||
stop = asyncio.Event()
|
|
||||||
queue: asyncio.Queue[str] = asyncio.Queue()
|
|
||||||
for idx in ids:
|
|
||||||
queue.put_nowait(idx)
|
|
||||||
|
|
||||||
workers = [
|
|
||||||
asyncio.create_task(worker(n, queue, args, counters, active, stop))
|
|
||||||
for n in range(min(args.concurrency, len(ids)))
|
|
||||||
]
|
|
||||||
await asyncio.gather(*workers)
|
|
||||||
|
|
||||||
print(
|
|
||||||
f"ALL FINISHED done={counters['done']} fail={counters['fail']} skip={counters['skip']}",
|
|
||||||
flush=True,
|
|
||||||
)
|
|
||||||
return 0 if counters["fail"] == 0 else 1
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(asyncio.run(main()))
|
|
||||||
|
|
@ -1,202 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Summarise the ``agentic_answer`` results across all LongMemEval samples.
|
|
||||||
|
|
||||||
Reports progress (how many of the 500 samples produced ``mem_answer.json``) and a
|
|
||||||
breakdown of answer *status*:
|
|
||||||
- answered — a non-empty answer that is not "not provided";
|
|
||||||
- not_provided — the agent gave up ("not provided");
|
|
||||||
- empty — ``mem_answer.json`` exists but the answer is blank;
|
|
||||||
- missing — no ``mem_answer.json`` yet.
|
|
||||||
|
|
||||||
Everything is broken down by ``question_type``. This script does NOT judge answer
|
|
||||||
correctness (there is no grader for ``mem_answer`` yet) — it only tracks progress
|
|
||||||
and collects predicted-vs-golden pairs. Tool-call statistics are read from the
|
|
||||||
aggregate written by ``run_agentic_answer.py`` when it is present.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/stats_agentic_answer.py
|
|
||||||
python benchmark/longmemeval/stats_agentic_answer.py --list-run-failed
|
|
||||||
python benchmark/longmemeval/stats_agentic_answer.py --list-unanswered
|
|
||||||
python benchmark/longmemeval/stats_agentic_answer.py --json
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import json
|
|
||||||
from collections import defaultdict
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGBASE = REPO / "logs" / "agentic_answer"
|
|
||||||
AGGREGATE = LOGBASE / "aggregate.json"
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument("--list-unanswered", action="store_true", help="list samples answered 'not provided' or empty")
|
|
||||||
p.add_argument("--list-run-failed", action="store_true", help="list launched samples with no readable output")
|
|
||||||
p.add_argument("--json", action="store_true", help="emit the summary as JSON")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs (numeric workspace dirs), numerically sorted."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def pct(num: int, den: int) -> str:
|
|
||||||
"""Format a percentage."""
|
|
||||||
return f"{(100.0 * num / den):.1f}%" if den else "n/a"
|
|
||||||
|
|
||||||
|
|
||||||
def logged_sample_ids() -> list[str]:
|
|
||||||
"""List sample IDs that have an agentic_answer launch log."""
|
|
||||||
logdir = LOGBASE / "agentic_answer"
|
|
||||||
if not logdir.exists():
|
|
||||||
return []
|
|
||||||
ids = [p.stem for p in logdir.glob("*.log") if p.stem.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def answer_status(pred: str, has_file: bool) -> str:
|
|
||||||
"""Classify an answer into answered / not_provided / empty / missing."""
|
|
||||||
if not has_file:
|
|
||||||
return "missing"
|
|
||||||
if not pred:
|
|
||||||
return "empty"
|
|
||||||
if "not provided" in pred.lower():
|
|
||||||
return "not_provided"
|
|
||||||
return "answered"
|
|
||||||
|
|
||||||
|
|
||||||
def load_tool_calls() -> dict[str, int]:
|
|
||||||
"""Map idx -> num_tool_calls from the aggregate, if it exists."""
|
|
||||||
if not AGGREGATE.exists():
|
|
||||||
return {}
|
|
||||||
try:
|
|
||||||
with AGGREGATE.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return {}
|
|
||||||
return {s["idx"]: s.get("num_tool_calls", 0) for s in data.get("samples", []) if "idx" in s}
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
"""Main entry point."""
|
|
||||||
args = parse_args()
|
|
||||||
ids = sample_ids()
|
|
||||||
total = len(ids)
|
|
||||||
tool_calls = load_tool_calls()
|
|
||||||
|
|
||||||
rows, unreadable = [], []
|
|
||||||
finished_ids = set()
|
|
||||||
for idx in ids:
|
|
||||||
query_path = DATA / idx / "query.json"
|
|
||||||
mem_path = DATA / idx / "mem_answer.json"
|
|
||||||
qtype = "(unknown)"
|
|
||||||
try:
|
|
||||||
with query_path.open(encoding="utf-8") as f:
|
|
||||||
qtype = json.load(f).get("question_type") or "(unknown)"
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
pass
|
|
||||||
|
|
||||||
has_file = mem_path.exists()
|
|
||||||
pred = ""
|
|
||||||
if has_file:
|
|
||||||
try:
|
|
||||||
with mem_path.open(encoding="utf-8") as f:
|
|
||||||
pred = str(json.load(f).get("answer") or "").strip()
|
|
||||||
finished_ids.add(idx)
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
unreadable.append(idx)
|
|
||||||
has_file = False
|
|
||||||
|
|
||||||
rows.append({"idx": idx, "type": qtype, "status": answer_status(pred, has_file)})
|
|
||||||
|
|
||||||
finished = [r for r in rows if r["status"] != "missing"]
|
|
||||||
n = len(finished)
|
|
||||||
launched = logged_sample_ids()
|
|
||||||
run_failed = [idx for idx in launched if idx not in finished_ids]
|
|
||||||
|
|
||||||
# Overall status tallies.
|
|
||||||
status_counts: dict[str, int] = defaultdict(int)
|
|
||||||
for r in rows:
|
|
||||||
status_counts[r["status"]] += 1
|
|
||||||
answered = status_counts["answered"]
|
|
||||||
unanswered = [r["idx"] for r in rows if r["status"] in ("not_provided", "empty")]
|
|
||||||
|
|
||||||
calls_vals = [tool_calls[i] for i in finished_ids if i in tool_calls]
|
|
||||||
avg_calls = sum(calls_vals) / len(calls_vals) if calls_vals else 0.0
|
|
||||||
|
|
||||||
# Per question_type breakdown.
|
|
||||||
by_type: dict[str, dict[str, int]] = defaultdict(lambda: {"n": 0, "answered": 0})
|
|
||||||
for r in finished:
|
|
||||||
by_type[r["type"]]["n"] += 1
|
|
||||||
by_type[r["type"]]["answered"] += 1 if r["status"] == "answered" else 0
|
|
||||||
|
|
||||||
if args.json:
|
|
||||||
print(
|
|
||||||
json.dumps(
|
|
||||||
{
|
|
||||||
"total": total,
|
|
||||||
"finished": n,
|
|
||||||
"pending": total - n - len(unreadable),
|
|
||||||
"unreadable": unreadable,
|
|
||||||
"launched": len(launched),
|
|
||||||
"run_failed": run_failed,
|
|
||||||
"status_counts": dict(status_counts),
|
|
||||||
"answered_rate": round(answered / n, 4) if n else None,
|
|
||||||
"avg_tool_calls": round(avg_calls, 2) if calls_vals else None,
|
|
||||||
"by_type": {
|
|
||||||
t: {**c, "answered_rate": round(c["answered"] / c["n"], 4)} for t, c in by_type.items()
|
|
||||||
},
|
|
||||||
"unanswered": unanswered,
|
|
||||||
"aggregate": str(AGGREGATE) if AGGREGATE.exists() else None,
|
|
||||||
},
|
|
||||||
ensure_ascii=False,
|
|
||||||
indent=2,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
print("=" * 60)
|
|
||||||
print("LongMemEval agentic_answer 统计")
|
|
||||||
print("=" * 60)
|
|
||||||
print(f"样例总数 : {total}")
|
|
||||||
print(f"已完成 (有产出) : {n} ({pct(n, total)})")
|
|
||||||
print(f"未完成 : {total - n - len(unreadable)}")
|
|
||||||
if unreadable:
|
|
||||||
print(f"损坏/无法解析 : {len(unreadable)} {unreadable}")
|
|
||||||
print(f"已启动过 (有 log) : {len(launched)}")
|
|
||||||
print(f"运行失败/无可读产出 : {len(run_failed)}")
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"已作答 (非 not provided): {answered} ({pct(answered, n)} of finished)")
|
|
||||||
print(f" 其中 not provided : {status_counts['not_provided']}")
|
|
||||||
print(f" 其中 空答案 : {status_counts['empty']}")
|
|
||||||
if calls_vals:
|
|
||||||
print(f"平均工具调用次数 : {avg_calls:.1f} (来自 {AGGREGATE.name})")
|
|
||||||
else:
|
|
||||||
print("平均工具调用次数 : n/a (先跑 run_agentic_answer.py 生成 aggregate.json)")
|
|
||||||
print("-" * 60)
|
|
||||||
print("按 question_type:")
|
|
||||||
print(f" {'type':<24} {'n':>4} {'已作答率':>12}")
|
|
||||||
for t in sorted(by_type):
|
|
||||||
c = by_type[t]
|
|
||||||
print(f" {t:<24} {c['n']:>4} {pct(c['answered'], c['n']):>12}")
|
|
||||||
|
|
||||||
if args.list_unanswered:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"not provided / 空答案的样例 ({len(unanswered)}): {unanswered}")
|
|
||||||
if args.list_run_failed:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"运行失败/无可读 mem_answer.json 的样例 ({len(run_failed)}): {run_failed}")
|
|
||||||
for idx in run_failed:
|
|
||||||
print(f" {idx}: {LOGBASE / 'agentic_answer' / f'{idx}.log'}")
|
|
||||||
print("=" * 60)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(main())
|
|
||||||
|
|
@ -1,344 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Summarise the ``check_golden.json`` verdicts across all LongMemEval samples.
|
|
||||||
|
|
||||||
Reports progress (how many of the 500 samples have finished) and accuracy:
|
|
||||||
- golden answer accuracy = share of finished samples whose golden answer the
|
|
||||||
auditor judged correct (``verdict.golden_answer_correct``);
|
|
||||||
- answer_session_ids accuracy = share whose claimed answer sessions the auditor
|
|
||||||
judged exactly correct (``verdict.answer_session_ids_correct``).
|
|
||||||
|
|
||||||
Everything is also broken down by ``question_type``. Use ``--list-bad`` to print
|
|
||||||
the samples whose golden answer was judged NOT correct.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/stats_golden_check.py
|
|
||||||
python benchmark/longmemeval/stats_golden_check.py --list-bad
|
|
||||||
python benchmark/longmemeval/stats_golden_check.py --list-run-failed
|
|
||||||
python benchmark/longmemeval/stats_golden_check.py --json
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import json
|
|
||||||
from collections import defaultdict
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGDIR = REPO / "logs" / "golden_check"
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__)
|
|
||||||
p.add_argument("--list-bad", action="store_true", help="list samples whose golden answer is NOT correct")
|
|
||||||
p.add_argument(
|
|
||||||
"--list-bad-sessions",
|
|
||||||
action="store_true",
|
|
||||||
help="list samples whose answer_session_ids is NOT correct",
|
|
||||||
)
|
|
||||||
p.add_argument(
|
|
||||||
"--list-run-failed",
|
|
||||||
action="store_true",
|
|
||||||
help="list launched samples that did not produce readable output",
|
|
||||||
)
|
|
||||||
p.add_argument("--json", action="store_true", help="emit the summary as JSON")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def pct(num: int, den: int) -> str:
|
|
||||||
"""Format a percentage."""
|
|
||||||
return f"{(100.0 * num / den):.1f}%" if den else "n/a"
|
|
||||||
|
|
||||||
|
|
||||||
def logged_sample_ids() -> list[str]:
|
|
||||||
"""List all sample IDs that have been launched but not finished."""
|
|
||||||
if not LOGDIR.exists():
|
|
||||||
return []
|
|
||||||
ids = [p.stem for p in LOGDIR.glob("*.log") if p.stem.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def load_json(path: Path) -> dict:
|
|
||||||
"""Load a JSON object, returning {} on any error."""
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
return data if isinstance(data, dict) else {}
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return {}
|
|
||||||
|
|
||||||
|
|
||||||
def question_type_for(idx: str, data: dict) -> str:
|
|
||||||
"""Return question_type from the output, session review, or query.json."""
|
|
||||||
question_type = str(data.get("question_type") or "").strip()
|
|
||||||
if question_type:
|
|
||||||
return question_type
|
|
||||||
|
|
||||||
review_path_raw = str(data.get("session_review_path") or "").strip()
|
|
||||||
review_path = Path(review_path_raw) if review_path_raw else DATA / idx / "session_review.json"
|
|
||||||
if not review_path.is_absolute():
|
|
||||||
review_path = REPO / review_path
|
|
||||||
review = load_json(review_path)
|
|
||||||
review_question_type = str((review.get("query") or {}).get("question_type") or "").strip()
|
|
||||||
if review_question_type:
|
|
||||||
return review_question_type
|
|
||||||
|
|
||||||
query = load_json(DATA / idx / "query.json")
|
|
||||||
return str(query.get("question_type") or "(unknown)").strip() or "(unknown)"
|
|
||||||
|
|
||||||
|
|
||||||
def question_id_for(idx: str, data: dict) -> str:
|
|
||||||
"""Return question_id from the output, session review, or query.json."""
|
|
||||||
question_id = str(data.get("question_id") or "").strip()
|
|
||||||
if question_id:
|
|
||||||
return question_id
|
|
||||||
|
|
||||||
review_path_raw = str(data.get("session_review_path") or "").strip()
|
|
||||||
review_path = Path(review_path_raw) if review_path_raw else DATA / idx / "session_review.json"
|
|
||||||
if not review_path.is_absolute():
|
|
||||||
review_path = REPO / review_path
|
|
||||||
review = load_json(review_path)
|
|
||||||
review_question_id = str((review.get("query") or {}).get("question_id") or "").strip()
|
|
||||||
if review_question_id:
|
|
||||||
return review_question_id
|
|
||||||
|
|
||||||
query = load_json(DATA / idx / "query.json")
|
|
||||||
return str(query.get("question_id") or "").strip()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_label(data: dict) -> str:
|
|
||||||
"""Format sample id as idx(question_id) when question_id is available."""
|
|
||||||
idx = str(data.get("_idx") or "")
|
|
||||||
qid = str(data.get("_question_id") or "").strip()
|
|
||||||
return f"{idx}({qid})" if qid else idx
|
|
||||||
|
|
||||||
|
|
||||||
def related_session_ids(data: dict) -> list[str]:
|
|
||||||
"""Return the best available session ids for a bad verdict record."""
|
|
||||||
verdict = data.get("verdict") if isinstance(data, dict) else None
|
|
||||||
if isinstance(verdict, dict):
|
|
||||||
true_ids = verdict.get("true_answer_session_ids")
|
|
||||||
if isinstance(true_ids, list):
|
|
||||||
ids = [str(session_id) for session_id in true_ids if str(session_id).strip()]
|
|
||||||
if ids:
|
|
||||||
return ids
|
|
||||||
|
|
||||||
summaries = data.get("session_summaries")
|
|
||||||
if isinstance(summaries, list):
|
|
||||||
return [
|
|
||||||
str(summary.get("session_id"))
|
|
||||||
for summary in summaries
|
|
||||||
if isinstance(summary, dict) and str(summary.get("session_id") or "").strip()
|
|
||||||
]
|
|
||||||
return []
|
|
||||||
|
|
||||||
|
|
||||||
def grouped_records(records: list[dict]) -> dict[str, list[dict]]:
|
|
||||||
"""Group records by question_type for human-readable list output."""
|
|
||||||
grouped: dict[str, list[dict]] = defaultdict(list)
|
|
||||||
for data in records:
|
|
||||||
question_type = str(data.get("_question_type") or "(unknown)")
|
|
||||||
grouped[question_type].append(
|
|
||||||
{
|
|
||||||
"index": str(data.get("_idx") or ""),
|
|
||||||
"question_id": str(data.get("_question_id") or ""),
|
|
||||||
"session_id": related_session_ids(data),
|
|
||||||
},
|
|
||||||
)
|
|
||||||
return dict(sorted(grouped.items()))
|
|
||||||
|
|
||||||
|
|
||||||
def verdict_bool(verdict: dict, new_key: str, old_key: str) -> bool:
|
|
||||||
"""Read a verdict boolean, accepting the old field name for compatibility."""
|
|
||||||
if verdict.get(new_key) is True:
|
|
||||||
return True
|
|
||||||
if verdict.get(new_key) is False:
|
|
||||||
return False
|
|
||||||
return verdict.get(old_key) is True
|
|
||||||
|
|
||||||
|
|
||||||
def has_current_verdict(data: dict) -> bool:
|
|
||||||
"""Return True when ``check_golden.json`` uses the current golden_check schema."""
|
|
||||||
verdict = data.get("verdict") if isinstance(data, dict) else None
|
|
||||||
if not isinstance(verdict, dict):
|
|
||||||
return False
|
|
||||||
return isinstance(verdict.get("golden_answer_correct"), bool) and isinstance(
|
|
||||||
verdict.get("answer_session_ids_correct"),
|
|
||||||
bool,
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def write_golden_check_list(done: list[dict], output_path: Path) -> None:
|
|
||||||
"""Write all readable check_golden records as JSONL."""
|
|
||||||
with output_path.open("w", encoding="utf-8") as f:
|
|
||||||
for data in done:
|
|
||||||
f.write(json.dumps(data, ensure_ascii=False))
|
|
||||||
f.write("\n")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
"""Main entry point."""
|
|
||||||
args = parse_args()
|
|
||||||
ids = sample_ids()
|
|
||||||
total = len(ids)
|
|
||||||
|
|
||||||
done, unreadable, stale = [], [], []
|
|
||||||
finished_ids = set()
|
|
||||||
for idx in ids:
|
|
||||||
path = DATA / idx / "check_golden.json"
|
|
||||||
if not path.exists():
|
|
||||||
continue
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
if not has_current_verdict(data):
|
|
||||||
stale.append(idx)
|
|
||||||
continue
|
|
||||||
data["_idx"] = idx
|
|
||||||
data["_question_type"] = question_type_for(idx, data)
|
|
||||||
data["_question_id"] = question_id_for(idx, data)
|
|
||||||
done.append(data)
|
|
||||||
finished_ids.add(idx)
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
unreadable.append(idx)
|
|
||||||
|
|
||||||
n = len(done)
|
|
||||||
output_path = Path.cwd() / "golden_check_list.jsonl"
|
|
||||||
write_golden_check_list(done, output_path)
|
|
||||||
launched = logged_sample_ids()
|
|
||||||
run_failed = [idx for idx in launched if idx not in finished_ids]
|
|
||||||
|
|
||||||
# Overall tallies.
|
|
||||||
golden_ok = sum(
|
|
||||||
1 for d in done if verdict_bool(d.get("verdict", {}), "golden_answer_correct", "golden_answer_reasonable")
|
|
||||||
)
|
|
||||||
sess_ok = sum(
|
|
||||||
1
|
|
||||||
for d in done
|
|
||||||
if verdict_bool(d.get("verdict", {}), "answer_session_ids_correct", "answer_session_ids_reasonable")
|
|
||||||
)
|
|
||||||
both_ok = sum(
|
|
||||||
1
|
|
||||||
for d in done
|
|
||||||
if verdict_bool(d.get("verdict", {}), "golden_answer_correct", "golden_answer_reasonable")
|
|
||||||
and verdict_bool(d.get("verdict", {}), "answer_session_ids_correct", "answer_session_ids_reasonable")
|
|
||||||
)
|
|
||||||
|
|
||||||
# Per question_type breakdown.
|
|
||||||
by_type: dict[str, dict[str, int]] = defaultdict(lambda: {"n": 0, "golden_ok": 0, "sess_ok": 0, "both_ok": 0})
|
|
||||||
for d in done:
|
|
||||||
v = d.get("verdict", {})
|
|
||||||
golden_is_ok = verdict_bool(v, "golden_answer_correct", "golden_answer_reasonable")
|
|
||||||
sess_is_ok = verdict_bool(v, "answer_session_ids_correct", "answer_session_ids_reasonable")
|
|
||||||
t = d.get("_question_type") or "(unknown)"
|
|
||||||
by_type[t]["n"] += 1
|
|
||||||
by_type[t]["golden_ok"] += 1 if golden_is_ok else 0
|
|
||||||
by_type[t]["sess_ok"] += 1 if sess_is_ok else 0
|
|
||||||
by_type[t]["both_ok"] += 1 if golden_is_ok and sess_is_ok else 0
|
|
||||||
|
|
||||||
bad_golden_records = [
|
|
||||||
d for d in done if not verdict_bool(d.get("verdict", {}), "golden_answer_correct", "golden_answer_reasonable")
|
|
||||||
]
|
|
||||||
bad_session_records = [
|
|
||||||
d
|
|
||||||
for d in done
|
|
||||||
if not verdict_bool(d.get("verdict", {}), "answer_session_ids_correct", "answer_session_ids_reasonable")
|
|
||||||
]
|
|
||||||
bad_golden = [d["_idx"] for d in bad_golden_records]
|
|
||||||
bad_sessions = [d["_idx"] for d in bad_session_records]
|
|
||||||
|
|
||||||
if args.json:
|
|
||||||
print(
|
|
||||||
json.dumps(
|
|
||||||
{
|
|
||||||
"total": total,
|
|
||||||
"finished": n,
|
|
||||||
"pending": total - n - len(unreadable),
|
|
||||||
"unreadable": unreadable,
|
|
||||||
"stale": stale,
|
|
||||||
"launched": len(launched),
|
|
||||||
"run_failed": run_failed,
|
|
||||||
"golden_answer_accuracy": round(golden_ok / n, 4) if n else None,
|
|
||||||
"answer_session_ids_accuracy": round(sess_ok / n, 4) if n else None,
|
|
||||||
"both_correct_rate": round(both_ok / n, 4) if n else None,
|
|
||||||
"golden_ok": golden_ok,
|
|
||||||
"sess_ok": sess_ok,
|
|
||||||
"both_ok": both_ok,
|
|
||||||
"by_type": {
|
|
||||||
t: {
|
|
||||||
**c,
|
|
||||||
"golden_bad": c["n"] - c["golden_ok"],
|
|
||||||
"session_bad": c["n"] - c["sess_ok"],
|
|
||||||
"both_bad": c["n"] - c["both_ok"],
|
|
||||||
"golden_acc": round(c["golden_ok"] / c["n"], 4),
|
|
||||||
"session_acc": round(c["sess_ok"] / c["n"], 4),
|
|
||||||
"both_acc": round(c["both_ok"] / c["n"], 4),
|
|
||||||
}
|
|
||||||
for t, c in by_type.items()
|
|
||||||
},
|
|
||||||
"bad_golden": bad_golden,
|
|
||||||
"bad_sessions": bad_sessions,
|
|
||||||
"golden_check_list": str(output_path),
|
|
||||||
},
|
|
||||||
ensure_ascii=False,
|
|
||||||
indent=2,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
print("=" * 60)
|
|
||||||
print("LongMemEval golden_check 统计")
|
|
||||||
print("=" * 60)
|
|
||||||
print(f"样例总数 : {total}")
|
|
||||||
print(f"已完成 (有产出) : {n} ({pct(n, total)})")
|
|
||||||
print(f"未完成 : {total - n - len(unreadable)}")
|
|
||||||
if unreadable:
|
|
||||||
print(f"损坏/无法解析 : {len(unreadable)} {unreadable}")
|
|
||||||
if stale:
|
|
||||||
print(f"旧格式待重跑 : {len(stale)} {stale}")
|
|
||||||
print(f"已合并 JSONL : {output_path}")
|
|
||||||
print(f"已启动过 (有 log) : {len(launched)}")
|
|
||||||
print(f"运行失败/无可读产出 : {len(run_failed)}")
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"golden answer 正确率 : {pct(golden_ok, n)} ({golden_ok}/{n})")
|
|
||||||
print(f"answer_session 正确率: {pct(sess_ok, n)} ({sess_ok}/{n})")
|
|
||||||
print(f"两者都正确 : {pct(both_ok, n)} ({both_ok}/{n})")
|
|
||||||
print("-" * 60)
|
|
||||||
print("按 question_type:")
|
|
||||||
print(
|
|
||||||
f" {'type':<24} {'n':>4} {'golden正确率':>14} {'golden错误':>10} "
|
|
||||||
f"{'session正确率':>14} {'session错误':>11} {'都正确':>10} {'都正确错误':>12}",
|
|
||||||
)
|
|
||||||
for t in sorted(by_type):
|
|
||||||
c = by_type[t]
|
|
||||||
print(
|
|
||||||
f" {t:<24} {c['n']:>4} {pct(c['golden_ok'], c['n']):>14} {c['n'] - c['golden_ok']:>10} "
|
|
||||||
f"{pct(c['sess_ok'], c['n']):>14} {c['n'] - c['sess_ok']:>11} "
|
|
||||||
f"{pct(c['both_ok'], c['n']):>10} {c['n'] - c['both_ok']:>12}",
|
|
||||||
)
|
|
||||||
|
|
||||||
if args.list_bad:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"golden answer 判为不正确的样例 ({len(bad_golden_records)}):")
|
|
||||||
print(json.dumps(grouped_records(bad_golden_records), ensure_ascii=False))
|
|
||||||
if args.list_bad_sessions:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"answer_session_ids 判为不正确的样例 ({len(bad_session_records)}):")
|
|
||||||
print(json.dumps(grouped_records(bad_session_records), ensure_ascii=False))
|
|
||||||
if args.list_run_failed:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"运行失败/无可读 check_golden.json 的样例 ({len(run_failed)}): {run_failed}")
|
|
||||||
for idx in run_failed:
|
|
||||||
print(f" {idx}: {LOGDIR / f'{idx}.log'}")
|
|
||||||
print("=" * 60)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(main())
|
|
||||||
|
|
@ -1,248 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""Summarise LongMemEval ``session_review.json`` artifacts.
|
|
||||||
|
|
||||||
This script is for upstream health checks before running ``golden_check``.
|
|
||||||
Samples with retryable per-session failures should be rerun as a whole; samples
|
|
||||||
with non-retryable fallback reviews are reported separately.
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
python benchmark/longmemeval/stats_session_review.py
|
|
||||||
python benchmark/longmemeval/stats_session_review.py --list-failed
|
|
||||||
python benchmark/longmemeval/stats_session_review.py --list-fallback
|
|
||||||
python benchmark/longmemeval/stats_session_review.py --json
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import json
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
REPO = Path(__file__).resolve().parents[2]
|
|
||||||
DATA = REPO / "datasets" / "longmemeval"
|
|
||||||
LOGDIR = REPO / "logs" / "session_review"
|
|
||||||
OUTPUT_FILENAME = "session_review.json"
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
|
||||||
"""Parse command-line arguments."""
|
|
||||||
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
||||||
p.add_argument("--list-failed", action="store_true", help="list samples with retryable failed per-session reviews")
|
|
||||||
p.add_argument("--list-fallback", action="store_true", help="list non-retryable fallback reviews")
|
|
||||||
p.add_argument("--list-missing", action="store_true", help="list samples missing session_review.json")
|
|
||||||
p.add_argument("--list-run-failed", action="store_true", help="list launched samples without a healthy output")
|
|
||||||
p.add_argument("--json", action="store_true", help="emit the summary as JSON")
|
|
||||||
return p.parse_args()
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids() -> list[str]:
|
|
||||||
"""List all numeric sample IDs."""
|
|
||||||
ids = [p.name for p in DATA.iterdir() if p.is_dir() and p.name.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def pct(num: int, den: int) -> str:
|
|
||||||
"""Format a percentage."""
|
|
||||||
return f"{(100.0 * num / den):.1f}%" if den else "n/a"
|
|
||||||
|
|
||||||
|
|
||||||
def load_json(path: Path) -> dict:
|
|
||||||
"""Load a JSON object, returning {} on any error."""
|
|
||||||
try:
|
|
||||||
with path.open(encoding="utf-8") as f:
|
|
||||||
data = json.load(f)
|
|
||||||
return data if isinstance(data, dict) else {}
|
|
||||||
except (OSError, json.JSONDecodeError):
|
|
||||||
return {}
|
|
||||||
|
|
||||||
|
|
||||||
def logged_sample_ids() -> list[str]:
|
|
||||||
"""List sample IDs that have a session_review runner log."""
|
|
||||||
if not LOGDIR.exists():
|
|
||||||
return []
|
|
||||||
ids = [p.stem for p in LOGDIR.glob("*.log") if p.stem.isdigit()]
|
|
||||||
return sorted(ids, key=int)
|
|
||||||
|
|
||||||
|
|
||||||
def review_block(data: dict) -> dict:
|
|
||||||
"""Return the review block when present."""
|
|
||||||
review = data.get("review") if isinstance(data, dict) else None
|
|
||||||
return review if isinstance(review, dict) else {}
|
|
||||||
|
|
||||||
|
|
||||||
def failure_details(data: dict) -> list[dict]:
|
|
||||||
"""Return retryable failed_reviews when present."""
|
|
||||||
failed_reviews = review_block(data).get("failed_reviews")
|
|
||||||
if not isinstance(failed_reviews, list):
|
|
||||||
return []
|
|
||||||
return [item for item in failed_reviews if isinstance(item, dict) and not item.get("fallback")]
|
|
||||||
|
|
||||||
|
|
||||||
def fallback_details(data: dict) -> list[dict]:
|
|
||||||
"""Return non-retryable fallback review details when present."""
|
|
||||||
review = review_block(data)
|
|
||||||
fallback_reviews = review.get("fallback_reviews")
|
|
||||||
if isinstance(fallback_reviews, list):
|
|
||||||
return [item for item in fallback_reviews if isinstance(item, dict)]
|
|
||||||
|
|
||||||
failed_reviews = review.get("failed_reviews")
|
|
||||||
if isinstance(failed_reviews, list):
|
|
||||||
return [item for item in failed_reviews if isinstance(item, dict) and item.get("fallback")]
|
|
||||||
return []
|
|
||||||
|
|
||||||
|
|
||||||
def failure_count(data: dict) -> int:
|
|
||||||
"""Return retryable failed review count."""
|
|
||||||
review = review_block(data)
|
|
||||||
raw = review.get("num_failed_reviews")
|
|
||||||
raw_fallback = review.get("num_fallback_reviews")
|
|
||||||
if isinstance(raw, int) and isinstance(raw_fallback, int):
|
|
||||||
return max(0, raw - raw_fallback)
|
|
||||||
return len(failure_details(data))
|
|
||||||
|
|
||||||
|
|
||||||
def fallback_count(data: dict) -> int:
|
|
||||||
"""Return non-retryable fallback review count."""
|
|
||||||
review = review_block(data)
|
|
||||||
raw = review.get("num_fallback_reviews")
|
|
||||||
if isinstance(raw, int):
|
|
||||||
return raw
|
|
||||||
return len(fallback_details(data))
|
|
||||||
|
|
||||||
|
|
||||||
def question_id(data: dict) -> str:
|
|
||||||
"""Return query.question_id when present."""
|
|
||||||
query = data.get("query") if isinstance(data, dict) else None
|
|
||||||
if not isinstance(query, dict):
|
|
||||||
return ""
|
|
||||||
return str(query.get("question_id") or "").strip()
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
"""Main entry point."""
|
|
||||||
args = parse_args()
|
|
||||||
ids = sample_ids()
|
|
||||||
total = len(ids)
|
|
||||||
|
|
||||||
healthy, failed, fallback, missing, unreadable = [], [], [], [], []
|
|
||||||
total_failed_sessions = 0
|
|
||||||
total_fallback_sessions = 0
|
|
||||||
failed_details_by_id: dict[str, list[dict]] = {}
|
|
||||||
fallback_details_by_id: dict[str, list[dict]] = {}
|
|
||||||
question_id_by_id: dict[str, str] = {}
|
|
||||||
|
|
||||||
for idx in ids:
|
|
||||||
path = DATA / idx / OUTPUT_FILENAME
|
|
||||||
if not path.exists():
|
|
||||||
missing.append(idx)
|
|
||||||
continue
|
|
||||||
data = load_json(path)
|
|
||||||
if not data:
|
|
||||||
unreadable.append(idx)
|
|
||||||
continue
|
|
||||||
question_id_by_id[idx] = question_id(data)
|
|
||||||
n_failed = failure_count(data)
|
|
||||||
n_fallback = fallback_count(data)
|
|
||||||
if n_failed:
|
|
||||||
failed.append(idx)
|
|
||||||
total_failed_sessions += n_failed
|
|
||||||
failed_details_by_id[idx] = failure_details(data)
|
|
||||||
if n_fallback:
|
|
||||||
fallback.append(idx)
|
|
||||||
total_fallback_sessions += n_fallback
|
|
||||||
fallback_details_by_id[idx] = fallback_details(data)
|
|
||||||
if not n_failed:
|
|
||||||
healthy.append(idx)
|
|
||||||
|
|
||||||
launched = logged_sample_ids()
|
|
||||||
healthy_set = set(healthy)
|
|
||||||
run_failed = [idx for idx in launched if idx not in healthy_set]
|
|
||||||
|
|
||||||
if args.json:
|
|
||||||
print(
|
|
||||||
json.dumps(
|
|
||||||
{
|
|
||||||
"total": total,
|
|
||||||
"healthy": len(healthy),
|
|
||||||
"failed_samples": failed,
|
|
||||||
"failed_sample_count": len(failed),
|
|
||||||
"failed_session_count": total_failed_sessions,
|
|
||||||
"fallback_samples": fallback,
|
|
||||||
"fallback_sample_count": len(fallback),
|
|
||||||
"fallback_session_count": total_fallback_sessions,
|
|
||||||
"missing": missing,
|
|
||||||
"unreadable": unreadable,
|
|
||||||
"launched": len(launched),
|
|
||||||
"run_failed_or_unhealthy": run_failed,
|
|
||||||
"failed_details": failed_details_by_id,
|
|
||||||
"fallback_details": fallback_details_by_id,
|
|
||||||
},
|
|
||||||
ensure_ascii=False,
|
|
||||||
indent=2,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
print("=" * 60)
|
|
||||||
print("LongMemEval session_review 统计")
|
|
||||||
print("=" * 60)
|
|
||||||
print(f"样例总数 : {total}")
|
|
||||||
print(f"可继续产出 : {len(healthy)} ({pct(len(healthy), total)})")
|
|
||||||
print(f"有可重试失败 : {len(failed)}")
|
|
||||||
print(f"可重试失败 session : {total_failed_sessions}")
|
|
||||||
print(f"有不可重试 fallback : {len(fallback)}")
|
|
||||||
print(f"fallback session : {total_fallback_sessions}")
|
|
||||||
print(f"缺少 session_review : {len(missing)}")
|
|
||||||
print(f"损坏/无法解析 : {len(unreadable)}")
|
|
||||||
print(f"已启动过 (有 log) : {len(launched)}")
|
|
||||||
print(f"运行失败/非健康产出 : {len(run_failed)}")
|
|
||||||
print("-" * 60)
|
|
||||||
print("有可重试 failed_reviews 的样例需要整体重跑:")
|
|
||||||
if failed:
|
|
||||||
print(" ".join(failed))
|
|
||||||
print("重跑命令示例:")
|
|
||||||
print(f"python benchmark/longmemeval/run_session_review.py --start {failed[0]} --end {failed[0]}")
|
|
||||||
else:
|
|
||||||
print("(none)")
|
|
||||||
if fallback:
|
|
||||||
print("-" * 60)
|
|
||||||
print("不可重试 fallback 的样例不用重跑:")
|
|
||||||
for idx in fallback:
|
|
||||||
details = fallback_details_by_id.get(idx) or []
|
|
||||||
session_ids = [str(item.get("session_id") or "(unknown)") for item in details]
|
|
||||||
qid = question_id_by_id.get(idx)
|
|
||||||
sample_label = f"{idx}({qid})" if qid else idx
|
|
||||||
print(f"{sample_label}: {' '.join(session_ids) if session_ids else '(unknown)'}")
|
|
||||||
|
|
||||||
if args.list_failed and failed:
|
|
||||||
print("-" * 60)
|
|
||||||
for idx in failed:
|
|
||||||
details = failed_details_by_id.get(idx) or []
|
|
||||||
print(f"{idx}: {DATA / idx / OUTPUT_FILENAME} failed_sessions={len(details)}")
|
|
||||||
for item in details:
|
|
||||||
session_id = item.get("session_id", "(unknown)")
|
|
||||||
error = str(item.get("error") or "").replace("\n", " ")
|
|
||||||
print(f" - {session_id}: {error}")
|
|
||||||
if args.list_fallback and fallback:
|
|
||||||
print("-" * 60)
|
|
||||||
for idx in fallback:
|
|
||||||
details = fallback_details_by_id.get(idx) or []
|
|
||||||
print(f"{idx}: {DATA / idx / OUTPUT_FILENAME} fallback_sessions={len(details)}")
|
|
||||||
for item in details:
|
|
||||||
session_id = item.get("session_id", "(unknown)")
|
|
||||||
reason = str(item.get("fallback_reason") or "fallback")
|
|
||||||
error = str(item.get("error") or "").replace("\n", " ")
|
|
||||||
raw_saved = "yes" if item.get("raw_session") else "no"
|
|
||||||
print(f" - {session_id}: reason={reason} raw_session_saved={raw_saved} error={error}")
|
|
||||||
if args.list_missing and missing:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"缺少 session_review.json 的样例 ({len(missing)}): {missing}")
|
|
||||||
if args.list_run_failed and run_failed:
|
|
||||||
print("-" * 60)
|
|
||||||
print(f"运行失败/非健康产出的样例 ({len(run_failed)}): {run_failed}")
|
|
||||||
for idx in run_failed:
|
|
||||||
print(f" {idx}: {LOGDIR / f'{idx}.log'}")
|
|
||||||
print("=" * 60)
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
raise SystemExit(main())
|
|
||||||
14
benchmark/pibench/.gitignore
vendored
Normal file
14
benchmark/pibench/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,14 @@
|
||||||
|
# 含真实 API key,绝不入库
|
||||||
|
env.sh
|
||||||
|
|
||||||
|
# 运行时产物(含对话内容,勿入库)
|
||||||
|
logs/
|
||||||
|
outputs/
|
||||||
|
reme_workspace/
|
||||||
|
nanobot_workspace/
|
||||||
|
|
||||||
|
# 数据符号链接(指向外部 π-Bench 仓库)
|
||||||
|
data
|
||||||
|
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
327
benchmark/pibench/README.md
Normal file
327
benchmark/pibench/README.md
Normal file
|
|
@ -0,0 +1,327 @@
|
||||||
|
[中文版 / Chinese version](./README_ZH.md)
|
||||||
|
|
||||||
|
# π-Bench Evaluation Suite
|
||||||
|
|
||||||
|
A glue layer that connects the **ReMe agent (with persistent memory)** to
|
||||||
|
**π-Bench** (Proactive Personal Assistant Benchmark). This directory contains
|
||||||
|
only the minimal code and configuration needed for the integration: the
|
||||||
|
π-Bench framework (`src/`), evaluation data (`data/`), the AppWorld tool
|
||||||
|
environment, and ReMe itself are all **external third-party dependencies**,
|
||||||
|
referenced in place via symlink and environment variables and never bundled
|
||||||
|
with this suite.
|
||||||
|
|
||||||
|
- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench (arXiv: 2605.14678)
|
||||||
|
- ReMe: the root of the ReMe repository this suite lives in (recommended
|
||||||
|
location: `ReMe/benchmark/pibench/`)
|
||||||
|
|
||||||
|
## 1. Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
π-Bench runner (src.main --mode run)
|
||||||
|
│ user_agent (simulated-user LLM) walks data/{persona}/episode.yaml
|
||||||
|
│ task by task, chatting with the agent over multiple turns and judging
|
||||||
|
│ hidden intents (PROC) during the run phase
|
||||||
|
▼
|
||||||
|
test server (π-Bench scripts/test_server.py, HTTP long-polling)
|
||||||
|
▲ /send │ /poll
|
||||||
|
│ ▼
|
||||||
|
bridge_reme.py ──────────────► ReMe Application (embedded as a library)
|
||||||
|
│ ├─ agent_wrapper: agent under test (AgentScope)
|
||||||
|
│ ├─ jobs: search / auto_memory / daily_write
|
||||||
|
│ └─ workspace: reme_workspace/{persona}/
|
||||||
|
│ (isolated persistent memory per persona)
|
||||||
|
└──── MCP ────► AppWorld MCP ────► AppWorld APIs (tool/app environment)
|
||||||
|
|
||||||
|
π-Bench runner (src.main --mode eval)
|
||||||
|
judger (judge LLM) reads the traces and scores each checklist item (COMP)
|
||||||
|
```
|
||||||
|
|
||||||
|
Key points:
|
||||||
|
- The bridge runs on **ReMe's own venv python** and uses ReMe as a library
|
||||||
|
(`resolve_app_config` + `Application`); **no ReMe source modification** is
|
||||||
|
required.
|
||||||
|
- Every incoming user message automatically triggers a ReMe memory `search`
|
||||||
|
and injects the matched memories (tuning knobs in §8); on task end (reset)
|
||||||
|
the session is distilled into daily notes by `auto_memory`.
|
||||||
|
- Tool calls executed by the agent (AppWorld MCP + ReMe job tools) are
|
||||||
|
captured per turn into the trace as `tool_steps`, so π-Bench
|
||||||
|
`tools_evaluation_path` scripts can score tool behavior (§7).
|
||||||
|
- π-Bench's `data/`, `src/` and AppWorld are not part of this suite; install
|
||||||
|
π-Bench first (§3.1).
|
||||||
|
|
||||||
|
## 2. Directory layout
|
||||||
|
|
||||||
|
```
|
||||||
|
pibench/
|
||||||
|
├── README.md / README_ZH.md # this document (English / Chinese)
|
||||||
|
├── env.sh.example # environment template (copy to env.sh, fill TODOs)
|
||||||
|
├── bridge_reme.py # ReMe ↔ test server bridge (memory inject/save,
|
||||||
|
│ # profile injection, tool-trace capture)
|
||||||
|
├── run_persona.sh # full pipeline for ONE persona (5 services + run + eval)
|
||||||
|
├── run_all.sh # batch over 5 personas (fresh/resume, default parallel=2)
|
||||||
|
├── resume.py # checkpoint resume: completion detection + surgical
|
||||||
|
│ # cleanup of interrupted tasks' residual memory
|
||||||
|
├── fix_trace_logs.py # run outputs → ~/.nanobot/trace_logs conversion,
|
||||||
|
│ # merging tool sidecars into turn files (pre-eval)
|
||||||
|
├── .gitignore # excludes env.sh and all runtime artifacts
|
||||||
|
└── config/
|
||||||
|
├── models/reme.yaml # runner model config (model_id=reme)
|
||||||
|
└── bench/evaluation/trace_history.yaml # trace render policy (shipped with
|
||||||
|
# the suite; passed via --history-config-path)
|
||||||
|
```
|
||||||
|
|
||||||
|
Generated at runtime (all git-ignored): `data` (symlink), `logs/`, `outputs/`,
|
||||||
|
`reme_workspace/`, `nanobot_workspace/`.
|
||||||
|
|
||||||
|
## 3. Prerequisites (third-party, install first)
|
||||||
|
|
||||||
|
### 3.1 π-Bench repository (with AppWorld)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://github.com/Simplified-Reasoning/Pi-Bench.git <pi-bench-dir>
|
||||||
|
cd <pi-bench-dir>
|
||||||
|
python3.11 -m venv .venv # scripts expect exactly this venv name
|
||||||
|
source .venv/bin/activate
|
||||||
|
pip install -e . # pibench runner (src.main)
|
||||||
|
bash scripts/setup_appworld.sh # install AppWorld and download its data (large)
|
||||||
|
```
|
||||||
|
|
||||||
|
Post-install sanity checks:
|
||||||
|
```bash
|
||||||
|
ls data/ # should contain researcher marketer pharmacist law_trainee Financier
|
||||||
|
.venv/bin/python -c "import src" && echo OK
|
||||||
|
.venv/bin/appworld --help >/dev/null && echo OK
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 ReMe repository
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd <reme-dir> # ReMe repository root (contains the reme/ package)
|
||||||
|
python3.11 -m venv .venv # scripts expect exactly this venv name
|
||||||
|
source .venv/bin/activate
|
||||||
|
pip install -e . # or ReMe's own install flow; `import reme` must work
|
||||||
|
```
|
||||||
|
|
||||||
|
Sanity check: `.venv/bin/python -c "import reme; print('ok')"`
|
||||||
|
|
||||||
|
## 4. Install this suite (step by step)
|
||||||
|
|
||||||
|
1. **Place the suite** (recommended inside the ReMe repo so `REME_DIR` is
|
||||||
|
inferred automatically):
|
||||||
|
```bash
|
||||||
|
cp -r pibench <reme-dir>/benchmark/pibench
|
||||||
|
cd <reme-dir>/benchmark/pibench
|
||||||
|
```
|
||||||
|
If placed elsewhere, set `REME_DIR` explicitly in env.sh later.
|
||||||
|
|
||||||
|
2. **Create the environment file and fill in the custom parameters**:
|
||||||
|
```bash
|
||||||
|
cp env.sh.example env.sh
|
||||||
|
```
|
||||||
|
Open `env.sh`; required items (marked TODO):
|
||||||
|
| Variable | Description |
|
||||||
|
|---|---|
|
||||||
|
| `PI_BENCH_ROOT` | π-Bench repo root (contains `src/` `data/` `.venv` `third_party/appworld`) |
|
||||||
|
| `USER_API_KEY` | API key of the simulated-user LLM (run phase, hidden-intent judging) |
|
||||||
|
| `JUDGER_API_KEY` | API key of the judger LLM (eval phase, checklist scoring) |
|
||||||
|
| `BRAVE_SEARCH_API_KEY` | optional; for the agent's web_search tool, `dummy` when unused |
|
||||||
|
|
||||||
|
Optional tuning: `REME_MODEL_NAME` (base model of the agent under test),
|
||||||
|
`REME_DIR`, `REME_LLM_BASE_URL` (default: DashScope OpenAI-compatible
|
||||||
|
endpoint).
|
||||||
|
|
||||||
|
3. **Link the evaluation data** (referenced in place, never copied):
|
||||||
|
```bash
|
||||||
|
ln -s "$PI_BENCH_ROOT/data" data
|
||||||
|
```
|
||||||
|
|
||||||
|
4. **(Optional) adjust model config** `config/models/reme.yaml`:
|
||||||
|
- `user_agent.model` / `judger.model`: model names for the simulated user
|
||||||
|
and the judger (literal values; π-Bench only expands `${ENV}` in
|
||||||
|
base_url/api_key).
|
||||||
|
- `run.turn_timeout`, `max_tool_iterations`, etc. as needed.
|
||||||
|
|
||||||
|
5. **Smoke check** (does not start the evaluation):
|
||||||
|
```bash
|
||||||
|
bash -n run_all.sh && bash -n run_persona.sh
|
||||||
|
source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. Run the evaluation
|
||||||
|
|
||||||
|
> ⚠️ For long runs use `screen`, **not nohup** (nohup loses the permission
|
||||||
|
> context in sandboxed/restricted environments and breaks child processes).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Full official run: wipe ALL personas' memory/outputs/traces first (default
|
||||||
|
# fresh mode, parallel=2)
|
||||||
|
mkdir -p logs # on a fresh deployment logs/ does not exist yet
|
||||||
|
screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
|
||||||
|
|
||||||
|
# Checkpoint continuation (after an interruption; no wipe, completed tasks skipped)
|
||||||
|
bash run_all.sh --resume
|
||||||
|
|
||||||
|
# Other usages
|
||||||
|
bash run_all.sh --parallel 1 # sequential
|
||||||
|
bash run_all.sh --resume --skip-eval # run phase only
|
||||||
|
bash run_persona.sh researcher # single persona (default --resume semantics)
|
||||||
|
bash run_persona.sh researcher --fresh
|
||||||
|
```
|
||||||
|
|
||||||
|
Time reference: 5 personas × 20 tasks, parallel=2, fresh full run ≈ 12–14 hours.
|
||||||
|
|
||||||
|
`run_all.sh` exits non-zero when any persona fails, so upstream automation
|
||||||
|
cannot mistake a partially failed suite run for a success.
|
||||||
|
|
||||||
|
## 6. Port allocation (parallel personas never collide)
|
||||||
|
|
||||||
|
| persona | AppWorld API | AppWorld MCP | Test Server | ReMe internal service |
|
||||||
|
|-------------|------|-------|------|-------|
|
||||||
|
| marketer | 9001 | 10001 | 9998 | 18766 |
|
||||||
|
| law_trainee | 9002 | 10002 | 9997 | 18767 |
|
||||||
|
| pharmacist | 9003 | 10003 | 9996 | 18768 |
|
||||||
|
| researcher | 9004 | 10004 | 9995 | 18765 |
|
||||||
|
| Financier | 9005 | 10005 | 9994 | 18769 |
|
||||||
|
|
||||||
|
## 7. Outputs and scores
|
||||||
|
|
||||||
|
- **Results**: `outputs/reme/{persona}/{task}/eval/results/*_result.json`
|
||||||
|
- `overall_average_score`: checklist completeness (COMP; the judger scores
|
||||||
|
each criterion YES/NO, weighted across dependency groups)
|
||||||
|
- `overall_proactiveness_average_score`: proactiveness (PROC; the
|
||||||
|
user_agent judges hidden-intent coverage during the run phase; each task
|
||||||
|
file also carries the global average)
|
||||||
|
- **Traces**: `~/.nanobot/trace_logs/reme/{persona}/{task}/...` (the scoring
|
||||||
|
input of the eval phase)
|
||||||
|
- **Logs**: `logs/` (`suite_<persona>.log` per persona; `bridge_*`,
|
||||||
|
`runner_run/eval_*`, `appworld_*`, `test_server_*` per service)
|
||||||
|
- **Memory store**: `reme_workspace/{persona}/` (daily/digest notes, raw
|
||||||
|
session dialogs, BM25 index, etc.; persistent across runs, wiped only in
|
||||||
|
fresh mode)
|
||||||
|
|
||||||
|
Score summary:
|
||||||
|
```bash
|
||||||
|
grep -h "overall_average_score\|overall_proactiveness" \
|
||||||
|
outputs/reme/*/*/eval/results/*_result.json | head
|
||||||
|
```
|
||||||
|
|
||||||
|
### Tool-trace capture (tools_evaluation support)
|
||||||
|
|
||||||
|
Some tasks define `objectives.tools_evaluation_path`: Python scripts that
|
||||||
|
score tool behavior (e.g. "the temporary Todoist board was created and
|
||||||
|
removed"). They need the executed tool calls in the trace. The pipeline:
|
||||||
|
|
||||||
|
1. During `reply()`, the bridge reads the persisted AgentScope session state
|
||||||
|
after each turn and extracts the new `tool_call` / `tool_result` blocks
|
||||||
|
(tool name, arguments, result).
|
||||||
|
2. Records are appended to
|
||||||
|
`outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`, tagged with the
|
||||||
|
turn number; AgentScope MCP names (`mcp__AppWorld__<tool>`) are normalized
|
||||||
|
to the π-Bench convention (`mcp_appworld_<tool>`).
|
||||||
|
3. `fix_trace_logs.py` pairs each `{ts}-messages.jsonl` run with the
|
||||||
|
temporally closest tools sidecar and merges the records into the generated
|
||||||
|
`turn_N.json` files under the `tool_steps` key — one of the two
|
||||||
|
tool-history formats understood by π-Bench's `collect_tool_history()`.
|
||||||
|
4. The eval phase then feeds `tool_steps` to both the tools_evaluation
|
||||||
|
scripts and the rendered `<tool_trace_extracts>` seen by the judger.
|
||||||
|
|
||||||
|
## 8. Memory mechanism (core design of this suite)
|
||||||
|
|
||||||
|
- **Persona isolation**: each persona has its own workspace
|
||||||
|
(`reme_workspace/{persona}/`); the bridge takes an exclusive
|
||||||
|
`.bridge.lock` on it at startup, so two bridges can never share one memory
|
||||||
|
store, and one persona's memory search can never reach another's memories.
|
||||||
|
- **Writes**: on task end (runner sends reset), the session is distilled by
|
||||||
|
the `auto_memory` job into daily notes and indexed by the background
|
||||||
|
watcher (BM25). Saves are non-blocking background tasks; the first message
|
||||||
|
of a new session waits for in-flight writes before searching.
|
||||||
|
- **Reads**: on every incoming user message the bridge runs one `search` and
|
||||||
|
injects matched memories (`[Relevant memories from previous sessions]`
|
||||||
|
prefix); without matches the message passes through unchanged. Retrieval
|
||||||
|
tuning (bridge CLI flags, adjustable in run_persona.sh):
|
||||||
|
- `--search-limit 3`: at most 3 memory chunks injected per message;
|
||||||
|
- `--search-min-score 2.0`: weak BM25 hits are filtered out;
|
||||||
|
- `tool_context_id` rotates per task: chunks already injected within the
|
||||||
|
same task are not re-injected (ReMe's seen-chunk dedup, 24h TTL); normal
|
||||||
|
recall resumes after task boundaries.
|
||||||
|
- **No self-leakage**: the in-progress session is not in the store yet
|
||||||
|
(saves happen on reset), so a task can never retrieve its own unfinished
|
||||||
|
content.
|
||||||
|
- The agent also holds `search`/`daily_write` tools and can retrieve/record
|
||||||
|
proactively.
|
||||||
|
- **System prompt**: `bridge_reme.py:build_system_prompt()` embeds the
|
||||||
|
HIDDEN-NEEDS protocol (proactiveness-oriented) and injects the persona
|
||||||
|
profile from `data/{persona}/profile.yaml` into every turn's system prompt.
|
||||||
|
|
||||||
|
## 9. Checkpoint resume and memory-cleanup semantics
|
||||||
|
|
||||||
|
- **Completion detection** (resume.py): scans
|
||||||
|
`outputs/reme/{persona}/**/history/*-log.jsonl` and
|
||||||
|
`outputs/reme/{persona}/run/*-log.jsonl` for
|
||||||
|
`Task finished task_id=X status=Y`. The status with the **newest event
|
||||||
|
timestamp** wins per task (record `timestamp`, falling back to
|
||||||
|
`timestamp_iso`, then to the timestamp embedded in the log file name) —
|
||||||
|
file category and read order alone can never override a newer record, so an
|
||||||
|
old run-level SUCCESS cannot mask a newer per-task ERROR. `SUCCESS /
|
||||||
|
MAX_TURNS / TIMEOUT` count as completed; `ERROR` and never-started tasks
|
||||||
|
are re-run (passed to the runner as repeated `--task-id` flags in episode
|
||||||
|
order).
|
||||||
|
- **Answer-leak prevention**: an interrupted task may already have been
|
||||||
|
distilled into daily notes during graceful shutdown; re-running it with
|
||||||
|
that memory injected would inflate scores. Before resuming,
|
||||||
|
`resume.py cleanup` therefore removes residual memory **only for tasks
|
||||||
|
about to be re-run** (daily/digest notes, session/dialog, mem_session;
|
||||||
|
matched via `session_id = pibench_{task}_*`). Completed tasks' memories are
|
||||||
|
never touched. Daily index files are refreshed **only for the dates that
|
||||||
|
lost notes**, by full workspace-relative wikilink path — and when the ReMe
|
||||||
|
package is importable, the refresh reuses ReMe's own daily-index rebuild
|
||||||
|
logic (`refresh_day_index`), so same-named notes on other dates are never
|
||||||
|
modified.
|
||||||
|
- **fresh vs resume are mutually exclusive**: a full memory wipe belongs to
|
||||||
|
fresh mode only (`run_all.sh` default, executed before any service starts);
|
||||||
|
resume never wipes.
|
||||||
|
|
||||||
|
## 10. Customization entry points
|
||||||
|
|
||||||
|
| Goal | Location |
|
||||||
|
|---|---|
|
||||||
|
| Base model of the agent under test | `REME_MODEL_NAME` in `env.sh` |
|
||||||
|
| user_agent / judger models | `config/models/reme.yaml` |
|
||||||
|
| Agent system prompt | `bridge_reme.py` `build_system_prompt()` |
|
||||||
|
| Memory retrieval limit/threshold | `--search-limit/--search-min-score` on the bridge command in `run_persona.sh` |
|
||||||
|
| ReMe internal parameters | **Do not modify ReMe source**; extend the built-in `benchmark` config and override via `resolve_app_config(config=...)` (see bridge `_init_reme_app`) |
|
||||||
|
| Turn timeout / tool iteration cap | `config/models/reme.yaml` `run.turn_timeout`, `model.max_tool_iterations` |
|
||||||
|
|
||||||
|
## 11. Troubleshooting
|
||||||
|
|
||||||
|
- **Port already in use**: the scripts auto-kill residual processes on the
|
||||||
|
four port groups above; if another suite (e.g. a different π-Bench
|
||||||
|
experiment) holds them, stop it first or change the port table in
|
||||||
|
run_persona.sh.
|
||||||
|
- **Bridge exits immediately with workspace locked**: another bridge already
|
||||||
|
holds the same workspace; make sure each persona uses its own
|
||||||
|
`--workspace-dir` (the scripts allocate one per persona).
|
||||||
|
- **Runner reports `${USER_API_KEY} ... empty`**: env.sh is unfilled or not
|
||||||
|
sourced; run_persona.sh sources env.sh automatically — when running the
|
||||||
|
runner manually, `source env.sh` first.
|
||||||
|
- **`Cannot import 'reme'`**: the bridge must run with
|
||||||
|
`${REME_DIR}/.venv/bin/python` (run_persona.sh already does); otherwise
|
||||||
|
check that `REME_DIR` points at the ReMe repository root.
|
||||||
|
- **AppWorld fails to start**: run `bash scripts/setup_appworld.sh` in the
|
||||||
|
π-Bench repo first (downloads data); inspect
|
||||||
|
`logs/appworld_*_<persona>.log`.
|
||||||
|
- **trace_history.yaml not found**: the runner needs
|
||||||
|
`config/bench/evaluation/trace_history.yaml`; this suite ships the file and
|
||||||
|
passes it explicitly via `--history-config-path`, and run_persona.sh fails
|
||||||
|
fast with a clear error if it is missing. Always launch run_persona.sh /
|
||||||
|
run_all.sh from the suite directory.
|
||||||
|
|
||||||
|
## 12. Privacy and security
|
||||||
|
|
||||||
|
- The suite code and config templates contain **no real API keys, user names
|
||||||
|
or absolute paths**; real keys live only in your local `env.sh`
|
||||||
|
(git-ignored).
|
||||||
|
- `logs/`, `outputs/`, `reme_workspace/` and `nanobot_workspace/` contain
|
||||||
|
full conversations and model outputs; never commit or share them.
|
||||||
|
- The `data` symlink points at the official π-Bench evaluation data; respect
|
||||||
|
its data license terms.
|
||||||
284
benchmark/pibench/README_ZH.md
Normal file
284
benchmark/pibench/README_ZH.md
Normal file
|
|
@ -0,0 +1,284 @@
|
||||||
|
# π-Bench 评测说明
|
||||||
|
|
||||||
|
[English version](./README.md)
|
||||||
|
|
||||||
|
将 **ReMe agent(带持久记忆)** 接入 **π-Bench**(Proactive Personal Assistant
|
||||||
|
Benchmark)的胶水层评测套件。只含对接所需的最小代码与配置;π-Bench 框架
|
||||||
|
(`src/`)、评测数据(`data/`)、AppWorld 工具环境、ReMe 本体均为**外部第三方
|
||||||
|
依赖**,通过符号链接与环境变量原位引用,不随本套件分发。
|
||||||
|
|
||||||
|
- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench (arXiv: 2605.14678)
|
||||||
|
- ReMe: 你所在 ReMe 仓库的根目录(本套件推荐放在 `ReMe/benchmark/pibench/`)
|
||||||
|
|
||||||
|
## 1. 架构总览
|
||||||
|
|
||||||
|
```
|
||||||
|
π-Bench runner (src.main --mode run)
|
||||||
|
│ user_agent(模拟用户 LLM)按 data/{persona}/episode.yaml 顺序
|
||||||
|
│ 逐任务、多轮地与 agent 对话,并在 run 阶段判定隐藏意图(PROC)
|
||||||
|
▼
|
||||||
|
test server (π-Bench scripts/test_server.py, HTTP 长轮询)
|
||||||
|
▲ /send │ /poll
|
||||||
|
│ ▼
|
||||||
|
bridge_reme.py ──────────────► ReMe Application(以库方式内嵌启动)
|
||||||
|
│ ├─ agent_wrapper: 被测 agent(AgentScope)
|
||||||
|
│ ├─ jobs: search / auto_memory / daily_write
|
||||||
|
│ └─ workspace: reme_workspace/{persona}/
|
||||||
|
│ (每 persona 独立持久记忆库,互不可见)
|
||||||
|
└──── MCP ────► AppWorld MCP ────► AppWorld API(工具/应用环境)
|
||||||
|
|
||||||
|
π-Bench runner (src.main --mode eval)
|
||||||
|
judger(裁判 LLM)读取 trace,按 checklist 逐条 YES/NO 打分(COMP)
|
||||||
|
```
|
||||||
|
|
||||||
|
要点:
|
||||||
|
- bridge 用 **ReMe 自己的 venv python** 运行,把 ReMe 当库用(`resolve_app_config`
|
||||||
|
+ `Application`),**ReMe 源码零改动**。
|
||||||
|
- 每条用户消息都会自动触发一次 ReMe memory `search` 并把命中记忆注入当前消息
|
||||||
|
(参数见 §8);任务结束(reset)时会话被 `auto_memory` 提炼为 daily 笔记落盘。
|
||||||
|
- agent 执行的每一轮工具调用(AppWorld MCP + ReMe job 工具)都会被采集并以
|
||||||
|
`tool_steps` 形式写入 trace,供 π-Bench 的 `tools_evaluation_path` 脚本
|
||||||
|
对工具行为评分(§7)。
|
||||||
|
- π-Bench 的 `data/`、`src/`、AppWorld 均不属于本套件,需先装好 π-Bench(§3.1)。
|
||||||
|
|
||||||
|
## 2. 目录结构
|
||||||
|
|
||||||
|
```
|
||||||
|
pibench/
|
||||||
|
├── README.md / README_ZH.md # 本文档(英文 / 中文)
|
||||||
|
├── env.sh.example # 环境配置模板(复制为 env.sh 后填写 TODO 项)
|
||||||
|
├── bridge_reme.py # ReMe ↔ test server 桥接(记忆注入/保存、
|
||||||
|
│ # profile 注入、工具调用轨迹采集)
|
||||||
|
├── run_persona.sh # 单 persona 全流程(5 个服务 + run + eval)
|
||||||
|
├── run_all.sh # 5 个 persona 批跑(fresh/resume,默认 2 并行)
|
||||||
|
├── resume.py # 断点续跑:完成判定 + 中断任务残留记忆的外科清理
|
||||||
|
├── fix_trace_logs.py # run 输出 → ~/.nanobot/trace_logs 转换,
|
||||||
|
│ # 并把工具轨迹合并进 turn 文件(eval 前置)
|
||||||
|
├── .gitignore # 排除 env.sh 与全部运行产物
|
||||||
|
└── config/
|
||||||
|
├── models/reme.yaml # runner 模型配置(model_id=reme)
|
||||||
|
└── bench/evaluation/trace_history.yaml # trace 渲染策略(随套件提供,
|
||||||
|
# 经 --history-config-path 显式传入)
|
||||||
|
```
|
||||||
|
|
||||||
|
运行时自动生成(均被 .gitignore 排除):`data`(符号链接)、`logs/`、
|
||||||
|
`outputs/`、`reme_workspace/`、`nanobot_workspace/`。
|
||||||
|
|
||||||
|
## 3. 前置依赖(第三方,先装好)
|
||||||
|
|
||||||
|
### 3.1 π-Bench 仓库(含 AppWorld)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://github.com/Simplified-Reasoning/Pi-Bench.git <pi-bench-dir>
|
||||||
|
cd <pi-bench-dir>
|
||||||
|
python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
|
||||||
|
source .venv/bin/activate
|
||||||
|
pip install -e . # pibench runner(src.main)
|
||||||
|
bash scripts/setup_appworld.sh # 安装 AppWorld 并下载其数据(体积较大,需网络)
|
||||||
|
```
|
||||||
|
|
||||||
|
装完自检:
|
||||||
|
```bash
|
||||||
|
ls data/ # 应含 researcher marketer pharmacist law_trainee Financier
|
||||||
|
.venv/bin/python -c "import src" && echo OK
|
||||||
|
.venv/bin/appworld --help >/dev/null && echo OK
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.2 ReMe 仓库
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd <reme-dir> # ReMe 仓库根目录(含 reme/ 包)
|
||||||
|
python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
|
||||||
|
source .venv/bin/activate
|
||||||
|
pip install -e . # 或按 ReMe 自身安装方式,保证 `import reme` 可用
|
||||||
|
```
|
||||||
|
|
||||||
|
自检:`.venv/bin/python -c "import reme; print('ok')"`
|
||||||
|
|
||||||
|
## 4. 安装本套件(逐步)
|
||||||
|
|
||||||
|
1. **放置套件**(推荐放进 ReMe 仓库,`REME_DIR` 可自动推断):
|
||||||
|
```bash
|
||||||
|
cp -r pibench <reme-dir>/benchmark/pibench
|
||||||
|
cd <reme-dir>/benchmark/pibench
|
||||||
|
```
|
||||||
|
若放在其他位置,稍后在 env.sh 中显式设置 `REME_DIR`。
|
||||||
|
|
||||||
|
2. **创建环境文件并填写自定义参数**:
|
||||||
|
```bash
|
||||||
|
cp env.sh.example env.sh
|
||||||
|
```
|
||||||
|
打开 `env.sh`,必填项(标 TODO 的):
|
||||||
|
| 变量 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `PI_BENCH_ROOT` | π-Bench 仓库根目录(含 `src/` `data/` `.venv` `third_party/appworld`) |
|
||||||
|
| `USER_API_KEY` | 模拟用户 LLM 的 API key(run 阶段判定隐藏意图) |
|
||||||
|
| `JUDGER_API_KEY` | 裁判 LLM 的 API key(eval 阶段 checklist 打分) |
|
||||||
|
| `BRAVE_SEARCH_API_KEY` | 可选;agent 的 web_search 工具用,不用填 `dummy` |
|
||||||
|
|
||||||
|
可选调整:`REME_MODEL_NAME`(被测 agent 基模)、`REME_DIR`、
|
||||||
|
`REME_LLM_BASE_URL`(默认 DashScope OpenAI 兼容端点)。
|
||||||
|
|
||||||
|
3. **链接评测数据**(π-Bench 数据原位引用,不复制):
|
||||||
|
```bash
|
||||||
|
ln -s "$PI_BENCH_ROOT/data" data
|
||||||
|
```
|
||||||
|
|
||||||
|
4. **(可选)调整模型配置** `config/models/reme.yaml`:
|
||||||
|
- `user_agent.model` / `judger.model`:模拟用户与裁判的模型名(字面量,
|
||||||
|
π-Bench 仅对 base_url/api_key 做 `${ENV}` 展开)。
|
||||||
|
- `run.turn_timeout`、`max_tool_iterations` 等按需。
|
||||||
|
|
||||||
|
5. **冒烟自检**(不启动评测):
|
||||||
|
```bash
|
||||||
|
bash -n run_all.sh && bash -n run_persona.sh
|
||||||
|
source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. 运行评测
|
||||||
|
|
||||||
|
> ⚠️ 长时间运行请放进 `screen`,**不要用 nohup**(nohup 在沙箱/受限环境下
|
||||||
|
> 会丢失权限上下文导致子进程异常)。
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 完整正式评测:先清空全部 persona 的记忆/输出/trace,再从头跑(默认 fresh,2 并行)
|
||||||
|
mkdir -p logs # 全新部署时 logs/ 尚不存在,先建再重定向
|
||||||
|
screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
|
||||||
|
|
||||||
|
# 断点续跑(中断后继续;不清记忆,跳过已完成任务)
|
||||||
|
bash run_all.sh --resume
|
||||||
|
|
||||||
|
# 其他用法
|
||||||
|
bash run_all.sh --parallel 1 # 串行
|
||||||
|
bash run_all.sh --resume --skip-eval # 只跑 run 阶段
|
||||||
|
bash run_persona.sh researcher # 单 persona(默认 --resume 语义)
|
||||||
|
bash run_persona.sh researcher --fresh
|
||||||
|
```
|
||||||
|
|
||||||
|
耗时参考:5 persona × 20 任务、2 并行,fresh 全量约 12–14 小时。
|
||||||
|
|
||||||
|
任一 persona 失败时 `run_all.sh` 以非零状态退出,上层自动化不会把部分失败
|
||||||
|
的评测误判为成功。
|
||||||
|
|
||||||
|
## 6. 端口分配(多 persona 并行互不冲突)
|
||||||
|
|
||||||
|
| persona | AppWorld API | AppWorld MCP | Test Server | ReMe 内部服务 |
|
||||||
|
|-------------|------|-------|------|-------|
|
||||||
|
| marketer | 9001 | 10001 | 9998 | 18766 |
|
||||||
|
| law_trainee | 9002 | 10002 | 9997 | 18767 |
|
||||||
|
| pharmacist | 9003 | 10003 | 9996 | 18768 |
|
||||||
|
| researcher | 9004 | 10004 | 9995 | 18765 |
|
||||||
|
| Financier | 9005 | 10005 | 9994 | 18769 |
|
||||||
|
|
||||||
|
## 7. 输出与分数
|
||||||
|
|
||||||
|
- **结果**:`outputs/reme/{persona}/{task}/eval/results/*_result.json`
|
||||||
|
- `overall_average_score`:checklist 完整度(COMP,judger 逐条 YES/NO 按依赖组加权)
|
||||||
|
- `overall_proactiveness_average_score`:主动性(PROC,run 阶段 user_agent
|
||||||
|
判定隐藏意图覆盖率;每个任务文件同时携带全局均值)
|
||||||
|
- **trace**:`~/.nanobot/trace_logs/reme/{persona}/{task}/...`(eval 的判分输入)
|
||||||
|
- **日志**:`logs/`(`suite_<persona>.log` 为每 persona 总日志,`bridge_*`、
|
||||||
|
`runner_run/eval_*`、`appworld_*`、`test_server_*` 分服务)
|
||||||
|
- **记忆库**:`reme_workspace/{persona}/`(daily/digest 笔记、session 原始对话、
|
||||||
|
BM25 索引等;跨运行持久,fresh 才清空)
|
||||||
|
|
||||||
|
查看汇总:
|
||||||
|
```bash
|
||||||
|
grep -h "overall_average_score\|overall_proactiveness" \
|
||||||
|
outputs/reme/*/*/eval/results/*_result.json | head
|
||||||
|
```
|
||||||
|
|
||||||
|
### 工具轨迹采集(tools_evaluation 支持)
|
||||||
|
|
||||||
|
部分任务定义了 `objectives.tools_evaluation_path`:用 Python 脚本对工具行为
|
||||||
|
打分(例如"临时 Todoist 看板已创建并被删除")。这些脚本需要 trace 里有真实
|
||||||
|
的工具调用记录。采集链路:
|
||||||
|
|
||||||
|
1. 每轮 `reply()` 之后,bridge 读取 AgentScope 落盘的会话状态,提取本轮新增
|
||||||
|
的 `tool_call` / `tool_result` 块(工具名、参数、结果)。
|
||||||
|
2. 记录按 turn 编号追加写入
|
||||||
|
`outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`;AgentScope 的
|
||||||
|
MCP 工具名(`mcp__AppWorld__<tool>`)会规范化为 π-Bench 约定
|
||||||
|
(`mcp_appworld_<tool>`)。
|
||||||
|
3. `fix_trace_logs.py` 将每个 `{ts}-messages.jsonl` 运行与时间上最接近的
|
||||||
|
tools 旁路文件配对,把记录合并进生成的 `turn_N.json` 的 `tool_steps`
|
||||||
|
字段——这是 π-Bench `collect_tool_history()` 支持的两种工具轨迹格式之一。
|
||||||
|
4. eval 阶段 `tool_steps` 既提供给 tools_evaluation 脚本,也会被渲染为
|
||||||
|
judger 可见的 `<tool_trace_extracts>`。
|
||||||
|
|
||||||
|
## 8. 记忆机制(本套件的核心设计)
|
||||||
|
|
||||||
|
- **persona 隔离**:每个 persona 独立 workspace(`reme_workspace/{persona}/`),
|
||||||
|
bridge 启动时对 workspace 加 `.bridge.lock` 排他锁,两个 bridge 不可能共用
|
||||||
|
同一记忆库;一个 persona 的 memory search 永远接触不到其他 persona 的记忆。
|
||||||
|
- **写入**:任务结束(runner 发送 reset)时,会话经 `auto_memory` job 提炼为
|
||||||
|
daily 笔记落盘,后台 watcher 建 BM25 索引。保存为非阻塞后台任务,
|
||||||
|
新会话首条消息会先等待在途写入完成再检索。
|
||||||
|
- **读取**:bridge 每收到一条用户消息自动 `search` 一次并注入命中记忆
|
||||||
|
(`[Relevant memories from previous sessions]` 前缀),无命中则原样透传。
|
||||||
|
检索参数(bridge 命令行,可在 run_persona.sh 中调整):
|
||||||
|
- `--search-limit 3`:每条消息最多注入 3 个记忆块;
|
||||||
|
- `--search-min-score 2.0`:过滤弱 BM25 命中;
|
||||||
|
- `tool_context_id` 按任务轮换:同一任务内已注入的记忆块不重复注入
|
||||||
|
(ReMe 自带 seen-chunk 去重,24h TTL),任务边界后恢复正常召回。
|
||||||
|
- **无自泄漏**:进行中的会话尚未入库(save 发生在 reset),任务不会检索到
|
||||||
|
自己未完成的内容。
|
||||||
|
- agent 同时持有 `search`/`daily_write` 工具,可主动检索/记录。
|
||||||
|
- **system prompt**:`bridge_reme.py:build_system_prompt()` 内置
|
||||||
|
HIDDEN-NEEDS 协议(面向 proactiveness),并把 `data/{persona}/profile.yaml`
|
||||||
|
的 persona profile 注入每轮 system prompt。
|
||||||
|
|
||||||
|
## 9. 断点续跑与记忆清理语义
|
||||||
|
|
||||||
|
- **完成判定**(resume.py):扫描 `outputs/reme/{persona}/**/history/*-log.jsonl`
|
||||||
|
与 `outputs/reme/{persona}/run/*-log.jsonl` 中的
|
||||||
|
`Task finished task_id=X status=Y`。每个任务以**事件时间最新**的记录为准
|
||||||
|
(优先取记录的 `timestamp`,回退 `timestamp_iso`,再回退日志文件名中的
|
||||||
|
时间戳)——文件类别与读取顺序本身不能覆盖更新的记录,因此旧的 run 级
|
||||||
|
SUCCESS 不会掩盖更新的 per-task ERROR。`SUCCESS/MAX_TURNS/TIMEOUT` 记为
|
||||||
|
完成,`ERROR`/未开始的任务重跑(按 episode 顺序以 `--task-id` 传给 runner)。
|
||||||
|
- **防答案泄漏**:被中断的任务可能已在优雅退出时提炼成 daily 笔记,直接重跑会
|
||||||
|
把答案注入、抬高分数。因此 resume 启动前 `resume.py cleanup` **只删除待重跑
|
||||||
|
任务**的残留记忆(daily/digest 笔记、session/dialog、mem_session,按
|
||||||
|
`session_id = pibench_{task}_*` 匹配),已完成任务的记忆一律不动。daily
|
||||||
|
索引**只刷新实际发生删除的日期**,按完整的 workspace 相对 wikilink 路径
|
||||||
|
匹配;当 ReMe 包可导入时,刷新直接复用 ReMe 自带的 daily 索引重建逻辑
|
||||||
|
(`refresh_day_index`),不会误改其他日期下的同名笔记条目。
|
||||||
|
- **fresh vs resume 互斥**:全量清记忆只属于 fresh 模式(`run_all.sh` 默认,
|
||||||
|
在任何服务启动前执行);resume 永不清全量。
|
||||||
|
|
||||||
|
## 10. 自定义与调优入口
|
||||||
|
|
||||||
|
| 目标 | 位置 |
|
||||||
|
|---|---|
|
||||||
|
| 被测 agent 基模 | `env.sh` 的 `REME_MODEL_NAME` |
|
||||||
|
| user_agent / judger 模型 | `config/models/reme.yaml` |
|
||||||
|
| agent system prompt | `bridge_reme.py` `build_system_prompt()` |
|
||||||
|
| 记忆检索条数/阈值 | `run_persona.sh` bridge 启动命令的 `--search-limit/--search-min-score` |
|
||||||
|
| ReMe 内部参数 | **不要改 ReMe 源码**;继承内置 `benchmark` 配置,并经 `resolve_app_config(config=...)` 覆盖(见 bridge `_init_reme_app`) |
|
||||||
|
| 轮超时/工具迭代上限 | `config/models/reme.yaml` `run.turn_timeout`、`model.max_tool_iterations` |
|
||||||
|
|
||||||
|
## 11. 故障排查
|
||||||
|
|
||||||
|
- **端口被占用**:脚本会自动 kill 上述 4 组端口上的残留进程;若与其他套件
|
||||||
|
(如别的 π-Bench 实验)冲突,请先停掉对方或改 run_persona.sh 的端口表。
|
||||||
|
- **bridge 启动即退出,提示 workspace locked**:另一个 bridge 正占用同一
|
||||||
|
workspace;确认每个 persona 用各自的 `--workspace-dir`(脚本已按 persona 分配)。
|
||||||
|
- **runner 报 `${USER_API_KEY} ... empty`**:env.sh 未填写或未生效;
|
||||||
|
run_persona.sh 会自动 source env.sh,手动运行 runner 时请先 `source env.sh`。
|
||||||
|
- **`Cannot import 'reme'`**:bridge 必须用 `${REME_DIR}/.venv/bin/python` 运行
|
||||||
|
(run_persona.sh 已如此),或检查 `REME_DIR` 是否指向 ReMe 仓库根目录。
|
||||||
|
- **AppWorld 启动失败**:先在 π-Bench 仓库执行 `bash scripts/setup_appworld.sh`
|
||||||
|
下载数据;查看 `logs/appworld_*_<persona>.log`。
|
||||||
|
- **trace_history.yaml 找不到**:runner 需要
|
||||||
|
`config/bench/evaluation/trace_history.yaml`;本套件已随附该文件并通过
|
||||||
|
`--history-config-path` 显式传入,run_persona.sh 启动前会做存在性检查,
|
||||||
|
缺失时立即报出清晰错误。请始终从套件目录启动 run_persona.sh / run_all.sh。
|
||||||
|
|
||||||
|
## 12. 隐私与安全
|
||||||
|
|
||||||
|
- 套件代码与配置模板中**不含任何真实 API key、用户名或绝对路径**;
|
||||||
|
真实 key 只存在于你本地的 `env.sh`(已被 .gitignore 排除)。
|
||||||
|
- `logs/`、`outputs/`、`reme_workspace/`、`nanobot_workspace/` 含完整对话内容
|
||||||
|
与模型输出,请勿提交仓库或外传。
|
||||||
|
- `data` 符号链接指向 π-Bench 官方评测数据,请遵守其数据许可条款。
|
||||||
1039
benchmark/pibench/bridge_reme.py
Executable file
1039
benchmark/pibench/bridge_reme.py
Executable file
File diff suppressed because it is too large
Load diff
53
benchmark/pibench/config/bench/evaluation/trace_history.yaml
Normal file
53
benchmark/pibench/config/bench/evaluation/trace_history.yaml
Normal file
|
|
@ -0,0 +1,53 @@
|
||||||
|
version: 1
|
||||||
|
|
||||||
|
format:
|
||||||
|
root_tag: trace
|
||||||
|
turn_tag: turn
|
||||||
|
message_tag: message
|
||||||
|
file_tag: file
|
||||||
|
tool_call_tag_prefix: tool_call
|
||||||
|
tool_result_tag_prefix: tool_result
|
||||||
|
|
||||||
|
text_policy:
|
||||||
|
default:
|
||||||
|
truncate_chars: 1200
|
||||||
|
mask_newlines: false
|
||||||
|
field_overrides:
|
||||||
|
files_read:
|
||||||
|
truncate_chars: 40000
|
||||||
|
assistant_content:
|
||||||
|
truncate_chars: 40000
|
||||||
|
tool_result_content:
|
||||||
|
truncate_chars: 40000
|
||||||
|
|
||||||
|
fields:
|
||||||
|
turn:
|
||||||
|
include_session_key: false
|
||||||
|
|
||||||
|
files:
|
||||||
|
enabled: true
|
||||||
|
|
||||||
|
messages:
|
||||||
|
enabled: true
|
||||||
|
include_message_role_attr: true
|
||||||
|
include_message_index_attr: false
|
||||||
|
include_system: false
|
||||||
|
include_user: true
|
||||||
|
include_assistant_thinking_content: false
|
||||||
|
include_assistant_thinking_reasoning: false
|
||||||
|
include_assistant_content: true
|
||||||
|
include_assistant_reasoning: false
|
||||||
|
include_assistant_tool_calls: false
|
||||||
|
require_matching_tool_call: true
|
||||||
|
|
||||||
|
tool_calls:
|
||||||
|
include_tool_call_id: false
|
||||||
|
tools:
|
||||||
|
web_fetch:
|
||||||
|
enabled: true
|
||||||
|
include_tool_call_keys: [url]
|
||||||
|
include_tool_result: false
|
||||||
|
web_search:
|
||||||
|
enabled: true
|
||||||
|
include_tool_call_keys: [query]
|
||||||
|
include_tool_result: false
|
||||||
40
benchmark/pibench/config/models/reme.yaml
Normal file
40
benchmark/pibench/config/models/reme.yaml
Normal file
|
|
@ -0,0 +1,40 @@
|
||||||
|
# ReMe model configuration for Pi-Bench
|
||||||
|
# Uses ReMe's AgentScope agent with Dashscope as the LLM backend
|
||||||
|
|
||||||
|
model:
|
||||||
|
model: reme
|
||||||
|
base_url: "http://localhost:8088"
|
||||||
|
api_key: "dummy"
|
||||||
|
provider: custom
|
||||||
|
max_tokens: 16384
|
||||||
|
max_tool_iterations: 120
|
||||||
|
memory_window: 100
|
||||||
|
|
||||||
|
user_agent:
|
||||||
|
model: qwen3.8-max
|
||||||
|
base_url: "${USER_BASE_URL}"
|
||||||
|
api_key: "${USER_API_KEY}"
|
||||||
|
temperature: 0.0
|
||||||
|
request_timeout: 360.0
|
||||||
|
|
||||||
|
judger:
|
||||||
|
model: qwen3.8-max
|
||||||
|
base_url: "${JUDGER_BASE_URL}"
|
||||||
|
api_key: "${JUDGER_API_KEY}"
|
||||||
|
temperature: 0.0
|
||||||
|
request_timeout: 360.0
|
||||||
|
|
||||||
|
tools:
|
||||||
|
brave_search_api_key: "${BRAVE_SEARCH_API_KEY}"
|
||||||
|
web_search_max_results: 10
|
||||||
|
|
||||||
|
nanobot:
|
||||||
|
trace_logs_dir: "~/.nanobot/trace_logs"
|
||||||
|
workspace_dir: "~/.nanobot/workspace"
|
||||||
|
copy_task_assets_to_workspace: true
|
||||||
|
|
||||||
|
run:
|
||||||
|
output_dir: outputs
|
||||||
|
log_level: INFO
|
||||||
|
user_mode: llm
|
||||||
|
turn_timeout: 2400.0
|
||||||
57
benchmark/pibench/env.sh.example
Normal file
57
benchmark/pibench/env.sh.example
Normal file
|
|
@ -0,0 +1,57 @@
|
||||||
|
#!/bin/bash
|
||||||
|
# ═══════════════════════════════════════════════════════════════════════
|
||||||
|
# pibench evaluation suite - environment configuration template
|
||||||
|
# Usage: cp env.sh.example env.sh, then fill in the TODO items below.
|
||||||
|
# ⚠️ env.sh contains real API keys; never commit or share it
|
||||||
|
# (already excluded via .gitignore).
|
||||||
|
# ═══════════════════════════════════════════════════════════════════════
|
||||||
|
|
||||||
|
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
|
||||||
|
# ─── TODO: π-Bench repository root ────────────────────────────────────
|
||||||
|
# Must contain src/, data/, scripts/test_server.py, third_party/appworld
|
||||||
|
# and .venv (see README setup).
|
||||||
|
export PI_BENCH_ROOT=""
|
||||||
|
|
||||||
|
# ─── ReMe repository ──────────────────────────────────────────────────
|
||||||
|
# Defaults to two levels above this directory (the layout this suite uses
|
||||||
|
# when placed at ReMe/benchmark/pibench); point it at the actual ReMe
|
||||||
|
# repository root if the suite lives elsewhere.
|
||||||
|
export REME_DIR="${REME_DIR:-$(cd "${SUITE_DIR}/../.." && pwd)}"
|
||||||
|
|
||||||
|
# ─── Base model of the agent under test (LLM used by the ReMe agent) ──
|
||||||
|
export REME_MODEL_NAME="${REME_MODEL_NAME:-qwen3.6-plus}"
|
||||||
|
|
||||||
|
# ─── LLM service endpoint (default: DashScope OpenAI-compatible; any
|
||||||
|
# OpenAI-compatible endpoint works) ────────────────────────────────
|
||||||
|
DASHSCOPE_BASE_URL="https://dashscope.aliyuncs.com/compatible-mode/v1"
|
||||||
|
export REME_LLM_BASE_URL="${REME_LLM_BASE_URL:-${DASHSCOPE_BASE_URL}}"
|
||||||
|
|
||||||
|
# ─── TODO: API keys ───────────────────────────────────────────────────
|
||||||
|
# USER_API_KEY : drives the simulated user LLM (run phase; judges whether
|
||||||
|
# hidden intents are satisfied and asks follow-ups)
|
||||||
|
# JUDGER_API_KEY: drives the judger LLM (eval phase; scores the checklist)
|
||||||
|
# The two may be identical; one strong model is recommended for both.
|
||||||
|
export USER_BASE_URL="${DASHSCOPE_BASE_URL}"
|
||||||
|
export USER_API_KEY="TODO-fill-in-user-agent-api-key"
|
||||||
|
|
||||||
|
export JUDGER_BASE_URL="${DASHSCOPE_BASE_URL}"
|
||||||
|
export JUDGER_API_KEY="TODO-fill-in-judger-api-key"
|
||||||
|
|
||||||
|
# The ReMe agent's key reuses USER_API_KEY by default (no need to repeat
|
||||||
|
# it when both use the same service and key).
|
||||||
|
export REME_LLM_API_KEY="${REME_LLM_API_KEY:-${USER_API_KEY}}"
|
||||||
|
|
||||||
|
# Brave Search (optional; used by the agent's web_search tool - use
|
||||||
|
# "dummy" when not needed).
|
||||||
|
export BRAVE_SEARCH_API_KEY="TODO-optional-brave-search-key-or-dummy"
|
||||||
|
|
||||||
|
# ─── Persistent memory workspaces (one subdirectory per persona,
|
||||||
|
# created automatically) ───────────────────────────────────────────
|
||||||
|
export REME_WORKSPACE_ROOT="${REME_WORKSPACE_ROOT:-${SUITE_DIR}/reme_workspace}"
|
||||||
|
|
||||||
|
# ─── Variables consumed by ReMe's default.yaml model config expansion;
|
||||||
|
# do not remove ────────────────────────────────────────────────────
|
||||||
|
export LLM_MODEL_NAME="${REME_MODEL_NAME}"
|
||||||
|
export LLM_BASE_URL="${REME_LLM_BASE_URL}"
|
||||||
|
export LLM_API_KEY="${REME_LLM_API_KEY}"
|
||||||
198
benchmark/pibench/fix_trace_logs.py
Executable file
198
benchmark/pibench/fix_trace_logs.py
Executable file
|
|
@ -0,0 +1,198 @@
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Convert reme_eval run outputs into eval-compatible trace logs.
|
||||||
|
|
||||||
|
outputs/{model_id}/{user_id}/{task_id}/history/{ts}-messages.jsonl
|
||||||
|
-> ~/.nanobot/trace_logs/{model_id}/{user_id}/{task_id}/{ts}/turn_N.json
|
||||||
|
|
||||||
|
The bridge additionally writes {ts}-tools.jsonl sidecar files next to the
|
||||||
|
message histories: one JSON object per executed tool call with fields
|
||||||
|
{turn, name, arguments, result}. Each messages run is paired with the
|
||||||
|
temporally closest sidecar, and the records are merged into the generated
|
||||||
|
turn files under the "tool_steps" key, which is one of the tool-history
|
||||||
|
formats π-Bench's collect_tool_history() understands. Without this step,
|
||||||
|
tools_evaluation scripts would see no tool evidence at all.
|
||||||
|
|
||||||
|
Usage: python fix_trace_logs.py [user_id ...] (no args = all users)
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
SUITE_DIR = Path(__file__).resolve().parent
|
||||||
|
OUTPUTS_DIR = SUITE_DIR / "outputs"
|
||||||
|
TRACE_LOGS_DIR = Path.home() / ".nanobot" / "trace_logs"
|
||||||
|
|
||||||
|
MESSAGES_FILE_RE = re.compile(r"^(\d{8}_\d{6})-messages\.jsonl$")
|
||||||
|
TOOLS_FILE_RE = re.compile(r"^(\d{8}_\d{6})-tools\.jsonl$")
|
||||||
|
TIME_FORMAT = "%Y%m%d_%H%M%S"
|
||||||
|
# A tool sidecar belongs to the messages run that started at most this many
|
||||||
|
# seconds earlier (the bridge stamps the sidecar when the task's first user
|
||||||
|
# message arrives, shortly after the runner opened the messages file).
|
||||||
|
MAX_PAIR_DELTA_SECONDS = 6 * 3600
|
||||||
|
|
||||||
|
|
||||||
|
def _to_epoch(timestamp: str) -> float:
|
||||||
|
"""Parse a YYYYMMDD_HHMMSS timestamp into epoch seconds."""
|
||||||
|
try:
|
||||||
|
return datetime.strptime(timestamp, TIME_FORMAT).timestamp()
|
||||||
|
except ValueError:
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
|
||||||
|
def load_tool_records(tools_file: Path) -> dict:
|
||||||
|
"""Group sidecar tool records by turn number."""
|
||||||
|
by_turn: dict = {}
|
||||||
|
try:
|
||||||
|
with open(tools_file, "r", encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
record = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
if not isinstance(record, dict) or not record.get("name"):
|
||||||
|
continue
|
||||||
|
turn = int(record.get("turn") or 0)
|
||||||
|
by_turn.setdefault(turn, []).append(
|
||||||
|
{
|
||||||
|
"name": record["name"],
|
||||||
|
"arguments": record.get("arguments", {}),
|
||||||
|
"result": record.get("result", ""),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
except OSError as exc:
|
||||||
|
print(f" WARNING: cannot read tool sidecar {tools_file}: {exc}")
|
||||||
|
return by_turn
|
||||||
|
|
||||||
|
|
||||||
|
def pair_tool_sidecars(message_runs: list, tool_runs: list) -> dict:
|
||||||
|
"""Pair each messages run with the temporally closest unused tool sidecar.
|
||||||
|
|
||||||
|
Fresh runs produce exactly one messages file and one sidecar per task;
|
||||||
|
re-runs append matching pairs, so sorted greedy nearest-timestamp
|
||||||
|
matching is stable. Sidecars farther away than MAX_PAIR_DELTA_SECONDS
|
||||||
|
(e.g. leftovers of a crashed bridge) stay unpaired.
|
||||||
|
"""
|
||||||
|
pairing: dict = {}
|
||||||
|
unused = list(tool_runs)
|
||||||
|
for msg_ts, _ in message_runs:
|
||||||
|
best_delta = None
|
||||||
|
best_item = None
|
||||||
|
for tool_ts, tool_path in unused:
|
||||||
|
delta = abs(_to_epoch(tool_ts) - _to_epoch(msg_ts))
|
||||||
|
if best_delta is None or delta < best_delta:
|
||||||
|
best_delta = delta
|
||||||
|
best_item = (tool_ts, tool_path)
|
||||||
|
if best_delta is not None and best_item is not None and best_delta <= MAX_PAIR_DELTA_SECONDS:
|
||||||
|
pairing[msg_ts] = best_item[1]
|
||||||
|
unused.remove(best_item)
|
||||||
|
return pairing
|
||||||
|
|
||||||
|
|
||||||
|
def build_turns(messages: list) -> list:
|
||||||
|
"""Split the flat message list into per-turn [user, assistant] groups."""
|
||||||
|
turns = []
|
||||||
|
i = 0
|
||||||
|
while i < len(messages):
|
||||||
|
turn_msgs = []
|
||||||
|
if messages[i]["role"] == "user":
|
||||||
|
turn_msgs.append({"role": "user", "content": messages[i]["message"]})
|
||||||
|
i += 1
|
||||||
|
if i < len(messages) and messages[i]["role"] == "assistant":
|
||||||
|
turn_msgs.append({"role": "assistant", "content": messages[i]["message"]})
|
||||||
|
i += 1
|
||||||
|
if not turn_msgs:
|
||||||
|
i += 1 # defensive: never spin on unexpected roles
|
||||||
|
continue
|
||||||
|
turns.append(turn_msgs)
|
||||||
|
return turns
|
||||||
|
|
||||||
|
|
||||||
|
def convert_task(model_id: str, user_id: str, task_dir: Path) -> None:
|
||||||
|
"""Convert one task's history dir into trace turn files with tool_steps."""
|
||||||
|
history_dir = task_dir / "history"
|
||||||
|
if not history_dir.is_dir():
|
||||||
|
return
|
||||||
|
|
||||||
|
message_runs = []
|
||||||
|
tool_runs = []
|
||||||
|
for msg_file in history_dir.glob("*-messages.jsonl"):
|
||||||
|
match = MESSAGES_FILE_RE.match(msg_file.name)
|
||||||
|
if match:
|
||||||
|
message_runs.append((match.group(1), msg_file))
|
||||||
|
for tools_file in history_dir.glob("*-tools.jsonl"):
|
||||||
|
match = TOOLS_FILE_RE.match(tools_file.name)
|
||||||
|
if match:
|
||||||
|
tool_runs.append((match.group(1), tools_file))
|
||||||
|
if not message_runs:
|
||||||
|
return
|
||||||
|
|
||||||
|
message_runs.sort(key=lambda item: item[0])
|
||||||
|
tool_runs.sort(key=lambda item: item[0])
|
||||||
|
pairing = pair_tool_sidecars(message_runs, tool_runs)
|
||||||
|
|
||||||
|
print(f"\n{model_id}/{user_id}/{task_dir.name}")
|
||||||
|
for timestamp, msg_file in message_runs:
|
||||||
|
trace_dir = TRACE_LOGS_DIR / model_id / user_id / task_dir.name / timestamp
|
||||||
|
trace_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
messages = []
|
||||||
|
with open(msg_file, "r", encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
msg = json.loads(line)
|
||||||
|
if msg.get("role") == "user" and msg.get("message") == "/new":
|
||||||
|
continue
|
||||||
|
messages.append(msg)
|
||||||
|
|
||||||
|
tools_file = pairing.get(timestamp)
|
||||||
|
tools_by_turn = load_tool_records(tools_file) if tools_file else {}
|
||||||
|
if tools_file is not None:
|
||||||
|
print(f" {timestamp}: paired tool sidecar {tools_file.name}")
|
||||||
|
|
||||||
|
turns = build_turns(messages)
|
||||||
|
for turn_idx, turn_msgs in enumerate(turns, start=1):
|
||||||
|
turn_data = {"messages": turn_msgs}
|
||||||
|
tool_steps = tools_by_turn.get(turn_idx)
|
||||||
|
if tool_steps:
|
||||||
|
turn_data["tool_steps"] = tool_steps
|
||||||
|
turn_file = trace_dir / f"turn_{turn_idx}.json"
|
||||||
|
with open(turn_file, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(turn_data, f, indent=2, ensure_ascii=False)
|
||||||
|
tool_total = sum(len(steps) for steps in tools_by_turn.values())
|
||||||
|
print(f" {timestamp}: {len(turns)} turns, {tool_total} tool step(s) -> {trace_dir}")
|
||||||
|
|
||||||
|
|
||||||
|
def convert_outputs(user_filter=None):
|
||||||
|
"""Convert message history JSONL files into per-turn trace JSON files."""
|
||||||
|
if not OUTPUTS_DIR.exists():
|
||||||
|
print(f"outputs dir not found: {OUTPUTS_DIR}")
|
||||||
|
return
|
||||||
|
|
||||||
|
for model_dir in sorted(OUTPUTS_DIR.iterdir()):
|
||||||
|
if not model_dir.is_dir():
|
||||||
|
continue
|
||||||
|
model_id = model_dir.name
|
||||||
|
|
||||||
|
for user_dir in sorted(model_dir.iterdir()):
|
||||||
|
if not user_dir.is_dir():
|
||||||
|
continue
|
||||||
|
user_id = user_dir.name
|
||||||
|
if user_filter and user_id not in user_filter:
|
||||||
|
continue
|
||||||
|
|
||||||
|
for task_dir in sorted(user_dir.iterdir()):
|
||||||
|
if task_dir.is_dir():
|
||||||
|
convert_task(model_id, user_id, task_dir)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
convert_outputs(set(sys.argv[1:]) or None)
|
||||||
|
print("\ndone")
|
||||||
332
benchmark/pibench/resume.py
Executable file
332
benchmark/pibench/resume.py
Executable file
|
|
@ -0,0 +1,332 @@
|
||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Checkpoint-resume support for the reme_eval suite.
|
||||||
|
|
||||||
|
Completion source of truth:
|
||||||
|
- outputs/reme/<persona>/<task_id>/history/*-log.jsonl (per-task logs,
|
||||||
|
flushed incrementally, survive mid-run kills)
|
||||||
|
- outputs/reme/<persona>/run/*-log.jsonl (run-level logs,
|
||||||
|
may be truncated if the process was killed before flush)
|
||||||
|
lines: "Task finished task_id=<id> status=<STATUS>"
|
||||||
|
A task counts as COMPLETED when its latest terminal status is one of
|
||||||
|
SUCCESS / MAX_TURNS / TIMEOUT. ERROR or never-started tasks stay pending.
|
||||||
|
|
||||||
|
"Latest" is decided by EVENT TIME, not by file category or read order:
|
||||||
|
each record's "timestamp" (epoch seconds, or "timestamp_iso" as fallback)
|
||||||
|
is compared across per-task and run-level logs alike, with the timestamp
|
||||||
|
embedded in the log file name as a last-resort fallback. This keeps an
|
||||||
|
old run-level SUCCESS from overriding a newer per-task ERROR when the
|
||||||
|
re-run died before the new run-level log captured the task.
|
||||||
|
|
||||||
|
Commands:
|
||||||
|
remaining <persona> [--json]
|
||||||
|
Print task_ids still to run, in data/<persona>/episode.yaml order
|
||||||
|
(one per line; --json prints {"completed": [...], "remaining": [...]}).
|
||||||
|
|
||||||
|
cleanup <persona> [--dry-run]
|
||||||
|
Surgically remove residual memory artifacts of tasks that are about
|
||||||
|
to be RE-RUN (i.e. pending tasks that left partial state because a
|
||||||
|
previous run was interrupted). This prevents answer leakage: an
|
||||||
|
interrupted task's conversation may already have been distilled into
|
||||||
|
daily notes during graceful shutdown, and re-running the task with
|
||||||
|
that memory injected would inflate scores.
|
||||||
|
|
||||||
|
Removed artifacts (only for pending tasks with residual state):
|
||||||
|
- daily/<date>/<note>.md whose frontmatter session_id matches
|
||||||
|
pibench_<task_id>_*, plus a refresh of ONLY the daily index of
|
||||||
|
the affected date(s) (daily/<date>.md), matched by the full
|
||||||
|
workspace-relative note path, never by bare file name
|
||||||
|
- digest notes with matching session_id
|
||||||
|
- session/dialog/pibench_<task_id>_*.jsonl
|
||||||
|
- mem_session/**.jsonl files containing pibench_<task_id>_
|
||||||
|
When the ReMe package is importable, the daily index refresh reuses
|
||||||
|
ReMe's own rebuild logic (reme.steps.file_io._daily_index.
|
||||||
|
refresh_day_index); otherwise index lines are dropped by exact
|
||||||
|
wikilink path match. Either way, indexes of other dates are never
|
||||||
|
touched. The ReMe watcher (init_changes_step) detects the deleted
|
||||||
|
daily notes on next bridge startup and removes them from the BM25
|
||||||
|
index itself.
|
||||||
|
|
||||||
|
Completed tasks' memories are NEVER touched by this command.
|
||||||
|
|
||||||
|
Design note (resume vs memory-wipe conflict):
|
||||||
|
A full memory wipe is a suite-level action of fresh mode (run_all.sh
|
||||||
|
without --resume) and happens before any service starts. Resume mode
|
||||||
|
never wipes; it only performs the surgical cleanup above. The two modes
|
||||||
|
are mutually exclusive, so a resumed run can never lose the cross-session
|
||||||
|
memory accumulated by completed tasks.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
try: # Reuse ReMe's daily-index rebuild when running inside the ReMe venv.
|
||||||
|
from reme.steps.file_io._daily_index import refresh_day_index
|
||||||
|
except ImportError: # pragma: no cover - depends on runtime venv
|
||||||
|
refresh_day_index = None
|
||||||
|
|
||||||
|
SUITE_DIR = Path(__file__).resolve().parent
|
||||||
|
DATA_DIR = Path(os.environ.get("REME_EVAL_DATA_DIR", SUITE_DIR / "data")).resolve()
|
||||||
|
OUTPUTS_DIR = Path(os.environ.get("REME_EVAL_OUTPUTS_DIR", SUITE_DIR / "outputs")) / "reme"
|
||||||
|
WORKSPACE_ROOT = Path(
|
||||||
|
os.environ.get("REME_WORKSPACE_ROOT", SUITE_DIR / "reme_workspace"),
|
||||||
|
).resolve()
|
||||||
|
|
||||||
|
COMPLETED_STATUSES = {"SUCCESS", "MAX_TURNS", "TIMEOUT"}
|
||||||
|
TASK_FINISHED_RE = re.compile(r"Task finished task_id=(\S+) status=(\S+)")
|
||||||
|
SESSION_ID_RE = re.compile(r"^session_id:\s*(\S+)", re.MULTILINE)
|
||||||
|
NOTE_COUNT_RE = re.compile(r"(description:\s*)\d+(\s*note\(s\) today)")
|
||||||
|
LOG_FILE_TS_RE = re.compile(r"^(\d{8}_\d{6})-log\.jsonl$")
|
||||||
|
TIME_FORMAT = "%Y%m%d_%H%M%S"
|
||||||
|
|
||||||
|
|
||||||
|
def log(msg: str) -> None:
|
||||||
|
"""Print a status message to stderr."""
|
||||||
|
print(msg, file=sys.stderr)
|
||||||
|
|
||||||
|
|
||||||
|
def episode_task_order(persona: str) -> list[str]:
|
||||||
|
"""Return the ordered task ids from the persona's episode.yaml."""
|
||||||
|
episode_path = DATA_DIR / persona / "episode.yaml"
|
||||||
|
with open(episode_path, "r", encoding="utf-8") as f:
|
||||||
|
episode = yaml.safe_load(f)
|
||||||
|
return [task["task_id"] for task in episode.get("tasks", [])]
|
||||||
|
|
||||||
|
|
||||||
|
def _event_time(record: dict, file_ts: str) -> float:
|
||||||
|
"""Best-effort event time (epoch seconds) of one log record.
|
||||||
|
|
||||||
|
Prefers the record's own timestamp fields; falls back to the timestamp
|
||||||
|
embedded in the log file name so that even stripped records keep a
|
||||||
|
meaningful order. Returns 0.0 when nothing is parseable.
|
||||||
|
"""
|
||||||
|
timestamp = record.get("timestamp")
|
||||||
|
if isinstance(timestamp, (int, float)) and not isinstance(timestamp, bool):
|
||||||
|
return float(timestamp)
|
||||||
|
iso = record.get("timestamp_iso")
|
||||||
|
if isinstance(iso, str):
|
||||||
|
try:
|
||||||
|
return datetime.fromisoformat(iso).timestamp()
|
||||||
|
except ValueError:
|
||||||
|
pass
|
||||||
|
if file_ts:
|
||||||
|
try:
|
||||||
|
return datetime.strptime(file_ts, TIME_FORMAT).timestamp()
|
||||||
|
except ValueError:
|
||||||
|
pass
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
|
||||||
|
def latest_task_statuses(persona: str) -> dict[str, str]:
|
||||||
|
"""Scan per-task and run-level logs; the newest EVENT TIME wins per task.
|
||||||
|
|
||||||
|
Every "Task finished" record across both log categories is keyed by
|
||||||
|
(event_time, file timestamp, file order, line number); the record with
|
||||||
|
the highest key decides the task's status. File category and read order
|
||||||
|
alone can never override a newer record from the other category.
|
||||||
|
"""
|
||||||
|
persona_dir = OUTPUTS_DIR / persona
|
||||||
|
if not persona_dir.is_dir():
|
||||||
|
return {}
|
||||||
|
|
||||||
|
log_files = sorted(persona_dir.glob("*/history/*-log.jsonl"))
|
||||||
|
log_files += sorted(persona_dir.glob("run/*-log.jsonl"))
|
||||||
|
|
||||||
|
best: dict[str, tuple[tuple, str]] = {}
|
||||||
|
for file_order, log_file in enumerate(log_files):
|
||||||
|
ts_match = LOG_FILE_TS_RE.match(log_file.name)
|
||||||
|
file_ts = ts_match.group(1) if ts_match else ""
|
||||||
|
try:
|
||||||
|
with open(log_file, "r", encoding="utf-8") as f:
|
||||||
|
for line_no, line in enumerate(f):
|
||||||
|
if "Task finished" not in line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
record = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
match = TASK_FINISHED_RE.search(str(record.get("message", "")))
|
||||||
|
if not match:
|
||||||
|
continue
|
||||||
|
task_id, status = match.group(1), match.group(2)
|
||||||
|
sort_key = (_event_time(record, file_ts), file_ts, file_order, line_no)
|
||||||
|
current = best.get(task_id)
|
||||||
|
if current is None or sort_key > current[0]:
|
||||||
|
best[task_id] = (sort_key, status)
|
||||||
|
except OSError:
|
||||||
|
continue
|
||||||
|
return {task_id: status for task_id, (_, status) in best.items()}
|
||||||
|
|
||||||
|
|
||||||
|
def split_tasks(persona: str) -> tuple[list[str], list[str]]:
|
||||||
|
"""Split the episode task order into completed and remaining tasks."""
|
||||||
|
order = episode_task_order(persona)
|
||||||
|
statuses = latest_task_statuses(persona)
|
||||||
|
completed = [t for t in order if statuses.get(t) in COMPLETED_STATUSES]
|
||||||
|
remaining = [t for t in order if t not in set(completed)]
|
||||||
|
return completed, remaining
|
||||||
|
|
||||||
|
|
||||||
|
def _daily_note_session_id(note_path: Path) -> str:
|
||||||
|
try:
|
||||||
|
text = note_path.read_text(encoding="utf-8")
|
||||||
|
except OSError:
|
||||||
|
return ""
|
||||||
|
match = SESSION_ID_RE.search(text)
|
||||||
|
return match.group(1) if match else ""
|
||||||
|
|
||||||
|
|
||||||
|
class _WorkspaceFileStoreShim:
|
||||||
|
"""Structural stand-in for ReMe's file store; only workspace_path is read."""
|
||||||
|
|
||||||
|
def __init__(self, workspace_path: Path):
|
||||||
|
self.workspace_path = workspace_path
|
||||||
|
|
||||||
|
|
||||||
|
def _refresh_daily_indexes(
|
||||||
|
workspace: Path,
|
||||||
|
removed_by_date: dict[str, set[str]],
|
||||||
|
removed: list[str],
|
||||||
|
) -> None:
|
||||||
|
"""Rebuild the daily index of each affected date via ReMe's own logic."""
|
||||||
|
for date in sorted(removed_by_date):
|
||||||
|
result = asyncio.run(
|
||||||
|
refresh_day_index(_WorkspaceFileStoreShim(workspace), date, "daily"),
|
||||||
|
)
|
||||||
|
if result.get("error"):
|
||||||
|
log(f"[resume] WARNING: daily index refresh failed for {date}: {result['error']}")
|
||||||
|
continue
|
||||||
|
removed.append(f"daily/{date}.md (refreshed, {len(removed_by_date[date])} note(s) removed)")
|
||||||
|
|
||||||
|
|
||||||
|
def _strip_index_lines(
|
||||||
|
workspace: Path,
|
||||||
|
removed_by_date: dict[str, set[str]],
|
||||||
|
removed: list[str],
|
||||||
|
dry_run: bool,
|
||||||
|
) -> None:
|
||||||
|
"""Fallback index edit: drop lines that reference removed notes by full
|
||||||
|
workspace-relative wikilink path, and fix the note count. Only the index
|
||||||
|
files of affected dates are touched."""
|
||||||
|
for date in sorted(removed_by_date):
|
||||||
|
index_path = workspace / "daily" / f"{date}.md"
|
||||||
|
if not index_path.is_file():
|
||||||
|
continue
|
||||||
|
wikilinks = [f"[[{rel_path}]]" for rel_path in sorted(removed_by_date[date])]
|
||||||
|
lines = index_path.read_text(encoding="utf-8").splitlines()
|
||||||
|
kept = [line for line in lines if not any(link in line for link in wikilinks)]
|
||||||
|
if len(kept) == len(lines):
|
||||||
|
continue
|
||||||
|
note_count = sum(1 for line in kept if line.startswith("- [[daily/"))
|
||||||
|
kept = [NOTE_COUNT_RE.sub(rf"\g<1>{note_count}\2", line) for line in kept]
|
||||||
|
removed.append(f"{index_path.relative_to(workspace)} (rewritten)")
|
||||||
|
if not dry_run:
|
||||||
|
index_path.write_text("\n".join(kept) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def cleanup_partial_memory(persona: str, remaining: list[str], dry_run: bool = False) -> list[str]:
|
||||||
|
"""Remove partial memory artifacts of remaining tasks so they can be re-run cleanly."""
|
||||||
|
workspace = WORKSPACE_ROOT / persona
|
||||||
|
removed: list[str] = []
|
||||||
|
if not workspace.is_dir() or not remaining:
|
||||||
|
return removed
|
||||||
|
|
||||||
|
prefixes = tuple(f"pibench_{task_id}_" for task_id in remaining)
|
||||||
|
|
||||||
|
def act(path: Path, label: str) -> None:
|
||||||
|
removed.append(label)
|
||||||
|
if not dry_run:
|
||||||
|
path.unlink()
|
||||||
|
|
||||||
|
# 1) daily / digest notes distilled from interrupted sessions. For daily
|
||||||
|
# notes, remember the full workspace-relative path grouped by date so only
|
||||||
|
# the affected daily indexes are refreshed below.
|
||||||
|
removed_by_date: dict[str, set[str]] = {}
|
||||||
|
for section in ("daily", "digest"):
|
||||||
|
section_root = workspace / section
|
||||||
|
if not section_root.is_dir():
|
||||||
|
continue
|
||||||
|
for note_path in section_root.rglob("*.md"):
|
||||||
|
if note_path.parent == section_root:
|
||||||
|
continue # index files handled below
|
||||||
|
session_id = _daily_note_session_id(note_path)
|
||||||
|
if session_id.startswith(prefixes):
|
||||||
|
rel_path = note_path.relative_to(workspace).as_posix()
|
||||||
|
act(note_path, rel_path)
|
||||||
|
if section == "daily":
|
||||||
|
removed_by_date.setdefault(note_path.parent.name, set()).add(rel_path)
|
||||||
|
|
||||||
|
# 2) daily index files: refresh only the dates that lost notes, matching
|
||||||
|
# notes by their full wikilink path instead of their bare file name.
|
||||||
|
if removed_by_date:
|
||||||
|
if dry_run:
|
||||||
|
for date in sorted(removed_by_date):
|
||||||
|
removed.append(f"daily/{date}.md (would refresh index)")
|
||||||
|
elif refresh_day_index is not None:
|
||||||
|
_refresh_daily_indexes(workspace, removed_by_date, removed)
|
||||||
|
else:
|
||||||
|
_strip_index_lines(workspace, removed_by_date, removed, dry_run)
|
||||||
|
|
||||||
|
# 3) raw dialog logs of interrupted sessions
|
||||||
|
dialog_dir = workspace / "session" / "dialog"
|
||||||
|
if dialog_dir.is_dir():
|
||||||
|
for task_id in remaining:
|
||||||
|
for dialog_path in dialog_dir.glob(f"pibench_{task_id}_*.jsonl"):
|
||||||
|
act(dialog_path, str(dialog_path.relative_to(workspace)))
|
||||||
|
|
||||||
|
# 4) agent-scope session states that contain interrupted-task sessions
|
||||||
|
mem_session_dir = workspace / "mem_session"
|
||||||
|
if mem_session_dir.is_dir():
|
||||||
|
for session_path in mem_session_dir.rglob("*.jsonl"):
|
||||||
|
try:
|
||||||
|
content = session_path.read_text(encoding="utf-8", errors="ignore")
|
||||||
|
except OSError:
|
||||||
|
continue
|
||||||
|
if any(prefix in content for prefix in prefixes):
|
||||||
|
act(session_path, str(session_path.relative_to(workspace)))
|
||||||
|
|
||||||
|
return removed
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
"""CLI entrypoint: run 'remaining' or 'cleanup' action for a persona."""
|
||||||
|
args = sys.argv[1:]
|
||||||
|
if len(args) < 2 or args[0] not in {"remaining", "cleanup"}:
|
||||||
|
print(__doc__, file=sys.stderr)
|
||||||
|
return 2
|
||||||
|
|
||||||
|
command, persona = args[0], args[1]
|
||||||
|
completed, remaining = split_tasks(persona)
|
||||||
|
|
||||||
|
if command == "remaining":
|
||||||
|
if "--json" in args:
|
||||||
|
print(json.dumps({"completed": completed, "remaining": remaining}))
|
||||||
|
else:
|
||||||
|
for task_id in remaining:
|
||||||
|
print(task_id)
|
||||||
|
log(
|
||||||
|
f"[resume] {persona}: completed={len(completed)} "
|
||||||
|
f"({', '.join(completed) if completed else '-'}) remaining={len(remaining)}",
|
||||||
|
)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
dry_run = "--dry-run" in args
|
||||||
|
removed = cleanup_partial_memory(persona, remaining, dry_run=dry_run)
|
||||||
|
if removed:
|
||||||
|
verb = "would remove" if dry_run else "removed"
|
||||||
|
log(f"[resume] {persona}: {verb} {len(removed)} partial-memory artifact(s):")
|
||||||
|
for item in removed:
|
||||||
|
log(f" - {item}")
|
||||||
|
else:
|
||||||
|
log(f"[resume] {persona}: no partial-memory artifacts to clean")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
119
benchmark/pibench/run_all.sh
Executable file
119
benchmark/pibench/run_all.sh
Executable file
|
|
@ -0,0 +1,119 @@
|
||||||
|
#!/bin/bash
|
||||||
|
# Run all 5 personas with the ReMe agent, PARALLEL at a time (default 2).
|
||||||
|
# Each persona's tasks follow data/{persona}/episode.yaml order.
|
||||||
|
#
|
||||||
|
# Usage:
|
||||||
|
# bash run_all.sh # FRESH official run: wipes ALL personas'
|
||||||
|
# # ReMe memory/outputs/trace logs first,
|
||||||
|
# # then runs everything from scratch.
|
||||||
|
# bash run_all.sh --resume # Checkpoint continuation: no wipe; every
|
||||||
|
# # persona skips already-completed tasks.
|
||||||
|
# bash run_all.sh --parallel 1 # sequential (original behavior)
|
||||||
|
# bash run_all.sh --skip-eval # run phase only
|
||||||
|
#
|
||||||
|
# Memory-wipe vs resume conflict resolution:
|
||||||
|
# The full ReMe memory wipe happens ONLY here, ONLY in fresh mode (the
|
||||||
|
# default), and ONLY before any service/bridge starts. --resume never
|
||||||
|
# wipes; run_persona.sh then additionally performs a surgical cleanup of
|
||||||
|
# residual memory belonging to interrupted (to-be-re-run) tasks, so a
|
||||||
|
# resumed run keeps all completed-task memory but never inherits a partial
|
||||||
|
# task's own answer. The two modes are mutually exclusive.
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
PERSONAS=(researcher marketer law_trainee pharmacist Financier)
|
||||||
|
TRACE_ROOT="${HOME}/.nanobot/trace_logs"
|
||||||
|
|
||||||
|
PARALLEL=2
|
||||||
|
MODE="fresh"
|
||||||
|
PASS_ARGS=()
|
||||||
|
while [[ $# -gt 0 ]]; do
|
||||||
|
case $1 in
|
||||||
|
--parallel)
|
||||||
|
PARALLEL="${2:-}"; shift 2 || true
|
||||||
|
case "$PARALLEL" in (""|*[!0-9]*) echo "--parallel needs a positive integer"; exit 2 ;; esac
|
||||||
|
[ "$PARALLEL" -lt 1 ] && PARALLEL=1
|
||||||
|
[ "$PARALLEL" -gt ${#PERSONAS[@]} ] && PARALLEL=${#PERSONAS[@]}
|
||||||
|
;;
|
||||||
|
--resume)
|
||||||
|
if [ "$MODE" = "fresh_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
|
||||||
|
MODE="resume"; shift ;;
|
||||||
|
--fresh)
|
||||||
|
if [ "$MODE" = "resume" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
|
||||||
|
MODE="fresh_set"; shift ;;
|
||||||
|
--skip-eval) PASS_ARGS+=(--skip-eval); shift ;;
|
||||||
|
*) echo "Unknown option: $1"; exit 1 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
[ "$MODE" = "fresh_set" ] && MODE="fresh"
|
||||||
|
|
||||||
|
START_TS=$(date +%Y%m%d_%H%M%S)
|
||||||
|
SUMMARY_LOG="${SUITE_DIR}/logs/run_all_${START_TS}.summary"
|
||||||
|
mkdir -p "${SUITE_DIR}/logs"
|
||||||
|
|
||||||
|
echo "############################################################"
|
||||||
|
echo "# reme_eval suite | mode=${MODE} parallel=${PARALLEL} | ${START_TS}"
|
||||||
|
echo "############################################################"
|
||||||
|
|
||||||
|
# ─── Fresh mode: suite-level wipe BEFORE anything starts ──────────────
|
||||||
|
if [ "$MODE" = "fresh" ]; then
|
||||||
|
echo "[fresh] wiping ALL personas' memory workspaces, outputs and trace logs..."
|
||||||
|
for persona in "${PERSONAS[@]}"; do
|
||||||
|
rm -rf "${SUITE_DIR}/reme_workspace/${persona}"
|
||||||
|
rm -rf "${SUITE_DIR}/outputs/reme/${persona}"
|
||||||
|
rm -rf "${TRACE_ROOT}/reme/${persona}"
|
||||||
|
rm -rf "${SUITE_DIR}/nanobot_workspace/${persona}"
|
||||||
|
done
|
||||||
|
echo "[fresh] wipe done."
|
||||||
|
else
|
||||||
|
echo "[resume] no memory wipe; personas resume after their last completed task."
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── Run personas in batches of PARALLEL ──────────────────────────────
|
||||||
|
STATUS_LIST=()
|
||||||
|
ANY_FAILED=0
|
||||||
|
OVERALL_START=$(date +%s)
|
||||||
|
TOTAL=${#PERSONAS[@]}
|
||||||
|
|
||||||
|
for ((i = 0; i < TOTAL; i += PARALLEL)); do
|
||||||
|
BATCH=("${PERSONAS[@]:i:PARALLEL}")
|
||||||
|
BATCH_PIDS=()
|
||||||
|
BATCH_NAMES=()
|
||||||
|
echo ""
|
||||||
|
echo "============================================================"
|
||||||
|
echo "# BATCH $(( i / PARALLEL + 1 )): ${BATCH[*]} started $(date '+%F %T')"
|
||||||
|
echo "============================================================"
|
||||||
|
for persona in "${BATCH[@]}"; do
|
||||||
|
bash "${SUITE_DIR}/run_persona.sh" "${persona}" --resume ${PASS_ARGS[@]+"${PASS_ARGS[@]}"} \
|
||||||
|
> "${SUITE_DIR}/logs/suite_${persona}.log" 2>&1 &
|
||||||
|
BATCH_PIDS+=($!)
|
||||||
|
BATCH_NAMES+=("$persona")
|
||||||
|
done
|
||||||
|
for j in $(seq 0 $(( ${#BATCH[@]} - 1 ))); do
|
||||||
|
pid=${BATCH_PIDS[$j]}
|
||||||
|
persona=${BATCH_NAMES[$j]}
|
||||||
|
if wait "$pid"; then
|
||||||
|
STATUS_LIST+=("${persona}: OK")
|
||||||
|
else
|
||||||
|
rc=$?
|
||||||
|
ANY_FAILED=1
|
||||||
|
STATUS_LIST+=("${persona}: FAILED rc=${rc}")
|
||||||
|
echo "[run_all] ${persona} FAILED (rc=${rc}); see logs/suite_${persona}.log"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
done
|
||||||
|
|
||||||
|
total=$(( $(date +%s) - OVERALL_START ))
|
||||||
|
echo ""
|
||||||
|
echo "================ FINAL SUMMARY (${total}s total) ================" | tee -a "${SUMMARY_LOG}"
|
||||||
|
for line in "${STATUS_LIST[@]}"; do
|
||||||
|
echo " ${line}" | tee -a "${SUMMARY_LOG}"
|
||||||
|
done
|
||||||
|
echo "Summary: ${SUMMARY_LOG}"
|
||||||
|
|
||||||
|
if [ "${ANY_FAILED}" -ne 0 ]; then
|
||||||
|
FAILED_COUNT=$(printf '%s\n' "${STATUS_LIST[@]}" | grep -c "FAILED")
|
||||||
|
echo "[run_all] ${FAILED_COUNT} persona(s) FAILED; suite run is marked as failed." | tee -a "${SUMMARY_LOG}"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
exit 0
|
||||||
301
benchmark/pibench/run_persona.sh
Executable file
301
benchmark/pibench/run_persona.sh
Executable file
|
|
@ -0,0 +1,301 @@
|
||||||
|
#!/bin/bash
|
||||||
|
# Run the full pi-bench evaluation for ONE persona with the ReMe agent.
|
||||||
|
# Tasks follow data/{persona}/episode.yaml order (runner-native).
|
||||||
|
#
|
||||||
|
# Usage: bash run_persona.sh <persona> [--fresh|--resume] [--skip-eval]
|
||||||
|
#
|
||||||
|
# Modes (default: --resume):
|
||||||
|
# --resume Checkpoint continuation. Never wipes memory. Tasks already
|
||||||
|
# finished (SUCCESS/MAX_TURNS/TIMEOUT in the task history logs)
|
||||||
|
# are skipped via repeated --task-id flags. Before starting, any
|
||||||
|
# residual memory of tasks that are about to be RE-RUN (partial
|
||||||
|
# sessions from an interrupted run) is surgically removed by
|
||||||
|
# resume.py cleanup, so re-runs don't inherit leaked answers.
|
||||||
|
# --fresh Wipes THIS persona's ReMe memory, outputs and trace logs first,
|
||||||
|
# then runs all tasks from scratch.
|
||||||
|
# The two flags are mutually exclusive. A full multi-persona memory wipe is a
|
||||||
|
# suite-level action of `run_all.sh` (fresh mode), never done here implicitly.
|
||||||
|
set -uo pipefail
|
||||||
|
|
||||||
|
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
TRACE_ROOT="${HOME}/.nanobot/trace_logs"
|
||||||
|
|
||||||
|
# ─── External dependencies (pi-bench / ReMe are NOT bundled; see README) ──
|
||||||
|
if [ ! -f "${SUITE_DIR}/env.sh" ]; then
|
||||||
|
echo "env.sh not found. Run: cp env.sh.example env.sh (then fill in the TODO items)"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
source "${SUITE_DIR}/env.sh"
|
||||||
|
|
||||||
|
PIBENCH_DIR="${PI_BENCH_ROOT:-}"
|
||||||
|
if [ -z "${PIBENCH_DIR}" ] || [ ! -f "${PIBENCH_DIR}/src/main.py" ]; then
|
||||||
|
echo "PI_BENCH_ROOT is unset or invalid (src/main.py not found). Set it in env.sh."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if [ ! -x "${PIBENCH_DIR}/.venv/bin/python" ] || [ ! -x "${PIBENCH_DIR}/.venv/bin/appworld" ]; then
|
||||||
|
echo "pi-bench venv incomplete: ${PIBENCH_DIR}/.venv must provide python + appworld (see README setup)."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if [ ! -x "${REME_DIR}/.venv/bin/python" ]; then
|
||||||
|
echo "ReMe venv not found: ${REME_DIR}/.venv/bin/python (check REME_DIR in env.sh)"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if [ ! -e "${SUITE_DIR}/data" ]; then
|
||||||
|
echo 'Benchmark data not linked. Run: ln -s "$PI_BENCH_ROOT/data" data'
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── Pre-flight: files the runner needs before any service starts ─────
|
||||||
|
MODEL_CONFIG="${SUITE_DIR}/config/models/reme.yaml"
|
||||||
|
HISTORY_CONFIG="${SUITE_DIR}/config/bench/evaluation/trace_history.yaml"
|
||||||
|
if [ ! -f "${MODEL_CONFIG}" ]; then
|
||||||
|
echo "Model config not found: ${MODEL_CONFIG} (see README directory layout)."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if [ ! -f "${HISTORY_CONFIG}" ]; then
|
||||||
|
echo "Trace history config not found: ${HISTORY_CONFIG}"
|
||||||
|
echo "pi-bench requires config/bench/evaluation/trace_history.yaml; see README."
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
APPWORLD_DIR="${PIBENCH_DIR}/third_party/appworld"
|
||||||
|
PI_PYTHON="${PIBENCH_DIR}/.venv/bin/python"
|
||||||
|
APPWORLD_BIN="${PIBENCH_DIR}/.venv/bin/appworld"
|
||||||
|
# resume.py runs on the ReMe venv so it can reuse ReMe's daily-index rebuild.
|
||||||
|
REME_PYTHON="${REME_DIR}/.venv/bin/python"
|
||||||
|
|
||||||
|
PERSONA="${1:-}"
|
||||||
|
if [ -z "$PERSONA" ]; then
|
||||||
|
echo "Usage: $0 <persona> [--fresh|--resume] [--skip-eval]"
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
shift
|
||||||
|
|
||||||
|
MODE="resume"
|
||||||
|
SKIP_EVAL=false
|
||||||
|
while [[ $# -gt 0 ]]; do
|
||||||
|
case $1 in
|
||||||
|
--fresh)
|
||||||
|
if [ "$MODE" = "resume_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
|
||||||
|
MODE="fresh"; shift ;;
|
||||||
|
--resume)
|
||||||
|
if [ "$MODE" = "fresh" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
|
||||||
|
MODE="resume_set"; shift ;;
|
||||||
|
--skip-eval) SKIP_EVAL=true; shift ;;
|
||||||
|
*) echo "Unknown option: $1"; exit 1 ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
[ "$MODE" = "resume_set" ] && MODE="resume"
|
||||||
|
|
||||||
|
# ─── Per-persona ports (pi-bench AGENTS.md convention) ────────────────
|
||||||
|
# REME_PORT: ReMe's internal HTTP service; must be unique per concurrent bridge.
|
||||||
|
case "$PERSONA" in
|
||||||
|
marketer) API_PORT=9001; MCP_PORT=10001; TEST_PORT=9998; REME_PORT=18766 ;;
|
||||||
|
law_trainee) API_PORT=9002; MCP_PORT=10002; TEST_PORT=9997; REME_PORT=18767 ;;
|
||||||
|
pharmacist) API_PORT=9003; MCP_PORT=10003; TEST_PORT=9996; REME_PORT=18768 ;;
|
||||||
|
researcher) API_PORT=9004; MCP_PORT=10004; TEST_PORT=9995; REME_PORT=18765 ;;
|
||||||
|
Financier) API_PORT=9005; MCP_PORT=10005; TEST_PORT=9994; REME_PORT=18769 ;;
|
||||||
|
*) echo "Unknown persona: $PERSONA"; exit 1 ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
API_URL="http://127.0.0.1:${API_PORT}"
|
||||||
|
MCP_URL="http://127.0.0.1:${MCP_PORT}/mcp"
|
||||||
|
TEST_URL="http://127.0.0.1:${TEST_PORT}"
|
||||||
|
LOG_DIR="${SUITE_DIR}/logs"
|
||||||
|
mkdir -p "${LOG_DIR}"
|
||||||
|
|
||||||
|
# ─── Environment (env.sh already sourced at the top) ──────────────────
|
||||||
|
WORKSPACE_DIR="${REME_WORKSPACE_ROOT}/${PERSONA}"
|
||||||
|
NANOBOT_WORKSPACE_DIR="${SUITE_DIR}/nanobot_workspace/${PERSONA}"
|
||||||
|
mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
|
||||||
|
|
||||||
|
echo "========================================="
|
||||||
|
echo "ReMe x Pi-Bench | persona=${PERSONA} | mode=${MODE}"
|
||||||
|
echo " api=${API_PORT} mcp=${MCP_PORT} test=${TEST_PORT} reme=${REME_PORT}"
|
||||||
|
echo " model=${REME_MODEL_NAME}"
|
||||||
|
echo " memory workspace=${WORKSPACE_DIR} (persistent)"
|
||||||
|
echo "========================================="
|
||||||
|
|
||||||
|
# ─── Fresh mode: wipe this persona's state ────────────────────────────
|
||||||
|
if [ "$MODE" = "fresh" ]; then
|
||||||
|
echo "[fresh] wiping persona state: memory workspace, outputs, trace logs"
|
||||||
|
rm -rf "${WORKSPACE_DIR}"
|
||||||
|
rm -rf "${SUITE_DIR}/outputs/reme/${PERSONA}"
|
||||||
|
rm -rf "${TRACE_ROOT}/reme/${PERSONA}"
|
||||||
|
rm -rf "${NANOBOT_WORKSPACE_DIR}"
|
||||||
|
mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── Resume: determine remaining tasks + clean partial memories ───────
|
||||||
|
TASK_ARGS=()
|
||||||
|
RUN_PHASE_NEEDED=true
|
||||||
|
if [ "$MODE" = "resume" ]; then
|
||||||
|
REMAINING_JSON="$("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" --json)"
|
||||||
|
if [ -z "$REMAINING_JSON" ]; then
|
||||||
|
echo "Failed to compute remaining tasks"; exit 1
|
||||||
|
fi
|
||||||
|
echo "[resume] ${REMAINING_JSON}"
|
||||||
|
REMAINING_TASKS=()
|
||||||
|
while IFS= read -r tid_line; do
|
||||||
|
[ -n "$tid_line" ] && REMAINING_TASKS+=("$tid_line")
|
||||||
|
done < <("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" 2>/dev/null)
|
||||||
|
if [ ${#REMAINING_TASKS[@]} -eq 0 ]; then
|
||||||
|
RUN_PHASE_NEEDED=false
|
||||||
|
echo "[resume] all tasks already completed; skipping run phase"
|
||||||
|
else
|
||||||
|
# Remove residual memory of interrupted (to-be-re-run) tasks so
|
||||||
|
# re-runs don't get their own partial answers injected.
|
||||||
|
"${REME_PYTHON}" "${SUITE_DIR}/resume.py" cleanup "${PERSONA}"
|
||||||
|
for tid in "${REMAINING_TASKS[@]}"; do
|
||||||
|
TASK_ARGS+=(--task-id "$tid")
|
||||||
|
done
|
||||||
|
echo "[resume] running ${#REMAINING_TASKS[@]} remaining task(s): ${REMAINING_TASKS[*]}"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── Port cleanup from previous runs ──────────────────────────────────
|
||||||
|
for port in ${API_PORT} ${MCP_PORT} ${TEST_PORT} ${REME_PORT}; do
|
||||||
|
pids=$(lsof -ti :${port} 2>/dev/null || true)
|
||||||
|
if [ -n "$pids" ]; then
|
||||||
|
echo "Killing stale processes on port ${port}: ${pids}"
|
||||||
|
kill -9 $pids 2>/dev/null || true
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
sleep 2
|
||||||
|
|
||||||
|
PIDS=()
|
||||||
|
cleanup() {
|
||||||
|
echo "[${PERSONA}] cleaning up services..."
|
||||||
|
for pid in "${PIDS[@]:-}"; do
|
||||||
|
kill "$pid" 2>/dev/null || true
|
||||||
|
done
|
||||||
|
wait 2>/dev/null || true
|
||||||
|
}
|
||||||
|
trap cleanup EXIT INT TERM
|
||||||
|
|
||||||
|
wait_for_service() {
|
||||||
|
local url="$1" name="$2" port="$3" timeout="${4:-180}"
|
||||||
|
echo -n " waiting for ${name}..."
|
||||||
|
local start=$(date +%s)
|
||||||
|
while true; do
|
||||||
|
if curl -sf --max-time 5 "${url}" > /dev/null 2>&1; then
|
||||||
|
echo " ready"; return 0
|
||||||
|
fi
|
||||||
|
if [ -n "$port" ] && lsof -ti :${port} > /dev/null 2>&1; then
|
||||||
|
local elapsed=$(( $(date +%s) - start ))
|
||||||
|
if [ "$elapsed" -ge 10 ]; then echo " ready (port)"; return 0; fi
|
||||||
|
fi
|
||||||
|
if [ $(( $(date +%s) - start )) -ge "$timeout" ]; then
|
||||||
|
echo " TIMEOUT"; return 1
|
||||||
|
fi
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
|
# ─── [1/5] AppWorld API ────────────────────────────────────────────────
|
||||||
|
echo "[1/5] AppWorld API (:${API_PORT})"
|
||||||
|
(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve apis --root . \
|
||||||
|
--port ${API_PORT}) > "${LOG_DIR}/appworld_api_${PERSONA}.log" 2>&1 &
|
||||||
|
PIDS+=($!)
|
||||||
|
if ! wait_for_service "${API_URL}/docs" "AppWorld API" "${API_PORT}" 180; then
|
||||||
|
tail -20 "${LOG_DIR}/appworld_api_${PERSONA}.log"; exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── [2/5] AppWorld MCP ────────────────────────────────────────────────
|
||||||
|
echo "[2/5] AppWorld MCP (:${MCP_PORT})"
|
||||||
|
TOOLS_CONFIG="${SUITE_DIR}/data/${PERSONA}/tools.yaml"
|
||||||
|
(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve mcp http --root . \
|
||||||
|
--remote-apis-url "${API_URL}" --port ${MCP_PORT} \
|
||||||
|
--tools-config-file "${TOOLS_CONFIG}") > "${LOG_DIR}/appworld_mcp_${PERSONA}.log" 2>&1 &
|
||||||
|
PIDS+=($!)
|
||||||
|
if ! wait_for_service "${MCP_URL}" "AppWorld MCP" "${MCP_PORT}" 180; then
|
||||||
|
tail -20 "${LOG_DIR}/appworld_mcp_${PERSONA}.log"; exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── [3/5] Test Server ─────────────────────────────────────────────────
|
||||||
|
echo "[3/5] Test Server (:${TEST_PORT})"
|
||||||
|
PORT=${TEST_PORT} "${PI_PYTHON}" "${PIBENCH_DIR}/scripts/test_server.py" \
|
||||||
|
> "${LOG_DIR}/test_server_${PERSONA}.log" 2>&1 &
|
||||||
|
PIDS+=($!)
|
||||||
|
if ! wait_for_service "${TEST_URL}/sent?after=-1" "Test Server" "${TEST_PORT}" 30; then
|
||||||
|
tail -20 "${LOG_DIR}/test_server_${PERSONA}.log"; exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── [4/5] ReMe Bridge (ReMe venv) ─────────────────────────────────────
|
||||||
|
echo "[4/5] ReMe Bridge (reme service port ${REME_PORT})"
|
||||||
|
"${REME_DIR}/.venv/bin/python" "${SUITE_DIR}/bridge_reme.py" \
|
||||||
|
--test-server-url "${TEST_URL}" \
|
||||||
|
--appworld-mcp-url "${MCP_URL}" \
|
||||||
|
--reme-dir "${REME_DIR}" \
|
||||||
|
--data-root "${SUITE_DIR}/data" \
|
||||||
|
--user-id "${PERSONA}" \
|
||||||
|
--workspace-dir "${WORKSPACE_DIR}" \
|
||||||
|
--reme-port "${REME_PORT}" \
|
||||||
|
--model-name "${REME_MODEL_NAME}" \
|
||||||
|
--model-base-url "${REME_LLM_BASE_URL}" \
|
||||||
|
--model-api-key "${REME_LLM_API_KEY}" \
|
||||||
|
> "${LOG_DIR}/bridge_${PERSONA}.log" 2>&1 &
|
||||||
|
BRIDGE_PID=$!
|
||||||
|
PIDS+=(${BRIDGE_PID})
|
||||||
|
sleep 5
|
||||||
|
if ! kill -0 "${BRIDGE_PID}" 2>/dev/null; then
|
||||||
|
echo "Bridge failed to start:"; tail -30 "${LOG_DIR}/bridge_${PERSONA}.log"; exit 1
|
||||||
|
fi
|
||||||
|
for i in $(seq 1 12); do
|
||||||
|
if grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null; then
|
||||||
|
echo " bridge initialized"; break
|
||||||
|
fi
|
||||||
|
sleep 5
|
||||||
|
done
|
||||||
|
grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null || {
|
||||||
|
echo "WARNING: bridge may not be ready:"; tail -20 "${LOG_DIR}/bridge_${PERSONA}.log"; }
|
||||||
|
|
||||||
|
# ─── [5/5] Runner (run phase) ──────────────────────────────────────────
|
||||||
|
if [ "$RUN_PHASE_NEEDED" = true ]; then
|
||||||
|
echo "[5/5] Runner: run phase (episode order from data/${PERSONA}/episode.yaml)"
|
||||||
|
cd "${SUITE_DIR}"
|
||||||
|
BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
|
||||||
|
"${PI_PYTHON}" -m src.main \
|
||||||
|
--model-config "${MODEL_CONFIG}" \
|
||||||
|
--history-config-path "${HISTORY_CONFIG}" \
|
||||||
|
--mode run --user-id "${PERSONA}" \
|
||||||
|
--workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
|
||||||
|
${TASK_ARGS[@]+"${TASK_ARGS[@]}"} \
|
||||||
|
2>&1 | tee "${LOG_DIR}/runner_run_${PERSONA}.log"
|
||||||
|
RUN_EXIT=${PIPESTATUS[0]}
|
||||||
|
if [ ${RUN_EXIT} -ne 0 ]; then
|
||||||
|
echo "Run phase failed (exit ${RUN_EXIT}). Logs: ${LOG_DIR}/"
|
||||||
|
exit ${RUN_EXIT}
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
echo "[5/5] Runner: run phase skipped (all tasks completed)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$SKIP_EVAL" = true ]; then
|
||||||
|
echo "Skipping eval (--skip-eval)"
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
|
||||||
|
# ─── Trace conversion + eval phase (always over all available traces) ──
|
||||||
|
echo "Converting trace logs..."
|
||||||
|
"${PI_PYTHON}" "${SUITE_DIR}/fix_trace_logs.py" "${PERSONA}"
|
||||||
|
|
||||||
|
echo "Runner: eval phase"
|
||||||
|
cd "${SUITE_DIR}"
|
||||||
|
BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
|
||||||
|
"${PI_PYTHON}" -m src.main \
|
||||||
|
--model-config "${MODEL_CONFIG}" \
|
||||||
|
--history-config-path "${HISTORY_CONFIG}" \
|
||||||
|
--mode eval --user-id "${PERSONA}" \
|
||||||
|
--workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
|
||||||
|
2>&1 | tee "${LOG_DIR}/runner_eval_${PERSONA}.log"
|
||||||
|
EVAL_EXIT=${PIPESTATUS[0]}
|
||||||
|
|
||||||
|
echo ""
|
||||||
|
echo "========================================="
|
||||||
|
echo "persona=${PERSONA} finished (eval exit=${EVAL_EXIT})"
|
||||||
|
echo " results : ${SUITE_DIR}/outputs/reme/${PERSONA}/"
|
||||||
|
echo " memory : ${WORKSPACE_DIR}/"
|
||||||
|
echo " logs : ${LOG_DIR}/"
|
||||||
|
echo "========================================="
|
||||||
|
exit ${EVAL_EXIT}
|
||||||
98
benchmark/toolmemory/README.md
Normal file
98
benchmark/toolmemory/README.md
Normal file
|
|
@ -0,0 +1,98 @@
|
||||||
|
## Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
|
||||||
|
|
||||||
|
**Language**: English (default) / [中文](./README_ZH.md)
|
||||||
|
|
||||||
|
> Paper: [arXiv:2608.03403](https://arxiv.org/abs/2608.03403)
|
||||||
|
> Code: [https://github.com/WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img src="./gitcha.png" alt="ExpG challenges and overview" width="85%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
### Overview
|
||||||
|
|
||||||
|
This folder archives **ExpG**, a tool-use enhancement built on [Agentscope ReMe](https://github.com/agentscope-ai/ReMe). ExpG mines, distills, and reuses experience from historical tool calls to provide **capability boundaries** and **best-practice guidance**, which helps agents:
|
||||||
|
|
||||||
|
- Select and invoke tools more robustly under dynamic or noisy environments;
|
||||||
|
- Let smaller models with guidance outperform larger, memoryless baselines;
|
||||||
|
- Improve consistently across tool selection, tool calling, and response generation.
|
||||||
|
|
||||||
|
**How ReMe is used:** Start the Tool Memory service; historical tool calls are written and evaluated via `add_tool_call_result`, distilled into tool-level guidance via `summary_tool_memory`, then retrieved and injected into later reasoning via `retrieve_tool_memory`. ReMe provides the vector store and service APIs; the acquisition / distillation / reuse strategy is implemented by ExpG. Full implementation and experiments are in [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### ExpG Mechanism
|
||||||
|
|
||||||
|
ExpG treats tool invocations as learnable experience and runs a three-stage pipeline:
|
||||||
|
|
||||||
|
1. **Experience Acquisition**
|
||||||
|
- Analyze invocation quality from historical trajectories (success/failure, cost, latency, etc.);
|
||||||
|
- Build structured experience units per tool, recording context, parameter patterns, and outcomes.
|
||||||
|
|
||||||
|
2. **Experience Distillation**
|
||||||
|
- Filter noisy or unhelpful experiences and keep representative patterns;
|
||||||
|
- Aggregate by equivalence classes to cover common and rare failure modes;
|
||||||
|
- Summarize with an LLM into generalizable textual guidance.
|
||||||
|
|
||||||
|
3. **Experience Reuse**
|
||||||
|
- Retrieve relevant experience / guidance for future tasks;
|
||||||
|
- Inject guidance into tool selection, argument generation, and response synthesis;
|
||||||
|
- Improve stability under dynamic environments and imperfect feedback.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Main Results
|
||||||
|
|
||||||
|
Performance comparison (%) across MetaTool, API-Bank, and BFCL-V3. **Bold** indicates the best results within each model.
|
||||||
|
|
||||||
|
| Model | Method | MetaTool Pass@1 | MetaTool Avg@3 | MetaTool Pass@3 | API-Bank Pass@1 | API-Bank Avg@3 | API-Bank Pass@3 | BFCL-V3 Pass@1 | BFCL-V3 Avg@3 | BFCL-V3 Pass@3 | Total Pass@1 | Total Avg@3 | Total Pass@3 |
|
||||||
|
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| GPT-5 nano | No Method | 72.62 | 72.76 | 78.49 | 82.96 | 83.46 | 86.97 | 53.80 | 53.00 | 60.95 | 70.82 | 70.62 | 76.63 |
|
||||||
|
| GPT-5 nano | Few-shot | 74.12 | 75.11 | 82.32 | 83.71 | 83.96 | **87.22** | 56.18 | 55.24 | 61.39 | 72.36 | 72.65 | 79.28 |
|
||||||
|
| GPT-5 nano | DRAFT | 73.94 | 73.04 | 78.97 | 84.21 | 83.46 | **87.22** | 57.27 | 57.27 | 62.26 | 72.52 | 71.58 | 77.23 |
|
||||||
|
| GPT-5 nano | Mem0 | 74.96 | 76.13 | 82.92 | 84.96 | 85.21 | **87.22** | 60.95 | 61.61 | 65.08 | 73.98 | 74.67 | 80.35 |
|
||||||
|
| GPT-5 nano | **ExpG** | **81.67** | **82.07** | **84.60** | **86.72** | **86.55** | **87.22** | **64.43** | **63.99** | **66.38** | **79.32** | **79.22** | **81.69** |
|
||||||
|
| DeepSeek-V3 | No Method | 83.10 | 82.94 | 84.66 | 84.71 | 84.38 | 85.46 | 58.79 | 59.65 | 65.94 | 78.92 | 78.66 | 81.37 |
|
||||||
|
| DeepSeek-V3 | Few-shot | 82.74 | 83.90 | 86.28 | 85.21 | 84.63 | 86.22 | 60.52 | 60.30 | 67.90 | 79.08 | 79.45 | 82.92 |
|
||||||
|
| DeepSeek-V3 | DRAFT | 80.23 | 80.79 | 82.44 | 84.96 | 85.63 | 86.47 | 62.26 | 61.61 | 68.55 | 77.70 | 77.80 | 80.54 |
|
||||||
|
| DeepSeek-V3 | Mem0 | 83.88 | 84.56 | 86.40 | 85.46 | 85.55 | 86.47 | 65.08 | 65.15 | 68.33 | 80.70 | 80.91 | 83.12 |
|
||||||
|
| DeepSeek-V3 | **ExpG** | **85.26** | **85.38** | **86.52** | **87.72** | **87.39** | **87.97** | **69.41** | **69.92** | **72.02** | **82.76** | **82.61** | **84.11** |
|
||||||
|
| Qwen3-8B | No Method | 76.51 | 76.97 | 77.71 | 83.96 | 83.88 | 84.21 | 58.79 | 58.28 | 60.30 | 74.46 | 74.41 | 75.56 |
|
||||||
|
| Qwen3-8B | Few-shot | 79.93 | 79.83 | 82.92 | 83.71 | 82.62 | 84.96 | 60.09 | 59.29 | 61.39 | 76.91 | 76.27 | 79.32 |
|
||||||
|
| Qwen3-8B | DRAFT | 78.19 | 77.33 | 77.89 | 85.71 | 84.96 | 85.46 | 60.74 | 60.30 | 62.91 | 76.20 | 75.18 | 76.35 |
|
||||||
|
| Qwen3-8B | Mem0 | 75.07 | 75.47 | 82.38 | 86.22 | 86.05 | 86.47 | 63.34 | 64.93 | 66.16 | 74.69 | 74.98 | 80.07 |
|
||||||
|
| Qwen3-8B | **ExpG** | **83.52** | **84.88** | **85.08** | **86.47** | **87.89** | **87.97** | **67.46** | **66.96** | **68.33** | **81.06** | **81.82** | **82.48** |
|
||||||
|
| Qwen3-32B | No Method | 80.05 | 79.43 | 80.17 | 84.71 | 84.88 | 85.21 | 65.15 | 65.08 | 66.16 | 78.05 | 77.55 | 78.41 |
|
||||||
|
| Qwen3-32B | **ExpG** | **84.68** | **85.02** | **86.28** | **86.97** | **87.30** | **87.72** | **70.72** | **71.01** | **73.32** | **82.48** | **82.56** | **84.14** |
|
||||||
|
| Qwen3-235B | No Method | 78.25 | 79.23 | 80.29 | 85.46 | 85.46 | 85.71 | 71.37 | 71.15 | 73.54 | 78.13 | 78.49 | 79.91 |
|
||||||
|
| Qwen3-235B | **ExpG** | **86.34** | **86.70** | **86.94** | **87.47** | **86.97** | **88.22** | **79.61** | **78.52** | **80.04** | **85.29** | **84.98** | **85.69** |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Reference Code
|
||||||
|
|
||||||
|
| Path | Role |
|
||||||
|
| --- | --- |
|
||||||
|
| [`tool_memory.py`](./tool_memory.py) | HTTP client for official ReMe Tool Memory APIs (`add_tool_call_result` / `summary_tool_memory` / `retrieve_tool_memory`) |
|
||||||
|
| [`parse_tool_call_result_prompt.yaml`](./parse_tool_call_result_prompt.yaml) | Prompt for multi-aspect evaluation of each tool call |
|
||||||
|
| [`summary_tool_memory_prompt.yaml`](./summary_tool_memory_prompt.yaml) | Prompt for summarizing tool call history into guidance |
|
||||||
|
| [`tool_memory_flows.yaml`](./tool_memory_flows.yaml) | Tool Memory flow / op config excerpt |
|
||||||
|
|
||||||
|
These are reference snippets. For the full runnable codebase, see [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Citation
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@misc{wang2026expg,
|
||||||
|
title = {Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance},
|
||||||
|
author = {Can Wang and Haoran Chen and Li Yu and Ding Hao and Bohai Zhao and Zhaoyang Liu and Zhiying Tu},
|
||||||
|
year = {2026},
|
||||||
|
eprint = {2608.03403},
|
||||||
|
archivePrefix = {arXiv},
|
||||||
|
primaryClass = {cs.AI},
|
||||||
|
url = {https://arxiv.org/abs/2608.03403},
|
||||||
|
howpublished = {\url{https://github.com/WangCan1178/ExpG}}
|
||||||
|
}
|
||||||
|
```
|
||||||
98
benchmark/toolmemory/README_ZH.md
Normal file
98
benchmark/toolmemory/README_ZH.md
Normal file
|
|
@ -0,0 +1,98 @@
|
||||||
|
## Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
|
||||||
|
|
||||||
|
**语言**:中文 / [English](./README.md)
|
||||||
|
|
||||||
|
> 论文:[arXiv:2608.03403](https://arxiv.org/abs/2608.03403)
|
||||||
|
> 代码:[https://github.com/WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img src="./gitcha.png" alt="ExpG 挑战与概览" width="85%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
### 简介
|
||||||
|
|
||||||
|
本目录归档基于 [Agentscope ReMe](https://github.com/agentscope-ai/ReMe) 的工具使用增强工作 **ExpG**:在 ReMe 记忆框架之上,从历史工具调用中挖掘、提炼并复用经验,为智能体提供工具的 **能力边界** 与 **最佳实践指导**,从而:
|
||||||
|
|
||||||
|
- 在动态或有噪环境下更鲁棒地选择和调用工具;
|
||||||
|
- 让较小模型在带有经验指导时超越更大、但无记忆的基线;
|
||||||
|
- 在工具选择、工具调用和响应生成等多个阶段带来一致收益。
|
||||||
|
|
||||||
|
**如何使用 ReMe:** 启动 Tool Memory 服务后,历史工具调用经 `add_tool_call_result` 写入并评估,经 `summary_tool_memory` 蒸馏成工具级指导,再经 `retrieve_tool_memory` 取回并注入后续推理。向量存储与服务接口由 ReMe 提供,经验获取 / 蒸馏 / 复用策略由 ExpG 实现。完整实现与实验见 [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### ExpG 机制概览
|
||||||
|
|
||||||
|
ExpG 将工具调用视为可学习经验,并通过三阶段流水线完成经验的获取、提炼与复用:
|
||||||
|
|
||||||
|
1. **经验获取(Experience Acquisition)**
|
||||||
|
- 从历史工具调用轨迹中分析调用质量(成功/失败、代价、时间等);
|
||||||
|
- 针对不同工具构建结构化的经验单元,记录调用上下文、参数模式和结果。
|
||||||
|
|
||||||
|
2. **经验蒸馏(Experience Distillation)**
|
||||||
|
- 过滤无效 / 噪声经验,保留具有代表性的调用模式;
|
||||||
|
- 基于“等价类”视角对经验进行聚合,覆盖常见模式与稀有失败模式;
|
||||||
|
- 使用 LLM 对经验进行总结,形成可泛化的文本化指导(guidance)。
|
||||||
|
|
||||||
|
3. **经验复用(Experience Reuse)**
|
||||||
|
- 在未来任务中,根据当前工具调用上下文检索相关经验 / 指导;
|
||||||
|
- 将经验引导融入到工具选择、参数生成和响应整理等环节;
|
||||||
|
- 使得代理在面对动态环境和不完美反馈时仍能保持稳定表现。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 主实验结果
|
||||||
|
|
||||||
|
MetaTool、API-Bank、BFCL-V3 上的性能对比(%)。**加粗**为各模型组内最优。
|
||||||
|
|
||||||
|
| Model | Method | MetaTool Pass@1 | MetaTool Avg@3 | MetaTool Pass@3 | API-Bank Pass@1 | API-Bank Avg@3 | API-Bank Pass@3 | BFCL-V3 Pass@1 | BFCL-V3 Avg@3 | BFCL-V3 Pass@3 | Total Pass@1 | Total Avg@3 | Total Pass@3 |
|
||||||
|
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| GPT-5 nano | No Method | 72.62 | 72.76 | 78.49 | 82.96 | 83.46 | 86.97 | 53.80 | 53.00 | 60.95 | 70.82 | 70.62 | 76.63 |
|
||||||
|
| GPT-5 nano | Few-shot | 74.12 | 75.11 | 82.32 | 83.71 | 83.96 | **87.22** | 56.18 | 55.24 | 61.39 | 72.36 | 72.65 | 79.28 |
|
||||||
|
| GPT-5 nano | DRAFT | 73.94 | 73.04 | 78.97 | 84.21 | 83.46 | **87.22** | 57.27 | 57.27 | 62.26 | 72.52 | 71.58 | 77.23 |
|
||||||
|
| GPT-5 nano | Mem0 | 74.96 | 76.13 | 82.92 | 84.96 | 85.21 | **87.22** | 60.95 | 61.61 | 65.08 | 73.98 | 74.67 | 80.35 |
|
||||||
|
| GPT-5 nano | **ExpG** | **81.67** | **82.07** | **84.60** | **86.72** | **86.55** | **87.22** | **64.43** | **63.99** | **66.38** | **79.32** | **79.22** | **81.69** |
|
||||||
|
| DeepSeek-V3 | No Method | 83.10 | 82.94 | 84.66 | 84.71 | 84.38 | 85.46 | 58.79 | 59.65 | 65.94 | 78.92 | 78.66 | 81.37 |
|
||||||
|
| DeepSeek-V3 | Few-shot | 82.74 | 83.90 | 86.28 | 85.21 | 84.63 | 86.22 | 60.52 | 60.30 | 67.90 | 79.08 | 79.45 | 82.92 |
|
||||||
|
| DeepSeek-V3 | DRAFT | 80.23 | 80.79 | 82.44 | 84.96 | 85.63 | 86.47 | 62.26 | 61.61 | 68.55 | 77.70 | 77.80 | 80.54 |
|
||||||
|
| DeepSeek-V3 | Mem0 | 83.88 | 84.56 | 86.40 | 85.46 | 85.55 | 86.47 | 65.08 | 65.15 | 68.33 | 80.70 | 80.91 | 83.12 |
|
||||||
|
| DeepSeek-V3 | **ExpG** | **85.26** | **85.38** | **86.52** | **87.72** | **87.39** | **87.97** | **69.41** | **69.92** | **72.02** | **82.76** | **82.61** | **84.11** |
|
||||||
|
| Qwen3-8B | No Method | 76.51 | 76.97 | 77.71 | 83.96 | 83.88 | 84.21 | 58.79 | 58.28 | 60.30 | 74.46 | 74.41 | 75.56 |
|
||||||
|
| Qwen3-8B | Few-shot | 79.93 | 79.83 | 82.92 | 83.71 | 82.62 | 84.96 | 60.09 | 59.29 | 61.39 | 76.91 | 76.27 | 79.32 |
|
||||||
|
| Qwen3-8B | DRAFT | 78.19 | 77.33 | 77.89 | 85.71 | 84.96 | 85.46 | 60.74 | 60.30 | 62.91 | 76.20 | 75.18 | 76.35 |
|
||||||
|
| Qwen3-8B | Mem0 | 75.07 | 75.47 | 82.38 | 86.22 | 86.05 | 86.47 | 63.34 | 64.93 | 66.16 | 74.69 | 74.98 | 80.07 |
|
||||||
|
| Qwen3-8B | **ExpG** | **83.52** | **84.88** | **85.08** | **86.47** | **87.89** | **87.97** | **67.46** | **66.96** | **68.33** | **81.06** | **81.82** | **82.48** |
|
||||||
|
| Qwen3-32B | No Method | 80.05 | 79.43 | 80.17 | 84.71 | 84.88 | 85.21 | 65.15 | 65.08 | 66.16 | 78.05 | 77.55 | 78.41 |
|
||||||
|
| Qwen3-32B | **ExpG** | **84.68** | **85.02** | **86.28** | **86.97** | **87.30** | **87.72** | **70.72** | **71.01** | **73.32** | **82.48** | **82.56** | **84.14** |
|
||||||
|
| Qwen3-235B | No Method | 78.25 | 79.23 | 80.29 | 85.46 | 85.46 | 85.71 | 71.37 | 71.15 | 73.54 | 78.13 | 78.49 | 79.91 |
|
||||||
|
| Qwen3-235B | **ExpG** | **86.34** | **86.70** | **86.94** | **87.47** | **86.97** | **88.22** | **79.61** | **78.52** | **80.04** | **85.29** | **84.98** | **85.69** |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 参考代码
|
||||||
|
|
||||||
|
| 路径 | 作用 |
|
||||||
|
| --- | --- |
|
||||||
|
| [`tool_memory.py`](./tool_memory.py) | 官方风格 ReMe Tool Memory HTTP 客户端(`add_tool_call_result` / `summary_tool_memory` / `retrieve_tool_memory`) |
|
||||||
|
| [`parse_tool_call_result_prompt.yaml`](./parse_tool_call_result_prompt.yaml) | 单次工具调用多维评估用的 prompt |
|
||||||
|
| [`summary_tool_memory_prompt.yaml`](./summary_tool_memory_prompt.yaml) | 将工具调用历史总结为 guidance 的 prompt |
|
||||||
|
| [`tool_memory_flows.yaml`](./tool_memory_flows.yaml) | Tool Memory 相关的 flow / op 配置摘录 |
|
||||||
|
|
||||||
|
以上为参考片段。完整可运行代码见 [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 引用
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@misc{wang2026expg,
|
||||||
|
title = {Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance},
|
||||||
|
author = {Can Wang and Haoran Chen and Li Yu and Ding Hao and Bohai Zhao and Zhaoyang Liu and Zhiying Tu},
|
||||||
|
year = {2026},
|
||||||
|
eprint = {2608.03403},
|
||||||
|
archivePrefix = {arXiv},
|
||||||
|
primaryClass = {cs.AI},
|
||||||
|
url = {https://arxiv.org/abs/2608.03403},
|
||||||
|
howpublished = {\url{https://github.com/WangCan1178/ExpG}}
|
||||||
|
}
|
||||||
|
```
|
||||||
BIN
benchmark/toolmemory/gitcha.png
Normal file
BIN
benchmark/toolmemory/gitcha.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 1.9 MiB |
49
benchmark/toolmemory/parse_tool_call_result_prompt.yaml
Normal file
49
benchmark/toolmemory/parse_tool_call_result_prompt.yaml
Normal file
|
|
@ -0,0 +1,49 @@
|
||||||
|
prompt: |
|
||||||
|
You are an expert in evaluating tool invocation process. The tool is invoked by an AI agent.
|
||||||
|
|
||||||
|
Tool invocation Information:
|
||||||
|
- Tool Name: {tool_name}
|
||||||
|
- Success Flag: {success_flag}
|
||||||
|
- Time Cost: {time_cost}s
|
||||||
|
- Token Cost: {token_cost} tokens
|
||||||
|
- Agent Context: {context}
|
||||||
|
- Input Parameters: {input_params}
|
||||||
|
- Tool Response: {response}
|
||||||
|
- Tool Schema: {schema}
|
||||||
|
|
||||||
|
Evaluation Method:
|
||||||
|
Start from a default score list of scores = [0, 0, 0, 0, 0, 0, 0, 0, 0, 0].
|
||||||
|
For each item below that is satisfied, assign 1 point to the corresponding index.
|
||||||
|
The final scores should be a list of 10 integers, each being either 0 or 1.
|
||||||
|
|
||||||
|
1. Use Quality (total 2 points. If context is provided, use it as an aid when evaluating):
|
||||||
|
- Index 1: Should the tool be invoked now? Consider whether all necessary information for the tool's invocation is ready, and whether the tool execution environment is correct. If it is a multi-round conversation, also consider the dependency relationships of the tool chain.
|
||||||
|
- Index 2: If should, is the chosen tool appropriate?
|
||||||
|
|
||||||
|
2. Input Quality (total 4 points. When evaluating, consider both the context and the tool schema):
|
||||||
|
- Index 3: Are all required parameters provided?
|
||||||
|
- Index 4: Are the input parameters valid and supported by the tool?
|
||||||
|
- Index 5: Are the input parameters in the correct format for their respective fields?
|
||||||
|
- Index 6: Does the value (content) of input parameter correctly reflect and match the given context?
|
||||||
|
|
||||||
|
3. Response Quality (total 4 points):
|
||||||
|
- Index 7: Does the response provide meaningful and useful information? Or are there any error messages or information that can be used as guidance for agent invoking tool better?
|
||||||
|
- Index 8: Does the response match the tool's intended purpose/function?
|
||||||
|
- Index 9: Does the response value correct (content appropriate) given the input parameters?
|
||||||
|
- Index 10: Does the response help accomplish the task within the given context?
|
||||||
|
|
||||||
|
Important:
|
||||||
|
1. Sometimes there is not enough information in the context or schema to make a complete evaluation. In such cases, make your best judgment based on the available information.
|
||||||
|
2. Some tools (commonly system tools such as mkdir, touch, echo, etc.) modify the external environment. Since these results cannot be obtained, they return "None" as the response. At this point, all the scores in the quality of the response should be obtained and should not be seen as a problem for the tool.
|
||||||
|
3. Evaluation independently from the success flag. The success_flag indicates whether the tool executed without technical errors. The evaluation should evaluate the quality of the tool invocation. A tool can execute successfully (Success Flag=1) but still produce low-quality or irrelevant responses, leading to a low evaluation score.
|
||||||
|
4. Sometimes an agent will execute multiple steps and invoke multiple tools to complete a task, but you only need to evaluate the use of one tool for one of the steps, not whether the final task is completed or not.
|
||||||
|
|
||||||
|
Answer Format:
|
||||||
|
Please provide your answer in the following JSON format:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"scores": [0,0,0,0,0,0,0,0,0,0],
|
||||||
|
"explanation": "A brief evaluation (2-3 sentences) explaining the quality of the tool invocation, based on your evaluation. Low-quality aspects need to be reified, especially the causes of tool invocation errors."
|
||||||
|
}
|
||||||
|
```
|
||||||
32
benchmark/toolmemory/summary_tool_memory_prompt.yaml
Normal file
32
benchmark/toolmemory/summary_tool_memory_prompt.yaml
Normal file
|
|
@ -0,0 +1,32 @@
|
||||||
|
prompt: |
|
||||||
|
You are an expert in analyzing tool usage patterns and generating practical usage guidance for agents.
|
||||||
|
|
||||||
|
Tool Information:
|
||||||
|
- Tool Name: {tool_name}
|
||||||
|
- Tool Schema: {tool_schema}
|
||||||
|
|
||||||
|
Recent Tool Invocation Experiences:
|
||||||
|
{experiences}
|
||||||
|
|
||||||
|
Important:
|
||||||
|
1. Assume the tool (tool schema) can't be changed, your task is to guide agent to use it better.
|
||||||
|
2. Your answer must be based on the information given, don't make it up. If not enough data, state "Not enough data to determine Core Function/Success Patterns/Common Issues/Best Practices."
|
||||||
|
3. Your answer will be used to guide the use of the tool in the future, so do not include content related to recent tool invocation experience such as "case #3" or "Call #2", but some values can be used as examples.
|
||||||
|
4. Pay attention to information not mentioned in the tool schema, such as the response upon successful tool invocation. It's also welcome to uncover insights, such as how tools can be used more effectively, and possible dependencies between tools. But if they aren't, don't make them up.
|
||||||
|
5. Finally, to avoid deriving incorrect guidance from individual invocation, check whether, if the agent follows the proposed guidance, it can perform better on all recent invocation histories. If not, revise the guidance until it can. Specifically:
|
||||||
|
- Don't write guidance in an absolute tone without a very deterministic message (meaning that all invocation histories are satisfied, otherwise it will result in failure).
|
||||||
|
- Sometimes there may be inconsistencies. Consider whether this is due to the context in which the tool is being used.
|
||||||
|
|
||||||
|
Your Task:
|
||||||
|
Based on the tool invocation history, generate a concise and logical tool usage guidance following this structure:
|
||||||
|
1. Core Function: What this tool does and when to use it.
|
||||||
|
2. Success Patterns: Parameter patterns and usage scenarios that work well.
|
||||||
|
3. Common Issues: Main pitfalls to avoid and why they fail.
|
||||||
|
4. Best Practices: 2-3 actionable recommendations.
|
||||||
|
|
||||||
|
Answer Format:
|
||||||
|
Provide a structured, concise guidance (max 200 words). Focus on actionable insights derived from actual usage data. Avoid generic advice and think step by step.
|
||||||
|
|
||||||
|
```txt
|
||||||
|
Your concise, data-driven tool usage guidance
|
||||||
|
```
|
||||||
234
benchmark/toolmemory/tool_memory.py
Normal file
234
benchmark/toolmemory/tool_memory.py
Normal file
|
|
@ -0,0 +1,234 @@
|
||||||
|
"""Official-style ReMe Tool Memory HTTP helpers.
|
||||||
|
|
||||||
|
Aligned with ReMe Tool Memory HTTP APIs (see ReMe cookbook
|
||||||
|
``use_tool_memory_demo.py`` and docs under ``docs/tool_memory/``):
|
||||||
|
|
||||||
|
- ``add_tool_call_result``
|
||||||
|
- ``summary_tool_memory``
|
||||||
|
- ``retrieve_tool_memory``
|
||||||
|
|
||||||
|
Response memories are read from ``metadata.memory_list[].content``.
|
||||||
|
This module does not use ExpG-only fields such as ``no_persist``,
|
||||||
|
``source_task``, or ``add_to``.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
DEFAULT_BASE_URL = "http://localhost:8002"
|
||||||
|
|
||||||
|
|
||||||
|
class ToolMemoryFetcher:
|
||||||
|
"""HTTP client for ReMe Tool Memory endpoints."""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
workspace_id: str,
|
||||||
|
base_url: str = DEFAULT_BASE_URL,
|
||||||
|
timeout: float = 60.0,
|
||||||
|
) -> None:
|
||||||
|
self.workspace_id = workspace_id
|
||||||
|
self.base_url = base_url.rstrip("/")
|
||||||
|
self.timeout = timeout
|
||||||
|
|
||||||
|
def _url(self, endpoint: str) -> str:
|
||||||
|
return f"{self.base_url}/{endpoint.lstrip('/')}"
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _join_tool_names(tool_names: List[str] | str) -> str:
|
||||||
|
if isinstance(tool_names, str):
|
||||||
|
return tool_names
|
||||||
|
return ",".join(tool_names)
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _memory_list(payload: Dict[str, Any]) -> List[Dict[str, Any]]:
|
||||||
|
metadata = payload.get("metadata") or {}
|
||||||
|
if not isinstance(metadata, dict):
|
||||||
|
return []
|
||||||
|
memory_list = metadata.get("memory_list") or []
|
||||||
|
return memory_list if isinstance(memory_list, list) else []
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def _content_by_tool(cls, payload: Dict[str, Any]) -> Dict[str, str]:
|
||||||
|
result: Dict[str, str] = {}
|
||||||
|
for memory in cls._memory_list(payload):
|
||||||
|
if not isinstance(memory, dict):
|
||||||
|
continue
|
||||||
|
tool_name = str(memory.get("when_to_use") or "").strip()
|
||||||
|
content = memory.get("content") or ""
|
||||||
|
if tool_name:
|
||||||
|
result[tool_name] = str(content)
|
||||||
|
return result
|
||||||
|
|
||||||
|
async def add_tool_call_result_async(
|
||||||
|
self,
|
||||||
|
tool_call_results: List[Dict[str, Any]],
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""Call ``add_tool_call_result``."""
|
||||||
|
async with httpx.AsyncClient() as client:
|
||||||
|
response = await client.post(
|
||||||
|
self._url("add_tool_call_result"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_call_results": tool_call_results,
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
async def summary_tool_memory_async(
|
||||||
|
self,
|
||||||
|
tool_names: List[str] | str,
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""Call ``summary_tool_memory``."""
|
||||||
|
async with httpx.AsyncClient() as client:
|
||||||
|
response = await client.post(
|
||||||
|
self._url("summary_tool_memory"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_names": self._join_tool_names(tool_names),
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
async def retrieve_tool_memory_async(
|
||||||
|
self,
|
||||||
|
tool_names: List[str] | str,
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""Call ``retrieve_tool_memory``."""
|
||||||
|
async with httpx.AsyncClient() as client:
|
||||||
|
response = await client.post(
|
||||||
|
self._url("retrieve_tool_memory"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_names": self._join_tool_names(tool_names),
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
async def collect_memory_async(
|
||||||
|
self,
|
||||||
|
tool_names: List[str],
|
||||||
|
) -> Dict[str, str]:
|
||||||
|
"""Summarize then retrieve guidance for tools.
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Mapping from tool name to memory ``content`` string.
|
||||||
|
"""
|
||||||
|
if not tool_names:
|
||||||
|
return {}
|
||||||
|
|
||||||
|
names = self._join_tool_names(tool_names)
|
||||||
|
try:
|
||||||
|
summary = await self.summary_tool_memory_async(names)
|
||||||
|
if not summary.get("success"):
|
||||||
|
logger.warning("summary_tool_memory failed for %s", names)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
logger.warning("summary_tool_memory error for %s: %s", names, exc)
|
||||||
|
|
||||||
|
try:
|
||||||
|
retrieved = await self.retrieve_tool_memory_async(names)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
logger.warning("retrieve_tool_memory error for %s: %s", names, exc)
|
||||||
|
return {}
|
||||||
|
|
||||||
|
if not retrieved.get("success"):
|
||||||
|
logger.warning("retrieve_tool_memory failed for %s", names)
|
||||||
|
return {}
|
||||||
|
|
||||||
|
return self._content_by_tool(retrieved)
|
||||||
|
|
||||||
|
def add_tool_call_result(
|
||||||
|
self,
|
||||||
|
tool_call_results: List[Dict[str, Any]],
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""Sync wrapper for ``add_tool_call_result``."""
|
||||||
|
with httpx.Client() as client:
|
||||||
|
response = client.post(
|
||||||
|
self._url("add_tool_call_result"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_call_results": tool_call_results,
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
def summary_tool_memory(self, tool_names: List[str] | str) -> Dict[str, Any]:
|
||||||
|
"""Sync wrapper for ``summary_tool_memory``."""
|
||||||
|
with httpx.Client() as client:
|
||||||
|
response = client.post(
|
||||||
|
self._url("summary_tool_memory"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_names": self._join_tool_names(tool_names),
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
def retrieve_tool_memory(self, tool_names: List[str] | str) -> Dict[str, Any]:
|
||||||
|
"""Sync wrapper for ``retrieve_tool_memory``."""
|
||||||
|
with httpx.Client() as client:
|
||||||
|
response = client.post(
|
||||||
|
self._url("retrieve_tool_memory"),
|
||||||
|
json={
|
||||||
|
"workspace_id": self.workspace_id,
|
||||||
|
"tool_names": self._join_tool_names(tool_names),
|
||||||
|
},
|
||||||
|
timeout=self.timeout,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
def collect_memory(self, tool_names: List[str]) -> Dict[str, str]:
|
||||||
|
"""Sync wrapper for summarize + retrieve.
|
||||||
|
|
||||||
|
Prefer ``collect_memory_async`` inside an existing event loop.
|
||||||
|
"""
|
||||||
|
if not tool_names:
|
||||||
|
return {}
|
||||||
|
|
||||||
|
names = self._join_tool_names(tool_names)
|
||||||
|
try:
|
||||||
|
summary = self.summary_tool_memory(names)
|
||||||
|
if not summary.get("success"):
|
||||||
|
logger.warning("summary_tool_memory failed for %s", names)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
logger.warning("summary_tool_memory error for %s: %s", names, exc)
|
||||||
|
|
||||||
|
try:
|
||||||
|
retrieved = self.retrieve_tool_memory(names)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
logger.warning("retrieve_tool_memory error for %s: %s", names, exc)
|
||||||
|
return {}
|
||||||
|
|
||||||
|
if not retrieved.get("success"):
|
||||||
|
logger.warning("retrieve_tool_memory failed for %s", names)
|
||||||
|
return {}
|
||||||
|
|
||||||
|
return self._content_by_tool(retrieved)
|
||||||
|
|
||||||
|
def get_memory_content(
|
||||||
|
self,
|
||||||
|
tool_names: List[str] | str,
|
||||||
|
) -> Optional[str]:
|
||||||
|
"""Retrieve and join memory contents for the given tools."""
|
||||||
|
payload = self.retrieve_tool_memory(tool_names)
|
||||||
|
if not payload.get("success"):
|
||||||
|
return None
|
||||||
|
contents = [content for content in self._content_by_tool(payload).values() if content]
|
||||||
|
return "\n\n".join(contents) if contents else None
|
||||||
45
benchmark/toolmemory/tool_memory_flows.yaml
Normal file
45
benchmark/toolmemory/tool_memory_flows.yaml
Normal file
|
|
@ -0,0 +1,45 @@
|
||||||
|
# Tool Memory flow / op config excerpt used by ExpG.
|
||||||
|
# Full runnable code: https://github.com/WangCan1178/ExpG
|
||||||
|
|
||||||
|
flow:
|
||||||
|
retrieve_tool_memory:
|
||||||
|
flow_content: retrieve_tool_memory_op
|
||||||
|
description: "Retrieves tool memories from the vector database based on tool names to provide tool usage patterns and best practices"
|
||||||
|
input_schema:
|
||||||
|
tool_names:
|
||||||
|
type: string
|
||||||
|
description: "Comma-separated tool names (e.g., 'tool_name1,tool_name2')"
|
||||||
|
required: true
|
||||||
|
|
||||||
|
add_tool_call_result:
|
||||||
|
flow_content: parse_tool_call_result_op >> update_vector_store_op
|
||||||
|
description: "Evaluates and adds tool call results to the tool memory database, creating new memory or updating existing memory for the specified tool"
|
||||||
|
input_schema:
|
||||||
|
tool_call_results:
|
||||||
|
type: array
|
||||||
|
description: "List of tool call result objects, each containing: tool_name, input, output, success, time_cost, token_cost, create_time"
|
||||||
|
required: true
|
||||||
|
|
||||||
|
summary_tool_memory:
|
||||||
|
flow_content: summary_tool_memory_op >> update_vector_store_op
|
||||||
|
description: "Analyzes tool call history and generates comprehensive usage patterns, best practices, and recommendations for the specified tools"
|
||||||
|
input_schema:
|
||||||
|
tool_names:
|
||||||
|
type: string
|
||||||
|
description: "Comma-separated tool names to summarize (e.g., 'tool_name1,tool_name2')"
|
||||||
|
required: true
|
||||||
|
|
||||||
|
op:
|
||||||
|
parse_tool_call_result_op:
|
||||||
|
backend: parse_tool_call_result_op
|
||||||
|
llm: default
|
||||||
|
params:
|
||||||
|
max_history_tool_call_cnt: 100
|
||||||
|
evaluation_sleep_interval: 1.0
|
||||||
|
|
||||||
|
summary_tool_memory_op:
|
||||||
|
backend: summary_tool_memory_op
|
||||||
|
llm: default
|
||||||
|
params:
|
||||||
|
data_from: '2025-09-10 10:56:58'
|
||||||
|
summary_sleep_interval: 1.0
|
||||||
18
deploy/docker/example.env
Normal file
18
deploy/docker/example.env
Normal file
|
|
@ -0,0 +1,18 @@
|
||||||
|
# Copy to the repository's .env for Docker Compose; keep real credentials private.
|
||||||
|
# File operations and BM25 search work without model credentials.
|
||||||
|
LLM_API_KEY=
|
||||||
|
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||||
|
LLM_MODEL_NAME=qwen3.7-plus
|
||||||
|
|
||||||
|
# Host settings: create this directory before starting Compose.
|
||||||
|
REME_DATA_DIR=./.reme
|
||||||
|
REME_PUBLISHED_PORT=2333
|
||||||
|
|
||||||
|
# On Linux, set these to the outputs of `id -u` and `id -g` so files remain yours.
|
||||||
|
REME_UID=1000
|
||||||
|
REME_GID=1000
|
||||||
|
|
||||||
|
# Optional container settings.
|
||||||
|
# REME_TIMEZONE=Asia/Shanghai
|
||||||
|
# REME_CONFIG=/etc/reme/config.yaml
|
||||||
|
# REME_IMAGE=ghcr.io/agentscope-ai/reme:main
|
||||||
100
deploy/docker/reme_container.py
Normal file
100
deploy/docker/reme_container.py
Normal file
|
|
@ -0,0 +1,100 @@
|
||||||
|
"""Container startup and health checks using ReMe's existing CLI contract."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from pathlib import Path
|
||||||
|
import sys
|
||||||
|
from urllib.request import ProxyHandler, Request, build_opener
|
||||||
|
|
||||||
|
HEALTH_STATE_PATH = Path("/tmp/reme-health.json")
|
||||||
|
_ENV_OVERRIDES = {
|
||||||
|
"REME_CONFIG": "config",
|
||||||
|
"REME_WORKSPACE_DIR": "workspace_dir",
|
||||||
|
"REME_HOST": "service.host",
|
||||||
|
"REME_PORT": "service.port",
|
||||||
|
"REME_TIMEZONE": "timezone",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def start_command(arguments: list[str], environment: dict[str, str]) -> list[str]:
|
||||||
|
"""Add container defaults only to `start`; explicit CLI arguments win."""
|
||||||
|
from reme.config import deep_merge_config, parse_kwargs
|
||||||
|
|
||||||
|
defaults = parse_kwargs(
|
||||||
|
"log_to_file=false",
|
||||||
|
*[f"{key}={json.dumps(environment[name])}" for name, key in _ENV_OVERRIDES.items() if environment.get(name)],
|
||||||
|
)
|
||||||
|
# Docker environment values are strings; the HTTP service expects an integer port.
|
||||||
|
if environment.get("REME_PORT"):
|
||||||
|
defaults["service"]["port"] = int(environment["REME_PORT"])
|
||||||
|
overrides = deep_merge_config(defaults, parse_kwargs(*arguments))
|
||||||
|
return ["reme", "start", *[f"{key}={json.dumps(value, ensure_ascii=False)}" for key, value in overrides.items()]]
|
||||||
|
|
||||||
|
|
||||||
|
def write_health_state(command: list[str], state_path: Path = HEALTH_STATE_PATH) -> None:
|
||||||
|
"""Record only the effective HTTP address, never credentials or user data."""
|
||||||
|
from reme.components.service.cli_service import prepare_start_config
|
||||||
|
from reme.config import parse_kwargs
|
||||||
|
from reme.constants import REME_DEFAULT_HOST, REME_DEFAULT_PORT, normalize_connect_host
|
||||||
|
from reme.plugin import resolve_plugin_runtime
|
||||||
|
|
||||||
|
config = resolve_plugin_runtime(prepare_start_config(parse_kwargs(*command[2:]))).config
|
||||||
|
service = config.get("service") or {}
|
||||||
|
if service.get("backend") != "http":
|
||||||
|
return
|
||||||
|
host = normalize_connect_host(service.get("host") or REME_DEFAULT_HOST)
|
||||||
|
if host == "::":
|
||||||
|
host = "::1"
|
||||||
|
port = int(service.get("port", REME_DEFAULT_PORT))
|
||||||
|
if not 1 <= port <= 65535:
|
||||||
|
raise ValueError("service.port must be between 1 and 65535")
|
||||||
|
host = f"[{host}]" if ":" in host else host
|
||||||
|
with state_path.open("w", encoding="utf-8") as state_file:
|
||||||
|
os.chmod(state_path, 0o600)
|
||||||
|
json.dump({"url": f"http://{host}:{port}/health_check"}, state_file)
|
||||||
|
|
||||||
|
|
||||||
|
def healthcheck(state_path: Path = HEALTH_STATE_PATH) -> int:
|
||||||
|
"""Require both a successful Job and a healthy component snapshot."""
|
||||||
|
try:
|
||||||
|
state = json.loads(state_path.read_text(encoding="utf-8"))
|
||||||
|
request = Request(state["url"], data=b"{}", headers={"Content-Type": "application/json"}, method="POST")
|
||||||
|
# A deployment's outbound proxy must not intercept its local probe.
|
||||||
|
with build_opener(ProxyHandler({})).open(request, timeout=4) as response:
|
||||||
|
payload = json.load(response)
|
||||||
|
healthy = payload.get("metadata", {}).get("health", {}).get("healthy")
|
||||||
|
if payload.get("success") is True and healthy is True:
|
||||||
|
return 0
|
||||||
|
except (OSError, ValueError, KeyError, TypeError, AttributeError):
|
||||||
|
pass
|
||||||
|
print("ReMe HTTP health check failed; inspect the container logs and POST /health_check", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
"""Prepare startup, then replace this process so signals reach ReMe."""
|
||||||
|
arguments = sys.argv[1:]
|
||||||
|
if arguments == ["--healthcheck"]:
|
||||||
|
return healthcheck()
|
||||||
|
|
||||||
|
HEALTH_STATE_PATH.unlink(missing_ok=True)
|
||||||
|
Path(os.environ.get("HOME", "/tmp/reme-home")).mkdir(parents=True, exist_ok=True)
|
||||||
|
if not arguments:
|
||||||
|
arguments = ["start"]
|
||||||
|
start_actions = {"start", "-start", "--start"}
|
||||||
|
if len(arguments) >= 2 and arguments[0] == "reme" and arguments[1] in start_actions:
|
||||||
|
arguments = arguments[1:]
|
||||||
|
if arguments[0] in start_actions:
|
||||||
|
from reme.utils import load_env
|
||||||
|
|
||||||
|
load_env()
|
||||||
|
arguments = start_command(arguments[1:], dict(os.environ))
|
||||||
|
write_health_state(arguments)
|
||||||
|
os.execvp(arguments[0], arguments)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
24
docker-compose.yml
Normal file
24
docker-compose.yml
Normal file
|
|
@ -0,0 +1,24 @@
|
||||||
|
services:
|
||||||
|
reme:
|
||||||
|
image: ${REME_IMAGE:-reme:local}
|
||||||
|
build:
|
||||||
|
context: .
|
||||||
|
init: false # The image already runs tini.
|
||||||
|
user: "${REME_UID:-1000}:${REME_GID:-1000}"
|
||||||
|
ports:
|
||||||
|
- "${REME_BIND_ADDRESS:-127.0.0.1}:${REME_PUBLISHED_PORT:-2333}:${REME_PORT:-2333}"
|
||||||
|
volumes:
|
||||||
|
- type: bind
|
||||||
|
source: ${REME_DATA_DIR:-./.reme}
|
||||||
|
target: /data
|
||||||
|
bind:
|
||||||
|
create_host_path: false
|
||||||
|
env_file:
|
||||||
|
- path: ${REME_ENV_FILE:-.env}
|
||||||
|
required: false
|
||||||
|
environment:
|
||||||
|
REME_HOST: 0.0.0.0
|
||||||
|
REME_PORT: ${REME_PORT:-2333}
|
||||||
|
REME_WORKSPACE_DIR: /data
|
||||||
|
restart: unless-stopped
|
||||||
|
stop_grace_period: 60s
|
||||||
396
docs/.vitepress/config.mts
Normal file
396
docs/.vitepress/config.mts
Normal file
|
|
@ -0,0 +1,396 @@
|
||||||
|
import fs from "node:fs";
|
||||||
|
import path from "node:path";
|
||||||
|
import { execFileSync } from "node:child_process";
|
||||||
|
import { fileURLToPath } from "node:url";
|
||||||
|
import { defineConfig, type DefaultTheme } from "vitepress";
|
||||||
|
import { legacyRoutes } from "./legacy-routes.mjs";
|
||||||
|
|
||||||
|
const sourceRoot = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..");
|
||||||
|
const repositoryRoot = path.resolve(sourceRoot, "../../..");
|
||||||
|
const repository = "https://github.com/agentscope-ai/ReMe";
|
||||||
|
const base = process.env.DOCS_BASE || "/";
|
||||||
|
|
||||||
|
function readSourceMap(): Record<string, string> {
|
||||||
|
try {
|
||||||
|
return JSON.parse(fs.readFileSync(path.join(sourceRoot, ".source-map.json"), "utf8"));
|
||||||
|
} catch {
|
||||||
|
return {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const sourceMap = readSourceMap();
|
||||||
|
|
||||||
|
function collectMarkdown(directory: string, root = directory): string[] {
|
||||||
|
const files: string[] = [];
|
||||||
|
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||||
|
if (entry.name.startsWith(".") || entry.name === "public" || entry.name === "figure") continue;
|
||||||
|
const absolute = path.join(directory, entry.name);
|
||||||
|
if (entry.isDirectory()) files.push(...collectMarkdown(absolute, root));
|
||||||
|
else if (entry.name.endsWith(".md")) files.push(path.relative(root, absolute).replaceAll(path.sep, "/"));
|
||||||
|
}
|
||||||
|
return files.sort();
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildLlmsFiles(outDir: string) {
|
||||||
|
const pages = collectMarkdown(sourceRoot);
|
||||||
|
const index = [
|
||||||
|
"# ReMe Documentation",
|
||||||
|
"",
|
||||||
|
"> Local-first, file-native memory for agents.",
|
||||||
|
"",
|
||||||
|
...pages.map((relativePath) => {
|
||||||
|
const source = fs.readFileSync(path.join(sourceRoot, relativePath), "utf8");
|
||||||
|
const title = source.match(/^#\s+(.+)$/m)?.[1]
|
||||||
|
|| source.match(/^title:\s*(.+)$/m)?.[1]
|
||||||
|
|| path.basename(relativePath, ".md");
|
||||||
|
const route = relativePath.replace(/(?:^|\/)index\.md$/, "").replace(/\.md$/, "");
|
||||||
|
return `- [${title}](https://reme.agentscope.io/${route})`;
|
||||||
|
}),
|
||||||
|
"",
|
||||||
|
];
|
||||||
|
fs.writeFileSync(path.join(outDir, "llms.txt"), index.join("\n"), "utf8");
|
||||||
|
|
||||||
|
const full = ["# ReMe Documentation", ""];
|
||||||
|
for (const relativePath of pages) {
|
||||||
|
const source = fs.readFileSync(path.join(sourceRoot, relativePath), "utf8");
|
||||||
|
full.push(`<!-- source: ${sourcePathFor(relativePath)} -->`, "", source, "", "---", "");
|
||||||
|
const pageDir = path.join(outDir, relativePath.replace(/\.md$/, ""));
|
||||||
|
fs.mkdirSync(pageDir, { recursive: true });
|
||||||
|
fs.writeFileSync(path.join(pageDir, "llms.txt"), source, "utf8");
|
||||||
|
}
|
||||||
|
fs.writeFileSync(path.join(outDir, "llms-full.txt"), full.join("\n"), "utf8");
|
||||||
|
}
|
||||||
|
|
||||||
|
function sourcePathFor(relativePath: string) {
|
||||||
|
return sourceMap[relativePath] || `docs/${relativePath}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function sourceLastUpdated(relativePath: string): number | undefined {
|
||||||
|
const sourcePath = sourcePathFor(relativePath);
|
||||||
|
try {
|
||||||
|
const timestamp = execFileSync("git", ["log", "-1", "--format=%ct", "--", sourcePath], {
|
||||||
|
cwd: repositoryRoot,
|
||||||
|
encoding: "utf8",
|
||||||
|
}).trim();
|
||||||
|
if (timestamp) return Number(timestamp) * 1000;
|
||||||
|
} catch {
|
||||||
|
// Fall back to the canonical file timestamp outside a Git checkout.
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
return fs.statSync(path.join(repositoryRoot, sourcePath)).mtimeMs;
|
||||||
|
} catch {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const legacyRedirectScript = `(() => {
|
||||||
|
const routes = ${JSON.stringify(legacyRoutes)};
|
||||||
|
const id = new URLSearchParams(window.location.search).get("doc");
|
||||||
|
const target = id && routes[id];
|
||||||
|
const base = ${JSON.stringify(base)};
|
||||||
|
if (target) {
|
||||||
|
const destination = /^https?:/.test(target)
|
||||||
|
? target
|
||||||
|
: base.replace(/\\/$/, "") + target;
|
||||||
|
window.location.replace(destination + window.location.hash);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const root = base.endsWith("/") ? base : base + "/";
|
||||||
|
if (window.location.pathname === root) {
|
||||||
|
window.location.replace(root + "zh/" + window.location.hash);
|
||||||
|
}
|
||||||
|
})();`;
|
||||||
|
|
||||||
|
function nav(language: "zh" | "en"): DefaultTheme.NavItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
return [
|
||||||
|
{ text: zh ? "首页" : "Home", link: `/${language}/` },
|
||||||
|
{ text: zh ? "文档" : "Docs", link: `/${language}/quick_start` },
|
||||||
|
{ text: zh ? "体验Studio" : "Try Studio", link: `/studio/?lang=${language}`, target: "_self" },
|
||||||
|
{ text: zh ? "集成" : "Integrations", link: `/${language}/integrations` },
|
||||||
|
{ text: zh ? "插件" : "Plugins", link: `/${language}/plugin_management` },
|
||||||
|
{ text: zh ? "评测" : "Benchmarks", link: `/${language}/benchmarks/longmemeval` },
|
||||||
|
{ text: zh ? "博客" : "Blog", link: `/${language}/reme-blog` },
|
||||||
|
{ text: zh ? "常见问题" : "FAQ", link: `/${language}/faq` },
|
||||||
|
];
|
||||||
|
}
|
||||||
|
|
||||||
|
function docsSidebar(language: "zh" | "en"): DefaultTheme.SidebarItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
text: zh ? "开始使用" : "Get Started",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "项目介绍" : "Introduction", link: `/${language}/overview` },
|
||||||
|
{ text: zh ? "快速开始" : "Quick Start", link: `/${language}/quick_start` },
|
||||||
|
{ text: zh ? "基础配置" : "Configuration", link: `/${language}/configuration` },
|
||||||
|
{ text: zh ? "服务与部署" : "Services and Deployment", link: `/${language}/services` },
|
||||||
|
{ text: zh ? "Docker 部署" : "Docker Deployment", link: `/${language}/docker` },
|
||||||
|
{ text: "ReMe Studio", link: `/${language}/workspace/studio` },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
text: zh ? "核心概念" : "Core Concepts",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "文件即记忆" : "Memory as File", link: `/${language}/memory_as_file` },
|
||||||
|
{ text: zh ? "记忆检索" : "Memory Search", link: `/${language}/memory_search` },
|
||||||
|
{ text: zh ? "自动关联" : "Auto Link", link: `/${language}/auto_link` },
|
||||||
|
{ text: zh ? "应用场景" : "Application Scenarios", link: `/${language}/reme_scene` },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
text: zh ? "记忆工作流" : "Memory Workflows",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: "Auto Memory", link: `/${language}/auto_memory` },
|
||||||
|
{ text: "Auto Resource", link: `/${language}/auto_resource` },
|
||||||
|
{ text: "Auto Dream", link: `/${language}/auto_dream` },
|
||||||
|
{ text: "Proactive", link: `/${language}/proactive` },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
text: zh ? "API 与运维" : "API and Operations",
|
||||||
|
collapsed: true,
|
||||||
|
items: [
|
||||||
|
{ text: "CLI", link: `/${language}/reference/cli` },
|
||||||
|
{ text: zh ? "Job API" : "Job API", link: `/${language}/reference/jobs` },
|
||||||
|
{ text: "HTTP / MCP", link: `/${language}/services#http-api` },
|
||||||
|
{ text: zh ? "诊断、备份与恢复" : "Diagnostics, Backup, and Recovery", link: `/${language}/operations` },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
text: zh ? "开发者" : "Development",
|
||||||
|
collapsed: true,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "代码框架" : "Framework", link: `/${language}/framework` },
|
||||||
|
{ text: zh ? "开源与贡献" : "Contributing", link: `/${language}/contributing` },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
];
|
||||||
|
}
|
||||||
|
|
||||||
|
function integrationsSidebar(language: "zh" | "en"): DefaultTheme.SidebarItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
return [{
|
||||||
|
text: zh ? "Agent 集成" : "Agent Integrations",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "集成总览" : "Overview", link: `/${language}/integrations` },
|
||||||
|
{ text: "Claude Code", link: `/${language}/integrations/claude-code` },
|
||||||
|
{ text: "Hermes Agent", link: `/${language}/integrations/hermes` },
|
||||||
|
{ text: "DeepSeek Harness", link: `/${language}/integrations/dsh` },
|
||||||
|
{ text: "OpenClaw", link: `/${language}/integrations/openclaw` },
|
||||||
|
],
|
||||||
|
}];
|
||||||
|
}
|
||||||
|
|
||||||
|
function pluginsSidebar(language: "zh" | "en"): DefaultTheme.SidebarItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
return [{
|
||||||
|
text: zh ? "插件" : "Plugins",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "插件管理" : "Plugin Management", link: `/${language}/plugin_management` },
|
||||||
|
{ text: zh ? "插件开发" : "Plugin Development", link: `/${language}/plugin_development` },
|
||||||
|
{ text: zh ? "每日论文" : "Daily Paper", link: `/${language}/plugins/daily-paper` },
|
||||||
|
{ text: "Auto Fin", link: `/${language}/plugins/auto-fin` },
|
||||||
|
{ text: "LME", link: `/${language}/plugins/lme` },
|
||||||
|
{ text: "BEAM", link: `/${language}/plugins/beam` },
|
||||||
|
],
|
||||||
|
}];
|
||||||
|
}
|
||||||
|
|
||||||
|
function benchmarksSidebar(language: "zh" | "en"): DefaultTheme.SidebarItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
return [{
|
||||||
|
text: zh ? "记忆能力评测" : "Memory Benchmarks",
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: "LongMemEval", link: `/${language}/benchmarks/longmemeval` },
|
||||||
|
{ text: "BEAM", link: `/${language}/benchmarks/beam` },
|
||||||
|
{ text: "π-Bench", link: `/${language}/benchmarks/pibench` },
|
||||||
|
{ text: "Tool Memory / ExpG", link: `/${language}/benchmarks/toolmemory` },
|
||||||
|
],
|
||||||
|
}];
|
||||||
|
}
|
||||||
|
|
||||||
|
function singlePageSidebar(language: "zh" | "en", page: "blog" | "faq"): DefaultTheme.SidebarItem[] {
|
||||||
|
const zh = language === "zh";
|
||||||
|
if (page === "blog") {
|
||||||
|
return [{
|
||||||
|
text: zh ? "ReMe 博客" : "ReMe Blog",
|
||||||
|
link: `/${language}/reme-blog`,
|
||||||
|
collapsed: false,
|
||||||
|
items: [
|
||||||
|
{ text: zh ? "ReMe介绍" : "About ReMe", link: `/${language}/reme-blog` },
|
||||||
|
{ text: zh ? "记忆标签" : "Memory Tags", link: `/${language}/blog_20260920` },
|
||||||
|
],
|
||||||
|
}];
|
||||||
|
}
|
||||||
|
return [{
|
||||||
|
text: zh ? "帮助" : "Help",
|
||||||
|
collapsed: false,
|
||||||
|
items: [{
|
||||||
|
text: zh ? "常见问题" : "Frequently Asked Questions",
|
||||||
|
link: `/${language}/faq`,
|
||||||
|
}],
|
||||||
|
}];
|
||||||
|
}
|
||||||
|
|
||||||
|
function sidebars(language: "zh" | "en"): DefaultTheme.SidebarMulti {
|
||||||
|
return {
|
||||||
|
[`/${language}/integrations`]: integrationsSidebar(language),
|
||||||
|
[`/${language}/workspace/`]: [],
|
||||||
|
[`/${language}/plugins/`]: pluginsSidebar(language),
|
||||||
|
[`/${language}/plugin_management`]: pluginsSidebar(language),
|
||||||
|
[`/${language}/plugin_development`]: pluginsSidebar(language),
|
||||||
|
[`/${language}/benchmarks/`]: benchmarksSidebar(language),
|
||||||
|
[`/${language}/reme-blog`]: singlePageSidebar(language, "blog"),
|
||||||
|
[`/${language}/blog_20260920`]: singlePageSidebar(language, "blog"),
|
||||||
|
[`/${language}/faq`]: singlePageSidebar(language, "faq"),
|
||||||
|
[`/${language}/`]: docsSidebar(language),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function configureRepositoryLinks(md: any) {
|
||||||
|
for (const ruleName of ["link_open", "image"] as const) {
|
||||||
|
const original = md.renderer.rules[ruleName];
|
||||||
|
md.renderer.rules[ruleName] = (tokens: any[], index: number, options: any, env: any, self: any) => {
|
||||||
|
const attribute = ruleName === "image" ? "src" : "href";
|
||||||
|
const token = tokens[index];
|
||||||
|
const attributeIndex = token.attrIndex(attribute);
|
||||||
|
const target = attributeIndex >= 0 ? token.attrs[attributeIndex][1] : "";
|
||||||
|
if (target && !/^(?:[a-z]+:|#|\/)/i.test(target)) {
|
||||||
|
const cleanTarget = target.split("#")[0].split("?")[0];
|
||||||
|
const generatedTarget = path.resolve(sourceRoot, path.dirname(env.relativePath), cleanTarget);
|
||||||
|
if (!fs.existsSync(generatedTarget)) {
|
||||||
|
const originalPage = sourcePathFor(env.relativePath);
|
||||||
|
const originalTarget = path.posix.normalize(path.posix.join(path.posix.dirname(originalPage), cleanTarget));
|
||||||
|
const suffix = target.slice(cleanTarget.length);
|
||||||
|
const url = ruleName === "image"
|
||||||
|
? `https://raw.githubusercontent.com/agentscope-ai/ReMe/main/${originalTarget}${suffix}`
|
||||||
|
: `${repository}/blob/main/${originalTarget}${suffix}`;
|
||||||
|
token.attrs[attributeIndex][1] = url;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return original ? original(tokens, index, options, env, self) : self.renderToken(tokens, index, options);
|
||||||
|
};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export default defineConfig({
|
||||||
|
lang: "zh-CN",
|
||||||
|
title: "ReMe",
|
||||||
|
description: "Local-first, file-native memory for agents",
|
||||||
|
base,
|
||||||
|
cleanUrls: true,
|
||||||
|
lastUpdated: true,
|
||||||
|
ignoreDeadLinks: [/^http:\/\/localhost(?::\d+)?(?:\/|$)/],
|
||||||
|
sitemap: {
|
||||||
|
hostname: "https://reme.agentscope.io",
|
||||||
|
transformItems(items) {
|
||||||
|
const isRoot = (url: string) => url.replace(/^\/+|\/+$/g, "") === "";
|
||||||
|
return items.filter((item) => !isRoot(item.url)).map((item) => {
|
||||||
|
const route = item.url.replace(/^\/+/, "");
|
||||||
|
const relativePath = !route || route.endsWith("/") ? `${route}index.md` : `${route}.md`;
|
||||||
|
const links = item.links?.filter((link) => !isRoot(link.url));
|
||||||
|
return { ...item, links, lastmod: sourceLastUpdated(relativePath) };
|
||||||
|
});
|
||||||
|
},
|
||||||
|
},
|
||||||
|
head: [
|
||||||
|
["link", { rel: "icon", type: "image/svg+xml", href: `${base}reme-icon.svg` }],
|
||||||
|
["meta", { name: "theme-color", content: "#087f6a", media: "(prefers-color-scheme: light)" }],
|
||||||
|
["meta", { name: "theme-color", content: "#0d1512", media: "(prefers-color-scheme: dark)" }],
|
||||||
|
["script", {
|
||||||
|
defer: "",
|
||||||
|
src: "https://cloud.umami.is/script.js",
|
||||||
|
"data-website-id": "8cafe9df-d883-4046-b5e9-36dfd21a4884",
|
||||||
|
"data-domains": "reme.agentscope.io",
|
||||||
|
}],
|
||||||
|
["script", {}, legacyRedirectScript],
|
||||||
|
],
|
||||||
|
markdown: {
|
||||||
|
config: configureRepositoryLinks,
|
||||||
|
},
|
||||||
|
transformPageData(pageData, { siteConfig }) {
|
||||||
|
const sourcePath = path.join(siteConfig.srcDir, pageData.relativePath);
|
||||||
|
pageData.frontmatter._sourcePath = sourcePathFor(pageData.relativePath);
|
||||||
|
pageData.lastUpdated = sourceLastUpdated(pageData.relativePath);
|
||||||
|
try {
|
||||||
|
pageData.frontmatter._rawMarkdown = fs.readFileSync(sourcePath, "utf8");
|
||||||
|
} catch {
|
||||||
|
pageData.frontmatter._rawMarkdown = "";
|
||||||
|
}
|
||||||
|
},
|
||||||
|
buildEnd(siteConfig) {
|
||||||
|
buildLlmsFiles(siteConfig.outDir);
|
||||||
|
},
|
||||||
|
themeConfig: {
|
||||||
|
logo: "/reme-icon.svg",
|
||||||
|
siteTitle: "ReMe",
|
||||||
|
nav: [
|
||||||
|
...nav("zh"),
|
||||||
|
{
|
||||||
|
text: "语言",
|
||||||
|
items: [
|
||||||
|
{ text: "简体中文", link: "/zh/" },
|
||||||
|
{ text: "English", link: "/en/" },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
],
|
||||||
|
outline: { label: "页面导航", level: [2, 3] },
|
||||||
|
search: {
|
||||||
|
provider: "local",
|
||||||
|
options: {
|
||||||
|
locales: {
|
||||||
|
zh: {
|
||||||
|
translations: {
|
||||||
|
button: { buttonText: "搜索文档", buttonAriaLabel: "搜索文档" },
|
||||||
|
modal: {
|
||||||
|
noResultsText: "没有找到相关内容",
|
||||||
|
resetButtonTitle: "清除查询",
|
||||||
|
footer: { selectText: "选择", navigateText: "切换", closeText: "关闭" },
|
||||||
|
},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
socialLinks: [{ icon: "github", link: repository }],
|
||||||
|
footer: {
|
||||||
|
message: "Released under the Apache-2.0 License.",
|
||||||
|
copyright: "Copyright ReMe contributors",
|
||||||
|
},
|
||||||
|
},
|
||||||
|
locales: {
|
||||||
|
zh: {
|
||||||
|
label: "简体中文",
|
||||||
|
lang: "zh-CN",
|
||||||
|
link: "/zh/",
|
||||||
|
themeConfig: {
|
||||||
|
nav: nav("zh"),
|
||||||
|
sidebar: sidebars("zh"),
|
||||||
|
outline: { label: "页面导航", level: [2, 3] },
|
||||||
|
docFooter: { prev: "上一页", next: "下一页" },
|
||||||
|
darkModeSwitchLabel: "外观",
|
||||||
|
sidebarMenuLabel: "菜单",
|
||||||
|
returnToTopLabel: "返回顶部",
|
||||||
|
langMenuLabel: "切换语言",
|
||||||
|
},
|
||||||
|
},
|
||||||
|
en: {
|
||||||
|
label: "English",
|
||||||
|
lang: "en-US",
|
||||||
|
link: "/en/",
|
||||||
|
themeConfig: {
|
||||||
|
nav: nav("en"),
|
||||||
|
sidebar: sidebars("en"),
|
||||||
|
outline: { label: "On this page", level: [2, 3] },
|
||||||
|
docFooter: { prev: "Previous page", next: "Next page" },
|
||||||
|
},
|
||||||
|
},
|
||||||
|
},
|
||||||
|
});
|
||||||
47
docs/.vitepress/legacy-routes.mjs
Normal file
47
docs/.vitepress/legacy-routes.mjs
Normal file
|
|
@ -0,0 +1,47 @@
|
||||||
|
export const legacyRoutes = {
|
||||||
|
"readme-zh": "/zh/",
|
||||||
|
"readme-en": "/en/",
|
||||||
|
"zh-quick_start": "/zh/quick_start",
|
||||||
|
"en-quick_start": "/en/quick_start",
|
||||||
|
"zh-plugin_management": "/zh/plugin_management",
|
||||||
|
"en-plugin_management": "/en/plugin_management",
|
||||||
|
"zh-memory_as_file": "/zh/memory_as_file",
|
||||||
|
"en-memory_as_file": "/en/memory_as_file",
|
||||||
|
"zh-memory_search": "/zh/memory_search",
|
||||||
|
"en-memory_search": "/en/memory_search",
|
||||||
|
"zh-auto_memory": "/zh/auto_memory",
|
||||||
|
"en-auto_memory": "/en/auto_memory",
|
||||||
|
"zh-auto_resource": "/zh/auto_resource",
|
||||||
|
"en-auto_resource": "/en/auto_resource",
|
||||||
|
"zh-auto_link": "/zh/auto_link",
|
||||||
|
"en-auto_link": "/en/auto_link",
|
||||||
|
"zh-auto_dream": "/zh/auto_dream",
|
||||||
|
"en-auto_dream": "/en/auto_dream",
|
||||||
|
"zh-proactive": "/zh/proactive",
|
||||||
|
"en-proactive": "/en/proactive",
|
||||||
|
"zh-reme_scene": "/zh/reme_scene",
|
||||||
|
"en-reme_scene": "/en/reme_scene",
|
||||||
|
"zh-framework": "/zh/framework",
|
||||||
|
"en-framework": "/en/framework",
|
||||||
|
"zh-reme-blog": "/zh/reme-blog",
|
||||||
|
"en-reme-blog": "/en/reme-blog",
|
||||||
|
"zh-contributing": "/zh/contributing",
|
||||||
|
"en-contributing": "/en/contributing",
|
||||||
|
"typescript-zh": "/zh/integrations",
|
||||||
|
"typescript-en": "/en/integrations",
|
||||||
|
"studio-zh": "/zh/workspace/studio",
|
||||||
|
"studio-en": "/en/workspace/studio",
|
||||||
|
"daily-paper-zh": "/zh/plugins/daily-paper",
|
||||||
|
"daily-paper-en": "/en/plugins/daily-paper",
|
||||||
|
"auto-fin-zh": "/zh/plugins/auto-fin",
|
||||||
|
"auto-fin-en": "/en/plugins/auto-fin",
|
||||||
|
"beam-zh": "/zh/benchmarks/beam",
|
||||||
|
"beam-en": "/en/benchmarks/beam",
|
||||||
|
"longmemeval-zh": "/zh/benchmarks/longmemeval",
|
||||||
|
"longmemeval-en": "/en/benchmarks/longmemeval",
|
||||||
|
"pibench-zh": "/zh/benchmarks/pibench",
|
||||||
|
"pibench-en": "/en/benchmarks/pibench",
|
||||||
|
"toolmemory-zh": "/zh/benchmarks/toolmemory",
|
||||||
|
"toolmemory-en": "/en/benchmarks/toolmemory",
|
||||||
|
"agents-guide": "https://github.com/agentscope-ai/ReMe/blob/main/AGENTS.md",
|
||||||
|
};
|
||||||
32
docs/.vitepress/theme/CopyMarkdownButton.vue
Normal file
32
docs/.vitepress/theme/CopyMarkdownButton.vue
Normal file
|
|
@ -0,0 +1,32 @@
|
||||||
|
<script setup lang="ts">
|
||||||
|
import { computed, ref } from "vue";
|
||||||
|
import { useData } from "vitepress";
|
||||||
|
|
||||||
|
const { frontmatter, lang } = useData();
|
||||||
|
const copied = ref(false);
|
||||||
|
const label = computed(() => {
|
||||||
|
if (copied.value) return lang.value.startsWith("zh") ? "已复制" : "Copied";
|
||||||
|
return lang.value.startsWith("zh") ? "复制 Markdown" : "Copy Markdown";
|
||||||
|
});
|
||||||
|
|
||||||
|
async function copyMarkdown() {
|
||||||
|
const markdown = String(frontmatter.value._rawMarkdown || "");
|
||||||
|
if (!markdown) return;
|
||||||
|
await navigator.clipboard.writeText(markdown);
|
||||||
|
copied.value = true;
|
||||||
|
window.setTimeout(() => { copied.value = false; }, 1800);
|
||||||
|
}
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<template>
|
||||||
|
<div class="copy-markdown-wrap">
|
||||||
|
<button class="copy-markdown" type="button" :class="{ copied }" @click="copyMarkdown">
|
||||||
|
<svg v-if="!copied" viewBox="0 0 24 24" aria-hidden="true">
|
||||||
|
<rect x="9" y="9" width="13" height="13" rx="2" />
|
||||||
|
<path d="M5 15H4a2 2 0 0 1-2-2V4a2 2 0 0 1 2-2h9a2 2 0 0 1 2 2v1" />
|
||||||
|
</svg>
|
||||||
|
<svg v-else viewBox="0 0 24 24" aria-hidden="true"><path d="m5 12 4 4L19 6" /></svg>
|
||||||
|
{{ label }}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
523
docs/.vitepress/theme/HomePage.vue
Normal file
523
docs/.vitepress/theme/HomePage.vue
Normal file
|
|
@ -0,0 +1,523 @@
|
||||||
|
<script setup lang="ts">
|
||||||
|
import { computed, onMounted, reactive } from "vue";
|
||||||
|
import { useData, withBase } from "vitepress";
|
||||||
|
|
||||||
|
const props = defineProps<{ lang: "zh" | "en" }>();
|
||||||
|
|
||||||
|
const repository = "https://github.com/agentscope-ai/ReMe";
|
||||||
|
const trafficShareBase = "https://cloud.umami.is/analytics/us/share/S1OZK1PSDLEpyiU5?date=30day&page=1";
|
||||||
|
const { isDark } = useData();
|
||||||
|
const stats = reactive({ stars: "3.4K+", forks: "293" });
|
||||||
|
|
||||||
|
const translations = {
|
||||||
|
zh: {
|
||||||
|
eyebrow: "LOCAL-FIRST · FILE-NATIVE",
|
||||||
|
title: "让 Agent 真正记住,\n让记忆始终属于你。",
|
||||||
|
lead: "ReMe 将对话、资料与经验沉淀为可读、可编辑、可检索、相互链接的本地文件,让不同 Agent 共享同一套长期记忆。",
|
||||||
|
quickStart: "快速开始",
|
||||||
|
learnMore: "了解 ReMe",
|
||||||
|
stars: "GitHub Stars",
|
||||||
|
forks: "Forks",
|
||||||
|
mapLabel: "OPEN ECOSYSTEM",
|
||||||
|
mapTitle: "ReMe 与开源生态",
|
||||||
|
capabilityTitle: "核心能力",
|
||||||
|
capabilities: [
|
||||||
|
{ mark: "▣", title: "文件即记忆", href: "/zh/memory_as_file", tone: "mint" },
|
||||||
|
{ mark: "✦", title: "自动记忆", href: "/zh/auto_memory", tone: "cyan" },
|
||||||
|
{ mark: "⌕", title: "混合检索", href: "/zh/memory_search", tone: "blue" },
|
||||||
|
{ mark: "⌁", title: "图谱关联", href: "/zh/auto_link", tone: "amber" },
|
||||||
|
],
|
||||||
|
agentTitle: "ReMe 接入 Agent",
|
||||||
|
agentGuide: "接入指南",
|
||||||
|
backendTitle: "检索引擎接入 ReMe",
|
||||||
|
backendGuide: "配置指南",
|
||||||
|
backendNote: "可选向量索引",
|
||||||
|
integrations: [
|
||||||
|
{ logo: "/ecosystem/qwenpaw.png", title: "QwenPaw", href: "https://github.com/agentscope-ai/QwenPaw", tone: "blue" },
|
||||||
|
{ logo: "/ecosystem/deepseek-harness.svg", title: "DeepSeek Harness", href: "https://github.com/deepseek-ai/deepseek-harness", tone: "mint" },
|
||||||
|
{ logo: "/ecosystem/openclaw.svg", title: "OpenClaw", href: "https://github.com/openclaw/openclaw", tone: "cyan" },
|
||||||
|
{ logo: "/ecosystem/claude-code.png", title: "Claude Code", href: "https://github.com/anthropics/claude-code", tone: "violet" },
|
||||||
|
{ logo: "/ecosystem/hermes.svg", title: "Hermes Agent", href: "https://github.com/NousResearch/hermes-agent", tone: "amber" },
|
||||||
|
],
|
||||||
|
backends: [
|
||||||
|
{ logo: "/ecosystem/zvec.ico", title: "Zvec", href: "https://github.com/alibaba/zvec", tone: "mint" },
|
||||||
|
{ logo: "/ecosystem/faiss.png", title: "FAISS", href: "https://github.com/facebookresearch/faiss", tone: "blue" },
|
||||||
|
],
|
||||||
|
benchmarkLabel: "02 / BENCHMARKS",
|
||||||
|
benchmarkTitle: "用真实评测,\n验证长期记忆",
|
||||||
|
benchmarkLead: "从跨会话检索到百万级上下文,ReMe 用可复现的公开基准验证长期记忆。",
|
||||||
|
benchmarkAction: "查看全部评测",
|
||||||
|
benchmarkNote: "仓库已发布参考结果 · Agentic score",
|
||||||
|
benchmarks: [
|
||||||
|
{ name: "LongMemEval", setting: "500 题 · cleaned-s", score: 89.4 },
|
||||||
|
{ name: "BEAM 100K", setting: "20 cases · 400 题", score: 66.1 },
|
||||||
|
{ name: "BEAM 1M", setting: "35 cases · 700 题", score: 65.0 },
|
||||||
|
],
|
||||||
|
piLabel: "π-Bench 主动性",
|
||||||
|
piDetail: "5 类用户画像的平均 PROC 得分",
|
||||||
|
piDelta: "较 NanoBot +2.4%",
|
||||||
|
sectionLabel: "03 / PRODUCTS & PLUGINS",
|
||||||
|
sectionTitle: "从记忆工作区,到自动研究",
|
||||||
|
sectionLead: "三个完整入口,把 ReMe 用到真实工作流中。",
|
||||||
|
products: [
|
||||||
|
{ mark: "▣", label: "WORKSPACE", title: "ReMe Studio", detail: "在本地 Web 工作区中浏览、编辑、搜索记忆,并探索 wikilink 图谱。", href: "/studio/?lang=zh", tone: "mint" },
|
||||||
|
{ mark: "◌", label: "DISCOVER", title: "Daily Paper", detail: "筛选值得阅读的论文,分析 PDF,并生成文件化笔记与五分钟简报。", href: "/zh/plugins/daily-paper", tone: "cyan" },
|
||||||
|
{ mark: "↗", label: "RESEARCH", title: "Auto Fin", detail: "连接最新财联社新闻与本地历史记忆,生成带 wikilink 的研究报告。", href: "/zh/plugins/auto-fin", tone: "amber" },
|
||||||
|
],
|
||||||
|
trafficLabel: "04 / OPEN METRICS",
|
||||||
|
trafficTitle: "公开、透明的访问趋势",
|
||||||
|
trafficDetail: "最近 30 天的页面浏览量与访问趋势,由 Umami 提供匿名统计。",
|
||||||
|
trafficAction: "打开完整数据页",
|
||||||
|
trafficFrameTitle: "ReMe 最近 30 天访问数据",
|
||||||
|
},
|
||||||
|
en: {
|
||||||
|
eyebrow: "LOCAL-FIRST · FILE-NATIVE",
|
||||||
|
title: "Memory for AI agents.\nFiles that remain yours.",
|
||||||
|
lead: "ReMe turns conversations, resources, and experience into readable, editable, searchable, interconnected local files—a shared long-term memory layer for every agent.",
|
||||||
|
quickStart: "Quick Start",
|
||||||
|
learnMore: "Meet ReMe",
|
||||||
|
stars: "GitHub Stars",
|
||||||
|
forks: "Forks",
|
||||||
|
mapLabel: "OPEN ECOSYSTEM",
|
||||||
|
mapTitle: "ReMe and the open ecosystem",
|
||||||
|
capabilityTitle: "CORE CAPABILITIES",
|
||||||
|
capabilities: [
|
||||||
|
{ mark: "▣", title: "Memory as files", href: "/en/memory_as_file", tone: "mint" },
|
||||||
|
{ mark: "✦", title: "Auto memory", href: "/en/auto_memory", tone: "cyan" },
|
||||||
|
{ mark: "⌕", title: "Hybrid search", href: "/en/memory_search", tone: "blue" },
|
||||||
|
{ mark: "⌁", title: "Linked graph", href: "/en/auto_link", tone: "amber" },
|
||||||
|
],
|
||||||
|
agentTitle: "ReMe for agents",
|
||||||
|
agentGuide: "Integration guide",
|
||||||
|
backendTitle: "Retrieval for ReMe",
|
||||||
|
backendGuide: "Configuration guide",
|
||||||
|
backendNote: "Optional vector indexes",
|
||||||
|
integrations: [
|
||||||
|
{ logo: "/ecosystem/qwenpaw.png", title: "QwenPaw", href: "https://github.com/agentscope-ai/QwenPaw", tone: "blue" },
|
||||||
|
{ logo: "/ecosystem/deepseek-harness.svg", title: "DeepSeek Harness", href: "https://github.com/deepseek-ai/deepseek-harness", tone: "mint" },
|
||||||
|
{ logo: "/ecosystem/openclaw.svg", title: "OpenClaw", href: "https://github.com/openclaw/openclaw", tone: "cyan" },
|
||||||
|
{ logo: "/ecosystem/claude-code.png", title: "Claude Code", href: "https://github.com/anthropics/claude-code", tone: "violet" },
|
||||||
|
{ logo: "/ecosystem/hermes.svg", title: "Hermes Agent", href: "https://github.com/NousResearch/hermes-agent", tone: "amber" },
|
||||||
|
],
|
||||||
|
backends: [
|
||||||
|
{ logo: "/ecosystem/zvec.ico", title: "Zvec", href: "https://github.com/alibaba/zvec", tone: "mint" },
|
||||||
|
{ logo: "/ecosystem/faiss.png", title: "FAISS", href: "https://github.com/facebookresearch/faiss", tone: "blue" },
|
||||||
|
],
|
||||||
|
benchmarkLabel: "02 / BENCHMARKS",
|
||||||
|
benchmarkTitle: "Memory that holds up\nunder pressure",
|
||||||
|
benchmarkLead: "From cross-session retrieval to million-token context, ReMe validates long-term memory with reproducible public benchmarks.",
|
||||||
|
benchmarkAction: "Explore all benchmarks",
|
||||||
|
benchmarkNote: "Published reference runs · Agentic score",
|
||||||
|
benchmarks: [
|
||||||
|
{ name: "LongMemEval", setting: "500 questions · cleaned-s", score: 89.4 },
|
||||||
|
{ name: "BEAM 100K", setting: "20 cases · 400 questions", score: 66.1 },
|
||||||
|
{ name: "BEAM 1M", setting: "35 cases · 700 questions", score: 65.0 },
|
||||||
|
],
|
||||||
|
piLabel: "π-Bench proactivity",
|
||||||
|
piDetail: "Average PROC score across five personas",
|
||||||
|
piDelta: "+2.4% over NanoBot",
|
||||||
|
sectionLabel: "03 / PRODUCTS & PLUGINS",
|
||||||
|
sectionTitle: "From memory workspace to automated research",
|
||||||
|
sectionLead: "Three complete paths for putting ReMe into real workflows.",
|
||||||
|
products: [
|
||||||
|
{ mark: "▣", label: "WORKSPACE", title: "ReMe Studio", detail: "Browse, edit, and search memory in a local web workspace, then explore its wikilink graph.", href: "/studio/?lang=en", tone: "mint" },
|
||||||
|
{ mark: "◌", label: "DISCOVER", title: "Daily Paper", detail: "Select useful papers, analyze PDFs, and create file-native notes plus a five-minute brief.", href: "/en/plugins/daily-paper", tone: "cyan" },
|
||||||
|
{ mark: "↗", label: "RESEARCH", title: "Auto Fin", detail: "Connect recent CLS news with local memory to create traceable, wikilink-backed reports.", href: "/en/plugins/auto-fin", tone: "amber" },
|
||||||
|
],
|
||||||
|
trafficLabel: "04 / OPEN METRICS",
|
||||||
|
trafficTitle: "Public, transparent traffic",
|
||||||
|
trafficDetail: "Page views and traffic trends from the last 30 days, measured anonymously with Umami.",
|
||||||
|
trafficAction: "Open the full report",
|
||||||
|
trafficFrameTitle: "ReMe traffic for the last 30 days",
|
||||||
|
},
|
||||||
|
} as const;
|
||||||
|
|
||||||
|
const text = computed(() => translations[props.lang]);
|
||||||
|
const trafficShareUrl = computed(() => `${trafficShareBase}&theme=${isDark.value ? "dark" : "light"}`);
|
||||||
|
const localLink = (href: string) => withBase(href);
|
||||||
|
|
||||||
|
function normalizeCompactCount(value: string) {
|
||||||
|
const normalized = value.trim().toUpperCase();
|
||||||
|
return /[KMB]$/.test(normalized) ? `${normalized}+` : normalized;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function readBadge(metric: "stars" | "forks") {
|
||||||
|
const response = await fetch(`https://img.shields.io/github/${metric}/agentscope-ai/ReMe.json`);
|
||||||
|
if (!response.ok) throw new Error(`Unable to load ${metric}`);
|
||||||
|
const payload = await response.json();
|
||||||
|
return normalizeCompactCount(String(payload.message || payload.value || ""));
|
||||||
|
}
|
||||||
|
|
||||||
|
onMounted(async () => {
|
||||||
|
const [stars, forks] = await Promise.allSettled([readBadge("stars"), readBadge("forks")]);
|
||||||
|
if (stars.status === "fulfilled" && stars.value) stats.stars = stars.value;
|
||||||
|
if (forks.status === "fulfilled" && forks.value) stats.forks = forks.value;
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
<template>
|
||||||
|
<div class="reme-home" :class="{ 'is-zh': lang === 'zh' }">
|
||||||
|
<section class="home-stage">
|
||||||
|
<div class="hero-copy">
|
||||||
|
<p class="eyebrow">{{ text.eyebrow }}</p>
|
||||||
|
<h1>{{ text.title }}</h1>
|
||||||
|
<p class="hero-lead">{{ text.lead }}</p>
|
||||||
|
|
||||||
|
<div class="hero-actions">
|
||||||
|
<a class="action primary" :href="localLink(`/${lang}/quick_start`)">{{ text.quickStart }} <span>→</span></a>
|
||||||
|
<a class="action secondary" :href="localLink(`/${lang}/overview`)">{{ text.learnMore }} <span>↗</span></a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="repo-stats" aria-live="polite">
|
||||||
|
<a :href="`${repository}/stargazers`" target="_blank" rel="noreferrer" :aria-label="`${stats.stars} ${text.stars}`">
|
||||||
|
<span class="stat-icon">☆</span>
|
||||||
|
<span><strong>{{ stats.stars }}</strong><small>{{ text.stars }}</small></span>
|
||||||
|
</a>
|
||||||
|
<a :href="`${repository}/forks`" target="_blank" rel="noreferrer" :aria-label="`${stats.forks} ${text.forks}`">
|
||||||
|
<span class="stat-icon fork-icon">⑂</span>
|
||||||
|
<span><strong>{{ stats.forks }}</strong><small>{{ text.forks }}</small></span>
|
||||||
|
</a>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="ecosystem-map">
|
||||||
|
<div class="map-heading">
|
||||||
|
<div>
|
||||||
|
<span>{{ text.mapLabel }}</span>
|
||||||
|
<strong>{{ text.mapTitle }}</strong>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="ecosystem-network">
|
||||||
|
<div class="network-brands">
|
||||||
|
<div class="network-group-label">
|
||||||
|
<strong>{{ text.agentTitle }}</strong>
|
||||||
|
<a :href="localLink(`/${lang}/integrations`)">{{ text.agentGuide }} ↗</a>
|
||||||
|
</div>
|
||||||
|
<div class="brand-viewport">
|
||||||
|
<div class="brand-reel">
|
||||||
|
<a
|
||||||
|
v-for="integration in text.integrations"
|
||||||
|
:key="integration.title"
|
||||||
|
class="brand-link"
|
||||||
|
:class="integration.tone"
|
||||||
|
:href="integration.href"
|
||||||
|
target="_blank"
|
||||||
|
rel="noopener noreferrer"
|
||||||
|
:aria-label="`${integration.title} GitHub`"
|
||||||
|
>
|
||||||
|
<span class="brand-mark" aria-hidden="true"><img :src="localLink(integration.logo)" alt="" /></span>
|
||||||
|
<strong>{{ integration.title }}</strong>
|
||||||
|
</a>
|
||||||
|
<div v-for="integration in text.integrations" :key="`${integration.title}-clone`" class="brand-link reel-clone" :class="integration.tone" aria-hidden="true">
|
||||||
|
<span class="brand-mark"><img :src="localLink(integration.logo)" alt="" /></span>
|
||||||
|
<strong>{{ integration.title }}</strong>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="network-group-label backend-label">
|
||||||
|
<strong>{{ text.backendTitle }}</strong>
|
||||||
|
<a :href="localLink(`/${lang}/memory_search#${lang === 'zh' ? '向量索引后端' : 'vector-index-backends'}`)">{{ text.backendGuide }} ↗</a>
|
||||||
|
</div>
|
||||||
|
<div class="backend-links">
|
||||||
|
<a v-for="backend in text.backends" :key="backend.title" class="brand-link" :class="backend.tone" :href="backend.href" target="_blank" rel="noopener noreferrer" :aria-label="`${backend.title} GitHub`">
|
||||||
|
<span class="brand-mark" aria-hidden="true"><img :src="localLink(backend.logo)" alt="" /></span>
|
||||||
|
<strong>{{ backend.title }}</strong>
|
||||||
|
</a>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="network-center" aria-hidden="true">
|
||||||
|
<span class="network-ring"></span>
|
||||||
|
<span class="network-core"><img :src="localLink('/reme-icon.svg')" alt="" /></span>
|
||||||
|
<strong>ReMe</strong>
|
||||||
|
</div>
|
||||||
|
<div class="network-capabilities">
|
||||||
|
<strong class="capability-label">{{ text.capabilityTitle }}</strong>
|
||||||
|
<a v-for="capability in text.capabilities" :key="capability.title" class="capability-link" :class="capability.tone" :href="localLink(capability.href)">
|
||||||
|
<span aria-hidden="true">{{ capability.mark }}</span>
|
||||||
|
<strong>{{ capability.title }}</strong>
|
||||||
|
<span aria-hidden="true">↗</span>
|
||||||
|
</a>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="benchmark-section page-panel">
|
||||||
|
<div class="benchmark-intro">
|
||||||
|
<p class="section-label">{{ text.benchmarkLabel }}</p>
|
||||||
|
<h2>{{ text.benchmarkTitle }}</h2>
|
||||||
|
<p>{{ text.benchmarkLead }}</p>
|
||||||
|
<a :href="localLink(`/${lang}/benchmarks/longmemeval`)">{{ text.benchmarkAction }} <span>→</span></a>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="benchmark-board">
|
||||||
|
<div class="benchmark-board-head">
|
||||||
|
<span>{{ text.benchmarkNote }}</span>
|
||||||
|
<span>0—100%</span>
|
||||||
|
</div>
|
||||||
|
<div class="benchmark-chart">
|
||||||
|
<div v-for="benchmark in text.benchmarks" :key="benchmark.name" class="benchmark-row">
|
||||||
|
<div class="benchmark-name">
|
||||||
|
<strong>{{ benchmark.name }}</strong>
|
||||||
|
<small>{{ benchmark.setting }}</small>
|
||||||
|
</div>
|
||||||
|
<div class="benchmark-track" aria-hidden="true">
|
||||||
|
<span :style="{ width: `${benchmark.score}%` }"></span>
|
||||||
|
</div>
|
||||||
|
<strong class="benchmark-score">{{ benchmark.score.toFixed(1) }}%</strong>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<a class="pi-score" :href="localLink(`/${lang}/benchmarks/pibench`)" :aria-label="`${text.piLabel}: 0.580`">
|
||||||
|
<span class="pi-symbol">π</span>
|
||||||
|
<span>
|
||||||
|
<small>{{ text.piLabel }}</small>
|
||||||
|
<strong>0.580</strong>
|
||||||
|
<em>{{ text.piDetail }}</em>
|
||||||
|
</span>
|
||||||
|
<b>{{ text.piDelta }} ↗</b>
|
||||||
|
</a>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="product-section">
|
||||||
|
<div class="section-heading">
|
||||||
|
<p class="section-label">{{ text.sectionLabel }}</p>
|
||||||
|
<div>
|
||||||
|
<h2>{{ text.sectionTitle }}</h2>
|
||||||
|
<p>{{ text.sectionLead }}</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="product-grid">
|
||||||
|
<a
|
||||||
|
v-for="product in text.products"
|
||||||
|
:key="product.title"
|
||||||
|
class="product-card"
|
||||||
|
:class="product.tone"
|
||||||
|
:href="localLink(product.href)"
|
||||||
|
:target="product.href.startsWith('/studio/') ? '_self' : undefined"
|
||||||
|
>
|
||||||
|
<span class="product-mark">{{ product.mark }}</span>
|
||||||
|
<span class="product-label">{{ product.label }}</span>
|
||||||
|
<strong>{{ product.title }}</strong>
|
||||||
|
<p>{{ product.detail }}</p>
|
||||||
|
<span class="product-arrow">→</span>
|
||||||
|
</a>
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="traffic-section page-panel">
|
||||||
|
<div class="traffic-heading">
|
||||||
|
<p class="section-label">{{ text.trafficLabel }}</p>
|
||||||
|
<h2>{{ text.trafficTitle }}</h2>
|
||||||
|
<p>{{ text.trafficDetail }}</p>
|
||||||
|
<a :href="localLink(`/${lang}/traffic`)">{{ text.trafficAction }} <span>→</span></a>
|
||||||
|
</div>
|
||||||
|
<div class="traffic-window">
|
||||||
|
<div class="traffic-window-bar" aria-hidden="true">
|
||||||
|
<span></span><span></span><span></span><b>reme.agentscope.io · 30 days</b>
|
||||||
|
</div>
|
||||||
|
<iframe :src="trafficShareUrl" :title="text.trafficFrameTitle" loading="lazy" referrerpolicy="no-referrer" />
|
||||||
|
</div>
|
||||||
|
</section>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<style scoped>
|
||||||
|
.reme-home {
|
||||||
|
--home-ink: #17231e;
|
||||||
|
--home-muted: #66736d;
|
||||||
|
--home-line: #d8e2dc;
|
||||||
|
--home-accent: #087f6a;
|
||||||
|
--home-surface: #ffffff;
|
||||||
|
--home-surface-soft: #f7f9f7;
|
||||||
|
--home-glass: rgba(255, 255, 255, 0.72);
|
||||||
|
--home-tile: rgba(255, 255, 255, 0.88);
|
||||||
|
--home-primary-bg: #17241e;
|
||||||
|
--home-primary-text: #ffffff;
|
||||||
|
--home-shadow: rgba(28, 57, 45, 0.12);
|
||||||
|
--section-light: #f8faf8;
|
||||||
|
--section-tint: #edf4f1;
|
||||||
|
max-width: 1720px;
|
||||||
|
margin: 0 auto;
|
||||||
|
padding: 0 clamp(24px, 4.5vw, 72px) 80px;
|
||||||
|
color: var(--home-ink);
|
||||||
|
}
|
||||||
|
.home-stage {
|
||||||
|
position: relative;
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: minmax(0, 1fr) 650px;
|
||||||
|
gap: clamp(42px, 4vw, 68px);
|
||||||
|
align-items: center;
|
||||||
|
min-height: calc(100vh - 64px);
|
||||||
|
padding: 72px 0 82px;
|
||||||
|
}
|
||||||
|
.home-stage::before {
|
||||||
|
position: absolute;
|
||||||
|
z-index: -1;
|
||||||
|
inset: 0 calc(50% - 50vw);
|
||||||
|
background:
|
||||||
|
radial-gradient(ellipse 70% 105% at -8% 18%, rgba(21, 158, 126, 0.14), transparent 72%),
|
||||||
|
radial-gradient(ellipse 68% 105% at 108% 10%, rgba(77, 103, 211, 0.13), transparent 73%),
|
||||||
|
linear-gradient(115deg, #f7fbf8 0%, #fbfaf6 49%, #f7f8fd 100%);
|
||||||
|
content: "";
|
||||||
|
}
|
||||||
|
.eyebrow, .section-label { margin: 0; color: var(--home-accent); font: 750 13px/1.4 var(--vp-font-family-mono); letter-spacing: 0.16em; }
|
||||||
|
.hero-copy h1 { max-width: 100%; margin: 23px 0 0; color: var(--home-ink); font: 760 clamp(52px, 4.2vw, 76px)/1.04 Georgia, "Times New Roman", serif; white-space: pre-wrap; letter-spacing: -0.052em; }
|
||||||
|
.is-zh .hero-copy h1 { max-width: 760px; font-size: clamp(52px, 3.6vw, 64px); white-space: pre-line; word-break: keep-all; }
|
||||||
|
.hero-lead { max-width: 650px; margin: 28px 0 0; color: var(--home-muted); font-size: clamp(17px, 1.3vw, 20px); line-height: 1.75; }
|
||||||
|
.hero-actions { display: flex; flex-wrap: wrap; gap: 12px; margin-top: 34px; }
|
||||||
|
.action { display: inline-flex; align-items: center; justify-content: space-between; gap: 28px; min-width: 166px; min-height: 54px; padding: 0 19px; border: 1px solid var(--home-line); border-radius: 12px; color: var(--home-ink); background: var(--home-glass); text-decoration: none; font-weight: 720; box-shadow: 0 8px 22px color-mix(in srgb, var(--home-shadow) 50%, transparent); transition: transform 160ms ease, box-shadow 160ms ease; }
|
||||||
|
.action.primary { border-color: var(--home-primary-bg); color: var(--home-primary-text); background: var(--home-primary-bg); box-shadow: 0 12px 26px color-mix(in srgb, var(--home-primary-bg) 28%, transparent); }
|
||||||
|
.action:hover { transform: translateY(-2px); box-shadow: 0 15px 28px var(--home-shadow); }
|
||||||
|
.repo-stats { display: flex; flex-wrap: wrap; gap: 34px; margin-top: 40px; }
|
||||||
|
.repo-stats a { display: flex; gap: 12px; align-items: flex-start; color: inherit; text-decoration: none; }
|
||||||
|
.stat-icon { color: var(--home-accent); font-size: 30px; line-height: 1; }
|
||||||
|
.fork-icon { transform: rotate(90deg); }
|
||||||
|
.repo-stats strong { display: block; font: 740 28px/1 var(--vp-font-family-mono); letter-spacing: -0.04em; }
|
||||||
|
.repo-stats small { display: block; margin-top: 8px; color: var(--home-muted); font-size: 13px; }
|
||||||
|
.ecosystem-map { position: relative; min-width: 0; }
|
||||||
|
.map-heading { margin-bottom: 25px; }
|
||||||
|
.map-heading > div { display: flex; min-width: 0; flex-direction: column; gap: 8px; }
|
||||||
|
.map-heading span { color: var(--home-accent); font: 700 10px/1.4 var(--vp-font-family-mono); letter-spacing: 0.12em; }
|
||||||
|
.map-heading strong { font-size: 18px; line-height: 1.3; white-space: nowrap; }
|
||||||
|
.ecosystem-network { position: relative; display: grid; grid-template-columns: minmax(0, 1.15fr) minmax(100px, 0.62fr) minmax(0, 1fr); gap: 12px; align-items: center; min-height: 370px; }
|
||||||
|
.ecosystem-network::before, .ecosystem-network::after { position: absolute; z-index: 0; top: 50%; width: 19%; border-top: 1px dashed color-mix(in srgb, var(--home-accent) 58%, var(--home-line)); content: ""; }
|
||||||
|
.ecosystem-network::before { left: 30%; }.ecosystem-network::after { right: 27%; }
|
||||||
|
.network-brands, .network-center, .network-capabilities { position: relative; z-index: 1; min-width: 0; }
|
||||||
|
.network-group-label { display: flex; align-items: baseline; justify-content: space-between; gap: 5px; margin-bottom: 7px; }
|
||||||
|
.network-group-label strong, .capability-label { color: var(--home-muted); font: 750 10px/1.3 var(--vp-font-family-mono); letter-spacing: 0.04em; }
|
||||||
|
.network-group-label a { flex: none; color: var(--home-accent); font-size: 10px; font-weight: 700; text-decoration: none; white-space: nowrap; }
|
||||||
|
.network-group-label a:hover { text-decoration: underline; }
|
||||||
|
.brand-viewport { height: 184px; overflow: hidden; mask-image: linear-gradient(transparent, #000 12%, #000 88%, transparent); }
|
||||||
|
.brand-reel { display: grid; grid-auto-rows: 46px; gap: 6px; animation: brand-scroll 19s linear infinite; }
|
||||||
|
.brand-viewport:hover .brand-reel { animation-play-state: paused; }
|
||||||
|
.brand-viewport:focus-within { overflow-y: auto; mask-image: none; }
|
||||||
|
.brand-viewport:focus-within .brand-reel { animation: none; }
|
||||||
|
.brand-link { display: flex; min-width: 0; height: 46px; align-items: center; gap: 7px; padding: 5px; border: 1px solid color-mix(in srgb, var(--card-accent) 26%, var(--home-line)); border-radius: 9px; color: var(--home-ink); background: var(--home-tile); text-decoration: none; transition: border-color 160ms ease, transform 160ms ease; }
|
||||||
|
.brand-link:hover { border-color: var(--card-accent); transform: translateX(2px); }
|
||||||
|
.brand-mark { display: grid; width: 32px; height: 32px; flex: none; place-items: center; overflow: hidden; border: 1px solid var(--home-line); border-radius: 7px; background: #fff; }
|
||||||
|
.brand-mark img { width: 27px; height: 27px; margin: 0; object-fit: contain; }
|
||||||
|
.brand-link strong { min-width: 0; font-size: 13px; line-height: 1.15; overflow-wrap: anywhere; }
|
||||||
|
.reel-clone { pointer-events: none; }
|
||||||
|
.backend-label { margin-top: 13px; }
|
||||||
|
.backend-links { display: grid; grid-template-columns: repeat(2, minmax(0, 1fr)); gap: 5px; }
|
||||||
|
.backend-links .brand-link { height: 42px; gap: 5px; }
|
||||||
|
.backend-links .brand-mark { width: 28px; height: 28px; }
|
||||||
|
.backend-links .brand-mark img { width: 23px; height: 23px; }
|
||||||
|
.network-center { display: flex; min-height: 164px; flex-direction: column; align-items: center; justify-content: center; gap: 16px; }
|
||||||
|
.network-ring { position: absolute; top: 50%; left: 50%; width: 108px; height: 108px; border: 1px solid color-mix(in srgb, var(--home-accent) 30%, var(--home-line)); border-radius: 50%; transform: translate(-50%, -64%); animation: hub-pulse 3.6s ease-in-out infinite; }
|
||||||
|
.network-ring::after { position: absolute; inset: 10px; border: 1px solid var(--home-line); border-radius: 50%; content: ""; }
|
||||||
|
.network-core { z-index: 1; display: grid; width: 64px; height: 64px; place-items: center; border: 1px solid var(--home-line); border-radius: 50%; background: var(--home-surface); box-shadow: 0 10px 28px var(--home-shadow); }
|
||||||
|
.network-core img { width: 42px; height: 42px; margin: 0; }
|
||||||
|
.network-center strong { z-index: 1; color: var(--home-accent); font: 750 12px/1 var(--vp-font-family-mono); }
|
||||||
|
.network-capabilities { display: grid; gap: 8px; }
|
||||||
|
.capability-label { display: block; margin-bottom: 2px; }
|
||||||
|
.capability-link { display: flex; min-height: 53px; align-items: center; gap: 7px; padding: 7px; border: 1px solid color-mix(in srgb, var(--card-accent) 30%, var(--home-line)); border-radius: 10px; color: var(--home-ink); background: radial-gradient(circle at 100% 0, color-mix(in srgb, var(--card-accent) 11%, transparent), transparent 70%), var(--home-tile); text-decoration: none; transition: transform 160ms ease, border-color 160ms ease; }
|
||||||
|
.capability-link:hover { border-color: var(--card-accent); transform: translateX(2px); }
|
||||||
|
.capability-link span:first-child { display: grid; width: 27px; height: 27px; flex: none; place-items: center; border-radius: 7px; color: var(--card-accent); background: color-mix(in srgb, var(--card-accent) 12%, var(--home-surface)); font-size: 16px; }
|
||||||
|
.capability-link strong { min-width: 0; font-size: 13px; line-height: 1.2; }
|
||||||
|
.capability-link span:last-child { margin-left: auto; color: var(--card-accent); font-size: 13px; }
|
||||||
|
.brand-link:focus-visible, .capability-link:focus-visible, .network-group-label a:focus-visible { outline: 2px solid var(--home-accent); outline-offset: 2px; }
|
||||||
|
.mint { --card-accent: #059b7f; }.cyan { --card-accent: #169cc4; }.blue { --card-accent: #536bd8; }.amber { --card-accent: #c48324; }.violet { --card-accent: #7654c2; }
|
||||||
|
@keyframes brand-scroll { to { transform: translateY(-260px); } }
|
||||||
|
@keyframes hub-pulse { 50% { transform: translate(-50%, -64%) scale(1.12); opacity: 0.55; } }
|
||||||
|
.page-panel { position: relative; min-height: calc(100vh - 64px); }
|
||||||
|
.benchmark-section { display: grid; grid-template-columns: minmax(310px, 0.72fr) minmax(560px, 1.28fr); gap: clamp(54px, 7vw, 110px); align-items: center; padding: 112px 0 120px; color: var(--home-ink); }
|
||||||
|
.benchmark-section::before { position: absolute; z-index: -1; inset: 0 calc(50% - 50vw); border-top: 1px solid var(--home-line); background: var(--section-tint); content: ""; }
|
||||||
|
.benchmark-intro .section-label { color: var(--home-accent); }
|
||||||
|
.benchmark-intro h2 { max-width: 620px; margin: 23px 0 0; color: var(--home-ink); font-size: clamp(42px, 4.2vw, 68px); line-height: 1.06; white-space: pre-line; letter-spacing: -0.052em; }
|
||||||
|
.is-zh .benchmark-intro h2 { word-break: keep-all; }
|
||||||
|
.benchmark-intro > p:not(.section-label) { max-width: 560px; margin: 24px 0 0; color: var(--home-muted); font-size: 17px; line-height: 1.75; }
|
||||||
|
.benchmark-intro > a, .traffic-heading > a { display: inline-flex; align-items: center; gap: 32px; min-height: 50px; margin-top: 34px; padding: 0 18px; border: 1px solid color-mix(in srgb, var(--home-ink) 24%, transparent); border-radius: 11px; color: var(--home-ink); text-decoration: none; font-weight: 720; transition: background 160ms ease, transform 160ms ease; }
|
||||||
|
.benchmark-intro > a:hover, .traffic-heading > a:hover { background: color-mix(in srgb, var(--home-ink) 6%, transparent); transform: translateY(-2px); }
|
||||||
|
.benchmark-board { padding: clamp(24px, 3vw, 38px); border: 1px solid rgba(255, 255, 255, 0.14); border-radius: 26px; color: white; background: #15352b; box-shadow: 0 28px 64px rgba(24, 58, 45, 0.2); }
|
||||||
|
.benchmark-board-head { display: flex; justify-content: space-between; gap: 20px; padding-bottom: 22px; border-bottom: 1px solid rgba(255, 255, 255, 0.13); color: #94aca3; font: 700 11px/1.4 var(--vp-font-family-mono); letter-spacing: 0.1em; text-transform: uppercase; }
|
||||||
|
.benchmark-chart { display: grid; gap: 28px; padding: 32px 0; }
|
||||||
|
.benchmark-row { display: grid; grid-template-columns: 150px minmax(140px, 1fr) 72px; gap: 20px; align-items: center; }
|
||||||
|
.benchmark-name strong { display: block; color: white; font-size: 16px; }
|
||||||
|
.benchmark-name small { display: block; margin-top: 5px; color: #8fa69d; font-size: 12px; }
|
||||||
|
.benchmark-track { height: 10px; overflow: hidden; border-radius: 99px; background: rgba(255, 255, 255, 0.09); }
|
||||||
|
.benchmark-track span { display: block; height: 100%; border-radius: inherit; background: linear-gradient(90deg, #38d3ae, #85ead2); box-shadow: 0 0 22px rgba(66, 220, 182, 0.28); }
|
||||||
|
.benchmark-score { color: white; font: 740 20px/1 var(--vp-font-family-mono); text-align: right; }
|
||||||
|
.pi-score { display: grid; grid-template-columns: 58px minmax(0, 1fr) auto; gap: 18px; align-items: center; padding: 20px; border: 1px solid rgba(142, 160, 255, 0.27); border-radius: 18px; color: white; background: linear-gradient(110deg, rgba(82, 105, 216, 0.23), rgba(82, 105, 216, 0.08)); text-decoration: none; }
|
||||||
|
.pi-symbol { display: grid; place-items: center; width: 58px; height: 58px; border-radius: 15px; color: #b9c6ff; background: rgba(107, 129, 236, 0.18); font: 700 31px/1 Georgia, serif; }
|
||||||
|
.pi-score small, .pi-score em { display: block; color: #aabbb4; font-size: 12px; font-style: normal; }
|
||||||
|
.pi-score strong { display: block; margin: 4px 0; font: 750 25px/1 var(--vp-font-family-mono); }
|
||||||
|
.pi-score b { color: #a9b7ff; font-size: 13px; white-space: nowrap; }
|
||||||
|
.product-section { position: relative; display: flex; min-height: calc(100vh - 64px); flex-direction: column; justify-content: center; padding: 112px 0 120px; }
|
||||||
|
.product-section::before { position: absolute; z-index: -1; inset: 0 calc(50% - 50vw); border-top: 1px solid var(--home-line); background: var(--section-light); content: ""; }
|
||||||
|
.section-heading { display: grid; grid-template-columns: minmax(190px, 0.38fr) minmax(0, 1fr); gap: 40px; align-items: start; margin-bottom: 36px; }
|
||||||
|
.section-heading h2, .traffic-heading h2 { max-width: 820px; margin: 0; color: var(--home-ink); font-size: clamp(38px, 3.6vw, 58px); line-height: 1.1; letter-spacing: -0.045em; }
|
||||||
|
.section-heading p:not(.section-label) { margin: 15px 0 0; color: var(--home-muted); font-size: 16px; }
|
||||||
|
.product-grid { display: grid; grid-template-columns: repeat(3, minmax(0, 1fr)); gap: 18px; }
|
||||||
|
.product-card { position: relative; display: flex; min-height: 360px; flex-direction: column; padding: 30px; overflow: hidden; border: 1px solid color-mix(in srgb, var(--card-accent) 30%, var(--home-line)); border-radius: 22px; color: var(--home-ink); background: radial-gradient(circle at 80% 0, color-mix(in srgb, var(--card-accent) 14%, transparent), transparent 50%), linear-gradient(150deg, var(--home-surface), color-mix(in srgb, var(--card-accent) 5%, var(--home-surface-soft))); text-decoration: none; transition: transform 170ms ease, box-shadow 170ms ease; }
|
||||||
|
.product-card:hover { transform: translateY(-4px); box-shadow: 0 20px 38px color-mix(in srgb, var(--card-accent) 14%, transparent); }
|
||||||
|
.product-mark { display: grid; place-items: center; width: 46px; height: 46px; border-radius: 13px; color: var(--card-accent); background: color-mix(in srgb, var(--card-accent) 13%, var(--home-surface)); font: 650 24px/1 var(--vp-font-family-mono); }
|
||||||
|
.product-label { margin-top: 50px; color: var(--card-accent); font: 720 11px/1.3 var(--vp-font-family-mono); letter-spacing: 0.12em; }
|
||||||
|
.product-card strong { margin-top: 10px; font-size: 23px; }
|
||||||
|
.product-card p { max-width: 390px; margin: 13px 0 0; color: var(--home-muted); font-size: 15px; line-height: 1.65; }
|
||||||
|
.product-arrow { position: absolute; right: 23px; bottom: 20px; color: var(--card-accent); font-size: 22px; }
|
||||||
|
.traffic-section { display: grid; grid-template-columns: minmax(300px, 0.64fr) minmax(600px, 1.36fr); gap: clamp(48px, 6vw, 94px); align-items: center; padding: 112px 0 116px; color: var(--home-ink); }
|
||||||
|
.traffic-section::before { position: absolute; z-index: -1; inset: 0 calc(50% - 50vw); border-top: 1px solid var(--home-line); background: var(--section-tint); content: ""; }
|
||||||
|
.traffic-heading .section-label { color: var(--home-accent); }
|
||||||
|
.traffic-heading h2 { margin-top: 22px; color: var(--home-ink); }
|
||||||
|
.traffic-heading > p:not(.section-label) { max-width: 460px; margin: 22px 0 0; color: var(--home-muted); font-size: 17px; line-height: 1.7; }
|
||||||
|
.traffic-window { overflow: hidden; height: min(650px, calc(100vh - 170px)); min-height: 520px; border: 1px solid rgba(255, 255, 255, 0.2); border-radius: 24px; background: white; box-shadow: 0 34px 80px rgba(0, 0, 0, 0.3); }
|
||||||
|
.traffic-window-bar { display: flex; align-items: center; gap: 8px; height: 46px; padding: 0 16px; border-bottom: 1px solid #e5e8e7; background: #f6f8f7; }
|
||||||
|
.traffic-window-bar span { width: 9px; height: 9px; border-radius: 50%; background: #a8b4af; }
|
||||||
|
.traffic-window-bar span:first-child { background: #f08d78; }
|
||||||
|
.traffic-window-bar span:nth-child(2) { background: #e6bf67; }
|
||||||
|
.traffic-window-bar span:nth-child(3) { background: #69bd9a; }
|
||||||
|
.traffic-window-bar b { margin-left: 8px; color: #7b8983; font: 650 11px/1 var(--vp-font-family-mono); }
|
||||||
|
.traffic-window iframe { width: 100%; height: calc(100% - 46px); border: 0; }
|
||||||
|
:global(html.dark .reme-home) { --home-ink: #edf7f3; --home-muted: #a8bbb3; --home-line: #2d4038; --home-accent: #57dfc3; --home-surface: #14201b; --home-surface-soft: #101a16; --home-glass: rgba(17, 28, 23, 0.78); --home-tile: rgba(20, 32, 27, 0.92); --home-primary-bg: #57dfc3; --home-primary-text: #07120e; --home-shadow: rgba(0, 0, 0, 0.3); --section-light: #0d1512; --section-tint: #14201b; color-scheme: dark; }
|
||||||
|
:global(html.dark .home-stage::before) { background: radial-gradient(ellipse 70% 105% at -8% 18%, rgba(24, 169, 143, 0.13), transparent 72%), radial-gradient(ellipse 68% 105% at 108% 10%, rgba(74, 100, 218, 0.15), transparent 73%), linear-gradient(115deg, #0d1713 0%, #101713 49%, #10131c 100%); }
|
||||||
|
:global(html.dark .benchmark-board) { background: #0b1712; box-shadow: 0 28px 64px rgba(0, 0, 0, 0.32); }
|
||||||
|
@media (max-width: 1680px) {
|
||||||
|
.home-stage { grid-template-columns: 1fr; min-height: auto; }
|
||||||
|
.hero-copy { max-width: 800px; padding-top: 26px; }
|
||||||
|
.hero-copy h1 { white-space: pre-wrap; }
|
||||||
|
.ecosystem-map { max-width: 850px; }
|
||||||
|
}
|
||||||
|
@media (max-width: 1320px) {
|
||||||
|
.benchmark-section, .traffic-section { grid-template-columns: 1fr; min-height: auto; }
|
||||||
|
.benchmark-intro, .traffic-heading { max-width: 720px; }
|
||||||
|
.traffic-window { width: 100%; max-width: 1000px; }
|
||||||
|
}
|
||||||
|
@media (max-width: 1000px) {
|
||||||
|
.product-grid { grid-template-columns: repeat(2, minmax(0, 1fr)); }
|
||||||
|
}
|
||||||
|
@media (max-width: 700px) {
|
||||||
|
.reme-home { padding-right: 20px; padding-left: 20px; }
|
||||||
|
.home-stage { gap: 42px; padding: 52px 0 62px; }
|
||||||
|
.hero-copy h1 { font-size: clamp(43px, 13vw, 62px); }
|
||||||
|
.hero-lead { font-size: 16px; }
|
||||||
|
.map-heading strong { white-space: normal; }
|
||||||
|
.section-heading { grid-template-columns: 1fr; gap: 18px; }
|
||||||
|
.product-grid { grid-template-columns: 1fr; }
|
||||||
|
.product-card { min-height: 280px; }
|
||||||
|
.benchmark-section, .product-section, .traffic-section { padding-top: 78px; padding-bottom: 84px; }
|
||||||
|
.benchmark-row { grid-template-columns: minmax(0, 1fr) auto; gap: 12px; }
|
||||||
|
.benchmark-track { grid-column: 1 / -1; grid-row: 2; }
|
||||||
|
.benchmark-score { grid-column: 2; grid-row: 1; }
|
||||||
|
.pi-score { grid-template-columns: 48px minmax(0, 1fr); }
|
||||||
|
.pi-symbol { width: 48px; height: 48px; }
|
||||||
|
.pi-score b { grid-column: 2; }
|
||||||
|
.traffic-window { height: 660px; min-height: 0; border-radius: 17px; }
|
||||||
|
}
|
||||||
|
@media (max-width: 520px) {
|
||||||
|
.ecosystem-network { grid-template-columns: minmax(0, 1.1fr) 62px minmax(0, 1fr); gap: 4px; }
|
||||||
|
.network-group-label { flex-wrap: wrap; }
|
||||||
|
.brand-link strong, .capability-link strong { font-size: 10px; }
|
||||||
|
.network-ring { width: 78px; height: 78px; }
|
||||||
|
.network-core { width: 52px; height: 52px; }
|
||||||
|
.network-core img { width: 34px; height: 34px; }
|
||||||
|
.capability-link { gap: 3px; padding: 5px; }
|
||||||
|
.capability-link span:first-child { width: 22px; height: 22px; font-size: 13px; }
|
||||||
|
.backend-links { grid-template-columns: 1fr; }
|
||||||
|
}
|
||||||
|
@media (prefers-reduced-motion: reduce) {
|
||||||
|
.brand-viewport { height: auto; mask-image: none; }
|
||||||
|
.brand-reel, .network-ring { animation: none; }
|
||||||
|
.reel-clone { display: none; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
15
docs/.vitepress/theme/SourceLink.vue
Normal file
15
docs/.vitepress/theme/SourceLink.vue
Normal file
|
|
@ -0,0 +1,15 @@
|
||||||
|
<script setup lang="ts">
|
||||||
|
import { computed } from "vue";
|
||||||
|
import { useData } from "vitepress";
|
||||||
|
|
||||||
|
const { frontmatter, lang } = useData();
|
||||||
|
const sourcePath = computed(() => String(frontmatter.value._sourcePath || ""));
|
||||||
|
const label = computed(() => lang.value.startsWith("zh") ? "在 GitHub 查看源文件" : "View source on GitHub");
|
||||||
|
const href = computed(() => `https://github.com/agentscope-ai/ReMe/blob/main/${sourcePath.value}`);
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<template>
|
||||||
|
<div v-if="sourcePath" class="source-link-wrap">
|
||||||
|
<a :href="href" target="_blank" rel="noreferrer">{{ label }} ↗</a>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
55
docs/.vitepress/theme/TrafficPage.vue
Normal file
55
docs/.vitepress/theme/TrafficPage.vue
Normal file
|
|
@ -0,0 +1,55 @@
|
||||||
|
<script setup lang="ts">
|
||||||
|
import { computed } from "vue";
|
||||||
|
import { useData } from "vitepress";
|
||||||
|
|
||||||
|
const props = defineProps<{ lang: "zh" | "en" }>();
|
||||||
|
const shareBase = "https://cloud.umami.is/analytics/us/share/S1OZK1PSDLEpyiU5?date=30day&page=1";
|
||||||
|
const { isDark } = useData();
|
||||||
|
const shareUrl = computed(() => `${shareBase}&theme=${isDark.value ? "dark" : "light"}`);
|
||||||
|
const text = computed(() => props.lang === "zh" ? {
|
||||||
|
eyebrow: "OPEN METRICS",
|
||||||
|
title: "ReMe 访问数据",
|
||||||
|
description: "最近 30 天的页面浏览量与访问趋势,由 Umami 提供隐私友好的匿名统计。",
|
||||||
|
action: "在 Umami 中打开完整页面",
|
||||||
|
frameTitle: "ReMe 最近 30 天访问数据",
|
||||||
|
} : {
|
||||||
|
eyebrow: "OPEN METRICS",
|
||||||
|
title: "ReMe traffic",
|
||||||
|
description: "Page views and traffic trends from the last 30 days, measured anonymously with privacy-friendly Umami analytics.",
|
||||||
|
action: "Open the full report in Umami",
|
||||||
|
frameTitle: "ReMe traffic for the last 30 days",
|
||||||
|
});
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<template>
|
||||||
|
<main class="traffic-page">
|
||||||
|
<header>
|
||||||
|
<div>
|
||||||
|
<p>{{ text.eyebrow }}</p>
|
||||||
|
<h1>{{ text.title }}</h1>
|
||||||
|
<span>{{ text.description }}</span>
|
||||||
|
</div>
|
||||||
|
<a :href="shareUrl" target="_blank" rel="noreferrer">{{ text.action }} ↗</a>
|
||||||
|
</header>
|
||||||
|
<div class="traffic-frame-wrap">
|
||||||
|
<iframe :src="shareUrl" :title="text.frameTitle" loading="eager" referrerpolicy="no-referrer" />
|
||||||
|
</div>
|
||||||
|
</main>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<style scoped>
|
||||||
|
.traffic-page { max-width: 1440px; margin: 0 auto; padding: clamp(54px, 7vw, 96px) clamp(22px, 5vw, 74px) 90px; }
|
||||||
|
.traffic-page header { display: flex; align-items: flex-end; justify-content: space-between; gap: 40px; margin-bottom: 34px; }
|
||||||
|
.traffic-page header p { margin: 0 0 15px; color: var(--vp-c-brand-1); font: 750 12px/1.4 var(--vp-font-family-mono); letter-spacing: 0.16em; }
|
||||||
|
.traffic-page h1 { margin: 0; color: var(--vp-c-text-1); font-size: clamp(42px, 5.5vw, 68px); line-height: 1.05; letter-spacing: -0.05em; }
|
||||||
|
.traffic-page header span { display: block; max-width: 720px; margin-top: 17px; color: var(--vp-c-text-2); font-size: 17px; line-height: 1.65; }
|
||||||
|
.traffic-page header a { flex: none; padding: 11px 15px; border: 1px solid var(--vp-c-divider); border-radius: 10px; color: var(--vp-c-text-1); background: var(--vp-c-bg-soft); text-decoration: none; font-size: 14px; font-weight: 700; }
|
||||||
|
.traffic-page header a:hover { border-color: var(--vp-c-brand-1); color: var(--vp-c-brand-1); }
|
||||||
|
.traffic-frame-wrap { height: min(880px, calc(100vh - 210px)); min-height: 650px; overflow: hidden; border: 1px solid var(--vp-c-divider); border-radius: 20px; background: white; box-shadow: 0 24px 58px rgba(26, 62, 47, 0.12); }
|
||||||
|
.traffic-frame-wrap iframe { width: 100%; height: 100%; border: 0; }
|
||||||
|
@media (max-width: 700px) {
|
||||||
|
.traffic-page { padding-top: 42px; }
|
||||||
|
.traffic-page header { align-items: flex-start; flex-direction: column; gap: 20px; }
|
||||||
|
.traffic-frame-wrap { height: 720px; min-height: 0; border-radius: 14px; }
|
||||||
|
}
|
||||||
|
</style>
|
||||||
370
docs/.vitepress/theme/custom.css
Normal file
370
docs/.vitepress/theme/custom.css
Normal file
|
|
@ -0,0 +1,370 @@
|
||||||
|
:root {
|
||||||
|
--vp-layout-max-width: 1560px;
|
||||||
|
--vp-c-brand-1: #087f6a;
|
||||||
|
--vp-c-brand-2: #086554;
|
||||||
|
--vp-c-brand-3: #19a98f;
|
||||||
|
--vp-c-brand-soft: rgba(8, 127, 106, 0.14);
|
||||||
|
--vp-c-bg: #ffffff;
|
||||||
|
--vp-c-bg-alt: #f4f7f5;
|
||||||
|
--vp-c-bg-elv: #ffffff;
|
||||||
|
--vp-c-bg-soft: #f1f6f3;
|
||||||
|
--vp-c-text-1: #17221d;
|
||||||
|
--vp-c-text-2: #526159;
|
||||||
|
--vp-c-text-3: #718078;
|
||||||
|
--vp-c-divider: #dce5e0;
|
||||||
|
--vp-font-family-base: Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif;
|
||||||
|
--vp-font-family-mono: "SFMono-Regular", Consolas, "Liberation Mono", monospace;
|
||||||
|
--reme-blue: #3156d9;
|
||||||
|
--reme-green: #087f6a;
|
||||||
|
}
|
||||||
|
|
||||||
|
.dark {
|
||||||
|
--vp-c-brand-1: #57dfc3;
|
||||||
|
--vp-c-brand-2: #35c6a9;
|
||||||
|
--vp-c-brand-3: #087f6a;
|
||||||
|
--vp-c-brand-soft: rgba(87, 223, 195, 0.14);
|
||||||
|
--vp-c-bg: #0d1512;
|
||||||
|
--vp-c-bg-alt: #09100d;
|
||||||
|
--vp-c-bg-elv: #14201b;
|
||||||
|
--vp-c-bg-soft: #17251f;
|
||||||
|
--vp-c-text-1: #edf7f3;
|
||||||
|
--vp-c-text-2: #bacbc4;
|
||||||
|
--vp-c-text-3: #91a49c;
|
||||||
|
--vp-c-divider: #283a33;
|
||||||
|
}
|
||||||
|
|
||||||
|
body {
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 8% 8%, rgba(25, 201, 176, 0.055), transparent 28rem),
|
||||||
|
var(--vp-c-bg);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNav {
|
||||||
|
border-bottom: 1px solid color-mix(in srgb, var(--vp-c-divider) 84%, transparent);
|
||||||
|
background: color-mix(in srgb, var(--vp-c-bg) 88%, transparent);
|
||||||
|
backdrop-filter: blur(18px) saturate(140%);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBarTitle .logo {
|
||||||
|
width: 30px;
|
||||||
|
height: 30px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBarTitle .title {
|
||||||
|
font-weight: 780;
|
||||||
|
letter-spacing: -0.02em;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBarSearch .DocSearch-Button,
|
||||||
|
.VPNavBarSearch button {
|
||||||
|
min-width: 190px;
|
||||||
|
border: 1px solid var(--vp-c-divider);
|
||||||
|
border-radius: 10px;
|
||||||
|
background: var(--vp-c-bg-alt);
|
||||||
|
}
|
||||||
|
|
||||||
|
@media (min-width: 960px) {
|
||||||
|
.VPNavBar.has-sidebar .container > .title {
|
||||||
|
width: 113px !important;
|
||||||
|
padding-right: 0 !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBar.has-sidebar .content {
|
||||||
|
padding-left: 113px !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBarSearch {
|
||||||
|
padding-left: 24px !important;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@media (min-width: 1496px) {
|
||||||
|
.VPNavBar.has-sidebar .container {
|
||||||
|
position: relative !important;
|
||||||
|
max-width: 1496px !important;
|
||||||
|
margin: 0 auto !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBar.has-sidebar .container > .title {
|
||||||
|
width: 81px !important;
|
||||||
|
padding-left: 0 !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPNavBar.has-sidebar .content {
|
||||||
|
padding-right: 0 !important;
|
||||||
|
padding-left: 81px !important;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPSidebar {
|
||||||
|
border-right: 1px solid var(--vp-c-divider);
|
||||||
|
background: color-mix(in srgb, var(--vp-c-bg-alt) 82%, transparent);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPSidebarItem .text {
|
||||||
|
font-size: 14px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPSidebarItem.level-0 > .item > .text {
|
||||||
|
color: var(--vp-c-text-1);
|
||||||
|
font-weight: 750;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPSidebarItem.is-active > .item .link > .text {
|
||||||
|
color: var(--vp-c-brand-1);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPDocAsideOutline {
|
||||||
|
border-left-color: var(--vp-c-divider);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPDoc .container > .content {
|
||||||
|
min-width: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPDoc .content-container {
|
||||||
|
max-width: 900px !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc {
|
||||||
|
color: var(--vp-c-text-1);
|
||||||
|
font-size: 16px;
|
||||||
|
line-height: 1.78;
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc h1 {
|
||||||
|
margin-bottom: 28px;
|
||||||
|
font-size: clamp(36px, 5vw, 52px);
|
||||||
|
line-height: 1.08;
|
||||||
|
letter-spacing: -0.045em;
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc h2 {
|
||||||
|
margin-top: 52px;
|
||||||
|
border-top-color: var(--vp-c-divider);
|
||||||
|
font-size: 27px;
|
||||||
|
letter-spacing: -0.025em;
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc h3 {
|
||||||
|
margin-top: 34px;
|
||||||
|
font-size: 20px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc :not(pre) > code {
|
||||||
|
border-radius: 5px;
|
||||||
|
color: color-mix(in srgb, var(--vp-c-brand-1) 82%, var(--vp-c-text-1));
|
||||||
|
}
|
||||||
|
|
||||||
|
.vp-doc div[class*="language-"] {
|
||||||
|
border: 1px solid var(--vp-c-divider);
|
||||||
|
border-radius: 12px;
|
||||||
|
box-shadow: inset 3px 0 0 color-mix(in srgb, var(--vp-c-brand-1) 62%, transparent);
|
||||||
|
}
|
||||||
|
|
||||||
|
.copy-markdown-wrap {
|
||||||
|
display: flex;
|
||||||
|
justify-content: flex-end;
|
||||||
|
margin-bottom: 18px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.copy-markdown {
|
||||||
|
display: inline-flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 7px;
|
||||||
|
min-height: 34px;
|
||||||
|
padding: 6px 12px;
|
||||||
|
border: 1px solid var(--vp-c-divider);
|
||||||
|
border-radius: 9px;
|
||||||
|
color: var(--vp-c-text-2);
|
||||||
|
background: var(--vp-c-bg-soft);
|
||||||
|
cursor: pointer;
|
||||||
|
font-size: 13px;
|
||||||
|
font-weight: 650;
|
||||||
|
}
|
||||||
|
|
||||||
|
.copy-markdown:hover,
|
||||||
|
.copy-markdown.copied {
|
||||||
|
border-color: var(--vp-c-brand-1);
|
||||||
|
color: var(--vp-c-brand-1);
|
||||||
|
}
|
||||||
|
|
||||||
|
.copy-markdown svg {
|
||||||
|
width: 15px;
|
||||||
|
height: 15px;
|
||||||
|
fill: none;
|
||||||
|
stroke: currentColor;
|
||||||
|
stroke-linecap: round;
|
||||||
|
stroke-linejoin: round;
|
||||||
|
stroke-width: 2;
|
||||||
|
}
|
||||||
|
|
||||||
|
.source-link-wrap {
|
||||||
|
margin-top: 48px;
|
||||||
|
padding-top: 20px;
|
||||||
|
border-top: 1px solid var(--vp-c-divider);
|
||||||
|
font-size: 13px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.source-link-wrap a {
|
||||||
|
color: var(--vp-c-text-3);
|
||||||
|
text-decoration: none;
|
||||||
|
}
|
||||||
|
|
||||||
|
.source-link-wrap a:hover {
|
||||||
|
color: var(--vp-c-brand-1);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHome {
|
||||||
|
overflow: hidden;
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 14% 14%, rgba(25, 201, 176, 0.14), transparent 30rem),
|
||||||
|
radial-gradient(circle at 86% 10%, rgba(49, 86, 217, 0.11), transparent 28rem);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .name {
|
||||||
|
background: linear-gradient(120deg, var(--reme-green), var(--reme-blue));
|
||||||
|
background-clip: text;
|
||||||
|
-webkit-background-clip: text;
|
||||||
|
-webkit-text-fill-color: transparent;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .name,
|
||||||
|
.VPHero .text {
|
||||||
|
letter-spacing: -0.05em;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-bg {
|
||||||
|
width: min(88%, 420px);
|
||||||
|
height: 220px;
|
||||||
|
border-radius: 42%;
|
||||||
|
background: linear-gradient(125deg, rgba(25, 201, 176, 0.34), rgba(49, 86, 217, 0.25));
|
||||||
|
filter: blur(52px);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-container {
|
||||||
|
isolation: isolate;
|
||||||
|
perspective: 900px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-container::before,
|
||||||
|
.VPHero .image-container::after {
|
||||||
|
position: absolute;
|
||||||
|
content: "";
|
||||||
|
pointer-events: none;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-container::before {
|
||||||
|
z-index: 0;
|
||||||
|
top: 50%;
|
||||||
|
left: 50%;
|
||||||
|
width: min(88%, 430px);
|
||||||
|
height: 210px;
|
||||||
|
border: 1px solid color-mix(in srgb, var(--vp-c-bg-elv) 64%, var(--reme-blue));
|
||||||
|
border-radius: 30px;
|
||||||
|
background:
|
||||||
|
linear-gradient(135deg, color-mix(in srgb, var(--vp-c-bg-elv) 92%, transparent), color-mix(in srgb, var(--vp-c-bg-soft) 76%, transparent)),
|
||||||
|
radial-gradient(circle at 15% 15%, rgba(25, 201, 176, 0.14), transparent 42%);
|
||||||
|
box-shadow:
|
||||||
|
0 30px 70px rgba(19, 70, 91, 0.16),
|
||||||
|
inset 0 1px 0 color-mix(in srgb, white 72%, transparent);
|
||||||
|
backdrop-filter: blur(22px) saturate(135%);
|
||||||
|
transform: translate(-50%, -50%) rotate(-1.5deg);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-container::after {
|
||||||
|
z-index: -1;
|
||||||
|
top: 50%;
|
||||||
|
left: 50%;
|
||||||
|
width: min(78%, 380px);
|
||||||
|
height: 210px;
|
||||||
|
border: 1px solid rgba(49, 86, 217, 0.17);
|
||||||
|
border-radius: 30px;
|
||||||
|
background: linear-gradient(135deg, rgba(25, 201, 176, 0.1), rgba(49, 86, 217, 0.11));
|
||||||
|
transform: translate(-46%, -48%) rotate(7deg);
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHero .image-src {
|
||||||
|
z-index: 1;
|
||||||
|
width: min(76%, 380px);
|
||||||
|
max-width: 380px !important;
|
||||||
|
max-height: 150px !important;
|
||||||
|
object-fit: contain;
|
||||||
|
filter: drop-shadow(0 12px 20px rgba(18, 78, 105, 0.16));
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .item:nth-child(1) { --feature-accent: #18b99e; }
|
||||||
|
.VPHomeFeatures .item:nth-child(2) { --feature-accent: #25a8dc; }
|
||||||
|
.VPHomeFeatures .item:nth-child(3) { --feature-accent: #6575e8; }
|
||||||
|
.VPHomeFeatures .item:nth-child(4) { --feature-accent: #e69a42; }
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature {
|
||||||
|
border-color: color-mix(in srgb, var(--feature-accent) 28%, var(--vp-c-divider));
|
||||||
|
border-radius: 16px;
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 10% 4%, color-mix(in srgb, var(--feature-accent) 16%, transparent), transparent 52%),
|
||||||
|
linear-gradient(150deg, var(--vp-c-bg-elv), color-mix(in srgb, var(--feature-accent) 7%, var(--vp-c-bg-soft)));
|
||||||
|
box-shadow:
|
||||||
|
inset 0 1px 0 color-mix(in srgb, white 76%, transparent),
|
||||||
|
0 8px 24px color-mix(in srgb, var(--feature-accent) 7%, transparent);
|
||||||
|
transition: transform 160ms ease, border-color 160ms ease, box-shadow 160ms ease;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature .box {
|
||||||
|
display: grid;
|
||||||
|
grid-template-columns: auto minmax(0, 1fr);
|
||||||
|
grid-template-rows: auto 1fr;
|
||||||
|
column-gap: 12px;
|
||||||
|
align-items: center;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature .icon {
|
||||||
|
grid-column: 1;
|
||||||
|
grid-row: 1;
|
||||||
|
width: auto;
|
||||||
|
height: auto;
|
||||||
|
margin: 0;
|
||||||
|
border: 0;
|
||||||
|
background: transparent;
|
||||||
|
box-shadow: none;
|
||||||
|
font-size: 25px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature .title {
|
||||||
|
grid-column: 2;
|
||||||
|
grid-row: 1;
|
||||||
|
color: color-mix(in srgb, var(--feature-accent) 22%, var(--vp-c-text-1));
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature .details {
|
||||||
|
grid-column: 1 / -1;
|
||||||
|
grid-row: 2;
|
||||||
|
align-self: start;
|
||||||
|
padding-top: 18px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature .link-text {
|
||||||
|
grid-column: 1 / -1;
|
||||||
|
}
|
||||||
|
|
||||||
|
.VPHomeFeatures .VPFeature:hover {
|
||||||
|
transform: translateY(-3px);
|
||||||
|
border-color: color-mix(in srgb, var(--feature-accent) 48%, var(--vp-c-divider));
|
||||||
|
box-shadow: 0 18px 38px color-mix(in srgb, var(--feature-accent) 15%, transparent);
|
||||||
|
}
|
||||||
|
|
||||||
|
.dark .VPHomeFeatures .VPFeature {
|
||||||
|
background:
|
||||||
|
radial-gradient(circle at 10% 4%, color-mix(in srgb, var(--feature-accent) 18%, transparent), transparent 54%),
|
||||||
|
linear-gradient(150deg, var(--vp-c-bg-elv), color-mix(in srgb, var(--feature-accent) 8%, var(--vp-c-bg-soft)));
|
||||||
|
box-shadow: inset 0 1px 0 rgba(255, 255, 255, 0.06);
|
||||||
|
}
|
||||||
|
|
||||||
|
@media (max-width: 768px) {
|
||||||
|
.vp-doc h1 { font-size: 34px; }
|
||||||
|
.vp-doc h2 { margin-top: 44px; font-size: 24px; }
|
||||||
|
.copy-markdown-wrap { margin-top: -8px; }
|
||||||
|
.VPHero .image-container::before,
|
||||||
|
.VPHero .image-container::after { height: 176px; border-radius: 24px; }
|
||||||
|
.VPHero .image-src { width: 72%; max-height: 120px !important; }
|
||||||
|
}
|
||||||
21
docs/.vitepress/theme/index.ts
Normal file
21
docs/.vitepress/theme/index.ts
Normal file
|
|
@ -0,0 +1,21 @@
|
||||||
|
import { h } from "vue";
|
||||||
|
import DefaultTheme from "vitepress/theme";
|
||||||
|
import CopyMarkdownButton from "./CopyMarkdownButton.vue";
|
||||||
|
import HomePage from "./HomePage.vue";
|
||||||
|
import SourceLink from "./SourceLink.vue";
|
||||||
|
import TrafficPage from "./TrafficPage.vue";
|
||||||
|
import "./custom.css";
|
||||||
|
|
||||||
|
export default {
|
||||||
|
extends: DefaultTheme,
|
||||||
|
enhanceApp({ app }) {
|
||||||
|
app.component("HomePage", HomePage);
|
||||||
|
app.component("TrafficPage", TrafficPage);
|
||||||
|
},
|
||||||
|
Layout() {
|
||||||
|
return h(DefaultTheme.Layout, null, {
|
||||||
|
"doc-before": () => h(CopyMarkdownButton),
|
||||||
|
"doc-after": () => h(SourceLink),
|
||||||
|
});
|
||||||
|
},
|
||||||
|
};
|
||||||
|
|
@ -1,23 +0,0 @@
|
||||||
# ReMe 仓库文档
|
|
||||||
|
|
||||||
本目录保存 ReMe 仓库 README 直接引用的中英文补充说明和图片资源,不作为文档站点的构建或部署来源。
|
|
||||||
|
|
||||||
面向用户发布的中英文文档位于 [agentscope-ai/docs](https://github.com/agentscope-ai/docs) 仓库,并由该仓库统一完成版本管理和 Mintlify 部署。
|
|
||||||
|
|
||||||
## 目录用途
|
|
||||||
|
|
||||||
```text
|
|
||||||
docs/
|
|
||||||
├── README.md 本目录的维护说明
|
|
||||||
├── doc.md 当前文档设计与维护边界
|
|
||||||
├── en/ README 引用的英文补充说明
|
|
||||||
├── zh/ README 引用的中文补充说明
|
|
||||||
└── figure/ ReMe README 使用的图片资源
|
|
||||||
```
|
|
||||||
|
|
||||||
## 维护原则
|
|
||||||
|
|
||||||
- `en/` 和 `zh/` 保持精简,服务 README 中需要进一步解释的功能与场景;修改路径时同步更新 README 链接。
|
|
||||||
- 具体实现以源码、schema、测试和运行时帮助为准,避免维护重复且容易过期的开发手册。
|
|
||||||
- README 引用的图片保留在 `figure/`;发布文档需要图片时,在统一文档仓库的 `images/reme/` 中维护对应副本。
|
|
||||||
- 网页文档、导航、版本和部署在统一文档仓库中维护。
|
|
||||||
85
docs/doc.md
85
docs/doc.md
|
|
@ -1,85 +0,0 @@
|
||||||
# ReMe 文档设计
|
|
||||||
|
|
||||||
本文定义 ReMe 文档的内容边界和维护方式。目标是让文档保持精简、稳定,并适合用户与 AI coding agent 快速理解。
|
|
||||||
|
|
||||||
## 两类文档,两种职责
|
|
||||||
|
|
||||||
| 位置 | 用途 | 是否部署 |
|
|
||||||
|---|---|---|
|
|
||||||
| `ReMe/docs/` | README 引用的中英文补充说明和图片 | 否 |
|
|
||||||
| `agentscope-ai/docs/reme/<version>/` | 面向用户的中英文产品文档 | 是 |
|
|
||||||
|
|
||||||
ReMe 仓库维护 `docs/en/`、`docs/zh/` 中供 README 直接引用的页面,但不把它们作为网页部署来源。网站内容、发布、版本选择、
|
|
||||||
导航和重定向都由统一文档仓库负责。
|
|
||||||
|
|
||||||
## 内容原则
|
|
||||||
|
|
||||||
### Concepts 只讲理念
|
|
||||||
|
|
||||||
Concepts 应解释 ReMe 为什么这样设计,而不是逐项描述组件和流水线实现。核心判断包括:
|
|
||||||
|
|
||||||
- **Memory as File**:记忆首先是用户拥有、可读写和可迁移的文件。
|
|
||||||
- **Memory from Experience**:长期记忆来自经验的提炼、修正和合并,而不是无限累积上下文。
|
|
||||||
- **Human-Agent Shared Memory**:用户和 Agent 共同读写同一份可见记忆。
|
|
||||||
- **Connected and Traceable**:长期结论可以通过链接回到来源和上下文。
|
|
||||||
|
|
||||||
算法、索引、Job、Step 和存储实现只有在帮助解释理念取舍时才进入 Concepts。
|
|
||||||
|
|
||||||
### Development 保持轻量
|
|
||||||
|
|
||||||
现代开发主要由 AI 直接阅读源码、schema 和测试完成。Development 只需要提供:
|
|
||||||
|
|
||||||
- 开发环境和最小验证命令;
|
|
||||||
- 代码目录入口;
|
|
||||||
- 兼容性与贡献要求;
|
|
||||||
- 哪些源码或 schema 是权威依据。
|
|
||||||
|
|
||||||
不为每个类、组件或扩展点编写重复的开发手册,也不维护 `generic_agent` 一类泛化教程。
|
|
||||||
|
|
||||||
### Reference 只记录稳定契约
|
|
||||||
|
|
||||||
Reference 记录 workspace、配置入口、CLI、HTTP、MCP 和文件格式的稳定语义。精确参数交给运行时帮助、Pydantic schema 和源码,避免文档复制一份容易失真的接口定义。
|
|
||||||
|
|
||||||
### Guides 只保留已验证路径
|
|
||||||
|
|
||||||
接入文档应对应真实、可验证的工作流。目前优先维护 Claude Code、QwenPaw,以及 Skill、CLI、MCP、HTTP、Python 的选择说明。没有可验证实现的框架不提前创建占位页。
|
|
||||||
|
|
||||||
## 发布文档结构
|
|
||||||
|
|
||||||
ReMe 参考 AgentScope 的版本目录和导航方式:
|
|
||||||
|
|
||||||
```text
|
|
||||||
agentscope-ai/docs/
|
|
||||||
├── reme/
|
|
||||||
│ └── 0.4.0.6/
|
|
||||||
│ ├── en/
|
|
||||||
│ └── zh/
|
|
||||||
└── images/
|
|
||||||
└── reme/
|
|
||||||
```
|
|
||||||
|
|
||||||
每个语言版本保持三组导航:
|
|
||||||
|
|
||||||
1. **Get Started / 快速开始**:Index、Overview、Quick Start、Concepts。
|
|
||||||
2. **Integrate / 接入**:接入选择、Claude Code、QwenPaw。
|
|
||||||
3. **Reference / 查阅与参与**:Reference、Support、Contributing。
|
|
||||||
|
|
||||||
ReMe 使用项目级别的 `/reme/latest/` 和 `/reme/stable/` 别名,不影响 AgentScope 自己的 `/latest/` 与 `/stable/`。
|
|
||||||
|
|
||||||
## 变更应该写在哪里
|
|
||||||
|
|
||||||
| 变更类型 | ReMe 仓库 | 统一文档仓库 |
|
|
||||||
|---|---|---|
|
|
||||||
| 产品理念或长期设计判断 | 更新 `docs/doc.md` 或相关设计记录 | 必要时同步 Concepts |
|
|
||||||
| 用户可见的安装、配置或行为 | 源码、schema、测试;影响 README 时同步 `docs/en/`、`docs/zh/` | 更新对应版本的用户文档 |
|
|
||||||
| 内部重构或组件调整 | 以代码和测试表达 | 稳定契约未变时无需更新 |
|
|
||||||
| README 图片 | 更新 `docs/figure/` | 发布页使用时同步到 `images/reme/` |
|
|
||||||
| 新版本发布 | 更新版本号和代码 | 新建版本目录、双语导航与 ReMe 别名 |
|
|
||||||
|
|
||||||
## 质量要求
|
|
||||||
|
|
||||||
- 每个用户流程必须能够在当前版本运行和验证。
|
|
||||||
- 文档不复制能够从代码可靠获得的细节。
|
|
||||||
- 删除过期内容优先于继续叠加补丁说明。
|
|
||||||
- 中英文页面保持信息等价,不要求逐句直译。
|
|
||||||
- 发布前在统一文档仓库运行 Mintlify 严格校验。
|
|
||||||
|
|
@ -1,16 +1,13 @@
|
||||||
# Auto Dream
|
# Auto Dream
|
||||||
|
|
||||||
`auto_dream` is ReMe's long-term memory distillation flow from daily to digest. It scans daily inputs for a specified date,
|
`auto_dream` is ReMe's long-term memory distillation flow from daily to digest. By default it scans the target date and
|
||||||
processes only files that changed since the previous dream, extracts content worth retaining as memory units, integrates those
|
the previous day, processes only files changed since the previous dream, extracts a small set of high-value memory units
|
||||||
units into `digest/`, and generates the day's `interests.yaml` for proactive use.
|
across that window, and integrates them into `digest/`.
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<img src="../figure/auto-dream-and-proactive.svg" alt="ReMe Auto Dream and Proactive flow from daily to digest to proactive" width="92%">
|
|
||||||
</p>
|
|
||||||
|
|
||||||
Its daily inputs usually come from [Auto Memory](./auto_memory.md) and [Auto Resource](./auto_resource.md). For the file
|
Its daily inputs usually come from [Auto Memory](./auto_memory.md) and [Auto Resource](./auto_resource.md). For the file
|
||||||
semantics of `digest/`, `derived_from::`, and wikilinks, see [Memory as File](./memory_as_file.md). For the linking strategy
|
semantics of `digest/`, Sources sections, and wikilinks, see [Memory as File](./memory_as_file.md). For the linking
|
||||||
used during Integrate, see [Auto Link](./auto_link.md). To read `interests.yaml`, use [Proactive](./proactive.md).
|
strategy used during Integrate, see [Auto Link](./auto_link.md). Proactive discovery is a separate flow; see
|
||||||
|
[Proactive](./proactive.md).
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
|
|
@ -26,53 +23,53 @@ auto_dream:
|
||||||
hint:
|
hint:
|
||||||
type: string
|
type: string
|
||||||
default: ""
|
default: ""
|
||||||
topic_count:
|
scan_days:
|
||||||
type: integer
|
type: integer
|
||||||
default: 3
|
default: 2
|
||||||
topic_diversity_days:
|
max_units:
|
||||||
type: integer
|
type: integer
|
||||||
default: 7
|
default: 5
|
||||||
steps:
|
steps:
|
||||||
- backend: dream_extract_step
|
- backend: dream_extract_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
topic_session_id: interests
|
scan_days: 2
|
||||||
|
max_units: 5
|
||||||
- backend: dream_integrate_step
|
- backend: dream_integrate_step
|
||||||
- backend: dream_topics_step
|
|
||||||
topic_count: 3
|
|
||||||
topic_diversity_days: 7
|
|
||||||
- backend: dream_finish_step
|
- backend: dream_finish_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
|
- backend: auto_tag_step
|
||||||
```
|
```
|
||||||
|
|
||||||
Parameters:
|
Parameters:
|
||||||
|
|
||||||
| Parameter | Purpose |
|
| Parameter | Purpose |
|
||||||
|---|---|
|
|------------------------|---------------------------------------------------------------------------------------------------------|
|
||||||
| `date` | Date to process in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
|
| `date` | Date to process in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
|
||||||
| `hint` | Additional guidance from the caller for the Extract and Integrate stages. |
|
| `hint` | Additional guidance from the caller for the Extract and Integrate stages. |
|
||||||
| `topic_count` | Maximum number of topics written to `interests.yaml`. Defaults to 3. |
|
| `scan_days` | Recent-date window ending at `date`; defaults to 2 and has a minimum of 1. |
|
||||||
| `topic_diversity_days` | Number of past days of `interests.yaml` files considered when avoiding duplicate topics. Defaults to 7. |
|
| `max_units` | Maximum reusable units extracted in one run; defaults to 5. |
|
||||||
|
|
||||||
## Inputs and Outputs
|
## Inputs and Outputs
|
||||||
|
|
||||||
Inputs are daily Markdown files for the specified date:
|
Inputs are daily Markdown files from the most recent `scan_days` ending at the specified date. For example,
|
||||||
|
`date=2026-06-20` with `scan_days=2` scans:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
daily/<date>.md
|
daily/2026-06-19.md
|
||||||
daily/<date>/**/*.md
|
daily/2026-06-19/**/*.md
|
||||||
|
daily/2026-06-20.md
|
||||||
|
daily/2026-06-20/**/*.md
|
||||||
```
|
```
|
||||||
|
|
||||||
`daily/<date>/interests.yaml` is excluded from extraction input so topics from the previous run do not feed back into the
|
Only Markdown day indexes and notes are scanned. Proactive state and `interests.yaml` are not Auto Dream inputs.
|
||||||
next extraction.
|
|
||||||
|
|
||||||
The main outputs are:
|
The main outputs are:
|
||||||
|
|
||||||
| Output | Description |
|
| Output | Description |
|
||||||
|---|---|
|
|--------------------------------|------------------------------------------------------------------------------|
|
||||||
| `digest/procedure/*.md` | Methods, workflows, runbooks, and executable experience. |
|
| `digest/procedure/*.md` | Methods, workflows, runbooks, and executable experience. |
|
||||||
| `digest/personal/*.md` | User-, team-, and project-related preferences, facts, and long-term context. |
|
| `digest/personal/*.md` | User-, team-, and project-related preferences, facts, and long-term context. |
|
||||||
| `digest/wiki/*.md` | General knowledge, concepts, observations, and decision precedents. |
|
| `digest/wiki/*.md` | General knowledge, concepts, observations, and decision precedents. |
|
||||||
| `daily/<date>/interests.yaml` | Topics worth proactive attention from the host agent that day. |
|
|
||||||
| `metadata/file_catalog/dream*` | Dream-specific catalog used to detect changes in daily inputs. |
|
| `metadata/file_catalog/dream*` | Dream-specific catalog used to detect changes in daily inputs. |
|
||||||
|
|
||||||
## Four Stages
|
## Four Stages
|
||||||
|
|
@ -81,88 +78,74 @@ The main outputs are:
|
||||||
|
|
||||||
`dream_extract_step` performs three tasks:
|
`dream_extract_step` performs three tasks:
|
||||||
|
|
||||||
1. Refresh the day's index page at `daily/<date>.md`.
|
1. Refresh each `daily/<date>.md` in the scan window.
|
||||||
2. Scan `daily/<date>.md` and `daily/<date>/**/*.md` and compare their mtimes with `file_catalog: dream`.
|
2. Scan those day indexes and `daily/<date>/**/*.md`, comparing mtimes with `file_catalog: dream`.
|
||||||
3. Send only changed files to the LLM and globally extract two structured result types: `units` and `topics`.
|
3. Send all changed files together to the LLM and globally extract structured memory `units`.
|
||||||
|
|
||||||
`units` are long-term memory units ready to be distilled into digest. Each has `name`, `bucket`, `summary`, and `paths`.
|
`units` are long-term memory units ready to be distilled into digest. Each has `name`, `bucket`, `summary`, and `paths`.
|
||||||
`bucket` may only be `procedure`, `personal`, or `wiki`; unknown values are routed to `wiki`.
|
A run returns at most `max_units`; extraction merges cross-file evidence for the same abstraction and drops passing
|
||||||
|
mentions, per-file summaries, and weak candidates without reusable value. `bucket` may only be `procedure`, `personal`,
|
||||||
|
or `wiki`; unknown values are routed to `wiki`.
|
||||||
|
|
||||||
`topics` are proactive-interest candidates for the day. They contain `title`, `reason`, `evidence`, `keywords`, and
|
If there are no changed files, Extract succeeds with no units; Integrate then has no unit work, and Finish still
|
||||||
`paths` and are filtered again in the Topics stage.
|
performs its normal catalog summary. If files changed but no LLM is
|
||||||
|
configured, Extract fails because extraction requires an LLM.
|
||||||
If there are no changed files, the flow ends early with success and skips later extraction work. If files changed but no LLM
|
|
||||||
is configured, Extract fails because extraction requires an LLM.
|
|
||||||
|
|
||||||
### 2. Integrate
|
### 2. Integrate
|
||||||
|
|
||||||
`dream_integrate_step` invokes an agent independently for each unit and integrates that unit into one digest node. It exposes
|
`dream_integrate_step` invokes an agent independently for each unit and integrates that unit into one digest node. It
|
||||||
these tools to the agent:
|
exposes these tools to the agent:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
node_search, read, frontmatter_read, write, edit, frontmatter_update
|
node_search, read, frontmatter_read, write, edit, frontmatter_update
|
||||||
```
|
```
|
||||||
|
|
||||||
This stage carries the core responsibility of `auto_link`. It first uses `node_search` to recall similar or related nodes at
|
This stage carries the core responsibility of `auto_link`. It first uses `node_search` to recall similar or related
|
||||||
digest-node granularity, decides whether to create or update a node, and finally writes sources and related digest nodes as
|
nodes at digest-node granularity, decides whether to create or update a node, and finally writes sources and related
|
||||||
wikilinks. See [Auto Link](./auto_link.md) for the recall, deduplication, and edge-writing rules.
|
digest nodes as wikilinks. See [Auto Link](./auto_link.md) for the recall, deduplication, and edge-writing rules.
|
||||||
|
|
||||||
|
Extract is the gate for deciding whether material is worth remembering, so Integrate has no `SKIP` action: each admitted
|
||||||
|
unit must land in exactly one digest node. Creates and updates must retain provenance and weave related digest links
|
||||||
|
into contextual sentences; bare wikilinks and standalone relationship fields are not valid output.
|
||||||
|
|
||||||
There are four integration actions:
|
There are four integration actions:
|
||||||
|
|
||||||
| Action | Meaning |
|
| Action | Meaning |
|
||||||
|---|---|
|
|---------------|--------------------------------------------------------------------------------|
|
||||||
| `CREATE` | No equivalent abstraction exists; create a new digest node. |
|
| `CREATE` | No equivalent abstraction exists; create a new digest node. |
|
||||||
| `CORROBORATE` | The same memory appeared again; append a source or strengthen the description. |
|
| `CORROBORATE` | The same memory appeared again; append a source or strengthen the description. |
|
||||||
| `REFINE` | New material adds boundaries, steps, prerequisites, applicability, or detail. |
|
| `REFINE` | New material adds boundaries, steps, prerequisites, applicability, or detail. |
|
||||||
| `CORRECT` | New material corrects errors, omissions, or conflicts in the existing node. |
|
| `CORRECT` | New material corrects errors, omissions, or conflicts in the existing node. |
|
||||||
|
|
||||||
Successfully integrated units are recorded in `integrate_results`. Failed units enter `failed_units`, and their source paths
|
Successfully integrated units are recorded in `integrate_results`. Failed units enter `failed_units`, and their source
|
||||||
enter `failed_paths`. The Finish stage does not checkpoint failed paths, ensuring that they can be retried later.
|
paths enter `failed_paths`. The Finish stage does not checkpoint failed paths, ensuring that they can be retried later.
|
||||||
|
|
||||||
### 3. Topics
|
### 3. Finish
|
||||||
|
|
||||||
`dream_topics_step` turns topic candidates from Extract into the final `daily/<date>/interests.yaml` for the day.
|
|
||||||
|
|
||||||
It reads:
|
|
||||||
|
|
||||||
```text
|
|
||||||
daily/<date>/interests.yaml
|
|
||||||
daily/<previous-date>/interests.yaml
|
|
||||||
```
|
|
||||||
|
|
||||||
Existing topics from the same day are preserved, while similar topics from the previous `topic_diversity_days` days are
|
|
||||||
deduplicated. At most three topics are written by default. With an LLM configured, the LLM selects topics that are more
|
|
||||||
specific, actionable, and non-repetitive. Without an LLM, the step falls back to local normalization and deduplication.
|
|
||||||
|
|
||||||
Example output format. See [Proactive](./proactive.md) for the interface that reads this file:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
date: 2026-06-20
|
|
||||||
topic_count: 3
|
|
||||||
diversity_days: 7
|
|
||||||
topics:
|
|
||||||
- title: Quality regression in the memory retrieval pipeline
|
|
||||||
reason: The user has recently made repeated changes to search, node_search, and dream integration.
|
|
||||||
evidence: daily/2026-06-20/session.md
|
|
||||||
keywords:
|
|
||||||
- memory search
|
|
||||||
- auto dream
|
|
||||||
paths:
|
|
||||||
- daily/2026-06-20/session.md
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. Finish
|
|
||||||
|
|
||||||
`dream_finish_step` completes the run:
|
`dream_finish_step` completes the run:
|
||||||
|
|
||||||
1. Write successfully processed changed paths to `file_catalog: dream`.
|
1. Write successfully processed changed paths to `file_catalog: dream`.
|
||||||
2. Also write `daily/<date>/interests.yaml` and `daily/<date>.md` to the catalog.
|
2. Also write every refreshed day-index page in the scan window to the catalog.
|
||||||
3. Persist the dream catalog if there were upserts or deletions.
|
3. Persist the dream catalog if there were upserts or deletions.
|
||||||
4. Return a summary containing counts for scanned, changed, integrated, topics, checkpoints, and related values.
|
4. Return a summary containing counts for scanned, changed, integrated, checkpoints, and related values.
|
||||||
|
|
||||||
|
Auto Dream neither reads nor writes proactive state or `interests.yaml`. Those files are owned by the proactive refresh
|
||||||
|
pipeline; see [Proactive](./proactive.md).
|
||||||
|
|
||||||
Failed paths are not checkpointed. The next `auto_dream` run therefore continues to treat them as changed inputs until
|
Failed paths are not checkpointed. The next `auto_dream` run therefore continues to treat them as changed inputs until
|
||||||
integration succeeds.
|
integration succeeds.
|
||||||
|
|
||||||
|
### 4. Auto Tag
|
||||||
|
|
||||||
|
After Finish, both `auto_dream` and `dream_cron` run `auto_tag_step` on Markdown digest files actually created or modified
|
||||||
|
during integration, including writes recovered after agent errors. Repeated writes to one file are tagged once. The
|
||||||
|
Step uses the same request-scoped `changes` contract as [Auto Memory](./auto_memory.md) and writes entity tags to the
|
||||||
|
configured frontmatter key, `memory_tags` by default. Unchanged files and daily source notes are not tagged by Dream.
|
||||||
|
|
||||||
|
Tagging diagnostics appear in `metadata.auto_tag`. Per-file tagging failures preserve the dream answer, success status,
|
||||||
|
and checkpoint decisions. A later run without file changes does not automatically retry failed tagging. Tag-index
|
||||||
|
updates follow the existing asynchronous file watcher.
|
||||||
|
|
||||||
## Running Auto Dream
|
## Running Auto Dream
|
||||||
|
|
||||||
CLI:
|
CLI:
|
||||||
|
|
@ -177,6 +160,12 @@ With caller guidance:
|
||||||
reme auto_dream date=2026-06-20 hint="Prioritize engineering decisions and long-term preferences"
|
reme auto_dream date=2026-06-20 hint="Prioritize engineering decisions and long-term preferences"
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Override the default scan window and unit cap:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme auto_dream date=2026-06-20 scan_days=3 max_units=8
|
||||||
|
```
|
||||||
|
|
||||||
The same set of steps can also be placed in a `cron` Job, for example to run every morning:
|
The same set of steps can also be placed in a `cron` Job, for example to run every morning:
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
|
|
@ -188,22 +177,22 @@ jobs:
|
||||||
- backend: dream_extract_step
|
- backend: dream_extract_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
- backend: dream_integrate_step
|
- backend: dream_integrate_step
|
||||||
- backend: dream_topics_step
|
|
||||||
- backend: dream_finish_step
|
- backend: dream_finish_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
|
- backend: auto_tag_step
|
||||||
```
|
```
|
||||||
|
|
||||||
## Important Boundaries
|
## Important Boundaries
|
||||||
|
|
||||||
`auto_dream` consumes only daily inputs and does not rewrite daily bodies. Daily preserves facts and the original situation;
|
`auto_dream` consumes only daily inputs and does not rewrite daily bodies. Daily preserves facts and the original
|
||||||
digest is the abstracted long-term memory layer.
|
situation; digest is the abstracted long-term memory layer.
|
||||||
|
|
||||||
`digest` is not a copy of the source text. Its body should preserve reusable abstractions, while details point back to sources
|
`digest` is not a copy of the source text. Its body should preserve reusable abstractions, while a Sources section
|
||||||
through `derived_from:: [[daily/<date>/...]]`. Links follow the workspace-relative wikilink semantics described in
|
points back with contextual sentences such as `The decision was recorded in [[daily/<date>/decision.md]].` Links follow
|
||||||
|
the workspace-relative wikilink semantics described in
|
||||||
[Memory as File](./memory_as_file.md).
|
[Memory as File](./memory_as_file.md).
|
||||||
|
|
||||||
`auto_dream` does not invent an overview from nothing. Only content that actually appears in daily input and is extracted as
|
`auto_dream` does not invent an overview from nothing. Only content that actually appears in daily input and is
|
||||||
a unit or topic can enter digest or `interests.yaml`.
|
extracted as a memory unit can enter digest.
|
||||||
|
|
||||||
The complete flow depends on an LLM for Extract and Integrate. Topics can perform local deduplication without an LLM, but that
|
The complete flow depends on an LLM for Extract, Integrate, and Auto Tag.
|
||||||
does not mean the full dream flow can run offline.
|
|
||||||
|
|
|
||||||
|
|
@ -1,11 +1,12 @@
|
||||||
# Auto Link
|
# Auto Link
|
||||||
|
|
||||||
In the current implementation, `auto_link` is not a separately registered Job. It is a capability of the Integrate stage in
|
In the current implementation, `auto_link` is not a separately registered Job. It is a capability of the Integrate stage
|
||||||
|
in
|
||||||
`auto_dream`: when `dream_integrate_step` writes a memory unit to `digest/`, it also recalls digest nodes, makes a
|
`auto_dream`: when `dream_integrate_step` writes a memory unit to `digest/`, it also recalls digest nodes, makes a
|
||||||
deduplication decision, links sources, and weaves wikilinks to related nodes into the result.
|
deduplication decision, links sources, and weaves wikilinks to related nodes into the result.
|
||||||
|
|
||||||
For the complete dream flow, see [Auto Dream](./auto_dream.md). For general wikilink, frontmatter, and workspace-relative
|
For the complete dream flow, see [Auto Dream](./auto_dream.md). For general wikilink, frontmatter, and
|
||||||
path semantics, see [Memory as File](./memory_as_file.md). For question-answering retrieval, see
|
workspace-relative path semantics, see [Memory as File](./memory_as_file.md). For question-answering retrieval, see
|
||||||
[Memory Search](./memory_search.md).
|
[Memory Search](./memory_search.md).
|
||||||
|
|
||||||
## Where It Runs
|
## Where It Runs
|
||||||
|
|
@ -17,22 +18,22 @@ auto_dream:
|
||||||
steps:
|
steps:
|
||||||
- dream_extract_step
|
- dream_extract_step
|
||||||
- dream_integrate_step # where auto_link actually happens
|
- dream_integrate_step # where auto_link actually happens
|
||||||
- dream_topics_step
|
|
||||||
- dream_finish_step
|
- dream_finish_step
|
||||||
|
- auto_tag_step
|
||||||
```
|
```
|
||||||
|
|
||||||
The Integrate stage processes each unit independently. A unit is written to exactly one target digest node, but that node may
|
The Integrate stage processes each unit independently. A unit is written to exactly one target digest node, but that
|
||||||
link to multiple sources and multiple related digest nodes.
|
node may link to multiple sources and multiple related digest nodes.
|
||||||
|
|
||||||
## Goals
|
## Goals
|
||||||
|
|
||||||
`auto_link` addresses graph quality at write time:
|
`auto_link` addresses graph quality at write time:
|
||||||
|
|
||||||
| Problem | Handling |
|
| Problem | Handling |
|
||||||
|---|---|
|
|------------------------------------------------|----------------------------------------------------------------------|
|
||||||
| The same memory already exists | Recall and update the existing node instead of creating a duplicate. |
|
| The same memory already exists | Recall and update the existing node instead of creating a duplicate. |
|
||||||
| New and existing material are related | Write workspace-relative wikilinks into the body. |
|
| New and existing material are related | Write workspace-relative wikilinks into the body. |
|
||||||
| A digest node is disconnected from its sources | Point back to daily/resource source material with `derived_from:: [[...]]`. |
|
| A digest node is disconnected from its sources | Add daily/resource links under a `## Sources` section. |
|
||||||
| A node contains only isolated prose | Add links to related digest nodes on both CREATE and UPDATE. |
|
| A node contains only isolated prose | Add links to related digest nodes on both CREATE and UPDATE. |
|
||||||
|
|
||||||
## Toolchain
|
## Toolchain
|
||||||
|
|
@ -48,24 +49,25 @@ edit
|
||||||
frontmatter_update
|
frontmatter_update
|
||||||
```
|
```
|
||||||
|
|
||||||
`node_search` is digest-only node retrieval designed for dream integration. It returns node-level signals such as the digest
|
`node_search` is digest-only node retrieval designed for dream integration. It returns node-level signals such as the
|
||||||
node's `path` and the `name` and `description` from frontmatter. It does not expand the body and does not perform the link
|
digest node's `path` and the `name` and `description` from frontmatter. It does not expand the body and does not perform
|
||||||
expansion used by ordinary search.
|
the link expansion used by ordinary search.
|
||||||
|
|
||||||
`read` and `frontmatter_read` are used only for candidates that may be relevant, avoiding expansion of every recalled result
|
`read` and `frontmatter_read` are used only for candidates that may be relevant, avoiding expansion of every recalled
|
||||||
into a large context.
|
result into a large context.
|
||||||
|
|
||||||
## Linking Flow
|
## Linking Flow
|
||||||
|
|
||||||
### 1. Recall candidate nodes
|
### 1. Recall candidate nodes
|
||||||
|
|
||||||
The agent first calls `node_search` with the unit's triggers, verbs, nouns, synonyms, and possible failure modes. Broad recall,
|
The agent first calls `node_search` with the unit's triggers, verbs, nouns, synonyms, and possible failure modes. Broad
|
||||||
for example `limit=20-30`, is recommended by default because this step serves both deduplication and link discovery.
|
recall, for example `limit=20-30`, is recommended by default because this step serves both deduplication and link
|
||||||
|
discovery.
|
||||||
|
|
||||||
Recalled results are internally classified into three groups:
|
Recalled results are internally classified into three groups:
|
||||||
|
|
||||||
| Classification | Meaning | Next action |
|
| Classification | Meaning | Next action |
|
||||||
|---|---|---|
|
|--------------------|---------------------------------------------------------------------------------------------------------|---------------------------|
|
||||||
| `same_abstraction` | The trigger or underlying abstraction is the same, with substantial content overlap. | Use as the UPDATE target. |
|
| `same_abstraction` | The trigger or underlying abstraction is the same, with substantial content overlap. | Use as the UPDATE target. |
|
||||||
| `related` | An adjacent process, prerequisite, failure mode, concept, preference, or upstream/downstream knowledge. | Write a body wikilink. |
|
| `related` | An adjacent process, prerequisite, failure mode, concept, preference, or upstream/downstream knowledge. | Write a body wikilink. |
|
||||||
| `unrelated` | Only superficially similar or unrelated. | Ignore. |
|
| `unrelated` | Only superficially similar or unrelated. | Ignore. |
|
||||||
|
|
@ -75,47 +77,47 @@ Recalled results are internally classified into three groups:
|
||||||
Every unit must select one action:
|
Every unit must select one action:
|
||||||
|
|
||||||
| Action | Linking semantics |
|
| Action | Linking semantics |
|
||||||
|---|---|
|
|---------------|-------------------------------------------------------------------------------------------------------------------------------|
|
||||||
| `CREATE` | Write a new `digest/<bucket>/<slug>.md` and add source and related-node links to its body. |
|
| `CREATE` | Write a new `digest/<bucket>/<slug>.md` and add source and related-node links to its body. |
|
||||||
| `CORROBORATE` | The same abstraction appeared again; append a new `derived_from:: [[...]]` and strengthen the description when needed. |
|
| `CORROBORATE` | The same abstraction appeared again; append its source link and strengthen the description when needed. |
|
||||||
| `REFINE` | New material extends the existing node; insert the additional content in the appropriate section and preserve existing links. |
|
| `REFINE` | New material extends the existing node; insert the additional content in the appropriate section and preserve existing links. |
|
||||||
| `CORRECT` | New material corrects the existing node; use source links to identify the basis for the correction. |
|
| `CORRECT` | New material corrects the existing node; use source links to identify the basis for the correction. |
|
||||||
|
|
||||||
An UPDATE should be additive whenever possible: do not delete existing wikilinks or `derived_from` entries. This prevents
|
An UPDATE should be additive whenever possible: do not delete existing wikilinks or source entries. This prevents later
|
||||||
later graph indexing and retrieval from losing edges.
|
graph indexing and retrieval from losing edges.
|
||||||
|
|
||||||
### 3. Write source edges
|
### 3. Write source edges
|
||||||
|
|
||||||
Source edges use Markdown wikilinks:
|
Source edges are ordinary wikilinks grouped under a Markdown heading:
|
||||||
|
|
||||||
```markdown
|
```markdown
|
||||||
derived_from:: [[daily/2026-06-20/session.md]]
|
## Sources
|
||||||
derived_from:: [[resource/2026-06-20/paper.md]]
|
|
||||||
|
The decision was recorded in [[daily/2026-06-20/session.md]], while the supporting technical evidence comes from
|
||||||
|
[[resource/2026-06-20/paper.md]].
|
||||||
```
|
```
|
||||||
|
|
||||||
These edges represent the evidence behind a digest node. Plain-text descriptions do not count as source edges because only
|
These edges represent the evidence behind a digest node. Plain-text descriptions do not count as source edges because
|
||||||
wikilinks can be parsed reliably by the file graph. For the complete parsing rules, see
|
only wikilinks can be parsed reliably by the file graph. The surrounding sentence must explain what each source
|
||||||
|
supports; a bare wikilink line is not valid Integrate output. For the complete parsing rules, see
|
||||||
[Memory as File](./memory_as_file.md#wikilink).
|
[Memory as File](./memory_as_file.md#wikilink).
|
||||||
|
|
||||||
### 4. Write relationships between digest nodes
|
### 4. Write relationships between digest nodes
|
||||||
|
|
||||||
Relationships between digest nodes also use complete workspace-relative paths:
|
Relationships between digest nodes use complete workspace-relative paths woven into natural prose:
|
||||||
|
|
||||||
```markdown
|
```markdown
|
||||||
relates_to:: [[digest/wiki/hybrid-search.md]]
|
This design extends [[digest/wiki/hybrid-search.md]] and uses
|
||||||
depends_on:: [[digest/procedure/rebuild-index.md]]
|
[[digest/procedure/rebuild-index.md]]. Follow
|
||||||
blocks_on:: [[digest/personal/team-review-preference.md]]
|
[[digest/personal/team-review-preference.md]] during review.
|
||||||
```
|
```
|
||||||
|
|
||||||
Predicates are open-ended. Common forms include `relates_to::`, `depends_on::`, and `blocks_on::`. The predicate sits outside
|
|
||||||
the brackets, while the target path goes inside `[[...]]` and should include the `.md` suffix.
|
|
||||||
|
|
||||||
## Bucket Differences
|
## Bucket Differences
|
||||||
|
|
||||||
`auto_link` adjusts the shape of its output according to the unit bucket:
|
`auto_link` adjusts the shape of its output according to the unit bucket:
|
||||||
|
|
||||||
| Bucket | Writing focus |
|
| Bucket | Writing focus |
|
||||||
|---|---|
|
|-------------|-----------------------------------------------------------------------------------------------------------------------------|
|
||||||
| `procedure` | Write a runbook with triggers, steps, inputs, and failure modes. Link prerequisites, substeps, and related preferences. |
|
| `procedure` | Write a runbook with triggers, steps, inputs, and failure modes. Link prerequisites, substeps, and related preferences. |
|
||||||
| `personal` | Write user-, team-, or project-specific facts and preferences. Link related projects, habits, and decision context. |
|
| `personal` | Write user-, team-, or project-specific facts and preferences. Link related projects, habits, and decision context. |
|
||||||
| `wiki` | Write general knowledge, principles, observations, and decision precedents. Link concepts, methods, and adjacent knowledge. |
|
| `wiki` | Write general knowledge, principles, observations, and decision precedents. Link concepts, methods, and adjacent knowledge. |
|
||||||
|
|
@ -127,18 +129,18 @@ Regardless of bucket, preserve source edges and weave recalled related digest no
|
||||||
`auto_link` uses `node_search`, not the question-answering `search`.
|
`auto_link` uses `node_search`, not the question-answering `search`.
|
||||||
|
|
||||||
| Capability | Purpose |
|
| Capability | Purpose |
|
||||||
|---|---|
|
|---------------|-----------------------------------------------------------------------------------------------------------|
|
||||||
| `search` | External question answering; returns chunks and can expand upstream/downstream link context. |
|
| `search` | External question answering; returns chunks and can expand upstream/downstream link context. |
|
||||||
| `node_search` | Dream integration; recalls only digest node-level summaries for deduplication and related-link decisions. |
|
| `node_search` | Dream integration; recalls only digest node-level summaries for deduplication and related-link decisions. |
|
||||||
|
|
||||||
This boundary matters. The Integrate stage needs to decide whether the same abstraction already exists and which nodes should
|
This boundary matters. The Integrate stage needs to decide whether the same abstraction already exists and which nodes
|
||||||
be linked; it should not load large numbers of body chunks into context. [Memory Search](./memory_search.md) handles
|
should be linked; it should not load large numbers of body chunks into context. [Memory Search](./memory_search.md)
|
||||||
question-oriented chunk retrieval, RRF fusion, and link expansion.
|
handles question-oriented chunk retrieval, RRF fusion, and link expansion.
|
||||||
|
|
||||||
## Failure and Retry
|
## Failure and Retry
|
||||||
|
|
||||||
If integration of a unit fails, `dream_integrate_step` records `failed_units` and `failed_paths`.
|
If integration of a unit fails, `dream_integrate_step` records `failed_units` and `failed_paths`.
|
||||||
`dream_finish_step` does not checkpoint those source paths, so the next `auto_dream` run processes them again.
|
`dream_finish_step` does not checkpoint those source paths, so the next `auto_dream` run processes them again.
|
||||||
|
|
||||||
This makes auto_link writes retryable: a failure does not mark the input as complete or silently discard digest edges that
|
This makes auto_link writes retryable: a failure does not mark the input as complete or silently discard digest edges
|
||||||
should have been created.
|
that should have been created.
|
||||||
|
|
|
||||||
|
|
@ -1,8 +1,9 @@
|
||||||
# Auto Memory
|
# Auto Memory
|
||||||
|
|
||||||
Auto Memory is ReMe's entry point for conversational memory. Each conversation is first distilled into a daily memory card
|
Auto Memory is ReMe's entry point for conversational memory. Within a target date, it uses `session_id` to find or update at
|
||||||
identified by `session_id`, and the day's `YYYY-MM-DD.md` page then indexes all of those cards. It turns "we talked about it"
|
most one daily memory card, whose filename is a concise topic or event name chosen by the Agent. The day's `YYYY-MM-DD.md`
|
||||||
into "it was remembered" while preserving the original conversation as evidence.
|
page indexes those cards. It turns "we talked about it" into "it was remembered" while retaining a source conversation record
|
||||||
|
as evidence.
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
|
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
|
||||||
|
|
@ -13,9 +14,9 @@ For the general file semantics of `daily/`, `session/`, frontmatter, and wikilin
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Conversation
|
Conversation
|
||||||
├─ step 1: daily/YYYY-MM-DD/<session_id>.md # one card per conversation
|
├─ step 1: daily/YYYY-MM-DD/<generated_name>.md # one topic-named card per session
|
||||||
├─ step 2: daily/YYYY-MM-DD.md # daily index linking the cards
|
├─ step 2: daily/YYYY-MM-DD.md # daily index linking the cards
|
||||||
└─ source: session/dialog/<session_id>.jsonl # original conversation
|
└─ source: session/dialog/<session_id>.jsonl # source conversation record
|
||||||
```
|
```
|
||||||
|
|
||||||
## What It Records
|
## What It Records
|
||||||
|
|
@ -39,29 +40,33 @@ workspace/
|
||||||
daily/
|
daily/
|
||||||
2026-06-20.md
|
2026-06-20.md
|
||||||
2026-06-20/
|
2026-06-20/
|
||||||
session-a.md
|
login-refactor-decision.md
|
||||||
session-b.md
|
retrieval-regression.md
|
||||||
```
|
```
|
||||||
|
|
||||||
`daily/2026-06-20/session-a.md` and `daily/2026-06-20/session-b.md` are memory cards distilled from different
|
The two files under the date directory are topic-named cards distilled from different conversations.
|
||||||
conversations. `daily/2026-06-20.md` is the index page for that day. Resource files enter the same daily memory layer; see
|
`daily/2026-06-20.md` is the index page for that day. Resource files enter the same daily memory layer; see
|
||||||
[Auto Resource](./auto_resource.md).
|
[Auto Resource](./auto_resource.md).
|
||||||
|
|
||||||
When a call includes `session_id`, Auto Memory records that conversation separately under the given ID:
|
When a call includes `session_id`, Auto Memory uses it to find the corresponding card through frontmatter, while the Agent
|
||||||
|
chooses a readable filename through `name`:
|
||||||
|
|
||||||
```text
|
```yaml
|
||||||
daily/2026-06-20/session-a.md
|
name: login-refactor-decision
|
||||||
|
session_id: session-a
|
||||||
|
source_conversation: "[[session/dialog/session-a.jsonl]]"
|
||||||
```
|
```
|
||||||
|
|
||||||
This keeps different conversations separate. A requirements discussion, a debugging session, and a documentation update can
|
This keeps different conversations separate without forcing opaque IDs into filenames. An update locates the existing note by
|
||||||
each have their own memory card. To see what happened on a particular day, start with `YYYY-MM-DD.md`. To inspect what was
|
`session_id` or `source_conversation`; if the Agent supplies a better frontmatter `name`, the system can rename the note and
|
||||||
distilled from one conversation, open the corresponding `<session_id>.md`.
|
retarget inbound wikilinks. To see what happened on a day, start with `YYYY-MM-DD.md`.
|
||||||
|
|
||||||
## Preserving the Original Information
|
## Preserving the Original Information
|
||||||
|
|
||||||
The distilled daily note is optimized for readability; the original conversation is retained for trust and verification.
|
The distilled daily note is optimized for readability; a filtered source conversation record is retained for trust and
|
||||||
|
verification.
|
||||||
|
|
||||||
While generating memory cards, Auto Memory also saves the raw sessions:
|
While generating memory cards, Auto Memory also saves the source messages:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
session/
|
session/
|
||||||
|
|
@ -70,12 +75,47 @@ session/
|
||||||
session-b.jsonl
|
session-b.jsonl
|
||||||
```
|
```
|
||||||
|
|
||||||
Each daily note points to its corresponding original conversation. When a memory needs verification, follow that link back to
|
Each daily note points to its corresponding conversation record. Saved messages omit tool-result blocks and base64 data
|
||||||
the complete context in which it was created.
|
blocks, preventing recalled memory and binary payloads from being mistaken for user-provided evidence later.
|
||||||
|
|
||||||
|
## Images in Conversations
|
||||||
|
|
||||||
|
Auto Memory can read images together with the surrounding conversation. Images are disabled by default; enable them for a
|
||||||
|
call with `include_images=true`.
|
||||||
|
|
||||||
|
Image input requires an `agentscope` wrapper with a vision-capable `as_llm` model and compatible formatter.
|
||||||
|
Auto Memory uses that model to read the conversation, without generating captions first. When images are disabled or no
|
||||||
|
image blocks are present, the existing text-only behavior is unchanged, including support for other wrappers.
|
||||||
|
|
||||||
|
Pass images as top-level AgentScope `DataBlock` values in `messages`, with an `image/` media type. Text and images stay in
|
||||||
|
their original order, with speaker and timestamp boundaries preserved. Base64 sources and HTTP(S) URLs pass unchanged to
|
||||||
|
the formatter; Auto Memory does not download or preprocess the images. URLs must be accessible to the model provider. For local
|
||||||
|
files, submit Base64 instead of a `file://` URL; other URL schemes are also unsupported.
|
||||||
|
|
||||||
|
The wrapper's `context_config.max_image_num` limits the number of images per call; Auto Memory rejects excess images rather
|
||||||
|
than increasing the limit. The AgentScope default is 5. To use a higher limit, set it when starting the service:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start components.agent_wrapper.default.context_config.max_image_num=20
|
||||||
|
```
|
||||||
|
|
||||||
|
Then call the running service from another terminal, using the same workspace:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme auto_memory session_id=session-a include_images=true messages='[...]'
|
||||||
|
```
|
||||||
|
|
||||||
|
Model and formatter limits still apply. When image input is enabled and images are present, Auto Memory checks the wrapper
|
||||||
|
backend, URL schemes and image count before saving the conversation. Later formatter or provider errors are returned
|
||||||
|
without retrying as text-only. As with text-only calls, those errors do not roll back an already saved conversation.
|
||||||
|
|
||||||
|
Source JSONL saving follows the filtering rules above, including the omission of Base64 blocks. To process those images
|
||||||
|
again, resubmit the original messages rather than the saved JSONL. No separate image files or caption cards are created,
|
||||||
|
though the wrapper's internal Agent state under `mem_session/agentscope` can contain image inputs.
|
||||||
|
|
||||||
## Message Timestamps
|
## Message Timestamps
|
||||||
|
|
||||||
Auto Memory preserves each message's `created_at` in both the prompt and the raw session JSONL. When importing historical
|
Auto Memory preserves each retained message's `created_at` in both the prompt and the source conversation JSONL. When importing historical
|
||||||
conversations or benchmark data, provide the actual occurrence time for every message so the model does not confuse event
|
conversations or benchmark data, provide the actual occurrence time for every message so the model does not confuse event
|
||||||
time with execution time:
|
time with execution time:
|
||||||
|
|
||||||
|
|
@ -92,8 +132,8 @@ For compatibility with common dataset schemas, `auto_memory` also checks `time_c
|
||||||
`timeCreated`, and `created_time` when `created_at` is absent. These fields may appear either at the top level of a message
|
`timeCreated`, and `created_time` when `created_at` is absent. These fields may appear either at the top level of a message
|
||||||
or inside `metadata`.
|
or inside `metadata`.
|
||||||
|
|
||||||
When a call does not explicitly provide `date`, Auto Memory uses the date of the earliest valid `created_at` value in the
|
When a call does not explicitly provide `date`, Auto Memory uses the latest valid `created_at` date in the messages. If no
|
||||||
messages. If no message contains a valid timestamp, it falls back to the current date. Historical imports may also specify the
|
message contains a valid timestamp, it falls back to the current date. Historical imports may also specify the
|
||||||
target date directly:
|
target date directly:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
|
@ -105,5 +145,13 @@ reme auto_memory \
|
||||||
|
|
||||||
## What Happens Next
|
## What Happens Next
|
||||||
|
|
||||||
|
The default `auto_memory` and `auto_memory_cc` jobs run `auto_tag_step` after recording memory. Only a daily note that
|
||||||
|
was actually created or modified is tagged, using its final path after any rename. Claude Code callers still pass only
|
||||||
|
`session_id`; repeated Stop events with no new messages skip both memory generation and tagging.
|
||||||
|
|
||||||
|
Tags describe the document's central entities and are stored in the configured frontmatter key (`memory_tags` by default).
|
||||||
|
Per-file tagging failures are reported in `metadata.auto_tag` while preserving the memory response. Calls without note
|
||||||
|
changes do not automatically retry failed tagging; the existing file watcher updates the tag index asynchronously.
|
||||||
|
|
||||||
Auto Memory only creates memory in the daily layer. To distill this material further into long-term `digest/` nodes, use
|
Auto Memory only creates memory in the daily layer. To distill this material further into long-term `digest/` nodes, use
|
||||||
[Auto Dream](./auto_dream.md). To search daily and digest content, use [Memory Search](./memory_search.md).
|
[Auto Dream](./auto_dream.md). To search daily and digest content, use [Memory Search](./memory_search.md).
|
||||||
|
|
|
||||||
|
|
@ -1,8 +1,8 @@
|
||||||
# Auto Resource `Beta`
|
# Auto Resource `Beta`
|
||||||
|
|
||||||
Auto Resource is ReMe's entry point for interpreting resources and is currently in **Beta**. Resource files first enter
|
Auto Resource is ReMe's entry point for interpreting resources and is currently in **Beta**. Resource files first enter
|
||||||
`resource/` by date and are then interpreted into daily resource cards. Each card's filename comes from the LLM-generated
|
`resource/`, preferably under a date directory, and are then interpreted into daily resource cards. Each card's filename
|
||||||
frontmatter `name`, and `source_resource` links the card back to its original file.
|
comes from the LLM-generated frontmatter `name`, and `source_resource` links the card back to its original file.
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
|
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
|
||||||
|
|
@ -13,7 +13,7 @@ For the general file semantics of workspace layers, `resource/`, and `daily/`, s
|
||||||
[Auto Memory](./auto_memory.md).
|
[Auto Memory](./auto_memory.md).
|
||||||
|
|
||||||
```text
|
```text
|
||||||
resource/YYYY-MM-DD/<resource_file>
|
resource/[YYYY-MM-DD/]<resource_file>
|
||||||
├─ step 1: daily/YYYY-MM-DD/<generated_name>.md # interpreted resource card
|
├─ step 1: daily/YYYY-MM-DD/<generated_name>.md # interpreted resource card
|
||||||
├─ step 2: source_resource points to the original resource
|
├─ step 2: source_resource points to the original resource
|
||||||
└─ step 3: daily/YYYY-MM-DD.md # daily index linking the cards
|
└─ step 3: daily/YYYY-MM-DD.md # daily index linking the cards
|
||||||
|
|
@ -21,8 +21,8 @@ resource/YYYY-MM-DD/<resource_file>
|
||||||
|
|
||||||
## What It Records
|
## What It Records
|
||||||
|
|
||||||
Auto Resource does more than copy file content. It extracts information that will make the resource easier to retrieve and
|
Auto Resource does more than copy file content. It extracts information that will make the resource easier to retrieve
|
||||||
understand later:
|
and understand later:
|
||||||
|
|
||||||
- Core content: what the resource is mainly about.
|
- Core content: what the resource is mainly about.
|
||||||
- Structure: its sections, tables, fields, and data organization.
|
- Structure: its sections, tables, fields, and data organization.
|
||||||
|
|
@ -34,26 +34,93 @@ In short, it turns "a file was archived" into "the resource is usable."
|
||||||
|
|
||||||
## Original Resource Entry Point
|
## Original Resource Entry Point
|
||||||
|
|
||||||
Auto Resource uses `resource/` as the entry point for source material. Resources must be placed under a date, which determines
|
Auto Resource uses `resource/` as the entry point for source material. Date directories are recommended, and their date
|
||||||
the day whose daily memory layer receives the interpreted card.
|
determines which daily memory layer receives the interpreted card. A file directly under `resource/` is also supported
|
||||||
|
and uses today in the application timezone when it is first processed. On later days, an exact `source_resource` match
|
||||||
|
keeps updates and deletion tied to that original daily card instead of creating a new card or leaving an orphan.
|
||||||
|
|
||||||
Example directory:
|
Example directory:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
workspace/
|
workspace/
|
||||||
resource/
|
resource/
|
||||||
|
quick-note.txt # enters today's daily layer
|
||||||
2026-06-20/
|
2026-06-20/
|
||||||
market-report.md
|
market-report.md
|
||||||
meeting-notes.csv
|
meeting-notes.csv
|
||||||
```
|
```
|
||||||
|
|
||||||
The current Beta version is best suited to text-based resources such as `md`, `txt`, `json`, `jsonl`, `csv`, `yaml`,
|
Text resources such as `md`, `txt`, `json`, `jsonl`, `csv`, `yaml`, and `html` are the primary fit. Image resources
|
||||||
and `html`.
|
(`png`, `jpg`, `jpeg`, `webp`, `gif`, `bmp`, `tiff`, `heic`) produce caption cards as described in
|
||||||
|
[Image Resources](#image-resources).
|
||||||
|
|
||||||
|
Internally, one `AutoResourceStep` receives each change batch and sends every item to the first configured processor
|
||||||
|
whose class-level matcher accepts it. `AutoImageResourceStep` handles image suffixes and `AutoTextResourceStep` is the
|
||||||
|
final fallback. A new modality can therefore add a registered processor, its prompt, and one `dispatch_steps` entry
|
||||||
|
without changing the router.
|
||||||
|
|
||||||
|
## Image Resources
|
||||||
|
|
||||||
|
Text and image resources share the same agent-wrapper and note-writing tools. Image inputs add a native AgentScope
|
||||||
|
image block alongside the interpretation instructions; the agent writes a caption card linked to the original image.
|
||||||
|
The card body starts with an `![[resource/...]]` embed link and the frontmatter carries `kind: image` and `media_type`,
|
||||||
|
so text search reaches image content through the caption.
|
||||||
|
|
||||||
|
Image processing is enabled by default (`include_images=true`) and requires an AgentScope wrapper bound to a compatible
|
||||||
|
model and formatter. Configure the model through `components.agent_wrapper.<name>.as_llm`, selecting the wrapper with
|
||||||
|
`agent_wrapper` on the resource Step. The former image-Step `as_llm` override and automatic `as_llm.vision` selection
|
||||||
|
are replaced by that binding. There is no separate caption model, schema-extraction call, or text-only retry after an
|
||||||
|
agent failure. An agent workflow can make multiple model requests while using its tools.
|
||||||
|
|
||||||
|
Each image interpretation starts a new session. Use the returned `agent_session_id` to find its processing record;
|
||||||
|
reprocessing the same image still updates the original card. The body should contain the image embed followed by a
|
||||||
|
description or transcription under `## Caption`, not an empty caption or a JSON response. Leave `status` to later
|
||||||
|
processing steps and keep its existing value when updating the card.
|
||||||
|
|
||||||
|
Customize image instructions with `prompt_dict.resource_instructions` (`resource_instructions_zh` for Chinese).
|
||||||
|
Rename existing `user_message` / `user_message_zh` settings accordingly.
|
||||||
|
Shared create/update templates insert these instructions at `{resource_instructions}`; older templates
|
||||||
|
without the placeholder receive them at the end.
|
||||||
|
|
||||||
|
Set `include_images=false` on an `auto_resource` call or as a Job default to skip **all** image events, including
|
||||||
|
deletions. Call-time values override Job defaults; when neither is set, image processing is enabled. For the watcher, use
|
||||||
|
`jobs.resource_watch_loop.include_images=false`; for manual calls, use `jobs.auto_resource.include_images=false`.
|
||||||
|
The image processor reports each skip in the existing result and warning log; text processing is unchanged. Existing
|
||||||
|
image cards are left untouched, even if their source image is deleted. Re-enabling images does not replay skipped
|
||||||
|
events; explicitly submit the affected paths to `auto_resource` when compensation is needed. The wrapper's configured
|
||||||
|
image-count limit is respected and must allow at least one image per resource call; it is not increased automatically.
|
||||||
|
|
||||||
|
Configure the wrapper when starting the persistent service. For example, to allow one image per agent context:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start components.agent_wrapper.default.context_config.max_image_num=1
|
||||||
|
```
|
||||||
|
|
||||||
|
The watcher processes resource changes automatically. To explicitly reprocess an existing `resource/photo.png`, run
|
||||||
|
the client in another terminal using the same workspace:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme auto_resource include_images=true changes='[{"path":"resource/photo.png","change":"modified"}]'
|
||||||
|
```
|
||||||
|
|
||||||
|
Images wider or taller than 2048px are downscaled,
|
||||||
|
and provider-unfriendly formats are re-encoded, in memory for the request only; the original file under
|
||||||
|
`resource/` is never modified. Before a full decode, image dimensions are checked against a default limit of 40,000,000
|
||||||
|
pixels; images over the limit and Pillow decompression-bomb warnings fail only that resource. EXIF orientation is
|
||||||
|
applied to the in-memory request copy before resizing or conversion. Oversized JPEGs first use decoder-level
|
||||||
|
downsampling, followed by a final thumbnail pass when needed. The VLM request MIME and the card's frontmatter
|
||||||
|
`media_type` use the format Pillow detects from the image bytes, rather than trusting the filename extension. When an
|
||||||
|
image changes, its card is updated in place; when the image is deleted, the card is removed with it, provided image
|
||||||
|
processing is enabled.
|
||||||
|
|
||||||
|
Image preprocessing uses Pillow from the `core` extra. HEIC resources additionally require the optional
|
||||||
|
`image-heif` extra: `pip install "reme-ai[image-heif]"`. Other supported image formats do not load or require the HEIF
|
||||||
|
plugin.
|
||||||
|
|
||||||
## Resource Cards
|
## Resource Cards
|
||||||
|
|
||||||
Each resource file produces one daily resource card. The system initially uses the resource file's stem as a temporary path.
|
Each resource file produces one daily resource card. The system initially uses the resource file's stem as a temporary
|
||||||
After the agent writes the card, the file is renamed according to its frontmatter `name`:
|
path. After the matching processor writes the card, the file is renamed according to its frontmatter `name`:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
resource/2026-06-20/market-report.md
|
resource/2026-06-20/market-report.md
|
||||||
|
|
@ -67,14 +134,19 @@ The resource card links to the original file through frontmatter:
|
||||||
source_resource: "[[resource/2026-06-20/market-report.md]]"
|
source_resource: "[[resource/2026-06-20/market-report.md]]"
|
||||||
```
|
```
|
||||||
|
|
||||||
When a resource changes, Auto Resource finds and updates the corresponding card through `source_resource`. When a resource is
|
When a resource changes, Auto Resource finds and updates the corresponding card through an exact `source_resource`
|
||||||
deleted, its daily note is also removed. The older `daily/YYYY-MM-DD/<resource_stem>.md` naming convention remains supported
|
match. When an enabled resource is deleted, only the explicitly linked daily note is removed. A same-stem note without that
|
||||||
as a fallback.
|
provenance marker is treated as user-owned and left untouched; new resource cards use a collision-free path instead.
|
||||||
|
|
||||||
|
A failed call may still have changed a card; `modified` records whether the file changed. If the agent writes the card
|
||||||
|
and then fails or is cancelled, the written content stays on disk. ReMe tries to complete metadata and update the day's
|
||||||
|
index for the card linked through `source_resource`, while preserving the original error or cancellation. A failed
|
||||||
|
image-note format check also leaves the written content in place. Failed calls are not retried automatically.
|
||||||
|
|
||||||
## Daily Index
|
## Daily Index
|
||||||
|
|
||||||
Resource cards enter the same daily memory layer as Auto Memory cards. The day's `YYYY-MM-DD.md` page acts as an index and
|
Resource cards enter the same daily memory layer as Auto Memory cards. The day's `YYYY-MM-DD.md` page acts as an index
|
||||||
organizes those resource cards:
|
and organizes those resource cards:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
daily/
|
daily/
|
||||||
|
|
@ -91,11 +163,12 @@ resource, open its corresponding resource card.
|
||||||
|
|
||||||
The interpreted daily note is optimized for readability; the original resource is retained for trust and verification.
|
The interpreted daily note is optimized for readability; the original resource is retained for trust and verification.
|
||||||
|
|
||||||
Auto Resource does not move the original file. It remains under `resource/YYYY-MM-DD/`. Text resources can therefore enter
|
Auto Resource does not move the original file. It remains at its original path under `resource/`. Resources can
|
||||||
the daily memory flow while their source files stay in their original location.
|
therefore enter the daily memory flow while their source files stay in their original location.
|
||||||
|
|
||||||
## What Happens Next
|
## What Happens Next
|
||||||
|
|
||||||
Auto Resource only creates resource interpretations in the daily layer. To distill long-term knowledge from resources into
|
Auto Resource only creates resource interpretations in the daily layer. To distill long-term knowledge from resources
|
||||||
`digest/`, use [Auto Dream](./auto_dream.md). To search original resources, daily cards, and digest nodes, use
|
into `digest/`, use [Auto Dream](./auto_dream.md). The default live index covers daily cards and digest nodes. Manual
|
||||||
[Memory Search](./memory_search.md).
|
`reindex` only rebuilds search indexes from chunks already accepted by an ingestion path; it does not add the original
|
||||||
|
resource files to search. See [Memory Search](./memory_search.md).
|
||||||
|
|
|
||||||
172
docs/en/blog_20260920.md
Normal file
172
docs/en/blog_20260920.md
Normal file
|
|
@ -0,0 +1,172 @@
|
||||||
|
# ReMe Memory Tags
|
||||||
|
|
||||||
|
Any memory system used over the long term eventually runs into a deceptively simple problem: **as memories accumulate, how do you search only the right subset?**
|
||||||
|
|
||||||
|
Suppose you and an agent have discussed three projects, all involving a launch, a budget, and an owner. Six months later, you ask:
|
||||||
|
|
||||||
|
> "What else do we need to confirm before launch?"
|
||||||
|
|
||||||
|
There is nothing wrong with the question, but it provides too few cues. Keyword search may retrieve every document that mentions "launch," while semantic search may blend experiences from several similar projects. Both find memories with similar content, but neither necessarily knows which project, company, or person you mean right now.
|
||||||
|
|
||||||
|
Human recall rarely works this way. We seldom run a full-text search across every experience at once. Instead, we begin with a few cues: **the ones about Alice, Project A, or that discussion from last year.** Once the scope narrows, the details begin to surface.
|
||||||
|
|
||||||
|
That is why ReMe adds memory tags. Each Markdown memory can express not only what it says, but also who or what it is mainly about—and that cue can participate directly in retrieval.
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img src="../figure/reme-blog/reme-blog-memory-tags.svg" alt="ReMe builds an index from Markdown tags and filters the search scope" width="100%">
|
||||||
|
</p>
|
||||||
|
|
||||||
|
## Why Memory Tags?
|
||||||
|
|
||||||
|
ReMe already uses BM25 for keyword search, optional embeddings for semantic similarity, and Wikilinks for traversing relationships between memories. Memory tags do not replace any of them. They add another dimension: **retrieval scope.**
|
||||||
|
|
||||||
|
Think of the three mechanisms as answering different questions:
|
||||||
|
|
||||||
|
- The query answers, "What am I looking for now?"
|
||||||
|
- A Wikilink answers, "Which memories are related to this one?"
|
||||||
|
- A memory tag answers, "Which memories should I search first?"
|
||||||
|
|
||||||
|
For example, "How did we handle the budget overrun?" may apply to many projects. If the search also includes `Project_A`, the agent can first narrow the scope to files related to Project A, then look for the specific details about the overrun.
|
||||||
|
|
||||||
|
Directories cannot fully solve this problem. A meeting note may concern Alice, Project A, and a customer at the same time, but a file normally occupies only one place on disk. Tags give the same memory multiple entry points without changing its original directory structure.
|
||||||
|
|
||||||
|
## Let Each Memory Say Who or What It Is About
|
||||||
|
|
||||||
|
ReMe memories remain plain Markdown. Tags live directly in YAML frontmatter, for example:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
---
|
||||||
|
name: Project A pre-launch checklist
|
||||||
|
description: Alice confirmed the launch window, rollback conditions, and customer notification order.
|
||||||
|
memory_tags:
|
||||||
|
- Alice
|
||||||
|
- Project_A
|
||||||
|
---
|
||||||
|
|
||||||
|
Project A is scheduled to launch on Thursday evening. Complete regression
|
||||||
|
testing first and have Alice confirm the customer notification. Roll back if
|
||||||
|
the error rate exceeds the agreed threshold.
|
||||||
|
```
|
||||||
|
|
||||||
|
The default field is named `memory_tags`. The name is intentional: this is not a loose collection of broad article keywords. It answers a more stable question:
|
||||||
|
|
||||||
|
> **Which real-world person or thing is this Markdown memory about?**
|
||||||
|
|
||||||
|
An entity can be a person, organization, company, project, or asset—for example, `Alice`, `CATL`, `Project_A`, or `Gold`. Compared with broad topics such as "work," "important," or "meeting," entities make better anchors for long-term memory because people, organizations, and projects tend to recur across many conversations.
|
||||||
|
|
||||||
|
In the default configuration, Auto Memory (`auto_memory`, `auto_memory_cc`) and Auto Dream (`auto_dream`, `dream_cron`) generate these tags for daily and digest Markdown files actually added or modified during the current run. Before tagging, the workflow reads the full document and its existing frontmatter, then checks tags already used in the workspace. It prefers an existing spelling for the same entity so that `Project_A`, `project a`, and `项目A` do not silently become three separate tags. Manual imports and edits do not trigger automatic tagging; existing `memory_tags` values are synchronized to the Tag Index by the file-watching workflow.
|
||||||
|
|
||||||
|
By default, a file receives only its most important entity. Multiple tags are used only when the document genuinely centers on multiple independent entities, and the total remains limited. A document without a clear core entity can use an empty list:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
memory_tags: []
|
||||||
|
```
|
||||||
|
|
||||||
|
This matters more than tagging for its own sake. More tags do not make a memory richer; too many broad tags only turn every filtered search back into a workspace-wide search.
|
||||||
|
|
||||||
|
Of course, `memory_tags` is only ReMe's default convention. The frontmatter field read by the tag index is configurable, and tag values remain under the user's control. Teams that already use `entities`, `people`, or another field can adapt the index to their files instead of migrating Markdown into a closed format.
|
||||||
|
|
||||||
|
## How Does the Tag Index Work?
|
||||||
|
|
||||||
|
After reading frontmatter, ReMe builds two simple cue maps: which files belong to a tag, and which tags belong to a file. For example:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Alice -> daily/project-a-launch.md
|
||||||
|
Project_A -> daily/project-a-launch.md
|
||||||
|
|
||||||
|
daily/project-a-launch.md -> Alice, Project_A
|
||||||
|
```
|
||||||
|
|
||||||
|
This is a bidirectional index derived from Markdown files. Relationships update when memories are created or modified, and stale relationships disappear when files are deleted. Tag comparison is case-insensitive and normalizes details such as whitespace, reducing accidental splits caused by spelling variations.
|
||||||
|
|
||||||
|
The index does not replace files or become a new source of truth. The real tags remain in user-visible, editable frontmatter. If the index is lost, it can be rebuilt from the Markdown metadata in the current file graph:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme reindex scope=tag
|
||||||
|
```
|
||||||
|
|
||||||
|
This follows ReMe's usual principle: **files belong to the user, indexes serve the files, and indexes are always rebuildable.**
|
||||||
|
|
||||||
|
To inspect the tags in a workspace and see how many files each tag covers, list them directly:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme list_tags order_by=file_count order=desc
|
||||||
|
```
|
||||||
|
|
||||||
|
Besides supporting search, this makes the structure of the memory workspace observable. You can quickly see that a project has accumulated many memories, or notice that one person's name has been split across several near-duplicate spellings.
|
||||||
|
|
||||||
|
## How Do Tags Participate in Search?
|
||||||
|
|
||||||
|
The most important role of memory tags is not display, but filtering.
|
||||||
|
|
||||||
|
Consider the earlier example. A natural-language query by itself looks like this:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme search query="What else do we need to confirm before launch?"
|
||||||
|
```
|
||||||
|
|
||||||
|
That searches the entire searchable memory scope. Add a tag:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme search \
|
||||||
|
query="What else do we need to confirm before launch?" \
|
||||||
|
tags='["Project_A"]'
|
||||||
|
```
|
||||||
|
|
||||||
|
ReMe first uses the Tag Index to find files tagged `Project_A`. BM25 and optional vector retrieval then produce direct matches only from those files, after which ranking fusion proceeds as usual.
|
||||||
|
|
||||||
|
There is one important boundary: tag filtering constrains direct retrieval hits, but it does not cut off Wikilink relationships. Default link expansion may still list the paths, names, and descriptions of neighboring memories outside the tag scope so the agent can decide whether to read further. Those neighbors do not become direct keyword or vector-search hits merely because they were listed.
|
||||||
|
|
||||||
|
The flow can be summarized as follows:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Natural-language question + tag cue
|
||||||
|
↓
|
||||||
|
Tag Index identifies candidate files
|
||||||
|
↓
|
||||||
|
Keyword / semantic search within those files
|
||||||
|
↓
|
||||||
|
Return direct matching passages and optionally list relationships
|
||||||
|
(related neighbors may fall outside the tag scope)
|
||||||
|
```
|
||||||
|
|
||||||
|
Tag filtering can also be combined with date conditions—for example, to inspect memories created for a project during the last month. Each condition narrows a separate dimension: the entity specifies who or what, the date specifies when, and the query specifies what you want to know.
|
||||||
|
|
||||||
|
When several tags are supplied, ReMe currently keeps files that match any of them. For example, `tags=[Alice, Project_A]` retrieves memories about Alice or Project A, then lets the query determine which results rank first. This lets an agent widen the candidate set with several plausible entity cues without returning to a workspace-wide search.
|
||||||
|
|
||||||
|
## What Changes in Practice?
|
||||||
|
|
||||||
|
Memory tags do not make a tag mandatory for every search. Searches without tags continue to work as before. The real change is that when a user or agent already knows part of the context, that context no longer has to remain hidden inside a vague query.
|
||||||
|
|
||||||
|
### 1. The Same Question Is Less Likely to Drift into Another Project
|
||||||
|
|
||||||
|
"Why was it delayed last time?", "Who approved the budget?", and "What remains before launch?" all depend heavily on context. Tags establish the project or person first, reducing the chance that memories with similar names or content enter the candidate set.
|
||||||
|
|
||||||
|
### 2. Memories About the Same Entity Can Accumulate Across Time
|
||||||
|
|
||||||
|
Alice may appear in meeting notes, project decisions, personal preferences, and retrospectives. Those files do not need to move into one directory. A shared tag creates an entity view across directories and dates.
|
||||||
|
|
||||||
|
### 3. Memory Structure Is Visible to Both People and Agents
|
||||||
|
|
||||||
|
Tags are not internal fields hidden in a specialized database. Users can open, edit, and review them in Markdown. An agent can inspect the tags that exist before deciding which entity cue to include in a search. Incorrect tags can be found, and naming can converge over time.
|
||||||
|
|
||||||
|
### 4. Search Becomes Easier to Explain
|
||||||
|
|
||||||
|
When a result is unexpected, the pipeline can be inspected step by step: does the document contain the right `memory_tags`, does the Tag Index include the path, or did keyword and semantic ranking fail to match it? This chain is easier to diagnose and correct than one opaque relevance score.
|
||||||
|
|
||||||
|
## Tags Are Retrieval Cues, Not a Taxonomy
|
||||||
|
|
||||||
|
The goal is not to turn a personal knowledge base into a carefully maintained classification tree. Real memories naturally overlap: one conversation may involve both a person and a project, while one decision may belong to today's meeting and shape a retrospective months later.
|
||||||
|
|
||||||
|
Memory tags are closer to the retrieval cues used by human memory. Seeing a person's name reminds us of shared experiences; thinking about a project brings related decisions, problems, and commitments to mind. A cue is not the memory itself, but it helps us enter the right context faster.
|
||||||
|
|
||||||
|
What ReMe does is deliberately simple:
|
||||||
|
|
||||||
|
- Preserve complete, readable memories in Markdown.
|
||||||
|
- Use `memory_tags` to express who or what a memory is about.
|
||||||
|
- Connect entities and files through a rebuildable Tag Index.
|
||||||
|
- Narrow the scope by tag before using keywords, semantics, and links to find the answer.
|
||||||
|
|
||||||
|
In this way, memory becomes more than a collection of full-text-searchable documents. It begins to acquire a structure that better matches how people associate ideas.
|
||||||
|
|
||||||
|
When you say, "That Alice project from last time," the agent receives more than a sentence. It receives a cue it can actually follow back into the past.
|
||||||
178
docs/en/configuration.md
Normal file
178
docs/en/configuration.md
Normal file
|
|
@ -0,0 +1,178 @@
|
||||||
|
---
|
||||||
|
title: Configuration
|
||||||
|
description: ReMe configuration files, environment expansion, command-line overrides, and core components.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Configuration
|
||||||
|
|
||||||
|
ReMe uses YAML or JSON to describe its Service, Jobs, and Components. The built-in default is `reme/config/default.yaml`. Select another configuration at startup and apply command-line overrides when needed.
|
||||||
|
|
||||||
|
## Precedence
|
||||||
|
|
||||||
|
Configuration is merged in this order, with later values winning:
|
||||||
|
|
||||||
|
1. `application_defaults` from enabled plugins.
|
||||||
|
2. The selected file; `default` is used when none is specified.
|
||||||
|
3. CLI dot-notation overrides.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start
|
||||||
|
reme start config=demo
|
||||||
|
reme start config=cookbook
|
||||||
|
reme start config=/absolute/path/to/app.yaml
|
||||||
|
reme start service.port=8181 workspace_dir=/data/reme
|
||||||
|
```
|
||||||
|
|
||||||
|
`config` accepts a built-in name or a `.yaml`, `.yml`, or `.json` file. Overrides are deep-merged, so changing `service.port` preserves sibling service settings.
|
||||||
|
The optional `cookbook` variant extends `default` and composes the separately installed Auto Fin, Daily Paper, and
|
||||||
|
DingTalk plugins. It requires the three DingTalk application credential environment variables before configuration
|
||||||
|
loading. It also enables `text-embedding-v4` vector retrieval, uses AgentScope with
|
||||||
|
`${LLM_MODEL_NAME:-qwen3.8-max}` by default, and runs the DingTalk bridge through Claude Code with the same
|
||||||
|
`LLM_MODEL_NAME` and `LLM_API_KEY`.
|
||||||
|
|
||||||
|
## CLI values
|
||||||
|
|
||||||
|
Arguments use `key=value`; leading `-` or `--` is accepted:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start --service.port=8181 --service.web_enabled=false
|
||||||
|
```
|
||||||
|
|
||||||
|
Values support null, booleans, numbers, JSON arrays and objects, quoted JSON strings, and plain strings. Numeric-looking values with leading zeroes, such as `007`, remain strings. Quote values such as `"true"` in JSON when they must remain strings.
|
||||||
|
|
||||||
|
## Environment variables
|
||||||
|
|
||||||
|
Configuration recursively expands:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
api_key: ${LLM_API_KEY}
|
||||||
|
base_url: ${LLM_BASE_URL:-https://example.com/v1}
|
||||||
|
```
|
||||||
|
|
||||||
|
`${VAR}` fails when undefined; `${VAR:-default}` uses its fallback. ReMe also searches for `.env` from the command's working directory through at most five parents.
|
||||||
|
|
||||||
|
Keep secrets in `.env` or the process environment, never in committed configuration.
|
||||||
|
|
||||||
|
## Application fields
|
||||||
|
|
||||||
|
| Field | Default | Purpose |
|
||||||
|
|---|---|---|
|
||||||
|
| `app_name` | `ReMe` | Display name |
|
||||||
|
| `workspace_dir` | `.reme` | User-owned workspace root, normalized to an absolute path |
|
||||||
|
| `metadata_dir` | `metadata` | Rebuildable indexes, graphs, and catalogs |
|
||||||
|
| `session_dir` | `session` | Agent sessions; standard transcripts use `session/dialog` |
|
||||||
|
| `mem_session_dir` | `mem_session` | Agent-wrapper sessions and configuration |
|
||||||
|
| `resource_dir` | `resource` | External resources |
|
||||||
|
| `daily_dir` | `daily` | Daily memory |
|
||||||
|
| `digest_dir` | `digest` | Consolidated long-term memory |
|
||||||
|
| `timezone` | `Asia/Shanghai` | IANA timezone used for dates and cron jobs |
|
||||||
|
| `language` | empty | Default language for LLM interactions |
|
||||||
|
| `plugins` | `[]` | Installed plugins enabled for this Application |
|
||||||
|
| `service` | HTTP | Service configuration |
|
||||||
|
| `jobs` | default Jobs | Job configurations by name |
|
||||||
|
| `components` | defaults | Components grouped by type and name |
|
||||||
|
|
||||||
|
`session_dir` must remain workspace-relative.
|
||||||
|
|
||||||
|
## Enabling jobs
|
||||||
|
|
||||||
|
`jobs.<name>.enabled` defaults to `true` for all Job types. Disabled jobs retain their configuration but do not
|
||||||
|
start, expose service interfaces, or accept calls through `Application.run_job()` / `run_stream_job()`.
|
||||||
|
For example, disable ReMe's Dream cron when a host plugin owns the schedule:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start jobs.dream_cron.enabled=false
|
||||||
|
```
|
||||||
|
|
||||||
|
Restart the service to apply the override. The separate `auto_dream` API remains available, and other jobs continue running.
|
||||||
|
`enable_serve` independently controls service exposure: `enabled=true, enable_serve=false` keeps a Job available for
|
||||||
|
local calls. Background and cron jobs are never service-exposed.
|
||||||
|
|
||||||
|
## LLM
|
||||||
|
|
||||||
|
The default LLM uses an OpenAI-compatible interface:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
components:
|
||||||
|
as_llm:
|
||||||
|
default:
|
||||||
|
backend: openai
|
||||||
|
model: qwen3.7-plus
|
||||||
|
context_size: 200000
|
||||||
|
credential:
|
||||||
|
api_key: ${LLM_API_KEY:-}
|
||||||
|
base_url: ${LLM_BASE_URL:-}
|
||||||
|
```
|
||||||
|
|
||||||
|
Built-in registrations include `openai`, `anthropic`, `dashscope`, `deepseek`, `gemini`, `moonshot`, `ollama`, and `xai`. Their detailed model fields follow the corresponding AgentScope wrappers.
|
||||||
|
|
||||||
|
File operations, BM25 search, wikilink traversal, and `proactive_read` do not require an LLM. Evolution workflows such
|
||||||
|
as `auto_memory`, `auto_resource`, `auto_dream`, and proactive refresh do.
|
||||||
|
|
||||||
|
## Embeddings
|
||||||
|
|
||||||
|
Vector retrieval is disabled by default. Credentials alone do not enable it: configure `as_embedding`, `embedding_store`, and connect the store to `file_store`.
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
components:
|
||||||
|
as_embedding:
|
||||||
|
default:
|
||||||
|
backend: openai
|
||||||
|
model: text-embedding-v4
|
||||||
|
dimensions: 1024
|
||||||
|
credential:
|
||||||
|
api_key: ${EMBEDDING_API_KEY}
|
||||||
|
base_url: ${EMBEDDING_BASE_URL:-https://dashscope.aliyuncs.com/compatible-mode/v1}
|
||||||
|
embedding_store:
|
||||||
|
default:
|
||||||
|
backend: local
|
||||||
|
as_embedding: default
|
||||||
|
file_store:
|
||||||
|
default:
|
||||||
|
backend: local
|
||||||
|
embedding_store: default
|
||||||
|
keyword_index: default
|
||||||
|
file_graph: default
|
||||||
|
```
|
||||||
|
|
||||||
|
Rebuild the embedding index after changing the model or dimensions.
|
||||||
|
|
||||||
|
## Service and Jobs
|
||||||
|
|
||||||
|
Minimal HTTP configuration:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
service:
|
||||||
|
backend: http
|
||||||
|
host: 127.0.0.1
|
||||||
|
port: 2333
|
||||||
|
web_enabled: true
|
||||||
|
mcp_enabled: true
|
||||||
|
mcp_path: /mcp
|
||||||
|
```
|
||||||
|
|
||||||
|
A Job declares a backend, parameter schema, and ordered Steps:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
jobs:
|
||||||
|
example:
|
||||||
|
backend: base
|
||||||
|
description: Example job
|
||||||
|
parameters:
|
||||||
|
type: object
|
||||||
|
properties:
|
||||||
|
text: { type: string }
|
||||||
|
required: [text]
|
||||||
|
steps:
|
||||||
|
- backend: example_step
|
||||||
|
```
|
||||||
|
|
||||||
|
Set `enable_serve: false` to keep a Job internal. Background and cron Jobs are never service-exposed.
|
||||||
|
|
||||||
|
## Inspect the effective configuration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme app_config
|
||||||
|
```
|
||||||
|
|
||||||
|
The result is the merged, validated configuration with secrets redacted. Use it when diagnosing plugin or override precedence. The authoritative contracts remain `reme/schema/application_config.py` and `reme/config/default.yaml`.
|
||||||
|
|
@ -8,12 +8,12 @@ ReMe is open source and hosted on GitHub:
|
||||||
|
|
||||||
## How to Contribute
|
## How to Contribute
|
||||||
|
|
||||||
Thank you for your interest in ReMe. ReMe is a file-first, self-evolving memory system for agents. Contributions are welcome
|
Thank you for your interest in ReMe. ReMe is a file-first, self-evolving memory system for agents. Contributions are
|
||||||
through issue reports, documentation improvements, additional tests, bug fixes, and new capabilities.
|
welcome through issue reports, documentation improvements, additional tests, bug fixes, and new capabilities.
|
||||||
|
|
||||||
If this is your first time running ReMe locally, start with [Quick Start](./quick_start.md). If your change affects runtime
|
If this is your first time running ReMe locally, start with [Quick Start](./quick_start.md). If your change affects
|
||||||
layers, Jobs, Steps, or components, read [ReMe Framework](./framework.md). If it affects workspace directories, frontmatter,
|
runtime layers, Jobs, Steps, or components, read [ReMe Framework](./framework.md). If it affects workspace directories,
|
||||||
wikilinks, or chunking, read [Memory as File](./memory_as_file.md).
|
frontmatter, wikilinks, or chunking, read [Memory as File](./memory_as_file.md).
|
||||||
|
|
||||||
### 1. Before You Begin
|
### 1. Before You Begin
|
||||||
|
|
||||||
|
|
@ -21,9 +21,10 @@ Before investing in an implementation:
|
||||||
|
|
||||||
- Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) for an existing issue or discussion.
|
- Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) for an existing issue or discussion.
|
||||||
- If a related issue is still open, comment that you would like to work on it to avoid duplicate effort.
|
- If a related issue is still open, comment that you would like to work on it to avoid duplicate effort.
|
||||||
- If no issue exists, create one describing the context, expected behavior, possible implementation, and scope of impact.
|
- If no issue exists, create one describing the context, expected behavior, possible implementation, and scope of
|
||||||
- For larger feature changes, align with maintainers on interfaces, configuration, compatibility, and test strategy before
|
impact.
|
||||||
submitting an implementation.
|
- For larger feature changes, align with maintainers on interfaces, configuration, compatibility, and test strategy
|
||||||
|
before submitting an implementation.
|
||||||
|
|
||||||
### 2. Local Development Environment
|
### 2. Local Development Environment
|
||||||
|
|
||||||
|
|
@ -38,7 +39,11 @@ The project requires Python 3.11 or later. A virtual environment is recommended:
|
||||||
```bash
|
```bash
|
||||||
python -m venv .venv
|
python -m venv .venv
|
||||||
source .venv/bin/activate
|
source .venv/bin/activate
|
||||||
pip install -e ".[dev,full]"
|
pip install -e reme_studio -e ".[dev,full]"
|
||||||
|
cd reme_studio
|
||||||
|
npm ci
|
||||||
|
npm run build:static
|
||||||
|
cd ..
|
||||||
pre-commit install
|
pre-commit install
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -53,8 +58,8 @@ CLI / Client -> Service -> Application -> Job -> Step -> Component / Workspace
|
||||||
|
|
||||||
In practice:
|
In practice:
|
||||||
|
|
||||||
- Capabilities exposed to users or external systems should normally be orchestrated by a Job, then exposed by a Service as a
|
- Capabilities exposed to users or external systems should normally be orchestrated by a Job, then exposed by a Service
|
||||||
CLI-, HTTP-, or MCP-callable interface.
|
as a CLI-, HTTP-, or MCP-callable interface.
|
||||||
- Reusable infrastructure belongs in `reme/components/`, with dependencies declared through `BaseComponent.bind()`.
|
- Reusable infrastructure belongs in `reme/components/`, with dependencies declared through `BaseComponent.bind()`.
|
||||||
- Atomic business operations belong in `reme/steps/` and access the file store, agent wrapper, catalog, LLM, and other
|
- Atomic business operations belong in `reme/steps/` and access the file store, agent wrapper, catalog, LLM, and other
|
||||||
components through `BaseStep.Ref`.
|
components through `BaseStep.Ref`.
|
||||||
|
|
@ -65,17 +70,19 @@ In practice:
|
||||||
|
|
||||||
When adding a Step or Job, pay particular attention to these conventions:
|
When adding a Step or Job, pay particular attention to these conventions:
|
||||||
|
|
||||||
- Register implementations with `@R.register("<backend_name>")`. Registration names should be stable, clear, and match the
|
- Register implementations with `@R.register("<backend_name>")`. Registration names should be stable, clear, and match
|
||||||
configured `backend`.
|
the configured `backend`.
|
||||||
- After adding a Step file, make sure its package `__init__.py` imports the module; otherwise, the registry will not load it.
|
- After adding a Step file, make sure its package `__init__.py` imports the module; otherwise, the registry will not
|
||||||
|
load it.
|
||||||
- A Step should perform one atomic business operation. Cross-step flows belong in Job configuration or a dedicated
|
- A Step should perform one atomic business operation. Cross-step flows belong in Job configuration or a dedicated
|
||||||
orchestration Step.
|
orchestration Step.
|
||||||
- A Job composes Steps and selects normal, streaming, background, or scheduled execution. `enable_serve` controls whether it
|
- A Job composes Steps and selects normal, streaming, background, or scheduled execution. `enable_serve` controls
|
||||||
is externally exposed.
|
whether it is externally exposed.
|
||||||
- When a Step needs components, prefer `BaseStep.Ref`. Do not reconstruct global components inside a Step or bypass
|
- When a Step needs components, prefer `BaseStep.Ref`. Do not reconstruct global components inside a Step or bypass
|
||||||
`ApplicationContext`.
|
`ApplicationContext`.
|
||||||
- File, index, graph, frontmatter, and wikilink behavior must preserve consistent workspace-relative path semantics.
|
- File, index, graph, frontmatter, and wikilink behavior must preserve consistent workspace-relative path semantics.
|
||||||
- Add fast tests under `tests/unit/` for new capabilities. Put cross-component, LLM, embedding, or service behavior under
|
- Add fast tests under `tests/unit/` for new capabilities. Put cross-component, LLM, embedding, or service behavior
|
||||||
|
under
|
||||||
`tests/integration/` when appropriate.
|
`tests/integration/` when appropriate.
|
||||||
|
|
||||||
### 4. Code and Documentation Changes
|
### 4. Code and Documentation Changes
|
||||||
|
|
@ -83,7 +90,7 @@ When adding a Step or Job, pay particular attention to these conventions:
|
||||||
Choose the appropriate entry point for the type of change:
|
Choose the appropriate entry point for the type of change:
|
||||||
|
|
||||||
| Change type | Primary location | Guidance |
|
| Change type | Primary location | Guidance |
|
||||||
|---|---|---|
|
|-----------------------------------|-------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------|
|
||||||
| Configuration or startup behavior | `reme/config/`, `reme/application.py`, `reme/reme.py` | Keep the default configuration runnable and avoid breaking existing CLI, HTTP, and MCP entry points. |
|
| Configuration or startup behavior | `reme/config/`, `reme/application.py`, `reme/reme.py` | Keep the default configuration runnable and avoid breaking existing CLI, HTTP, and MCP entry points. |
|
||||||
| Component capability | `reme/components/` | Reuse `BaseComponent`, the registry, and context objects. |
|
| Component capability | `reme/components/` | Reuse `BaseComponent`, the registry, and context objects. |
|
||||||
| Job or Step | `reme/components/job/`, `reme/steps/` | Follow the Job -> Step model in [ReMe Framework](./framework.md), keep request and response schemas clear, and add corresponding tests. |
|
| Job or Step | `reme/components/job/`, `reme/steps/` | Follow the Job -> Step model in [ReMe Framework](./framework.md), keep request and response schemas clear, and add corresponding tests. |
|
||||||
|
|
@ -165,15 +172,16 @@ pytest tests/unit/test_reme_cli.py
|
||||||
|
|
||||||
If `pre-commit` modifies files automatically, commit those changes and rerun the checks until everything passes.
|
If `pre-commit` modifies files automatically, commit those changes and rerun the checks until everything passes.
|
||||||
|
|
||||||
The current pre-commit configuration includes YAML/TOML/JSON validation, private-key detection, trailing-whitespace checks,
|
The current pre-commit configuration includes YAML/TOML/JSON validation, private-key detection, trailing-whitespace
|
||||||
|
checks,
|
||||||
`black`, `flake8`, `pylint`, and `pyroma`. The main formatting rules are:
|
`black`, `flake8`, `pylint`, and `pyroma`. The main formatting rules are:
|
||||||
|
|
||||||
- `black --line-length=120`
|
- `black --line-length=120`
|
||||||
- `flake8 --max-line-length=120`
|
- `flake8 --max-line-length=120`
|
||||||
- `pylint --max-line-length=120`
|
- `pylint --max-line-length=120`
|
||||||
|
|
||||||
Some integration tests may require an LLM, embeddings, or external service configuration. If you cannot run them locally,
|
Some integration tests may require an LLM, embeddings, or external service configuration. If you cannot run them
|
||||||
state why they were skipped and what alternative validation you completed in the PR description.
|
locally, state why they were skipped and what alternative validation you completed in the PR description.
|
||||||
|
|
||||||
### 8. Testing Requirements
|
### 8. Testing Requirements
|
||||||
|
|
||||||
|
|
@ -183,7 +191,8 @@ Add tests according to the risk of the change:
|
||||||
- For a new Step, Job, or component, cover at least the main path and a failure path.
|
- For a new Step, Job, or component, cover at least the main path and a failure path.
|
||||||
- For changes to shared logic such as indexes, graphs, wikilinks, frontmatter, or file operations, add edge cases.
|
- For changes to shared logic such as indexes, graphs, wikilinks, frontmatter, or file operations, add edge cases.
|
||||||
- For changes to the CLI, services, or configuration parsing, cover the user-visible entry point.
|
- For changes to the CLI, services, or configuration parsing, cover the user-visible entry point.
|
||||||
- Documentation-only changes usually do not require new tests, but running `pre-commit run --all-files` is still recommended.
|
- Documentation-only changes usually do not require new tests, but running `pre-commit run --all-files` is still
|
||||||
|
recommended.
|
||||||
|
|
||||||
Place tests according to the existing structure:
|
Place tests according to the existing structure:
|
||||||
|
|
||||||
|
|
@ -200,6 +209,13 @@ Documentation lives under:
|
||||||
docs/
|
docs/
|
||||||
```
|
```
|
||||||
|
|
||||||
|
User guides should have matching `docs/zh/` and `docs/en/` versions and appear in the corresponding navigation in
|
||||||
|
`docs/.vitepress/config.mts`. The ReMe Studio, TypeScript, plugin, and benchmark READMEs remain canonical in their own
|
||||||
|
directories; `github-pages/scripts/generate-content.mjs` mirrors them during builds. Never edit `.generated/` or `dist/`.
|
||||||
|
|
||||||
|
The Job API reference is generated from `reme/config/default.yaml`. Update that YAML and its tests when a default Job
|
||||||
|
contract changes rather than editing generated pages.
|
||||||
|
|
||||||
Documentation should:
|
Documentation should:
|
||||||
|
|
||||||
- Use clear titles that directly identify a capability or flow.
|
- Use clear titles that directly identify a capability or flow.
|
||||||
|
|
@ -207,15 +223,24 @@ Documentation should:
|
||||||
- Use real repository paths such as `reme/config/default.yaml`, `reme/steps/`, and `tests/unit/`.
|
- Use real repository paths such as `reme/config/default.yaml`, `reme/steps/`, and `tests/unit/`.
|
||||||
- Describe default behavior according to the current code, `pyproject.toml`, and default configuration.
|
- Describe default behavior according to the current code, `pyproject.toml`, and default configuration.
|
||||||
|
|
||||||
|
Validate the documentation site with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd github-pages
|
||||||
|
npm ci
|
||||||
|
npm test
|
||||||
|
npm run build
|
||||||
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Getting Help
|
## Getting Help
|
||||||
|
|
||||||
- Bugs and feature requests: [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues)
|
- Bugs and feature requests: [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues)
|
||||||
- Project home: [GitHub Repository](https://github.com/agentscope-ai/ReMe)
|
- Project home: [GitHub Repository](https://github.com/agentscope-ai/ReMe)
|
||||||
- Documentation site: [https://reme.agentscope.io/](https://reme.agentscope.io/)
|
- Documentation site: [https://reme.agentscope.io](https://reme.agentscope.io)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
Thank you for contributing to ReMe. Your improvements help make long-term memory for agents more readable, controllable, and
|
Thank you for contributing to ReMe. Your improvements help make long-term memory for agents more readable, controllable,
|
||||||
maintainable.
|
and maintainable.
|
||||||
|
|
|
||||||
156
docs/en/docker.md
Normal file
156
docs/en/docker.md
Normal file
|
|
@ -0,0 +1,156 @@
|
||||||
|
---
|
||||||
|
title: Docker Deployment
|
||||||
|
description: Run ReMe and Studio in Docker with a user-owned persistent workspace.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Docker Deployment
|
||||||
|
|
||||||
|
The image includes ReMe's `core` and HEIF image dependencies and the Studio static frontend. One HTTP process serves the API,
|
||||||
|
Studio at `/`, and MCP at `/mcp`. The default configuration keeps embeddings disabled; file operations and BM25 search do
|
||||||
|
not require model credentials.
|
||||||
|
|
||||||
|
## Build and run with Compose
|
||||||
|
|
||||||
|
Use Docker Engine or Docker Desktop with Compose **2.24.0 or newer**. From the repository root:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p .reme
|
||||||
|
docker compose up --build -d
|
||||||
|
docker compose logs -f reme
|
||||||
|
```
|
||||||
|
|
||||||
|
Open <http://127.0.0.1:2333>. Compose binds the host port to loopback and mounts `./.reme` at `/data`. The source checkout and
|
||||||
|
Studio assets are not mounted over the installed application.
|
||||||
|
|
||||||
|
On Linux, if your workspace is not owned by UID/GID 1000, set the process identity before starting:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export REME_UID=$(id -u)
|
||||||
|
export REME_GID=$(id -g)
|
||||||
|
docker compose up --build -d
|
||||||
|
```
|
||||||
|
|
||||||
|
For model-powered memory evolution, copy `deploy/docker/example.env` to `.env` if you do not already have one, then fill in
|
||||||
|
your model credentials. Compose injects this optional file at runtime; Docker builds exclude `.env` files. Compose's
|
||||||
|
`--env-file` controls variable interpolation; `REME_ENV_FILE` selects the file injected into the container.
|
||||||
|
|
||||||
|
## Use a published image
|
||||||
|
|
||||||
|
The Docker workflow publishes `ghcr.io/agentscope-ai/reme:main` after successful main-branch checks. Stable GitHub releases
|
||||||
|
publish their package version and `latest`; prereleases do not update `latest`. Both Linux amd64 and arm64 images are tested
|
||||||
|
before their combined tags are published. Publication starts when the workflow is enabled in the repository.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mkdir -p "$HOME/.reme"
|
||||||
|
docker run -d --name reme \
|
||||||
|
--user "$(id -u):$(id -g)" \
|
||||||
|
-p 127.0.0.1:2333:2333 \
|
||||||
|
--mount "type=bind,source=$HOME/.reme,target=/data" \
|
||||||
|
--restart unless-stopped \
|
||||||
|
ghcr.io/agentscope-ai/reme:main
|
||||||
|
```
|
||||||
|
|
||||||
|
Add `--env-file /path/to/model.env` before the image name when using model credentials. For reproducible deployments,
|
||||||
|
replace `main` with a released version or image digest. To use the image with Compose, set `REME_IMAGE`, then run
|
||||||
|
`docker compose pull` and `docker compose up -d --no-build`.
|
||||||
|
|
||||||
|
## Paths, configuration, and ports
|
||||||
|
|
||||||
|
The image runs as UID/GID 1000 by default. `/data` contains the entire workspace: source sessions, resources, daily notes,
|
||||||
|
digest notes, and rebuildable metadata. Create the host directory yourself and make it writable by the configured user.
|
||||||
|
Mounting only `metadata/` does not preserve the source memories. Paths in a custom configuration refer to the container's
|
||||||
|
filesystem; additional paths need additional mounts.
|
||||||
|
|
||||||
|
| Setting | Meaning |
|
||||||
|
|---|---|
|
||||||
|
| `REME_WORKSPACE_DIR` | Container workspace; image default `/data`, fixed to `/data` by Compose |
|
||||||
|
| `REME_CONFIG` | Existing config name or mounted YAML/JSON path; unset uses the built-in default |
|
||||||
|
| `REME_HOST` | HTTP bind address; image and Compose use `0.0.0.0` |
|
||||||
|
| `REME_PORT` | Container port override; Compose defaults to `2333` |
|
||||||
|
| `REME_TIMEZONE` | Optional application timezone override; otherwise the application default applies |
|
||||||
|
| `REME_DATA_DIR` | Compose host workspace directory; default `./.reme` |
|
||||||
|
| `REME_PUBLISHED_PORT` | Compose host port; default `2333`, independent of the container port |
|
||||||
|
| `REME_BIND_ADDRESS` | Compose host bind address; default `127.0.0.1` |
|
||||||
|
| `REME_UID`, `REME_GID` | Compose process identity; both default to `1000` |
|
||||||
|
| `REME_ENV_FILE` | Optional Compose runtime environment file; default `.env` |
|
||||||
|
|
||||||
|
Explicit `start key=value` arguments override container environment settings, which override the loaded configuration for
|
||||||
|
those keys. Other keys retain ReMe's normal deep merge behavior. File logging defaults to off in the image; use container
|
||||||
|
logs. `log_to_file=true` explicitly enables file logs under `/app/logs`, which requires a separate mount to persist them.
|
||||||
|
The temporary home directory `/tmp/reme-home` and probe address file are disposable, not workspace storage.
|
||||||
|
|
||||||
|
To customize the full job/component configuration, copy `reme/config/default.yaml` to `reme.yaml`, edit it, set
|
||||||
|
`REME_CONFIG=/etc/reme/config.yaml` in `.env`, and add `compose.override.yaml`:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
services:
|
||||||
|
reme:
|
||||||
|
volumes:
|
||||||
|
- ./reme.yaml:/etc/reme/config.yaml:ro
|
||||||
|
```
|
||||||
|
|
||||||
|
Keep the `health_check` Job enabled and in `service.jobs` if you use an allowlist. The image's probe uses HTTP; when
|
||||||
|
overriding the service to CLI or MCP stdio, disable the Docker health check with `--no-healthcheck` (or Compose
|
||||||
|
`healthcheck: {disable: true}`).
|
||||||
|
|
||||||
|
To override startup without changing the image:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --rm -p 127.0.0.1:2444:2444 \
|
||||||
|
--mount "type=bind,source=$HOME/.reme,target=/data" \
|
||||||
|
reme:local start service.port=2444 timezone=UTC
|
||||||
|
docker compose exec reme reme health_check
|
||||||
|
docker compose exec reme reme status
|
||||||
|
```
|
||||||
|
|
||||||
|
Other commands pass through unchanged, including `reme start job=version` for a one-shot job or `python` for diagnostics.
|
||||||
|
From a host CLI, supply the published address explicitly, for example `reme health_check host=127.0.0.1 port=2444`.
|
||||||
|
Host process discovery cannot reconstruct a container's startup arguments. Configure agent integrations to use the
|
||||||
|
published HTTP or MCP endpoint instead of starting a second native ReMe on the same workspace.
|
||||||
|
|
||||||
|
## Networking and optional tools
|
||||||
|
|
||||||
|
`127.0.0.1` inside a container refers to that container. A model server on the host needs a reachable host address, such as
|
||||||
|
`host.docker.internal` on Docker Desktop. On Linux, add `extra_hosts: ["host.docker.internal:host-gateway"]` to the service
|
||||||
|
and configure the model URL accordingly. Another Compose service is reachable by its service name.
|
||||||
|
|
||||||
|
The HTTP action API has no built-in authentication and includes write/delete operations. Keep the default loopback port
|
||||||
|
publication. For access from another machine, place an authenticated TLS proxy in front of the service and restrict direct
|
||||||
|
access to its port; the same restriction must cover Studio, HTTP Jobs, and MCP.
|
||||||
|
|
||||||
|
The image includes the configured agent SDK dependencies, but host OAuth files, transcripts, plugins, external MCP
|
||||||
|
executables, and host workspace paths are not automatically available. Mount required data explicitly and install extra
|
||||||
|
plugins/tools in a derived image so that recreating the container preserves the installation. Keep credentials out of
|
||||||
|
Docker build arguments and layers. Optional FAISS/zvec backends also depend on the capabilities of the target machine.
|
||||||
|
|
||||||
|
## Health, upgrade, and recovery
|
||||||
|
|
||||||
|
Docker posts to the existing `/health_check` Job and requires both `success=true` and `metadata.health.healthy=true`.
|
||||||
|
The probe address follows effective configuration and CLI port overrides, and bypasses outbound proxy settings.
|
||||||
|
Initialization has a 120-second health grace period; larger workspaces may need a longer Compose `healthcheck.start_period`.
|
||||||
|
This reports component health, not whether a remote model will accept a future request. An unhealthy Docker status alone
|
||||||
|
does not trigger `restart: unless-stopped`; that policy restarts exited processes.
|
||||||
|
|
||||||
|
Before upgrading, stop writes and back up the **complete** host workspace and your deployment configuration. Then:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker compose stop
|
||||||
|
# Back up the configured host workspace here.
|
||||||
|
docker compose pull
|
||||||
|
docker compose up -d --no-build
|
||||||
|
docker compose exec reme reme health_check
|
||||||
|
```
|
||||||
|
|
||||||
|
For a locally built deployment, replace `pull` and `up --no-build` with `docker compose up --build -d`. Compose allows
|
||||||
|
60 seconds for orderly shutdown. Recreating containers leaves the bind-mounted workspace intact; use one ReMe writer
|
||||||
|
process per workspace. See [backup and recovery](./operations.md) for restoring derived state without deleting memory.
|
||||||
|
|
||||||
|
For container validation after a local build:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker build -t reme:local .
|
||||||
|
python scripts/test_docker_image.py --image reme:local
|
||||||
|
```
|
||||||
|
|
||||||
|
The smoke check uses a disposable workspace and no model credentials. It verifies Studio, HTTP and MCP, file containment,
|
||||||
|
non-root execution, graceful shutdown, and memory search after replacing a container on a different port.
|
||||||
72
docs/en/faq.md
Normal file
72
docs/en/faq.md
Normal file
|
|
@ -0,0 +1,72 @@
|
||||||
|
---
|
||||||
|
title: Frequently Asked Questions
|
||||||
|
description: Quick answers for ReMe installation, services, models, retrieval, files, and plugins.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Frequently Asked Questions
|
||||||
|
|
||||||
|
## Do basic file operations require a model API key?
|
||||||
|
|
||||||
|
No. `write`, `read`, `list`, `stat`, BM25 search, wikilink traversal, and `proactive_read` work without model
|
||||||
|
credentials. `auto_memory`, `auto_resource`, `auto_dream`, and proactive refresh require an LLM.
|
||||||
|
|
||||||
|
## Why is search still BM25-only after setting an embedding key?
|
||||||
|
|
||||||
|
Embeddings are disabled by default. Configure `as_embedding` and `embedding_store`, then connect `file_store.default.embedding_store` to that component. See [Configuration](./configuration.md#embeddings).
|
||||||
|
|
||||||
|
## Why did `reme reindex` not discover a new file?
|
||||||
|
|
||||||
|
`reindex` rebuilds indexes from current `file_chunks`; it does not scan the workspace. Check `index_update_loop`, the watched directory and extension, and `health_check`.
|
||||||
|
|
||||||
|
## How do I use another workspace?
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start workspace_dir=/absolute/path/to/memory
|
||||||
|
```
|
||||||
|
|
||||||
|
Ordinary CLI calls discover the running service, so they do not need the workspace argument again.
|
||||||
|
|
||||||
|
## What if port 2333 is occupied?
|
||||||
|
|
||||||
|
Do not stop an unknown listener. Select another port:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start service.port=8181
|
||||||
|
```
|
||||||
|
|
||||||
|
Then confirm it with `reme find_reme`.
|
||||||
|
|
||||||
|
## Why is an installed plugin missing its Jobs?
|
||||||
|
|
||||||
|
Installation only makes the distribution discoverable in the active Python environment. Enable it for the Application:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme start plugins='["auto-fin"]'
|
||||||
|
```
|
||||||
|
|
||||||
|
Restart a running service after changing package or enablement state.
|
||||||
|
|
||||||
|
## May I edit workspace Markdown directly?
|
||||||
|
|
||||||
|
Yes. Files are the source of truth and watchers ingest changes. Keep frontmatter valid, use complete workspace-relative wikilinks, and avoid unconditional concurrent saves.
|
||||||
|
|
||||||
|
## May I expose ReMe publicly?
|
||||||
|
|
||||||
|
Not with the default configuration alone. Jobs can write and delete, HTTP CORS is permissive, and there is no general authentication layer. Use a controlled network or authenticated TLS reverse proxy and restrict `service.jobs`.
|
||||||
|
|
||||||
|
## How should I back up and migrate memory?
|
||||||
|
|
||||||
|
Stop writes and back up the complete workspace. `session/`, `resource/`, `daily/`, and `digest/` are the key sources; `metadata/` can be backed up or rebuilt. See [Diagnostics, Backup, and Recovery](./operations.md).
|
||||||
|
|
||||||
|
## Why is Studio unavailable?
|
||||||
|
|
||||||
|
The base `reme-ai` package has no frontend assets. Install `reme-ai[web]` or `reme-ai[core]`, or set `service.web_static_dir`. Missing Studio assets do not disable the Job API.
|
||||||
|
|
||||||
|
## Which capabilities does the running service expose?
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme help
|
||||||
|
reme app_config
|
||||||
|
```
|
||||||
|
|
||||||
|
Static documentation describes defaults; plugins and custom configuration may change the active service.
|
||||||
|
|
@ -16,13 +16,13 @@ To run and use ReMe first, see [Quick Start](./quick_start.md). For workspace fi
|
||||||
|
|
||||||
### Capability Boundary
|
### Capability Boundary
|
||||||
|
|
||||||
ReMe v4 focuses on long-term memory: it distills conversations and resources into `daily/`, organizes them into `digest/`,
|
ReMe v4 focuses on long-term memory: it distills conversations and resources into `daily/`, organizes them into
|
||||||
and exposes write, retrieval, and proactive-read capabilities through the CLI, HTTP, and MCP.
|
`digest/`, and exposes write, retrieval, and proactive-read capabilities through the CLI, HTTP, and MCP.
|
||||||
|
|
||||||
Single-session context-window management is outside the scope of ReMe v4. This includes compressing the current conversation,
|
Single-session context-window management is outside the scope of ReMe v4. This includes compressing the current
|
||||||
injecting summaries, trimming tool output, or providing an independent `/compact` interface. Those capabilities belong in
|
conversation, injecting summaries, trimming tool output, or providing an independent `/compact` interface. Those
|
||||||
the host agent framework. ReMe accepts conversations, resources, and file changes that have already occurred and persists the
|
capabilities belong in the host agent framework. ReMe accepts conversations, resources, and file changes that have
|
||||||
information with long-term value.
|
already occurred and persists the information with long-term value.
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
|
|
@ -39,7 +39,7 @@ flowchart LR
|
||||||
Core layers:
|
Core layers:
|
||||||
|
|
||||||
| Layer | Main location | Responsibility |
|
| Layer | Main location | Responsibility |
|
||||||
|---|---|---|
|
|-------------|----------------------------|---------------------------------------------------------------------------------------------------|
|
||||||
| CLI | `reme/reme.py` | Parse commands; `start` launches the service; other actions call the service through a client. |
|
| CLI | `reme/reme.py` | Parse commands; `start` launches the service; other actions call the service through a client. |
|
||||||
| Service | `reme/components/service/` | Register Jobs as HTTP endpoints or MCP tools. |
|
| Service | `reme/components/service/` | Register Jobs as HTTP endpoints or MCP tools. |
|
||||||
| Application | `reme/application.py` | Assemble configured objects, start them in dependency order, close them, and invoke Jobs. |
|
| Application | `reme/application.py` | Assemble configured objects, start them in dependency order, close them, and invoke Jobs. |
|
||||||
|
|
@ -55,11 +55,13 @@ Core layers:
|
||||||
reme/
|
reme/
|
||||||
reme.py # CLI entry point
|
reme.py # CLI entry point
|
||||||
application.py # Application assembly and lifecycle
|
application.py # Application assembly and lifecycle
|
||||||
|
plugin.py # installed plugin contract and entry-point loader
|
||||||
config/
|
config/
|
||||||
default.yaml # default service / jobs / components
|
default.yaml # default service / jobs / components
|
||||||
|
cookbook.yaml # Auto Fin + Daily Paper + DingTalk composition
|
||||||
config_parser.py # config=, dot notation, and env placeholder parsing
|
config_parser.py # config=, dot notation, and env placeholder parsing
|
||||||
components/
|
components/
|
||||||
component_registry.py # global registry R
|
component_registry.py # backend registry and application-local copies
|
||||||
base_component.py # ComponentMixin / BaseComponent / bind dependency declarations
|
base_component.py # ComponentMixin / BaseComponent / bind dependency declarations
|
||||||
runtime_context.py # context for one Job execution
|
runtime_context.py # context for one Job execution
|
||||||
job/ # BaseJob / StreamJob / BackgroundJob / CronJob
|
job/ # BaseJob / StreamJob / BackgroundJob / CronJob
|
||||||
|
|
@ -68,18 +70,24 @@ reme/
|
||||||
file_store/ # file-index coordination layer
|
file_store/ # file-index coordination layer
|
||||||
file_graph/ # wikilink graph
|
file_graph/ # wikilink graph
|
||||||
keyword_index/ # BM25 and other keyword indexes
|
keyword_index/ # BM25 and other keyword indexes
|
||||||
file_chunker/ # Markdown / default text chunking
|
file_chunker/ # Markdown / JSON / JSONL / generic text chunking
|
||||||
file_catalog/ # change checkpoints
|
file_catalog/ # change checkpoints
|
||||||
as_llm/, as_embedding/ # model wrappers
|
as_llm/, as_embedding/ # model wrappers
|
||||||
agent_wrapper/ # AgentScope / Claude Code wrappers
|
agent_wrapper/ # AgentScope / Claude Code / Codex wrappers
|
||||||
steps/
|
steps/
|
||||||
base_step.py # BaseStep, Ref, dispatch_steps
|
base_step.py # BaseStep, Ref, dispatch_steps
|
||||||
common/ # version, help, health_check, demo
|
common/ # version, help, health_check, status, chat
|
||||||
file_io/ # read/write/edit/delete/move/frontmatter/daily
|
file_io/ # read/write/edit/delete/move/frontmatter/daily
|
||||||
index/ # watch/init/update/search/traverse
|
index/ # watch/init/update/search/traverse
|
||||||
evolve/ # auto_memory, auto_resource, auto_dream, proactive
|
evolve/ # auto_memory, auto_resource, auto_dream, proactive
|
||||||
transfer/ # upload/download/ingest
|
transfer/ # upload/download
|
||||||
channel/ # MCP channel tools
|
plugins/
|
||||||
|
dingtalk/ # independent DingTalk integration plugin distribution
|
||||||
|
auto-fin/ # independent example plugin distribution
|
||||||
|
daily_paper/ # independent paper-research plugin distribution
|
||||||
|
integrations/
|
||||||
|
claude_code/ # Claude Code adapter and marketplace
|
||||||
|
hermes_agent/ # Hermes Agent memory-provider adapter
|
||||||
```
|
```
|
||||||
|
|
||||||
The default workspace directories are defined by `ApplicationConfig`:
|
The default workspace directories are defined by `ApplicationConfig`:
|
||||||
|
|
@ -87,7 +95,8 @@ The default workspace directories are defined by `ApplicationConfig`:
|
||||||
```text
|
```text
|
||||||
<workspace_dir>/
|
<workspace_dir>/
|
||||||
metadata/ # persistent file_store, file_graph, keyword_index, file_catalog, and related state
|
metadata/ # persistent file_store, file_graph, keyword_index, file_catalog, and related state
|
||||||
session/ # agent sessions and original conversations
|
session/ # source conversations used by memory workflows
|
||||||
|
mem_session/ # generated Agent wrapper sessions and configuration
|
||||||
resource/ # external resources
|
resource/ # external resources
|
||||||
daily/ # lightly processed memory
|
daily/ # lightly processed memory
|
||||||
digest/ # long-term digest memory
|
digest/ # long-term digest memory
|
||||||
|
|
@ -128,7 +137,7 @@ reme search query="memory" backend=mcp
|
||||||
Configuration parsing supports:
|
Configuration parsing supports:
|
||||||
|
|
||||||
| Capability | Source | Description |
|
| Capability | Source | Description |
|
||||||
|---|---|---|
|
|------------------------|-------------------------|--------------------------------------------------------------------------|
|
||||||
| Default configuration | `resolve_app_config()` | Load `reme/config/default.yaml` when `config` is not specified. |
|
| Default configuration | `resolve_app_config()` | Load `reme/config/default.yaml` when `config` is not specified. |
|
||||||
| Explicit configuration | `config=<name-or-path>` | Accept a built-in configuration name or a YAML/JSON file path. |
|
| Explicit configuration | `config=<name-or-path>` | Accept a built-in configuration name or a YAML/JSON file path. |
|
||||||
| Dot notation | `parse_dot_notation()` | For example, `service.port=8181`. |
|
| Dot notation | `parse_dot_notation()` | For example, `service.port=8181`. |
|
||||||
|
|
@ -139,10 +148,14 @@ Configuration parsing supports:
|
||||||
|
|
||||||
`BaseService.run_app()` executes in this order:
|
`BaseService.run_app()` executes in this order:
|
||||||
|
|
||||||
|
Set the optional `service.jobs` list to restrict HTTP or MCP exposure to those job names. If omitted, all jobs with
|
||||||
|
`enable_serve: true` remain eligible; an empty list exposes none. The whitelist does not override `enable_serve: false`.
|
||||||
|
When the list is configured, a missing, disabled, unsupported, or invalid selected job fails service startup.
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
A["Service.build_service(app)"] --> B["read app.context.jobs"]
|
A["Service.build_service(app)"] --> B["read app.context.jobs"]
|
||||||
B --> C{"job.enable_serve == true?"}
|
B --> C{"enabled and selected by service.jobs?"}
|
||||||
C -->|yes| D["Service.add_job(job)"]
|
C -->|yes| D["Service.add_job(job)"]
|
||||||
C -->|no| E["skip registration"]
|
C -->|no| E["skip registration"]
|
||||||
D --> F["Service.start_service(app)"]
|
D --> F["Service.start_service(app)"]
|
||||||
|
|
@ -154,19 +167,28 @@ flowchart LR
|
||||||
HTTP service behavior:
|
HTTP service behavior:
|
||||||
|
|
||||||
| Job type | HTTP exposure |
|
| Job type | HTTP exposure |
|
||||||
|---|---|
|
|-------------------------------------------|---------------------------------------------------|
|
||||||
| Non-`StreamJob` with `enable_serve: true` | `POST /<job.name>` returning `Response` JSON. |
|
| Non-`StreamJob` with `enable_serve: true` | `POST /<job.name>` returning `Response` JSON. |
|
||||||
| `StreamJob` | `POST /<job.name>` returning `text/event-stream`. |
|
| `StreamJob` | `POST /<job.name>` returning `text/event-stream`. |
|
||||||
| `enable_serve: false` | No endpoint is registered. |
|
| `enable_serve: false` | No endpoint is registered. |
|
||||||
|
|
||||||
|
After registering Job endpoints, the HTTP service can also mount the ReMe Studio single-page application. The default is
|
||||||
|
`service.web_enabled=true`. Builds are resolved from `service.web_static_dir`, `REME_WEB_STATIC_DIR`, the optional
|
||||||
|
`reme_studio` package installed by the `web` and `core` extras, and source-tree locations such as
|
||||||
|
`reme_studio/dist-static`. If no `index.html` is found, only the frontend is skipped and the Job API remains available. The
|
||||||
|
Studio `GET` fallback does not replace existing `POST /<job.name>` routes.
|
||||||
|
|
||||||
MCP service behavior:
|
MCP service behavior:
|
||||||
|
|
||||||
| Job type | MCP exposure |
|
| Job type | MCP exposure |
|
||||||
|---|---|
|
|-------------------------------------------|-------------------------------------------------------------------|
|
||||||
| Non-`StreamJob` with `enable_serve: true` | Registered as an MCP tool. |
|
| Non-`StreamJob` with `enable_serve: true` | Registered as an MCP tool. |
|
||||||
| `StreamJob` | Currently skipped and not registered. |
|
| `StreamJob` | Currently skipped and not registered. |
|
||||||
| `BackgroundJob` | Forces `enable_serve=False` at construction and is never exposed. |
|
| `BackgroundJob` | Forces `enable_serve=False` at construction and is never exposed. |
|
||||||
|
|
||||||
|
MCP services can inject server-owned arguments with `injected_job_kwargs`; callers cannot override those arguments. Set
|
||||||
|
`tool_error_on_failure: true` to expose an unsuccessful ReMe `Response` as an MCP tool error.
|
||||||
|
|
||||||
## 4. Registry and Dependency Injection
|
## 4. Registry and Dependency Injection
|
||||||
|
|
||||||
### 4.1 Global Registry R
|
### 4.1 Global Registry R
|
||||||
|
|
@ -192,23 +214,55 @@ The registry key is:
|
||||||
`component_type` comes from a class attribute:
|
`component_type` comes from a class attribute:
|
||||||
|
|
||||||
| Type | Class attribute |
|
| Type | Class attribute |
|
||||||
|---|---|
|
|-----------|-----------------------------------------------------------|
|
||||||
| Step | `BaseStep.component_type = ComponentEnum.STEP` |
|
| Step | `BaseStep.component_type = ComponentEnum.STEP` |
|
||||||
| Job | `BaseJob.component_type = ComponentEnum.JOB` |
|
| Job | `BaseJob.component_type = ComponentEnum.JOB` |
|
||||||
| Service | `BaseService.component_type = ComponentEnum.SERVICE` |
|
| Service | `BaseService.component_type = ComponentEnum.SERVICE` |
|
||||||
| FileStore | `BaseFileStore.component_type = ComponentEnum.FILE_STORE` |
|
| FileStore | `BaseFileStore.component_type = ComponentEnum.FILE_STORE` |
|
||||||
|
|
||||||
The same backend name can therefore exist under different component types. For example, `http` can be both a service backend
|
The same backend name can therefore exist under different component types. For example, `http` can be both a service
|
||||||
and a client backend.
|
backend and a client backend.
|
||||||
|
|
||||||
### 4.2 Registration Through Module Imports
|
`ComponentEnum` provides the built-in identifiers, but installed plugins may declare a new type with a namespaced
|
||||||
|
string such as `example.reranker`. Custom identifiers use lowercase letters and numbers separated by `.`, `_`, or `-`.
|
||||||
|
They are configured under `components` and participate in the same dependency ordering and lifecycle as built-ins.
|
||||||
|
|
||||||
Registration happens when a module is imported. `reme/components/__init__.py` imports component packages, while
|
### 4.2 Built-in and Plugin Registration
|
||||||
`reme/steps/__init__.py` imports `channel/common/evolve/file_io/index/transfer`. Each package's `__init__.py` then imports
|
|
||||||
its concrete modules, causing `@R.register(...)` to execute.
|
|
||||||
|
|
||||||
After adding a Step file, make sure the package's `__init__.py` imports it. Otherwise, the backend will not appear in the
|
Built-in implementations populate the built-in registry through package imports. ReMe freezes that template after
|
||||||
registry.
|
bootstrap, and each `Application` receives a mutable copy. Runtime code resolves backends through the application's
|
||||||
|
registry rather than changing the process-wide template. ReMe then loads only the installed plugins explicitly named by
|
||||||
|
`plugins` in the resolved configuration. A plugin exposes its package through the `reme.plugins` Python entry-point
|
||||||
|
group. The package's `plugin.yaml` has two optional mappings: `backends` maps registration names to
|
||||||
|
`module:Class` targets, and `application_defaults` contributes a low-priority `ApplicationConfig` fragment. The
|
||||||
|
entry-point name is the plugin's identity.
|
||||||
|
Plugins are enabled explicitly through the application config's `plugins` list or a `plugins=[...]` CLI override.
|
||||||
|
Plugin registration therefore stays local to one application;
|
||||||
|
duplicate `(component_type, backend)` providers fail during assembly instead of overwriting each other.
|
||||||
|
|
||||||
|
The legacy Python `Plugin` descriptor and `reme.configs` entry points remain accepted during migration. Configuration
|
||||||
|
files can use `extends` to inherit another built-in, legacy plugin, or file-based configuration. See the independently
|
||||||
|
packaged [DingTalk](../../plugins/dingtalk/README.md), [Auto Fin](../../plugins/auto-fin/README.md), and
|
||||||
|
[Daily Paper](../../plugins/daily_paper/README.md) plugins.
|
||||||
|
|
||||||
|
Plugin packages are managed locally and remain separate from per-application activation:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
reme plugins list
|
||||||
|
reme plugins install plugins/dingtalk
|
||||||
|
reme plugins install plugins/auto-fin
|
||||||
|
reme plugins install plugins/daily_paper
|
||||||
|
reme plugins show daily-paper
|
||||||
|
reme plugins validate daily-paper
|
||||||
|
reme plugins uninstall daily-paper
|
||||||
|
|
||||||
|
reme start config=cookbook
|
||||||
|
```
|
||||||
|
|
||||||
|
These management commands use the current Python interpreter's pip and never run through an HTTP or MCP service.
|
||||||
|
The built-in `cookbook` configuration composes the three plugins, adds DingTalk delivery to the two report pipelines,
|
||||||
|
and starts the DingTalk Agent bridge as a background Job. Enabling Auto Fin or Daily Paper alone keeps it independent
|
||||||
|
from DingTalk.
|
||||||
|
|
||||||
### 4.3 Component.bind
|
### 4.3 Component.bind
|
||||||
|
|
||||||
|
|
@ -228,7 +282,7 @@ flowchart LR
|
||||||
Rules for `BaseComponent.bind(name, BaseClass, optional=True)`:
|
Rules for `BaseComponent.bind(name, BaseClass, optional=True)`:
|
||||||
|
|
||||||
| Scenario | Behavior |
|
| Scenario | Behavior |
|
||||||
|---|---|
|
|-----------------------------------------|------------------------------------------------------------|
|
||||||
| `name` is empty | Return `None` and skip the dependency. |
|
| `name` is empty | Return `None` and skip the dependency. |
|
||||||
| `app_context` exists | Look up `app_context.components[ctype][name]`. |
|
| `app_context` exists | Look up `app_context.components[ctype][name]`. |
|
||||||
| Dependency missing and `optional=True` | Resolve to `None`. |
|
| Dependency missing and `optional=True` | Resolve to `None`. |
|
||||||
|
|
@ -237,8 +291,8 @@ Rules for `BaseComponent.bind(name, BaseClass, optional=True)`:
|
||||||
|
|
||||||
### 4.4 Step.Ref
|
### 4.4 Step.Ref
|
||||||
|
|
||||||
Steps do not participate in component topological startup. They are created temporarily for each Job invocation. Steps access
|
Steps do not participate in component topological startup. They are created temporarily for each Job invocation. Steps
|
||||||
components primarily through `BaseStep.Ref`:
|
access components primarily through `BaseStep.Ref`:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
file_store: BaseFileStore = Ref(BaseFileStore, ComponentEnum.FILE_STORE)
|
file_store: BaseFileStore = Ref(BaseFileStore, ComponentEnum.FILE_STORE)
|
||||||
|
|
@ -294,11 +348,13 @@ flowchart LR
|
||||||
F --> G["start CronJob"]
|
F --> G["start CronJob"]
|
||||||
```
|
```
|
||||||
|
|
||||||
During shutdown, objects in `_started_components` are closed in reverse order so dependents close before their dependencies.
|
During shutdown, objects in `_started_components` are closed in reverse order so dependents close before their
|
||||||
|
dependencies.
|
||||||
|
|
||||||
## 6. Job Model
|
## 6. Job Model
|
||||||
|
|
||||||
A Job is the orchestration unit for an externally callable capability or background task. Jobs are configured under `jobs:`
|
A Job is the orchestration unit for an externally callable capability or background task. Jobs are configured under
|
||||||
|
`jobs:`
|
||||||
in `reme/config/default.yaml`.
|
in `reme/config/default.yaml`.
|
||||||
|
|
||||||
### 6.1 BaseJob
|
### 6.1 BaseJob
|
||||||
|
|
@ -320,7 +376,7 @@ flowchart LR
|
||||||
Important source behavior:
|
Important source behavior:
|
||||||
|
|
||||||
| Source | Behavior |
|
| Source | Behavior |
|
||||||
|---|---|
|
|--------------------|----------------------------------------------------------------------------------|
|
||||||
| `_start()` | Parse each Step config from YAML into `(step_cls, params)`. |
|
| `_start()` | Parse each Step config from YAML into `(step_cls, params)`. |
|
||||||
| `_build_steps()` | Create new Step instances for every call, avoiding state shared across requests. |
|
| `_build_steps()` | Create new Step instances for every call, avoiding state shared across requests. |
|
||||||
| `__call__()` | Create a `RuntimeContext` and execute Steps sequentially. |
|
| `__call__()` | Create a `RuntimeContext` and execute Steps sequentially. |
|
||||||
|
|
@ -331,7 +387,7 @@ Important source behavior:
|
||||||
`StreamJob` extends `BaseJob` but returns streaming chunks:
|
`StreamJob` extends `BaseJob` but returns streaming chunks:
|
||||||
|
|
||||||
| Behavior | Description |
|
| Behavior | Description |
|
||||||
|---|---|
|
|-------------|------------------------------------------------------------|
|
||||||
| Context | Includes `stream_queue`. |
|
| Context | Includes `stream_queue`. |
|
||||||
| Step output | Call `context.add_stream_string(text, ChunkEnum.CONTENT)`. |
|
| Step output | Call `context.add_stream_string(text, ChunkEnum.CONTENT)`. |
|
||||||
| Exception | Write `ChunkEnum.ERROR`. |
|
| Exception | Write `ChunkEnum.ERROR`. |
|
||||||
|
|
@ -355,8 +411,8 @@ flowchart LR
|
||||||
J --> K["wait close_timeout; cancel on timeout"]
|
J --> K["wait close_timeout; cancel on timeout"]
|
||||||
```
|
```
|
||||||
|
|
||||||
The default `BackgroundJob.__call__()` also executes configured Steps in sequence, but it does not swallow exceptions, which
|
The default `BackgroundJob.__call__()` also executes configured Steps in sequence, but it does not swallow exceptions,
|
||||||
allows the supervisor to restart the task.
|
which allows the supervisor to restart the task.
|
||||||
|
|
||||||
### 6.4 CronJob
|
### 6.4 CronJob
|
||||||
|
|
||||||
|
|
@ -370,7 +426,6 @@ jobs:
|
||||||
steps:
|
steps:
|
||||||
- backend: dream_extract_step
|
- backend: dream_extract_step
|
||||||
- backend: dream_integrate_step
|
- backend: dream_integrate_step
|
||||||
- backend: dream_topics_step
|
|
||||||
- backend: dream_finish_step
|
- backend: dream_finish_step
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -382,7 +437,9 @@ The current implementation uses `croniter` to calculate the next trigger time. T
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
Jobs["default.yaml jobs"] --> BG["background<br/>index_update_loop<br/>resource_watch_loop<br/>digest_watch_loop"]
|
Jobs["default.yaml jobs"] --> BG["background<br/>index_update_loop<br/>resource_watch_loop<br/>digest_watch_loop"]
|
||||||
Jobs --> Base["base<br/>version / help / health_check<br/>search / node_search / traverse / reindex<br/>read / write / edit / delete / move / list / stat<br/>daily_list / daily_reindex / daily_write<br/>auto_memory / auto_resource / auto_dream / proactive"]
|
Jobs --> Cron["cron<br/>dream_cron<br/>proactive_refresh_cron<br/>optimize_index_cron"]
|
||||||
|
Jobs --> Stream["stream<br/>chat"]
|
||||||
|
Jobs --> Base["base<br/>version / help / health_check / status / app_config<br/>search / node_search / traverse / graph_snapshot / reindex<br/>read / load / read_image / write / save / edit / delete / move / list / stat / frontmatter_*<br/>daily_list / daily_reindex / daily_write<br/>auto_memory / auto_memory_cc / auto_resource / auto_dream / proactive_refresh / proactive_read"]
|
||||||
```
|
```
|
||||||
|
|
||||||
## 7. Step Model
|
## 7. Step Model
|
||||||
|
|
@ -407,7 +464,7 @@ flowchart LR
|
||||||
`RuntimeContext` is shared by all Steps within one Job invocation:
|
`RuntimeContext` is shared by all Steps within one Job invocation:
|
||||||
|
|
||||||
| Field | Description |
|
| Field | Description |
|
||||||
|---|---|
|
|----------------|----------------------------------------------------------------------------|
|
||||||
| `response` | Final `Response(answer, success, metadata)`. |
|
| `response` | Final `Response(answer, success, metadata)`. |
|
||||||
| `data` | Free-form dictionary containing input parameters and intermediate results. |
|
| `data` | Free-form dictionary containing input parameters and intermediate results. |
|
||||||
| `stream_queue` | Output queue for streaming Jobs. |
|
| `stream_queue` | Output queue for streaming Jobs. |
|
||||||
|
|
@ -476,18 +533,19 @@ flowchart LR
|
||||||
Current default components in `reme/config/default.yaml`:
|
Current default components in `reme/config/default.yaml`:
|
||||||
|
|
||||||
| ComponentEnum | Name | Backend | Description |
|
| ComponentEnum | Name | Backend | Description |
|
||||||
|---|---|---|---|
|
|-------------------|---------------------------------|--------------------------------------------------|--------------------------------------------------------------------------------|
|
||||||
| `service` | singleton | `http` | Default HTTP service. |
|
| `service` | singleton | `http` | Default HTTP service. |
|
||||||
| `tokenizer` | `default` | `regex` | BM25 tokenizer. |
|
| `tokenizer` | `default` | `regex` | BM25 tokenizer. |
|
||||||
| `as_embedding` | `default` | `${EMBEDDING_BACKEND:-openai}` | Embedding model wrapper. |
|
| `as_embedding` | `default` | Not configured by default; example uses `openai` | Provides the embedding model wrapper after uncommenting the example config. |
|
||||||
| `embedding_store` | `default` | `local` | Embedding store depending on `as_embedding: default`. |
|
| `embedding_store` | `default` | Not configured by default; example uses `local` | Depends on `as_embedding: default` after uncommenting the example config. |
|
||||||
| `as_llm` | `default` | `${LLM_BACKEND:-openai}` | LLM model wrapper. |
|
| `as_llm` | `default` | `${LLM_BACKEND:-openai}` | LLM model wrapper. |
|
||||||
| `agent_wrapper` | `default` | `agentscope` | AgentScope wrapper. |
|
| `agent_wrapper` | `default` | `agentscope` | AgentScope wrapper. |
|
||||||
| `agent_wrapper` | `claude_code` | `claude_code` | Claude Code wrapper. |
|
| `agent_wrapper` | `claude_code` | `claude_code` | Claude Code wrapper. |
|
||||||
|
| `agent_wrapper` | `codex/codex_oauth` | `codex` | Codex wrappers for API-key and OAuth authentication. |
|
||||||
| `file_graph` | `default` | `local` | Wikilink graph. |
|
| `file_graph` | `default` | `local` | Wikilink graph. |
|
||||||
| `file_catalog` | `default/resource/digest/dream` | `local` | File-change checkpoints. |
|
| `file_catalog` | `default/resource/digest/dream` | `local` | File-change checkpoints. |
|
||||||
| `file_chunker` | `markdown` | `markdown` | Markdown AST chunking. |
|
| `file_chunker` | `markdown` | `markdown` | Markdown AST chunking. |
|
||||||
| `file_chunker` | `default` | `default` | Default text chunking, currently supporting `jsonl`. |
|
| `file_chunker` | `json/jsonl/default` | `json/jsonl/default` | JSON, JSONL, and generic text chunkers; generic text supports `txt` and `log`. |
|
||||||
| `keyword_index` | `default` | `bm25` | BM25 keyword index. |
|
| `keyword_index` | `default` | `bm25` | BM25 keyword index. |
|
||||||
| `file_store` | `default` | `local` | Combines file_graph and keyword_index; defaults to `embedding_store: ""`. |
|
| `file_store` | `default` | `local` | Combines file_graph and keyword_index; defaults to `embedding_store: ""`. |
|
||||||
|
|
||||||
|
|
@ -548,7 +606,7 @@ class MySearchStep(BaseStep):
|
||||||
Common attributes available directly:
|
Common attributes available directly:
|
||||||
|
|
||||||
| Attribute | Component resolved by default |
|
| Attribute | Component resolved by default |
|
||||||
|---|---|
|
|----------------------|-------------------------------------|
|
||||||
| `self.as_llm` | `.model` from `as_llm: default`. |
|
| `self.as_llm` | `.model` from `as_llm: default`. |
|
||||||
| `self.agent_wrapper` | `agent_wrapper: default`; optional. |
|
| `self.agent_wrapper` | `agent_wrapper: default`; optional. |
|
||||||
| `self.file_catalog` | `file_catalog: default`; optional. |
|
| `self.file_catalog` | `file_catalog: default`; optional. |
|
||||||
|
|
@ -565,7 +623,7 @@ steps:
|
||||||
### 9.4 Step Design Guidance
|
### 9.4 Step Design Guidance
|
||||||
|
|
||||||
| Guidance | Reason |
|
| Guidance | Reason |
|
||||||
|---|---|
|
|---------------------------------------------------------------------------------|-------------------------------------------------------------------------------|
|
||||||
| Read input from `context` and write intermediate results to `context`. | A multi-Step Job passes data through the same context. |
|
| Read input from `context` and write intermediate results to `context`. | A multi-Step Job passes data through the same context. |
|
||||||
| Write the final result to `context.response`. | Services and clients consume the standard `Response`. |
|
| Write the final result to `context.response`. | Services and clients consume the standard `Response`. |
|
||||||
| Do not store request-scoped state on a Step instance. | A Step is rebuilt for every Job call, and stateless Steps are easier to test. |
|
| Do not store request-scoped state on a Step instance. | A Step is rebuilt for every Job call, and stateless Steps are easier to test. |
|
||||||
|
|
@ -593,8 +651,8 @@ async def test_uppercase_step():
|
||||||
|
|
||||||
## 10. Adding a Job
|
## 10. Adding a Job
|
||||||
|
|
||||||
A Job usually requires no new Python class; configure existing Steps instead. Add a new Job backend only when a new execution
|
A Job usually requires no new Python class; configure existing Steps instead. Add a new Job backend only when a new
|
||||||
model is required.
|
execution model is required.
|
||||||
|
|
||||||
### 10.1 Adding a Normal Request Job
|
### 10.1 Adding a Normal Request Job
|
||||||
|
|
||||||
|
|
@ -730,7 +788,7 @@ jobs:
|
||||||
Characteristics of a background Job:
|
Characteristics of a background Job:
|
||||||
|
|
||||||
| Characteristic | Description |
|
| Characteristic | Description |
|
||||||
|---|---|
|
|---------------------------------|--------------------------------------------------------------------|
|
||||||
| Not externally exposed | `BackgroundJob.__init__()` forces `enable_serve=False`. |
|
| Not externally exposed | `BackgroundJob.__init__()` forces `enable_serve=False`. |
|
||||||
| Has a supervisor | Restarts with exponential backoff after an exception by default. |
|
| Has a supervisor | Restarts with exponential backoff after an exception by default. |
|
||||||
| Has a stop event | Notifies the loop to exit during close. |
|
| Has a stop event | Notifies the loop to exit during close. |
|
||||||
|
|
@ -749,7 +807,6 @@ jobs:
|
||||||
- backend: dream_extract_step
|
- backend: dream_extract_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
- backend: dream_integrate_step
|
- backend: dream_integrate_step
|
||||||
- backend: dream_topics_step
|
|
||||||
- backend: dream_finish_step
|
- backend: dream_finish_step
|
||||||
file_catalog: dream
|
file_catalog: dream
|
||||||
```
|
```
|
||||||
|
|
@ -761,7 +818,7 @@ An invalid `cron` expression fails at startup.
|
||||||
Most use cases require only a new Step plus a YAML Job. Consider adding `reme/components/job/*.py` only in these cases:
|
Most use cases require only a new Step plus a YAML Job. Consider adding `reme/components/job/*.py` only in these cases:
|
||||||
|
|
||||||
| Requirement | New Job class? |
|
| Requirement | New Job class? |
|
||||||
|---|---|
|
|---------------------------------------------------------------------|--------------------------------|
|
||||||
| Add a business command | No; use `backend: base`. |
|
| Add a business command | No; use `backend: base`. |
|
||||||
| Chain existing steps | No; use `steps:`. |
|
| Chain existing steps | No; use `steps:`. |
|
||||||
| Need SSE/streaming output | No; use `backend: stream`. |
|
| Need SSE/streaming output | No; use `backend: stream`. |
|
||||||
|
|
|
||||||
9
docs/en/index.md
Normal file
9
docs/en/index.md
Normal file
|
|
@ -0,0 +1,9 @@
|
||||||
|
---
|
||||||
|
layout: home
|
||||||
|
markdownStyles: false
|
||||||
|
title: ReMe
|
||||||
|
titleTemplate: false
|
||||||
|
description: ReMe is a local-first, file-native memory system for agents.
|
||||||
|
---
|
||||||
|
|
||||||
|
<HomePage lang="en" />
|
||||||
100
docs/en/integrations.md
Normal file
100
docs/en/integrations.md
Normal file
|
|
@ -0,0 +1,100 @@
|
||||||
|
---
|
||||||
|
title: Agent Integrations
|
||||||
|
description: Connect ReMe to agents through the CLI, HTTP, MCP, Skills, and host adapters.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Agent Integrations
|
||||||
|
|
||||||
|
ReMe keeps memory in an independent service and a user-owned workspace. Multiple agents can call the same memory system without binding storage to one model or host.
|
||||||
|
|
||||||
|
## Choose an interface
|
||||||
|
|
||||||
|
| Scenario | Recommended interface |
|
||||||
|
|---|---|
|
||||||
|
| Local script or hook | ReMe CLI |
|
||||||
|
| Application backend | HTTP Client |
|
||||||
|
| Tool-protocol host | MCP |
|
||||||
|
| DeepSeek Harness | [`@agentscope-ai/reme-dsh-plugin`](./integrations/dsh.md) profile bundle |
|
||||||
|
| OpenClaw | [`@agentscope-ai/reme-openclaw-plugin`](./integrations/openclaw.md) |
|
||||||
|
| Claude Code | [Shared HTTP MCP + Skill + Stop Hook](./integrations/claude-code.md) |
|
||||||
|
| Hermes Agent | Memory provider adapter |
|
||||||
|
| Codex or another coding agent | `reme_memory` Skill or MCP |
|
||||||
|
|
||||||
|
## General memory loop
|
||||||
|
|
||||||
|
1. Before answering, call `search` for relevant memory.
|
||||||
|
2. Use `read` on high-value results and `traverse` when relationships matter.
|
||||||
|
3. Retain workspace-relative source paths in the answer.
|
||||||
|
4. At session end, pass source messages to `auto_memory`.
|
||||||
|
5. Let background or scheduled workflows consolidate daily notes into digest memory.
|
||||||
|
|
||||||
|
An empty search result must remain empty; do not present model inference as recalled history.
|
||||||
|
|
||||||
|
## MCP
|
||||||
|
|
||||||
|
The default HTTP service exposes streamable HTTP MCP at `http://127.0.0.1:2333/mcp`. Common tools include `search`,
|
||||||
|
`read`, `traverse`, `list`, `auto_memory`, and `proactive_read`.
|
||||||
|
|
||||||
|
Use `service.jobs` to expose a read-only subset or keep write tools in a separate configuration.
|
||||||
|
|
||||||
|
## CLI and Skill
|
||||||
|
|
||||||
|
`skills/reme_memory/SKILL.md` defines a general workflow for agents that can run local commands: installation checks, service discovery, retrieval, reading, and persistence boundaries.
|
||||||
|
|
||||||
|
It deliberately avoids silently modifying Python environments, stopping unknown processes on port conflicts, writing recalled tool output back as conversation source, or persisting credentials.
|
||||||
|
|
||||||
|
## DeepSeek Harness
|
||||||
|
|
||||||
|
Install the self-contained [DeepSeek Harness plugin](./integrations/dsh.md):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
dsh plugin --profile web add @agentscope-ai/reme-dsh-plugin
|
||||||
|
```
|
||||||
|
|
||||||
|
Release links: [Awesome DSH Plugin](https://awesome-dsh-plugin.com/p/agentscope-ai/ReMe--integrations-dsh/) and
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-dsh-plugin).
|
||||||
|
|
||||||
|
It injects long-term-memory usage guidance into new root-agent sessions and exposes the read-only `reme_search` tool;
|
||||||
|
it does not preload the full memory history into the prompt. Completed user/assistant turns can be submitted to
|
||||||
|
`auto_memory` in background batches, while a timezone-aware schedule runs `auto_dream` to consolidate daily notes.
|
||||||
|
|
||||||
|
DSH settings configure the endpoint, guidance language, search limits, capture interval, root-agent filtering, and
|
||||||
|
consolidation schedule. The ReMe Status page exposes Overview, Auto Memory, Memory Consolidation, Components, Journal,
|
||||||
|
and Personal Knowledge Base views. Runtime counters are diagnostic state; workspace Markdown remains the durable source
|
||||||
|
of truth.
|
||||||
|
|
||||||
|
## OpenClaw
|
||||||
|
|
||||||
|
Install the independently published [OpenClaw plugin](./integrations/openclaw.md):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
openclaw plugins install clawhub:@agentscope-ai/reme-openclaw-plugin
|
||||||
|
```
|
||||||
|
|
||||||
|
Release links: [ClawHub](https://clawhub.ai/agentscope-ai/plugins/reme-openclaw-plugin) and
|
||||||
|
[npm](https://www.npmjs.com/package/@agentscope-ai/reme-openclaw-plugin). The plugin provides its own host-specific
|
||||||
|
ReMe HTTP boundary and release lifecycle.
|
||||||
|
|
||||||
|
## Claude Code
|
||||||
|
|
||||||
|
The [Claude Code plugin](./integrations/claude-code.md) connects every Claude Code window to one ReMe HTTP process at
|
||||||
|
`http://127.0.0.1:2333/mcp` by default. The `reme-memory` Skill selects among semantic `search`, topological `traverse`,
|
||||||
|
and state-oriented `daily_list` / `frontmatter_read`, then reads and cites the relevant workspace paths.
|
||||||
|
|
||||||
|
On Stop, the hook passes only the Claude Code `session_id` to the server-side `auto_memory_cc` job. On POSIX systems it
|
||||||
|
detaches the potentially long model call so Claude Code can stop immediately; unreachable-service and other best-effort
|
||||||
|
failures are written to the plugin log instead of blocking the host. ReMe resolves the local transcript, and repeated
|
||||||
|
Stop events with no new messages do not create duplicate memory.
|
||||||
|
|
||||||
|
## Hermes Agent
|
||||||
|
|
||||||
|
`integrations/hermes_agent/` provides a memory provider with HTTP and embedded modes. It recalls context before model calls and asynchronously invokes `auto_memory` after each turn. Its `config_schema.py` is rendered by Hermes' generic memory settings UI.
|
||||||
|
|
||||||
|
## Production guidance
|
||||||
|
|
||||||
|
- choose a stable absolute `workspace_dir`;
|
||||||
|
- reuse a service discovered by `reme find_reme`;
|
||||||
|
- treat `reme help` as the active Job contract;
|
||||||
|
- apply timeouts and failure logging to writes;
|
||||||
|
- do not block the host's core response path when memory is temporarily unavailable;
|
||||||
|
- use authentication, TLS, and a minimal Job allowlist for remote access.
|
||||||
|
|
@ -6,37 +6,39 @@ ReMe's core idea is **Memory as File, File as Memory**.
|
||||||
<img src="../figure/memory-as-file.svg" alt="ReMe Memory as File model" width="92%">
|
<img src="../figure/memory-as-file.svg" alt="ReMe Memory as File model" width="92%">
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
**Memory as File**: long-term memory is not hidden in a black-box database. It lives in Markdown files, resource files, and
|
**Memory as File**: long-term memory is not hidden in a black-box database. Its source material and readable memories
|
||||||
index snapshots under the workspace. Users and agents can directly read, write, move, and delete those files.
|
live in user-owned files under the workspace. Users and agents can directly read, write, move, and delete those files;
|
||||||
|
indexes and snapshots under `metadata/` are derived state that can be rebuilt.
|
||||||
|
|
||||||
**File as Memory**: each file is more than ordinary text. It is an indexable, linkable, and evolvable memory node. ReMe parses
|
**File as Memory**: each file is more than ordinary text. It is an indexable, linkable, and evolvable memory node. ReMe
|
||||||
frontmatter, body chunks, and wikilink edges from files and organizes them into retrieval indexes and a graph.
|
parses frontmatter, body chunks, and wikilink edges from files and organizes them into retrieval indexes and a graph.
|
||||||
|
|
||||||
In other words, files are both a human-readable interface and an operational interface for agents. Directory structure
|
In other words, files are both a human-readable interface and an operational interface for agents. Directory structure
|
||||||
carries the memory layers, while Markdown syntax expresses content, metadata, and relationships.
|
carries the memory layers, while Markdown syntax expresses content, metadata, and relationships.
|
||||||
|
|
||||||
## Design Goals
|
## Design Goals
|
||||||
|
|
||||||
ReMe represents memory as files not merely for convenient storage, but to give long-term memory several essential properties:
|
ReMe represents memory as files not merely for convenient storage, but to give long-term memory several essential
|
||||||
|
properties:
|
||||||
|
|
||||||
| Goal | Meaning |
|
| Goal | Meaning |
|
||||||
|---|---|
|
|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||||
| Readable | Users can open the workspace directly and read daily notes, digest nodes, and source material like ordinary notes. |
|
| Readable | Users can open the workspace directly and read daily notes, digest nodes, and source material like ordinary notes. |
|
||||||
| Editable | Users and agents can correct, extend, move, or delete memory with file operations, without a specialized database client. |
|
| Editable | Users and agents can correct, extend, move, or delete memory with file operations, without a specialized database client. |
|
||||||
| Traceable | Long-term conclusions in digest can point back to daily, resource, or session sources through `derived_from:: [[...]]`. |
|
| Traceable | Long-term conclusions in digest can point back to daily, resource, or session files from a Sources section. |
|
||||||
| Portable | The workspace is an ordinary directory. Markdown, JSONL, YAML, and resource files can be backed up, synchronized, versioned, or moved to other tools. |
|
| Portable | The workspace is an ordinary directory. Markdown, JSONL, YAML, and resource files can be backed up, synchronized, versioned, or moved to other tools. |
|
||||||
| Indexable | Although the files are plain text, ReMe parses frontmatter, chunks, and wikilinks to build a retrieval index and file graph. |
|
| Indexable | Although the files are plain text, ReMe parses frontmatter, chunks, and wikilinks to build a retrieval index and file graph. |
|
||||||
| Collaborative | Humans judge and correct; agents organize, link, and retrieve. Both operate on the same files. |
|
| Collaborative | Humans judge and correct; agents organize, link, and retrieve. Both operate on the same files. |
|
||||||
|
|
||||||
ReMe memory is therefore neither a hidden database record nor a prompt fragment visible only to an LLM. It is first a file
|
ReMe memory is therefore neither a hidden database record nor a prompt fragment visible only to an LLM. It is first a
|
||||||
owned by the user and only then indexed by the system for retrieval.
|
file owned by the user and only then indexed by the system for retrieval.
|
||||||
|
|
||||||
## Memory Layers
|
## Memory Layers
|
||||||
|
|
||||||
A ReMe workspace divides memory into four layers:
|
A ReMe workspace divides memory into four layers:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
raw input -> session/ + resource/
|
source records -> session/ + resource/
|
||||||
working memory -> daily/
|
working memory -> daily/
|
||||||
long memory -> digest/
|
long memory -> digest/
|
||||||
system state -> metadata/
|
system state -> metadata/
|
||||||
|
|
@ -44,19 +46,23 @@ system state -> metadata/
|
||||||
|
|
||||||
Each layer solves a different problem.
|
Each layer solves a different problem.
|
||||||
|
|
||||||
`session/` and `resource/` preserve raw input. Their purpose is to retain the original situation: conversations, agent
|
`session/` and `resource/` preserve source records. Files under `resource/` remain unchanged at their original path.
|
||||||
sessions, uploaded material, web pages, and reports remain intact as evidence for later verification.
|
Standard Auto Memory records retain conversation messages while intentionally omitting tool-result and base64 data
|
||||||
|
blocks; this keeps recalled output and binary payloads from masquerading as user-provided evidence. Generated Agent
|
||||||
|
runtime state instead lives under `mem_session/`.
|
||||||
|
|
||||||
`daily/` is the lightly processed layer. It organizes the day's conversations and resources into more readable daily notes:
|
`daily/` is the lightly processed layer. It organizes the day's conversations and resources into more readable daily
|
||||||
what happened, which conclusions were reached, which follow-up tasks remain, and where the source material lives. Daily does
|
notes:
|
||||||
not aim for final abstraction; it is closer to a workbench for the day.
|
what happened, which conclusions were reached, which follow-up tasks remain, and where the source material lives. Daily
|
||||||
|
does not aim for final abstraction; it is closer to a workbench for the day.
|
||||||
|
|
||||||
`digest/` is the deeply processed layer. It stores memory nodes that can be reused over time, such as user preferences,
|
`digest/` is the deeply processed layer. It stores memory nodes that can be reused over time, such as user preferences,
|
||||||
project background, procedural experience, conceptual knowledge, and decision precedents. Digest should not merely copy
|
project background, procedural experience, conceptual knowledge, and decision precedents. Digest should not merely copy
|
||||||
daily. It should merge recurring facts, methods, and relationships into more stable descriptions.
|
daily. It should merge recurring facts, methods, and relationships into more stable descriptions.
|
||||||
|
|
||||||
`metadata/` is the system index layer. It stores runtime state such as the file catalog, chunk index, and graph snapshots.
|
`metadata/` is the system index layer. It stores runtime state such as the file catalog, chunk index, and graph
|
||||||
Users normally do not edit this content manually. The actual editing surface is `daily/`, `digest/`, and, when necessary,
|
snapshots. Users normally do not edit this content manually. The actual editing surface is `daily/`, `digest/`, and,
|
||||||
|
when necessary,
|
||||||
`resource/`.
|
`resource/`.
|
||||||
|
|
||||||
These layers let ReMe preserve both the original situation and its abstraction: daily reconstructs what happened, while
|
These layers let ReMe preserve both the original situation and its abstraction: daily reconstructs what happened, while
|
||||||
|
|
@ -73,22 +79,24 @@ The corresponding automatic flows are [Auto Memory](./auto_memory.md), [Auto Res
|
||||||
```text
|
```text
|
||||||
<workspace_dir>/
|
<workspace_dir>/
|
||||||
├── metadata/ # system index layer; persistent indexes, graph, catalogs; not a manual editing surface
|
├── metadata/ # system index layer; persistent indexes, graph, catalogs; not a manual editing surface
|
||||||
├── session/ # raw input layer; original conversations and agent sessions
|
├── session/ # source-record layer; source conversations
|
||||||
│ ├── dialog/
|
│ ├── dialog/
|
||||||
│ │ └── <session_id>.jsonl # conversation messages saved by auto_memory
|
│ │ └── <session_id>.jsonl # source messages saved by auto_memory
|
||||||
│ ├── agentscope/
|
|
||||||
│ │ └── <session_id>.jsonl
|
|
||||||
│ └── claude_code/
|
│ └── claude_code/
|
||||||
│ └── <session_id>.jsonl
|
│ └── <session_id>.jsonl # ReMe copy used by auto_memory_cc
|
||||||
├── resource/ # raw input layer; original external material
|
├── mem_session/ # generated Agent wrapper sessions/config, not user memory
|
||||||
|
│ ├── agentscope/
|
||||||
|
│ ├── claude_config/
|
||||||
|
│ └── codex/
|
||||||
|
├── resource/ # source-record layer; original external material
|
||||||
|
│ ├── <resource>.<ext> # root-level input uses today's date
|
||||||
│ └── YYYY-MM-DD/
|
│ └── YYYY-MM-DD/
|
||||||
│ └── <resource>.<ext>
|
│ └── <resource>.<ext> # dated input uses the directory date
|
||||||
├── daily/ # lightly processed layer; facts, conversation summaries, and resource interpretations by date
|
├── daily/ # lightly processed layer; facts, conversation summaries, and resource interpretations by date
|
||||||
│ ├── YYYY-MM-DD.md # index page for the day
|
│ ├── YYYY-MM-DD.md # index page for the day
|
||||||
│ └── YYYY-MM-DD/
|
│ └── YYYY-MM-DD/
|
||||||
│ ├── <session_id>.md # daily note distilled from a conversation
|
│ ├── <generated_name>.md # topic-named conversation or resource card
|
||||||
│ ├── <resource_stem>.md # daily note distilled from a resource
|
│ └── interests.yaml # proactive interest topics generated by proactive refresh
|
||||||
│ └── interests.yaml # proactive interest topics generated by auto_dream
|
|
||||||
└── digest/ # deeply processed layer; reusable personal facts, procedures, and knowledge nodes
|
└── digest/ # deeply processed layer; reusable personal facts, procedures, and knowledge nodes
|
||||||
├── personal/
|
├── personal/
|
||||||
│ └── <memory>.md # user profile, preferences, and durable personal facts
|
│ └── <memory>.md # user profile, preferences, and durable personal facts
|
||||||
|
|
@ -103,17 +111,21 @@ Typical flows:
|
||||||
```text
|
```text
|
||||||
conversation
|
conversation
|
||||||
-> session/dialog/<session_id>.jsonl
|
-> session/dialog/<session_id>.jsonl
|
||||||
-> daily/YYYY-MM-DD/<session_id>.md
|
-> daily/YYYY-MM-DD/<generated_name>.md
|
||||||
-> digest/personal | digest/procedure | digest/wiki
|
-> digest/personal | digest/procedure | digest/wiki
|
||||||
|
|
||||||
external resource
|
external resource
|
||||||
-> resource/YYYY-MM-DD/<resource>.<ext>
|
-> resource/[YYYY-MM-DD/]<resource>.<ext>
|
||||||
-> daily/YYYY-MM-DD/<resource_stem>.md
|
-> daily/YYYY-MM-DD/<generated_name>.md
|
||||||
-> digest/wiki | digest/procedure
|
-> digest/wiki | digest/procedure
|
||||||
```
|
```
|
||||||
|
|
||||||
The first two steps focus on recording and organizing; the final step focuses on long-term distillation. `auto_memory` and
|
The first two steps focus on recording and organizing; the final step focuses on long-term distillation. `auto_memory`
|
||||||
`auto_resource` generate daily notes from raw input, and `auto_dream` extracts and integrates digest nodes from daily.
|
and
|
||||||
|
`auto_resource` generate daily notes from source input, and `auto_dream` extracts and integrates digest nodes from
|
||||||
|
daily. The generated daily filename comes from validated frontmatter `name`; `session_id`, `source_conversation`, and
|
||||||
|
`source_resource`
|
||||||
|
provide stable provenance and lookup identity instead of determining the filename.
|
||||||
|
|
||||||
## Markdown Format
|
## Markdown Format
|
||||||
|
|
||||||
|
|
@ -131,9 +143,7 @@ tags: [new energy, solar]
|
||||||
# Conclusions
|
# Conclusions
|
||||||
|
|
||||||
The solar supply chain consists of [[digest/wiki/polysilicon.md]], wafers, cells, and modules.
|
The solar supply chain consists of [[digest/wiki/polysilicon.md]], wafers, cells, and modules.
|
||||||
|
One major producer is [[digest/wiki/longi.md|LONGi]].
|
||||||
upstream:: [[digest/wiki/polysilicon.md]]
|
|
||||||
[company:: [[digest/wiki/longi.md|LONGi]]]
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### Frontmatter
|
### Frontmatter
|
||||||
|
|
@ -148,8 +158,8 @@ source_conversation: [[session/dialog/abc.jsonl]]
|
||||||
---
|
---
|
||||||
```
|
```
|
||||||
|
|
||||||
The current code recognizes `name` and `description` explicitly. Other fields are preserved as additional metadata. The write
|
The current code recognizes `name` and `description` explicitly. Other fields are preserved as additional metadata. The
|
||||||
interface merges `name`, `description`, and `metadata` into frontmatter.
|
write interface merges `name`, `description`, and `metadata` into frontmatter.
|
||||||
|
|
||||||
Treat frontmatter as a node-level summary and the body as evidence, explanation, and relationships. For example:
|
Treat frontmatter as a node-level summary and the body as evidence, explanation, and relationships. For example:
|
||||||
|
|
||||||
|
|
@ -163,28 +173,31 @@ confidence: observed
|
||||||
|
|
||||||
The user repeatedly asks documentation to explain motivation, boundaries, and examples while avoiding marketing language.
|
The user repeatedly asks documentation to explain motivation, boundaries, and examples while avoiding marketing language.
|
||||||
|
|
||||||
derived_from:: [[daily/2026-06-20/session-a.md]]
|
Apply this preference when following [[digest/procedure/technical-documentation.md]].
|
||||||
related:: [[digest/procedure/technical-documentation.md]]
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
This preference was recorded in [[daily/2026-06-20/documentation-style.md]], which captures the user's repeated guidance.
|
||||||
```
|
```
|
||||||
|
|
||||||
This has three benefits:
|
This has three benefits:
|
||||||
|
|
||||||
1. `name` and `description` serve as lightweight summaries in lists, recall results, and agent decisions.
|
1. `name` and `description` serve as lightweight summaries in lists, recall results, and agent decisions.
|
||||||
2. The body can carry fuller facts, conditions, counterexamples, and sources.
|
2. The body can carry fuller facts, conditions, counterexamples, and sources.
|
||||||
3. Typed wikilinks such as `derived_from::` and `related::` can be parsed by the graph and maintained when files move.
|
3. Ordinary wikilinks can be parsed by the graph and maintained when files move.
|
||||||
|
|
||||||
Frontmatter is best for stable, short, structured fields; the body is best for explanations meant for people. Do not put long
|
Frontmatter is best for stable, short, structured fields; the body is best for explanations meant for people. Do not put
|
||||||
body text into YAML fields.
|
long body text into YAML fields.
|
||||||
|
|
||||||
### Wikilink
|
### Wikilink
|
||||||
|
|
||||||
Wikilinks express relationships between files with `[[...]]`:
|
Wikilinks express relationships between files with `[[...]]`:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
[[digest/wiki/solar.md]]
|
[[daily/2026-06-20/session.md]]
|
||||||
[[digest/wiki/solar.md#supply-chain]]
|
[[notes/example.md#L9]]
|
||||||
[[digest/wiki/solar.md|solar]]
|
[[notes/example.md#L9-L10]]
|
||||||
![[resource/2026-06-01/report.md]]
|
[[notes/example.md#L9-L10,L15-L20]]
|
||||||
```
|
```
|
||||||
|
|
||||||
ReMe wikilinks use **literal path semantics**:
|
ReMe wikilinks use **literal path semantics**:
|
||||||
|
|
@ -196,69 +209,74 @@ ReMe wikilinks use **literal path semantics**:
|
||||||
ReMe does not append `.md` automatically, search by filename, or automatically resolve folder notes. Use complete
|
ReMe does not append `.md` automatically, search by filename, or automatically resolve folder notes. Use complete
|
||||||
workspace-relative paths with their extensions.
|
workspace-relative paths with their extensions.
|
||||||
|
|
||||||
|
Ordinary Markdown links such as `[label](../wiki/example.md)` do not create `FileLink` edges and are not rewritten by
|
||||||
|
move or retarget operations.
|
||||||
|
|
||||||
|
Anchors such as `#L9`, `#L9-L10`, and `#L9-L10,L15-L20` remain ordinary `target_anchor` strings in the graph. The graph
|
||||||
|
parser does not validate line-anchor syntax, so values such as `#L0`, `#L10-L9`, and `#L9,` are also stored. The `read`
|
||||||
|
job does not interpret an anchor appended to `path`; use the separate 1-based, inclusive `start_line` and `end_line`
|
||||||
|
arguments to read a range, for example `read(path="digest/wiki/solar.md", start_line=9, end_line=10)`.
|
||||||
|
|
||||||
Wikilinks support these behaviors:
|
Wikilinks support these behaviors:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
body link -> create a FileLink
|
body link -> create a FileLink
|
||||||
predicate:: link -> create a FileLink with a relationship name
|
|
||||||
move a file -> rewrite [[old path]] in inbound edges by default
|
move a file -> rewrite [[old path]] in inbound edges by default
|
||||||
delete a file -> return remaining inbound edges so references can be cleaned up
|
delete a file -> return remaining inbound edges so references can be cleaned up
|
||||||
search match -> expand inbound and outbound links to provide context
|
search match -> expand inbound and outbound links to provide context
|
||||||
```
|
```
|
||||||
|
|
||||||
Supported relationship forms:
|
|
||||||
|
|
||||||
```markdown
|
|
||||||
industry:: [[digest/wiki/new-energy.md]]
|
|
||||||
[competitor:: [[digest/wiki/byd.md]]]
|
|
||||||
```
|
|
||||||
|
|
||||||
Parsed result:
|
Parsed result:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
FileLink
|
FileLink
|
||||||
source_path = current file
|
source_path = current file
|
||||||
target_path = digest/wiki/new-energy.md
|
target_path = notes/example.md
|
||||||
predicate = industry
|
target_anchor = L9-L10,L15-L20
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Older documents containing wrappers such as `related:: [[path]]`,
|
||||||
|
`- related:: [[path]]`, or `[related:: [[path]]]` remain readable. ReMe ignores the surrounding text and indexes the
|
||||||
|
inner `[[path]]` as an ordinary link. Graph changes are applied when source files pass through the normal ingestion
|
||||||
|
path. `reme reindex` only rebuilds BM25 and embedding indexes from existing chunks; it does not reparse files or rebuild
|
||||||
|
the derived graph.
|
||||||
|
|
||||||
### Sources and Relationships
|
### Sources and Relationships
|
||||||
|
|
||||||
The two most important link types in ReMe are source links and conceptual relationship links.
|
The two most important link types in ReMe are source links and conceptual relationship links.
|
||||||
|
|
||||||
A source link explains where a long-term memory came from:
|
A Sources section records where a long-term memory came from:
|
||||||
|
|
||||||
```markdown
|
```markdown
|
||||||
derived_from:: [[daily/2026-06-20/session-a.md]]
|
## Sources
|
||||||
derived_from:: [[resource/2026-06-20/report.pdf]]
|
|
||||||
|
The preference was observed in [[daily/2026-06-20/documentation-style.md]], and the supporting report evidence is retained in
|
||||||
|
[[resource/2026-06-20/report.pdf]].
|
||||||
```
|
```
|
||||||
|
|
||||||
A conceptual relationship link explains which other long-term memories relate to the node:
|
A conceptual relationship link explains which other long-term memories relate to the node. Weave it into natural prose:
|
||||||
|
|
||||||
```markdown
|
```markdown
|
||||||
related:: [[digest/wiki/solar-supply-chain.md]]
|
This analysis extends [[digest/wiki/solar-supply-chain.md]], follows
|
||||||
depends_on:: [[digest/procedure/research-report-analysis.md]]
|
[[digest/procedure/research-report-analysis.md]], and contrasts with
|
||||||
contrasts_with:: [[digest/wiki/central-inverter.md]]
|
[[digest/wiki/central-inverter.md]].
|
||||||
```
|
```
|
||||||
|
|
||||||
Ordinary body wikilinks also create graph edges, but when the relationship itself has semantic value, prefer
|
|
||||||
`predicate:: [[path]]`. This makes the meaning of links clearer to search, graph traversal, and later agent integration.
|
|
||||||
|
|
||||||
## Human and Agent Editing
|
## Human and Agent Editing
|
||||||
|
|
||||||
Because memory is stored as files, users can edit the workspace directly, while agents can read and write the same files
|
Because memory is stored as files, users can edit the workspace directly, while agents can read and write the same files
|
||||||
through ReMe's file tools. Both follow the same conventions:
|
through ReMe's file tools. Both follow the same conventions:
|
||||||
|
|
||||||
| Operation | Guidance |
|
| Operation | Guidance |
|
||||||
|---|---|
|
|---------------|---------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||||
| Add memory | Write to the appropriate directory, use frontmatter for Markdown, and prefer complete workspace-relative wikilinks. |
|
| Add memory | Write to the appropriate directory, use frontmatter for Markdown, and prefer complete workspace-relative wikilinks. |
|
||||||
| Edit a body | Preserve existing sources and important wikilinks. When correcting an old conclusion, explain how the new material changes the previous judgment. |
|
| Edit a body | Preserve existing sources and important wikilinks. When correcting an old conclusion, explain how the new material changes the previous judgment. |
|
||||||
| Move a file | ReMe's move tool rewrites old paths in inbound edges by default. After a manual move, inspect inbound links again. |
|
| Move a file | ReMe's move tool rewrites old paths in inbound edges by default. After a manual move, inspect inbound links again. |
|
||||||
| Delete a file | Check inbound links first. ReMe's delete tool returns source files that still point to the target, making dangling references easier to clean up. |
|
| Delete a file | Check inbound links first. ReMe's delete tool returns source files that still point to the target, making dangling references easier to clean up. |
|
||||||
| Edit metadata | Use frontmatter for short fields. When the body changes substantially, update `description` as well. |
|
| Edit metadata | Use frontmatter for short fields. When the body changes substantially, update `description` as well. |
|
||||||
|
|
||||||
A practical rule is: **an agent may rewrite the wording, but it must not lose evidence edges**. In particular,
|
A practical rule is: **an agent may rewrite the wording, but it must not lose evidence edges**. In particular, Sources
|
||||||
`derived_from:: [[...]]` and existing digest-to-digest wikilinks are the basis for traceable and extensible long-term memory.
|
entries and existing digest-to-digest wikilinks are the basis for traceable and extensible long-term memory.
|
||||||
|
|
||||||
## Path Semantics
|
## Path Semantics
|
||||||
|
|
||||||
|
|
@ -266,7 +284,7 @@ All file tools and wikilinks use workspace-relative paths as their basic unit:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
digest/wiki/solar.md
|
digest/wiki/solar.md
|
||||||
daily/2026-06-20/session-a.md
|
daily/2026-06-20/documentation-style.md
|
||||||
resource/2026-06-20/report.pdf
|
resource/2026-06-20/report.pdf
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
@ -278,16 +296,16 @@ Recommended practices:
|
||||||
1. Include `.md` when linking a Markdown file.
|
1. Include `.md` when linking a Markdown file.
|
||||||
2. Use the complete source path when linking from digest to daily or resource.
|
2. Use the complete source path when linking from digest to daily or resource.
|
||||||
3. Rename or move files through ReMe's move tool whenever possible to avoid stale paths.
|
3. Rename or move files through ReMe's move tool whenever possible to avoid stale paths.
|
||||||
4. Put external source material under `resource/YYYY-MM-DD/...` and long-term abstractions under `digest/...`. Do not put
|
4. Put external source material under `resource/YYYY-MM-DD/...` and long-term abstractions under `digest/...`. Do not
|
||||||
raw source material directly into digest.
|
put raw source material directly into digest.
|
||||||
|
|
||||||
Explicit path semantics sacrifice a little convenience when writing by hand, but provide predictability, portability, and
|
Explicit path semantics sacrifice a little convenience when writing by hand, but provide predictability, portability,
|
||||||
automatic maintainability.
|
and automatic maintainability.
|
||||||
|
|
||||||
## Memory Chunking
|
## Memory Chunking
|
||||||
|
|
||||||
Memory chunking divides a file into retrievable fragments. ReMe does not split Markdown at fixed lengths by default; it tries
|
Memory chunking divides a file into retrievable fragments. ReMe does not split Markdown at fixed lengths by default; it
|
||||||
to preserve semantic structure.
|
tries to preserve semantic structure.
|
||||||
|
|
||||||
This section explains how files become retrieval chunks. For index updates, BM25, vector recall, and link expansion, see
|
This section explains how files become retrieval chunks. For index updates, BM25, vector recall, and link expansion, see
|
||||||
[Memory Search](./memory_search.md).
|
[Memory Search](./memory_search.md).
|
||||||
|
|
@ -302,8 +320,8 @@ Document
|
||||||
chunk 1 | chunk 2 | chunk 3 | ...
|
chunk 1 | chunk 2 | chunk 3 | ...
|
||||||
```
|
```
|
||||||
|
|
||||||
This is simple, but it can cut headings, tables, code blocks, lists, and `[[wikilinks]]` in the middle. After a match, the
|
This is simple, but it can cut headings, tables, code blocks, lists, and `[[wikilinks]]` in the middle. After a match,
|
||||||
agent often sees only an isolated fragment without knowing its section or relationship to other memory nodes.
|
the agent often sees only an isolated fragment without knowing its section or relationship to other memory nodes.
|
||||||
|
|
||||||
ReMe chunking is closer to splitting memory by file structure:
|
ReMe chunking is closer to splitting memory by file structure:
|
||||||
|
|
||||||
|
|
@ -369,5 +387,10 @@ Matched body fragment
|
||||||
|
|
||||||
This lets the agent see not only an isolated paragraph but also its structural position in the source file.
|
This lets the agent see not only an isolated paragraph but also its structural position in the source file.
|
||||||
|
|
||||||
Non-Markdown files use `DefaultFileChunker` by default. It splits by byte size and preserves a small overlap. For Markdown,
|
Non-Markdown files use `DefaultFileChunker` by default. It splits by byte size and preserves a small overlap. For
|
||||||
the chunker also avoids cutting `[[wikilinks]]` in the middle.
|
Markdown, the chunker also avoids cutting `[[wikilinks]]` in the middle.
|
||||||
|
|
||||||
|
`DefaultFileChunker` and `MarkdownFileChunker` decode files with their configured `encoding` and normalize platform
|
||||||
|
newlines to LF before indexing. Their default `invalid_encoding_policy: replace` keeps decodable content searchable
|
||||||
|
when a source contains invalid bytes, without modifying the source file. Set `invalid_encoding_policy: strict` on a
|
||||||
|
chunker component to reject such files instead.
|
||||||
|
|
|
||||||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Add table
Reference in a new issue