diff --git a/.github/ISSUE_TEMPLATE/bug_report.yml b/.github/ISSUE_TEMPLATE/bug_report.yml
new file mode 100644
index 00000000..3e4f9d54
--- /dev/null
+++ b/.github/ISSUE_TEMPLATE/bug_report.yml
@@ -0,0 +1,97 @@
+name: Bug report
+description: Report reproducible incorrect or unexpected ReMe behavior
+title: "[Bug]: "
+labels: [bug]
+body:
+ - type: markdown
+ attributes:
+ value: |
+ Thanks for helping improve ReMe. Please remove secrets, API keys, and private memory content before submitting.
+
+ - type: textarea
+ id: description
+ attributes:
+ label: Description
+ description: What happened, and what did you expect instead?
+ placeholder: Describe the observed and expected behavior.
+ validations:
+ required: true
+
+ - type: textarea
+ id: reproduce
+ attributes:
+ label: Steps to reproduce
+ description: Provide the smallest configuration and command sequence that reproduces the problem.
+ placeholder: |
+ 1. Configure ...
+ 2. Run ...
+ 3. Observe ...
+ validations:
+ required: true
+
+ - type: textarea
+ id: config
+ attributes:
+ label: Relevant configuration
+ description: Include only relevant values and redact credentials, tokens, endpoints, and private paths.
+ render: yaml
+
+ - type: textarea
+ id: logs
+ attributes:
+ label: Logs or traceback
+ description: Paste relevant output after removing secrets and private workspace content.
+ render: shell
+
+ - type: input
+ id: reme-version
+ attributes:
+ label: ReMe version
+ placeholder: e.g. 0.4.1.8 or a commit SHA
+ validations:
+ required: true
+
+ - type: input
+ id: python-version
+ attributes:
+ label: Python version
+ placeholder: e.g. 3.11.9
+ validations:
+ required: true
+
+ - type: dropdown
+ id: os
+ attributes:
+ label: Operating system
+ options:
+ - Linux
+ - macOS
+ - Windows
+ - Other
+ validations:
+ required: true
+
+ - type: dropdown
+ id: area
+ attributes:
+ label: Affected area
+ options:
+ - CLI or configuration
+ - HTTP, MCP, or local service
+ - Memory or workspace files
+ - Search, catalog, graph, or index
+ - Model or agent integration
+ - ReMe Studio
+ - Plugin or external integration
+ - Packaging or installation
+ - Other
+ validations:
+ required: true
+
+ - type: checkboxes
+ id: safety
+ attributes:
+ label: Data safety
+ options:
+ - label: I removed credentials and private memory content from this report.
+ required: true
diff --git a/.github/ISSUE_TEMPLATE/config.yml b/.github/ISSUE_TEMPLATE/config.yml
new file mode 100644
index 00000000..16d229f1
--- /dev/null
+++ b/.github/ISSUE_TEMPLATE/config.yml
@@ -0,0 +1,8 @@
+blank_issues_enabled: false
+contact_links:
+ - name: ReMe documentation
+ url: https://reme.agentscope.io
+ about: Read the installation, configuration, and usage guides.
+ - name: Existing issues
+ url: https://github.com/agentscope-ai/ReMe/issues
+ about: Search for existing reports and discussions before opening a new issue.
diff --git a/.github/ISSUE_TEMPLATE/feature_request.yml b/.github/ISSUE_TEMPLATE/feature_request.yml
new file mode 100644
index 00000000..12e38e78
--- /dev/null
+++ b/.github/ISSUE_TEMPLATE/feature_request.yml
@@ -0,0 +1,64 @@
+name: Feature request
+description: Propose a focused enhancement to ReMe
+title: "[Feature]: "
+labels: [enhancement]
+body:
+ - type: textarea
+ id: problem
+ attributes:
+ label: Problem
+ description: What user problem or limitation should this change address?
+ validations:
+ required: true
+
+ - type: textarea
+ id: proposal
+ attributes:
+ label: Proposed behavior
+ description: Describe the desired behavior and its user-visible contract.
+ validations:
+ required: true
+
+ - type: dropdown
+ id: area
+ attributes:
+ label: Area
+ options:
+ - CLI or configuration
+ - Jobs or steps
+ - Memory or workspace files
+ - Search, catalog, graph, or index
+ - Service or client
+ - Model or agent integration
+ - ReMe Studio
+ - Plugin or external integration
+ - Documentation
+ - Other
+ validations:
+ required: true
+
+ - type: textarea
+ id: ownership
+ attributes:
+ label: Local-first and compatibility considerations
+ description: Explain any effect on user-owned files, rebuildable state, configuration, schemas, or service interfaces.
+
+ - type: textarea
+ id: alternatives
+ attributes:
+ label: Alternatives considered
+ description: Describe workarounds or alternative designs you considered.
+
+ - type: textarea
+ id: examples
+ attributes:
+ label: Example usage
+ description: Show the proposed CLI, configuration, API, or UI behavior when useful.
+ render: shell
+
+ - type: checkboxes
+ id: contribution
+ attributes:
+ label: Contribution
+ options:
+ - label: I am willing to help implement or test this feature.
diff --git a/.github/ISSUE_TEMPLATE/question.yml b/.github/ISSUE_TEMPLATE/question.yml
new file mode 100644
index 00000000..70d5d553
--- /dev/null
+++ b/.github/ISSUE_TEMPLATE/question.yml
@@ -0,0 +1,53 @@
+name: Usage question
+description: Ask for help using or configuring ReMe
+title: "[Question]: "
+labels: [question]
+body:
+ - type: markdown
+ attributes:
+ value: Please check the documentation and existing issues before asking a new question.
+
+ - type: textarea
+ id: goal
+ attributes:
+ label: What are you trying to achieve?
+ validations:
+ required: true
+
+ - type: textarea
+ id: attempted
+ attributes:
+ label: What have you tried?
+ description: Include relevant commands or configuration, with secrets and private memory content removed.
+ validations:
+ required: true
+
+ - type: input
+ id: reme-version
+ attributes:
+ label: ReMe version
+ placeholder: e.g. 0.4.1.8 or a commit SHA
+
+ - type: dropdown
+ id: area
+ attributes:
+ label: Area
+ options:
+ - Installation
+ - Configuration
+ - CLI or service usage
+ - Memory and workspace management
+ - Search and retrieval
+ - ReMe Studio
+ - Plugin or integration
+ - Other
+
+ - type: checkboxes
+ id: checked
+ attributes:
+ label: Before submitting
+ options:
+ - label: I checked the [ReMe documentation](https://reme.agentscope.io) and searched existing issues.
+ required: true
+ - label: I removed credentials and private memory content.
+ required: true
diff --git a/.github/PULL_REQUEST_TEMPLATE.md b/.github/PULL_REQUEST_TEMPLATE.md
new file mode 100644
index 00000000..aa3e4af2
--- /dev/null
+++ b/.github/PULL_REQUEST_TEMPLATE.md
@@ -0,0 +1,35 @@
+## Summary
+
+
+
+## Related issue
+
+
+
+## Contract and data impact
+
+- [ ] No public configuration, schema, CLI, endpoint, streaming, or workspace-layout contract changes
+- [ ] No user-owned memory files are deleted or rewritten
+- [ ] Derived indexes, catalogs, graphs, caches, and metadata remain rebuildable
+
+
+
+## Validation
+
+
+
+- [ ] Focused tests pass
+- [ ] Unit tests pass, or omitted tests are explained below
+- [ ] `pre-commit run --all-files` passes, or omitted checks are explained below
+- [ ] Frontend checks were run when `website/` changed
+
+## Checklist
+
+- [ ] I reviewed the diff for unrelated changes and sensitive data
+- [ ] Tests cover intentional behavior changes
+- [ ] Defaults, schemas, and concise documentation were updated together when required
+- [ ] Long-lived clients, tasks, services, and executors follow the application lifecycle
+
+## Screenshots or additional notes
+
+
diff --git a/.github/workflows/_build-docs.yml b/.github/workflows/_build-docs.yml
new file mode 100644
index 00000000..060abd09
--- /dev/null
+++ b/.github/workflows/_build-docs.yml
@@ -0,0 +1,56 @@
+name: _Build documentation
+
+on:
+ workflow_call:
+ inputs:
+ run_tests:
+ description: Run the documentation test suite before building
+ required: false
+ default: true
+ type: boolean
+ upload_pages_artifact:
+ description: Upload the build for a later GitHub Pages deployment job
+ required: false
+ default: false
+ type: boolean
+
+permissions:
+ contents: read
+
+jobs:
+ build:
+ name: Build documentation
+ runs-on: ubuntu-latest
+ defaults:
+ run:
+ working-directory: github-pages
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Set up Node
+ uses: actions/setup-node@v6
+ with:
+ node-version: '22.13'
+ cache: npm
+ cache-dependency-path: github-pages/package-lock.json
+
+ - name: Install dependencies
+ run: npm ci
+
+ - name: Run tests
+ if: inputs.run_tests
+ run: npm test
+
+ - name: Build documentation
+ run: npm run build
+
+ - name: Configure Pages
+ if: inputs.upload_pages_artifact
+ uses: actions/configure-pages@v6
+
+ - name: Upload Pages artifact
+ if: inputs.upload_pages_artifact
+ uses: actions/upload-pages-artifact@v4
+ with:
+ path: github-pages/dist
diff --git a/.github/workflows/_build-python-packages.yml b/.github/workflows/_build-python-packages.yml
new file mode 100644
index 00000000..21a91dad
--- /dev/null
+++ b/.github/workflows/_build-python-packages.yml
@@ -0,0 +1,104 @@
+name: _Build Python packages
+
+on:
+ workflow_call:
+ inputs:
+ expected_version:
+ description: Expected release version; omit for a consistency-only check
+ required: false
+ default: ''
+ type: string
+ upload_artifacts:
+ description: Upload distributions for later publish jobs
+ required: false
+ default: false
+ type: boolean
+
+permissions:
+ contents: read
+
+jobs:
+ distributions:
+ name: Build Python distributions
+ runs-on: ubuntu-latest
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Set up Node
+ uses: actions/setup-node@v6
+ with:
+ node-version: '22.13'
+ cache: npm
+ cache-dependency-path: website/package-lock.json
+
+ - name: Set up Python
+ uses: actions/setup-python@v6
+ with:
+ python-version: '3.11'
+
+ - name: Install build dependencies
+ run: |
+ python -m pip install --upgrade pip
+ python -m pip install build packaging pytest twine
+
+ - name: Validate package versions
+ if: inputs.expected_version == ''
+ run: python scripts/bump_version.py --check
+
+ - name: Validate release version
+ if: inputs.expected_version != ''
+ env:
+ EXPECTED_VERSION: ${{ inputs.expected_version }}
+ run: python scripts/bump_version.py --check --expected-version "${EXPECTED_VERSION}"
+
+ - name: Run package tests
+ run: PYTHONPATH=. python -m pytest tests/unit/test_package_versions.py -q
+
+ - name: Build Studio static workspace
+ working-directory: website
+ run: |
+ npm ci
+ npm run build:static
+
+ - name: Build and check distributions
+ run: |
+ python scripts/package_studio.py
+ mkdir -p dist/reme dist/studio
+ python -m build --outdir dist/reme
+ python -m build packages/reme_ai_studio --outdir dist/studio
+ python -m twine check dist/reme/* dist/studio/*
+
+ - name: Verify distributions and isolated installation
+ run: |
+ REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
+ STUDIO_WHEEL="$(pwd)/$(ls dist/studio/reme_ai_studio-*.whl)"
+ STUDIO_SDIST="$(pwd)/$(ls dist/studio/reme_ai_studio-*.tar.gz)"
+ python -m zipfile -l "${REME_WHEEL}" | (! grep 'reme/web/')
+ python -m zipfile -l "${STUDIO_WHEEL}" | grep 'reme_ai_studio/static/index.html'
+ python -m zipfile -l "${STUDIO_WHEEL}" | grep 'dist-info/licenses/LICENSE'
+ python -m tarfile -l "${STUDIO_SDIST}" | grep '/LICENSE'
+ python -m venv "${RUNNER_TEMP}/reme-package-smoke"
+ "${RUNNER_TEMP}/reme-package-smoke/bin/python" -m pip install \
+ --find-links "$(pwd)/dist/studio" "${REME_WHEEL}[core]"
+ cd "${RUNNER_TEMP}"
+ "${RUNNER_TEMP}/reme-package-smoke/bin/python" -c \
+ "import reme; from reme_ai_studio import static_dir; assert (static_dir() / 'index.html').is_file()"
+ "${RUNNER_TEMP}/reme-package-smoke/bin/python" -c \
+ "from reme.utils import resolve_web_static_dir; assert (resolve_web_static_dir() / 'index.html').is_file()"
+
+ - name: Upload ReMe Studio distributions
+ if: inputs.upload_artifacts
+ uses: actions/upload-artifact@v4
+ with:
+ name: reme-studio-distributions
+ path: dist/studio/
+ if-no-files-found: error
+
+ - name: Upload ReMe distributions
+ if: inputs.upload_artifacts
+ uses: actions/upload-artifact@v4
+ with:
+ name: reme-distributions
+ path: dist/reme/
+ if-no-files-found: error
diff --git a/.github/workflows/ci-docs.yml b/.github/workflows/ci-docs.yml
new file mode 100644
index 00000000..5650d860
--- /dev/null
+++ b/.github/workflows/ci-docs.yml
@@ -0,0 +1,48 @@
+name: CI / Documentation
+
+on:
+ push:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-docs.yml'
+ - '.github/workflows/_build-docs.yml'
+ - 'AGENTS.md'
+ - 'README.md'
+ - 'README_ZH.md'
+ - 'docs/**'
+ - 'github-pages/**'
+ - 'website/README*.md'
+ - 'website/public/og.jpg'
+ - 'plugins/*/README*.md'
+ - 'benchmark/*/README*.md'
+ - 'skills/reme_memory/SKILL.md'
+ pull_request:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-docs.yml'
+ - '.github/workflows/_build-docs.yml'
+ - 'AGENTS.md'
+ - 'README.md'
+ - 'README_ZH.md'
+ - 'docs/**'
+ - 'github-pages/**'
+ - 'website/README*.md'
+ - 'website/public/og.jpg'
+ - 'plugins/*/README*.md'
+ - 'benchmark/*/README*.md'
+ - 'skills/reme_memory/SKILL.md'
+ workflow_dispatch:
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+permissions:
+ contents: read
+
+jobs:
+ documentation:
+ name: Test and build documentation
+ uses: ./.github/workflows/_build-docs.yml
+ with:
+ run_tests: true
diff --git a/.github/workflows/ci-packages.yml b/.github/workflows/ci-packages.yml
new file mode 100644
index 00000000..ea34f701
--- /dev/null
+++ b/.github/workflows/ci-packages.yml
@@ -0,0 +1,46 @@
+name: CI / Python packages
+
+on:
+ push:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-packages.yml'
+ - '.github/workflows/_build-python-packages.yml'
+ - '.github/workflows/release-python.yml'
+ - 'packages/reme_ai_studio/**'
+ - 'pyproject.toml'
+ - 'reme/__init__.py'
+ - 'reme/utils/web_static.py'
+ - 'scripts/bump_version.py'
+ - 'scripts/package_studio.py'
+ - 'tests/unit/test_package_versions.py'
+ - 'website/**'
+ - 'LICENSE'
+ pull_request:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-packages.yml'
+ - '.github/workflows/_build-python-packages.yml'
+ - '.github/workflows/release-python.yml'
+ - 'packages/reme_ai_studio/**'
+ - 'pyproject.toml'
+ - 'reme/__init__.py'
+ - 'reme/utils/web_static.py'
+ - 'scripts/bump_version.py'
+ - 'scripts/package_studio.py'
+ - 'tests/unit/test_package_versions.py'
+ - 'website/**'
+ - 'LICENSE'
+ workflow_dispatch:
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+permissions:
+ contents: read
+
+jobs:
+ distributions:
+ name: Build and verify distributions
+ uses: ./.github/workflows/_build-python-packages.yml
diff --git a/.github/workflows/ci-python-quality.yml b/.github/workflows/ci-python-quality.yml
new file mode 100644
index 00000000..7b6a5437
--- /dev/null
+++ b/.github/workflows/ci-python-quality.yml
@@ -0,0 +1,37 @@
+name: CI / Python quality
+
+on:
+ push:
+ pull_request:
+ workflow_dispatch:
+
+permissions:
+ contents: read
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+jobs:
+ pre-commit:
+ name: Pre-commit
+ runs-on: ubuntu-latest
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Setup Python
+ uses: actions/setup-python@v6
+ with:
+ python-version: '3.11'
+ cache: pip
+
+ - name: Update setuptools
+ run: |
+ pip install -U setuptools wheel
+
+ - name: Install
+ run: |
+ pip install -q -e packages/reme_ai_studio -e ".[dev,core]"
+
+ - name: Pre-commit starts
+ run: pre-commit run --all-files
diff --git a/.github/workflows/unittest.yml b/.github/workflows/ci-python-tests.yml
similarity index 76%
rename from .github/workflows/unittest.yml
rename to .github/workflows/ci-python-tests.yml
index 0eb9e664..acae7f6a 100644
--- a/.github/workflows/unittest.yml
+++ b/.github/workflows/ci-python-tests.yml
@@ -1,4 +1,4 @@
-name: Tests ReMe
+name: CI / Python tests
on:
push:
@@ -11,6 +11,9 @@ concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
+permissions:
+ contents: read
+
jobs:
unit-tests:
name: Unit Tests - py${{ matrix.python-version }}
@@ -21,10 +24,10 @@ jobs:
python-version: ["3.11", "3.12", "3.13"]
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@v6
- name: Set up Python ${{ matrix.python-version }}
- uses: actions/setup-python@v5
+ uses: actions/setup-python@v6
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
@@ -32,12 +35,13 @@ jobs:
- name: Install dependencies
run: |
python -m pip install --upgrade pip setuptools wheel
- pip install -e ".[dev,core]"
+ pip install -e packages/reme_ai_studio -e ".[dev,core]"
+ pip install --no-deps -e plugins/auto-fin
pip install coverage
- name: Run unit tests
run: |
- coverage run -m pytest tests/unit \
+ coverage run -m pytest tests/unit plugins/auto-fin \
-v \
--tb=long \
-s \
diff --git a/.github/workflows/ci-typescript.yml b/.github/workflows/ci-typescript.yml
new file mode 100644
index 00000000..12f056c0
--- /dev/null
+++ b/.github/workflows/ci-typescript.yml
@@ -0,0 +1,47 @@
+name: CI / TypeScript integrations
+
+on:
+ push:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-typescript.yml'
+ - '.github/workflows/release-typescript.yml'
+ - 'packages/typescript/**'
+ pull_request:
+ branches: [main, master, dev, develop]
+ paths:
+ - '.github/workflows/ci-typescript.yml'
+ - '.github/workflows/release-typescript.yml'
+ - 'packages/typescript/**'
+ workflow_dispatch:
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+permissions:
+ contents: read
+
+jobs:
+ package:
+ name: Type-check, test, and pack
+ runs-on: ubuntu-latest
+ defaults:
+ run:
+ working-directory: packages/typescript
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - uses: actions/setup-node@v6
+ with:
+ node-version: '22.19'
+ cache: npm
+ cache-dependency-path: packages/typescript/package-lock.json
+
+ - run: npm ci
+ - run: npm run format:check
+ - run: npm run lint
+ - run: npm run typecheck
+ - run: npm test
+ - run: npm run test:package
diff --git a/.github/workflows/ci-website.yml b/.github/workflows/ci-website.yml
new file mode 100644
index 00000000..957adaf0
--- /dev/null
+++ b/.github/workflows/ci-website.yml
@@ -0,0 +1,50 @@
+name: CI / Website
+
+on:
+ push:
+ paths:
+ - "website/**"
+ - ".github/workflows/ci-website.yml"
+ pull_request:
+ paths:
+ - "website/**"
+ - ".github/workflows/ci-website.yml"
+
+ workflow_dispatch:
+
+permissions:
+ contents: read
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+jobs:
+ website:
+ name: Website checks
+ runs-on: ubuntu-latest
+ defaults:
+ run:
+ working-directory: website
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Setup Node
+ uses: actions/setup-node@v6
+ with:
+ node-version: "22"
+ cache: npm
+ cache-dependency-path: website/package-lock.json
+
+ - name: Install dependencies
+ run: npm ci
+
+ - name: Run format check
+ run: npm run format:check
+
+ - name: Run lint
+ run: npm run lint
+
+ - name: Run tests
+ run: npm test
diff --git a/.github/workflows/windows-smoke.yml b/.github/workflows/ci-windows.yml
similarity index 56%
rename from .github/workflows/windows-smoke.yml
rename to .github/workflows/ci-windows.yml
index 8a1f787a..6c7b6986 100644
--- a/.github/workflows/windows-smoke.yml
+++ b/.github/workflows/ci-windows.yml
@@ -1,4 +1,4 @@
-name: Windows Smoke
+name: CI / Windows
on:
push:
@@ -11,6 +11,9 @@ concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
+permissions:
+ contents: read
+
jobs:
cli-smoke:
name: CLI smoke - py${{ matrix.python-version }}
@@ -21,10 +24,17 @@ jobs:
python-version: ["3.11"]
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@v6
+
+ - name: Set up Node
+ uses: actions/setup-node@v6
+ with:
+ node-version: '22'
+ cache: npm
+ cache-dependency-path: website/package-lock.json
- name: Set up Python ${{ matrix.python-version }}
- uses: actions/setup-python@v5
+ uses: actions/setup-python@v6
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
@@ -32,10 +42,23 @@ jobs:
- name: Install package
run: |
python -m pip install --upgrade pip setuptools wheel
- pip install -e ".[dev,core]"
+ pip install -e packages/reme_ai_studio -e ".[dev,core]"
+
+ - name: Build Studio static workspace
+ working-directory: website
+ run: |
+ npm ci
+ npm run build:static
+
+ - name: Verify editable source installation serves Studio
+ shell: pwsh
+ run: |
+ Push-Location $env:RUNNER_TEMP
+ python -c "from reme.utils import resolve_web_static_dir; assert (resolve_web_static_dir() / 'index.html').is_file()"
+ Pop-Location
- name: Run version job
- run: reme start service.backend=cli job=version
+ run: reme start config=tests/fixtures/config/version-smoke.yaml job=version
- name: Run Windows path tests
run: |
diff --git a/.github/workflows/deploy-docs.yml b/.github/workflows/deploy-docs.yml
new file mode 100644
index 00000000..859030b5
--- /dev/null
+++ b/.github/workflows/deploy-docs.yml
@@ -0,0 +1,51 @@
+name: Deploy / Documentation
+
+on:
+ push:
+ branches: [main]
+ paths:
+ - "github-pages/**"
+ - "docs/**"
+ - "README.md"
+ - "README_ZH.md"
+ - "website/README*.md"
+ - "website/public/og.jpg"
+ - "plugins/*/README*.md"
+ - "benchmark/*/README*.md"
+ - "skills/reme_memory/SKILL.md"
+ - "AGENTS.md"
+ - ".github/workflows/deploy-docs.yml"
+ - ".github/workflows/_build-docs.yml"
+ workflow_dispatch:
+
+permissions:
+ contents: read
+ pages: write
+ id-token: write
+
+concurrency:
+ group: pages
+ cancel-in-progress: true
+
+jobs:
+ build:
+ name: Build documentation
+ uses: ./.github/workflows/_build-docs.yml
+ with:
+ run_tests: false
+ upload_pages_artifact: true
+ permissions:
+ contents: read
+ pages: write
+ id-token: write
+
+ deploy:
+ environment:
+ name: github-pages
+ url: ${{ steps.deployment.outputs.page_url }}
+ runs-on: ubuntu-latest
+ needs: build
+ steps:
+ - name: Deploy
+ id: deployment
+ uses: actions/deploy-pages@v5
diff --git a/.github/workflows/pr-title-check.yml b/.github/workflows/policy-pr-title.yml
similarity index 92%
rename from .github/workflows/pr-title-check.yml
rename to .github/workflows/policy-pr-title.yml
index e928aed2..a75c7f4a 100644
--- a/.github/workflows/pr-title-check.yml
+++ b/.github/workflows/policy-pr-title.yml
@@ -1,10 +1,14 @@
-name: PR Title Check
+name: Policy / PR title
on:
pull_request:
branches: [main, master, dev, develop]
types: [opened, edited, synchronize, reopened]
+permissions:
+ contents: read
+ pull-requests: read
+
jobs:
check-pr-title:
runs-on: ubuntu-latest
diff --git a/.github/workflows/pre-commit.yml b/.github/workflows/pre-commit.yml
deleted file mode 100644
index 7444fe65..00000000
--- a/.github/workflows/pre-commit.yml
+++ /dev/null
@@ -1,38 +0,0 @@
-name: Pre-commit
-
-on: [ push, pull_request ]
-
-jobs:
- run:
- runs-on: ${{ matrix.os }}
- strategy:
- fail-fast: True
- matrix:
- os: [ ubuntu-latest ]
- env:
- OS: ${{ matrix.os }}
- PYTHON: '3.11'
- steps:
- - uses: actions/checkout@v4
- - name: Setup Python
- uses: actions/setup-python@v5
- with:
- python-version: '3.11'
- - name: Update setuptools
- run: |
- pip install -U setuptools wheel
- - name: Install
- run: |
- pip install -q -e ".[dev,core]"
- - name: Install pre-commit
- run: |
- pre-commit install
- - name: Pre-commit starts
- run: |
- pre-commit run --all-files > pre-commit.log 2>&1 || true
- cat pre-commit.log
- if grep -q Failed pre-commit.log; then
- echo -e "\e[41m [**FAIL**] Please install pre-commit and format your code first. \e[0m"
- exit 1
- fi
- echo -e "\e[46m ********************************Passed******************************** \e[0m"
diff --git a/.github/workflows/python-publish.yml b/.github/workflows/python-publish.yml
deleted file mode 100644
index 88de942c..00000000
--- a/.github/workflows/python-publish.yml
+++ /dev/null
@@ -1,45 +0,0 @@
-# This workflow will upload a Python Package using Twine when a release is created
-# For more information see: https://docs.github.com/en/actions/automating-builds-and-tests/building-and-testing-python#publishing-to-package-registries
-
-# This workflow uses actions that are not certified by GitHub.
-# They are provided by a third-party and are governed by
-# separate terms of service, privacy policy, and support
-# documentation.
-
-name: Publish Python Package to Pypi
-
-on:
- workflow_dispatch:
- release:
- types: [published]
-
-permissions:
- contents: read
-
-jobs:
- deploy:
-
- runs-on: ubuntu-latest
-
- steps:
- - uses: actions/checkout@v6
- - name: Set up Python
- uses: actions/setup-python@v6
- with:
- python-version: '3.11'
- - name: Install dependencies
- run: |
- python -m pip install --upgrade pip
- pip install setuptools wheel build
- - name: Build package
- run: python -m build
- - name: Test installation
- run: |
- WHEEL="$(ls dist/*.whl)"
- pip install "${WHEEL}[core]"
- python -c "import reme; print(reme.__version__)"
- - name: Publish package to PyPI
- uses: pypa/gh-action-pypi-publish@release/v1
- with:
- user: __token__
- password: ${{ secrets.PYPI_API_TOKEN }}
diff --git a/.github/workflows/release-auto-fin.yml b/.github/workflows/release-auto-fin.yml
new file mode 100644
index 00000000..03b7e19a
--- /dev/null
+++ b/.github/workflows/release-auto-fin.yml
@@ -0,0 +1,138 @@
+# 发布操作手册:
+# 1. 先将 plugins/auto-fin/pyproject.toml 中的 project.version 更新为待发布版本并合入目标分支。
+# 2. 确认插件依赖的 reme-ai 版本已经发布到 PyPI;本工作流会在构建阶段验证该依赖可下载。
+# 3. 确认仓库 Actions Secret 已配置 PYPI_API_TOKEN,且 PyPI 上不存在相同版本。
+# 4. 在 GitHub 仓库的 Actions 页面选择“Release / Auto Fin plugin”,点击“Run workflow”。
+# 5. 输入与 project.version 完全一致的版本号(例如 0.1.0)后运行;版本也可以带 v 前缀。
+#
+# 推荐发布顺序:reme-ai -> reme-auto-fin -> QwenPaw 更新依赖并通过 plugins: [auto-fin] 启用。
+# 当前仅支持 workflow_dispatch 手动触发,不会因 push、tag 或 release 自动发布。
+
+name: Release / Auto Fin plugin
+
+run-name: Publish reme-auto-fin ${{ inputs.version }}
+
+on:
+ workflow_dispatch:
+ inputs:
+ version:
+ description: Version from plugins/auto-fin/pyproject.toml (for example, 0.1.0)
+ required: true
+ type: string
+
+permissions:
+ contents: read
+
+concurrency:
+ group: publish-reme-auto-fin
+ cancel-in-progress: false
+
+jobs:
+ build:
+ runs-on: ubuntu-latest
+ env:
+ RELEASE_VERSION: ${{ inputs.version }}
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Set up Python
+ uses: actions/setup-python@v6
+ with:
+ python-version: '3.11'
+
+ - name: Install test and build dependencies
+ run: |
+ python -m pip install --upgrade pip
+ python -m pip install build packaging pytest pytest-asyncio twine
+ python -m pip install -e ".[core]"
+ python -m pip install --no-deps -e plugins/auto-fin
+
+ - name: Validate package name and release version
+ id: package
+ run: |
+ python - "${RELEASE_VERSION}" <<'PY'
+ import os
+ import sys
+ import tomllib
+ from pathlib import Path
+
+ from packaging.requirements import Requirement
+ from packaging.version import Version
+
+ project = tomllib.loads(Path("plugins/auto-fin/pyproject.toml").read_text(encoding="utf-8"))["project"]
+ expected = Version(sys.argv[1].removeprefix("v"))
+ actual = Version(project["version"])
+ if project["name"] != "reme-auto-fin":
+ raise SystemExit(f"Expected project name 'reme-auto-fin', found {project['name']!r}")
+ if actual != expected:
+ raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
+ requirements = [requirement for requirement in project["dependencies"] if requirement.startswith("reme-ai")]
+ if len(requirements) != 1:
+ raise SystemExit(f"Expected one reme-ai dependency, found {requirements!r}")
+ reme_requirement = Requirement(requirements[0])
+ if reme_requirement.name != "reme-ai" or set(reme_requirement.extras) != {"core"}:
+ raise SystemExit(f"Expected a reme-ai[core] dependency, found {requirements[0]!r}")
+ with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
+ print(f"reme_requirement={reme_requirement}", file=output)
+ print(f"Publishing {project['name']} {actual}")
+ PY
+
+ - name: Run Auto Fin tests
+ run: python -m pytest plugins/auto-fin -q
+
+ - name: Require the plugin-enabled ReMe release on PyPI
+ run: |
+ python -m pip download --no-deps \
+ --dest "${RUNNER_TEMP}/reme-auto-fin-core" \
+ "${{ steps.package.outputs.reme_requirement }}"
+
+ - name: Build and check distributions
+ run: |
+ mkdir -p dist/auto-fin
+ python -m build plugins/auto-fin --outdir dist/auto-fin
+ python -m twine check dist/auto-fin/*
+
+ - name: Verify distributions and isolated installation
+ run: |
+ AUTO_FIN_WHEEL="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.whl)"
+ AUTO_FIN_SDIST="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.tar.gz)"
+ python -m zipfile -l "${AUTO_FIN_WHEEL}" | grep 'dist-info/licenses/LICENSE'
+ python -m tarfile -l "${AUTO_FIN_SDIST}" | grep '/LICENSE'
+ python -m venv "${RUNNER_TEMP}/reme-auto-fin-smoke"
+ "${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" -m pip install "${AUTO_FIN_WHEEL}"
+ cd "${RUNNER_TEMP}"
+ "${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" - <<'PY'
+ from importlib.metadata import distribution
+
+ package = distribution("reme-auto-fin")
+ plugins = {entry.name: entry for entry in package.entry_points if entry.group == "reme.plugins"}
+ configs = {entry.name: entry for entry in package.entry_points if entry.group == "reme.configs"}
+ assert plugins["auto-fin"].load().name == "auto-fin"
+ assert configs["auto-fin"].load().is_file()
+ PY
+
+ - name: Upload distributions
+ uses: actions/upload-artifact@v4
+ with:
+ name: reme-auto-fin-${{ inputs.version }}
+ path: dist/auto-fin/
+ if-no-files-found: error
+
+ publish:
+ needs: build
+ runs-on: ubuntu-latest
+
+ steps:
+ - name: Download distributions
+ uses: actions/download-artifact@v4
+ with:
+ name: reme-auto-fin-${{ inputs.version }}
+ path: dist/auto-fin
+
+ - name: Publish reme-auto-fin
+ uses: pypa/gh-action-pypi-publish@release/v1
+ with:
+ user: __token__
+ password: ${{ secrets.PYPI_API_TOKEN }}
+ packages-dir: dist/auto-fin
diff --git a/.github/workflows/release-python.yml b/.github/workflows/release-python.yml
new file mode 100644
index 00000000..6380b3e9
--- /dev/null
+++ b/.github/workflows/release-python.yml
@@ -0,0 +1,58 @@
+name: Release / Python packages
+
+on:
+ workflow_dispatch:
+ inputs:
+ version:
+ description: Release version
+ required: true
+ type: string
+ release:
+ types: [published]
+
+permissions:
+ contents: read
+
+jobs:
+ build:
+ name: Build and verify distributions
+ uses: ./.github/workflows/_build-python-packages.yml
+ with:
+ expected_version: ${{ github.event_name == 'release' && github.event.release.tag_name || inputs.version }}
+ upload_artifacts: true
+
+ publish-studio:
+ needs: build
+ runs-on: ubuntu-latest
+ steps:
+ - name: Download ReMe Studio distributions
+ uses: actions/download-artifact@v4
+ with:
+ name: reme-studio-distributions
+ path: dist/studio
+
+ - name: Publish ReMe Studio
+ uses: pypa/gh-action-pypi-publish@release/v1
+ with:
+ user: __token__
+ password: ${{ secrets.PYPI_API_TOKEN }}
+ packages-dir: dist/studio
+ skip-existing: true
+
+ publish-reme:
+ needs: publish-studio
+ runs-on: ubuntu-latest
+ steps:
+ - name: Download ReMe distributions
+ uses: actions/download-artifact@v4
+ with:
+ name: reme-distributions
+ path: dist/reme
+
+ - name: Publish ReMe
+ uses: pypa/gh-action-pypi-publish@release/v1
+ with:
+ user: __token__
+ password: ${{ secrets.PYPI_API_TOKEN }}
+ packages-dir: dist/reme
+ skip-existing: true
diff --git a/.github/workflows/release-typescript.yml b/.github/workflows/release-typescript.yml
new file mode 100644
index 00000000..75eba88a
--- /dev/null
+++ b/.github/workflows/release-typescript.yml
@@ -0,0 +1,132 @@
+# Release checklist:
+# 1. Update packages/typescript/package.json and package-lock.json to the release version and merge them.
+# 2. Configure npm Trusted Publishing for agentscope-ai/ReMe and this workflow file.
+# 3. Run this workflow manually with the exact package version (an optional v prefix is accepted).
+# 4. Use the `next` tag for prereleases and `latest` only for stable releases.
+
+name: Release / TypeScript integrations
+
+run-name: Publish @agentscope-ai/reme ${{ inputs.version }} (${{ inputs.npm_tag }})
+
+on:
+ workflow_dispatch:
+ inputs:
+ version:
+ description: Version from packages/typescript/package.json (for example, 0.1.0)
+ required: true
+ type: string
+ npm_tag:
+ description: npm distribution tag
+ required: true
+ default: latest
+ type: choice
+ options:
+ - next
+ - latest
+
+permissions:
+ contents: read
+
+concurrency:
+ group: publish-agentscope-ai-reme
+ cancel-in-progress: false
+
+jobs:
+ build:
+ runs-on: ubuntu-latest
+ env:
+ RELEASE_VERSION: ${{ inputs.version }}
+ NPM_TAG: ${{ inputs.npm_tag }}
+
+ steps:
+ - uses: actions/checkout@v6
+
+ - name: Set up Node
+ uses: actions/setup-node@v6
+ with:
+ node-version: '22.19'
+
+ - name: Validate package name and release version
+ working-directory: packages/typescript
+ run: |
+ node --input-type=module <<'JS'
+ import { readFileSync } from 'node:fs';
+
+ const manifest = JSON.parse(readFileSync('package.json', 'utf8'));
+ const expected = process.env.RELEASE_VERSION.replace(/^v/, '');
+ if (manifest.name !== '@agentscope-ai/reme') {
+ throw new Error(`Unexpected package name: ${manifest.name}`);
+ }
+ if (manifest.version !== expected) {
+ throw new Error(`package.json is ${manifest.version}, workflow input is ${expected}`);
+ }
+ const prerelease = manifest.version.includes('-');
+ const npmTag = process.env.NPM_TAG;
+ if (prerelease !== (npmTag === 'next')) {
+ throw new Error(prerelease
+ ? 'Prerelease versions must use the next npm tag'
+ : 'Stable versions must use the latest npm tag');
+ }
+ console.log(`Preparing ${manifest.name}@${manifest.version}`);
+ JS
+
+ - name: Install dependencies
+ working-directory: packages/typescript
+ run: npm ci
+
+ - name: Type-check and test
+ working-directory: packages/typescript
+ run: |
+ npm run format:check
+ npm run lint
+ npm run typecheck
+ npm test
+ npm run test:package
+
+ - name: Pack npm tarball
+ working-directory: packages/typescript
+ run: |
+ mkdir -p "${RUNNER_TEMP}/reme-typescript-package"
+ npm pack --pack-destination "${RUNNER_TEMP}/reme-typescript-package"
+
+ - name: Upload npm tarball
+ uses: actions/upload-artifact@v4
+ with:
+ name: agentscope-ai-reme-${{ inputs.version }}
+ path: ${{ runner.temp }}/reme-typescript-package/*.tgz
+ if-no-files-found: error
+
+ publish:
+ needs: build
+ runs-on: ubuntu-latest
+ permissions:
+ contents: read
+ id-token: write
+
+ steps:
+ - name: Set up Node for npm
+ uses: actions/setup-node@v6
+ with:
+ node-version: '24'
+ registry-url: https://registry.npmjs.org
+
+ - name: Download npm tarball
+ uses: actions/download-artifact@v4
+ with:
+ name: agentscope-ai-reme-${{ inputs.version }}
+ path: dist/typescript
+
+ - name: Reject an existing package version
+ env:
+ PACKAGE_VERSION: ${{ inputs.version }}
+ run: |
+ PACKAGE_VERSION="${PACKAGE_VERSION#v}"
+ if npm view "@agentscope-ai/reme@${PACKAGE_VERSION}" version >/dev/null 2>&1; then
+ echo "@agentscope-ai/reme@${PACKAGE_VERSION} already exists" >&2
+ exit 1
+ fi
+
+ - name: Publish to npm
+ env:
+ NPM_TAG: ${{ inputs.npm_tag }}
+ run: npm publish dist/typescript/*.tgz --access public --tag "${NPM_TAG}" --provenance
diff --git a/.github/workflows/security-codeql.yml b/.github/workflows/security-codeql.yml
new file mode 100644
index 00000000..2e558a46
--- /dev/null
+++ b/.github/workflows/security-codeql.yml
@@ -0,0 +1,44 @@
+name: Security / CodeQL
+
+on:
+ push:
+ branches: [main]
+ pull_request:
+ branches: [main]
+ schedule:
+ - cron: '0 1 * * 1'
+ workflow_dispatch:
+
+permissions:
+ actions: read
+ contents: read
+ packages: read
+ security-events: write
+
+concurrency:
+ group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
+ cancel-in-progress: true
+
+jobs:
+ analyze:
+ name: Analyze ${{ matrix.language }}
+ runs-on: ubuntu-latest
+ strategy:
+ fail-fast: false
+ matrix:
+ language: [python, javascript-typescript]
+
+ steps:
+ - name: Checkout repository
+ uses: actions/checkout@v6
+
+ - name: Initialize CodeQL
+ uses: github/codeql-action/init@v4
+ with:
+ languages: ${{ matrix.language }}
+ build-mode: none
+
+ - name: Perform CodeQL analysis
+ uses: github/codeql-action/analyze@v4
+ with:
+ category: /language:${{ matrix.language }}
diff --git a/.gitignore b/.gitignore
index d9c98bac..350817dc 100644
--- a/.gitignore
+++ b/.gitignore
@@ -30,8 +30,13 @@ htmlcov/
# Packaging / build outputs
build/
dist/
+node_modules/
*.egg-info/
+# Website build integration source (not generated output)
+!website/build/
+!website/build/**
+
# Logs / temporary files
*.log
nohup.out
@@ -46,6 +51,7 @@ temp*/
# ReMe runtime data
.reme/
reme_workspace/
+reme_workspace_auto_fin_real_test*/
vault/
*.db
*.sqlite
@@ -56,6 +62,10 @@ docs/_build/
site/
evaluation/
+# The pi-Bench suite ships its own trace-history render config, which must
+# stay in git even though it lives under an evaluation/ directory.
+!benchmark/pibench/config/bench/evaluation/
+!benchmark/pibench/config/bench/evaluation/**
datasets/
# Claude Code skills (local only)
diff --git a/AGENTS.md b/AGENTS.md
index f1bc7da7..8e0aa211 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -1,20 +1,19 @@
# AGENTS.md
-This file guides coding agents working in the ReMe repository. Keep changes small,
-testable, and consistent with the contracts already expressed by the code.
+This file guides coding agents working in the ReMe repository. Keep changes small, testable, and consistent with the
+contracts expressed by the current code.
## Project Principles
ReMe is a local-first, file-native memory system for agents.
-- User-owned memory files are the source of truth.
-- Indexes, caches, metadata, and generated state must be rebuildable.
-- Prefer transparent formats and behavior over hidden state.
-- Preserve user control over storage, configuration, and service boundaries.
+- User-owned workspace files are the durable source of truth.
+- Indexes, catalogs, graphs, caches, and generated metadata must remain rebuildable.
+- Prefer transparent formats and predictable behavior over hidden state.
+- Preserve user control over workspace paths, configuration, and service boundaries.
- Keep concepts focused on project intent; let code and schemas describe implementation.
-When a proposed convenience conflicts with these principles, favor data ownership,
-recoverability, and predictable behavior.
+When convenience conflicts with these principles, favor data ownership, recoverability, and explicit behavior.
## Sources of Truth
@@ -22,166 +21,190 @@ Use this order when documentation and implementation disagree:
1. Current code and public Pydantic schemas.
2. Tests that describe supported behavior.
-3. CLI help and the built-in configuration.
-4. Development documentation and historical notes.
+3. CLI behavior and the built-in configuration.
+4. README files and other development documentation.
-Do not copy large implementation descriptions into documentation. Link to the relevant
-module or express the stable contract instead. If behavior changes intentionally, update
-the code, schema, tests, configuration, and concise documentation together as needed.
+Do not duplicate large implementation descriptions in documentation. Express the stable contract and link to the
+relevant module where useful. When behavior changes intentionally, update the implementation, schemas, tests, defaults,
+and concise documentation together.
## Repository Map
-- `reme/reme.py`: CLI entry point and client/server dispatch.
-- `reme/application.py`: application assembly, dependency ordering, and lifecycle.
-- `reme/components/application_context.py`: application-wide wiring and shared in-memory metadata.
-- `reme/components/runtime_context.py`: scratch state shared by steps within one execution.
-- `reme/config/default.yaml`: built-in jobs, components, and defaults.
-- `reme/schema/`: public and runtime Pydantic contracts.
-- `reme/components/`: services, stores, clients, jobs, and component registration.
-- `reme/steps/`: executable job steps.
-- `tests/unit/`: primary fast validation suite.
-- `tests/integration/`: tests that may require real credentials or services.
-- `tests/vector/` and `tests/light/`: specialized suites.
-- `plugins/reme/`: Claude Code integration.
-- `skills/reme_memory/`: skill that communicates with the ReMe service.
-- `skills/qwenpaw_memory/`: separate direct-file memory convention; it does not call ReMe.
-- `docs/`: pages and assets that support the repository README; not the deployed docs site.
+- `reme/reme.py`: CLI entry point; dispatches `start`, `find_reme`, and client calls.
+- `reme/application.py`: application assembly, dependency ordering, job execution, and lifecycle.
+- `reme/config/config_parser.py`: YAML/JSON loading, environment expansion, dot-notation parsing, and deep config
+ merging.
+- `reme/config/default.yaml`: default service, jobs, steps, and components. Other files in
+ `reme/config/` are named configuration variants.
+- `reme/schema/application_config.py`: typed application, component, and job configuration.
+- `reme/schema/`: request, response, streaming, memory, graph, and file contracts.
+- `reme/components/application_context.py`: application-wide wiring and in-memory shared state.
+- `reme/components/runtime_context.py`: request-scoped data, response, streaming queue, and stop event.
+- `reme/components/base_component.py`: component lifecycle, dependency binding, and workspace helpers.
+- `reme/components/component_registry.py`: the frozen built-in registry template and application-local registry factory.
+- `reme/components/job/`: base, stream, background, and cron job implementations.
+- `reme/components/service/`: local CLI, HTTP, and MCP service backends.
+- `reme/components/`: agent wrappers, model adapters, stores, catalogs, graphs, indexes, clients, tokenizers, and
+ outbound proxies.
+- `reme/steps/`: registered job steps grouped by common, file I/O, index, evolve, cookbook, benchmark, and transfer
+ concerns.
+- `reme/utils/`: shared utilities, including service discovery, logging, web-static resolution, session I/O, token
+ accounting, and wikilink handling.
+- `tests/unit/`: primary fast, isolated validation suite.
+- `tests/integration/`: service/model tests that may need credentials or external processes.
+- `website/`: ReMe Workspace frontend source; its static build can be served by the HTTP service.
+- `plugins/`: installable ReMe extensions, such as Auto Fin.
+- `integrations/`: adapters that connect ReMe to external agent hosts, such as Claude Code, DSH, and Hermes Agent.
+- `skills/`: standalone skills; `reme_memory` calls ReMe, while other skills may use separate tools or direct-file
+ conventions.
+- `benchmark/` and `cookbook/`: runnable evaluations and example workflows.
+- `docs/`: README-linked supporting pages and figures.
## Development Setup
-ReMe requires Python 3.11 or newer.
+ReMe requires Python 3.11 or newer. Install the editable development environment with:
```bash
-pip install -e ".[dev,core]"
+pip install -e packages/reme_ai_studio -e ".[dev,core]"
```
-Before changing behavior, inspect the adjacent implementation, schemas, configuration,
-and focused tests. Follow existing patterns unless the task explicitly calls for a new
-contract or architecture.
+Before changing behavior, inspect the adjacent implementation, schema, built-in config, and focused tests. Follow
+existing async and typing patterns unless the task explicitly requires a new contract.
-## Change Workflow
+## Configuration and CLI Contracts
-1. Identify the narrowest supported contract affected by the request.
-2. Read the relevant implementation and tests before editing.
-3. Make the smallest coherent change; avoid unrelated cleanup.
-4. Update related schemas, defaults, registrations, and imports when required.
-5. Add or adjust focused tests for observable behavior.
-6. Run proportionate validation and report anything not run.
+- CLI syntax is `reme ACTION key=value ...`; leading `-` or `--` on arguments is accepted.
+- Nested overrides use dot notation. Values support null, booleans, numbers, JSON collections, and quoted JSON strings;
+ leading-zero numeric-looking values remain strings.
+- `config=` loads a discovered config name or a `.yaml`, `.yml`, or `.json` file. With no explicit config
+ path, `default` is loaded when available.
+- Config files expand `${VAR}` and `${VAR:-default}` recursively. An undefined variable without a default is an error.
+- CLI/config overrides are deep-merged over the loaded file. Do not silently change this merge behavior or stable
+ configuration keys.
+- `ApplicationConfig` normalizes `workspace_dir` to an expanded absolute path. `session_dir`
+ must remain workspace-relative; standard transcripts live under `{session_dir}/dialog`.
+- `reme start` runs the configured service. `reme start job= ...` switches to the one-shot CLI service and runs
+ the job through the normal application lifecycle.
+- Other actions use a client selected from the running service configuration when discoverable, otherwise from local
+ config. Client-selection arguments must not leak into the job payload.
-Component and step discovery depends on registration imports:
+## Registration and Application Lifecycle
-- Components use `R.register(...)` in `reme/components/component_registry.py`.
-- Component packages must be reachable through `reme/components/__init__.py`.
-- Step modules must be reachable through `reme/steps/__init__.py`.
+Component and Step discovery is import-driven:
-Adding an implementation without its registration import can leave it undiscoverable at
-runtime. Treat the implementation, registry entry, and import side effect as one change.
+- Implementations declare a non-`BASE` `component_type` and register with `@R.register("backend")`
+ or `R.register(Class, "backend")`.
+- Component packages must be imported through `reme/components/__init__.py`.
+- Step packages/modules must be reachable through their package `__init__.py` chain and ultimately
+ `reme/steps/__init__.py`.
+- Adding an implementation without its registration import leaves it undiscoverable at runtime. Treat implementation,
+ registration, import side effect, defaults, and tests as one change.
-Do not silently change stable CLI flags, configuration keys, workspace layouts, serialized
-schemas, or service interfaces. When such a change is required, preserve compatibility
-where practical and make the migration explicit.
+`Application` validates config through `ApplicationContext`, creates workspace directories, instantiates the service,
+configured components, and jobs, and then manages lifecycle as follows:
-## Step State Model
+- Components start in topological dependency order. Missing required dependencies and cycles fail explicitly; optional
+ dependencies may resolve to `None`.
+- Jobs start after components in this order: base jobs, stream jobs, background jobs, then cron jobs.
+- Shutdown closes everything in reverse start order and then shuts down the optional thread pool.
+- If startup fails, already-started resources are closed.
+- `BaseComponent.start()` and `close()` are lock-protected and idempotent. Dependencies created by a standalone
+ `default_factory` are owned and closed by the parent component.
-Treat every Step as stateless. `BaseJob` stores Step specifications and builds fresh Step
-instances for each Job invocation. A Step instance must not use `self` or class variables to
-retain mutable runtime state between calls.
+Keep async clients, tasks, executors, and services under this lifecycle. Do not introduce an untracked long-lived
+resource.
-Place state according to its lifetime:
+## Jobs, Steps, and State
-- Constructor fields on `self`: immutable Step configuration and resolved dependencies only.
-- `self.context` (`RuntimeContext`): request data and intermediate results for one Job
- execution; sequential Steps share this context.
-- `self.app_context.metadata`: in-memory state that must be shared across Step or Job
- invocations for the lifetime of the Application.
-- Workspace files or a dedicated Component/store: durable state that must survive an
- Application restart.
+`BaseJob` resolves configured Step classes during job startup and constructs fresh Step instances for every invocation.
+Job-level kwargs are merged into each `RuntimeContext`, with call-time kwargs taking precedence. Sequential Steps in one
+invocation share the same `RuntimeContext` and `Response`.
-Use narrow, namespaced keys in `app_context.metadata`, following existing patterns such as
-`tool_contexts`. The ApplicationContext is shared, so account for
-concurrent access when values are mutable. New Step code must not fall back to `self.kwargs`
-or another Step field to emulate shared state when `app_context` is absent; tests of shared
-state should construct an `ApplicationContext`. If shared state grows into a stable
-service-level contract or needs its own lifecycle, locking, or persistence, promote it to a
-typed ApplicationContext field or a dedicated Component instead of expanding an ad hoc
-metadata bucket.
+Treat Step instances as invocation-scoped:
-Do not use `Response.metadata` as a state store. It is request-scoped output for callers and
-diagnostics, distinct from `ApplicationContext.metadata`.
+- Constructor fields and `self.kwargs` hold Step configuration and resolved dependencies. They may be cached or adjusted
+ during that one invocation, but must not be relied on across Job calls.
+- `self.context.data` holds request inputs and intermediate values shared by sequential Steps.
+- `self.context.response.answer`, `success`, and `metadata` are request-scoped output. Because the same response travels
+ through the Step chain, later Steps may consume metadata produced earlier, but it is not application-lifetime or
+ durable storage.
+- `self.app_context.metadata` holds in-memory state shared across Job/Step invocations for the life of one
+ `Application`, such as counters, tool-context state, session maps, or locks.
+- Workspace files or a dedicated Component/store hold durable state that must survive restart.
+
+Use narrow, namespaced keys in `app_context.metadata` and protect shared mutable values against concurrent access. The
+search/draft helpers intentionally mirror tool-context state into
+`self.kwargs` only when no `ApplicationContext` exists for standalone use and unit tests; do not generalize that
+compatibility fallback into persistent runtime state. If shared state becomes a stable service contract or needs
+dedicated lifecycle, locking, or persistence, promote it to a typed context field or Component.
+
+Additional Step contracts:
+
+- `Ref` dependencies resolve in this order: Step kwargs, current `RuntimeContext`, then the named application component.
+ The value is cached only on the current Step instance and cleared before each call.
+- `input_mapping` and `output_mapping` copy keys within `RuntimeContext.data`; missing sources are ignored.
+- Dispatched Steps receive the current `RuntimeContext`, so their data and response are shared.
+- Base jobs convert uncaught Step errors into `Response(success=False)`; stream jobs emit an error chunk and always a
+ terminal `DONE`; background jobs let errors reach their supervisor.
+- Background jobs are never service-exposed. MCP also skips stream jobs. Respect `enable_serve`
+ and any configured service job allowlist.
+
+## Workspace and File Safety
+
+- Application startup creates the workspace plus configured metadata, session, memory-session, resource, daily, and
+ digest directories.
+- File-operation paths are resolved against the workspace and must stay inside it. Home-relative paths are unsupported,
+ traversal escapes are rejected, and `_allowed_paths` restrictions fail closed when invalid.
+- Preserve per-path locking, encoding detection, byte limits, truncation behavior, and optimistic
+ `expected_mtime` checks when modifying file operations.
+- Do not bypass the existing file steps or stores in a way that weakens workspace containment.
+- Never write test state into the repository's `.reme/`; use `tmp_path` or another isolated workspace.
+- Do not delete or rewrite user memory to repair an index or make a test pass. Rebuild derived state from source files
+ instead.
## Validation
Use the narrowest useful check while iterating, then broaden it according to risk.
-Run a focused test:
+Focused test:
```bash
pytest tests/unit/path/to/test_file.py -v
```
-Run the main unit suite:
+Main unit suite:
```bash
pytest tests/unit -v --tb=long -s --log-cli-level=WARNING
```
-Run repository formatting and lint checks when the change warrants it:
+Repository formatting and lint checks:
```bash
pre-commit run --all-files
```
-Formatting and lint configuration is authoritative. Python code currently uses a maximum
-line length of 120 for Black and Flake8, with Pylint also run by pre-commit.
+Black and Flake8 use a 120-character line limit and Python 3.11 formatting; Pylint is also run by pre-commit. If
+`website/` changes, use its Node 22.13+ scripts and run the proportionate checks from that directory, such as
+`npm run format:check`, `npm run lint`, or `npm test`.
-Integration tests may contact real services and require credentials such as
-`LLM_API_KEY` or `EMBEDDING_API_KEY`. Do not run credentialed or externally mutating tests
-automatically. Run them only when the task requires them and the user has supplied or
-authorized the necessary environment.
+Integration tests may contact real model providers, services, or agent subprocesses and can require credentials. Do not
+run credentialed or externally mutating tests automatically; run them only when the task requires them and the necessary
+environment has been supplied or authorized. Mock network, model, and subprocess boundaries in unit tests.
-## Coding and Test Conventions
-
-- Target Python 3.11+ and follow the surrounding typing and async style.
-- Steps are stateless. If a step needs to persist state, store it in
- `self.app_context.metadata` rather than on the step instance.
-- Keep public schemas explicit and backward-compatible where practical.
-- Close async clients, services, tasks, and other lifecycle resources deterministically.
-- Prefer clear failures over silently falling back to corrupt or ambiguous state.
-- Keep indexes and caches derivable from user-owned source files.
-- Use `tmp_path` or another isolated temporary workspace in tests.
-- Never write test state into the repository's `.reme/` directory.
-- Mock network or model boundaries in unit tests.
-- Do not commit `.env` files, credentials, runtime memory, logs, indexes, or caches.
-
-## Documentation Boundaries
-
-ReMe's local docs and the deployed documentation site have separate responsibilities.
-
-- Keep `docs/` focused on content and assets used by `README.md` and `README_ZH.md`.
-- Preserve README-linked pages under `docs/en/` and `docs/zh/`, including their relative
- paths, unless the README is updated in the same change.
-- Keep README-required images under `docs/figure/`.
-- Keep the README's main documentation index pointed at `docs.agentscope.io` or the
- `agentscope-ai/docs` repository, following the existing link style.
-- Do not treat local README-supporting pages as the source for the deployed website.
-
-The separate `agentscope-ai/docs` repository owns website content, navigation, versioning,
-and deployment. Public ReMe pages live there under `reme//`. Make website changes
-in that repository and follow its existing version-management conventions.
-
-Do not add website build configuration or deployment workflows to ReMe unless the task
-explicitly changes this repository boundary.
-
-## Agent Guardrails
+## Change Guardrails
- Preserve unrelated user changes in a dirty working tree.
-- Do not edit generated output when the source can be changed instead.
-- Do not delete or rewrite user data to make a test pass.
-- Avoid broad refactors unless they are necessary for the requested outcome.
-- Do not introduce dependencies without a concrete need and repository-level justification.
-- Treat network access, real credentials, and external service mutations as opt-in.
-- State which validations passed and which were not run in the final handoff.
+- Make the smallest coherent change and avoid unrelated cleanup or broad refactors.
+- Do not edit generated output when the source can be changed instead. The publish workflow builds
+ `website/dist-static` and copies it into `reme/web`; change `website/` source for frontend work.
+- Do not silently change CLI flags, configuration keys, workspace layouts, serialized schemas, endpoint shapes,
+ streaming termination, or service interfaces. Preserve compatibility where practical and document intentional
+ migrations.
+- Do not introduce dependencies without a concrete repository-level need.
+- Do not commit `.env` files, credentials, runtime memory, logs, indexes, caches, benchmark outputs, or generated
+ website distributions.
+- State which validations passed and which relevant checks were not run in the final handoff.
-If a requirement is ambiguous, first infer intent from nearby code, tests, and schemas. Ask
-the user only when the remaining choice would materially alter a public contract, user data,
-or external system.
+If a requirement is ambiguous, infer intent from nearby code, schemas, defaults, and tests. Ask the user only when the
+remaining choice would materially alter a public contract, user data, or an external system.
diff --git a/README.md b/README.md
index aa54b69d..ae7bcb98 100644
--- a/README.md
+++ b/README.md
@@ -8,7 +8,7 @@
-
+
@@ -20,26 +20,28 @@
- An agent memory layer that turns conversations and resources into readable, editable, searchable Markdown memory.
+ A local-first, self-evolving personal knowledge base for AI agents.
> Previous versions: [0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
-🧠 ReMe is a local-first memory layer for **AI agents**. It turns conversations and resources into file-based long-term
-memory, then continuously indexes, links, and consolidates that memory for future recall.
+🧠 ReMe turns conversations and resources into readable, editable, searchable, and interconnected Markdown memory. It
+works alongside agents such as QwenPaw, OpenClaw, Hermes, and Claude Code, continuously organizing what they learn while
+keeping the files under the user's control.
## ✨ Core Ideas
-- **Memory as File**: Markdown files with frontmatter and wikilinks serve as memory nodes that both users and agents can
- read and write directly.
+- **Memory as File, File as Memory**: Markdown files with frontmatter and wikilinks serve as memory nodes that both
+ users and agents can inspect, edit, move, and back up directly.
- **Self-evolving knowledge base**: Auto Memory, Auto Resource, and Auto Dream progressively transform conversations and
- resources into long-term memories, while automatically building wikilink relationships.
+ resources into daily notes and long-term knowledge, while Auto Link writes relationships and sources back into the
+ files.
- **Progressive hybrid search**: ReMe combines wikilinks, BM25, and embeddings for hybrid retrieval across keyword
- matching, semantic recall, and relationship expansion.
+ matching, optional semantic recall, and relationship expansion without loading every neighboring file into context.
- **Agent-friendly integration**: SKILL.md + CLI integration makes it easy for different agents to read, write,
- maintain, and reuse memory.
+ maintain, and reuse the same local workspace. HTTP, MCP, and Python integrations are also available.
@@ -51,7 +53,7 @@ memory, then continuously indexes, links, and consolidates that memory for futur
[QwenPaw](https://github.com/agentscope-ai/QwenPaw), [OpenClaw](https://github.com/openclaw/openclaw), and
[Hermes](https://github.com/nousresearch/hermes-agent) a user-editable long-term memory layer.
- **Coding agents**: Preserve coding style, project background, repository decisions, and workflow experience across
- sessions when integrating with coding agents such as [Claude Code](plugins/claude_code/reme).
+ sessions when integrating with coding agents such as [Claude Code](integrations/claude_code/reme).
- **LLM Wiki**: Turn conversations, notes, and resources into a searchable, traceable, and linked Markdown knowledge
base that both users and agents can maintain.
- **Self-evolving agents**: Support agents that learn from experience by saving successful paths, failed attempts,
@@ -59,11 +61,15 @@ memory, then continuously indexes, links, and consolidates that memory for futur
## 📰 News
-- [2026.08] - [Experience-driven enhancement method](benchmark/toolmemory/README.md) of agent tool-use
- execution built on ReMe is available on [arXiv:2608.03403](https://arxiv.org/abs/2608.03403).
-- [2026.07] - Introduced optional Cookbooks: [Daily Paper](cookbook/daily_paper/README.md) for paper discovery and
- analysis, and [Auto Fin](cookbook/auto-fin/README.md) for file-native ETF event research based on CLS news and
- historical market reactions.
+- [2026.08] - Published [`@agentscope-ai/reme`](https://www.npmjs.com/package/@agentscope-ai/reme), providing a native
+ ReMe memory integration for DeepSeek Harness.
+- [2026.08] - Published the [ReMe blog](https://agentscope-ai.github.io/ReMe/?doc=en-reme-blog), an end-to-end introduction to its local-first memory
+ architecture, self-evolving workflows, hybrid search, proactive discovery, and benchmark results.
+- [2026.08] - [Experience-driven enhancement method](https://reme.agentscope.io/?doc=toolmemory-en) of agent tool-use execution built
+ on ReMe is available on [arXiv:2608.03403](https://arxiv.org/abs/2608.03403).
+- [2026.07] - Introduced optional Cookbooks: [Daily Paper](https://reme.agentscope.io/?doc=daily-paper-en) for paper discovery and
+ analysis, and [Auto Fin](https://reme.agentscope.io/?doc=auto-fin-en) for researching the latest 24 hours of topic-related CLS news
+ with local-memory search and validated historical wikilinks.
- [2026.07] - Our
paper [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
has been accepted to Findings of ACL 2026.
@@ -85,9 +91,26 @@ Install from source:
```bash
git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
-pip install -e ".[core]"
+pip install -e packages/reme_ai_studio -e ".[core]"
+cd website
+npm ci
+npm run build:static
+cd ..
```
+The static build requires Node.js 22.13 or newer and makes Studio available from the source tree.
+
+### DeepSeek Harness Integration
+
+With the ReMe service running, install the npm package into the DeepSeek Harness Web profile:
+
+```bash
+dsh plugin --profile web add @agentscope-ai/reme
+```
+
+The plugin recalls relevant ReMe memory before agent steps and submits completed main-agent turns for automatic memory
+capture. See the [TypeScript integration guide](packages/typescript/README.md#deepseek-harness) for configuration.
+
### Environment Variables
Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are
@@ -126,13 +149,19 @@ reme start service.port=8181
# reme start workspace_dir=/tmp/reme-demo service.port=8181
```
-After startup, check the service status. If you use a custom port, replace `2333` in the URL below with that port.
-
```bash
reme version
+reme health_check
+reme help
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
```
+### ReMe Studio (Optional)
+
+The `core` installation above includes Studio. After starting ReMe, open to browse, edit, and
+search the workspace. To add Studio to a base installation, use `pip install "reme-ai[web]"`. See the
+[ReMe Studio guide](https://reme.agentscope.io/?doc=studio-en) for source builds, configuration, and development.
+
### 5-Minute Memory Demo
With the service running, write a memory node, let ReMe index it, then retrieve it:
@@ -167,42 +196,65 @@ ReMe stores agent memory as readable Markdown.
Related: [[digest/wiki/memory-as-file.md]]
```
-## 🧑🍳 Cookbooks
+## 📚 Usage Guides
-Cookbooks are optional, end-to-end workflows assembled from ReMe jobs and steps. They are not enabled by the default
-configuration; select the cookbook's standalone configuration when starting ReMe. Each new cookbook will be added as
-another row in this table.
+These Markdown guides cover the main user workflows and the runtime contracts implemented by the current code.
-| Cookbook | Capability |
+| Guide | What you will learn |
+|-------|---------------------|
+| [Quick Start](docs/en/quick_start.md) | Install ReMe, start the service, and run the first file and memory operations. |
+| [Plugin Management](docs/en/plugin_management.md) | Install, inspect, validate, enable, and uninstall local ReMe plugins. |
+| [Memory as File](docs/en/memory_as_file.md) | Understand workspace layers, frontmatter, wikilinks, chunks, and the file-as-source-of-truth model. |
+| [Auto Memory](docs/en/auto_memory.md) | Preserve source conversations and distill reusable daily memory cards. |
+| [Auto Resource](docs/en/auto_resource.md) | Import supported text resources and turn them into source-linked daily cards. |
+| [Auto Dream](docs/en/auto_dream.md) and [Auto Link](docs/en/auto_link.md) | Consolidate daily notes into evolving digest nodes and readable wikilink relationships. |
+| [Memory Search](docs/en/memory_search.md) | Use BM25, optional vectors, RRF fusion, line-range recall, and progressive link expansion. |
+| [Proactive](docs/en/proactive.md) | Read interest topics safely and integrate them into a host agent's decision flow. |
+| [Agent Integration Scenarios](docs/en/reme_scene.md) | Choose among CLI/SKILL.md, HTTP, MCP, and embedded Python integration. |
+| [Framework](docs/en/framework.md) | Understand Application, Job, Step, Component, service, configuration, and lifecycle boundaries. |
+| [ReMe Blog](https://agentscope-ai.github.io/ReMe/?doc=en-reme-blog) | Read the product story, design rationale, examples, and benchmark summary. |
+
+## 🔌 Plugins
+
+Plugins are optional Python distributions that contribute Component, Step, or Job backends and configuration. They are
+installed separately and enabled explicitly by configuration. Auto Fin is the complete external-plugin example; Daily
+Paper remains an optional research workflow while it is migrated to the same packaging model.
+
+| Plugin / workflow | Capability |
|-----------------------------------------------|---------------------------------------------------------------------------------------------------------------|
-| [Daily Paper](cookbook/daily_paper/README.md) | Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief. |
-| [Auto Fin](cookbook/auto-fin/README.md) | Match CLS events to liquid ETFs, study historical reactions, and generate file-native research reports. |
+| [Daily Paper](https://reme.agentscope.io/?doc=daily-paper-en) | Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief. |
+| [Auto Fin](https://reme.agentscope.io/?doc=auto-fin-en) | Fetch topic-related CLS news, search ReMe history, and generate wikilink-backed Markdown reports. |
## 📁 Memory System
> Memory as File, File as Memory.
-ReMe treats **memory as files**, progressively processing raw conversations and external resources from `session/` and
-`resource/` into `daily/`, then consolidating them into reusable long-term memory nodes under `digest/`.
+ReMe treats **memory as files**, progressively processing filtered conversation source records and external resources
+from `session/` and `resource/` into `daily/`, then `digest/`. The default workspace is `.reme/` under the current
+directory; `workspace_dir=...` selects a different user-owned location.
### Directory Structure
```text
/
-├── metadata/ # Persistent system state such as indexes, graphs, and catalogs
-├── session/ # Raw conversations and agent sessions
+├── metadata/ # Rebuildable indexes, graphs, catalogs, and caches
+├── session/ # Conversation source records and agent sessions
│ ├── dialog/
-│ │ └── .jsonl
-│ ├── agentscope/
+│ │ └── .jsonl # Source messages saved by auto_memory
│ └── claude_code/
+│ └── .jsonl # ReMe copy used by auto_memory_cc
+├── mem_session/ # Generated agent-wrapper sessions/config, not user memory
+│ ├── agentscope/
+│ ├── claude_config/
+│ └── codex/
├── resource/ # External raw materials
+│ ├── . # Root-level files enter today's daily layer
│ └── YYYY-MM-DD/
│ └── .
├── daily/ # Lightly processed memory: daily facts, conversation summaries, resource readings
│ ├── YYYY-MM-DD.md
│ └── YYYY-MM-DD/
-│ ├── .md
-│ ├── .md
+│ ├── .md # Topic-named conversation or resource card
│ └── interests.yaml
└── digest/ # Long-term memory: personal facts, procedural experience, knowledge nodes
├── personal/
@@ -219,23 +271,16 @@ ReMe treats **memory as files**, progressively processing raw conversations and
## 🧭 Memory Design Philosophy
-> Capture raw dialogs and resources, refine them into long-term preferences, reusable experience, and valuable
-> knowledge,
-> while keeping the result editable by humans and agents.
+ReMe follows a capture → index → consolidate → recall loop. Workspace files remain the durable source of truth;
+everything under `metadata/` is rebuildable.
-### Automatic Memory Flow
-
-ReMe follows a capture → index → consolidate → recall loop. Conversations and resources first become daily memory cards;
-background jobs keep files searchable; `auto_dream` distills stable knowledge into `digest/`; agents recall memory
-through search, wikilinks, or proactive topics.
-
-| Capability | Entry point | What it does | Output |
-|---------------------------------------------|-------------------------------------------------|-------------------------------------------------------------------------------------------------|---------------------------------------------------------|
-| [`auto_memory`](docs/en/auto_memory.md) | Agent hook or `reme auto_memory` | Distills useful conversation facts while preserving the raw session. | `session/dialog/*.jsonl`, `daily//.md` |
-| [`auto_resource`](docs/en/auto_resource.md) | Resource watcher or `reme auto_resource` | Turns files under `resource//` into source-linked daily cards. | `daily//.md` |
-| [`auto_index`](docs/en/memory_search.md) | Background watcher or `reme reindex` | Maintains chunks, the BM25 index, the wikilink graph, and the optional embedding index. | Searchable `daily/`, `digest/`, and `resource/` content |
-| [`auto_dream`](docs/en/auto_dream.md) | `dream_cron` or `reme auto_dream` | Consolidates changed daily cards into long-term personal, procedure, and wiki memory. | `digest/**`, `daily//interests.yaml` |
-| [`proactive`](docs/en/proactive.md) | `reme proactive` before an agent decides to act | Reads topics generated by `auto_dream`; the host agent decides whether and how to mention them. | Structured topics from `daily//interests.yaml` |
+| Capability | Entry point | What it does | Output |
+|---------------------------------------------|-------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------------|
+| [`auto_memory`](docs/en/auto_memory.md) | Agent hook or `reme auto_memory` | Distills useful conversation facts while preserving a filtered conversation source record. | `session/dialog/*.jsonl`, `daily//.md` |
+| [`auto_resource`](docs/en/auto_resource.md) | Resource watcher or `reme auto_resource` | Turns files under `resource/` into source-linked, content-named daily cards. | `daily//.md` |
+| [`auto_index`](docs/en/memory_search.md) | Background watcher or `reme reindex` | Live-indexes Markdown in `daily/` and `digest/`; a full rebuild also scans `resource/` and JSONL. | Searchable chunks, BM25, wikilink graph, and optional vectors |
+| [`auto_dream`](docs/en/auto_dream.md) | `dream_cron` or `reme auto_dream` | By default, extracts up to five reusable units from changed files in the latest two-day window, then creates, corroborates, refines, or corrects digest nodes. | `digest/**`, `daily//interests.yaml` |
+| [`proactive`](docs/en/proactive.md) | `reme proactive` before an agent decides to act | Reads topics generated by `auto_dream`; the host agent decides whether and how to mention them. | Structured topics from `daily//interests.yaml` |
@@ -256,18 +301,41 @@ through search, wikilinks, or proactive topics.
+Search returns matching chunks with line ranges and bounded wikilink neighbors. Optional vector results are fused with
+BM25 through reciprocal rank fusion (RRF).
+
+> [!IMPORTANT]
+> `proactive` only reads and exposes interest topics produced by Auto Dream. It does not independently browse the web,
+> send notifications, or rewrite the knowledge base; the host agent decides whether and how to act on a topic.
+
+## 📊 Performance
+
+ReMe evaluates multi-session and long-context memory with agentic search-and-read workflows. The figures below are the
+published reference runs in this repository; model, prompt, dataset, and judging details are documented with each
+benchmark.
+
+| Benchmark | Setting | Sample size | Agentic score | Focus |
+|--------------------------------------------------------------|--------------|-------------------------:|--------------:|--------------------------------------------------------------------|
+| **[LongMemEval cleaned-s](https://reme.agentscope.io/?doc=longmemeval-en)** | **Overall** | **500 questions** | **89.4%** | Cross-session retrieval, knowledge updates, and temporal reasoning |
+| [BEAM](https://reme.agentscope.io/?doc=beam-en) | 100K context | 20 cases / 400 questions | 66.1% | Ten types of long-context memory tasks |
+| [BEAM](https://reme.agentscope.io/?doc=beam-en) | 1M context | 35 cases / 700 questions | 65.0% | Ultra-long conversation settings |
+
+ReMe also achieved a **0.580 PROC score across five user personas** in the repository's
+[π-Bench evaluation](https://reme.agentscope.io/?doc=pibench-en), 2.4% above NanoBot under the same test-model configuration. PROC
+measures proactive handling of hidden intent, clarification, cross-session preferences and conventions, task
+dependencies, and underspecified requests.
+
## 🤝 Agent-friendly Integration
ReMe can run as a local memory service accessed through the CLI, HTTP API, or MCP server, or it can be embedded in the
-host process through its Python API. Agents can choose the path that fits their runtime and share a local memory workspace
-when appropriate.
+host process through its Python API.
-| Agents | Recommended path | Available after integration |
-|---------------------------------------------|--------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------|
-| **QwenPaw** | Embed ReMe in-process through its Python API. | Reuse the host application's lifecycle and model config while keeping memory local and file-based. |
-| **Claude Code** | Start the streamable HTTP MCP service and install [plugins/claude_code/reme](plugins/claude_code/reme). | MCP recall tools, a `reme-memory` skill, and a Stop hook that records sessions automatically. |
-| **Hermes** | Start the HTTP service and install [plugins/hermes_agent](plugins/hermes_agent). | Recall relevant memory before model calls and enqueue `auto_memory` after each completed turn. |
-| **Other CLI-capable agents (OpenClaw/Codex)** | Copy or install [skills/reme_memory/SKILL.md](skills/reme_memory/SKILL.md). | Search, read, and write memory via the CLI; automatic recording requires explicit host lifecycle hooks. |
+| Agents | Recommended path | Available after integration |
+|-----------------------------------------------|---------------------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------|
+| **QwenPaw** | Embed ReMe in-process through its Python API. | Reuse the host application's lifecycle and model config while keeping memory local and file-based. |
+| **Claude Code** | Start the streamable HTTP MCP service and install [integrations/claude_code/reme](integrations/claude_code/reme). | MCP recall tools, a `reme-memory` skill, and a Stop hook that records sessions automatically. |
+| **Hermes** | Start the HTTP service and install [integrations/hermes_agent](integrations/hermes_agent). | Recall relevant memory before model calls and enqueue `auto_memory` after each completed turn. |
+| **Other CLI-capable agents (OpenClaw/Codex)** | Copy or install [skills/reme_memory/SKILL.md](skills/reme_memory/SKILL.md). | Search, read, and write memory via the CLI; automatic recording requires explicit host lifecycle hooks. |
Integration demos
@@ -299,26 +367,21 @@ when appropriate.
## 🛠️ ReMe Operations
-ReMe operates the workspace through a unified job interface exposed by the CLI. Agents usually only need retrieval,
-reading, writing, editing, and automatic memory commands. Lower-level indexing, frontmatter, and file operation commands
-are mainly for maintenance, debugging, or advanced integration. Run `reme help` for the full job list.
+Run `reme help` for the full job list. Common workspace and maintenance commands are:
| Command | Purpose |
|-------------------------------------------|----------------------------------------------------------------------------------------|
-| `reme start` | Start the local ReMe service. |
-| `reme version` / `reme health_check` | Check package and component status. |
| `reme status` | Show stateful data-component memory estimates and process RSS. |
| [`reme search`](docs/en/memory_search.md) | Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled. |
| `reme read` / `reme write` / `reme edit` | Inspect and maintain Markdown memory files. |
-| `reme auto_memory` | Turn conversation messages into daily memory cards. Requires LLM credentials. |
-| `reme auto_resource` | Interpret files under `resource/` into daily resource cards. Requires LLM credentials. |
-| `reme auto_dream` / `reme proactive` | Consolidate daily memory into long-term digest and surface topics worth attention. |
+| `reme traverse` / `reme graph_snapshot` | Explore wikilink neighborhoods or the category-rooted digest graph. |
+| `reme chat` | Stream a read-only, workspace-aware agent conversation. Requires LLM credentials. |
| `reme reindex` | Rebuild search and wikilink indexes from existing files. |
## 🤝 Community and Support
-- **Issues and requests**: Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) first. If there is no
- related discussion, open a new issue with background, expected behavior, and impact scope.
+- **Issues, requests, and help**: Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) first. If there is no
+ related discussion, open one with the background, expected behavior, and impact scope.
- **Code contributions**: Before making changes, read
the [contribution guide](https://docs.agentscope.io/reme/latest/en/contribution). Source, schemas, and tests are the
authoritative architecture and extension guide.
@@ -328,8 +391,7 @@ are mainly for maintenance, debugging, or advanced integration. Run `reme help`
`docs(zh): update quick start`.
- **Pre-submit checks**: Before submitting a PR, try to run `pre-commit run --all-files` and `pytest`. If tests that
depend on LLMs, embeddings, or external services cannot run, explain that in the PR.
-- **Get help**: Use [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues) for bugs and feature requests. Project
- documentation is available at [https://docs.agentscope.io/reme](https://docs.agentscope.io/reme).
+- **Documentation**: Visit [reme.agentscope.io](https://reme.agentscope.io).
### Contributors
diff --git a/README_ZH.md b/README_ZH.md
index 212566cb..8887749b 100644
--- a/README_ZH.md
+++ b/README_ZH.md
@@ -8,7 +8,7 @@
-
+
@@ -20,22 +20,23 @@
@@ -287,24 +352,20 @@ ReMe 既可以作为本地记忆服务,通过 CLI、HTTP API 或 MCP server
## 🛠️ ReMe Operations
-ReMe 通过 CLI 暴露的统一 job interface 操作 workspace。Agent 通常只需要使用检索、读取、写入、编辑和自动记忆相关命令;更底层的索引、
-frontmatter 和文件操作接口主要用于维护、调试或高级集成。完整 job 列表可以运行 `reme help` 查看。
+运行 `reme help` 可查看完整 job 列表。常用 workspace 与维护命令如下:
| 命令 | 作用 |
|-------------------------------------------|---------------------------------------------------------------|
-| `reme start` | 启动本地 ReMe 服务。 |
-| `reme version` / `reme health_check` | 检查包版本和组件状态。 |
| `reme status` | 查看有状态数据组件的内存估算及进程 RSS。 |
| [`reme search`](docs/zh/memory_search.md) | 默认使用 BM25 和 wikilink 检索,启用后增加向量检索。 |
| `reme read` / `reme write` / `reme edit` | 检查和维护 Markdown 记忆文件。 |
-| `reme auto_memory` | 将对话 messages 转为 daily 记忆卡片;需要 LLM 凭证。 |
-| `reme auto_resource` | 将 `resource/` 下的文件解读为 daily 资料卡片;需要 LLM 凭证。 |
-| `reme auto_dream` / `reme proactive` | 将 daily 记忆整理为长期 digest,并暴露值得关注的主题。 |
+| `reme traverse` / `reme graph_snapshot` | 浏览 wikilink 邻域或按类别组织的 digest 图。 |
+| `reme chat` | 与可感知 workspace 的只读 Agent 进行流式对话;需要 LLM 凭证。 |
| `reme reindex` | 基于已有文件重建检索和 wikilink 索引。 |
## 🤝 社区与支持
-- **问题反馈与需求**:请先查看 [Open Issues](https://github.com/agentscope-ai/ReMe/issues);如无相关讨论,可新建 Issue
+- **问题反馈、需求与帮助**:请先查看 [Open Issues](https://github.com/agentscope-ai/ReMe/issues);如无相关讨论,可新建 Issue
说明背景、目标行为和影响范围。
- **代码贡献**:改动前建议阅读 [贡献指南](https://docs.agentscope.io/reme/latest/zh/contribution)。架构与扩展方式以源码、schema
和测试为准。
@@ -313,8 +374,7 @@ frontmatter 和文件操作接口主要用于维护、调试或高级集成。
`docs(zh): update quick start`。
- **提交前检查**:提交 PR 前请尽量运行 `pre-commit run --all-files` 和 `pytest`;如有依赖 LLM、embedding 或外部服务的测试无法运行,请在
PR 中说明。
-- **获取帮助**:如需反馈 Bug 或功能请求,请使用 [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues);项目文档见
- [https://docs.agentscope.io/reme](https://docs.agentscope.io/reme)。
+- **项目文档**:访问 [reme.agentscope.io](https://reme.agentscope.io)。
### 贡献者
diff --git a/benchmark/pibench/.gitignore b/benchmark/pibench/.gitignore
new file mode 100644
index 00000000..6af97fd4
--- /dev/null
+++ b/benchmark/pibench/.gitignore
@@ -0,0 +1,14 @@
+# 含真实 API key,绝不入库
+env.sh
+
+# 运行时产物(含对话内容,勿入库)
+logs/
+outputs/
+reme_workspace/
+nanobot_workspace/
+
+# 数据符号链接(指向外部 π-Bench 仓库)
+data
+
+__pycache__/
+*.pyc
diff --git a/benchmark/pibench/README.md b/benchmark/pibench/README.md
new file mode 100644
index 00000000..c62249a0
--- /dev/null
+++ b/benchmark/pibench/README.md
@@ -0,0 +1,327 @@
+[中文版 / Chinese version](./README_ZH.md)
+
+# π-Bench Evaluation Suite
+
+A glue layer that connects the **ReMe agent (with persistent memory)** to
+**π-Bench** (Proactive Personal Assistant Benchmark). This directory contains
+only the minimal code and configuration needed for the integration: the
+π-Bench framework (`src/`), evaluation data (`data/`), the AppWorld tool
+environment, and ReMe itself are all **external third-party dependencies**,
+referenced in place via symlink and environment variables and never bundled
+with this suite.
+
+- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench (arXiv: 2605.14678)
+- ReMe: the root of the ReMe repository this suite lives in (recommended
+ location: `ReMe/benchmark/pibench/`)
+
+## 1. Architecture
+
+```
+π-Bench runner (src.main --mode run)
+ │ user_agent (simulated-user LLM) walks data/{persona}/episode.yaml
+ │ task by task, chatting with the agent over multiple turns and judging
+ │ hidden intents (PROC) during the run phase
+ ▼
+test server (π-Bench scripts/test_server.py, HTTP long-polling)
+ ▲ /send │ /poll
+ │ ▼
+bridge_reme.py ──────────────► ReMe Application (embedded as a library)
+ │ ├─ agent_wrapper: agent under test (AgentScope)
+ │ ├─ jobs: search / auto_memory / daily_write
+ │ └─ workspace: reme_workspace/{persona}/
+ │ (isolated persistent memory per persona)
+ └──── MCP ────► AppWorld MCP ────► AppWorld APIs (tool/app environment)
+
+π-Bench runner (src.main --mode eval)
+ judger (judge LLM) reads the traces and scores each checklist item (COMP)
+```
+
+Key points:
+- The bridge runs on **ReMe's own venv python** and uses ReMe as a library
+ (`resolve_app_config` + `Application`); **no ReMe source modification** is
+ required.
+- Every incoming user message automatically triggers a ReMe memory `search`
+ and injects the matched memories (tuning knobs in §8); on task end (reset)
+ the session is distilled into daily notes by `auto_memory`.
+- Tool calls executed by the agent (AppWorld MCP + ReMe job tools) are
+ captured per turn into the trace as `tool_steps`, so π-Bench
+ `tools_evaluation_path` scripts can score tool behavior (§7).
+- π-Bench's `data/`, `src/` and AppWorld are not part of this suite; install
+ π-Bench first (§3.1).
+
+## 2. Directory layout
+
+```
+pibench/
+├── README.md / README_ZH.md # this document (English / Chinese)
+├── env.sh.example # environment template (copy to env.sh, fill TODOs)
+├── bridge_reme.py # ReMe ↔ test server bridge (memory inject/save,
+│ # profile injection, tool-trace capture)
+├── run_persona.sh # full pipeline for ONE persona (5 services + run + eval)
+├── run_all.sh # batch over 5 personas (fresh/resume, default parallel=2)
+├── resume.py # checkpoint resume: completion detection + surgical
+│ # cleanup of interrupted tasks' residual memory
+├── fix_trace_logs.py # run outputs → ~/.nanobot/trace_logs conversion,
+│ # merging tool sidecars into turn files (pre-eval)
+├── .gitignore # excludes env.sh and all runtime artifacts
+└── config/
+ ├── models/reme.yaml # runner model config (model_id=reme)
+ └── bench/evaluation/trace_history.yaml # trace render policy (shipped with
+ # the suite; passed via --history-config-path)
+```
+
+Generated at runtime (all git-ignored): `data` (symlink), `logs/`, `outputs/`,
+`reme_workspace/`, `nanobot_workspace/`.
+
+## 3. Prerequisites (third-party, install first)
+
+### 3.1 π-Bench repository (with AppWorld)
+
+```bash
+git clone https://github.com/Simplified-Reasoning/Pi-Bench.git
+cd
+python3.11 -m venv .venv # scripts expect exactly this venv name
+source .venv/bin/activate
+pip install -e . # pibench runner (src.main)
+bash scripts/setup_appworld.sh # install AppWorld and download its data (large)
+```
+
+Post-install sanity checks:
+```bash
+ls data/ # should contain researcher marketer pharmacist law_trainee Financier
+.venv/bin/python -c "import src" && echo OK
+.venv/bin/appworld --help >/dev/null && echo OK
+```
+
+### 3.2 ReMe repository
+
+```bash
+cd # ReMe repository root (contains the reme/ package)
+python3.11 -m venv .venv # scripts expect exactly this venv name
+source .venv/bin/activate
+pip install -e . # or ReMe's own install flow; `import reme` must work
+```
+
+Sanity check: `.venv/bin/python -c "import reme; print('ok')"`
+
+## 4. Install this suite (step by step)
+
+1. **Place the suite** (recommended inside the ReMe repo so `REME_DIR` is
+ inferred automatically):
+ ```bash
+ cp -r pibench /benchmark/pibench
+ cd /benchmark/pibench
+ ```
+ If placed elsewhere, set `REME_DIR` explicitly in env.sh later.
+
+2. **Create the environment file and fill in the custom parameters**:
+ ```bash
+ cp env.sh.example env.sh
+ ```
+ Open `env.sh`; required items (marked TODO):
+ | Variable | Description |
+ |---|---|
+ | `PI_BENCH_ROOT` | π-Bench repo root (contains `src/` `data/` `.venv` `third_party/appworld`) |
+ | `USER_API_KEY` | API key of the simulated-user LLM (run phase, hidden-intent judging) |
+ | `JUDGER_API_KEY` | API key of the judger LLM (eval phase, checklist scoring) |
+ | `BRAVE_SEARCH_API_KEY` | optional; for the agent's web_search tool, `dummy` when unused |
+
+ Optional tuning: `REME_MODEL_NAME` (base model of the agent under test),
+ `REME_DIR`, `REME_LLM_BASE_URL` (default: DashScope OpenAI-compatible
+ endpoint).
+
+3. **Link the evaluation data** (referenced in place, never copied):
+ ```bash
+ ln -s "$PI_BENCH_ROOT/data" data
+ ```
+
+4. **(Optional) adjust model config** `config/models/reme.yaml`:
+ - `user_agent.model` / `judger.model`: model names for the simulated user
+ and the judger (literal values; π-Bench only expands `${ENV}` in
+ base_url/api_key).
+ - `run.turn_timeout`, `max_tool_iterations`, etc. as needed.
+
+5. **Smoke check** (does not start the evaluation):
+ ```bash
+ bash -n run_all.sh && bash -n run_persona.sh
+ source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
+ ```
+
+## 5. Run the evaluation
+
+> ⚠️ For long runs use `screen`, **not nohup** (nohup loses the permission
+> context in sandboxed/restricted environments and breaks child processes).
+
+```bash
+# Full official run: wipe ALL personas' memory/outputs/traces first (default
+# fresh mode, parallel=2)
+mkdir -p logs # on a fresh deployment logs/ does not exist yet
+screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
+
+# Checkpoint continuation (after an interruption; no wipe, completed tasks skipped)
+bash run_all.sh --resume
+
+# Other usages
+bash run_all.sh --parallel 1 # sequential
+bash run_all.sh --resume --skip-eval # run phase only
+bash run_persona.sh researcher # single persona (default --resume semantics)
+bash run_persona.sh researcher --fresh
+```
+
+Time reference: 5 personas × 20 tasks, parallel=2, fresh full run ≈ 12–14 hours.
+
+`run_all.sh` exits non-zero when any persona fails, so upstream automation
+cannot mistake a partially failed suite run for a success.
+
+## 6. Port allocation (parallel personas never collide)
+
+| persona | AppWorld API | AppWorld MCP | Test Server | ReMe internal service |
+|-------------|------|-------|------|-------|
+| marketer | 9001 | 10001 | 9998 | 18766 |
+| law_trainee | 9002 | 10002 | 9997 | 18767 |
+| pharmacist | 9003 | 10003 | 9996 | 18768 |
+| researcher | 9004 | 10004 | 9995 | 18765 |
+| Financier | 9005 | 10005 | 9994 | 18769 |
+
+## 7. Outputs and scores
+
+- **Results**: `outputs/reme/{persona}/{task}/eval/results/*_result.json`
+ - `overall_average_score`: checklist completeness (COMP; the judger scores
+ each criterion YES/NO, weighted across dependency groups)
+ - `overall_proactiveness_average_score`: proactiveness (PROC; the
+ user_agent judges hidden-intent coverage during the run phase; each task
+ file also carries the global average)
+- **Traces**: `~/.nanobot/trace_logs/reme/{persona}/{task}/...` (the scoring
+ input of the eval phase)
+- **Logs**: `logs/` (`suite_.log` per persona; `bridge_*`,
+ `runner_run/eval_*`, `appworld_*`, `test_server_*` per service)
+- **Memory store**: `reme_workspace/{persona}/` (daily/digest notes, raw
+ session dialogs, BM25 index, etc.; persistent across runs, wiped only in
+ fresh mode)
+
+Score summary:
+```bash
+grep -h "overall_average_score\|overall_proactiveness" \
+ outputs/reme/*/*/eval/results/*_result.json | head
+```
+
+### Tool-trace capture (tools_evaluation support)
+
+Some tasks define `objectives.tools_evaluation_path`: Python scripts that
+score tool behavior (e.g. "the temporary Todoist board was created and
+removed"). They need the executed tool calls in the trace. The pipeline:
+
+1. During `reply()`, the bridge reads the persisted AgentScope session state
+ after each turn and extracts the new `tool_call` / `tool_result` blocks
+ (tool name, arguments, result).
+2. Records are appended to
+ `outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`, tagged with the
+ turn number; AgentScope MCP names (`mcp__AppWorld__`) are normalized
+ to the π-Bench convention (`mcp_appworld_`).
+3. `fix_trace_logs.py` pairs each `{ts}-messages.jsonl` run with the
+ temporally closest tools sidecar and merges the records into the generated
+ `turn_N.json` files under the `tool_steps` key — one of the two
+ tool-history formats understood by π-Bench's `collect_tool_history()`.
+4. The eval phase then feeds `tool_steps` to both the tools_evaluation
+ scripts and the rendered `` seen by the judger.
+
+## 8. Memory mechanism (core design of this suite)
+
+- **Persona isolation**: each persona has its own workspace
+ (`reme_workspace/{persona}/`); the bridge takes an exclusive
+ `.bridge.lock` on it at startup, so two bridges can never share one memory
+ store, and one persona's memory search can never reach another's memories.
+- **Writes**: on task end (runner sends reset), the session is distilled by
+ the `auto_memory` job into daily notes and indexed by the background
+ watcher (BM25). Saves are non-blocking background tasks; the first message
+ of a new session waits for in-flight writes before searching.
+- **Reads**: on every incoming user message the bridge runs one `search` and
+ injects matched memories (`[Relevant memories from previous sessions]`
+ prefix); without matches the message passes through unchanged. Retrieval
+ tuning (bridge CLI flags, adjustable in run_persona.sh):
+ - `--search-limit 3`: at most 3 memory chunks injected per message;
+ - `--search-min-score 2.0`: weak BM25 hits are filtered out;
+ - `tool_context_id` rotates per task: chunks already injected within the
+ same task are not re-injected (ReMe's seen-chunk dedup, 24h TTL); normal
+ recall resumes after task boundaries.
+- **No self-leakage**: the in-progress session is not in the store yet
+ (saves happen on reset), so a task can never retrieve its own unfinished
+ content.
+- The agent also holds `search`/`daily_write` tools and can retrieve/record
+ proactively.
+- **System prompt**: `bridge_reme.py:build_system_prompt()` embeds the
+ HIDDEN-NEEDS protocol (proactiveness-oriented) and injects the persona
+ profile from `data/{persona}/profile.yaml` into every turn's system prompt.
+
+## 9. Checkpoint resume and memory-cleanup semantics
+
+- **Completion detection** (resume.py): scans
+ `outputs/reme/{persona}/**/history/*-log.jsonl` and
+ `outputs/reme/{persona}/run/*-log.jsonl` for
+ `Task finished task_id=X status=Y`. The status with the **newest event
+ timestamp** wins per task (record `timestamp`, falling back to
+ `timestamp_iso`, then to the timestamp embedded in the log file name) —
+ file category and read order alone can never override a newer record, so an
+ old run-level SUCCESS cannot mask a newer per-task ERROR. `SUCCESS /
+ MAX_TURNS / TIMEOUT` count as completed; `ERROR` and never-started tasks
+ are re-run (passed to the runner as repeated `--task-id` flags in episode
+ order).
+- **Answer-leak prevention**: an interrupted task may already have been
+ distilled into daily notes during graceful shutdown; re-running it with
+ that memory injected would inflate scores. Before resuming,
+ `resume.py cleanup` therefore removes residual memory **only for tasks
+ about to be re-run** (daily/digest notes, session/dialog, mem_session;
+ matched via `session_id = pibench_{task}_*`). Completed tasks' memories are
+ never touched. Daily index files are refreshed **only for the dates that
+ lost notes**, by full workspace-relative wikilink path — and when the ReMe
+ package is importable, the refresh reuses ReMe's own daily-index rebuild
+ logic (`refresh_day_index`), so same-named notes on other dates are never
+ modified.
+- **fresh vs resume are mutually exclusive**: a full memory wipe belongs to
+ fresh mode only (`run_all.sh` default, executed before any service starts);
+ resume never wipes.
+
+## 10. Customization entry points
+
+| Goal | Location |
+|---|---|
+| Base model of the agent under test | `REME_MODEL_NAME` in `env.sh` |
+| user_agent / judger models | `config/models/reme.yaml` |
+| Agent system prompt | `bridge_reme.py` `build_system_prompt()` |
+| Memory retrieval limit/threshold | `--search-limit/--search-min-score` on the bridge command in `run_persona.sh` |
+| ReMe internal parameters | **Do not modify ReMe source**; write a dedicated config modeled on `reme/config/beam.yaml` and override via `resolve_app_config(config=...)` (see bridge `_init_reme_app`) |
+| Turn timeout / tool iteration cap | `config/models/reme.yaml` `run.turn_timeout`, `model.max_tool_iterations` |
+
+## 11. Troubleshooting
+
+- **Port already in use**: the scripts auto-kill residual processes on the
+ four port groups above; if another suite (e.g. a different π-Bench
+ experiment) holds them, stop it first or change the port table in
+ run_persona.sh.
+- **Bridge exits immediately with workspace locked**: another bridge already
+ holds the same workspace; make sure each persona uses its own
+ `--workspace-dir` (the scripts allocate one per persona).
+- **Runner reports `${USER_API_KEY} ... empty`**: env.sh is unfilled or not
+ sourced; run_persona.sh sources env.sh automatically — when running the
+ runner manually, `source env.sh` first.
+- **`Cannot import 'reme'`**: the bridge must run with
+ `${REME_DIR}/.venv/bin/python` (run_persona.sh already does); otherwise
+ check that `REME_DIR` points at the ReMe repository root.
+- **AppWorld fails to start**: run `bash scripts/setup_appworld.sh` in the
+ π-Bench repo first (downloads data); inspect
+ `logs/appworld_*_.log`.
+- **trace_history.yaml not found**: the runner needs
+ `config/bench/evaluation/trace_history.yaml`; this suite ships the file and
+ passes it explicitly via `--history-config-path`, and run_persona.sh fails
+ fast with a clear error if it is missing. Always launch run_persona.sh /
+ run_all.sh from the suite directory.
+
+## 12. Privacy and security
+
+- The suite code and config templates contain **no real API keys, user names
+ or absolute paths**; real keys live only in your local `env.sh`
+ (git-ignored).
+- `logs/`, `outputs/`, `reme_workspace/` and `nanobot_workspace/` contain
+ full conversations and model outputs; never commit or share them.
+- The `data` symlink points at the official π-Bench evaluation data; respect
+ its data license terms.
diff --git a/benchmark/pibench/README_ZH.md b/benchmark/pibench/README_ZH.md
new file mode 100644
index 00000000..0a8b58d6
--- /dev/null
+++ b/benchmark/pibench/README_ZH.md
@@ -0,0 +1,284 @@
+# π-Bench 评测说明
+
+[English version](./README.md)
+
+将 **ReMe agent(带持久记忆)** 接入 **π-Bench**(Proactive Personal Assistant
+Benchmark)的胶水层评测套件。只含对接所需的最小代码与配置;π-Bench 框架
+(`src/`)、评测数据(`data/`)、AppWorld 工具环境、ReMe 本体均为**外部第三方
+依赖**,通过符号链接与环境变量原位引用,不随本套件分发。
+
+- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench (arXiv: 2605.14678)
+- ReMe: 你所在 ReMe 仓库的根目录(本套件推荐放在 `ReMe/benchmark/pibench/`)
+
+## 1. 架构总览
+
+```
+π-Bench runner (src.main --mode run)
+ │ user_agent(模拟用户 LLM)按 data/{persona}/episode.yaml 顺序
+ │ 逐任务、多轮地与 agent 对话,并在 run 阶段判定隐藏意图(PROC)
+ ▼
+test server (π-Bench scripts/test_server.py, HTTP 长轮询)
+ ▲ /send │ /poll
+ │ ▼
+bridge_reme.py ──────────────► ReMe Application(以库方式内嵌启动)
+ │ ├─ agent_wrapper: 被测 agent(AgentScope)
+ │ ├─ jobs: search / auto_memory / daily_write
+ │ └─ workspace: reme_workspace/{persona}/
+ │ (每 persona 独立持久记忆库,互不可见)
+ └──── MCP ────► AppWorld MCP ────► AppWorld API(工具/应用环境)
+
+π-Bench runner (src.main --mode eval)
+ judger(裁判 LLM)读取 trace,按 checklist 逐条 YES/NO 打分(COMP)
+```
+
+要点:
+- bridge 用 **ReMe 自己的 venv python** 运行,把 ReMe 当库用(`resolve_app_config`
+ + `Application`),**ReMe 源码零改动**。
+- 每条用户消息都会自动触发一次 ReMe memory `search` 并把命中记忆注入当前消息
+ (参数见 §8);任务结束(reset)时会话被 `auto_memory` 提炼为 daily 笔记落盘。
+- agent 执行的每一轮工具调用(AppWorld MCP + ReMe job 工具)都会被采集并以
+ `tool_steps` 形式写入 trace,供 π-Bench 的 `tools_evaluation_path` 脚本
+ 对工具行为评分(§7)。
+- π-Bench 的 `data/`、`src/`、AppWorld 均不属于本套件,需先装好 π-Bench(§3.1)。
+
+## 2. 目录结构
+
+```
+pibench/
+├── README.md / README_ZH.md # 本文档(英文 / 中文)
+├── env.sh.example # 环境配置模板(复制为 env.sh 后填写 TODO 项)
+├── bridge_reme.py # ReMe ↔ test server 桥接(记忆注入/保存、
+│ # profile 注入、工具调用轨迹采集)
+├── run_persona.sh # 单 persona 全流程(5 个服务 + run + eval)
+├── run_all.sh # 5 个 persona 批跑(fresh/resume,默认 2 并行)
+├── resume.py # 断点续跑:完成判定 + 中断任务残留记忆的外科清理
+├── fix_trace_logs.py # run 输出 → ~/.nanobot/trace_logs 转换,
+│ # 并把工具轨迹合并进 turn 文件(eval 前置)
+├── .gitignore # 排除 env.sh 与全部运行产物
+└── config/
+ ├── models/reme.yaml # runner 模型配置(model_id=reme)
+ └── bench/evaluation/trace_history.yaml # trace 渲染策略(随套件提供,
+ # 经 --history-config-path 显式传入)
+```
+
+运行时自动生成(均被 .gitignore 排除):`data`(符号链接)、`logs/`、
+`outputs/`、`reme_workspace/`、`nanobot_workspace/`。
+
+## 3. 前置依赖(第三方,先装好)
+
+### 3.1 π-Bench 仓库(含 AppWorld)
+
+```bash
+git clone https://github.com/Simplified-Reasoning/Pi-Bench.git
+cd
+python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
+source .venv/bin/activate
+pip install -e . # pibench runner(src.main)
+bash scripts/setup_appworld.sh # 安装 AppWorld 并下载其数据(体积较大,需网络)
+```
+
+装完自检:
+```bash
+ls data/ # 应含 researcher marketer pharmacist law_trainee Financier
+.venv/bin/python -c "import src" && echo OK
+.venv/bin/appworld --help >/dev/null && echo OK
+```
+
+### 3.2 ReMe 仓库
+
+```bash
+cd # ReMe 仓库根目录(含 reme/ 包)
+python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
+source .venv/bin/activate
+pip install -e . # 或按 ReMe 自身安装方式,保证 `import reme` 可用
+```
+
+自检:`.venv/bin/python -c "import reme; print('ok')"`
+
+## 4. 安装本套件(逐步)
+
+1. **放置套件**(推荐放进 ReMe 仓库,`REME_DIR` 可自动推断):
+ ```bash
+ cp -r pibench /benchmark/pibench
+ cd /benchmark/pibench
+ ```
+ 若放在其他位置,稍后在 env.sh 中显式设置 `REME_DIR`。
+
+2. **创建环境文件并填写自定义参数**:
+ ```bash
+ cp env.sh.example env.sh
+ ```
+ 打开 `env.sh`,必填项(标 TODO 的):
+ | 变量 | 说明 |
+ |---|---|
+ | `PI_BENCH_ROOT` | π-Bench 仓库根目录(含 `src/` `data/` `.venv` `third_party/appworld`) |
+ | `USER_API_KEY` | 模拟用户 LLM 的 API key(run 阶段判定隐藏意图) |
+ | `JUDGER_API_KEY` | 裁判 LLM 的 API key(eval 阶段 checklist 打分) |
+ | `BRAVE_SEARCH_API_KEY` | 可选;agent 的 web_search 工具用,不用填 `dummy` |
+
+ 可选调整:`REME_MODEL_NAME`(被测 agent 基模)、`REME_DIR`、
+ `REME_LLM_BASE_URL`(默认 DashScope OpenAI 兼容端点)。
+
+3. **链接评测数据**(π-Bench 数据原位引用,不复制):
+ ```bash
+ ln -s "$PI_BENCH_ROOT/data" data
+ ```
+
+4. **(可选)调整模型配置** `config/models/reme.yaml`:
+ - `user_agent.model` / `judger.model`:模拟用户与裁判的模型名(字面量,
+ π-Bench 仅对 base_url/api_key 做 `${ENV}` 展开)。
+ - `run.turn_timeout`、`max_tool_iterations` 等按需。
+
+5. **冒烟自检**(不启动评测):
+ ```bash
+ bash -n run_all.sh && bash -n run_persona.sh
+ source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
+ ```
+
+## 5. 运行评测
+
+> ⚠️ 长时间运行请放进 `screen`,**不要用 nohup**(nohup 在沙箱/受限环境下
+> 会丢失权限上下文导致子进程异常)。
+
+```bash
+# 完整正式评测:先清空全部 persona 的记忆/输出/trace,再从头跑(默认 fresh,2 并行)
+mkdir -p logs # 全新部署时 logs/ 尚不存在,先建再重定向
+screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
+
+# 断点续跑(中断后继续;不清记忆,跳过已完成任务)
+bash run_all.sh --resume
+
+# 其他用法
+bash run_all.sh --parallel 1 # 串行
+bash run_all.sh --resume --skip-eval # 只跑 run 阶段
+bash run_persona.sh researcher # 单 persona(默认 --resume 语义)
+bash run_persona.sh researcher --fresh
+```
+
+耗时参考:5 persona × 20 任务、2 并行,fresh 全量约 12–14 小时。
+
+任一 persona 失败时 `run_all.sh` 以非零状态退出,上层自动化不会把部分失败
+的评测误判为成功。
+
+## 6. 端口分配(多 persona 并行互不冲突)
+
+| persona | AppWorld API | AppWorld MCP | Test Server | ReMe 内部服务 |
+|-------------|------|-------|------|-------|
+| marketer | 9001 | 10001 | 9998 | 18766 |
+| law_trainee | 9002 | 10002 | 9997 | 18767 |
+| pharmacist | 9003 | 10003 | 9996 | 18768 |
+| researcher | 9004 | 10004 | 9995 | 18765 |
+| Financier | 9005 | 10005 | 9994 | 18769 |
+
+## 7. 输出与分数
+
+- **结果**:`outputs/reme/{persona}/{task}/eval/results/*_result.json`
+ - `overall_average_score`:checklist 完整度(COMP,judger 逐条 YES/NO 按依赖组加权)
+ - `overall_proactiveness_average_score`:主动性(PROC,run 阶段 user_agent
+ 判定隐藏意图覆盖率;每个任务文件同时携带全局均值)
+- **trace**:`~/.nanobot/trace_logs/reme/{persona}/{task}/...`(eval 的判分输入)
+- **日志**:`logs/`(`suite_.log` 为每 persona 总日志,`bridge_*`、
+ `runner_run/eval_*`、`appworld_*`、`test_server_*` 分服务)
+- **记忆库**:`reme_workspace/{persona}/`(daily/digest 笔记、session 原始对话、
+ BM25 索引等;跨运行持久,fresh 才清空)
+
+查看汇总:
+```bash
+grep -h "overall_average_score\|overall_proactiveness" \
+ outputs/reme/*/*/eval/results/*_result.json | head
+```
+
+### 工具轨迹采集(tools_evaluation 支持)
+
+部分任务定义了 `objectives.tools_evaluation_path`:用 Python 脚本对工具行为
+打分(例如"临时 Todoist 看板已创建并被删除")。这些脚本需要 trace 里有真实
+的工具调用记录。采集链路:
+
+1. 每轮 `reply()` 之后,bridge 读取 AgentScope 落盘的会话状态,提取本轮新增
+ 的 `tool_call` / `tool_result` 块(工具名、参数、结果)。
+2. 记录按 turn 编号追加写入
+ `outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`;AgentScope 的
+ MCP 工具名(`mcp__AppWorld__`)会规范化为 π-Bench 约定
+ (`mcp_appworld_`)。
+3. `fix_trace_logs.py` 将每个 `{ts}-messages.jsonl` 运行与时间上最接近的
+ tools 旁路文件配对,把记录合并进生成的 `turn_N.json` 的 `tool_steps`
+ 字段——这是 π-Bench `collect_tool_history()` 支持的两种工具轨迹格式之一。
+4. eval 阶段 `tool_steps` 既提供给 tools_evaluation 脚本,也会被渲染为
+ judger 可见的 ``。
+
+## 8. 记忆机制(本套件的核心设计)
+
+- **persona 隔离**:每个 persona 独立 workspace(`reme_workspace/{persona}/`),
+ bridge 启动时对 workspace 加 `.bridge.lock` 排他锁,两个 bridge 不可能共用
+ 同一记忆库;一个 persona 的 memory search 永远接触不到其他 persona 的记忆。
+- **写入**:任务结束(runner 发送 reset)时,会话经 `auto_memory` job 提炼为
+ daily 笔记落盘,后台 watcher 建 BM25 索引。保存为非阻塞后台任务,
+ 新会话首条消息会先等待在途写入完成再检索。
+- **读取**:bridge 每收到一条用户消息自动 `search` 一次并注入命中记忆
+ (`[Relevant memories from previous sessions]` 前缀),无命中则原样透传。
+ 检索参数(bridge 命令行,可在 run_persona.sh 中调整):
+ - `--search-limit 3`:每条消息最多注入 3 个记忆块;
+ - `--search-min-score 2.0`:过滤弱 BM25 命中;
+ - `tool_context_id` 按任务轮换:同一任务内已注入的记忆块不重复注入
+ (ReMe 自带 seen-chunk 去重,24h TTL),任务边界后恢复正常召回。
+- **无自泄漏**:进行中的会话尚未入库(save 发生在 reset),任务不会检索到
+ 自己未完成的内容。
+- agent 同时持有 `search`/`daily_write` 工具,可主动检索/记录。
+- **system prompt**:`bridge_reme.py:build_system_prompt()` 内置
+ HIDDEN-NEEDS 协议(面向 proactiveness),并把 `data/{persona}/profile.yaml`
+ 的 persona profile 注入每轮 system prompt。
+
+## 9. 断点续跑与记忆清理语义
+
+- **完成判定**(resume.py):扫描 `outputs/reme/{persona}/**/history/*-log.jsonl`
+ 与 `outputs/reme/{persona}/run/*-log.jsonl` 中的
+ `Task finished task_id=X status=Y`。每个任务以**事件时间最新**的记录为准
+ (优先取记录的 `timestamp`,回退 `timestamp_iso`,再回退日志文件名中的
+ 时间戳)——文件类别与读取顺序本身不能覆盖更新的记录,因此旧的 run 级
+ SUCCESS 不会掩盖更新的 per-task ERROR。`SUCCESS/MAX_TURNS/TIMEOUT` 记为
+ 完成,`ERROR`/未开始的任务重跑(按 episode 顺序以 `--task-id` 传给 runner)。
+- **防答案泄漏**:被中断的任务可能已在优雅退出时提炼成 daily 笔记,直接重跑会
+ 把答案注入、抬高分数。因此 resume 启动前 `resume.py cleanup` **只删除待重跑
+ 任务**的残留记忆(daily/digest 笔记、session/dialog、mem_session,按
+ `session_id = pibench_{task}_*` 匹配),已完成任务的记忆一律不动。daily
+ 索引**只刷新实际发生删除的日期**,按完整的 workspace 相对 wikilink 路径
+ 匹配;当 ReMe 包可导入时,刷新直接复用 ReMe 自带的 daily 索引重建逻辑
+ (`refresh_day_index`),不会误改其他日期下的同名笔记条目。
+- **fresh vs resume 互斥**:全量清记忆只属于 fresh 模式(`run_all.sh` 默认,
+ 在任何服务启动前执行);resume 永不清全量。
+
+## 10. 自定义与调优入口
+
+| 目标 | 位置 |
+|---|---|
+| 被测 agent 基模 | `env.sh` 的 `REME_MODEL_NAME` |
+| user_agent / judger 模型 | `config/models/reme.yaml` |
+| agent system prompt | `bridge_reme.py` `build_system_prompt()` |
+| 记忆检索条数/阈值 | `run_persona.sh` bridge 启动命令的 `--search-limit/--search-min-score` |
+| ReMe 内部参数 | **不要改 ReMe 源码**;仿照 `reme/config/beam.yaml` 写专有配置,经 `resolve_app_config(config=...)` 覆盖(见 bridge `_init_reme_app`) |
+| 轮超时/工具迭代上限 | `config/models/reme.yaml` `run.turn_timeout`、`model.max_tool_iterations` |
+
+## 11. 故障排查
+
+- **端口被占用**:脚本会自动 kill 上述 4 组端口上的残留进程;若与其他套件
+ (如别的 π-Bench 实验)冲突,请先停掉对方或改 run_persona.sh 的端口表。
+- **bridge 启动即退出,提示 workspace locked**:另一个 bridge 正占用同一
+ workspace;确认每个 persona 用各自的 `--workspace-dir`(脚本已按 persona 分配)。
+- **runner 报 `${USER_API_KEY} ... empty`**:env.sh 未填写或未生效;
+ run_persona.sh 会自动 source env.sh,手动运行 runner 时请先 `source env.sh`。
+- **`Cannot import 'reme'`**:bridge 必须用 `${REME_DIR}/.venv/bin/python` 运行
+ (run_persona.sh 已如此),或检查 `REME_DIR` 是否指向 ReMe 仓库根目录。
+- **AppWorld 启动失败**:先在 π-Bench 仓库执行 `bash scripts/setup_appworld.sh`
+ 下载数据;查看 `logs/appworld_*_.log`。
+- **trace_history.yaml 找不到**:runner 需要
+ `config/bench/evaluation/trace_history.yaml`;本套件已随附该文件并通过
+ `--history-config-path` 显式传入,run_persona.sh 启动前会做存在性检查,
+ 缺失时立即报出清晰错误。请始终从套件目录启动 run_persona.sh / run_all.sh。
+
+## 12. 隐私与安全
+
+- 套件代码与配置模板中**不含任何真实 API key、用户名或绝对路径**;
+ 真实 key 只存在于你本地的 `env.sh`(已被 .gitignore 排除)。
+- `logs/`、`outputs/`、`reme_workspace/`、`nanobot_workspace/` 含完整对话内容
+ 与模型输出,请勿提交仓库或外传。
+- `data` 符号链接指向 π-Bench 官方评测数据,请遵守其数据许可条款。
diff --git a/benchmark/pibench/bridge_reme.py b/benchmark/pibench/bridge_reme.py
new file mode 100755
index 00000000..dadd7877
--- /dev/null
+++ b/benchmark/pibench/bridge_reme.py
@@ -0,0 +1,1039 @@
+#!/usr/bin/env python3
+"""
+Bridge script: Connects ReMe agent to Pi-Bench Test Server.
+
+Uses ReMe's AgentScope-based agent wrapper directly as a library,
+with MCP integration to AppWorld and cross-session memory support.
+
+Flow:
+1. Poll Test Server /poll for user messages
+2. Forward to ReMe agent (via AgentScope)
+3. Extract reply text
+4. Send reply back to Test Server POST /send
+5. On session end (reset), save conversation as ReMe daily memory (non-blocking)
+6. On every incoming user message, trigger a ReMe memory search and inject
+ the relevant memories retrieved from previous sessions
+7. After every agent reply, capture the turn's tool calls (tool name,
+ arguments, result) from the persisted AgentScope session state and append
+ them to outputs////history/-tools.jsonl;
+ fix_trace_logs.py merges these into the per-turn traces as tool_steps so
+ π-Bench tools_evaluation scripts can score tool behavior.
+
+Key design decisions:
+- Memory saves are non-blocking (fire-and-forget asyncio tasks) so reset
+ acknowledgments are sent immediately and don't time out.
+- A pending-save tracker ensures the first message of a new session waits
+ for any in-flight memory writes to complete before searching.
+- User profile is loaded from data/{user_id}/profile.yaml and injected
+ into every turn's system prompt.
+- AgentScope session state is maintained via `resume` within a task,
+ and cleared on reset for cross-task isolation.
+- Memory search tuning: each search is capped at `--search-limit`
+ results (default 3), weak BM25 hits below `--search-min-score`
+ (default 2.0) are filtered, and a per-task `tool_context_id`
+ enables ReMe's seen-chunk dedup so the same memory chunk is not
+ re-injected on every turn of the same task.
+- Persona isolation: the workspace defaults to a per-user subdirectory
+ and an exclusive lock file guarantees that no two bridges can share
+ one memory store at runtime.
+
+Usage:
+ python bridge_reme.py [--test-server-url URL] [--reme-dir DIR]
+"""
+
+import argparse
+import asyncio
+import fcntl
+import json
+import logging
+import os
+import re
+import signal
+import sys
+from datetime import datetime
+from pathlib import Path
+from typing import Any, Dict, List, Optional
+
+import httpx
+import yaml
+
+logging.basicConfig(
+ level=logging.INFO,
+ format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
+ datefmt="%Y-%m-%d %H:%M:%S",
+)
+logger = logging.getLogger("bridge_reme")
+
+
+# ─── User Profile Loading ─────────────────────────────────────────────
+
+
+def load_user_profile(data_root: str, user_id: str) -> str:
+ """Load user profile YAML and return as formatted text for the system prompt.
+
+ Handles the full Pi-Bench profile schema: role (with sub-sections),
+ preferences, and long_term_goals.
+ """
+ profile_path = Path(data_root) / user_id / "profile.yaml"
+ if not profile_path.exists():
+ logger.warning("User profile not found: %s", profile_path)
+ return ""
+
+ with open(profile_path, "r", encoding="utf-8") as f:
+ profile = yaml.safe_load(f)
+
+ if not profile:
+ return ""
+
+ parts = []
+
+ # Role section: contains the full persona description
+ if "role" in profile and profile["role"]:
+ role_text = str(profile["role"]).strip()
+ if role_text:
+ parts.append(f"## User Profile\n{role_text}")
+
+ # Preferences section
+ if "preferences" in profile and profile["preferences"]:
+ prefs = profile["preferences"]
+ if isinstance(prefs, dict):
+ pref_lines = []
+ for k, v in prefs.items():
+ if v is not None and str(v).strip():
+ pref_lines.append(f"- {k}: {v}")
+ if pref_lines:
+ parts.append("## Preferences\n" + "\n".join(pref_lines))
+ elif isinstance(prefs, str):
+ parts.append(f"## Preferences\n{prefs}")
+
+ # Long-term goals
+ if "long_term_goals" in profile and profile["long_term_goals"]:
+ goals = profile["long_term_goals"]
+ if isinstance(goals, list):
+ goal_lines = [f"- {g}" for g in goals if g]
+ if goal_lines:
+ parts.append("## Long-term Goals\n" + "\n".join(goal_lines))
+ elif isinstance(goals, str):
+ parts.append(f"## Long-term Goals\n{goals}")
+
+ result = "\n\n".join(parts)
+ logger.info(
+ "Loaded profile for %s: %d chars, sections: %s",
+ user_id,
+ len(result),
+ [k for k in ["role", "preferences", "long_term_goals"] if k in profile],
+ )
+ return result
+
+
+def build_system_prompt(user_profile: str) -> str:
+ """Build the system prompt for the ReMe agent with profile context."""
+ base_prompt = """\
+You are a proactive personal assistant agent in a long-horizon evaluation. Be thorough, anticipatory,
+detail-oriented; use the user's profile, memory and tools proactively
+(AppWorld via MCP; memory `search`/`daily_write`; file tools).
+
+## HIDDEN-NEEDS PROTOCOL (MANDATORY)
+Every task carries implicit needs the user does not state. Before each substantive response:
+1. Derive the implicit needs of THIS task (method below), plus what the user's profile and past sessions imply.
+2. Cover EVERY need explicitly and specifically in this response.
+3. Anything you cannot cover now, you MUST still raise explicitly: one precise question or a concrete
+next step targeting exactly that need. Generic closers do not count.
+
+## HOW TO DERIVE IMPLICIT NEEDS
+- Entities: for every item the task involves (a paper, product, person, account, case, event), cover the
+attributes this user would need: what it is + key details, availability or cost, suitability/evaluation,
+how to proceed, risks, and alternatives.
+- Action completeness: if the task implies an action chain (prepare → execute → verify), cover every
+stage, including verification and closing the loop.
+- Context: apply everything the user's profile, constraints and past sessions imply (budget, size, format,
+style, tools, deadlines) without being reminded.
+- Structure: provide the format or verdict the user would expect (table, overall rating, pass/fail,
+conclusion-first) whenever applicable.
+
+## DELIVERABLE STRUCTURE
+What (conclusion first) → Why → How → Risks (limits, fallbacks) → Next steps.
+
+## STRICTNESS
+An implicit need counts only with specific, detailed content or a concrete action — vague or generic
+scores nothing. Deliver specifics in your FIRST response.
+"""
+
+ if user_profile:
+ base_prompt += f"\n\n---\n\n{user_profile}\n"
+
+ base_prompt += (
+ "\n\n---\n\nAlways respond in the same language as the user's message. Use tools proactively to help the user."
+ )
+ return base_prompt
+
+
+# ─── ReMe Bridge ──────────────────────────────────────────────────────
+
+
+class ReMeBridge:
+ """Bridge between Pi-Bench Test Server and ReMe agent."""
+
+ def __init__(
+ self,
+ test_server_url: str = "http://localhost:9999",
+ appworld_mcp_url: str = "http://localhost:10000/mcp",
+ reme_dir: str = "",
+ data_root: str = "data",
+ user_id: str = "researcher",
+ poll_timeout: int = 30,
+ workspace_dir: str = "",
+ model_name: str = "qwen3.6-plus",
+ model_base_url: str = "",
+ model_api_key: str = "",
+ reme_port: int = 18765,
+ search_limit: int = 3,
+ search_min_score: float = 2.0,
+ outputs_dir: str = "",
+ model_id: str = "reme",
+ ):
+ self.test_server_url = test_server_url.rstrip("/")
+ self.appworld_mcp_url = appworld_mcp_url
+ self.reme_dir = Path(reme_dir).resolve() if reme_dir else None
+ self.data_root = Path(data_root)
+ self.user_id = user_id
+ self.poll_timeout = poll_timeout
+ if workspace_dir:
+ self.workspace_dir = Path(workspace_dir)
+ else:
+ # Per-persona default so two bridges can never share a memory store.
+ root = os.environ.get("REME_WORKSPACE_ROOT", "/tmp/reme_pibench_workspaces")
+ self.workspace_dir = Path(root) / user_id
+ self.model_name = model_name
+ self.model_base_url = model_base_url
+ self.model_api_key = model_api_key
+ self.reme_port = reme_port
+ self.search_limit = search_limit
+ self.search_min_score = search_min_score
+ # Runner outputs root; the tool-trace sidecar files are written next
+ # to the runner's *-messages.jsonl history files.
+ if outputs_dir:
+ self.outputs_dir = Path(outputs_dir).resolve()
+ else:
+ self.outputs_dir = (self.data_root.parent / "outputs").resolve()
+ self.model_id = model_id
+ # Task generation counter: rotated on every reset so the search dedup
+ # context (tool_context_id) is scoped to a single task.
+ self.task_seq = 0
+ self._workspace_lock_fd: Optional[int] = None
+
+ # Tool-trace capture state (per bridge lifetime):
+ # - turn counter per chat (each user message = one π-Bench turn)
+ # - already-seen session content block ids (tool_call / tool_result)
+ # - tool_call blocks waiting for their tool_result block
+ # - sidecar file timestamp per chat (fixed at first capture)
+ self._turn_by_chat: Dict[str, int] = {}
+ self._seen_tool_block_ids: set = set()
+ self._pending_tool_calls: Dict[str, Dict] = {}
+ self._tools_file_ts: Dict[str, str] = {}
+
+ self.client: Optional[httpx.AsyncClient] = None
+ self.running = False
+
+ # ReMe components
+ self.app = None
+ self.agent_wrapper = None
+ self.auto_memory_job = None
+ self.search_job = None
+
+ # Session state
+ self.user_profile_text = ""
+ self.session_messages: Dict[str, List[Dict]] = {}
+ self.agent_session_id: Optional[str] = None
+
+ # Non-blocking memory save tracking
+ self._pending_memory_tasks: List[asyncio.Task] = []
+
+ async def start(self):
+ """Initialize the ReMe application and bridge components."""
+ self.client = httpx.AsyncClient(timeout=300.0, trust_env=False)
+ self.running = True
+
+ # Enforce per-persona workspace isolation before anything else: an
+ # exclusive lock guarantees no other bridge can use this memory store.
+ self._acquire_workspace_lock()
+ if self.workspace_dir.name != self.user_id:
+ logger.warning(
+ "Workspace basename %r != user_id %r; cross-persona isolation "
+ "relies on each bridge having its own workspace_dir",
+ self.workspace_dir.name,
+ self.user_id,
+ )
+ logger.info(
+ "Memory isolation: user=%s workspace=%s reme_port=%d",
+ self.user_id,
+ self.workspace_dir,
+ self.reme_port,
+ )
+
+ # Load user profile
+ self.user_profile_text = load_user_profile(str(self.data_root), self.user_id)
+ logger.info("User profile loaded: %d chars", len(self.user_profile_text))
+
+ # Initialize ReMe application
+ await self._init_reme_app()
+
+ logger.info(
+ "Bridge started: test_server=%s appworld_mcp=%s user=%s model=%s",
+ self.test_server_url,
+ self.appworld_mcp_url,
+ self.user_id,
+ self.model_name,
+ )
+
+ def _acquire_workspace_lock(self):
+ """Take an exclusive lock on the workspace (persona isolation guard)."""
+ self.workspace_dir.mkdir(parents=True, exist_ok=True)
+ lock_path = self.workspace_dir / ".bridge.lock"
+ fd = os.open(str(lock_path), os.O_CREAT | os.O_RDWR)
+ try:
+ fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB)
+ except BlockingIOError as exc:
+ os.close(fd)
+ raise SystemExit(
+ f"Workspace {self.workspace_dir} is already locked by another "
+ f"bridge process; each persona needs its own workspace_dir.",
+ ) from exc
+ os.ftruncate(fd, 0)
+ os.write(fd, f"pid={os.getpid()} user={self.user_id}\n".encode())
+ self._workspace_lock_fd = fd
+
+ def _release_workspace_lock(self):
+ if self._workspace_lock_fd is not None:
+ try:
+ fcntl.flock(self._workspace_lock_fd, fcntl.LOCK_UN)
+ os.close(self._workspace_lock_fd)
+ except OSError:
+ pass
+ self._workspace_lock_fd = None
+
+ async def _init_reme_app(self):
+ """Initialize the ReMe application with proper configuration."""
+ # Add reme to Python path so imports work
+ if self.reme_dir:
+ reme_str = str(self.reme_dir)
+ if reme_str not in sys.path:
+ sys.path.insert(0, reme_str)
+
+ try:
+ from reme.config import resolve_app_config
+ from reme.application import Application
+ except ImportError as exc:
+ raise RuntimeError(
+ "Cannot import 'reme'. Run the bridge with the ReMe venv "
+ "python, or pass --reme-dir pointing to the ReMe repo root.",
+ ) from exc
+
+ # Set environment variables for ReMe LLM config expansion
+ os.environ["LLM_MODEL_NAME"] = self.model_name
+ if self.model_base_url:
+ os.environ["LLM_BASE_URL"] = self.model_base_url
+ if self.model_api_key:
+ os.environ["LLM_API_KEY"] = self.model_api_key
+ # Ensure BRAVE_SEARCH_API_KEY is set (required by some tools)
+ if not os.environ.get("BRAVE_SEARCH_API_KEY"):
+ os.environ["BRAVE_SEARCH_API_KEY"] = "dummy"
+
+ # Ensure workspace exists
+ self.workspace_dir.mkdir(parents=True, exist_ok=True)
+
+ # Load .env from reme dir if available
+ environment = {}
+ if self.reme_dir:
+ env_path = self.reme_dir / ".env"
+ if env_path.exists():
+ with open(env_path, "r", encoding="utf-8") as f:
+ for line in f:
+ line = line.strip()
+ if line and not line.startswith("#") and "=" in line:
+ key, _, value = line.partition("=")
+ environment[key.strip()] = value.strip()
+
+ # Override with explicit values
+ if self.model_base_url:
+ environment["LLM_BASE_URL"] = self.model_base_url
+ if self.model_api_key:
+ environment["LLM_API_KEY"] = self.model_api_key
+ environment["LLM_MODEL_NAME"] = self.model_name
+
+ # Use resolve_app_config to load default.yaml and merge overrides
+ reme_config = resolve_app_config(
+ log_config=False,
+ workspace_dir=str(self.workspace_dir),
+ service={"backend": "http", "host": "127.0.0.1", "port": self.reme_port},
+ environment=environment,
+ )
+
+ try:
+ self.app = Application(**reme_config)
+ await self.app.start()
+
+ # Get components
+ self.agent_wrapper = self.app.context.components.get("agent_wrapper", {}).get("default")
+ if self.agent_wrapper is None:
+ raise RuntimeError("agent_wrapper component 'default' not found")
+
+ # Get jobs for memory operations
+ self.auto_memory_job = self.app.context.jobs.get("auto_memory")
+ self.search_job = self.app.context.jobs.get("search")
+
+ logger.info("ReMe initialized OK")
+ logger.info(" agent_wrapper: %s", getattr(self.agent_wrapper, "name", "default"))
+ logger.info(" auto_memory: %s", "yes" if self.auto_memory_job else "no")
+ logger.info(" search: %s", "yes" if self.search_job else "no")
+ logger.info(" total jobs: %d", len(self.app.context.jobs))
+
+ except Exception as e:
+ logger.exception("Failed to initialize ReMe: %s", e)
+ raise
+
+ async def stop(self):
+ """Stop the bridge and cleanup."""
+ self.running = False
+ self._release_workspace_lock()
+
+ # Wait for pending memory saves
+ if self._pending_memory_tasks:
+ logger.info("Waiting for %d pending memory saves...", len(self._pending_memory_tasks))
+ for task in self._pending_memory_tasks:
+ try:
+ await asyncio.wait_for(task, timeout=60.0)
+ except (asyncio.TimeoutError, Exception) as e:
+ logger.warning("Pending memory save timed out or failed: %s", e)
+
+ if self.app:
+ try:
+ await self.app.close()
+ except Exception as e:
+ logger.warning("Error closing ReMe app: %s", e)
+ if self.client:
+ await self.client.aclose()
+ self.client = None
+ logger.info("Bridge stopped")
+
+ def _create_mcp_client(self):
+ """Create an MCP client for AppWorld."""
+ from agentscope.mcp import MCPClient, HttpMCPConfig
+
+ return MCPClient(
+ name="AppWorld",
+ is_stateful=False,
+ mcp_config=HttpMCPConfig(
+ url=self.appworld_mcp_url,
+ timeout=120.0,
+ ),
+ )
+
+ def _build_agent_toolkit(self, mcp_client, job_tool_names: List[str]):
+ """Build the agent toolkit so AppWorld MCP tools are really registered.
+
+ agent_wrapper.reply() accepts a prebuilt ``toolkit`` kwarg but does
+ not wire a bare ``mcps`` kwarg into the agent, so the toolkit is
+ assembled here: ReMe job tools (search / auto_memory / daily_write)
+ plus the AppWorld MCP client. Returns None when the wrapper lacks
+ the required hooks; the caller then falls back to plain kwargs.
+ """
+ try:
+ from agentscope.tool import Toolkit
+ except ImportError as exc:
+ logger.warning("Cannot import agentscope Toolkit: %s", exc)
+ return None
+
+ make_tool = getattr(type(self.agent_wrapper), "_make_tool", None)
+ if make_tool is None:
+ logger.warning(
+ "agent_wrapper %s cannot wrap jobs as tools; falling back to kwargs (MCP tools may be unavailable)",
+ type(self.agent_wrapper).__name__,
+ )
+ return None
+
+ tools = []
+ for name in job_tool_names:
+ job = self.app.context.jobs.get(name) if self.app is not None else None
+ if job is None:
+ continue
+ try:
+ tools.append(make_tool(job, None, None))
+ except Exception as exc:
+ logger.warning("Failed to wrap job '%s' as agent tool: %s", name, exc)
+
+ # AgentScope builtin file tools (bash/read/write/edit/glob/grep).
+ # A prebuilt toolkit bypasses _build_agent's builtin-tools branch,
+ # and since ReMe commit e05b201d builtins are opt-in, they must be
+ # added here explicitly — the agent needs them to read task asset
+ # files (e.g. instrument_booking_brief.md) from its workspace.
+ builtin = getattr(self.agent_wrapper, "_builtin_tools", None)
+ if builtin is not None:
+ try:
+ tools.extend(builtin("all", sequential_tool_calls=True))
+ except Exception as exc:
+ logger.warning("Failed to add builtin tools: %s", exc)
+
+ try:
+ return Toolkit(tools=tools, mcps=[mcp_client])
+ except Exception as exc:
+ logger.warning("Failed to build agent toolkit: %s", exc)
+ return None
+
+ # ─── Memory Operations ──────────────────────────────────────────
+
+ async def _wait_for_pending_memory_saves(self):
+ """Wait for all in-flight memory save tasks to complete."""
+ if not self._pending_memory_tasks:
+ return
+ logger.info(
+ "Waiting for %d pending memory saves before search...",
+ len(self._pending_memory_tasks),
+ )
+ tasks = self._pending_memory_tasks[:]
+ self._pending_memory_tasks.clear()
+ for task in tasks:
+ try:
+ await asyncio.wait_for(task, timeout=120.0)
+ except asyncio.TimeoutError:
+ logger.warning("Memory save task timed out (120s)")
+ except Exception as e:
+ logger.warning("Memory save task failed: %s", e)
+
+ async def _search_memory(self, query: str, tool_context_id: str = "") -> str:
+ """Search ReMe memory for relevant context from previous sessions."""
+ if not self.search_job:
+ return ""
+ try:
+ response = await self.search_job(
+ query=query,
+ limit=self.search_limit,
+ min_score=self.search_min_score,
+ tool_context_id=tool_context_id,
+ )
+ if response.success and response.answer:
+ returned = response.metadata.get("counts", {}).get("returned", "?")
+ logger.info(
+ "Memory search: hits=%s limit=%d min_score=%s ctx=%s",
+ returned,
+ self.search_limit,
+ self.search_min_score,
+ tool_context_id or "-",
+ )
+ return response.answer
+ except Exception as e:
+ logger.warning("Memory search failed: %s", e)
+ return ""
+
+ async def _do_save_session_memory(self, chat_id: str, messages: List[Dict]):
+ """Actually perform the session memory save (runs as background task)."""
+ if not self.auto_memory_job or not messages:
+ return
+
+ session_id = f"pibench_{chat_id}_{datetime.now().strftime('%Y%m%d_%H%M%S')}"
+
+ # Convert messages to auto_memory format
+ memory_messages = []
+ for msg in messages:
+ role = msg.get("role", "user")
+ memory_messages.append(
+ {
+ "role": role,
+ "name": msg.get("name", "user" if role == "user" else "assistant"),
+ "content": msg.get("content", ""),
+ "created_at": msg.get("timestamp", datetime.now().isoformat()),
+ },
+ )
+
+ try:
+ logger.info(
+ "Saving session memory: chat_id=%s messages=%d",
+ chat_id,
+ len(memory_messages),
+ )
+ response = await self.auto_memory_job(
+ messages=memory_messages,
+ session_id=session_id,
+ memory_hint=(
+ f"Pi-Bench evaluation session for task {chat_id}. "
+ f"User persona: {self.user_id}. "
+ f"Save key decisions, actions taken, important outcomes, "
+ f"and any user preferences or context that may be useful "
+ f"for future sessions."
+ ),
+ )
+ if response.success:
+ preview = (response.answer or "OK")[:200]
+ logger.info("Session memory saved: %s", preview)
+ else:
+ logger.warning("Memory save returned unsuccessful: %s", response.answer)
+ except Exception as e:
+ logger.exception("Error saving session memory: %s", e)
+
+ def _schedule_memory_save(self, chat_id: str, messages: List[Dict]):
+ """Schedule a non-blocking memory save task."""
+ if not self.auto_memory_job or not messages:
+ return
+
+ task = asyncio.create_task(
+ self._do_save_session_memory(chat_id, messages),
+ name=f"memory_save_{chat_id}",
+ )
+ self._pending_memory_tasks.append(task)
+
+ # Clean up completed tasks from the tracking list
+ self._pending_memory_tasks = [t for t in self._pending_memory_tasks if not t.done()]
+
+ # ─── Tool Trace Capture ─────────────────────────────────────────
+
+ MCP_TOOL_NAME_RE = re.compile(r"^mcp__(?P[A-Za-z0-9_-]+?)__(?P.+)$")
+
+ @classmethod
+ def _normalize_tool_name(cls, name: str) -> str:
+ """Map AgentScope MCP tool names to the π-Bench / nanobot convention.
+
+ AgentScope registers MCP tools as ``mcp____`` while
+ π-Bench task objectives and tools_evaluation scripts expect
+ ``mcp__`` (lower-case client, single underscores).
+ """
+ match = cls.MCP_TOOL_NAME_RE.match(name)
+ if match:
+ return f"mcp_{match.group('client').lower()}_{match.group('tool')}"
+ return name
+
+ @staticmethod
+ def _tool_result_text(output: Any) -> str:
+ """Flatten an AgentScope tool-result payload into plain text."""
+ if isinstance(output, str):
+ return output
+ if isinstance(output, list):
+ parts = [
+ str(item.get("text") or "") for item in output if isinstance(item, dict) and item.get("type") == "text"
+ ]
+ return "\n".join(parts)
+ return ""
+
+ def _tools_file_for(self, chat_id: str) -> Path:
+ """Return (and lazily name) the tool-trace sidecar file of a task."""
+ task_dir = self.outputs_dir / self.model_id / self.user_id / chat_id / "history"
+ task_dir.mkdir(parents=True, exist_ok=True)
+ if chat_id not in self._tools_file_ts:
+ self._tools_file_ts[chat_id] = datetime.now().strftime("%Y%m%d_%H%M%S")
+ return task_dir / f"{self._tools_file_ts[chat_id]}-tools.jsonl"
+
+ def _capture_tool_calls(self, chat_id: str, session_id: str) -> None:
+ """Record the current turn's tool calls from the AgentScope session.
+
+ After every reply, the agent wrapper dumps the full session context to
+ ``/mem_session/agentscope/.jsonl``. This method
+ scans that dump for tool_call / tool_result content blocks that were
+ not seen before and appends the completed pairs to the per-task
+ sidecar file consumed by fix_trace_logs.py.
+ """
+ if not session_id:
+ return
+ mem_session_dir = "mem_session"
+ app_config = getattr(getattr(self.app, "context", None), "app_config", None)
+ if app_config is not None and getattr(app_config, "mem_session_dir", None):
+ mem_session_dir = app_config.mem_session_dir
+ state_path = self.workspace_dir / mem_session_dir / "agentscope" / f"{session_id}.jsonl"
+ if not state_path.is_file():
+ return
+ try:
+ lines = state_path.read_text(encoding="utf-8").splitlines()
+ except OSError as exc:
+ logger.warning("Cannot read agent session state %s: %s", state_path, exc)
+ return
+
+ turn = self._turn_by_chat.get(chat_id, 0)
+ records: List[Dict] = []
+ for line in lines[1:]: # line 1 is the state header, not a message
+ try:
+ msg = json.loads(line)
+ except json.JSONDecodeError:
+ continue
+ content = msg.get("content")
+ if not isinstance(content, list):
+ continue
+ for block in content:
+ if not isinstance(block, dict):
+ continue
+ block_id = str(block.get("id") or "")
+ block_type = block.get("type")
+ if block_type not in ("tool_call", "tool_result"):
+ continue
+ # A tool_result block reuses its tool_call's id, so dedup
+ # must be keyed on (type, id), not id alone.
+ seen_key = (block_type, block_id)
+ if not block_id or seen_key in self._seen_tool_block_ids:
+ continue
+ if block_type == "tool_call":
+ self._seen_tool_block_ids.add(seen_key)
+ arguments = block.get("input") or ""
+ if isinstance(arguments, str):
+ try:
+ arguments = json.loads(arguments)
+ except json.JSONDecodeError:
+ arguments = {"raw_input": arguments}
+ self._pending_tool_calls[block_id] = {
+ "turn": turn,
+ "name": self._normalize_tool_name(str(block.get("name") or "")),
+ "arguments": arguments,
+ }
+ elif block_type == "tool_result":
+ self._seen_tool_block_ids.add(seen_key)
+ call = self._pending_tool_calls.pop(block_id, None)
+ if call is None:
+ continue
+ call["result"] = self._tool_result_text(block.get("output"))
+ records.append(call)
+
+ if records:
+ tools_path = self._tools_file_for(chat_id)
+ with open(tools_path, "a", encoding="utf-8") as f:
+ for record in records:
+ f.write(json.dumps(record, ensure_ascii=False) + "\n")
+ logger.info(
+ "Tool trace: chat=%s turn=%d captured=%d -> %s",
+ chat_id,
+ turn,
+ len(records),
+ tools_path.name,
+ )
+
+ # ─── Message Processing ─────────────────────────────────────────
+
+ async def process_message(self, _sender_id: str, chat_id: str, content: str) -> Optional[str]:
+ """Process a user message through the ReMe agent."""
+ # Track session messages for later memory save
+ if chat_id not in self.session_messages:
+ self.session_messages[chat_id] = []
+
+ self.session_messages[chat_id].append(
+ {
+ "role": "user",
+ "name": "user",
+ "content": content,
+ "timestamp": datetime.now().isoformat(),
+ },
+ )
+
+ # Each user message is one π-Bench turn; tool records captured after
+ # the reply below are tagged with this turn number.
+ self._turn_by_chat[chat_id] = self._turn_by_chat.get(chat_id, 0) + 1
+
+ # Build system prompt with user profile
+ system_prompt = build_system_prompt(self.user_profile_text)
+
+ # Create MCP client for AppWorld
+ mcp_client = self._create_mcp_client()
+
+ # Determine which reme jobs to expose as tools
+ job_tools = []
+ if self.search_job:
+ job_tools.append("search")
+ if self.auto_memory_job:
+ job_tools.extend(["auto_memory", "daily_write"])
+
+ # On EVERY incoming user message, automatically trigger a ReMe
+ # memory search and inject the relevant memories retrieved from
+ # previous sessions. Memory is only surfaced through search
+ # (relevance-filtered), never dumped wholesale. The in-progress
+ # session is not in the store yet (saves happen on reset), so a
+ # task can never retrieve its own partial content.
+ memory_context = ""
+ # Wait for any in-flight memory saves so the store is complete
+ # before searching (no-op when nothing is pending).
+ await self._wait_for_pending_memory_saves()
+ if self.search_job:
+ try:
+ memory_context = await self._search_memory(
+ content,
+ tool_context_id=f"pibench_{self.user_id}_task_{self.task_seq}",
+ )
+ if memory_context:
+ logger.info("Found relevant memory: %d chars", len(memory_context))
+ except Exception as e:
+ logger.warning("Memory search failed: %s", e)
+
+ try:
+ # Prepend memory context if available
+ user_message = content
+ if memory_context:
+ user_message = (
+ f"[Relevant memories from previous sessions]\n"
+ f"{memory_context}\n\n"
+ f"[Current user message]\n{content}"
+ )
+
+ # Call ReMe agent with MCP tools and memory tools
+ reply_kwargs = {
+ "system_prompt": system_prompt,
+ "permission_mode": "bypass",
+ }
+ toolkit = self._build_agent_toolkit(mcp_client, job_tools)
+ if toolkit is not None:
+ # Prebuilt toolkit: registers AppWorld MCP tools AND job tools.
+ reply_kwargs["toolkit"] = toolkit
+ else:
+ # Fallback path (kept for wrappers without toolkit support).
+ reply_kwargs["mcps"] = [mcp_client]
+ if job_tools:
+ reply_kwargs["job_tools"] = job_tools
+
+ # Resume existing session for multi-turn continuity within same task
+ if self.agent_session_id:
+ reply_kwargs["resume"] = self.agent_session_id
+
+ result = await self.agent_wrapper.reply(user_message, **reply_kwargs)
+
+ reply_text = result.get("result", "")
+ session_id = result.get("session_id", "")
+
+ if session_id:
+ self.agent_session_id = session_id
+
+ # Capture the turn's tool calls for π-Bench tools_evaluation.
+ self._capture_tool_calls(chat_id, session_id or self.agent_session_id or "")
+
+ # Track the assistant reply
+ self.session_messages[chat_id].append(
+ {
+ "role": "assistant",
+ "name": "assistant",
+ "content": reply_text,
+ "timestamp": datetime.now().isoformat(),
+ },
+ )
+
+ logger.info("Reply: %d chars, session=%s", len(reply_text), session_id)
+ return reply_text
+
+ except Exception as e:
+ logger.exception("Error processing message: %s", e)
+ return None
+
+ async def handle_reset(self, chat_id: str):
+ """Handle session reset: schedule non-blocking memory save and clear state."""
+ messages = self.session_messages.pop(chat_id, [])
+ if messages:
+ self._schedule_memory_save(chat_id, messages)
+ # Clear agent session for cross-task isolation
+ self.agent_session_id = None
+ # New task boundary: rotate the search dedup context so memories can be
+ # recalled again in the next task while repeats within a task are filtered.
+ self.task_seq += 1
+
+ # ─── Test Server Communication ──────────────────────────────────
+
+ async def poll_test_server(self) -> Optional[List[Dict[str, Any]]]:
+ """Poll Test Server for pending messages.
+
+ Returns None on connection/response errors so the caller can back off;
+ an empty list means a successful poll with no pending messages.
+ """
+ try:
+ resp = await self.client.get(
+ f"{self.test_server_url}/poll",
+ params={"timeout": self.poll_timeout},
+ )
+ if resp.is_success:
+ data = resp.json()
+ messages = data.get("messages", [])
+ if messages:
+ logger.info("Received %d messages", len(messages))
+ return messages
+ logger.warning("Poll returned HTTP %s", resp.status_code)
+ except Exception as e:
+ logger.warning("Poll error: %s", e)
+ return None
+
+ async def send_to_test_server(self, chat_id: str, content: str) -> bool:
+ """Send reply back to Test Server."""
+ payload = {
+ "chat_id": chat_id,
+ "content": content,
+ "media": [],
+ "meta": {},
+ }
+ try:
+ resp = await self.client.post(
+ f"{self.test_server_url}/send",
+ json=payload,
+ )
+ if resp.is_success:
+ logger.info("Sent reply: chat_id=%s len=%d", chat_id, len(content))
+ return True
+ logger.error("Failed to send: %s", resp.status_code)
+ except Exception:
+ logger.exception("Error sending to Test Server")
+ return False
+
+ # ─── Main Loop ──────────────────────────────────────────────────
+
+ def _install_signal_handlers(self):
+ """Install SIGTERM/SIGINT handlers for graceful shutdown."""
+ loop = asyncio.get_running_loop()
+ for sig_name in ("SIGTERM", "SIGINT"):
+ sig = getattr(signal, sig_name, None)
+ if sig is not None:
+ loop.add_signal_handler(sig, self._handle_shutdown_signal, sig_name)
+
+ def _handle_shutdown_signal(self, sig_name: str):
+ """Handle shutdown signal: stop the bridge loop gracefully."""
+ logger.info("Received %s, initiating graceful shutdown...", sig_name)
+ self.running = False
+
+ async def run(self):
+ """Main bridge loop."""
+ await self.start()
+ self._install_signal_handlers()
+
+ consecutive_poll_errors = 0
+ try:
+ while self.running:
+ messages = await self.poll_test_server()
+
+ if messages is None:
+ # Back off on poll failure to avoid a tight error loop.
+ consecutive_poll_errors += 1
+ if consecutive_poll_errors in (1, 5, 20) or consecutive_poll_errors % 50 == 0:
+ logger.warning(
+ "Poll failure #%d, backing off",
+ consecutive_poll_errors,
+ )
+ await asyncio.sleep(min(2 ** min(consecutive_poll_errors, 5), 30))
+ continue
+ consecutive_poll_errors = 0
+
+ for msg in messages:
+ sender_id = msg.get("sender_id", "unknown")
+ chat_id = msg.get("chat_id", "default")
+ content = msg.get("content", "")
+
+ if not content:
+ continue
+
+ logger.info(
+ "Processing: sender=%s chat=%s len=%d",
+ sender_id,
+ chat_id,
+ len(content),
+ )
+
+ # Check for reset/new-session signal
+ if content.strip().lower() in ("reset", "new session", "/new"):
+ logger.info("Reset signal: chat_id=%s", chat_id)
+ await self.handle_reset(chat_id)
+ await self.send_to_test_server(chat_id, "New session started")
+ continue
+
+ # Forward to ReMe agent
+ reply = await self.process_message(sender_id, chat_id, content)
+
+ if reply:
+ await self.send_to_test_server(chat_id, reply)
+ else:
+ logger.warning("No reply for chat_id=%s", chat_id)
+ await self.send_to_test_server(
+ chat_id,
+ "[Error: Agent failed to generate response]",
+ )
+
+ except KeyboardInterrupt:
+ logger.info("Interrupted by user")
+ finally:
+ # Save any remaining session memories
+ for cid in list(self.session_messages.keys()):
+ messages = self.session_messages.pop(cid, [])
+ if messages:
+ self._schedule_memory_save(cid, messages)
+ # Wait for all pending memory saves to complete
+ if self._pending_memory_tasks:
+ logger.info(
+ "Waiting for %d pending memory saves to complete...",
+ len(self._pending_memory_tasks),
+ )
+ await self._wait_for_pending_memory_saves()
+ await self.stop()
+
+
+async def main():
+ """CLI entrypoint: parse arguments and run the ReMe bridge."""
+ # Ignore SIGHUP to prevent bridge from being killed (same fix as qwenpaw)
+ signal.signal(signal.SIGHUP, signal.SIG_IGN)
+
+ parser = argparse.ArgumentParser(
+ description="Bridge between Pi-Bench Test Server and ReMe agent",
+ )
+ parser.add_argument("--test-server-url", default="http://localhost:9999")
+ parser.add_argument("--appworld-mcp-url", default="http://localhost:10000/mcp")
+ parser.add_argument(
+ "--reme-dir",
+ default="",
+ help="ReMe repo root. Optional when 'reme' is already importable "
+ "(e.g. running inside the ReMe repo with its own venv).",
+ )
+ parser.add_argument("--data-root", default="data")
+ parser.add_argument("--user-id", default="researcher")
+ parser.add_argument("--poll-timeout", type=int, default=30)
+ parser.add_argument("--workspace-dir", default="")
+ parser.add_argument("--model-name", default="qwen3.6-plus")
+ parser.add_argument("--model-base-url", default="")
+ parser.add_argument("--model-api-key", default="")
+ parser.add_argument(
+ "--reme-port",
+ type=int,
+ default=18765,
+ help="Port for ReMe's internal HTTP service (must be unique per concurrently running bridge).",
+ )
+ parser.add_argument(
+ "--search-limit",
+ type=int,
+ default=3,
+ help="Max memory chunks injected per user message.",
+ )
+ parser.add_argument(
+ "--search-min-score",
+ type=float,
+ default=2.0,
+ help="Min BM25 score for injected memory chunks.",
+ )
+ parser.add_argument(
+ "--outputs-dir",
+ default="",
+ help="Runner outputs root for tool-trace sidecar files "
+ "(default: /../outputs, matching the runner layout).",
+ )
+ parser.add_argument(
+ "--model-id",
+ default="reme",
+ help="model_id used under outputs//...; must match "
+ "config/models/reme.yaml so traces align with the runner.",
+ )
+
+ args = parser.parse_args()
+
+ bridge = ReMeBridge(
+ test_server_url=args.test_server_url,
+ appworld_mcp_url=args.appworld_mcp_url,
+ reme_dir=args.reme_dir,
+ data_root=args.data_root,
+ user_id=args.user_id,
+ poll_timeout=args.poll_timeout,
+ workspace_dir=args.workspace_dir,
+ model_name=args.model_name,
+ model_base_url=args.model_base_url,
+ model_api_key=args.model_api_key,
+ reme_port=args.reme_port,
+ search_limit=args.search_limit,
+ search_min_score=args.search_min_score,
+ outputs_dir=args.outputs_dir,
+ model_id=args.model_id,
+ )
+
+ await bridge.run()
+
+
+if __name__ == "__main__":
+ asyncio.run(main())
diff --git a/benchmark/pibench/config/bench/evaluation/trace_history.yaml b/benchmark/pibench/config/bench/evaluation/trace_history.yaml
new file mode 100644
index 00000000..2a42c6ea
--- /dev/null
+++ b/benchmark/pibench/config/bench/evaluation/trace_history.yaml
@@ -0,0 +1,53 @@
+version: 1
+
+format:
+ root_tag: trace
+ turn_tag: turn
+ message_tag: message
+ file_tag: file
+ tool_call_tag_prefix: tool_call
+ tool_result_tag_prefix: tool_result
+
+text_policy:
+ default:
+ truncate_chars: 1200
+ mask_newlines: false
+ field_overrides:
+ files_read:
+ truncate_chars: 40000
+ assistant_content:
+ truncate_chars: 40000
+ tool_result_content:
+ truncate_chars: 40000
+
+fields:
+ turn:
+ include_session_key: false
+
+ files:
+ enabled: true
+
+ messages:
+ enabled: true
+ include_message_role_attr: true
+ include_message_index_attr: false
+ include_system: false
+ include_user: true
+ include_assistant_thinking_content: false
+ include_assistant_thinking_reasoning: false
+ include_assistant_content: true
+ include_assistant_reasoning: false
+ include_assistant_tool_calls: false
+ require_matching_tool_call: true
+
+ tool_calls:
+ include_tool_call_id: false
+ tools:
+ web_fetch:
+ enabled: true
+ include_tool_call_keys: [url]
+ include_tool_result: false
+ web_search:
+ enabled: true
+ include_tool_call_keys: [query]
+ include_tool_result: false
diff --git a/benchmark/pibench/config/models/reme.yaml b/benchmark/pibench/config/models/reme.yaml
new file mode 100644
index 00000000..4573dffc
--- /dev/null
+++ b/benchmark/pibench/config/models/reme.yaml
@@ -0,0 +1,40 @@
+# ReMe model configuration for Pi-Bench
+# Uses ReMe's AgentScope agent with Dashscope as the LLM backend
+
+model:
+ model: reme
+ base_url: "http://localhost:8088"
+ api_key: "dummy"
+ provider: custom
+ max_tokens: 16384
+ max_tool_iterations: 120
+ memory_window: 100
+
+user_agent:
+ model: qwen3.8-max
+ base_url: "${USER_BASE_URL}"
+ api_key: "${USER_API_KEY}"
+ temperature: 0.0
+ request_timeout: 360.0
+
+judger:
+ model: qwen3.8-max
+ base_url: "${JUDGER_BASE_URL}"
+ api_key: "${JUDGER_API_KEY}"
+ temperature: 0.0
+ request_timeout: 360.0
+
+tools:
+ brave_search_api_key: "${BRAVE_SEARCH_API_KEY}"
+ web_search_max_results: 10
+
+nanobot:
+ trace_logs_dir: "~/.nanobot/trace_logs"
+ workspace_dir: "~/.nanobot/workspace"
+ copy_task_assets_to_workspace: true
+
+run:
+ output_dir: outputs
+ log_level: INFO
+ user_mode: llm
+ turn_timeout: 2400.0
diff --git a/benchmark/pibench/env.sh.example b/benchmark/pibench/env.sh.example
new file mode 100644
index 00000000..62c35575
--- /dev/null
+++ b/benchmark/pibench/env.sh.example
@@ -0,0 +1,57 @@
+#!/bin/bash
+# ═══════════════════════════════════════════════════════════════════════
+# pibench evaluation suite - environment configuration template
+# Usage: cp env.sh.example env.sh, then fill in the TODO items below.
+# ⚠️ env.sh contains real API keys; never commit or share it
+# (already excluded via .gitignore).
+# ═══════════════════════════════════════════════════════════════════════
+
+SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
+
+# ─── TODO: π-Bench repository root ────────────────────────────────────
+# Must contain src/, data/, scripts/test_server.py, third_party/appworld
+# and .venv (see README setup).
+export PI_BENCH_ROOT=""
+
+# ─── ReMe repository ──────────────────────────────────────────────────
+# Defaults to two levels above this directory (the layout this suite uses
+# when placed at ReMe/benchmark/pibench); point it at the actual ReMe
+# repository root if the suite lives elsewhere.
+export REME_DIR="${REME_DIR:-$(cd "${SUITE_DIR}/../.." && pwd)}"
+
+# ─── Base model of the agent under test (LLM used by the ReMe agent) ──
+export REME_MODEL_NAME="${REME_MODEL_NAME:-qwen3.6-plus}"
+
+# ─── LLM service endpoint (default: DashScope OpenAI-compatible; any
+# OpenAI-compatible endpoint works) ────────────────────────────────
+DASHSCOPE_BASE_URL="https://dashscope.aliyuncs.com/compatible-mode/v1"
+export REME_LLM_BASE_URL="${REME_LLM_BASE_URL:-${DASHSCOPE_BASE_URL}}"
+
+# ─── TODO: API keys ───────────────────────────────────────────────────
+# USER_API_KEY : drives the simulated user LLM (run phase; judges whether
+# hidden intents are satisfied and asks follow-ups)
+# JUDGER_API_KEY: drives the judger LLM (eval phase; scores the checklist)
+# The two may be identical; one strong model is recommended for both.
+export USER_BASE_URL="${DASHSCOPE_BASE_URL}"
+export USER_API_KEY="TODO-fill-in-user-agent-api-key"
+
+export JUDGER_BASE_URL="${DASHSCOPE_BASE_URL}"
+export JUDGER_API_KEY="TODO-fill-in-judger-api-key"
+
+# The ReMe agent's key reuses USER_API_KEY by default (no need to repeat
+# it when both use the same service and key).
+export REME_LLM_API_KEY="${REME_LLM_API_KEY:-${USER_API_KEY}}"
+
+# Brave Search (optional; used by the agent's web_search tool - use
+# "dummy" when not needed).
+export BRAVE_SEARCH_API_KEY="TODO-optional-brave-search-key-or-dummy"
+
+# ─── Persistent memory workspaces (one subdirectory per persona,
+# created automatically) ───────────────────────────────────────────
+export REME_WORKSPACE_ROOT="${REME_WORKSPACE_ROOT:-${SUITE_DIR}/reme_workspace}"
+
+# ─── Variables consumed by ReMe's default.yaml model config expansion;
+# do not remove ────────────────────────────────────────────────────
+export LLM_MODEL_NAME="${REME_MODEL_NAME}"
+export LLM_BASE_URL="${REME_LLM_BASE_URL}"
+export LLM_API_KEY="${REME_LLM_API_KEY}"
diff --git a/benchmark/pibench/fix_trace_logs.py b/benchmark/pibench/fix_trace_logs.py
new file mode 100755
index 00000000..2a855e17
--- /dev/null
+++ b/benchmark/pibench/fix_trace_logs.py
@@ -0,0 +1,198 @@
+#!/usr/bin/env python3
+"""Convert reme_eval run outputs into eval-compatible trace logs.
+
+outputs/{model_id}/{user_id}/{task_id}/history/{ts}-messages.jsonl
+ -> ~/.nanobot/trace_logs/{model_id}/{user_id}/{task_id}/{ts}/turn_N.json
+
+The bridge additionally writes {ts}-tools.jsonl sidecar files next to the
+message histories: one JSON object per executed tool call with fields
+{turn, name, arguments, result}. Each messages run is paired with the
+temporally closest sidecar, and the records are merged into the generated
+turn files under the "tool_steps" key, which is one of the tool-history
+formats π-Bench's collect_tool_history() understands. Without this step,
+tools_evaluation scripts would see no tool evidence at all.
+
+Usage: python fix_trace_logs.py [user_id ...] (no args = all users)
+"""
+
+import json
+import re
+import sys
+from datetime import datetime
+from pathlib import Path
+
+SUITE_DIR = Path(__file__).resolve().parent
+OUTPUTS_DIR = SUITE_DIR / "outputs"
+TRACE_LOGS_DIR = Path.home() / ".nanobot" / "trace_logs"
+
+MESSAGES_FILE_RE = re.compile(r"^(\d{8}_\d{6})-messages\.jsonl$")
+TOOLS_FILE_RE = re.compile(r"^(\d{8}_\d{6})-tools\.jsonl$")
+TIME_FORMAT = "%Y%m%d_%H%M%S"
+# A tool sidecar belongs to the messages run that started at most this many
+# seconds earlier (the bridge stamps the sidecar when the task's first user
+# message arrives, shortly after the runner opened the messages file).
+MAX_PAIR_DELTA_SECONDS = 6 * 3600
+
+
+def _to_epoch(timestamp: str) -> float:
+ """Parse a YYYYMMDD_HHMMSS timestamp into epoch seconds."""
+ try:
+ return datetime.strptime(timestamp, TIME_FORMAT).timestamp()
+ except ValueError:
+ return 0.0
+
+
+def load_tool_records(tools_file: Path) -> dict:
+ """Group sidecar tool records by turn number."""
+ by_turn: dict = {}
+ try:
+ with open(tools_file, "r", encoding="utf-8") as f:
+ for line in f:
+ line = line.strip()
+ if not line:
+ continue
+ try:
+ record = json.loads(line)
+ except json.JSONDecodeError:
+ continue
+ if not isinstance(record, dict) or not record.get("name"):
+ continue
+ turn = int(record.get("turn") or 0)
+ by_turn.setdefault(turn, []).append(
+ {
+ "name": record["name"],
+ "arguments": record.get("arguments", {}),
+ "result": record.get("result", ""),
+ },
+ )
+ except OSError as exc:
+ print(f" WARNING: cannot read tool sidecar {tools_file}: {exc}")
+ return by_turn
+
+
+def pair_tool_sidecars(message_runs: list, tool_runs: list) -> dict:
+ """Pair each messages run with the temporally closest unused tool sidecar.
+
+ Fresh runs produce exactly one messages file and one sidecar per task;
+ re-runs append matching pairs, so sorted greedy nearest-timestamp
+ matching is stable. Sidecars farther away than MAX_PAIR_DELTA_SECONDS
+ (e.g. leftovers of a crashed bridge) stay unpaired.
+ """
+ pairing: dict = {}
+ unused = list(tool_runs)
+ for msg_ts, _ in message_runs:
+ best_delta = None
+ best_item = None
+ for tool_ts, tool_path in unused:
+ delta = abs(_to_epoch(tool_ts) - _to_epoch(msg_ts))
+ if best_delta is None or delta < best_delta:
+ best_delta = delta
+ best_item = (tool_ts, tool_path)
+ if best_delta is not None and best_item is not None and best_delta <= MAX_PAIR_DELTA_SECONDS:
+ pairing[msg_ts] = best_item[1]
+ unused.remove(best_item)
+ return pairing
+
+
+def build_turns(messages: list) -> list:
+ """Split the flat message list into per-turn [user, assistant] groups."""
+ turns = []
+ i = 0
+ while i < len(messages):
+ turn_msgs = []
+ if messages[i]["role"] == "user":
+ turn_msgs.append({"role": "user", "content": messages[i]["message"]})
+ i += 1
+ if i < len(messages) and messages[i]["role"] == "assistant":
+ turn_msgs.append({"role": "assistant", "content": messages[i]["message"]})
+ i += 1
+ if not turn_msgs:
+ i += 1 # defensive: never spin on unexpected roles
+ continue
+ turns.append(turn_msgs)
+ return turns
+
+
+def convert_task(model_id: str, user_id: str, task_dir: Path) -> None:
+ """Convert one task's history dir into trace turn files with tool_steps."""
+ history_dir = task_dir / "history"
+ if not history_dir.is_dir():
+ return
+
+ message_runs = []
+ tool_runs = []
+ for msg_file in history_dir.glob("*-messages.jsonl"):
+ match = MESSAGES_FILE_RE.match(msg_file.name)
+ if match:
+ message_runs.append((match.group(1), msg_file))
+ for tools_file in history_dir.glob("*-tools.jsonl"):
+ match = TOOLS_FILE_RE.match(tools_file.name)
+ if match:
+ tool_runs.append((match.group(1), tools_file))
+ if not message_runs:
+ return
+
+ message_runs.sort(key=lambda item: item[0])
+ tool_runs.sort(key=lambda item: item[0])
+ pairing = pair_tool_sidecars(message_runs, tool_runs)
+
+ print(f"\n{model_id}/{user_id}/{task_dir.name}")
+ for timestamp, msg_file in message_runs:
+ trace_dir = TRACE_LOGS_DIR / model_id / user_id / task_dir.name / timestamp
+ trace_dir.mkdir(parents=True, exist_ok=True)
+
+ messages = []
+ with open(msg_file, "r", encoding="utf-8") as f:
+ for line in f:
+ line = line.strip()
+ if not line:
+ continue
+ msg = json.loads(line)
+ if msg.get("role") == "user" and msg.get("message") == "/new":
+ continue
+ messages.append(msg)
+
+ tools_file = pairing.get(timestamp)
+ tools_by_turn = load_tool_records(tools_file) if tools_file else {}
+ if tools_file is not None:
+ print(f" {timestamp}: paired tool sidecar {tools_file.name}")
+
+ turns = build_turns(messages)
+ for turn_idx, turn_msgs in enumerate(turns, start=1):
+ turn_data = {"messages": turn_msgs}
+ tool_steps = tools_by_turn.get(turn_idx)
+ if tool_steps:
+ turn_data["tool_steps"] = tool_steps
+ turn_file = trace_dir / f"turn_{turn_idx}.json"
+ with open(turn_file, "w", encoding="utf-8") as f:
+ json.dump(turn_data, f, indent=2, ensure_ascii=False)
+ tool_total = sum(len(steps) for steps in tools_by_turn.values())
+ print(f" {timestamp}: {len(turns)} turns, {tool_total} tool step(s) -> {trace_dir}")
+
+
+def convert_outputs(user_filter=None):
+ """Convert message history JSONL files into per-turn trace JSON files."""
+ if not OUTPUTS_DIR.exists():
+ print(f"outputs dir not found: {OUTPUTS_DIR}")
+ return
+
+ for model_dir in sorted(OUTPUTS_DIR.iterdir()):
+ if not model_dir.is_dir():
+ continue
+ model_id = model_dir.name
+
+ for user_dir in sorted(model_dir.iterdir()):
+ if not user_dir.is_dir():
+ continue
+ user_id = user_dir.name
+ if user_filter and user_id not in user_filter:
+ continue
+
+ for task_dir in sorted(user_dir.iterdir()):
+ if task_dir.is_dir():
+ convert_task(model_id, user_id, task_dir)
+
+
+if __name__ == "__main__":
+ convert_outputs(set(sys.argv[1:]) or None)
+ print("\ndone")
diff --git a/benchmark/pibench/resume.py b/benchmark/pibench/resume.py
new file mode 100755
index 00000000..d40b4317
--- /dev/null
+++ b/benchmark/pibench/resume.py
@@ -0,0 +1,332 @@
+#!/usr/bin/env python3
+"""Checkpoint-resume support for the reme_eval suite.
+
+Completion source of truth:
+ - outputs/reme///history/*-log.jsonl (per-task logs,
+ flushed incrementally, survive mid-run kills)
+ - outputs/reme//run/*-log.jsonl (run-level logs,
+ may be truncated if the process was killed before flush)
+ lines: "Task finished task_id= status="
+ A task counts as COMPLETED when its latest terminal status is one of
+ SUCCESS / MAX_TURNS / TIMEOUT. ERROR or never-started tasks stay pending.
+
+ "Latest" is decided by EVENT TIME, not by file category or read order:
+ each record's "timestamp" (epoch seconds, or "timestamp_iso" as fallback)
+ is compared across per-task and run-level logs alike, with the timestamp
+ embedded in the log file name as a last-resort fallback. This keeps an
+ old run-level SUCCESS from overriding a newer per-task ERROR when the
+ re-run died before the new run-level log captured the task.
+
+Commands:
+ remaining [--json]
+ Print task_ids still to run, in data//episode.yaml order
+ (one per line; --json prints {"completed": [...], "remaining": [...]}).
+
+ cleanup [--dry-run]
+ Surgically remove residual memory artifacts of tasks that are about
+ to be RE-RUN (i.e. pending tasks that left partial state because a
+ previous run was interrupted). This prevents answer leakage: an
+ interrupted task's conversation may already have been distilled into
+ daily notes during graceful shutdown, and re-running the task with
+ that memory injected would inflate scores.
+
+ Removed artifacts (only for pending tasks with residual state):
+ - daily//.md whose frontmatter session_id matches
+ pibench__*, plus a refresh of ONLY the daily index of
+ the affected date(s) (daily/.md), matched by the full
+ workspace-relative note path, never by bare file name
+ - digest notes with matching session_id
+ - session/dialog/pibench__*.jsonl
+ - mem_session/**.jsonl files containing pibench__
+ When the ReMe package is importable, the daily index refresh reuses
+ ReMe's own rebuild logic (reme.steps.file_io._daily_index.
+ refresh_day_index); otherwise index lines are dropped by exact
+ wikilink path match. Either way, indexes of other dates are never
+ touched. The ReMe watcher (init_changes_step) detects the deleted
+ daily notes on next bridge startup and removes them from the BM25
+ index itself.
+
+ Completed tasks' memories are NEVER touched by this command.
+
+Design note (resume vs memory-wipe conflict):
+ A full memory wipe is a suite-level action of fresh mode (run_all.sh
+ without --resume) and happens before any service starts. Resume mode
+ never wipes; it only performs the surgical cleanup above. The two modes
+ are mutually exclusive, so a resumed run can never lose the cross-session
+ memory accumulated by completed tasks.
+"""
+
+import asyncio
+import json
+import os
+import re
+import sys
+from datetime import datetime
+from pathlib import Path
+
+import yaml
+
+try: # Reuse ReMe's daily-index rebuild when running inside the ReMe venv.
+ from reme.steps.file_io._daily_index import refresh_day_index
+except ImportError: # pragma: no cover - depends on runtime venv
+ refresh_day_index = None
+
+SUITE_DIR = Path(__file__).resolve().parent
+DATA_DIR = Path(os.environ.get("REME_EVAL_DATA_DIR", SUITE_DIR / "data")).resolve()
+OUTPUTS_DIR = Path(os.environ.get("REME_EVAL_OUTPUTS_DIR", SUITE_DIR / "outputs")) / "reme"
+WORKSPACE_ROOT = Path(
+ os.environ.get("REME_WORKSPACE_ROOT", SUITE_DIR / "reme_workspace"),
+).resolve()
+
+COMPLETED_STATUSES = {"SUCCESS", "MAX_TURNS", "TIMEOUT"}
+TASK_FINISHED_RE = re.compile(r"Task finished task_id=(\S+) status=(\S+)")
+SESSION_ID_RE = re.compile(r"^session_id:\s*(\S+)", re.MULTILINE)
+NOTE_COUNT_RE = re.compile(r"(description:\s*)\d+(\s*note\(s\) today)")
+LOG_FILE_TS_RE = re.compile(r"^(\d{8}_\d{6})-log\.jsonl$")
+TIME_FORMAT = "%Y%m%d_%H%M%S"
+
+
+def log(msg: str) -> None:
+ """Print a status message to stderr."""
+ print(msg, file=sys.stderr)
+
+
+def episode_task_order(persona: str) -> list[str]:
+ """Return the ordered task ids from the persona's episode.yaml."""
+ episode_path = DATA_DIR / persona / "episode.yaml"
+ with open(episode_path, "r", encoding="utf-8") as f:
+ episode = yaml.safe_load(f)
+ return [task["task_id"] for task in episode.get("tasks", [])]
+
+
+def _event_time(record: dict, file_ts: str) -> float:
+ """Best-effort event time (epoch seconds) of one log record.
+
+ Prefers the record's own timestamp fields; falls back to the timestamp
+ embedded in the log file name so that even stripped records keep a
+ meaningful order. Returns 0.0 when nothing is parseable.
+ """
+ timestamp = record.get("timestamp")
+ if isinstance(timestamp, (int, float)) and not isinstance(timestamp, bool):
+ return float(timestamp)
+ iso = record.get("timestamp_iso")
+ if isinstance(iso, str):
+ try:
+ return datetime.fromisoformat(iso).timestamp()
+ except ValueError:
+ pass
+ if file_ts:
+ try:
+ return datetime.strptime(file_ts, TIME_FORMAT).timestamp()
+ except ValueError:
+ pass
+ return 0.0
+
+
+def latest_task_statuses(persona: str) -> dict[str, str]:
+ """Scan per-task and run-level logs; the newest EVENT TIME wins per task.
+
+ Every "Task finished" record across both log categories is keyed by
+ (event_time, file timestamp, file order, line number); the record with
+ the highest key decides the task's status. File category and read order
+ alone can never override a newer record from the other category.
+ """
+ persona_dir = OUTPUTS_DIR / persona
+ if not persona_dir.is_dir():
+ return {}
+
+ log_files = sorted(persona_dir.glob("*/history/*-log.jsonl"))
+ log_files += sorted(persona_dir.glob("run/*-log.jsonl"))
+
+ best: dict[str, tuple[tuple, str]] = {}
+ for file_order, log_file in enumerate(log_files):
+ ts_match = LOG_FILE_TS_RE.match(log_file.name)
+ file_ts = ts_match.group(1) if ts_match else ""
+ try:
+ with open(log_file, "r", encoding="utf-8") as f:
+ for line_no, line in enumerate(f):
+ if "Task finished" not in line:
+ continue
+ try:
+ record = json.loads(line)
+ except json.JSONDecodeError:
+ continue
+ match = TASK_FINISHED_RE.search(str(record.get("message", "")))
+ if not match:
+ continue
+ task_id, status = match.group(1), match.group(2)
+ sort_key = (_event_time(record, file_ts), file_ts, file_order, line_no)
+ current = best.get(task_id)
+ if current is None or sort_key > current[0]:
+ best[task_id] = (sort_key, status)
+ except OSError:
+ continue
+ return {task_id: status for task_id, (_, status) in best.items()}
+
+
+def split_tasks(persona: str) -> tuple[list[str], list[str]]:
+ """Split the episode task order into completed and remaining tasks."""
+ order = episode_task_order(persona)
+ statuses = latest_task_statuses(persona)
+ completed = [t for t in order if statuses.get(t) in COMPLETED_STATUSES]
+ remaining = [t for t in order if t not in set(completed)]
+ return completed, remaining
+
+
+def _daily_note_session_id(note_path: Path) -> str:
+ try:
+ text = note_path.read_text(encoding="utf-8")
+ except OSError:
+ return ""
+ match = SESSION_ID_RE.search(text)
+ return match.group(1) if match else ""
+
+
+class _WorkspaceFileStoreShim:
+ """Structural stand-in for ReMe's file store; only workspace_path is read."""
+
+ def __init__(self, workspace_path: Path):
+ self.workspace_path = workspace_path
+
+
+def _refresh_daily_indexes(
+ workspace: Path,
+ removed_by_date: dict[str, set[str]],
+ removed: list[str],
+) -> None:
+ """Rebuild the daily index of each affected date via ReMe's own logic."""
+ for date in sorted(removed_by_date):
+ result = asyncio.run(
+ refresh_day_index(_WorkspaceFileStoreShim(workspace), date, "daily"),
+ )
+ if result.get("error"):
+ log(f"[resume] WARNING: daily index refresh failed for {date}: {result['error']}")
+ continue
+ removed.append(f"daily/{date}.md (refreshed, {len(removed_by_date[date])} note(s) removed)")
+
+
+def _strip_index_lines(
+ workspace: Path,
+ removed_by_date: dict[str, set[str]],
+ removed: list[str],
+ dry_run: bool,
+) -> None:
+ """Fallback index edit: drop lines that reference removed notes by full
+ workspace-relative wikilink path, and fix the note count. Only the index
+ files of affected dates are touched."""
+ for date in sorted(removed_by_date):
+ index_path = workspace / "daily" / f"{date}.md"
+ if not index_path.is_file():
+ continue
+ wikilinks = [f"[[{rel_path}]]" for rel_path in sorted(removed_by_date[date])]
+ lines = index_path.read_text(encoding="utf-8").splitlines()
+ kept = [line for line in lines if not any(link in line for link in wikilinks)]
+ if len(kept) == len(lines):
+ continue
+ note_count = sum(1 for line in kept if line.startswith("- [[daily/"))
+ kept = [NOTE_COUNT_RE.sub(rf"\g<1>{note_count}\2", line) for line in kept]
+ removed.append(f"{index_path.relative_to(workspace)} (rewritten)")
+ if not dry_run:
+ index_path.write_text("\n".join(kept) + "\n", encoding="utf-8")
+
+
+def cleanup_partial_memory(persona: str, remaining: list[str], dry_run: bool = False) -> list[str]:
+ """Remove partial memory artifacts of remaining tasks so they can be re-run cleanly."""
+ workspace = WORKSPACE_ROOT / persona
+ removed: list[str] = []
+ if not workspace.is_dir() or not remaining:
+ return removed
+
+ prefixes = tuple(f"pibench_{task_id}_" for task_id in remaining)
+
+ def act(path: Path, label: str) -> None:
+ removed.append(label)
+ if not dry_run:
+ path.unlink()
+
+ # 1) daily / digest notes distilled from interrupted sessions. For daily
+ # notes, remember the full workspace-relative path grouped by date so only
+ # the affected daily indexes are refreshed below.
+ removed_by_date: dict[str, set[str]] = {}
+ for section in ("daily", "digest"):
+ section_root = workspace / section
+ if not section_root.is_dir():
+ continue
+ for note_path in section_root.rglob("*.md"):
+ if note_path.parent == section_root:
+ continue # index files handled below
+ session_id = _daily_note_session_id(note_path)
+ if session_id.startswith(prefixes):
+ rel_path = note_path.relative_to(workspace).as_posix()
+ act(note_path, rel_path)
+ if section == "daily":
+ removed_by_date.setdefault(note_path.parent.name, set()).add(rel_path)
+
+ # 2) daily index files: refresh only the dates that lost notes, matching
+ # notes by their full wikilink path instead of their bare file name.
+ if removed_by_date:
+ if dry_run:
+ for date in sorted(removed_by_date):
+ removed.append(f"daily/{date}.md (would refresh index)")
+ elif refresh_day_index is not None:
+ _refresh_daily_indexes(workspace, removed_by_date, removed)
+ else:
+ _strip_index_lines(workspace, removed_by_date, removed, dry_run)
+
+ # 3) raw dialog logs of interrupted sessions
+ dialog_dir = workspace / "session" / "dialog"
+ if dialog_dir.is_dir():
+ for task_id in remaining:
+ for dialog_path in dialog_dir.glob(f"pibench_{task_id}_*.jsonl"):
+ act(dialog_path, str(dialog_path.relative_to(workspace)))
+
+ # 4) agent-scope session states that contain interrupted-task sessions
+ mem_session_dir = workspace / "mem_session"
+ if mem_session_dir.is_dir():
+ for session_path in mem_session_dir.rglob("*.jsonl"):
+ try:
+ content = session_path.read_text(encoding="utf-8", errors="ignore")
+ except OSError:
+ continue
+ if any(prefix in content for prefix in prefixes):
+ act(session_path, str(session_path.relative_to(workspace)))
+
+ return removed
+
+
+def main() -> int:
+ """CLI entrypoint: run 'remaining' or 'cleanup' action for a persona."""
+ args = sys.argv[1:]
+ if len(args) < 2 or args[0] not in {"remaining", "cleanup"}:
+ print(__doc__, file=sys.stderr)
+ return 2
+
+ command, persona = args[0], args[1]
+ completed, remaining = split_tasks(persona)
+
+ if command == "remaining":
+ if "--json" in args:
+ print(json.dumps({"completed": completed, "remaining": remaining}))
+ else:
+ for task_id in remaining:
+ print(task_id)
+ log(
+ f"[resume] {persona}: completed={len(completed)} "
+ f"({', '.join(completed) if completed else '-'}) remaining={len(remaining)}",
+ )
+ return 0
+
+ dry_run = "--dry-run" in args
+ removed = cleanup_partial_memory(persona, remaining, dry_run=dry_run)
+ if removed:
+ verb = "would remove" if dry_run else "removed"
+ log(f"[resume] {persona}: {verb} {len(removed)} partial-memory artifact(s):")
+ for item in removed:
+ log(f" - {item}")
+ else:
+ log(f"[resume] {persona}: no partial-memory artifacts to clean")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/benchmark/pibench/run_all.sh b/benchmark/pibench/run_all.sh
new file mode 100755
index 00000000..e6ecd3b1
--- /dev/null
+++ b/benchmark/pibench/run_all.sh
@@ -0,0 +1,119 @@
+#!/bin/bash
+# Run all 5 personas with the ReMe agent, PARALLEL at a time (default 2).
+# Each persona's tasks follow data/{persona}/episode.yaml order.
+#
+# Usage:
+# bash run_all.sh # FRESH official run: wipes ALL personas'
+# # ReMe memory/outputs/trace logs first,
+# # then runs everything from scratch.
+# bash run_all.sh --resume # Checkpoint continuation: no wipe; every
+# # persona skips already-completed tasks.
+# bash run_all.sh --parallel 1 # sequential (original behavior)
+# bash run_all.sh --skip-eval # run phase only
+#
+# Memory-wipe vs resume conflict resolution:
+# The full ReMe memory wipe happens ONLY here, ONLY in fresh mode (the
+# default), and ONLY before any service/bridge starts. --resume never
+# wipes; run_persona.sh then additionally performs a surgical cleanup of
+# residual memory belonging to interrupted (to-be-re-run) tasks, so a
+# resumed run keeps all completed-task memory but never inherits a partial
+# task's own answer. The two modes are mutually exclusive.
+set -uo pipefail
+
+SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
+PERSONAS=(researcher marketer law_trainee pharmacist Financier)
+TRACE_ROOT="${HOME}/.nanobot/trace_logs"
+
+PARALLEL=2
+MODE="fresh"
+PASS_ARGS=()
+while [[ $# -gt 0 ]]; do
+ case $1 in
+ --parallel)
+ PARALLEL="${2:-}"; shift 2 || true
+ case "$PARALLEL" in (""|*[!0-9]*) echo "--parallel needs a positive integer"; exit 2 ;; esac
+ [ "$PARALLEL" -lt 1 ] && PARALLEL=1
+ [ "$PARALLEL" -gt ${#PERSONAS[@]} ] && PARALLEL=${#PERSONAS[@]}
+ ;;
+ --resume)
+ if [ "$MODE" = "fresh_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
+ MODE="resume"; shift ;;
+ --fresh)
+ if [ "$MODE" = "resume" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
+ MODE="fresh_set"; shift ;;
+ --skip-eval) PASS_ARGS+=(--skip-eval); shift ;;
+ *) echo "Unknown option: $1"; exit 1 ;;
+ esac
+done
+[ "$MODE" = "fresh_set" ] && MODE="fresh"
+
+START_TS=$(date +%Y%m%d_%H%M%S)
+SUMMARY_LOG="${SUITE_DIR}/logs/run_all_${START_TS}.summary"
+mkdir -p "${SUITE_DIR}/logs"
+
+echo "############################################################"
+echo "# reme_eval suite | mode=${MODE} parallel=${PARALLEL} | ${START_TS}"
+echo "############################################################"
+
+# ─── Fresh mode: suite-level wipe BEFORE anything starts ──────────────
+if [ "$MODE" = "fresh" ]; then
+ echo "[fresh] wiping ALL personas' memory workspaces, outputs and trace logs..."
+ for persona in "${PERSONAS[@]}"; do
+ rm -rf "${SUITE_DIR}/reme_workspace/${persona}"
+ rm -rf "${SUITE_DIR}/outputs/reme/${persona}"
+ rm -rf "${TRACE_ROOT}/reme/${persona}"
+ rm -rf "${SUITE_DIR}/nanobot_workspace/${persona}"
+ done
+ echo "[fresh] wipe done."
+else
+ echo "[resume] no memory wipe; personas resume after their last completed task."
+fi
+
+# ─── Run personas in batches of PARALLEL ──────────────────────────────
+STATUS_LIST=()
+ANY_FAILED=0
+OVERALL_START=$(date +%s)
+TOTAL=${#PERSONAS[@]}
+
+for ((i = 0; i < TOTAL; i += PARALLEL)); do
+ BATCH=("${PERSONAS[@]:i:PARALLEL}")
+ BATCH_PIDS=()
+ BATCH_NAMES=()
+ echo ""
+ echo "============================================================"
+ echo "# BATCH $(( i / PARALLEL + 1 )): ${BATCH[*]} started $(date '+%F %T')"
+ echo "============================================================"
+ for persona in "${BATCH[@]}"; do
+ bash "${SUITE_DIR}/run_persona.sh" "${persona}" --resume ${PASS_ARGS[@]+"${PASS_ARGS[@]}"} \
+ > "${SUITE_DIR}/logs/suite_${persona}.log" 2>&1 &
+ BATCH_PIDS+=($!)
+ BATCH_NAMES+=("$persona")
+ done
+ for j in $(seq 0 $(( ${#BATCH[@]} - 1 ))); do
+ pid=${BATCH_PIDS[$j]}
+ persona=${BATCH_NAMES[$j]}
+ if wait "$pid"; then
+ STATUS_LIST+=("${persona}: OK")
+ else
+ rc=$?
+ ANY_FAILED=1
+ STATUS_LIST+=("${persona}: FAILED rc=${rc}")
+ echo "[run_all] ${persona} FAILED (rc=${rc}); see logs/suite_${persona}.log"
+ fi
+ done
+done
+
+total=$(( $(date +%s) - OVERALL_START ))
+echo ""
+echo "================ FINAL SUMMARY (${total}s total) ================" | tee -a "${SUMMARY_LOG}"
+for line in "${STATUS_LIST[@]}"; do
+ echo " ${line}" | tee -a "${SUMMARY_LOG}"
+done
+echo "Summary: ${SUMMARY_LOG}"
+
+if [ "${ANY_FAILED}" -ne 0 ]; then
+ FAILED_COUNT=$(printf '%s\n' "${STATUS_LIST[@]}" | grep -c "FAILED")
+ echo "[run_all] ${FAILED_COUNT} persona(s) FAILED; suite run is marked as failed." | tee -a "${SUMMARY_LOG}"
+ exit 1
+fi
+exit 0
diff --git a/benchmark/pibench/run_persona.sh b/benchmark/pibench/run_persona.sh
new file mode 100755
index 00000000..78f7d5ca
--- /dev/null
+++ b/benchmark/pibench/run_persona.sh
@@ -0,0 +1,301 @@
+#!/bin/bash
+# Run the full pi-bench evaluation for ONE persona with the ReMe agent.
+# Tasks follow data/{persona}/episode.yaml order (runner-native).
+#
+# Usage: bash run_persona.sh [--fresh|--resume] [--skip-eval]
+#
+# Modes (default: --resume):
+# --resume Checkpoint continuation. Never wipes memory. Tasks already
+# finished (SUCCESS/MAX_TURNS/TIMEOUT in the task history logs)
+# are skipped via repeated --task-id flags. Before starting, any
+# residual memory of tasks that are about to be RE-RUN (partial
+# sessions from an interrupted run) is surgically removed by
+# resume.py cleanup, so re-runs don't inherit leaked answers.
+# --fresh Wipes THIS persona's ReMe memory, outputs and trace logs first,
+# then runs all tasks from scratch.
+# The two flags are mutually exclusive. A full multi-persona memory wipe is a
+# suite-level action of `run_all.sh` (fresh mode), never done here implicitly.
+set -uo pipefail
+
+SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
+TRACE_ROOT="${HOME}/.nanobot/trace_logs"
+
+# ─── External dependencies (pi-bench / ReMe are NOT bundled; see README) ──
+if [ ! -f "${SUITE_DIR}/env.sh" ]; then
+ echo "env.sh not found. Run: cp env.sh.example env.sh (then fill in the TODO items)"
+ exit 1
+fi
+source "${SUITE_DIR}/env.sh"
+
+PIBENCH_DIR="${PI_BENCH_ROOT:-}"
+if [ -z "${PIBENCH_DIR}" ] || [ ! -f "${PIBENCH_DIR}/src/main.py" ]; then
+ echo "PI_BENCH_ROOT is unset or invalid (src/main.py not found). Set it in env.sh."
+ exit 1
+fi
+if [ ! -x "${PIBENCH_DIR}/.venv/bin/python" ] || [ ! -x "${PIBENCH_DIR}/.venv/bin/appworld" ]; then
+ echo "pi-bench venv incomplete: ${PIBENCH_DIR}/.venv must provide python + appworld (see README setup)."
+ exit 1
+fi
+if [ ! -x "${REME_DIR}/.venv/bin/python" ]; then
+ echo "ReMe venv not found: ${REME_DIR}/.venv/bin/python (check REME_DIR in env.sh)"
+ exit 1
+fi
+if [ ! -e "${SUITE_DIR}/data" ]; then
+ echo 'Benchmark data not linked. Run: ln -s "$PI_BENCH_ROOT/data" data'
+ exit 1
+fi
+
+# ─── Pre-flight: files the runner needs before any service starts ─────
+MODEL_CONFIG="${SUITE_DIR}/config/models/reme.yaml"
+HISTORY_CONFIG="${SUITE_DIR}/config/bench/evaluation/trace_history.yaml"
+if [ ! -f "${MODEL_CONFIG}" ]; then
+ echo "Model config not found: ${MODEL_CONFIG} (see README directory layout)."
+ exit 1
+fi
+if [ ! -f "${HISTORY_CONFIG}" ]; then
+ echo "Trace history config not found: ${HISTORY_CONFIG}"
+ echo "pi-bench requires config/bench/evaluation/trace_history.yaml; see README."
+ exit 1
+fi
+
+APPWORLD_DIR="${PIBENCH_DIR}/third_party/appworld"
+PI_PYTHON="${PIBENCH_DIR}/.venv/bin/python"
+APPWORLD_BIN="${PIBENCH_DIR}/.venv/bin/appworld"
+# resume.py runs on the ReMe venv so it can reuse ReMe's daily-index rebuild.
+REME_PYTHON="${REME_DIR}/.venv/bin/python"
+
+PERSONA="${1:-}"
+if [ -z "$PERSONA" ]; then
+ echo "Usage: $0 [--fresh|--resume] [--skip-eval]"
+ exit 1
+fi
+shift
+
+MODE="resume"
+SKIP_EVAL=false
+while [[ $# -gt 0 ]]; do
+ case $1 in
+ --fresh)
+ if [ "$MODE" = "resume_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
+ MODE="fresh"; shift ;;
+ --resume)
+ if [ "$MODE" = "fresh" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
+ MODE="resume_set"; shift ;;
+ --skip-eval) SKIP_EVAL=true; shift ;;
+ *) echo "Unknown option: $1"; exit 1 ;;
+ esac
+done
+[ "$MODE" = "resume_set" ] && MODE="resume"
+
+# ─── Per-persona ports (pi-bench AGENTS.md convention) ────────────────
+# REME_PORT: ReMe's internal HTTP service; must be unique per concurrent bridge.
+case "$PERSONA" in
+ marketer) API_PORT=9001; MCP_PORT=10001; TEST_PORT=9998; REME_PORT=18766 ;;
+ law_trainee) API_PORT=9002; MCP_PORT=10002; TEST_PORT=9997; REME_PORT=18767 ;;
+ pharmacist) API_PORT=9003; MCP_PORT=10003; TEST_PORT=9996; REME_PORT=18768 ;;
+ researcher) API_PORT=9004; MCP_PORT=10004; TEST_PORT=9995; REME_PORT=18765 ;;
+ Financier) API_PORT=9005; MCP_PORT=10005; TEST_PORT=9994; REME_PORT=18769 ;;
+ *) echo "Unknown persona: $PERSONA"; exit 1 ;;
+esac
+
+API_URL="http://127.0.0.1:${API_PORT}"
+MCP_URL="http://127.0.0.1:${MCP_PORT}/mcp"
+TEST_URL="http://127.0.0.1:${TEST_PORT}"
+LOG_DIR="${SUITE_DIR}/logs"
+mkdir -p "${LOG_DIR}"
+
+# ─── Environment (env.sh already sourced at the top) ──────────────────
+WORKSPACE_DIR="${REME_WORKSPACE_ROOT}/${PERSONA}"
+NANOBOT_WORKSPACE_DIR="${SUITE_DIR}/nanobot_workspace/${PERSONA}"
+mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
+
+echo "========================================="
+echo "ReMe x Pi-Bench | persona=${PERSONA} | mode=${MODE}"
+echo " api=${API_PORT} mcp=${MCP_PORT} test=${TEST_PORT} reme=${REME_PORT}"
+echo " model=${REME_MODEL_NAME}"
+echo " memory workspace=${WORKSPACE_DIR} (persistent)"
+echo "========================================="
+
+# ─── Fresh mode: wipe this persona's state ────────────────────────────
+if [ "$MODE" = "fresh" ]; then
+ echo "[fresh] wiping persona state: memory workspace, outputs, trace logs"
+ rm -rf "${WORKSPACE_DIR}"
+ rm -rf "${SUITE_DIR}/outputs/reme/${PERSONA}"
+ rm -rf "${TRACE_ROOT}/reme/${PERSONA}"
+ rm -rf "${NANOBOT_WORKSPACE_DIR}"
+ mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
+fi
+
+# ─── Resume: determine remaining tasks + clean partial memories ───────
+TASK_ARGS=()
+RUN_PHASE_NEEDED=true
+if [ "$MODE" = "resume" ]; then
+ REMAINING_JSON="$("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" --json)"
+ if [ -z "$REMAINING_JSON" ]; then
+ echo "Failed to compute remaining tasks"; exit 1
+ fi
+ echo "[resume] ${REMAINING_JSON}"
+ REMAINING_TASKS=()
+ while IFS= read -r tid_line; do
+ [ -n "$tid_line" ] && REMAINING_TASKS+=("$tid_line")
+ done < <("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" 2>/dev/null)
+ if [ ${#REMAINING_TASKS[@]} -eq 0 ]; then
+ RUN_PHASE_NEEDED=false
+ echo "[resume] all tasks already completed; skipping run phase"
+ else
+ # Remove residual memory of interrupted (to-be-re-run) tasks so
+ # re-runs don't get their own partial answers injected.
+ "${REME_PYTHON}" "${SUITE_DIR}/resume.py" cleanup "${PERSONA}"
+ for tid in "${REMAINING_TASKS[@]}"; do
+ TASK_ARGS+=(--task-id "$tid")
+ done
+ echo "[resume] running ${#REMAINING_TASKS[@]} remaining task(s): ${REMAINING_TASKS[*]}"
+ fi
+fi
+
+# ─── Port cleanup from previous runs ──────────────────────────────────
+for port in ${API_PORT} ${MCP_PORT} ${TEST_PORT} ${REME_PORT}; do
+ pids=$(lsof -ti :${port} 2>/dev/null || true)
+ if [ -n "$pids" ]; then
+ echo "Killing stale processes on port ${port}: ${pids}"
+ kill -9 $pids 2>/dev/null || true
+ fi
+done
+sleep 2
+
+PIDS=()
+cleanup() {
+ echo "[${PERSONA}] cleaning up services..."
+ for pid in "${PIDS[@]:-}"; do
+ kill "$pid" 2>/dev/null || true
+ done
+ wait 2>/dev/null || true
+}
+trap cleanup EXIT INT TERM
+
+wait_for_service() {
+ local url="$1" name="$2" port="$3" timeout="${4:-180}"
+ echo -n " waiting for ${name}..."
+ local start=$(date +%s)
+ while true; do
+ if curl -sf --max-time 5 "${url}" > /dev/null 2>&1; then
+ echo " ready"; return 0
+ fi
+ if [ -n "$port" ] && lsof -ti :${port} > /dev/null 2>&1; then
+ local elapsed=$(( $(date +%s) - start ))
+ if [ "$elapsed" -ge 10 ]; then echo " ready (port)"; return 0; fi
+ fi
+ if [ $(( $(date +%s) - start )) -ge "$timeout" ]; then
+ echo " TIMEOUT"; return 1
+ fi
+ sleep 2
+ done
+}
+
+# ─── [1/5] AppWorld API ────────────────────────────────────────────────
+echo "[1/5] AppWorld API (:${API_PORT})"
+(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve apis --root . \
+ --port ${API_PORT}) > "${LOG_DIR}/appworld_api_${PERSONA}.log" 2>&1 &
+PIDS+=($!)
+if ! wait_for_service "${API_URL}/docs" "AppWorld API" "${API_PORT}" 180; then
+ tail -20 "${LOG_DIR}/appworld_api_${PERSONA}.log"; exit 1
+fi
+
+# ─── [2/5] AppWorld MCP ────────────────────────────────────────────────
+echo "[2/5] AppWorld MCP (:${MCP_PORT})"
+TOOLS_CONFIG="${SUITE_DIR}/data/${PERSONA}/tools.yaml"
+(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve mcp http --root . \
+ --remote-apis-url "${API_URL}" --port ${MCP_PORT} \
+ --tools-config-file "${TOOLS_CONFIG}") > "${LOG_DIR}/appworld_mcp_${PERSONA}.log" 2>&1 &
+PIDS+=($!)
+if ! wait_for_service "${MCP_URL}" "AppWorld MCP" "${MCP_PORT}" 180; then
+ tail -20 "${LOG_DIR}/appworld_mcp_${PERSONA}.log"; exit 1
+fi
+
+# ─── [3/5] Test Server ─────────────────────────────────────────────────
+echo "[3/5] Test Server (:${TEST_PORT})"
+PORT=${TEST_PORT} "${PI_PYTHON}" "${PIBENCH_DIR}/scripts/test_server.py" \
+ > "${LOG_DIR}/test_server_${PERSONA}.log" 2>&1 &
+PIDS+=($!)
+if ! wait_for_service "${TEST_URL}/sent?after=-1" "Test Server" "${TEST_PORT}" 30; then
+ tail -20 "${LOG_DIR}/test_server_${PERSONA}.log"; exit 1
+fi
+
+# ─── [4/5] ReMe Bridge (ReMe venv) ─────────────────────────────────────
+echo "[4/5] ReMe Bridge (reme service port ${REME_PORT})"
+"${REME_DIR}/.venv/bin/python" "${SUITE_DIR}/bridge_reme.py" \
+ --test-server-url "${TEST_URL}" \
+ --appworld-mcp-url "${MCP_URL}" \
+ --reme-dir "${REME_DIR}" \
+ --data-root "${SUITE_DIR}/data" \
+ --user-id "${PERSONA}" \
+ --workspace-dir "${WORKSPACE_DIR}" \
+ --reme-port "${REME_PORT}" \
+ --model-name "${REME_MODEL_NAME}" \
+ --model-base-url "${REME_LLM_BASE_URL}" \
+ --model-api-key "${REME_LLM_API_KEY}" \
+ > "${LOG_DIR}/bridge_${PERSONA}.log" 2>&1 &
+BRIDGE_PID=$!
+PIDS+=(${BRIDGE_PID})
+sleep 5
+if ! kill -0 "${BRIDGE_PID}" 2>/dev/null; then
+ echo "Bridge failed to start:"; tail -30 "${LOG_DIR}/bridge_${PERSONA}.log"; exit 1
+fi
+for i in $(seq 1 12); do
+ if grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null; then
+ echo " bridge initialized"; break
+ fi
+ sleep 5
+done
+grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null || {
+ echo "WARNING: bridge may not be ready:"; tail -20 "${LOG_DIR}/bridge_${PERSONA}.log"; }
+
+# ─── [5/5] Runner (run phase) ──────────────────────────────────────────
+if [ "$RUN_PHASE_NEEDED" = true ]; then
+ echo "[5/5] Runner: run phase (episode order from data/${PERSONA}/episode.yaml)"
+ cd "${SUITE_DIR}"
+ BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
+ "${PI_PYTHON}" -m src.main \
+ --model-config "${MODEL_CONFIG}" \
+ --history-config-path "${HISTORY_CONFIG}" \
+ --mode run --user-id "${PERSONA}" \
+ --workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
+ ${TASK_ARGS[@]+"${TASK_ARGS[@]}"} \
+ 2>&1 | tee "${LOG_DIR}/runner_run_${PERSONA}.log"
+ RUN_EXIT=${PIPESTATUS[0]}
+ if [ ${RUN_EXIT} -ne 0 ]; then
+ echo "Run phase failed (exit ${RUN_EXIT}). Logs: ${LOG_DIR}/"
+ exit ${RUN_EXIT}
+ fi
+else
+ echo "[5/5] Runner: run phase skipped (all tasks completed)"
+fi
+
+if [ "$SKIP_EVAL" = true ]; then
+ echo "Skipping eval (--skip-eval)"
+ exit 0
+fi
+
+# ─── Trace conversion + eval phase (always over all available traces) ──
+echo "Converting trace logs..."
+"${PI_PYTHON}" "${SUITE_DIR}/fix_trace_logs.py" "${PERSONA}"
+
+echo "Runner: eval phase"
+cd "${SUITE_DIR}"
+BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
+ "${PI_PYTHON}" -m src.main \
+ --model-config "${MODEL_CONFIG}" \
+ --history-config-path "${HISTORY_CONFIG}" \
+ --mode eval --user-id "${PERSONA}" \
+ --workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
+ 2>&1 | tee "${LOG_DIR}/runner_eval_${PERSONA}.log"
+EVAL_EXIT=${PIPESTATUS[0]}
+
+echo ""
+echo "========================================="
+echo "persona=${PERSONA} finished (eval exit=${EVAL_EXIT})"
+echo " results : ${SUITE_DIR}/outputs/reme/${PERSONA}/"
+echo " memory : ${WORKSPACE_DIR}/"
+echo " logs : ${LOG_DIR}/"
+echo "========================================="
+exit ${EVAL_EXIT}
diff --git a/cookbook/auto-fin/README.md b/cookbook/auto-fin/README.md
deleted file mode 100644
index cf92ccc9..00000000
--- a/cookbook/auto-fin/README.md
+++ /dev/null
@@ -1,215 +0,0 @@
-# Auto Fin Cookbook
-
-[中文](README_ZH.md)
-
-Auto Fin is a local-first, file-native ETF event-research workflow. It collects CLS news and market data through
-Tushare, identifies current news related to a configured ETF list, retrieves comparable events from local ReMe memory,
-calculates observed post-event returns, and writes a Chinese research report.
-
-> Auto Fin is for event research and holding-period reference only. It is not investment advice, does not connect to a
-> broker, and does not place or simulate trades.
-
-The workflow is assembled by
-[`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml). Its public schemas are in
-[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py), and its four steps are in
-[`reme/steps/cookbook/auto_fin/`](../../reme/steps/cookbook/auto_fin/).
-
-## Quick start
-
-Auto Fin requires Python 3.11 or newer, the `core` dependencies, a Tushare token, and an available AgentScope LLM.
-
-```bash
-python -m pip install -e ".[core]"
-export TUSHARE_TOKEN="your-tushare-token"
-export LLM_API_KEY="your-api-key"
-reme start config=daily_cookbook job=auto_fin
-```
-
-The built-in LLM component defaults to `qwen3.7-plus`. `LLM_BASE_URL` has no built-in value, so set it when your
-provider requires a custom OpenAI-compatible endpoint. The model and endpoint can be overridden with
-`LLM_MODEL_NAME` and `LLM_BASE_URL`.
-
-The default workspace is `reme_workspace/` beneath the process working directory. Override it with
-`DAILY_PAPER_WORKSPACE_DIR`; Auto Fin and Daily Paper share this setting.
-
-Dates and times use `Asia/Shanghai`. An explicit `date` must be today's date:
-
-```bash
-reme start config=daily_cookbook job=auto_fin date=2026-08-07
-```
-
-Auto Fin checks the SSE trading calendar first and skips the whole workflow on a closed market day.
-
-## Pipeline
-
-```text
-Tushare trade calendar
- │
- ├─ closed day ──► skip
- ▼
-Collect CLS news + configured ETF history
- ▼
-Update the ReMe index
- ▼
-Select ETF/news relationships with an agent
- ▼
-Search local memory for comparable historical news
- ▼
-Select same/opposite events with an agent + calculate D1/D2/D3/D5 returns in code
- ▼
-Generate report with an agent ──► refresh day index ──► DingTalk (optional)
-```
-
-| Step | Responsibility | Agent |
-|---|---|---|
-| `auto_fin_data_step` | Check the trading day, maintain news, and cache configured ETF market history | No |
-| `auto_fin_topic_step` | Select direct relationships between today's news and configured ETFs | Yes |
-| `auto_fin_history_step` | Retrieve comparable news, validate selections, and calculate observed returns | Yes |
-| `auto_fin_merge_step` | Combine prepared evidence and the previous report into the final Markdown | Yes |
-
-All three model-facing steps use structured Pydantic output. Agents make semantic judgments; code owns identifier
-validation, source resolution, market calculations, and file writes.
-
-## Data and selection boundaries
-
-### News
-
-`auto_fin_data_step` calls Tushare `major_news` with `src="财联社"`. The default lookback is 60 calendar days including
-today. Existing files for earlier days are reused, while today's file is always overwritten with news from 00:00
-through the current decision time. Large responses are recursively split when a request returns at least 400 rows.
-
-Each item is stored in `daily/YYYY-MM-DD/auto_fin_news.md` with a stable ID made from its publication timestamp and a
-short content hash. The current-event set used by the Topic step is today's complete file, not an increment since an
-earlier run.
-
-### Configured ETFs
-
-The built-in configuration currently enables:
-
-- `518880.SH`
-- `159530.SZ`
-- `512760.SH`
-
-Other examples remain commented out in `daily_cookbook.yaml`. For each enabled code, the Data step resolves its name
-through `etf_basic`, then pages backward through `fund_daily` and `fund_adj` and rewrites its complete local JSONL
-history. A missing ETF name fails the run.
-
-The Topic agent receives only the configured ETF code/name pairs and today's locally stored news. It may retain up to
-`current_news_limit_per_etf` valid, unique news references per ETF (10 by default). Unknown ETF codes, unknown news IDs,
-empty reasons, duplicates, and ETFs with no accepted event are removed by code.
-
-## Historical comparison and returns
-
-For every accepted current ETF/news pair, `auto_fin_history_step` calls the configured `memory_search` job over the
-60-day news window, ending yesterday. `historical_search_limit` controls the maximum search results requested per
-current event. Only search hits whose path is named `auto_fin_news.md` contribute candidate IDs; the step rereads the
-source Markdown and resolves those IDs before calling the History agent.
-
-The History agent may select at most five candidates by default and labels each relationship `same` or `opposite`.
-Code discards unknown or duplicate IDs and empty reasons, then calculates adjusted cumulative returns for D1, D2, D3,
-and D5:
-
-- For an event before 15:00 on a trading day, the adjusted same-day close is the entry; D1 is the next trading close.
-- For an event at or after 15:00, the adjusted next-trading-day open is the entry; D1 is that day's close.
-- If an entry or horizon cannot be calculated from valid positive prices and adjustment factors, that value is `null`.
-
-The final agent receives the fixed ETF list, all current and historical evidence, `same`/`opposite` directions, computed
-returns, and the most recent earlier `auto_fin.md`. It decides whether the evidence supports a recommendation or an
-explicit wait-and-see conclusion; the code does not calculate a score, expected return, or mandatory holding period.
-
-## Outputs
-
-```text
-reme_workspace/
-├── daily/
-│ ├── YYYY-MM-DD.md
-│ └── YYYY-MM-DD/
-│ ├── auto_fin_news.md
-│ └── auto_fin.md
-└── resource/
- ├── fin/
- │ ├── etfs.json
- │ ├── 518880.SH.jsonl
- │ └── .jsonl
- └── YYYY-MM-DD/
- ├── auto_fin_topic_output.json
- ├── auto_fin_history_001_output.json
- ├── ...
- ├── auto_fin_analysis.jsonl
- └── auto_fin_merge_output.json
-```
-
-The daily news and report are user-owned Markdown. `resource/fin/` contains the market cache used for deterministic
-return calculations. Date-scoped JSON/JSONL files preserve structured agent replies and prepared analyses. Writes use
-same-directory temporary files and atomic replacement; the day index is refreshed after the report is written.
-
-## Parameters and defaults
-
-Public job parameters:
-
-| Parameter | Default | Purpose |
-|---|---:|---|
-| `date` | `""` | Empty uses today in `Asia/Shanghai`; a value must be strict `YYYY-MM-DD` and equal today |
-| `historical_search_limit` | `10` | Maximum `memory_search` results requested for each current event; minimum 1 |
-
-Relevant job settings in `daily_cookbook.yaml`:
-
-| Setting | Default | Purpose |
-|---|---:|---|
-| `etf_codes` | three enabled codes above | Fixed ETF research universe |
-| `news_lookback_days` | `60` | Local news and historical-search window |
-| `current_news_limit_per_etf` | `10` | Maximum accepted current events per ETF |
-| `historical_news_limit` | `5` | Maximum comparable events retained per current event |
-
-There is no public `force` parameter. Earlier news files are reused, today's news and all configured ETF market files
-are refreshed, and same-day report/resource paths are overwritten on each successful run.
-
-## Environment and scheduling
-
-| Variable | Required | Purpose |
-|---|---|---|
-| `TUSHARE_TOKEN` | Yes | Trading calendar, CLS news, ETF metadata, prices, and adjustment factors |
-| `LLM_API_KEY` | Provider-dependent | Shared AgentScope LLM credentials; config defaults to an empty value |
-| `LLM_MODEL_NAME` | No | Defaults to `qwen3.7-plus` |
-| `LLM_BASE_URL` | Provider-dependent | OpenAI-compatible endpoint; no built-in default |
-| `TUSHARE_MIRROR_URL` | No | Replaces the Tushare SDK HTTP URL after trimming a trailing slash |
-| `DAILY_PAPER_WORKSPACE_DIR` | No | Shared standalone cookbook workspace |
-| `DINGTALK_*` | No | Optional DingTalk application, robot, and group settings |
-
-The optional mirror can be configured, for example, as:
-
-```bash
-export TUSHARE_MIRROR_URL="http://112.124.63.173:4000/tushare"
-```
-
-`auto_fin_0930_cron`, `auto_fin_1130_cron`, and `auto_fin_1800_cron` run every day at 09:30, 11:30, and 18:00 in
-`Asia/Shanghai`. The crons fire on weekends and holidays, but the Data step then skips the remaining workflow when
-Tushare reports that the date is not an SSE trading day. Same-day reruns refine the existing report.
-
-To send a completed report, configure `DINGTALK_APP_KEY`, `DINGTALK_APP_SECRET`, `DINGTALK_ROBOT_CODE`, and the
-comma-separated `DINGTALK_CONVERSATION_IDS`. With no conversation IDs, delivery is a no-op.
-
-## Agent and failure boundaries
-
-Auto Fin and Daily Paper share the tool-free `default` AgentScope wrapper. Built-in and configured job tools are not
-exposed to their model calls. Auto Fin itself invokes `memory_search` in deterministic step code; this is not an agent
-tool call. The separate interactive `dingtalk_wait` step has its own `bash` and ReMe job-tool allowlist.
-
-The standalone config has no embedding store enabled by default, so `memory_search` uses the available BM25 path;
-vector/BM25 fusion requires enabling the commented embedding components.
-
-Invalid dates, missing credentials or services, invalid structured model output, unknown configured ETFs, missing
-market files, and failed memory search stop the job. A market holiday is a successful skip. The workflow has no global
-same-date execution lock or cross-file transaction, and repeated successful runs can resend DingTalk notifications.
-
-## Tests
-
-Focused unit tests mock model and market-data boundaries:
-
-```bash
-python -m pip install -e ".[dev,core]"
-pytest tests/unit/test_auto_fin.py -v
-```
-
-Tests requiring real Tushare, LLM, or DingTalk credentials should be run separately and only with explicit
-authorization.
diff --git a/cookbook/auto-fin/README_ZH.md b/cookbook/auto-fin/README_ZH.md
deleted file mode 100644
index 38cf5f9e..00000000
--- a/cookbook/auto-fin/README_ZH.md
+++ /dev/null
@@ -1,201 +0,0 @@
-# Auto Fin Cookbook
-
-[English](README.md)
-
-Auto Fin 是一个 local-first、file-native 的 ETF 事件研究工作流。它通过 Tushare 获取财联社新闻和行情数据,从固定
-ETF 列表中识别与当日新闻相关的标的,利用 ReMe 本地记忆检索可比历史事件,计算事件后的实际收益,并生成中文研究报告。
-
-> Auto Fin 只提供事件研究和持有时间参考,不构成投资建议,不连接券商,也不会执行或模拟交易。
-
-工作流由 [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml) 装配;公开 schema 位于
-[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py),四个 Step 位于
-[`reme/steps/cookbook/auto_fin/`](../../reme/steps/cookbook/auto_fin/)。
-
-## 快速开始
-
-要求 Python 3.11 或更高版本、`core` 依赖、Tushare token 和可用的 AgentScope LLM。
-
-```bash
-python -m pip install -e ".[core]"
-export TUSHARE_TOKEN="your-tushare-token"
-export LLM_API_KEY="your-api-key"
-reme start config=daily_cookbook job=auto_fin
-```
-
-内置 LLM 组件默认使用 `qwen3.7-plus`。`LLM_BASE_URL` 没有内置默认值;如果服务商要求自定义 OpenAI 兼容
-endpoint,需要显式设置。可通过 `LLM_MODEL_NAME` 和 `LLM_BASE_URL` 覆盖模型与 endpoint。
-
-默认 workspace 是进程启动目录下的 `reme_workspace/`。可通过 `DAILY_PAPER_WORKSPACE_DIR` 覆盖;Auto Fin 与
-Daily Paper 共用该设置。
-
-日期和时间使用 `Asia/Shanghai`。显式传入的 `date` 必须是当天:
-
-```bash
-reme start config=daily_cookbook job=auto_fin date=2026-08-07
-```
-
-Auto Fin 首先检查上交所交易日历;休市日会跳过整个工作流。
-
-## 工作流
-
-```text
-Tushare 交易日历
- │
- ├─ 休市 ──► 跳过
- ▼
-采集财联社新闻 + 固定 ETF 行情历史
- ▼
-更新 ReMe 索引
- ▼
-Agent 筛选 ETF/当日新闻关系
- ▼
-从本地记忆检索可比历史新闻
- ▼
-Agent 选择 same/opposite 事件 + 代码计算 D1/D2/D3/D5 收益
- ▼
-Agent 生成报告 ──► 刷新当日索引 ──► 钉钉(可选)
-```
-
-| Step | 职责 | Agent |
-|---|---|---|
-| `auto_fin_data_step` | 检查交易日、维护新闻并缓存固定 ETF 的完整行情历史 | 否 |
-| `auto_fin_topic_step` | 筛选当日新闻与固定 ETF 的直接关系 | 是 |
-| `auto_fin_history_step` | 检索可比新闻、校验选择并计算实际收益 | 是 |
-| `auto_fin_merge_step` | 汇总证据和上一份报告,生成最终 Markdown | 是 |
-
-三个模型 Step 都使用 Pydantic 结构化输出。Agent 负责语义判断;标识校验、来源解析、行情计算和文件写入由代码负责。
-
-## 数据与筛选边界
-
-### 新闻
-
-`auto_fin_data_step` 调用 Tushare `major_news`,并固定传入 `src="财联社"`。默认回看 60 个自然日(包含当天)。
-更早日期已有的文件会复用;当天文件始终覆盖为 00:00 至当前决策时刻的新闻。单次请求返回至少 400 条时,时间区间会递归拆分。
-
-每条新闻写入 `daily/YYYY-MM-DD/auto_fin_news.md`,其稳定 ID 由发布时间和短内容哈希组成。Topic Step 使用当天
-完整文件,不是从上一次运行到本次运行之间的增量。
-
-### 固定 ETF
-
-内置配置当前启用:
-
-- `518880.SH`
-- `159530.SZ`
-- `512760.SH`
-
-`daily_cookbook.yaml` 中还保留了其他被注释的示例。Data Step 通过 `etf_basic` 解析每个启用代码的名称,然后对
-`fund_daily` 和 `fund_adj` 向前分页,并覆盖写入完整本地 JSONL 行情历史。任一 ETF 无法解析名称都会终止运行。
-
-Topic Agent 只接收固定 ETF 的 code/name 和当天本地新闻。每只 ETF 默认最多保留
-`current_news_limit_per_etf=10` 条有效且唯一的新闻引用。未知 ETF、未知 news ID、空理由和重复项会被代码移除;
-没有有效事件的 ETF 不进入后续步骤。
-
-## 历史比较与收益
-
-对每个有效的 ETF/当日新闻组合,`auto_fin_history_step` 会在 60 日新闻窗口内调用配置中的 `memory_search`,
-结束日期为昨天。`historical_search_limit` 控制每个当前事件最多请求多少条检索结果。只有路径名为
-`auto_fin_news.md` 的命中才会贡献候选 ID;Step 会重新读取源 Markdown 并解析 ID,再调用 History Agent。
-
-History Agent 默认最多选择五条候选,并将关系标记为 `same` 或 `opposite`。代码会移除未知或重复 ID 以及空理由,
-随后计算 D1、D2、D3、D5 的复权累计收益:
-
-- 交易日 15:00 前发生的事件,以当日复权收盘价为入场价,D1 是下一交易日收盘价;
-- 15:00 或之后发生的事件,以下一交易日复权开盘价为入场价,D1 是该日收盘价;
-- 如果无法从有效正价格和复权因子计算入场价或某个期限,该值为 `null`。
-
-最终 Agent 接收固定 ETF 列表、所有当前/历史证据、`same`/`opposite` 方向、代码计算的收益,以及此前最近一份
-`auto_fin.md`。它自行判断证据是否支持推荐或应明确观望;代码不会计算评分、期望收益,也不强制给出持有期限。
-
-## 产物
-
-```text
-reme_workspace/
-├── daily/
-│ ├── YYYY-MM-DD.md
-│ └── YYYY-MM-DD/
-│ ├── auto_fin_news.md
-│ └── auto_fin.md
-└── resource/
- ├── fin/
- │ ├── etfs.json
- │ ├── 518880.SH.jsonl
- │ └── <其他固定 ETF>.jsonl
- └── YYYY-MM-DD/
- ├── auto_fin_topic_output.json
- ├── auto_fin_history_001_output.json
- ├── ...
- ├── auto_fin_analysis.jsonl
- └── auto_fin_merge_output.json
-```
-
-每日新闻和报告是用户拥有的 Markdown。`resource/fin/` 是确定性收益计算所用的行情缓存;日期目录下的 JSON/JSONL
-保留结构化 Agent 回复和整理后的分析。写入通过同目录临时文件原子替换;报告写完后会刷新当日索引。
-
-## 参数与默认值
-
-公开 Job 参数:
-
-| 参数 | 默认值 | 作用 |
-|---|---:|---|
-| `date` | `""` | 空值使用 `Asia/Shanghai` 当天;非空值必须是严格 `YYYY-MM-DD` 且等于当天 |
-| `historical_search_limit` | `10` | 每个当前事件请求的 `memory_search` 结果上限;最小值为 1 |
-
-`daily_cookbook.yaml` 中相关的 Job 级配置:
-
-| 配置 | 默认值 | 作用 |
-|---|---:|---|
-| `etf_codes` | 上述三个启用代码 | 固定 ETF 研究范围 |
-| `news_lookback_days` | `60` | 本地新闻及历史检索窗口 |
-| `current_news_limit_per_etf` | `10` | 每只 ETF 最多保留的当前事件数 |
-| `historical_news_limit` | `5` | 每个当前事件最多保留的可比历史事件数 |
-
-当前没有公开 `force` 参数。更早的新闻文件会复用;当天新闻和所有固定 ETF 行情文件会刷新;同一天再次成功运行会覆盖
-当天报告和 resource 产物。
-
-## 环境变量与定时任务
-
-| 变量 | 必需 | 作用 |
-|---|---|---|
-| `TUSHARE_TOKEN` | 是 | 交易日历、财联社新闻、ETF 元数据、价格与复权因子 |
-| `LLM_API_KEY` | 取决于服务商 | 共享 AgentScope LLM 凭据;配置默认值为空 |
-| `LLM_MODEL_NAME` | 否 | 默认 `qwen3.7-plus` |
-| `LLM_BASE_URL` | 取决于服务商 | OpenAI 兼容 endpoint;无内置默认值 |
-| `TUSHARE_MIRROR_URL` | 否 | 去掉末尾 `/` 后替换 Tushare SDK HTTP URL |
-| `DAILY_PAPER_WORKSPACE_DIR` | 否 | standalone cookbook 的共享 workspace |
-| `DINGTALK_*` | 否 | 可选的钉钉应用、机器人和群设置 |
-
-镜像可按需配置,例如:
-
-```bash
-export TUSHARE_MIRROR_URL="http://112.124.63.173:4000/tushare"
-```
-
-`auto_fin_0930_cron`、`auto_fin_1130_cron` 和 `auto_fin_1800_cron` 按 `Asia/Shanghai` 时区每天 09:30、11:30 和
-18:00 触发。Cron 在周末和节假日仍会启动,但如果 Tushare 返回当天不是上交所交易日,Data Step 会跳过后续工作流;
-同一天的后续运行会在已有报告基础上继续完善。
-
-要发送完成的报告,需要配置 `DINGTALK_APP_KEY`、`DINGTALK_APP_SECRET`、`DINGTALK_ROBOT_CODE` 和逗号分隔的
-`DINGTALK_CONVERSATION_IDS`。没有会话 ID 时发送步骤无副作用。
-
-## Agent 与失败边界
-
-Auto Fin 和 Daily Paper 共用无工具的 `default` AgentScope wrapper,其模型调用不会暴露内置工具或配置型 Job
-工具。Auto Fin 由确定性的 Step 代码主动调用 `memory_search`,这不是 Agent 工具调用。独立的交互式
-`dingtalk_wait` Step 才有自己的 `bash` 和 ReMe Job tool allowlist。
-
-standalone 配置默认未启用 embedding store,因此 `memory_search` 使用可用的 BM25 路径;只有启用被注释的
-embedding 组件后才有向量/BM25 融合。
-
-非法日期、缺少凭据或服务、模型结构化输出无效、固定 ETF 未知、行情文件缺失、记忆检索失败都会终止 Job;休市日是
-成功跳过。工作流没有同日期全局执行锁或跨文件事务;重复成功运行也可能重复发送钉钉通知。
-
-## 测试
-
-聚焦单元测试会 mock 模型和行情数据边界:
-
-```bash
-python -m pip install -e ".[dev,core]"
-pytest tests/unit/test_auto_fin.py -v
-```
-
-需要真实 Tushare、LLM 或钉钉凭据的测试应单独运行,且需要显式授权。
diff --git a/docs/en/auto_dream.md b/docs/en/auto_dream.md
index 749ebe46..bb6177f2 100644
--- a/docs/en/auto_dream.md
+++ b/docs/en/auto_dream.md
@@ -1,16 +1,17 @@
# Auto Dream
-`auto_dream` is ReMe's long-term memory distillation flow from daily to digest. It scans daily inputs for a specified date,
-processes only files that changed since the previous dream, extracts content worth retaining as memory units, integrates those
-units into `digest/`, and generates the day's `interests.yaml` for proactive use.
+`auto_dream` is ReMe's long-term memory distillation flow from daily to digest. By default it scans the target date and
+the previous day, processes only files changed since the previous dream, extracts a small set of high-value memory units
+across that window, integrates them into `digest/`, and writes the target day's `interests.yaml` for proactive use.
Its daily inputs usually come from [Auto Memory](./auto_memory.md) and [Auto Resource](./auto_resource.md). For the file
-semantics of `digest/`, Sources sections, and wikilinks, see [Memory as File](./memory_as_file.md). For the linking strategy
-used during Integrate, see [Auto Link](./auto_link.md). To read `interests.yaml`, use [Proactive](./proactive.md).
+semantics of `digest/`, Sources sections, and wikilinks, see [Memory as File](./memory_as_file.md). For the linking
+strategy used during Integrate, see [Auto Link](./auto_link.md). To read `interests.yaml`,
+use [Proactive](./proactive.md).
## Configuration
@@ -26,6 +27,12 @@ auto_dream:
hint:
type: string
default: ""
+ scan_days:
+ type: integer
+ default: 2
+ max_units:
+ type: integer
+ default: 5
topic_count:
type: integer
default: 3
@@ -36,6 +43,8 @@ auto_dream:
- backend: dream_extract_step
file_catalog: dream
topic_session_id: interests
+ scan_days: 2
+ max_units: 5
- backend: dream_integrate_step
- backend: dream_topics_step
topic_count: 3
@@ -46,34 +55,39 @@ auto_dream:
Parameters:
-| Parameter | Purpose |
-|---|---|
-| `date` | Date to process in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
-| `hint` | Additional guidance from the caller for the Extract and Integrate stages. |
-| `topic_count` | Maximum number of topics written to `interests.yaml`. Defaults to 3. |
+| Parameter | Purpose |
+|------------------------|---------------------------------------------------------------------------------------------------------|
+| `date` | Date to process in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
+| `hint` | Additional guidance from the caller for the Extract and Integrate stages. |
+| `scan_days` | Recent-date window ending at `date`; defaults to 2 and has a minimum of 1. |
+| `max_units` | Maximum reusable units extracted in one run; defaults to 5. |
+| `topic_count` | Maximum number of topics written to `interests.yaml`. Defaults to 3. |
| `topic_diversity_days` | Number of past days of `interests.yaml` files considered when avoiding duplicate topics. Defaults to 7. |
## Inputs and Outputs
-Inputs are daily Markdown files for the specified date:
+Inputs are daily Markdown files from the most recent `scan_days` ending at the specified date. For example,
+`date=2026-06-20` with `scan_days=2` scans:
```text
-daily/.md
-daily//**/*.md
+daily/2026-06-19.md
+daily/2026-06-19/**/*.md
+daily/2026-06-20.md
+daily/2026-06-20/**/*.md
```
-`daily//interests.yaml` is excluded from extraction input so topics from the previous run do not feed back into the
-next extraction.
+Every `daily//interests.yaml` in the scan window is excluded from extraction so previous proactive output cannot
+feed back into the next run. Final topics are written only for the target date.
The main outputs are:
-| Output | Description |
-|---|---|
-| `digest/procedure/*.md` | Methods, workflows, runbooks, and executable experience. |
-| `digest/personal/*.md` | User-, team-, and project-related preferences, facts, and long-term context. |
-| `digest/wiki/*.md` | General knowledge, concepts, observations, and decision precedents. |
-| `daily//interests.yaml` | Topics worth proactive attention from the host agent that day. |
-| `metadata/file_catalog/dream*` | Dream-specific catalog used to detect changes in daily inputs. |
+| Output | Description |
+|--------------------------------|------------------------------------------------------------------------------|
+| `digest/procedure/*.md` | Methods, workflows, runbooks, and executable experience. |
+| `digest/personal/*.md` | User-, team-, and project-related preferences, facts, and long-term context. |
+| `digest/wiki/*.md` | General knowledge, concepts, observations, and decision precedents. |
+| `daily//interests.yaml` | Topics worth proactive attention from the host agent that day. |
+| `metadata/file_catalog/dream*` | Dream-specific catalog used to detect changes in daily inputs. |
## Four Stages
@@ -81,43 +95,50 @@ The main outputs are:
`dream_extract_step` performs three tasks:
-1. Refresh the day's index page at `daily/.md`.
-2. Scan `daily/.md` and `daily//**/*.md` and compare their mtimes with `file_catalog: dream`.
-3. Send only changed files to the LLM and globally extract two structured result types: `units` and `topics`.
+1. Refresh each `daily/.md` in the scan window.
+2. Scan those day indexes and `daily//**/*.md`, comparing mtimes with `file_catalog: dream`.
+3. Send all changed files together to the LLM and globally extract two structured result types: `units` and `topics`.
`units` are long-term memory units ready to be distilled into digest. Each has `name`, `bucket`, `summary`, and `paths`.
-`bucket` may only be `procedure`, `personal`, or `wiki`; unknown values are routed to `wiki`.
+A run returns at most `max_units`; extraction merges cross-file evidence for the same abstraction and drops passing
+mentions, per-file summaries, and weak candidates without reusable value. `bucket` may only be `procedure`, `personal`,
+or `wiki`; unknown values are routed to `wiki`.
`topics` are proactive-interest candidates for the day. They contain `title`, `reason`, `evidence`, `keywords`, and
`paths` and are filtered again in the Topics stage.
-If there are no changed files, the flow ends early with success and skips later extraction work. If files changed but no LLM
-is configured, Extract fails because extraction requires an LLM.
+If there are no changed files, Extract succeeds with no units; Integrate then has no unit work, Topics preserves any
+existing target-day topics, and Finish still performs its normal catalog summary. If files changed but no LLM is
+configured, Extract fails because extraction requires an LLM.
### 2. Integrate
-`dream_integrate_step` invokes an agent independently for each unit and integrates that unit into one digest node. It exposes
-these tools to the agent:
+`dream_integrate_step` invokes an agent independently for each unit and integrates that unit into one digest node. It
+exposes these tools to the agent:
```text
node_search, read, frontmatter_read, write, edit, frontmatter_update
```
-This stage carries the core responsibility of `auto_link`. It first uses `node_search` to recall similar or related nodes at
-digest-node granularity, decides whether to create or update a node, and finally writes sources and related digest nodes as
-wikilinks. See [Auto Link](./auto_link.md) for the recall, deduplication, and edge-writing rules.
+This stage carries the core responsibility of `auto_link`. It first uses `node_search` to recall similar or related
+nodes at digest-node granularity, decides whether to create or update a node, and finally writes sources and related
+digest nodes as wikilinks. See [Auto Link](./auto_link.md) for the recall, deduplication, and edge-writing rules.
+
+Extract is the gate for deciding whether material is worth remembering, so Integrate has no `SKIP` action: each admitted
+unit must land in exactly one digest node. Creates and updates must retain provenance and weave related digest links
+into contextual sentences; bare wikilinks and standalone relationship fields are not valid output.
There are four integration actions:
-| Action | Meaning |
-|---|---|
-| `CREATE` | No equivalent abstraction exists; create a new digest node. |
+| Action | Meaning |
+|---------------|--------------------------------------------------------------------------------|
+| `CREATE` | No equivalent abstraction exists; create a new digest node. |
| `CORROBORATE` | The same memory appeared again; append a source or strengthen the description. |
-| `REFINE` | New material adds boundaries, steps, prerequisites, applicability, or detail. |
-| `CORRECT` | New material corrects errors, omissions, or conflicts in the existing node. |
+| `REFINE` | New material adds boundaries, steps, prerequisites, applicability, or detail. |
+| `CORRECT` | New material corrects errors, omissions, or conflicts in the existing node. |
-Successfully integrated units are recorded in `integrate_results`. Failed units enter `failed_units`, and their source paths
-enter `failed_paths`. The Finish stage does not checkpoint failed paths, ensuring that they can be retried later.
+Successfully integrated units are recorded in `integrate_results`. Failed units enter `failed_units`, and their source
+paths enter `failed_paths`. The Finish stage does not checkpoint failed paths, ensuring that they can be retried later.
### 3. Topics
@@ -127,7 +148,7 @@ It reads:
```text
daily//interests.yaml
-daily//interests.yaml
+daily//interests.yaml
```
Existing topics from the same day are preserved, while similar topics from the previous `topic_diversity_days` days are
@@ -156,7 +177,8 @@ topics:
`dream_finish_step` completes the run:
1. Write successfully processed changed paths to `file_catalog: dream`.
-2. Also write `daily//interests.yaml` and `daily/.md` to the catalog.
+2. Also write the target `daily//interests.yaml` and every refreshed day-index page in the scan window to the
+ catalog.
3. Persist the dream catalog if there were upserts or deletions.
4. Return a summary containing counts for scanned, changed, integrated, topics, checkpoints, and related values.
@@ -177,6 +199,12 @@ With caller guidance:
reme auto_dream date=2026-06-20 hint="Prioritize engineering decisions and long-term preferences"
```
+Override the default scan window and unit cap:
+
+```bash
+reme auto_dream date=2026-06-20 scan_days=3 max_units=8
+```
+
The same set of steps can also be placed in a `cron` Job, for example to run every morning:
```yaml
@@ -195,15 +223,16 @@ jobs:
## Important Boundaries
-`auto_dream` consumes only daily inputs and does not rewrite daily bodies. Daily preserves facts and the original situation;
-digest is the abstracted long-term memory layer.
+`auto_dream` consumes only daily inputs and does not rewrite daily bodies. Daily preserves facts and the original
+situation; digest is the abstracted long-term memory layer.
-`digest` is not a copy of the source text. Its body should preserve reusable abstractions, while a Sources section points
-back with entries such as `- [[daily//...]]`. Links follow the workspace-relative wikilink semantics described in
+`digest` is not a copy of the source text. Its body should preserve reusable abstractions, while a Sources section
+points back with contextual sentences such as `The decision was recorded in [[daily//decision.md]].` Links follow
+the workspace-relative wikilink semantics described in
[Memory as File](./memory_as_file.md).
-`auto_dream` does not invent an overview from nothing. Only content that actually appears in daily input and is extracted as
-a unit or topic can enter digest or `interests.yaml`.
+`auto_dream` does not invent an overview from nothing. Only content that actually appears in daily input and is
+extracted as a unit or topic can enter digest or `interests.yaml`.
-The complete flow depends on an LLM for Extract and Integrate. Topics can perform local deduplication without an LLM, but that
-does not mean the full dream flow can run offline.
+The complete flow depends on an LLM for Extract and Integrate. Topics can perform local deduplication without an LLM,
+but that does not mean the full dream flow can run offline.
diff --git a/docs/en/auto_link.md b/docs/en/auto_link.md
index 0a97976a..dc016e9e 100644
--- a/docs/en/auto_link.md
+++ b/docs/en/auto_link.md
@@ -1,11 +1,12 @@
# Auto Link
-In the current implementation, `auto_link` is not a separately registered Job. It is a capability of the Integrate stage in
+In the current implementation, `auto_link` is not a separately registered Job. It is a capability of the Integrate stage
+in
`auto_dream`: when `dream_integrate_step` writes a memory unit to `digest/`, it also recalls digest nodes, makes a
deduplication decision, links sources, and weaves wikilinks to related nodes into the result.
-For the complete dream flow, see [Auto Dream](./auto_dream.md). For general wikilink, frontmatter, and workspace-relative
-path semantics, see [Memory as File](./memory_as_file.md). For question-answering retrieval, see
+For the complete dream flow, see [Auto Dream](./auto_dream.md). For general wikilink, frontmatter, and
+workspace-relative path semantics, see [Memory as File](./memory_as_file.md). For question-answering retrieval, see
[Memory Search](./memory_search.md).
## Where It Runs
@@ -21,19 +22,19 @@ auto_dream:
- dream_finish_step
```
-The Integrate stage processes each unit independently. A unit is written to exactly one target digest node, but that node may
-link to multiple sources and multiple related digest nodes.
+The Integrate stage processes each unit independently. A unit is written to exactly one target digest node, but that
+node may link to multiple sources and multiple related digest nodes.
## Goals
`auto_link` addresses graph quality at write time:
-| Problem | Handling |
-|---|---|
-| The same memory already exists | Recall and update the existing node instead of creating a duplicate. |
-| New and existing material are related | Write workspace-relative wikilinks into the body. |
-| A digest node is disconnected from its sources | Add daily/resource links under a `## Sources` section. |
-| A node contains only isolated prose | Add links to related digest nodes on both CREATE and UPDATE. |
+| Problem | Handling |
+|------------------------------------------------|----------------------------------------------------------------------|
+| The same memory already exists | Recall and update the existing node instead of creating a duplicate. |
+| New and existing material are related | Write workspace-relative wikilinks into the body. |
+| A digest node is disconnected from its sources | Add daily/resource links under a `## Sources` section. |
+| A node contains only isolated prose | Add links to related digest nodes on both CREATE and UPDATE. |
## Toolchain
@@ -48,41 +49,42 @@ edit
frontmatter_update
```
-`node_search` is digest-only node retrieval designed for dream integration. It returns node-level signals such as the digest
-node's `path` and the `name` and `description` from frontmatter. It does not expand the body and does not perform the link
-expansion used by ordinary search.
+`node_search` is digest-only node retrieval designed for dream integration. It returns node-level signals such as the
+digest node's `path` and the `name` and `description` from frontmatter. It does not expand the body and does not perform
+the link expansion used by ordinary search.
-`read` and `frontmatter_read` are used only for candidates that may be relevant, avoiding expansion of every recalled result
-into a large context.
+`read` and `frontmatter_read` are used only for candidates that may be relevant, avoiding expansion of every recalled
+result into a large context.
## Linking Flow
### 1. Recall candidate nodes
-The agent first calls `node_search` with the unit's triggers, verbs, nouns, synonyms, and possible failure modes. Broad recall,
-for example `limit=20-30`, is recommended by default because this step serves both deduplication and link discovery.
+The agent first calls `node_search` with the unit's triggers, verbs, nouns, synonyms, and possible failure modes. Broad
+recall, for example `limit=20-30`, is recommended by default because this step serves both deduplication and link
+discovery.
Recalled results are internally classified into three groups:
-| Classification | Meaning | Next action |
-|---|---|---|
-| `same_abstraction` | The trigger or underlying abstraction is the same, with substantial content overlap. | Use as the UPDATE target. |
-| `related` | An adjacent process, prerequisite, failure mode, concept, preference, or upstream/downstream knowledge. | Write a body wikilink. |
-| `unrelated` | Only superficially similar or unrelated. | Ignore. |
+| Classification | Meaning | Next action |
+|--------------------|---------------------------------------------------------------------------------------------------------|---------------------------|
+| `same_abstraction` | The trigger or underlying abstraction is the same, with substantial content overlap. | Use as the UPDATE target. |
+| `related` | An adjacent process, prerequisite, failure mode, concept, preference, or upstream/downstream knowledge. | Write a body wikilink. |
+| `unrelated` | Only superficially similar or unrelated. | Ignore. |
### 2. Choose a write action
Every unit must select one action:
-| Action | Linking semantics |
-|---|---|
-| `CREATE` | Write a new `digest//.md` and add source and related-node links to its body. |
-| `CORROBORATE` | The same abstraction appeared again; append its source link and strengthen the description when needed. |
-| `REFINE` | New material extends the existing node; insert the additional content in the appropriate section and preserve existing links. |
-| `CORRECT` | New material corrects the existing node; use source links to identify the basis for the correction. |
+| Action | Linking semantics |
+|---------------|-------------------------------------------------------------------------------------------------------------------------------|
+| `CREATE` | Write a new `digest//.md` and add source and related-node links to its body. |
+| `CORROBORATE` | The same abstraction appeared again; append its source link and strengthen the description when needed. |
+| `REFINE` | New material extends the existing node; insert the additional content in the appropriate section and preserve existing links. |
+| `CORRECT` | New material corrects the existing node; use source links to identify the basis for the correction. |
-An UPDATE should be additive whenever possible: do not delete existing wikilinks or source entries. This prevents
-later graph indexing and retrieval from losing edges.
+An UPDATE should be additive whenever possible: do not delete existing wikilinks or source entries. This prevents later
+graph indexing and retrieval from losing edges.
### 3. Write source edges
@@ -91,12 +93,13 @@ Source edges are ordinary wikilinks grouped under a Markdown heading:
```markdown
## Sources
-- [[daily/2026-06-20/session.md]]
-- [[resource/2026-06-20/paper.md]]
+The decision was recorded in [[daily/2026-06-20/session.md]], while the supporting technical evidence comes from
+[[resource/2026-06-20/paper.md]].
```
-These edges represent the evidence behind a digest node. Plain-text descriptions do not count as source edges because only
-wikilinks can be parsed reliably by the file graph. For the complete parsing rules, see
+These edges represent the evidence behind a digest node. Plain-text descriptions do not count as source edges because
+only wikilinks can be parsed reliably by the file graph. The surrounding sentence must explain what each source
+supports; a bare wikilink line is not valid Integrate output. For the complete parsing rules, see
[Memory as File](./memory_as_file.md#wikilink).
### 4. Write relationships between digest nodes
@@ -113,11 +116,11 @@ This design extends [[digest/wiki/hybrid-search.md]] and uses
`auto_link` adjusts the shape of its output according to the unit bucket:
-| Bucket | Writing focus |
-|---|---|
-| `procedure` | Write a runbook with triggers, steps, inputs, and failure modes. Link prerequisites, substeps, and related preferences. |
-| `personal` | Write user-, team-, or project-specific facts and preferences. Link related projects, habits, and decision context. |
-| `wiki` | Write general knowledge, principles, observations, and decision precedents. Link concepts, methods, and adjacent knowledge. |
+| Bucket | Writing focus |
+|-------------|-----------------------------------------------------------------------------------------------------------------------------|
+| `procedure` | Write a runbook with triggers, steps, inputs, and failure modes. Link prerequisites, substeps, and related preferences. |
+| `personal` | Write user-, team-, or project-specific facts and preferences. Link related projects, habits, and decision context. |
+| `wiki` | Write general knowledge, principles, observations, and decision precedents. Link concepts, methods, and adjacent knowledge. |
Regardless of bucket, preserve source edges and weave recalled related digest nodes into the body whenever possible.
@@ -125,19 +128,19 @@ Regardless of bucket, preserve source edges and weave recalled related digest no
`auto_link` uses `node_search`, not the question-answering `search`.
-| Capability | Purpose |
-|---|---|
-| `search` | External question answering; returns chunks and can expand upstream/downstream link context. |
+| Capability | Purpose |
+|---------------|-----------------------------------------------------------------------------------------------------------|
+| `search` | External question answering; returns chunks and can expand upstream/downstream link context. |
| `node_search` | Dream integration; recalls only digest node-level summaries for deduplication and related-link decisions. |
-This boundary matters. The Integrate stage needs to decide whether the same abstraction already exists and which nodes should
-be linked; it should not load large numbers of body chunks into context. [Memory Search](./memory_search.md) handles
-question-oriented chunk retrieval, RRF fusion, and link expansion.
+This boundary matters. The Integrate stage needs to decide whether the same abstraction already exists and which nodes
+should be linked; it should not load large numbers of body chunks into context. [Memory Search](./memory_search.md)
+handles question-oriented chunk retrieval, RRF fusion, and link expansion.
## Failure and Retry
If integration of a unit fails, `dream_integrate_step` records `failed_units` and `failed_paths`.
`dream_finish_step` does not checkpoint those source paths, so the next `auto_dream` run processes them again.
-This makes auto_link writes retryable: a failure does not mark the input as complete or silently discard digest edges that
-should have been created.
+This makes auto_link writes retryable: a failure does not mark the input as complete or silently discard digest edges
+that should have been created.
diff --git a/docs/en/auto_memory.md b/docs/en/auto_memory.md
index c29c1e4a..ffe3db95 100644
--- a/docs/en/auto_memory.md
+++ b/docs/en/auto_memory.md
@@ -1,8 +1,9 @@
# Auto Memory
-Auto Memory is ReMe's entry point for conversational memory. Each conversation is first distilled into a daily memory card
-identified by `session_id`, and the day's `YYYY-MM-DD.md` page then indexes all of those cards. It turns "we talked about it"
-into "it was remembered" while preserving the original conversation as evidence.
+Auto Memory is ReMe's entry point for conversational memory. Within a target date, it uses `session_id` to find or update at
+most one daily memory card, whose filename is a concise topic or event name chosen by the Agent. The day's `YYYY-MM-DD.md`
+page indexes those cards. It turns "we talked about it" into "it was remembered" while retaining a source conversation record
+as evidence.
@@ -13,9 +14,9 @@ For the general file semantics of `daily/`, `session/`, frontmatter, and wikilin
```text
Conversation
- ├─ step 1: daily/YYYY-MM-DD/.md # one card per conversation
- ├─ step 2: daily/YYYY-MM-DD.md # daily index linking the cards
- └─ source: session/dialog/.jsonl # original conversation
+ ├─ step 1: daily/YYYY-MM-DD/.md # one topic-named card per session
+ ├─ step 2: daily/YYYY-MM-DD.md # daily index linking the cards
+ └─ source: session/dialog/.jsonl # source conversation record
```
## What It Records
@@ -39,29 +40,33 @@ workspace/
daily/
2026-06-20.md
2026-06-20/
- session-a.md
- session-b.md
+ login-refactor-decision.md
+ retrieval-regression.md
```
-`daily/2026-06-20/session-a.md` and `daily/2026-06-20/session-b.md` are memory cards distilled from different
-conversations. `daily/2026-06-20.md` is the index page for that day. Resource files enter the same daily memory layer; see
+The two files under the date directory are topic-named cards distilled from different conversations.
+`daily/2026-06-20.md` is the index page for that day. Resource files enter the same daily memory layer; see
[Auto Resource](./auto_resource.md).
-When a call includes `session_id`, Auto Memory records that conversation separately under the given ID:
+When a call includes `session_id`, Auto Memory uses it to find the corresponding card through frontmatter, while the Agent
+chooses a readable filename through `name`:
-```text
-daily/2026-06-20/session-a.md
+```yaml
+name: login-refactor-decision
+session_id: session-a
+source_conversation: "[[session/dialog/session-a.jsonl]]"
```
-This keeps different conversations separate. A requirements discussion, a debugging session, and a documentation update can
-each have their own memory card. To see what happened on a particular day, start with `YYYY-MM-DD.md`. To inspect what was
-distilled from one conversation, open the corresponding `.md`.
+This keeps different conversations separate without forcing opaque IDs into filenames. An update locates the existing note by
+`session_id` or `source_conversation`; if the Agent supplies a better frontmatter `name`, the system can rename the note and
+retarget inbound wikilinks. To see what happened on a day, start with `YYYY-MM-DD.md`.
## Preserving the Original Information
-The distilled daily note is optimized for readability; the original conversation is retained for trust and verification.
+The distilled daily note is optimized for readability; a filtered source conversation record is retained for trust and
+verification.
-While generating memory cards, Auto Memory also saves the raw sessions:
+While generating memory cards, Auto Memory also saves the source messages:
```text
session/
@@ -70,12 +75,12 @@ session/
session-b.jsonl
```
-Each daily note points to its corresponding original conversation. When a memory needs verification, follow that link back to
-the complete context in which it was created.
+Each daily note points to its corresponding conversation record. Saved messages omit tool-result blocks and base64 data
+blocks, preventing recalled memory and binary payloads from being mistaken for user-provided evidence later.
## Message Timestamps
-Auto Memory preserves each message's `created_at` in both the prompt and the raw session JSONL. When importing historical
+Auto Memory preserves each retained message's `created_at` in both the prompt and the source conversation JSONL. When importing historical
conversations or benchmark data, provide the actual occurrence time for every message so the model does not confuse event
time with execution time:
@@ -92,8 +97,8 @@ For compatibility with common dataset schemas, `auto_memory` also checks `time_c
`timeCreated`, and `created_time` when `created_at` is absent. These fields may appear either at the top level of a message
or inside `metadata`.
-When a call does not explicitly provide `date`, Auto Memory uses the date of the earliest valid `created_at` value in the
-messages. If no message contains a valid timestamp, it falls back to the current date. Historical imports may also specify the
+When a call does not explicitly provide `date`, Auto Memory uses the latest valid `created_at` date in the messages. If no
+message contains a valid timestamp, it falls back to the current date. Historical imports may also specify the
target date directly:
```bash
diff --git a/docs/en/auto_resource.md b/docs/en/auto_resource.md
index a3ae3f19..9d9495b1 100644
--- a/docs/en/auto_resource.md
+++ b/docs/en/auto_resource.md
@@ -1,8 +1,8 @@
# Auto Resource `Beta`
Auto Resource is ReMe's entry point for interpreting resources and is currently in **Beta**. Resource files first enter
-`resource/` by date and are then interpreted into daily resource cards. Each card's filename comes from the LLM-generated
-frontmatter `name`, and `source_resource` links the card back to its original file.
+`resource/`, preferably under a date directory, and are then interpreted into daily resource cards. Each card's filename
+comes from the LLM-generated frontmatter `name`, and `source_resource` links the card back to its original file.
@@ -13,7 +13,7 @@ For the general file semantics of workspace layers, `resource/`, and `daily/`, s
[Auto Memory](./auto_memory.md).
```text
-resource/YYYY-MM-DD/
+resource/[YYYY-MM-DD/]
├─ step 1: daily/YYYY-MM-DD/.md # interpreted resource card
├─ step 2: source_resource points to the original resource
└─ step 3: daily/YYYY-MM-DD.md # daily index linking the cards
@@ -21,8 +21,8 @@ resource/YYYY-MM-DD/
## What It Records
-Auto Resource does more than copy file content. It extracts information that will make the resource easier to retrieve and
-understand later:
+Auto Resource does more than copy file content. It extracts information that will make the resource easier to retrieve
+and understand later:
- Core content: what the resource is mainly about.
- Structure: its sections, tables, fields, and data organization.
@@ -34,26 +34,28 @@ In short, it turns "a file was archived" into "the resource is usable."
## Original Resource Entry Point
-Auto Resource uses `resource/` as the entry point for source material. Resources must be placed under a date, which determines
-the day whose daily memory layer receives the interpreted card.
+Auto Resource uses `resource/` as the entry point for source material. Date directories are recommended, and their date
+determines which daily memory layer receives the interpreted card. A file directly under `resource/` is also supported
+and uses today in the application timezone.
Example directory:
```text
workspace/
resource/
+ quick-note.txt # enters today's daily layer
2026-06-20/
market-report.md
meeting-notes.csv
```
-The current Beta version is best suited to text-based resources such as `md`, `txt`, `json`, `jsonl`, `csv`, `yaml`,
-and `html`.
+The current Beta version is best suited to text-based resources such as `md`, `txt`, `json`, `jsonl`, `csv`, `yaml`, and
+`html`.
## Resource Cards
-Each resource file produces one daily resource card. The system initially uses the resource file's stem as a temporary path.
-After the agent writes the card, the file is renamed according to its frontmatter `name`:
+Each resource file produces one daily resource card. The system initially uses the resource file's stem as a temporary
+path. After the agent writes the card, the file is renamed according to its frontmatter `name`:
```text
resource/2026-06-20/market-report.md
@@ -67,14 +69,14 @@ The resource card links to the original file through frontmatter:
source_resource: "[[resource/2026-06-20/market-report.md]]"
```
-When a resource changes, Auto Resource finds and updates the corresponding card through `source_resource`. When a resource is
-deleted, its daily note is also removed. The older `daily/YYYY-MM-DD/.md` naming convention remains supported
-as a fallback.
+When a resource changes, Auto Resource finds and updates the corresponding card through `source_resource`. When a
+resource is deleted, its daily note is also removed. The older `daily/YYYY-MM-DD/.md` naming convention
+remains supported as a fallback.
## Daily Index
-Resource cards enter the same daily memory layer as Auto Memory cards. The day's `YYYY-MM-DD.md` page acts as an index and
-organizes those resource cards:
+Resource cards enter the same daily memory layer as Auto Memory cards. The day's `YYYY-MM-DD.md` page acts as an index
+and organizes those resource cards:
```text
daily/
@@ -91,11 +93,13 @@ resource, open its corresponding resource card.
The interpreted daily note is optimized for readability; the original resource is retained for trust and verification.
-Auto Resource does not move the original file. It remains under `resource/YYYY-MM-DD/`. Text resources can therefore enter
-the daily memory flow while their source files stay in their original location.
+Auto Resource does not move the original file. It remains at its original path under `resource/`. Text resources can
+therefore enter the daily memory flow while their source files stay in their original location.
## What Happens Next
-Auto Resource only creates resource interpretations in the daily layer. To distill long-term knowledge from resources into
-`digest/`, use [Auto Dream](./auto_dream.md). To search original resources, daily cards, and digest nodes, use
-[Memory Search](./memory_search.md).
+Auto Resource only creates resource interpretations in the daily layer. To distill long-term knowledge from resources
+into
+`digest/`, use [Auto Dream](./auto_dream.md). The default live index covers daily cards and digest nodes. Run
+`reme reindex`
+when original resource files must also be directly searchable; see [Memory Search](./memory_search.md).
diff --git a/docs/en/contributing.md b/docs/en/contributing.md
index b4fea323..e9fe9c3e 100644
--- a/docs/en/contributing.md
+++ b/docs/en/contributing.md
@@ -8,12 +8,12 @@ ReMe is open source and hosted on GitHub:
## How to Contribute
-Thank you for your interest in ReMe. ReMe is a file-first, self-evolving memory system for agents. Contributions are welcome
-through issue reports, documentation improvements, additional tests, bug fixes, and new capabilities.
+Thank you for your interest in ReMe. ReMe is a file-first, self-evolving memory system for agents. Contributions are
+welcome through issue reports, documentation improvements, additional tests, bug fixes, and new capabilities.
-If this is your first time running ReMe locally, start with [Quick Start](./quick_start.md). If your change affects runtime
-layers, Jobs, Steps, or components, read [ReMe Framework](./framework.md). If it affects workspace directories, frontmatter,
-wikilinks, or chunking, read [Memory as File](./memory_as_file.md).
+If this is your first time running ReMe locally, start with [Quick Start](./quick_start.md). If your change affects
+runtime layers, Jobs, Steps, or components, read [ReMe Framework](./framework.md). If it affects workspace directories,
+frontmatter, wikilinks, or chunking, read [Memory as File](./memory_as_file.md).
### 1. Before You Begin
@@ -21,9 +21,10 @@ Before investing in an implementation:
- Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) for an existing issue or discussion.
- If a related issue is still open, comment that you would like to work on it to avoid duplicate effort.
-- If no issue exists, create one describing the context, expected behavior, possible implementation, and scope of impact.
-- For larger feature changes, align with maintainers on interfaces, configuration, compatibility, and test strategy before
- submitting an implementation.
+- If no issue exists, create one describing the context, expected behavior, possible implementation, and scope of
+ impact.
+- For larger feature changes, align with maintainers on interfaces, configuration, compatibility, and test strategy
+ before submitting an implementation.
### 2. Local Development Environment
@@ -38,7 +39,11 @@ The project requires Python 3.11 or later. A virtual environment is recommended:
```bash
python -m venv .venv
source .venv/bin/activate
-pip install -e ".[dev,full]"
+pip install -e packages/reme_ai_studio -e ".[dev,full]"
+cd website
+npm ci
+npm run build:static
+cd ..
pre-commit install
```
@@ -53,8 +58,8 @@ CLI / Client -> Service -> Application -> Job -> Step -> Component / Workspace
In practice:
-- Capabilities exposed to users or external systems should normally be orchestrated by a Job, then exposed by a Service as a
- CLI-, HTTP-, or MCP-callable interface.
+- Capabilities exposed to users or external systems should normally be orchestrated by a Job, then exposed by a Service
+ as a CLI-, HTTP-, or MCP-callable interface.
- Reusable infrastructure belongs in `reme/components/`, with dependencies declared through `BaseComponent.bind()`.
- Atomic business operations belong in `reme/steps/` and access the file store, agent wrapper, catalog, LLM, and other
components through `BaseStep.Ref`.
@@ -65,31 +70,33 @@ In practice:
When adding a Step or Job, pay particular attention to these conventions:
-- Register implementations with `@R.register("")`. Registration names should be stable, clear, and match the
- configured `backend`.
-- After adding a Step file, make sure its package `__init__.py` imports the module; otherwise, the registry will not load it.
+- Register implementations with `@R.register("")`. Registration names should be stable, clear, and match
+ the configured `backend`.
+- After adding a Step file, make sure its package `__init__.py` imports the module; otherwise, the registry will not
+ load it.
- A Step should perform one atomic business operation. Cross-step flows belong in Job configuration or a dedicated
orchestration Step.
-- A Job composes Steps and selects normal, streaming, background, or scheduled execution. `enable_serve` controls whether it
- is externally exposed.
+- A Job composes Steps and selects normal, streaming, background, or scheduled execution. `enable_serve` controls
+ whether it is externally exposed.
- When a Step needs components, prefer `BaseStep.Ref`. Do not reconstruct global components inside a Step or bypass
`ApplicationContext`.
- File, index, graph, frontmatter, and wikilink behavior must preserve consistent workspace-relative path semantics.
-- Add fast tests under `tests/unit/` for new capabilities. Put cross-component, LLM, embedding, or service behavior under
+- Add fast tests under `tests/unit/` for new capabilities. Put cross-component, LLM, embedding, or service behavior
+ under
`tests/integration/` when appropriate.
### 4. Code and Documentation Changes
Choose the appropriate entry point for the type of change:
-| Change type | Primary location | Guidance |
-|---|---|---|
-| Configuration or startup behavior | `reme/config/`, `reme/application.py`, `reme/reme.py` | Keep the default configuration runnable and avoid breaking existing CLI, HTTP, and MCP entry points. |
-| Component capability | `reme/components/` | Reuse `BaseComponent`, the registry, and context objects. |
-| Job or Step | `reme/components/job/`, `reme/steps/` | Follow the Job -> Step model in [ReMe Framework](./framework.md), keep request and response schemas clear, and add corresponding tests. |
-| Data structure | `reme/schema/`, `reme/enumeration/` | Preserve serialization compatibility and existing frontmatter and wikilink semantics. |
-| Utility | `reme/utils/` | Keep function boundaries small and cover edge cases with unit tests. |
-| User documentation | `docs/en/`, `README.md` | Update documentation when user-visible behavior changes. |
+| Change type | Primary location | Guidance |
+|-----------------------------------|-------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------|
+| Configuration or startup behavior | `reme/config/`, `reme/application.py`, `reme/reme.py` | Keep the default configuration runnable and avoid breaking existing CLI, HTTP, and MCP entry points. |
+| Component capability | `reme/components/` | Reuse `BaseComponent`, the registry, and context objects. |
+| Job or Step | `reme/components/job/`, `reme/steps/` | Follow the Job -> Step model in [ReMe Framework](./framework.md), keep request and response schemas clear, and add corresponding tests. |
+| Data structure | `reme/schema/`, `reme/enumeration/` | Preserve serialization compatibility and existing frontmatter and wikilink semantics. |
+| Utility | `reme/utils/` | Keep function boundaries small and cover edge cases with unit tests. |
+| User documentation | `docs/en/`, `README.md` | Update documentation when user-visible behavior changes. |
If a change involves an LLM, embeddings, an external service, file watching, or a background task, also describe its
dependencies, failure behavior, and local validation method.
@@ -165,15 +172,16 @@ pytest tests/unit/test_reme_cli.py
If `pre-commit` modifies files automatically, commit those changes and rerun the checks until everything passes.
-The current pre-commit configuration includes YAML/TOML/JSON validation, private-key detection, trailing-whitespace checks,
+The current pre-commit configuration includes YAML/TOML/JSON validation, private-key detection, trailing-whitespace
+checks,
`black`, `flake8`, `pylint`, and `pyroma`. The main formatting rules are:
- `black --line-length=120`
- `flake8 --max-line-length=120`
- `pylint --max-line-length=120`
-Some integration tests may require an LLM, embeddings, or external service configuration. If you cannot run them locally,
-state why they were skipped and what alternative validation you completed in the PR description.
+Some integration tests may require an LLM, embeddings, or external service configuration. If you cannot run them
+locally, state why they were skipped and what alternative validation you completed in the PR description.
### 8. Testing Requirements
@@ -183,7 +191,8 @@ Add tests according to the risk of the change:
- For a new Step, Job, or component, cover at least the main path and a failure path.
- For changes to shared logic such as indexes, graphs, wikilinks, frontmatter, or file operations, add edge cases.
- For changes to the CLI, services, or configuration parsing, cover the user-visible entry point.
-- Documentation-only changes usually do not require new tests, but running `pre-commit run --all-files` is still recommended.
+- Documentation-only changes usually do not require new tests, but running `pre-commit run --all-files` is still
+ recommended.
Place tests according to the existing structure:
@@ -213,9 +222,9 @@ Documentation should:
- Bugs and feature requests: [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues)
- Project home: [GitHub Repository](https://github.com/agentscope-ai/ReMe)
-- Documentation site: [https://reme.agentscope.io/](https://reme.agentscope.io/)
+- Documentation site: [https://reme.agentscope.io](https://reme.agentscope.io)
---
-Thank you for contributing to ReMe. Your improvements help make long-term memory for agents more readable, controllable, and
-maintainable.
+Thank you for contributing to ReMe. Your improvements help make long-term memory for agents more readable, controllable,
+and maintainable.
diff --git a/docs/en/framework.md b/docs/en/framework.md
index ec3836c1..c0f0ce52 100644
--- a/docs/en/framework.md
+++ b/docs/en/framework.md
@@ -16,13 +16,13 @@ To run and use ReMe first, see [Quick Start](./quick_start.md). For workspace fi
### Capability Boundary
-ReMe v4 focuses on long-term memory: it distills conversations and resources into `daily/`, organizes them into `digest/`,
-and exposes write, retrieval, and proactive-read capabilities through the CLI, HTTP, and MCP.
+ReMe v4 focuses on long-term memory: it distills conversations and resources into `daily/`, organizes them into
+`digest/`, and exposes write, retrieval, and proactive-read capabilities through the CLI, HTTP, and MCP.
-Single-session context-window management is outside the scope of ReMe v4. This includes compressing the current conversation,
-injecting summaries, trimming tool output, or providing an independent `/compact` interface. Those capabilities belong in
-the host agent framework. ReMe accepts conversations, resources, and file changes that have already occurred and persists the
-information with long-term value.
+Single-session context-window management is outside the scope of ReMe v4. This includes compressing the current
+conversation, injecting summaries, trimming tool output, or providing an independent `/compact` interface. Those
+capabilities belong in the host agent framework. ReMe accepts conversations, resources, and file changes that have
+already occurred and persists the information with long-term value.
```mermaid
flowchart LR
@@ -38,16 +38,16 @@ flowchart LR
Core layers:
-| Layer | Main location | Responsibility |
-|---|---|---|
-| CLI | `reme/reme.py` | Parse commands; `start` launches the service; other actions call the service through a client. |
-| Service | `reme/components/service/` | Register Jobs as HTTP endpoints or MCP tools. |
-| Application | `reme/application.py` | Assemble configured objects, start them in dependency order, close them, and invoke Jobs. |
-| Job | `reme/components/job/` | Orchestrate Steps and select normal, streaming, background, or scheduled execution. |
-| Step | `reme/steps/` | Atomic business operations such as file I/O, retrieval, indexing, and self-evolution. |
-| Component | `reme/components/` | Reusable infrastructure such as file_store, file_graph, keyword_index, and agent_wrapper. |
-| Schema | `reme/schema/` | Data structures such as `Request`, `Response`, `FileChunk`, `FileNode`, and configuration models. |
-| Config | `reme/config/` | Default YAML configuration and command-line override parsing. |
+| Layer | Main location | Responsibility |
+|-------------|----------------------------|---------------------------------------------------------------------------------------------------|
+| CLI | `reme/reme.py` | Parse commands; `start` launches the service; other actions call the service through a client. |
+| Service | `reme/components/service/` | Register Jobs as HTTP endpoints or MCP tools. |
+| Application | `reme/application.py` | Assemble configured objects, start them in dependency order, close them, and invoke Jobs. |
+| Job | `reme/components/job/` | Orchestrate Steps and select normal, streaming, background, or scheduled execution. |
+| Step | `reme/steps/` | Atomic business operations such as file I/O, retrieval, indexing, and self-evolution. |
+| Component | `reme/components/` | Reusable infrastructure such as file_store, file_graph, keyword_index, and agent_wrapper. |
+| Schema | `reme/schema/` | Data structures such as `Request`, `Response`, `FileChunk`, `FileNode`, and configuration models. |
+| Config | `reme/config/` | Default YAML configuration and command-line override parsing. |
## 2. Directory Structure
@@ -55,11 +55,12 @@ Core layers:
reme/
reme.py # CLI entry point
application.py # Application assembly and lifecycle
+ plugin.py # installed plugin contract and entry-point loader
config/
default.yaml # default service / jobs / components
config_parser.py # config=, dot notation, and env placeholder parsing
components/
- component_registry.py # global registry R
+ component_registry.py # backend registry and application-local copies
base_component.py # ComponentMixin / BaseComponent / bind dependency declarations
runtime_context.py # context for one Job execution
job/ # BaseJob / StreamJob / BackgroundJob / CronJob
@@ -68,18 +69,24 @@ reme/
file_store/ # file-index coordination layer
file_graph/ # wikilink graph
keyword_index/ # BM25 and other keyword indexes
- file_chunker/ # Markdown / default text chunking
+ file_chunker/ # Markdown / JSON / JSONL / generic text chunking
file_catalog/ # change checkpoints
as_llm/, as_embedding/ # model wrappers
- agent_wrapper/ # AgentScope / Claude Code wrappers
+ agent_wrapper/ # AgentScope / Claude Code / Codex wrappers
steps/
base_step.py # BaseStep, Ref, dispatch_steps
- common/ # version, help, health_check, demo
+ common/ # version, help, health_check, status, chat
+ benchmark/ # LongMemEval / BEAM evaluation steps
+ cookbook/ # optional research workflow steps
file_io/ # read/write/edit/delete/move/frontmatter/daily
index/ # watch/init/update/search/traverse
evolve/ # auto_memory, auto_resource, auto_dream, proactive
- transfer/ # upload/download/ingest
- channel/ # MCP channel tools
+ transfer/ # upload/download
+plugins/
+ auto-fin/ # independent example plugin distribution
+integrations/
+ claude_code/ # Claude Code adapter and marketplace
+ hermes_agent/ # Hermes Agent memory-provider adapter
```
The default workspace directories are defined by `ApplicationConfig`:
@@ -87,7 +94,8 @@ The default workspace directories are defined by `ApplicationConfig`:
```text
/
metadata/ # persistent file_store, file_graph, keyword_index, file_catalog, and related state
- session/ # agent sessions and original conversations
+ session/ # source conversations used by memory workflows
+ mem_session/ # generated Agent wrapper sessions and configuration
resource/ # external resources
daily/ # lightly processed memory
digest/ # long-term digest memory
@@ -127,13 +135,13 @@ reme search query="memory" backend=mcp
Configuration parsing supports:
-| Capability | Source | Description |
-|---|---|---|
-| Default configuration | `resolve_app_config()` | Load `reme/config/default.yaml` when `config` is not specified. |
-| Explicit configuration | `config=` | Accept a built-in configuration name or a YAML/JSON file path. |
-| Dot notation | `parse_dot_notation()` | For example, `service.port=8181`. |
-| Environment variables | `_expand_env_vars()` | Support `${VAR}` and `${VAR:-default}`. |
-| Value conversion | `_convert_value()` | Convert bool, int, float, JSON list/dict, and null values automatically. |
+| Capability | Source | Description |
+|------------------------|-------------------------|--------------------------------------------------------------------------|
+| Default configuration | `resolve_app_config()` | Load `reme/config/default.yaml` when `config` is not specified. |
+| Explicit configuration | `config=` | Accept a built-in configuration name or a YAML/JSON file path. |
+| Dot notation | `parse_dot_notation()` | For example, `service.port=8181`. |
+| Environment variables | `_expand_env_vars()` | Support `${VAR}` and `${VAR:-default}`. |
+| Value conversion | `_convert_value()` | Convert bool, int, float, JSON list/dict, and null values automatically. |
### 3.2 Service
@@ -157,22 +165,28 @@ flowchart LR
HTTP service behavior:
-| Job type | HTTP exposure |
-|---|---|
-| Non-`StreamJob` with `enable_serve: true` | `POST /` returning `Response` JSON. |
-| `StreamJob` | `POST /` returning `text/event-stream`. |
-| `enable_serve: false` | No endpoint is registered. |
+| Job type | HTTP exposure |
+|-------------------------------------------|---------------------------------------------------|
+| Non-`StreamJob` with `enable_serve: true` | `POST /` returning `Response` JSON. |
+| `StreamJob` | `POST /` returning `text/event-stream`. |
+| `enable_serve: false` | No endpoint is registered. |
+
+After registering Job endpoints, the HTTP service can also mount the ReMe Studio single-page application. The default is
+`service.web_enabled=true`. Builds are resolved from `service.web_static_dir`, `REME_WEB_STATIC_DIR`, the optional
+`reme-ai-studio` package installed by the `web` and `core` extras, and source-tree locations such as
+`website/dist-static`. If no `index.html` is found, only the frontend is skipped and the Job API remains available. The
+Studio `GET` fallback does not replace existing `POST /` routes.
MCP service behavior:
-| Job type | MCP exposure |
-|---|---|
-| Non-`StreamJob` with `enable_serve: true` | Registered as an MCP tool. |
-| `StreamJob` | Currently skipped and not registered. |
-| `BackgroundJob` | Forces `enable_serve=False` at construction and is never exposed. |
+| Job type | MCP exposure |
+|-------------------------------------------|-------------------------------------------------------------------|
+| Non-`StreamJob` with `enable_serve: true` | Registered as an MCP tool. |
+| `StreamJob` | Currently skipped and not registered. |
+| `BackgroundJob` | Forces `enable_serve=False` at construction and is never exposed. |
-MCP services can inject server-owned arguments with `injected_job_kwargs`; callers cannot override those arguments.
-Set `tool_error_on_failure: true` to expose an unsuccessful ReMe `Response` as an MCP tool error.
+MCP services can inject server-owned arguments with `injected_job_kwargs`; callers cannot override those arguments. Set
+`tool_error_on_failure: true` to expose an unsuccessful ReMe `Response` as an MCP tool error.
## 4. Registry and Dependency Injection
@@ -198,24 +212,50 @@ The registry key is:
`component_type` comes from a class attribute:
-| Type | Class attribute |
-|---|---|
-| Step | `BaseStep.component_type = ComponentEnum.STEP` |
-| Job | `BaseJob.component_type = ComponentEnum.JOB` |
-| Service | `BaseService.component_type = ComponentEnum.SERVICE` |
+| Type | Class attribute |
+|-----------|-----------------------------------------------------------|
+| Step | `BaseStep.component_type = ComponentEnum.STEP` |
+| Job | `BaseJob.component_type = ComponentEnum.JOB` |
+| Service | `BaseService.component_type = ComponentEnum.SERVICE` |
| FileStore | `BaseFileStore.component_type = ComponentEnum.FILE_STORE` |
-The same backend name can therefore exist under different component types. For example, `http` can be both a service backend
-and a client backend.
+The same backend name can therefore exist under different component types. For example, `http` can be both a service
+backend and a client backend.
-### 4.2 Registration Through Module Imports
+`ComponentEnum` provides the built-in identifiers, but installed plugins may declare a new type with a namespaced
+string such as `example.reranker`. Custom identifiers use lowercase letters and numbers separated by `.`, `_`, or `-`.
+They are configured under `components` and participate in the same dependency ordering and lifecycle as built-ins.
-Registration happens when a module is imported. `reme/components/__init__.py` imports component packages, while
-`reme/steps/__init__.py` imports `channel/common/evolve/file_io/index/transfer`. Each package's `__init__.py` then imports
-its concrete modules, causing `@R.register(...)` to execute.
+### 4.2 Built-in and Plugin Registration
-After adding a Step file, make sure the package's `__init__.py` imports it. Otherwise, the backend will not appear in the
-registry.
+Built-in implementations populate the built-in registry through package imports. ReMe freezes that template after
+bootstrap, and each `Application` receives a mutable copy. Runtime code resolves backends through the application's
+registry rather than changing the process-wide template. ReMe then loads only the installed plugins explicitly named by
+`plugins` in the resolved configuration. A plugin exposes its package through the `reme.plugins` Python entry-point
+group. The package's `plugin.yaml` has two optional mappings: `backends` maps registration names to
+`module:Class` targets, and `application_defaults` contributes a low-priority `ApplicationConfig` fragment. The
+entry-point name is the plugin's identity.
+Plugins are enabled explicitly through the application config's `plugins` list or a `plugins=[...]` CLI override.
+Plugin registration therefore stays local to one application;
+duplicate `(component_type, backend)` providers fail during assembly instead of overwriting each other.
+
+The legacy Python `Plugin` descriptor and `reme.configs` entry points remain accepted during migration. Configuration
+files can use `extends` to inherit another built-in, legacy plugin, or file-based configuration. The
+[Auto Fin plugin](../../plugins/auto-fin/README.md) is the current packaging example.
+
+Plugin packages are managed locally and remain separate from per-application activation:
+
+```bash
+reme plugins list
+reme plugins install reme-auto-fin
+reme plugins show auto-fin
+reme plugins validate auto-fin
+reme plugins uninstall auto-fin
+
+reme start plugins='["auto-fin"]'
+```
+
+These management commands use the current Python interpreter's pip and never run through an HTTP or MCP service.
### 4.3 Component.bind
@@ -234,18 +274,18 @@ flowchart LR
Rules for `BaseComponent.bind(name, BaseClass, optional=True)`:
-| Scenario | Behavior |
-|---|---|
-| `name` is empty | Return `None` and skip the dependency. |
-| `app_context` exists | Look up `app_context.components[ctype][name]`. |
-| Dependency missing and `optional=True` | Resolve to `None`. |
-| Dependency missing and `optional=False` | Fail at startup. |
-| Standalone mode | A private component can be created with `default_factory`. |
+| Scenario | Behavior |
+|-----------------------------------------|------------------------------------------------------------|
+| `name` is empty | Return `None` and skip the dependency. |
+| `app_context` exists | Look up `app_context.components[ctype][name]`. |
+| Dependency missing and `optional=True` | Resolve to `None`. |
+| Dependency missing and `optional=False` | Fail at startup. |
+| Standalone mode | A private component can be created with `default_factory`. |
### 4.4 Step.Ref
-Steps do not participate in component topological startup. They are created temporarily for each Job invocation. Steps access
-components primarily through `BaseStep.Ref`:
+Steps do not participate in component topological startup. They are created temporarily for each Job invocation. Steps
+access components primarily through `BaseStep.Ref`:
```python
file_store: BaseFileStore = Ref(BaseFileStore, ComponentEnum.FILE_STORE)
@@ -301,11 +341,13 @@ flowchart LR
F --> G["start CronJob"]
```
-During shutdown, objects in `_started_components` are closed in reverse order so dependents close before their dependencies.
+During shutdown, objects in `_started_components` are closed in reverse order so dependents close before their
+dependencies.
## 6. Job Model
-A Job is the orchestration unit for an externally callable capability or background task. Jobs are configured under `jobs:`
+A Job is the orchestration unit for an externally callable capability or background task. Jobs are configured under
+`jobs:`
in `reme/config/default.yaml`.
### 6.1 BaseJob
@@ -326,23 +368,23 @@ flowchart LR
Important source behavior:
-| Source | Behavior |
-|---|---|
-| `_start()` | Parse each Step config from YAML into `(step_cls, params)`. |
-| `_build_steps()` | Create new Step instances for every call, avoiding state shared across requests. |
-| `__call__()` | Create a `RuntimeContext` and execute Steps sequentially. |
-| Exception handling | Catch the exception, set `response.success=False`, and set `answer=str(e)`. |
+| Source | Behavior |
+|--------------------|----------------------------------------------------------------------------------|
+| `_start()` | Parse each Step config from YAML into `(step_cls, params)`. |
+| `_build_steps()` | Create new Step instances for every call, avoiding state shared across requests. |
+| `__call__()` | Create a `RuntimeContext` and execute Steps sequentially. |
+| Exception handling | Catch the exception, set `response.success=False`, and set `answer=str(e)`. |
### 6.2 StreamJob
`StreamJob` extends `BaseJob` but returns streaming chunks:
-| Behavior | Description |
-|---|---|
-| Context | Includes `stream_queue`. |
+| Behavior | Description |
+|-------------|------------------------------------------------------------|
+| Context | Includes `stream_queue`. |
| Step output | Call `context.add_stream_string(text, ChunkEnum.CONTENT)`. |
-| Exception | Write `ChunkEnum.ERROR`. |
-| Completion | Always send a `DONE` chunk. |
+| Exception | Write `ChunkEnum.ERROR`. |
+| Completion | Always send a `DONE` chunk. |
### 6.3 BackgroundJob
@@ -362,8 +404,8 @@ flowchart LR
J --> K["wait close_timeout; cancel on timeout"]
```
-The default `BackgroundJob.__call__()` also executes configured Steps in sequence, but it does not swallow exceptions, which
-allows the supervisor to restart the task.
+The default `BackgroundJob.__call__()` also executes configured Steps in sequence, but it does not swallow exceptions,
+which allows the supervisor to restart the task.
### 6.4 CronJob
@@ -389,7 +431,9 @@ The current implementation uses `croniter` to calculate the next trigger time. T
```mermaid
flowchart LR
Jobs["default.yaml jobs"] --> BG["background index_update_loop resource_watch_loop digest_watch_loop"]
- Jobs --> Base["base version / help / health_check search / node_search / traverse / reindex read / write / edit / delete / move / list / stat daily_list / daily_reindex / daily_write auto_memory / auto_resource / auto_dream / proactive"]
+ Jobs --> Cron["cron dream_cron optimize_index_cron"]
+ Jobs --> Stream["stream chat"]
+ Jobs --> Base["base version / help / health_check / status / app_config search / node_search / traverse / graph_snapshot / reindex read / load / read_image / write / save / edit / delete / move / list / stat / frontmatter_* daily_list / daily_reindex / daily_write auto_memory / auto_memory_cc / auto_resource / auto_dream / proactive"]
```
## 7. Step Model
@@ -413,12 +457,12 @@ flowchart LR
`RuntimeContext` is shared by all Steps within one Job invocation:
-| Field | Description |
-|---|---|
-| `response` | Final `Response(answer, success, metadata)`. |
-| `data` | Free-form dictionary containing input parameters and intermediate results. |
-| `stream_queue` | Output queue for streaming Jobs. |
-| `stop_event` | Stop signal for background Jobs. |
+| Field | Description |
+|----------------|----------------------------------------------------------------------------|
+| `response` | Final `Response(answer, success, metadata)`. |
+| `data` | Free-form dictionary containing input parameters and intermediate results. |
+| `stream_queue` | Output queue for streaming Jobs. |
+| `stop_event` | Stop signal for background Jobs. |
Common Step code:
@@ -482,21 +526,22 @@ flowchart LR
Current default components in `reme/config/default.yaml`:
-| ComponentEnum | Name | Backend | Description |
-|---|---|---|---|
-| `service` | singleton | `http` | Default HTTP service. |
-| `tokenizer` | `default` | `regex` | BM25 tokenizer. |
-| `as_embedding` | `default` | `${EMBEDDING_BACKEND:-openai}` | Embedding model wrapper. |
-| `embedding_store` | `default` | `local` | Embedding store depending on `as_embedding: default`. |
-| `as_llm` | `default` | `${LLM_BACKEND:-openai}` | LLM model wrapper. |
-| `agent_wrapper` | `default` | `agentscope` | AgentScope wrapper. |
-| `agent_wrapper` | `claude_code` | `claude_code` | Claude Code wrapper. |
-| `file_graph` | `default` | `local` | Wikilink graph. |
-| `file_catalog` | `default/resource/digest/dream` | `local` | File-change checkpoints. |
-| `file_chunker` | `markdown` | `markdown` | Markdown AST chunking. |
-| `file_chunker` | `default` | `default` | Default text chunking, currently supporting `jsonl`. |
-| `keyword_index` | `default` | `bm25` | BM25 keyword index. |
-| `file_store` | `default` | `local` | Combines file_graph and keyword_index; defaults to `embedding_store: ""`. |
+| ComponentEnum | Name | Backend | Description |
+|-------------------|---------------------------------|--------------------------------------------------|--------------------------------------------------------------------------------|
+| `service` | singleton | `http` | Default HTTP service. |
+| `tokenizer` | `default` | `regex` | BM25 tokenizer. |
+| `as_embedding` | `default` | Not configured by default; example uses `openai` | Provides the embedding model wrapper after uncommenting the example config. |
+| `embedding_store` | `default` | Not configured by default; example uses `local` | Depends on `as_embedding: default` after uncommenting the example config. |
+| `as_llm` | `default` | `${LLM_BACKEND:-openai}` | LLM model wrapper. |
+| `agent_wrapper` | `default` | `agentscope` | AgentScope wrapper. |
+| `agent_wrapper` | `claude_code` | `claude_code` | Claude Code wrapper. |
+| `agent_wrapper` | `codex/codex_oauth` | `codex` | Codex wrappers for API-key and OAuth authentication. |
+| `file_graph` | `default` | `local` | Wikilink graph. |
+| `file_catalog` | `default/resource/digest/dream` | `local` | File-change checkpoints. |
+| `file_chunker` | `markdown` | `markdown` | Markdown AST chunking. |
+| `file_chunker` | `json/jsonl/default` | `json/jsonl/default` | JSON, JSONL, and generic text chunkers; generic text supports `txt` and `log`. |
+| `keyword_index` | `default` | `bm25` | BM25 keyword index. |
+| `file_store` | `default` | `local` | Combines file_graph and keyword_index; defaults to `embedding_store: ""`. |
Note that the `search` Step configuration contains `vector_weight`, but `file_store.default.embedding_store` is empty by
default. Vector retrieval is available only when the runtime configuration enables an embedding store.
@@ -554,12 +599,12 @@ class MySearchStep(BaseStep):
Common attributes available directly:
-| Attribute | Component resolved by default |
-|---|---|
-| `self.as_llm` | `.model` from `as_llm: default`. |
+| Attribute | Component resolved by default |
+|----------------------|-------------------------------------|
+| `self.as_llm` | `.model` from `as_llm: default`. |
| `self.agent_wrapper` | `agent_wrapper: default`; optional. |
-| `self.file_catalog` | `file_catalog: default`; optional. |
-| `self.file_store` | `file_store: default`. |
+| `self.file_catalog` | `file_catalog: default`; optional. |
+| `self.file_store` | `file_store: default`. |
To select a non-default component from Job configuration:
@@ -571,13 +616,13 @@ steps:
### 9.4 Step Design Guidance
-| Guidance | Reason |
-|---|---|
-| Read input from `context` and write intermediate results to `context`. | A multi-Step Job passes data through the same context. |
-| Write the final result to `context.response`. | Services and clients consume the standard `Response`. |
-| Do not store request-scoped state on a Step instance. | A Step is rebuilt for every Job call, and stateless Steps are easier to test. |
-| A background loop that supports interruption should check `context.stop_event`. | `BackgroundJob.close()` relies on the stop event for graceful shutdown. |
-| Call `add_stream_string()` only from a StreamJob. | A normal Job has no stream queue. |
+| Guidance | Reason |
+|---------------------------------------------------------------------------------|-------------------------------------------------------------------------------|
+| Read input from `context` and write intermediate results to `context`. | A multi-Step Job passes data through the same context. |
+| Write the final result to `context.response`. | Services and clients consume the standard `Response`. |
+| Do not store request-scoped state on a Step instance. | A Step is rebuilt for every Job call, and stateless Steps are easier to test. |
+| A background loop that supports interruption should check `context.stop_event`. | `BackgroundJob.close()` relies on the stop event for graceful shutdown. |
+| Call `add_stream_string()` only from a StreamJob. | A normal Job has no stream queue. |
### 9.5 Unit Test Example
@@ -600,8 +645,8 @@ async def test_uppercase_step():
## 10. Adding a Job
-A Job usually requires no new Python class; configure existing Steps instead. Add a new Job backend only when a new execution
-model is required.
+A Job usually requires no new Python class; configure existing Steps instead. Add a new Job backend only when a new
+execution model is required.
### 10.1 Adding a Normal Request Job
@@ -736,11 +781,11 @@ jobs:
Characteristics of a background Job:
-| Characteristic | Description |
-|---|---|
-| Not externally exposed | `BackgroundJob.__init__()` forces `enable_serve=False`. |
-| Has a supervisor | Restarts with exponential backoff after an exception by default. |
-| Has a stop event | Notifies the loop to exit during close. |
+| Characteristic | Description |
+|---------------------------------|--------------------------------------------------------------------|
+| Not externally exposed | `BackgroundJob.__init__()` forces `enable_serve=False`. |
+| Has a supervisor | Restarts with exponential backoff after an exception by default. |
+| Has a stop event | Notifies the loop to exit during close. |
| Suitable for watching/consuming | File watching, queue consumption, and periodic long-running loops. |
### 10.5 Adding a Cron Job
@@ -767,14 +812,14 @@ An invalid `cron` expression fails at startup.
Most use cases require only a new Step plus a YAML Job. Consider adding `reme/components/job/*.py` only in these cases:
-| Requirement | New Job class? |
-|---|---|
-| Add a business command | No; use `backend: base`. |
-| Chain existing steps | No; use `steps:`. |
-| Need SSE/streaming output | No; use `backend: stream`. |
-| Need a background loop | No; use `backend: background`. |
-| Need cron scheduling | No; use `backend: cron`. |
-| Need entirely new scheduling, concurrency, or transaction semantics | Yes; add a Job backend. |
+| Requirement | New Job class? |
+|---------------------------------------------------------------------|--------------------------------|
+| Add a business command | No; use `backend: base`. |
+| Chain existing steps | No; use `steps:`. |
+| Need SSE/streaming output | No; use `backend: stream`. |
+| Need a background loop | No; use `backend: background`. |
+| Need cron scheduling | No; use `backend: cron`. |
+| Need entirely new scheduling, concurrency, or transaction semantics | Yes; add a Job backend. |
Minimal shape of a new Job backend:
diff --git a/docs/en/memory_as_file.md b/docs/en/memory_as_file.md
index da41506c..5202887a 100644
--- a/docs/en/memory_as_file.md
+++ b/docs/en/memory_as_file.md
@@ -6,37 +6,39 @@ ReMe's core idea is **Memory as File, File as Memory**.
-**Memory as File**: long-term memory is not hidden in a black-box database. It lives in Markdown files, resource files, and
-index snapshots under the workspace. Users and agents can directly read, write, move, and delete those files.
+**Memory as File**: long-term memory is not hidden in a black-box database. Its source material and readable memories
+live in user-owned files under the workspace. Users and agents can directly read, write, move, and delete those files;
+indexes and snapshots under `metadata/` are derived state that can be rebuilt.
-**File as Memory**: each file is more than ordinary text. It is an indexable, linkable, and evolvable memory node. ReMe parses
-frontmatter, body chunks, and wikilink edges from files and organizes them into retrieval indexes and a graph.
+**File as Memory**: each file is more than ordinary text. It is an indexable, linkable, and evolvable memory node. ReMe
+parses frontmatter, body chunks, and wikilink edges from files and organizes them into retrieval indexes and a graph.
In other words, files are both a human-readable interface and an operational interface for agents. Directory structure
carries the memory layers, while Markdown syntax expresses content, metadata, and relationships.
## Design Goals
-ReMe represents memory as files not merely for convenient storage, but to give long-term memory several essential properties:
+ReMe represents memory as files not merely for convenient storage, but to give long-term memory several essential
+properties:
-| Goal | Meaning |
-|---|---|
-| Readable | Users can open the workspace directly and read daily notes, digest nodes, and source material like ordinary notes. |
-| Editable | Users and agents can correct, extend, move, or delete memory with file operations, without a specialized database client. |
-| Traceable | Long-term conclusions in digest can point back to daily, resource, or session files from a Sources section. |
-| Portable | The workspace is an ordinary directory. Markdown, JSONL, YAML, and resource files can be backed up, synchronized, versioned, or moved to other tools. |
-| Indexable | Although the files are plain text, ReMe parses frontmatter, chunks, and wikilinks to build a retrieval index and file graph. |
-| Collaborative | Humans judge and correct; agents organize, link, and retrieve. Both operate on the same files. |
+| Goal | Meaning |
+|---------------|-------------------------------------------------------------------------------------------------------------------------------------------------------|
+| Readable | Users can open the workspace directly and read daily notes, digest nodes, and source material like ordinary notes. |
+| Editable | Users and agents can correct, extend, move, or delete memory with file operations, without a specialized database client. |
+| Traceable | Long-term conclusions in digest can point back to daily, resource, or session files from a Sources section. |
+| Portable | The workspace is an ordinary directory. Markdown, JSONL, YAML, and resource files can be backed up, synchronized, versioned, or moved to other tools. |
+| Indexable | Although the files are plain text, ReMe parses frontmatter, chunks, and wikilinks to build a retrieval index and file graph. |
+| Collaborative | Humans judge and correct; agents organize, link, and retrieve. Both operate on the same files. |
-ReMe memory is therefore neither a hidden database record nor a prompt fragment visible only to an LLM. It is first a file
-owned by the user and only then indexed by the system for retrieval.
+ReMe memory is therefore neither a hidden database record nor a prompt fragment visible only to an LLM. It is first a
+file owned by the user and only then indexed by the system for retrieval.
## Memory Layers
A ReMe workspace divides memory into four layers:
```text
-raw input -> session/ + resource/
+source records -> session/ + resource/
working memory -> daily/
long memory -> digest/
system state -> metadata/
@@ -44,19 +46,23 @@ system state -> metadata/
Each layer solves a different problem.
-`session/` and `resource/` preserve raw input. Their purpose is to retain the original situation: conversations, agent
-sessions, uploaded material, web pages, and reports remain intact as evidence for later verification.
+`session/` and `resource/` preserve source records. Files under `resource/` remain unchanged at their original path.
+Standard Auto Memory records retain conversation messages while intentionally omitting tool-result and base64 data
+blocks; this keeps recalled output and binary payloads from masquerading as user-provided evidence. Generated Agent
+runtime state instead lives under `mem_session/`.
-`daily/` is the lightly processed layer. It organizes the day's conversations and resources into more readable daily notes:
-what happened, which conclusions were reached, which follow-up tasks remain, and where the source material lives. Daily does
-not aim for final abstraction; it is closer to a workbench for the day.
+`daily/` is the lightly processed layer. It organizes the day's conversations and resources into more readable daily
+notes:
+what happened, which conclusions were reached, which follow-up tasks remain, and where the source material lives. Daily
+does not aim for final abstraction; it is closer to a workbench for the day.
`digest/` is the deeply processed layer. It stores memory nodes that can be reused over time, such as user preferences,
project background, procedural experience, conceptual knowledge, and decision precedents. Digest should not merely copy
daily. It should merge recurring facts, methods, and relationships into more stable descriptions.
-`metadata/` is the system index layer. It stores runtime state such as the file catalog, chunk index, and graph snapshots.
-Users normally do not edit this content manually. The actual editing surface is `daily/`, `digest/`, and, when necessary,
+`metadata/` is the system index layer. It stores runtime state such as the file catalog, chunk index, and graph
+snapshots. Users normally do not edit this content manually. The actual editing surface is `daily/`, `digest/`, and,
+when necessary,
`resource/`.
These layers let ReMe preserve both the original situation and its abstraction: daily reconstructs what happened, while
@@ -73,21 +79,23 @@ The corresponding automatic flows are [Auto Memory](./auto_memory.md), [Auto Res
```text
/
├── metadata/ # system index layer; persistent indexes, graph, catalogs; not a manual editing surface
-├── session/ # raw input layer; original conversations and agent sessions
+├── session/ # source-record layer; source conversations
│ ├── dialog/
-│ │ └── .jsonl # conversation messages saved by auto_memory
-│ ├── agentscope/
-│ │ └── .jsonl
+│ │ └── .jsonl # source messages saved by auto_memory
│ └── claude_code/
-│ └── .jsonl
-├── resource/ # raw input layer; original external material
+│ └── .jsonl # ReMe copy used by auto_memory_cc
+├── mem_session/ # generated Agent wrapper sessions/config, not user memory
+│ ├── agentscope/
+│ ├── claude_config/
+│ └── codex/
+├── resource/ # source-record layer; original external material
+│ ├── . # root-level input uses today's date
│ └── YYYY-MM-DD/
-│ └── .
+│ └── . # dated input uses the directory date
├── daily/ # lightly processed layer; facts, conversation summaries, and resource interpretations by date
│ ├── YYYY-MM-DD.md # index page for the day
│ └── YYYY-MM-DD/
-│ ├── .md # daily note distilled from a conversation
-│ ├── .md # daily note distilled from a resource
+│ ├── .md # topic-named conversation or resource card
│ └── interests.yaml # proactive interest topics generated by auto_dream
└── digest/ # deeply processed layer; reusable personal facts, procedures, and knowledge nodes
├── personal/
@@ -103,17 +111,21 @@ Typical flows:
```text
conversation
-> session/dialog/.jsonl
- -> daily/YYYY-MM-DD/.md
+ -> daily/YYYY-MM-DD/.md
-> digest/personal | digest/procedure | digest/wiki
external resource
- -> resource/YYYY-MM-DD/.
- -> daily/YYYY-MM-DD/.md
+ -> resource/[YYYY-MM-DD/].
+ -> daily/YYYY-MM-DD/.md
-> digest/wiki | digest/procedure
```
-The first two steps focus on recording and organizing; the final step focuses on long-term distillation. `auto_memory` and
-`auto_resource` generate daily notes from raw input, and `auto_dream` extracts and integrates digest nodes from daily.
+The first two steps focus on recording and organizing; the final step focuses on long-term distillation. `auto_memory`
+and
+`auto_resource` generate daily notes from source input, and `auto_dream` extracts and integrates digest nodes from
+daily. The generated daily filename comes from validated frontmatter `name`; `session_id`, `source_conversation`, and
+`source_resource`
+provide stable provenance and lookup identity instead of determining the filename.
## Markdown Format
@@ -146,8 +158,8 @@ source_conversation: [[session/dialog/abc.jsonl]]
---
```
-The current code recognizes `name` and `description` explicitly. Other fields are preserved as additional metadata. The write
-interface merges `name`, `description`, and `metadata` into frontmatter.
+The current code recognizes `name` and `description` explicitly. Other fields are preserved as additional metadata. The
+write interface merges `name`, `description`, and `metadata` into frontmatter.
Treat frontmatter as a node-level summary and the body as evidence, explanation, and relationships. For example:
@@ -165,7 +177,7 @@ Apply this preference when following [[digest/procedure/technical-documentation.
## Sources
-- [[daily/2026-06-20/session-a.md]]
+This preference was recorded in [[daily/2026-06-20/documentation-style.md]], which captures the user's repeated guidance.
```
This has three benefits:
@@ -174,8 +186,8 @@ This has three benefits:
2. The body can carry fuller facts, conditions, counterexamples, and sources.
3. Ordinary wikilinks can be parsed by the graph and maintained when files move.
-Frontmatter is best for stable, short, structured fields; the body is best for explanations meant for people. Do not put long
-body text into YAML fields.
+Frontmatter is best for stable, short, structured fields; the body is best for explanations meant for people. Do not put
+long body text into YAML fields.
### Wikilink
@@ -197,13 +209,13 @@ ReMe wikilinks use **literal path semantics**:
ReMe does not append `.md` automatically, search by filename, or automatically resolve folder notes. Use complete
workspace-relative paths with their extensions.
-Ordinary Markdown links such as `[label](../wiki/example.md)` do not create `FileLink` edges and are not rewritten by move or
-retarget operations.
+Ordinary Markdown links such as `[label](../wiki/example.md)` do not create `FileLink` edges and are not rewritten by
+move or retarget operations.
Anchors such as `#L9`, `#L9-L10`, and `#L9-L10,L15-L20` remain ordinary `target_anchor` strings in the graph. The graph
-parser does not validate line-anchor syntax, so values such as `#L0`, `#L10-L9`, and `#L9,` are also stored. The `read` job
-does not interpret an anchor appended to `path`; use the separate 1-based, inclusive `start_line` and `end_line` arguments to
-read a range, for example `read(path="digest/wiki/solar.md", start_line=9, end_line=10)`.
+parser does not validate line-anchor syntax, so values such as `#L0`, `#L10-L9`, and `#L9,` are also stored. The `read`
+job does not interpret an anchor appended to `path`; use the separate 1-based, inclusive `start_line` and `end_line`
+arguments to read a range, for example `read(path="digest/wiki/solar.md", start_line=9, end_line=10)`.
Wikilinks support these behaviors:
@@ -224,9 +236,8 @@ FileLink
```
Older documents containing wrappers such as `related:: [[path]]`,
-`- related:: [[path]]`, or `[related:: [[path]]]` remain readable. ReMe
-ignores the surrounding text and indexes the inner `[[path]]` as an ordinary
-link. After upgrading from a version that stored typed links, run `reme reindex`
+`- related:: [[path]]`, or `[related:: [[path]]]` remain readable. ReMe ignores the surrounding text and indexes the
+inner `[[path]]` as an ordinary link. After upgrading from a version that stored typed links, run `reme reindex`
once to rebuild the derived graph without the removed relationship field.
### Sources and Relationships
@@ -238,8 +249,8 @@ A Sources section records where a long-term memory came from:
```markdown
## Sources
-- [[daily/2026-06-20/session-a.md]]
-- [[resource/2026-06-20/report.pdf]]
+The preference was observed in [[daily/2026-06-20/documentation-style.md]], and the supporting report evidence is retained in
+[[resource/2026-06-20/report.pdf]].
```
A conceptual relationship link explains which other long-term memories relate to the node. Weave it into natural prose:
@@ -255,16 +266,16 @@ This analysis extends [[digest/wiki/solar-supply-chain.md]], follows
Because memory is stored as files, users can edit the workspace directly, while agents can read and write the same files
through ReMe's file tools. Both follow the same conventions:
-| Operation | Guidance |
-|---|---|
-| Add memory | Write to the appropriate directory, use frontmatter for Markdown, and prefer complete workspace-relative wikilinks. |
-| Edit a body | Preserve existing sources and important wikilinks. When correcting an old conclusion, explain how the new material changes the previous judgment. |
-| Move a file | ReMe's move tool rewrites old paths in inbound edges by default. After a manual move, inspect inbound links again. |
+| Operation | Guidance |
+|---------------|---------------------------------------------------------------------------------------------------------------------------------------------------|
+| Add memory | Write to the appropriate directory, use frontmatter for Markdown, and prefer complete workspace-relative wikilinks. |
+| Edit a body | Preserve existing sources and important wikilinks. When correcting an old conclusion, explain how the new material changes the previous judgment. |
+| Move a file | ReMe's move tool rewrites old paths in inbound edges by default. After a manual move, inspect inbound links again. |
| Delete a file | Check inbound links first. ReMe's delete tool returns source files that still point to the target, making dangling references easier to clean up. |
-| Edit metadata | Use frontmatter for short fields. When the body changes substantially, update `description` as well. |
+| Edit metadata | Use frontmatter for short fields. When the body changes substantially, update `description` as well. |
-A practical rule is: **an agent may rewrite the wording, but it must not lose evidence edges**. In particular, Sources entries
-and existing digest-to-digest wikilinks are the basis for traceable and extensible long-term memory.
+A practical rule is: **an agent may rewrite the wording, but it must not lose evidence edges**. In particular, Sources
+entries and existing digest-to-digest wikilinks are the basis for traceable and extensible long-term memory.
## Path Semantics
@@ -272,7 +283,7 @@ All file tools and wikilinks use workspace-relative paths as their basic unit:
```text
digest/wiki/solar.md
-daily/2026-06-20/session-a.md
+daily/2026-06-20/documentation-style.md
resource/2026-06-20/report.pdf
```
@@ -284,16 +295,16 @@ Recommended practices:
1. Include `.md` when linking a Markdown file.
2. Use the complete source path when linking from digest to daily or resource.
3. Rename or move files through ReMe's move tool whenever possible to avoid stale paths.
-4. Put external source material under `resource/YYYY-MM-DD/...` and long-term abstractions under `digest/...`. Do not put
- raw source material directly into digest.
+4. Put external source material under `resource/YYYY-MM-DD/...` and long-term abstractions under `digest/...`. Do not
+ put raw source material directly into digest.
-Explicit path semantics sacrifice a little convenience when writing by hand, but provide predictability, portability, and
-automatic maintainability.
+Explicit path semantics sacrifice a little convenience when writing by hand, but provide predictability, portability,
+and automatic maintainability.
## Memory Chunking
-Memory chunking divides a file into retrievable fragments. ReMe does not split Markdown at fixed lengths by default; it tries
-to preserve semantic structure.
+Memory chunking divides a file into retrievable fragments. ReMe does not split Markdown at fixed lengths by default; it
+tries to preserve semantic structure.
This section explains how files become retrieval chunks. For index updates, BM25, vector recall, and link expansion, see
[Memory Search](./memory_search.md).
@@ -308,8 +319,8 @@ Document
chunk 1 | chunk 2 | chunk 3 | ...
```
-This is simple, but it can cut headings, tables, code blocks, lists, and `[[wikilinks]]` in the middle. After a match, the
-agent often sees only an isolated fragment without knowing its section or relationship to other memory nodes.
+This is simple, but it can cut headings, tables, code blocks, lists, and `[[wikilinks]]` in the middle. After a match,
+the agent often sees only an isolated fragment without knowing its section or relationship to other memory nodes.
ReMe chunking is closer to splitting memory by file structure:
@@ -375,5 +386,5 @@ Matched body fragment
This lets the agent see not only an isolated paragraph but also its structural position in the source file.
-Non-Markdown files use `DefaultFileChunker` by default. It splits by byte size and preserves a small overlap. For Markdown,
-the chunker also avoids cutting `[[wikilinks]]` in the middle.
+Non-Markdown files use `DefaultFileChunker` by default. It splits by byte size and preserves a small overlap. For
+Markdown, the chunker also avoids cutting `[[wikilinks]]` in the middle.
diff --git a/docs/en/memory_search.md b/docs/en/memory_search.md
index 5f2c3784..bdc4692e 100644
--- a/docs/en/memory_search.md
+++ b/docs/en/memory_search.md
@@ -1,8 +1,10 @@
# Memory Search
-Memory Search is ReMe's memory retrieval entry point. It continuously builds files under `daily/`, `digest/`, and `resource/`
-into a searchable chunk index and wikilink graph. At query time, it first recalls the most relevant fragments and then expands
-context along the bidirectional links of the files containing those fragments.
+Memory Search is ReMe's memory retrieval entry point. The default background loop continuously builds Markdown under
+`daily/` and `digest/` into a searchable chunk index and wikilink graph. At query time, it first recalls the most
+relevant fragments and then expands context along the bidirectional links of the files containing those fragments.
+`reme reindex` has a broader rebuild scope that also scans `resource/` and JSONL; it is intentionally different from the
+live watcher.
@@ -21,14 +23,17 @@ workspace files
## What It Searches
-The default `index_update_loop` watches three memory directories:
+The default `index_update_loop` watches two memory directories:
- `daily_dir`: daily working memory and session memory cards generated by Auto Memory.
- `digest_dir`: long-term distilled digest nodes.
-- `resource_dir`: external resources or imported material.
-The default suffixes are `md` and `jsonl`. Markdown uses the `markdown` chunker, which parses frontmatter, heading structure,
-and `[[wikilinks]]`. JSONL uses the `default` chunker and creates overlapping chunks by byte size.
+The live watcher handles only the `md` suffix. A separate `resource_watch_loop` watches `resource_dir`, and Auto
+Resource turns those inputs into daily cards that enter the live index. When `reme reindex` is run manually, its
+configuration scans
+`daily_dir`, `digest_dir`, and `resource_dir` for `md` and `jsonl`; Markdown uses the `markdown` chunker and JSONL uses
+the
+`jsonl` chunker.
## How the Index Is Built
@@ -39,8 +44,8 @@ The background Job `index_update_loop` maintains the index using configuration f
```yaml
index_update_loop:
backend: background
- watch_dirs: [ daily_dir, digest_dir, resource_dir ]
- watch_suffixes: [ md, jsonl ]
+ watch_dirs: [daily_dir, digest_dir]
+ watch_suffixes: [md]
steps:
- backend: init_changes_step
monitor_type: file_store
@@ -54,9 +59,9 @@ index_update_loop:
`FileNode.st_mtime` values already stored in `file_store`, calculates added, modified, and deleted changes, and passes
`context["changes"]` to `update_index_step`.
-While the service is running, `watch_changes_step` takes over. It uses `watchfiles.awatch()` to watch the same directories,
-groups file events within a quiet window, and uses `coalesce_changes()` to collapse repeated events for the same path into one
-stable batch of changes.
+While the service is running, `watch_changes_step` takes over. It uses `watchfiles.awatch()` to watch the same
+directories, groups file events within a quiet window, and uses `coalesce_changes()` to collapse repeated events for the
+same path into one stable batch of changes.
`update_index_step` performs the actual index writes:
@@ -66,13 +71,14 @@ stable batch of changes.
4. For a deleted file, remove its records from `file_store`, `keyword_index`, and `file_graph`.
5. When changes exist, dump state to `metadata/` so it can be restored on the next startup.
-The Markdown chunker parses YAML frontmatter, heading structure, and wikilinks into `FileNode`, `FileChunk`, and `FileLink`
+The Markdown chunker parses YAML frontmatter, heading structure, and wikilinks into `FileNode`, `FileChunk`, and
+`FileLink`
objects. For detailed chunking rules, see [Memory as File](./memory_as_file.md#memory-chunking).
### Index Optimization
-Both BM25 and the FAISS HNSW vector index use tombstone markers instead of physical removal when deleting nodes;
-too many tombstones degrade search performance. An idle-time optimization mechanism is built in—the `optimize_index_cron`
+Both BM25 and the FAISS HNSW vector index use tombstone markers instead of physical removal when deleting nodes; too
+many tombstones degrade search performance. An idle-time optimization mechanism is built in—the `optimize_index_cron`
scheduled job compacts tombstones and rebuilds indexes during off-peak hours:
```yaml
@@ -100,17 +106,25 @@ file_store:
It combines three kinds of capability:
-| Part | Default state | Purpose |
-|---|---|---|
-| `file_chunks` | Enabled | Store `FileChunk` text, line numbers, scores, and optional embeddings. |
-| `keyword_index.default` | Enabled | BM25 inverted index where chunk ID is the document ID. |
-| `file_graph.default` | Enabled | Store `FileNode` objects and wikilink edges. |
-| `embedding_store` | Disabled | When enabled, generate embeddings for chunks and support vector recall. |
+| Part | Default state | Purpose |
+|-------------------------|---------------|-------------------------------------------------------------------------|
+| `file_chunks` | Enabled | Store `FileChunk` text, line numbers, scores, and optional embeddings. |
+| `keyword_index.default` | Enabled | BM25 inverted index where chunk ID is the document ID. |
+| `file_graph.default` | Enabled | Store `FileNode` objects and wikilink edges. |
+| `embedding_store` | Disabled | When enabled, generate embeddings for chunks and support vector recall. |
Out of the box, search therefore uses primarily BM25 plus link expansion. After setting `embedding_store: default`,
`SearchStep` runs vector and keyword recall together. Additionally, switching the `file_store` `backend` from `local` to
`faiss` upgrades vector retrieval from a linear scan to a FAISS HNSW index, offering faster recall at scale.
+The embedding store accepts `health_check_timeout` for its startup probe. A temporary failure skips the current vector
+backfill while keeping BM25 available; a later successful provider request resumes the missing-vector backfill
+automatically.
+
+Embedded integrations that have already verified a provider can call `resume_embedding(verified=True)`. When changing
+the embedding vector space, pass `rebuild=True`; persisted vectors are invalidated before a serial background rebuild,
+and vector search remains unavailable until the rebuilt vectors are safely persisted.
+
## How to Search
The `search` Job is also configured in `default.yaml`:
@@ -123,10 +137,12 @@ search:
query: string
limit: integer
min_score: number
+ start_date: string
+ end_date: string
steps:
- backend: search_step
vector_weight: 0.7
- candidate_multiplier: 3.0
+ candidate_multiplier: 5.0
expand_links: true
max_links_per_direction: 10
```
@@ -137,11 +153,17 @@ Call it with:
reme search query="recent discussions about indexing" limit=5
```
+Use `start_date` and `end_date` for inclusive `YYYY-MM-DD` filtering:
+
+```bash
+reme search query="index regression" start_date=2026-06-01 end_date=2026-06-20 limit=10
+```
+
`search_step` executes in this order:
```mermaid
flowchart LR
- A["query + limit"] --> B["candidates = limit * candidate_multiplier"]
+ A["query + limit"] --> B["candidates = min(200, limit * candidate_multiplier)"]
B --> C["file_store.vector_search(...)"]
B --> D["file_store.keyword_search(...)"]
C --> E["RRF fusion"]
@@ -152,17 +174,17 @@ flowchart LR
H --> I["Response.answer + metadata"]
```
-If only BM25 has results, the BM25 ranking is returned directly. If only vector search has results, the vector ranking is
-returned directly. When both have results, they are fused with RRF. RRF does not compare BM25 and cosine scores directly; it
-compares ranks in the two result lists:
+If only BM25 has results, the BM25 ranking is returned directly. If only vector search has results, the vector ranking
+is returned directly. When both have results, they are fused with RRF. RRF does not compare BM25 and cosine scores
+directly; it compares ranks in the two result lists:
```text
fused_score = vector_weight / (60 + vector_rank)
+ keyword_weight / (60 + keyword_rank)
```
-The default `vector_weight=0.7` gives semantic recall more weight when embeddings are enabled, while keyword search can still
-promote chunks with exact term matches.
+The default `vector_weight=0.7` gives semantic recall more weight when embeddings are enabled, while keyword search can
+still promote chunks with exact term matches.
## How BM25 Works
@@ -174,12 +196,14 @@ promote chunks with exact term matches.
- The inverted index records which chunks contain each token and its term frequency within each chunk.
- A query scores only the posting lists matching its tokens and returns the highest-scoring chunk IDs.
-When a file changes, `LocalFileStore.upsert()` first removes the BM25 documents corresponding to the file's old `chunk_ids`
+When a file changes, `LocalFileStore.upsert()` first removes the BM25 documents corresponding to the file's old
+`chunk_ids`
and then adds the new chunk text. Deletion is lazy; the index can later be compacted with optimize.
## Progressive Expansion
-"Progressive" in Memory Search does not mean putting the entire repository into one result. Retrieval expands in three layers:
+"Progressive" in Memory Search does not mean putting the entire repository into one result. Retrieval expands in three
+layers:
1. Chunk recall: return only the `limit` most relevant text fragments.
2. File location: each result includes `path:start_line-end_line`. Pass the path and line bounds separately as `path`,
@@ -198,8 +222,8 @@ matched chunk
-> render neighbor path, name, description, and anchor
```
-This keeps search results short while still showing which long-term nodes, resources, or other daily notes a memory connects
-to. If a result is worth pursuing, use `read path=...` to open the source or
+This keeps search results short while still showing which long-term nodes, resources, or other daily notes a memory
+connects to. If a result is worth pursuing, use `read path=...` to open the source or
`traverse path=... depth=2` to continue along the wikilink graph.
## Return Format
@@ -213,7 +237,7 @@ to. If a result is worth pursuing, use `read path=...` to open the source or
Typical text structure:
```text
-========== daily/2026-06-20/session-a.md:12-28 [score=0.0317 keyword=4.8120] ==========
+========== daily/2026-06-20/retrieval-regression.md:12-28 [score=0.0317 keyword=4.8120] ==========
...matched memory fragment...
outlinks (2):
-> digest/indexing.md name="Indexing" description="..."
@@ -221,5 +245,5 @@ Typical text structure:
<- daily/2026-06-19.md name="..."
```
-`counts` reports how many vector and keyword candidates were recalled and how many results were ultimately returned. With
-embeddings disabled by default, `vector` is usually `0` and `hybrid` is `false`.
+`counts` reports how many vector and keyword candidates were recalled and how many results were ultimately returned.
+With embeddings disabled by default, `vector` is usually `0` and `hybrid` is `false`.
diff --git a/docs/en/plugin_management.md b/docs/en/plugin_management.md
new file mode 100644
index 00000000..9575aae8
--- /dev/null
+++ b/docs/en/plugin_management.md
@@ -0,0 +1,228 @@
+# Plugin Management
+
+ReMe plugins are ordinary Python distributions discovered through the `reme.plugins` entry-point group. Installing a
+plugin makes it available to the current Python environment; it does not enable the plugin in every ReMe application.
+
+Keep these two operations separate:
+
+```text
+reme plugins install ... install a package into the current Python environment
+plugins: [auto-fin] enable an installed plugin for one Application
+```
+
+Plugin package management is local-only. It does not run through a ReMe HTTP or MCP service and never edits application
+configuration files automatically.
+
+A typical plugin workflow has three stages:
+
+1. Install ReMe and the plugin distribution.
+2. Configure the plugin's runtime environment as described in the
+ [ReMe environment-variable guide](../../README.md#environment-variables).
+3. Start an Application with the plugin explicitly enabled, for example
+ `reme start plugins='["auto-fin"]'`.
+
+## List installed plugins
+
+```bash
+reme plugins list
+```
+
+The table shows the plugin entry-point name, Python distribution, version, and plugin contract:
+
+```text
+PLUGIN DISTRIBUTION VERSION FORMAT
+-------- ------------- ------- --------
+auto-fin reme-auto-fin 0.1.0 manifest
+```
+
+`manifest` plugins use the current package-level `plugin.yaml` contract. `legacy` plugins use the compatible Python
+descriptor contract.
+
+A manifest separates backend registration from application configuration:
+
+```yaml
+backends:
+ example_step: example_plugin.steps:ExampleStep
+
+application_defaults:
+ jobs:
+ example:
+ backend: base
+ steps:
+ - backend: example_step
+```
+
+`application_defaults` is a partial `ApplicationConfig`. It is kept below the manifest's `backends` namespace because
+backend import declarations are part of plugin discovery and are not application configuration.
+
+Use JSON when another local tool needs structured output:
+
+```bash
+reme plugins list --json
+```
+
+To compare installed plugins with one application config:
+
+```bash
+reme plugins list --config daily_cookbook
+```
+
+The optional `ENABLED` column reflects only the `plugins` list resolved from that config. A command-line override used
+by another running process is not a global enable state.
+
+## Install a plugin package
+
+Install a published distribution:
+
+```bash
+reme plugins install reme-auto-fin
+```
+
+Install or upgrade a pinned version:
+
+```bash
+reme plugins install 'reme-auto-fin==0.1.0'
+reme plugins install reme-auto-fin --upgrade
+```
+
+Install a local plugin project:
+
+```bash
+reme plugins install ./plugins/auto-fin
+```
+
+Use editable mode while developing it:
+
+```bash
+reme plugins install ./plugins/auto-fin --editable
+```
+
+ReMe invokes pip through the same Python interpreter that runs the `reme` command. Pip remains responsible for package
+resolution, downloads, dependency changes, and build execution. Install only packages and local projects you trust.
+
+After installation, confirm the discovered plugin name:
+
+```bash
+reme plugins list
+reme plugins validate auto-fin
+```
+
+## Inspect a plugin
+
+```bash
+reme plugins show auto-fin
+```
+
+For a manifest plugin, the result includes its registered backend names and default Job names. JSON output is also
+available:
+
+```bash
+reme plugins show auto-fin --json
+```
+
+`show` identifies the package contract without constructing a ReMe Application.
+
+## Validate a plugin
+
+Validate an installed plugin:
+
+```bash
+reme plugins validate auto-fin
+```
+
+Validate a local project before installation:
+
+```bash
+reme plugins validate ./plugins/auto-fin
+```
+
+Validation checks the entry point, `plugin.yaml`, backend imports and component types, registry collisions, merged
+`application_defaults`, and the resulting `ApplicationConfig`. Validation imports plugin backend modules, so run it
+only for trusted code.
+
+## Enable a plugin in a service
+
+Installation alone does not load plugin code into an Application. Enable plugins explicitly in configuration:
+
+```yaml
+plugins:
+ - auto-fin
+```
+
+Or add them for one service launch:
+
+```bash
+reme start plugins='["auto-fin"]'
+```
+
+When `config` is omitted, ReMe loads `default.yaml`. The plugin's `application_defaults` are merged below that config,
+so explicit config values and CLI overrides win. This mapping is an `ApplicationConfig` fragment, not a separate
+configuration schema. The plugin backends are registered only in that Application's local registry.
+
+After the default HTTP service starts, access plugin Jobs through ReMe's CLI client or HTTP:
+
+```bash
+reme auto_fin topics="黄金,AI,存储芯片"
+```
+
+```bash
+curl -s http://127.0.0.1:2333/auto_fin \
+ -H 'Content-Type: application/json' \
+ -d '{"topics":"黄金,AI,存储芯片"}'
+```
+
+When the application uses an MCP service, service-enabled plugin Jobs appear as MCP tools instead.
+
+To add the plugin to another application config, select it explicitly:
+
+```bash
+reme start config=daily_cookbook plugins='["auto-fin"]'
+```
+
+## Uninstall a plugin
+
+Use the plugin entry-point name, not necessarily the distribution name:
+
+```bash
+reme plugins uninstall auto-fin
+```
+
+Skip pip's confirmation prompt when needed:
+
+```bash
+reme plugins uninstall auto-fin --yes
+```
+
+ReMe resolves `auto-fin` to the distribution that provides it, such as `reme-auto-fin`. If one distribution provides
+multiple plugin entry points, the command lists the other plugins that will also be removed.
+
+Uninstallation does not rewrite user configuration. Remove the plugin from relevant `plugins` lists yourself;
+otherwise the next Application startup fails explicitly because the configured plugin is no longer installed. Restart
+already-running ReMe processes after installing, upgrading, or uninstalling packages.
+
+## Troubleshooting
+
+### Plugin is installed but unavailable
+
+Check that the `reme` command and pip package share one Python interpreter:
+
+```bash
+reme plugins list
+python -c 'import sys; print(sys.executable)'
+```
+
+Using `reme plugins install` avoids the most common interpreter mismatch because it runs `python -m pip` with ReMe's
+own interpreter.
+
+### Plugin is installed but not loaded
+
+Add its entry-point name to the Application's `plugins` list. ReMe intentionally has no global enable/disable state.
+
+### Startup reports that the plugin is not installed
+
+The active config still enables a missing plugin. Reinstall it or remove the corresponding name from `plugins`.
+
+### Changes are not visible in a running service
+
+Plugin discovery and backend registration happen during Application construction. Restart the service after changing
+installed packages.
diff --git a/docs/en/proactive.md b/docs/en/proactive.md
index 7b13bd6d..19c0b38d 100644
--- a/docs/en/proactive.md
+++ b/docs/en/proactive.md
@@ -1,17 +1,17 @@
# Proactive
-`proactive` is ReMe's interface for reading proactive memory. It does not reanalyze daily notes or call an LLM. It only reads
-the current day's interest topics written by `auto_dream`:
+`proactive` is ReMe's interface for reading proactive memory. It does not reanalyze daily notes or call an LLM. It only
+reads the current day's interest topics written by `auto_dream`:
```text
daily//interests.yaml
```
-A host agent can use it to learn "what is worth proactive attention today," then decide whether to remind the user, ask a
-follow-up question, recommend a next step, or produce a proactive insight.
+A host agent can use it to learn "what is worth proactive attention today," then decide whether to remind the user, ask
+a follow-up question, recommend a next step, or produce a proactive insight.
-`interests.yaml` is generated by the Topics stage of [Auto Dream](./auto_dream.md). `proactive` only reads and exposes the
-result.
+`interests.yaml` is generated by the Topics stage of [Auto Dream](./auto_dream.md). `proactive` only reads and exposes
+the result.
## Configuration
@@ -34,10 +34,10 @@ proactive:
Parameters:
-| Parameter | Purpose |
-|---|---|
-| `date` | Date to read in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
-| `include_content` | Whether to return the raw YAML in the answer and metadata. Defaults to `true`. |
+| Parameter | Purpose |
+|-------------------|-------------------------------------------------------------------------------------------|
+| `date` | Date to read in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
+| `include_content` | Whether to return the raw YAML in the answer and metadata. Defaults to `true`. |
## Input Contract
@@ -67,15 +67,15 @@ When the file is read successfully, `proactive_step` returns `summary` and `topi
`include_content=true`, the answer also contains `content`. The same result fields remain available in standard response
metadata:
-| Field | Description |
-|---|---|
-| `date` | The date actually read. |
-| `path` | `daily//interests.yaml`. |
-| `topics` | Parsed topic list. |
+| Field | Description |
+|-----------|------------------------------------------------------|
+| `date` | The date actually read. |
+| `path` | `daily//interests.yaml`. |
+| `topics` | Parsed topic list. |
| `content` | Raw YAML; returned only when `include_content=true`. |
-| `skipped` | `true` when the file does not exist. |
-| `error` | Read or parse error. |
-| `summary` | Short summary. |
+| `skipped` | `true` when the file does not exist. |
+| `error` | Read or parse error. |
+| `summary` | Short summary. |
When the file exists and parses successfully, the answer is structured data. For example:
@@ -135,21 +135,21 @@ daily notes
The responsibilities are divided as follows. For the complete Extract, Integrate, Topics, and Finish flow, see
[Auto Dream](./auto_dream.md):
-| Module | Responsibility |
-|---|---|
-| `dream_extract_step` | Extract topic candidates from changed daily inputs. |
-| `dream_topics_step` | Deduplicate, select, and write `interests.yaml`. |
-| `proactive_step` | Read `interests.yaml` and expose it to the host agent. |
+| Module | Responsibility |
+|----------------------|--------------------------------------------------------|
+| `dream_extract_step` | Extract topic candidates from changed daily inputs. |
+| `dream_topics_step` | Deduplicate, select, and write `interests.yaml`. |
+| `proactive_step` | Read `interests.yaml` and expose it to the host agent. |
-`proactive` does not modify files, update a catalog, or decide whether the user should be interrupted. It only provides the
-day's topic material. The caller's product policy determines whether, when, and in what tone to push it to the user.
+`proactive` does not modify files, update a catalog, or decide whether the user should be interrupted. It only provides
+the day's topic material. The caller's product policy determines whether, when, and in what tone to push it to the user.
## Failure Modes
-| Scenario | Behavior |
-|---|---|
-| `interests.yaml` does not exist | `success=true`, `skipped=true`, `topics=[]`. |
-| YAML cannot be read or parsed | `success=false`; the answer contains an error summary. |
-| YAML exists but has no valid topics | `success=true`, `topics=[]`. |
+| Scenario | Behavior |
+|-------------------------------------|--------------------------------------------------------|
+| `interests.yaml` does not exist | `success=true`, `skipped=true`, `topics=[]`. |
+| YAML cannot be read or parsed | `success=false`; the answer contains an error summary. |
+| YAML exists but has no valid topics | `success=true`, `topics=[]`. |
Callers should therefore check `success` first, then `skipped`, and finally whether `topics` is empty.
diff --git a/docs/en/quick_start.md b/docs/en/quick_start.md
index d0e31036..bd73f91a 100644
--- a/docs/en/quick_start.md
+++ b/docs/en/quick_start.md
@@ -15,11 +15,17 @@ Install from source:
```bash
git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
-pip install -e ".[core]"
+pip install -e packages/reme_ai_studio -e ".[core]"
+cd website
+npm ci
+npm run build:static
+cd ..
```
-Installing the `core` extra is recommended. The current code imports the AgentScope wrapper, and self-evolving memory also
-depends on it.
+The static build step requires Node.js 22.13 or newer and makes Studio available when running ReMe from the source tree.
+
+Installing the `core` extra is recommended. The current code imports the AgentScope wrapper, and self-evolving memory
+also depends on it.
To use agent workflows such as `auto_memory`, `auto_resource`, and `auto_dream`, configure an LLM:
@@ -51,10 +57,16 @@ reme start service.port=8181
```bash
reme version
reme health_check
-reme list
+reme help
```
-`reme list` lists server actions. Ordinary commands invoke server Jobs over HTTP.
+`reme help` lists server actions. Ordinary commands invoke server Jobs over HTTP.
+
+The base `reme-ai` package does not include frontend assets. Install `reme-ai[web]` or `reme-ai[core]`, then open
+ for ReMe Studio. It uses the same service to
+browse, edit, and search the workspace and inspect the digest wikilink graph. Disable it with
+`service.web_enabled=false`, or provide a custom build with `service.web_static_dir` / `REME_WEB_STATIC_DIR`. The Job
+API still starts if no web build is found.
---
@@ -65,7 +77,8 @@ The default workspace is `.reme/` under the current directory. It is created aut
```text
.reme/
├── metadata/ # persistent indexes, graph, catalogs, and related state
-├── session/ # agent sessions and original conversations
+├── session/ # source conversation records
+├── mem_session/ # generated Agent wrapper sessions/config
├── resource/ # external resources
├── daily/ # daily notes
└── digest/ # long-term memory
@@ -91,12 +104,13 @@ reme write \
description="Example memory for the quick start" \
content="# Quick Start Demo
-ReMe indexes Markdown under the daily, digest, and resource directories.
+The default live watcher indexes Markdown under the daily and digest directories.
Related link: [[digest/wiki/search-demo.md]]"
```
-`path` is relative to the workspace. A missing suffix is automatically completed with `.md`. For Markdown files, `name` and
+`path` is relative to the workspace. A missing suffix is automatically completed with `.md`. For Markdown files, `name`
+and
`description` are written to frontmatter.
The background watcher builds the index automatically. You can also rebuild it manually:
@@ -117,8 +131,8 @@ Read:
reme read path=digest/wiki/quick-start-demo start_line=1 end_line=20
```
-With the default configuration, retrieval is primarily BM25 plus wikilink graph expansion. Vector retrieval is supported by
-the code, but the embedding store is disabled by default. For the full retrieval flow, see
+With the default configuration, retrieval is primarily BM25 plus wikilink graph expansion. Vector retrieval is supported
+by the code, but the embedding store is disabled by default. For the full retrieval flow, see
[Memory Search](./memory_search.md).
---
@@ -132,7 +146,13 @@ reme frontmatter_read path=digest/wiki/quick-start-demo
reme frontmatter_update path=digest/wiki/quick-start-demo metadata='{"tags":["demo"]}'
```
-The name `list` is used by the CLI to list actions, so the file-listing Job must be called over HTTP:
+The file-listing Job can be called directly from the CLI:
+
+```bash
+reme list path=digest recursive=true limit=50
+```
+
+The equivalent HTTP call is:
```bash
curl -s http://127.0.0.1:2333/list \
@@ -161,7 +181,8 @@ reme auto_memory \
memory_hint="Record the user's preference"
```
-After placing external material under `resource/YYYY-MM-DD/`, the default background task watches
+After placing external material under `resource/YYYY-MM-DD/` or directly under `resource/`, the default background task
+watches
`md/txt/json/jsonl/csv/yaml/html`. You can also trigger processing manually:
```bash
@@ -175,7 +196,8 @@ reme auto_dream date=2026-06-20
reme proactive date=2026-06-20
```
-These flows require a working LLM. Without an LLM configuration, start with basic capabilities such as `write`, `read`, and
+These flows require a working LLM. Without an LLM configuration, start with basic capabilities such as `write`, `read`,
+and
`search`.
For more detail, see [Auto Memory](./auto_memory.md), [Auto Resource](./auto_resource.md),
diff --git a/docs/en/reme-blog.md b/docs/en/reme-blog.md
index 516bd7d4..d52b5f8f 100644
--- a/docs/en/reme-blog.md
+++ b/docs/en/reme-blog.md
@@ -14,7 +14,7 @@ That is exactly what ReMe sets out to do.
GitHub: [https://github.com/agentscope-ai/ReMe](https://github.com/agentscope-ai/ReMe)
-Documentation: [https://docs.agentscope.io/reme](https://docs.agentscope.io/reme)
+Documentation: [https://reme.agentscope.io](https://reme.agentscope.io)
@@ -71,7 +71,7 @@ When writing an article, refer to [[digest/procedure/Technical content writing p
## Sources
-- [[daily/2026-08-07/content-discussion.md]]
+This preference was observed in [[daily/2026-08-07/content-discussion.md]], which records the user's writing guidance.
```
Months later, even if you have forgotten the conversation, the agent can still read the preference, find the related process, and follow `Sources` back to the original context.
@@ -90,14 +90,17 @@ For example, you might say in a conversation:
> “Let's not refactor the login module this week. We can do it after the customer demo. Upgrading dependencies directly caused compatibility issues last time, so let's add regression tests first.”
-This short passage contains project status, a time constraint, a lesson from a previous failure, and a next action. Auto Memory extracts these details from the conversation stream and writes them into a daily memory card, while preserving the original conversation in `session/dialog/`.
+This short passage contains project status, a time constraint, a lesson from a previous failure, and a next action. Auto Memory extracts these details from the conversation stream and writes them into a daily memory card, while retaining a source conversation record in `session/dialog/`.
```text
-session/dialog/project-a.jsonl Original conversation, preserving what happened
-daily/2026-08-07/project-a.md Memory card, optimized for reading
-daily/2026-08-07.md Daily index, providing an overview
+session/dialog/project-a.jsonl Source conversation record
+daily/2026-08-07/login-refactor-decision.md Content-named memory card
+daily/2026-08-07.md Daily index, providing an overview
```
+`session_id` remains in the card's frontmatter for stable lookup and provenance; the filename comes from the Agent-generated
+topic/event `name`, so it does not have to match the session ID.
+
The next time the login module comes up, the agent does not need to search through the entire chat history. It can immediately see why the refactor was postponed, what went wrong before, and what should happen next.
It is like having a recorder who is always present—not one that mechanically transcribes every word, but one that organizes what will still matter later.
@@ -136,7 +139,9 @@ Suppose conversations and external materials give you three pieces of informatio
- A project document later confirmed that insufficient Node.js memory was the root cause;
- A third note added that the issue occurs more often in large TypeScript projects.
-Auto Dream scans all changed daily files, merges evidence that points to the same abstraction, keeps only reusable memory units, and writes them into three categories of long-term memory:
+By default, Auto Dream looks at the two most recent days ending at the target date and sends only daily files changed since
+the previous run to extraction. It merges cross-file evidence for the same abstraction and keeps only the strongest reusable
+memories within a default cap of five units, then writes them into three categories of long-term memory:
- `Personal`: preferences, conventions, and constraints specific to a user, team, or project;
- `Procedure`: repeatable processes, methods, and troubleshooting guides;
@@ -159,7 +164,8 @@ follow the “add regression tests first” convention in [[digest/personal/Team
## Sources
-- [[daily/2026-08-07/build-debug.md|Build troubleshooting record]] provides the root cause and applicable scenarios.
+The root cause and applicable scenarios were documented in
+[[daily/2026-08-07/build-debug.md|Build troubleshooting record]].
```
Knowledge evolves and links are created in the same workflow. Relationships are not invisible edges hidden in a graph database; they are readable, editable content in the files themselves. The files can rebuild the graph—the graph never takes control of the files.
@@ -170,7 +176,10 @@ Knowledge evolves and links are created in the same workflow. Relationships are
-Markdown is easy for people to read, but if files are merely piled into directories, agents still struggle to find them quickly. ReMe continuously watches `daily/`, `digest/`, and `resource/`, synchronizing additions, changes, and deletions to a rebuildable index.
+Markdown is easy for people to read, but if files are merely piled into directories, agents still struggle to find them
+quickly. The default live index watches Markdown under `daily/` and `digest/`. A separate resource workflow watches
+`resource/` and turns those files into daily cards that enter the same index. For a full rebuild from existing files,
+`reme reindex` also scans `resource/` and JSONL.
A Markdown file is parsed into:
@@ -299,13 +308,16 @@ That is what ReMe sets out to do: **make memory not only persistent, but continu
## Integrate ReMe with the Agents You Already Use
-ReMe can run as a local memory service accessed through its CLI, HTTP API, or MCP Server, or it can be embedded in a host process through its Python API. Different agents can choose the integration that best fits their runtime environment and share the same local memory workspace when needed.
+ReMe can run as a local memory service accessed through its CLI, HTTP API, or MCP Server, or it can be embedded in a host
+process through its Python API. The default HTTP service can also serve ReMe Studio at the same address for browsing,
+editing, and searching the workspace and inspecting the digest wikilink graph. Different agents can choose the integration
+that best fits their runtime environment and share the same local memory workspace when needed.
| Agent | Recommended integration | Capabilities after integration |
|-------|-------------------------|--------------------------------|
| **QwenPaw** | Embed ReMe in-process through the Python API. | Reuse the host application's lifecycle and model configuration while keeping memories local and file-based. |
-| **Claude Code** | Start the streamable HTTP MCP Service and install [`plugins/claude_code/reme`](../../plugins/claude_code/reme). | MCP memory-recall tools, the `reme-memory` skill, and a Stop hook that automatically records sessions. |
-| **Hermes** | Start the HTTP Service and install [`plugins/hermes_agent`](../../plugins/hermes_agent). | Automatically recall relevant memories before model calls and invoke `auto_memory` asynchronously after each conversation turn. |
+| **Claude Code** | Start the streamable HTTP MCP Service and install [`integrations/claude_code/reme`](../../integrations/claude_code/reme). | MCP memory-recall tools, the `reme-memory` skill, and a Stop hook that automatically records sessions. |
+| **Hermes** | Start the HTTP Service and install [`integrations/hermes_agent`](../../integrations/hermes_agent). | Automatically recall relevant memories before model calls and invoke `auto_memory` asynchronously after each conversation turn. |
| **OpenClaw, Codex, and other CLI-capable agents** | Copy or install [`skills/reme_memory/SKILL.md`](../../skills/reme_memory/SKILL.md). | Search, read, and write memories through the CLI; automatic recording requires the host agent to integrate explicitly with the conversation lifecycle. |
For installation, configuration, and integration demos, see the [README](../../README.md).
diff --git a/docs/en/reme_scene.md b/docs/en/reme_scene.md
index d7ce86d5..c0dcf2a2 100644
--- a/docs/en/reme_scene.md
+++ b/docs/en/reme_scene.md
@@ -54,20 +54,22 @@ session/
daily/
├── 2026-05-18.md
└── 2026-05-18/
- ├── 2026-05-18-close.md
- ├── glencore-q3.md
- ├── cobalt-policy.md
- ├── cathode-trend.md
+ ├── cobalt-supply-risk.md
+ ├── glencore-output-update.md
+ ├── drc-cobalt-policy.md
+ ├── high-nickel-cathode-trend.md
└── interests.yaml # generated after auto_dream
```
The corresponding flow is:
-- `auto_memory` saves the original conversation to `session/dialog/.jsonl`, then asks the agent to write
- important facts to `daily//.md`.
-- `resource_watch_loop` watches text-file changes under `resource/` and triggers `auto_resource_step` to write a
- same-named daily note.
-- `daily_create` maintains `daily/.md` as the index page for that day.
+- `auto_memory` saves a filtered source conversation record to `session/dialog/.jsonl`, then asks the agent to write
+ important facts to a topic-named `daily//.md`. The note keeps `session_id` and
+ `source_conversation` in frontmatter for stable lookup and provenance.
+- `resource_watch_loop` watches text-file changes under `resource/` and triggers `auto_resource_step` to write a daily note
+ with `source_resource`. The agent suggests a content-based filename, which the system sanitizes and de-duplicates; it is
+ not guaranteed to match the resource filename.
+- Auto Memory, Auto Resource, and Auto Dream refresh `daily/.md` after writing.
### Day 1 evening: Auto Dream writes to Digest
@@ -81,8 +83,8 @@ reme auto_dream date=2026-05-18
```text
dream_extract_step
- scan daily/2026-05-18.md and changed files under daily/2026-05-18/
- output units and topics
+ scan the daily window from 2026-05-17 through 2026-05-18 by default
+ output at most 5 units plus topics from changed files
dream_integrate_step
recall existing digest nodes with node_search for each unit
decide CREATE / CORROBORATE / REFINE / CORRECT
@@ -122,7 +124,7 @@ Changes to mining-rights policy in the DRC may affect KFM mine operations and sh
## Sources
-- [[daily/2026-05-18/2026-05-18-close.md]]
+The production decline and policy risk were recorded in [[daily/2026-05-18/cobalt-supply-risk.md]].
```
Note that wikilinks use literal path semantics. Prefer complete workspace-relative paths with the `.md` extension. ReMe
@@ -235,7 +237,7 @@ topics:
reason: The user repeatedly mentioned KFM and cobalt-price risk today
keywords: [cobalt, DRC, CMOC, KFM]
paths:
- - daily/2026-05-18/2026-05-18-close.md
+ - daily/2026-05-18/cobalt-supply-risk.md
```
Call:
@@ -321,7 +323,7 @@ The build stalls near the end. CPU usage is low, but memory keeps growing.
## Sources
-- [[daily/2026-03-10/build-oom-2026-03-10.md]]
+The failed attempts and successful memory adjustment were recorded in [[daily/2026-03-10/build-oom-2026-03-10.md]].
```
Example `digest/personal/code-style.md`:
@@ -374,7 +376,7 @@ and upgrading the minification plugin did not help last time.
- `digest/procedure/` stores both "how to do it" and "which paths failed," letting the agent reuse diagnostic experience.
- `digest/personal/` stores user preferences so the agent can follow the same engineering style across sessions.
-- The original conversation remains under `session/dialog/`; daily records stay traceable, and digest is only the
+- The source conversation record remains under `session/dialog/`; daily records stay traceable, and digest is only the
long-term distilled result.
## Scenario 3: A Personal Second Brain
@@ -423,7 +425,7 @@ At lunch on 2026-04-20, Alice recommended [[digest/wiki/deep-work.md]], a book a
## Sources
-- [[daily/2026-04-20/lunch-with-alice.md]]
+The recommendation was recorded in [[daily/2026-04-20/lunch-with-alice.md]].
```
### An associative recall
diff --git a/docs/figure/auto-dream-and-proactive.svg b/docs/figure/auto-dream-and-proactive.svg
index 0206f496..8ebe99cd 100644
--- a/docs/figure/auto-dream-and-proactive.svg
+++ b/docs/figure/auto-dream-and-proactive.svg
@@ -1,129 +1,117 @@
-