Compare commits

..

No commits in common. "main" and "v0.1.9" have entirely different histories.
main ... v0.1.9

747 changed files with 17480 additions and 124391 deletions

View file

@ -1,97 +0,0 @@
name: Bug report
description: Report reproducible incorrect or unexpected ReMe behavior
title: "[Bug]: "
labels: [bug]
body:
- type: markdown
attributes:
value: |
Thanks for helping improve ReMe. Please remove secrets, API keys, and private memory content before submitting.
- type: textarea
id: description
attributes:
label: Description
description: What happened, and what did you expect instead?
placeholder: Describe the observed and expected behavior.
validations:
required: true
- type: textarea
id: reproduce
attributes:
label: Steps to reproduce
description: Provide the smallest configuration and command sequence that reproduces the problem.
placeholder: |
1. Configure ...
2. Run ...
3. Observe ...
validations:
required: true
- type: textarea
id: config
attributes:
label: Relevant configuration
description: Include only relevant values and redact credentials, tokens, endpoints, and private paths.
render: yaml
- type: textarea
id: logs
attributes:
label: Logs or traceback
description: Paste relevant output after removing secrets and private workspace content.
render: shell
- type: input
id: reme-version
attributes:
label: ReMe version
placeholder: e.g. 0.4.1.8 or a commit SHA
validations:
required: true
- type: input
id: python-version
attributes:
label: Python version
placeholder: e.g. 3.11.9
validations:
required: true
- type: dropdown
id: os
attributes:
label: Operating system
options:
- Linux
- macOS
- Windows
- Other
validations:
required: true
- type: dropdown
id: area
attributes:
label: Affected area
options:
- CLI or configuration
- HTTP, MCP, or local service
- Memory or workspace files
- Search, catalog, graph, or index
- Model or agent integration
- ReMe Studio
- Plugin or external integration
- Packaging or installation
- Other
validations:
required: true
- type: checkboxes
id: safety
attributes:
label: Data safety
options:
- label: I removed credentials and private memory content from this report.
required: true

View file

@ -1,8 +0,0 @@
blank_issues_enabled: false
contact_links:
- name: ReMe documentation
url: https://reme.agentscope.io
about: Read the installation, configuration, and usage guides.
- name: Existing issues
url: https://github.com/agentscope-ai/ReMe/issues
about: Search for existing reports and discussions before opening a new issue.

View file

@ -1,64 +0,0 @@
name: Feature request
description: Propose a focused enhancement to ReMe
title: "[Feature]: "
labels: [enhancement]
body:
- type: textarea
id: problem
attributes:
label: Problem
description: What user problem or limitation should this change address?
validations:
required: true
- type: textarea
id: proposal
attributes:
label: Proposed behavior
description: Describe the desired behavior and its user-visible contract.
validations:
required: true
- type: dropdown
id: area
attributes:
label: Area
options:
- CLI or configuration
- Jobs or steps
- Memory or workspace files
- Search, catalog, graph, or index
- Service or client
- Model or agent integration
- ReMe Studio
- Plugin or external integration
- Documentation
- Other
validations:
required: true
- type: textarea
id: ownership
attributes:
label: Local-first and compatibility considerations
description: Explain any effect on user-owned files, rebuildable state, configuration, schemas, or service interfaces.
- type: textarea
id: alternatives
attributes:
label: Alternatives considered
description: Describe workarounds or alternative designs you considered.
- type: textarea
id: examples
attributes:
label: Example usage
description: Show the proposed CLI, configuration, API, or UI behavior when useful.
render: shell
- type: checkboxes
id: contribution
attributes:
label: Contribution
options:
- label: I am willing to help implement or test this feature.

View file

@ -1,53 +0,0 @@
name: Usage question
description: Ask for help using or configuring ReMe
title: "[Question]: "
labels: [question]
body:
- type: markdown
attributes:
value: Please check the documentation and existing issues before asking a new question.
- type: textarea
id: goal
attributes:
label: What are you trying to achieve?
validations:
required: true
- type: textarea
id: attempted
attributes:
label: What have you tried?
description: Include relevant commands or configuration, with secrets and private memory content removed.
validations:
required: true
- type: input
id: reme-version
attributes:
label: ReMe version
placeholder: e.g. 0.4.1.8 or a commit SHA
- type: dropdown
id: area
attributes:
label: Area
options:
- Installation
- Configuration
- CLI or service usage
- Memory and workspace management
- Search and retrieval
- ReMe Studio
- Plugin or integration
- Other
- type: checkboxes
id: checked
attributes:
label: Before submitting
options:
- label: I checked the [ReMe documentation](https://reme.agentscope.io) and searched existing issues.
required: true
- label: I removed credentials and private memory content.
required: true

View file

@ -1,35 +0,0 @@
## Summary
<!-- Explain the problem and the smallest coherent change that addresses it. -->
## Related issue
<!-- Use "Fixes #123" when applicable. -->
## Contract and data impact
- [ ] No public configuration, schema, CLI, endpoint, streaming, or workspace-layout contract changes
- [ ] No user-owned memory files are deleted or rewritten
- [ ] Derived indexes, catalogs, graphs, caches, and metadata remain rebuildable
<!-- If any item is unchecked, describe the impact and migration or recovery path. -->
## Validation
<!-- List the exact checks run and their results. Explain relevant checks that were not run. -->
- [ ] Focused tests pass
- [ ] Unit tests pass, or omitted tests are explained below
- [ ] `pre-commit run --all-files` passes, or omitted checks are explained below
- [ ] Frontend checks were run when `reme_studio/` changed
## Checklist
- [ ] I reviewed the diff for unrelated changes and sensitive data
- [ ] Tests cover intentional behavior changes
- [ ] Defaults, schemas, and concise documentation were updated together when required
- [ ] Long-lived clients, tasks, services, and executors follow the application lifecycle
## Screenshots or additional notes
<!-- Include UI screenshots, compatibility notes, or follow-up work when relevant. -->

View file

@ -1,58 +0,0 @@
name: _Build documentation
on:
workflow_call:
inputs:
run_tests:
description: Run the documentation test suite before building
required: false
default: true
type: boolean
upload_pages_artifact:
description: Upload the build for a later GitHub Pages deployment job
required: false
default: false
type: boolean
permissions:
contents: read
jobs:
build:
name: Build documentation
runs-on: ubuntu-latest
defaults:
run:
working-directory: github-pages
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Node
uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: '22.22.3'
cache: npm
cache-dependency-path: github-pages/package-lock.json
- name: Install dependencies
run: npm ci
- name: Run tests
if: inputs.run_tests
run: npm test
- name: Build documentation
run: npm run build
- name: Configure Pages
if: inputs.upload_pages_artifact
uses: actions/configure-pages@45bfe0192ca1faeb007ade9deae92b16b8254a0d # v6
- name: Upload Pages artifact
if: inputs.upload_pages_artifact
uses: actions/upload-pages-artifact@7b1f4a764d45c48632c6b24a0339c27f5614fb0b # v4
with:
path: github-pages/dist

View file

@ -1,88 +0,0 @@
name: _Build Python packages
on:
workflow_call:
inputs:
expected_version:
description: Expected release version; omit for a consistency-only check
required: false
default: ''
type: string
upload_artifacts:
description: Upload distributions for later publish jobs
required: false
default: false
type: boolean
permissions:
contents: read
jobs:
distributions:
name: Build Python distributions
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: '3.11'
- name: Install build dependencies
run: |
python -m pip install --upgrade pip
python -m pip install build packaging pytest twine
- name: Validate package versions
if: inputs.expected_version == ''
run: python scripts/bump_version.py --check
- name: Validate release version
if: inputs.expected_version != ''
env:
EXPECTED_VERSION: ${{ inputs.expected_version }}
run: python scripts/bump_version.py --check --expected-version "${EXPECTED_VERSION}"
- name: Run package tests
run: PYTHONPATH=. python -m pytest tests/unit/test_package_versions.py -q
- name: Build and check distributions
run: |
mkdir -p dist/reme
python -m build --outdir dist/reme
python -m twine check dist/reme/*
- name: Verify distributions and isolated installation
run: |
REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
python -m zipfile -l "${REME_WHEEL}" | (! grep 'reme/web/')
python -m zipfile -l "${REME_WHEEL}" | (! grep 'reme_studio/')
python -m venv "${RUNNER_TEMP}/reme-package-smoke"
"${RUNNER_TEMP}/reme-package-smoke/bin/python" -m pip install "${REME_WHEEL}[as]"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-package-smoke/bin/python" -c "import reme"
- name: Verify released core dependencies
if: inputs.expected_version != ''
run: |
REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
python -m venv "${RUNNER_TEMP}/reme-core-package-smoke"
"${RUNNER_TEMP}/reme-core-package-smoke/bin/python" -m pip install "${REME_WHEEL}[core]"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-core-package-smoke/bin/python" - <<'PY'
from reme_studio import static_dir
assert (static_dir() / "index.html").is_file()
PY
- name: Upload ReMe distributions
if: inputs.upload_artifacts
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: reme-distributions
path: dist/reme/
if-no-files-found: error

View file

@ -1,48 +0,0 @@
name: CI / Documentation
on:
push:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-docs.yml'
- '.github/workflows/_build-docs.yml'
- 'AGENTS.md'
- 'README.md'
- 'README_ZH.md'
- 'docs/**'
- 'github-pages/**'
- 'reme_studio/README*.md'
- 'reme_studio/public/og.jpg'
- 'typescript/README*.md'
- 'plugins/*/README*.md'
- 'benchmark/*/README*.md'
pull_request:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-docs.yml'
- '.github/workflows/_build-docs.yml'
- 'AGENTS.md'
- 'README.md'
- 'README_ZH.md'
- 'docs/**'
- 'github-pages/**'
- 'reme_studio/README*.md'
- 'reme_studio/public/og.jpg'
- 'typescript/README*.md'
- 'plugins/*/README*.md'
- 'benchmark/*/README*.md'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
documentation:
name: Test and build documentation
uses: ./.github/workflows/_build-docs.yml
with:
run_tests: true

View file

@ -1,40 +0,0 @@
name: CI / Python packages
on:
push:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-packages.yml'
- '.github/workflows/_build-python-packages.yml'
- '.github/workflows/release-python.yml'
- 'pyproject.toml'
- 'README.md'
- 'reme/**'
- 'scripts/bump_version.py'
- 'tests/unit/test_package_versions.py'
- 'LICENSE'
pull_request:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-packages.yml'
- '.github/workflows/_build-python-packages.yml'
- '.github/workflows/release-python.yml'
- 'pyproject.toml'
- 'README.md'
- 'reme/**'
- 'scripts/bump_version.py'
- 'tests/unit/test_package_versions.py'
- 'LICENSE'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
distributions:
name: Build and verify distributions
uses: ./.github/workflows/_build-python-packages.yml

View file

@ -1,40 +0,0 @@
name: CI / Python quality
on:
push:
pull_request:
workflow_dispatch:
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
pre-commit:
name: Pre-commit
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Setup Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: '3.11'
cache: pip
- name: Update setuptools
run: |
pip install -U setuptools wheel
- name: Install
run: |
pip install -q -e reme_studio -e ".[dev,core]"
pip install -q --no-deps -e plugins/auto-fin -e plugins/daily_paper
- name: Pre-commit starts
run: pre-commit run --all-files

View file

@ -1,54 +0,0 @@
name: CI / Python tests
on:
push:
branches: [main, master, dev, develop]
pull_request:
branches: [main, master, dev, develop]
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
unit-tests:
name: Unit Tests - py${{ matrix.python-version }}
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
- name: Install dependencies
run: |
python -m pip install --upgrade pip setuptools wheel
pip install -e reme_studio -e ".[dev,core]"
pip install --no-deps -e plugins/auto-fin
pip install -e plugins/daily_paper
pip install coverage
- name: Run unit tests
run: |
coverage run -m pytest tests/unit plugins/auto-fin plugins/daily_paper \
-v \
--tb=long \
-s \
--log-cli-level=WARNING
- name: Generate coverage report
run: coverage report -m

View file

@ -1,90 +0,0 @@
name: CI / ReMe Studio
on:
push:
paths:
- "reme_studio/**"
- ".github/workflows/ci-reme-studio.yml"
- ".github/workflows/release-reme-studio.yml"
- "scripts/package_studio.py"
- "tests/unit/test_package_versions.py"
- "pyproject.toml"
- "LICENSE"
pull_request:
paths:
- "reme_studio/**"
- ".github/workflows/ci-reme-studio.yml"
- ".github/workflows/release-reme-studio.yml"
- "scripts/package_studio.py"
- "tests/unit/test_package_versions.py"
- "pyproject.toml"
- "LICENSE"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
studio:
name: Studio checks
runs-on: ubuntu-latest
defaults:
run:
working-directory: reme_studio
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Setup Node
uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: "22.22.3"
cache: npm
cache-dependency-path: reme_studio/package-lock.json
- name: Install dependencies
run: npm ci
- name: Run format check
run: npm run format:check
- name: Run lint
run: npm run lint
- name: Run tests
run: npm test
- name: Verify npm package
run: |
npm pack --pack-destination "${RUNNER_TEMP}"
tar -tzf "${RUNNER_TEMP}"/agentscope-ai-reme_studio-*.tgz | grep '^package/dist-static/index.html$'
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.11"
- name: Build and verify Python package
working-directory: .
run: |
python -m pip install build packaging pytest twine
PYTHONPATH=. python -m pytest tests/unit/test_package_versions.py -q
python scripts/package_studio.py
python -m build reme_studio --outdir dist/studio
python -m twine check dist/studio/*
STUDIO_WHEEL="$(pwd)/$(ls dist/studio/reme_studio-*.whl)"
python -m venv "${RUNNER_TEMP}/reme-studio-package-smoke"
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" -m pip install "${STUDIO_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" - <<'PY'
from reme_studio import static_dir
assert (static_dir() / "index.html").is_file()
PY

View file

@ -1,51 +0,0 @@
name: CI / TypeScript integrations
on:
push:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-typescript.yml'
- '.github/workflows/release-typescript.yml'
- 'typescript/**'
pull_request:
branches: [main, master, dev, develop]
paths:
- '.github/workflows/ci-typescript.yml'
- '.github/workflows/release-typescript.yml'
- 'typescript/**'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
package:
name: Type-check, test, and pack
runs-on: ubuntu-latest
defaults:
run:
working-directory: typescript
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: '22.22.3'
cache: npm
cache-dependency-path: typescript/package-lock.json
- run: npm ci
- run: npm run format:check
- run: npm run lint
- run: npm run typecheck
- run: npm test
- run: npm run test:package
- name: Validate OpenClaw package contract
run: npx --yes clawhub@0.23.3 package validate . --json

View file

@ -1,51 +0,0 @@
name: CI / Windows
on:
push:
branches: [main, master, dev, develop]
pull_request:
branches: [main, master, dev, develop]
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
cli-smoke:
name: CLI smoke - py${{ matrix.python-version }}
runs-on: windows-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.11"]
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: ${{ matrix.python-version }}
cache: 'pip'
- name: Install package
run: |
python -m pip install --upgrade pip setuptools wheel
pip install -e ".[dev,as]"
- name: Run version job
run: reme start config=tests/fixtures/config/version-smoke.yaml job=version
- name: Run Windows path tests
run: |
python -m pytest `
tests/unit/test_auto_dream.py::test_scan_day_files_includes_nested_md_and_excludes_interests `
tests/unit/test_auto_dream.py::test_dream_extract_matches_posix_catalog_paths `
tests/unit/test_read_with_neighbors.py::test_read_with_neighbors_uses_posix_nested_path `
-v

View file

@ -1,52 +0,0 @@
name: Deploy / Documentation
on:
push:
branches: [main]
paths:
- "github-pages/**"
- "docs/**"
- "README.md"
- "README_ZH.md"
- "reme_studio/README*.md"
- "reme_studio/public/og.jpg"
- "typescript/README*.md"
- "plugins/*/README*.md"
- "benchmark/*/README*.md"
- "AGENTS.md"
- ".github/workflows/deploy-docs.yml"
- ".github/workflows/_build-docs.yml"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: pages
cancel-in-progress: true
jobs:
build:
name: Build documentation
uses: ./.github/workflows/_build-docs.yml
with:
run_tests: true
upload_pages_artifact: true
permissions:
contents: read
pages: write
id-token: write
deploy:
environment:
name: github-pages
url: ${{ steps.deployment.outputs.page_url }}
runs-on: ubuntu-latest
needs: build
permissions:
pages: write
id-token: write
steps:
- name: Deploy
id: deployment
uses: actions/deploy-pages@cd2ce8fcbc39b97be8ca5fce6e763baed58fa128 # v5

View file

@ -1,40 +0,0 @@
name: Policy / PR title
on:
pull_request:
branches: [main, master, dev, develop]
types: [opened, edited, synchronize, reopened]
permissions:
contents: read
pull-requests: read
jobs:
check-pr-title:
runs-on: ubuntu-latest
steps:
- name: Check PR title format
uses: amannn/action-semantic-pull-request@48f256284bd46cdaab1048c3721360e808335d50 # v6.1.1
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
with:
types: |
feat
fix
docs
ci
refactor
test
chore
perf
style
build
revert
requireScope: false
scopePattern: ^[a-z0-9_-]+$
scopePatternError: |
The scope must contain only lowercase letters, numbers, hyphens, and underscores.
Example: "feat(memory): add redis cache support"
validateSingleCommit: false
ignoreLabels: |
ignore-semantic-pull-request

40
.github/workflows/python-publish.yml vendored Normal file
View file

@ -0,0 +1,40 @@
# This workflow will upload a Python Package using Twine when a release is created
# For more information see: https://docs.github.com/en/actions/automating-builds-and-tests/building-and-testing-python#publishing-to-package-registries
# This workflow uses actions that are not certified by GitHub.
# They are provided by a third-party and are governed by
# separate terms of service, privacy policy, and support
# documentation.
name: Publish Python Package to Pypi
on:
workflow_dispatch:
release:
types: [published]
permissions:
contents: read
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install setuptools wheel build
- name: Build package
run: python -m build
- name: Publish package to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
with:
user: __token__
password: ${{ secrets.PYPI_API_TOKEN }}

View file

@ -1,157 +0,0 @@
# 发布操作手册:
# 1. 先将 plugins/auto-fin/pyproject.toml 中的 project.version 更新为待发布版本并合入目标分支。
# 2. 确认插件依赖的 reme-ai 版本已经发布到 PyPI本工作流会在构建阶段验证该依赖可下载。
# 3. 确认 PyPI Trusted Publisher 已绑定本仓库、此工作流和 pypi environment且 PyPI 上不存在相同版本。
# 4. 在 GitHub 仓库的 Actions 页面选择“Release / Auto Fin plugin”点击“Run workflow”。
# 5. 输入与 project.version 完全一致的版本号(例如 0.1.0)后运行;版本也可以带 v 前缀。
#
# 推荐发布顺序reme-ai -> reme-auto-fin -> QwenPaw 更新依赖并通过 plugins: [auto-fin] 启用。
# 当前仅支持 workflow_dispatch 手动触发,不会因 push、tag 或 release 自动发布。
name: Release / Auto Fin plugin
run-name: Publish reme-auto-fin ${{ inputs.version }}
on:
workflow_dispatch:
inputs:
version:
description: Version from plugins/auto-fin/pyproject.toml (for example, 0.1.0)
required: true
type: string
permissions:
contents: read
concurrency:
group: publish-reme-auto-fin
cancel-in-progress: false
jobs:
build:
runs-on: ubuntu-latest
env:
RELEASE_VERSION: ${{ inputs.version }}
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: '3.11'
- name: Install test and build dependencies
run: |
python -m pip install --upgrade pip
python -m pip install build packaging pytest pytest-asyncio twine
python -m pip install -e ".[core]"
python -m pip install --no-deps -e plugins/auto-fin
- name: Validate package name and release version
id: package
run: |
python - "${RELEASE_VERSION}" <<'PY'
import os
import sys
import tomllib
from pathlib import Path
from packaging.requirements import Requirement
from packaging.version import Version
project = tomllib.loads(Path("plugins/auto-fin/pyproject.toml").read_text(encoding="utf-8"))["project"]
expected = Version(sys.argv[1].removeprefix("v"))
actual = Version(project["version"])
if project["name"] != "reme-auto-fin":
raise SystemExit(f"Expected project name 'reme-auto-fin', found {project['name']!r}")
if actual != expected:
raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
requirements = [requirement for requirement in project["dependencies"] if requirement.startswith("reme-ai")]
if len(requirements) != 1:
raise SystemExit(f"Expected one reme-ai dependency, found {requirements!r}")
reme_requirement = Requirement(requirements[0])
if reme_requirement.name != "reme-ai" or reme_requirement.extras:
raise SystemExit(f"Expected a base reme-ai dependency, found {requirements[0]!r}")
if Version("0.4.1.8") in reme_requirement.specifier or Version("0.4.1.9") not in reme_requirement.specifier:
raise SystemExit(f"Expected reme-ai>=0.4.1.9, found {requirements[0]!r}")
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
print(f"reme_requirement={reme_requirement}", file=output)
print(f"Publishing {project['name']} {actual}")
PY
- name: Run Auto Fin tests
run: python -m pytest plugins/auto-fin -q
- name: Require the plugin-enabled ReMe release on PyPI
env:
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
run: |
python -m pip download --no-deps \
--dest "${RUNNER_TEMP}/reme-auto-fin-base" \
"${REME_REQUIREMENT}"
- name: Build and check distributions
run: |
mkdir -p dist/auto-fin
python -m build plugins/auto-fin --outdir dist/auto-fin
python -m twine check dist/auto-fin/*
- name: Verify distributions and isolated installation
run: |
AUTO_FIN_WHEEL="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.whl)"
AUTO_FIN_SDIST="$(pwd)/$(ls dist/auto-fin/reme_auto_fin-*.tar.gz)"
python -m zipfile -l "${AUTO_FIN_WHEEL}" | grep 'dist-info/licenses/LICENSE'
python -m tarfile -l "${AUTO_FIN_SDIST}" | grep '/LICENSE'
python -m venv "${RUNNER_TEMP}/reme-auto-fin-smoke"
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" -m pip install \
"agentscope[model-ollama]==2.0.7" "${AUTO_FIN_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" - <<'PY'
from importlib.metadata import distribution
from reme.plugin_manifest import load_package_manifest
package = distribution("reme-auto-fin")
plugins = {entry.name: entry for entry in package.entry_points if entry.group == "reme.plugins"}
assert plugins["auto-fin"].value == "reme_auto_fin"
manifest = load_package_manifest("reme_auto_fin", plugin_name="auto-fin")
assert set(manifest.backends) == {
"auto_fin_data_step",
"auto_fin_topic_step",
"auto_fin_merge_step",
}
assert set(manifest.application_defaults["jobs"]) == {
"auto_fin",
"auto_fin_cron",
}
PY
- name: Upload distributions
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: reme-auto-fin-${{ inputs.version }}
path: dist/auto-fin/
if-no-files-found: error
publish:
needs: build
runs-on: ubuntu-latest
environment: pypi
permissions:
contents: read
id-token: write
steps:
- name: Download distributions
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: reme-auto-fin-${{ inputs.version }}
path: dist/auto-fin
- name: Publish reme-auto-fin
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
with:
packages-dir: dist/auto-fin

View file

@ -1,157 +0,0 @@
# Release checklist:
# 1. Update project.version in plugins/daily_paper/pyproject.toml and merge it into the target branch.
# 2. Publish the required reme-ai version before this plugin; the build verifies that dependency on PyPI.
# 3. Configure PyPI Trusted Publishing for this repository/workflow and its pypi environment.
# 4. Run "Release / Daily Paper plugin" from GitHub Actions with the exact project version (a v prefix is accepted).
#
# Recommended order: reme-ai -> reme-daily-paper -> downstream applications enabling plugins: [daily-paper].
# This workflow is intentionally manual and never publishes from a push, tag, or GitHub release event.
name: Release / Daily Paper plugin
run-name: Publish reme-daily-paper ${{ inputs.version }}
on:
workflow_dispatch:
inputs:
version:
description: Version from plugins/daily_paper/pyproject.toml (for example, 0.1.0)
required: true
type: string
permissions:
contents: read
concurrency:
group: publish-reme-daily-paper
cancel-in-progress: false
jobs:
build:
runs-on: ubuntu-latest
env:
RELEASE_VERSION: ${{ inputs.version }}
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: '3.11'
- name: Install test and build dependencies
run: |
python -m pip install --upgrade pip
python -m pip install build packaging pytest pytest-asyncio twine
python -m pip install -e ".[core]"
python -m pip install -e plugins/daily_paper
- name: Validate package name, dependencies, and release version
id: package
run: |
python - "${RELEASE_VERSION}" <<'PY'
import os
import sys
import tomllib
from pathlib import Path
from packaging.requirements import Requirement
from packaging.version import Version
project = tomllib.loads(Path("plugins/daily_paper/pyproject.toml").read_text(encoding="utf-8"))["project"]
expected = Version(sys.argv[1].removeprefix("v"))
actual = Version(project["version"])
if project["name"] != "reme-daily-paper":
raise SystemExit(f"Expected project name 'reme-daily-paper', found {project['name']!r}")
if actual != expected:
raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
requirements = [Requirement(value) for value in project["dependencies"]]
reme_requirements = [requirement for requirement in requirements if requirement.name == "reme-ai"]
if len(reme_requirements) != 1 or reme_requirements[0].extras:
raise SystemExit(f"Expected one base reme-ai dependency, found {reme_requirements!r}")
if Version("0.4.1.8") in reme_requirements[0].specifier or Version("0.4.1.9") not in reme_requirements[0].specifier:
raise SystemExit(f"Expected reme-ai>=0.4.1.9, found {reme_requirements!r}")
if sum(requirement.name == "pypdf" for requirement in requirements) != 1:
raise SystemExit("Expected exactly one pypdf dependency")
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
print(f"reme_requirement={reme_requirements[0]}", file=output)
print(f"Publishing {project['name']} {actual}")
PY
- name: Run Daily Paper tests
run: python -m pytest plugins/daily_paper -q
- name: Require the plugin-enabled ReMe release on PyPI
env:
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
run: |
python -m pip download --no-deps \
--dest "${RUNNER_TEMP}/reme-daily-paper-base" \
"${REME_REQUIREMENT}"
- name: Build and check distributions
run: |
mkdir -p dist/daily-paper
python -m build plugins/daily_paper --outdir dist/daily-paper
python -m twine check dist/daily-paper/*
- name: Verify distributions and isolated installation
run: |
DAILY_PAPER_WHEEL="$(pwd)/$(ls dist/daily-paper/reme_daily_paper-*.whl)"
DAILY_PAPER_SDIST="$(pwd)/$(ls dist/daily-paper/reme_daily_paper-*.tar.gz)"
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'reme_daily_paper/plugin.yaml'
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'reme_daily_paper/analyze.yaml'
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'dist-info/licenses/LICENSE'
python -m tarfile -l "${DAILY_PAPER_SDIST}" | grep '/LICENSE'
python -m venv "${RUNNER_TEMP}/reme-daily-paper-smoke"
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" -m pip install \
"agentscope[model-ollama]==2.0.7" "${DAILY_PAPER_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" - <<'PY'
from importlib.metadata import distribution
from reme.plugin_manifest import load_package_manifest
package = distribution("reme-daily-paper")
plugins = {entry.name: entry for entry in package.entry_points if entry.group == "reme.plugins"}
assert plugins["daily-paper"].value == "reme_daily_paper"
manifest = load_package_manifest("reme_daily_paper", plugin_name="daily-paper")
assert set(manifest.backends) == {
"daily_paper_collect_step",
"daily_paper_rank_step",
"daily_paper_select_step",
"daily_paper_analyze_step",
"daily_paper_digest_step",
}
assert set(manifest.application_defaults["jobs"]) == {"daily_paper", "daily_paper_cron"}
PY
- name: Upload distributions
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: reme-daily-paper-${{ inputs.version }}
path: dist/daily-paper/
if-no-files-found: error
publish:
needs: build
runs-on: ubuntu-latest
environment: pypi
permissions:
contents: read
id-token: write
steps:
- name: Download distributions
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: reme-daily-paper-${{ inputs.version }}
path: dist/daily-paper
- name: Publish reme-daily-paper
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
with:
packages-dir: dist/daily-paper

View file

@ -1,47 +0,0 @@
name: Release / Python packages
# Configure a PyPI Trusted Publisher for this repository, workflow, and its
# pypi environment before running the manual release.
on:
workflow_dispatch:
inputs:
version:
description: Release version
required: true
type: string
permissions:
contents: read
concurrency:
group: publish-reme-ai
cancel-in-progress: false
jobs:
build:
name: Build and verify distributions
uses: ./.github/workflows/_build-python-packages.yml
with:
expected_version: ${{ inputs.version }}
upload_artifacts: true
publish-reme:
needs: build
runs-on: ubuntu-latest
environment: pypi
permissions:
contents: read
id-token: write
steps:
- name: Download ReMe distributions
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: reme-distributions
path: dist/reme
- name: Publish ReMe
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
with:
packages-dir: dist/reme
skip-existing: true

View file

@ -1,158 +0,0 @@
# Release checklist:
# 1. Update reme_studio/pyproject.toml, package.json, and package-lock.json to the same Studio version.
# 2. Configure npm Trusted Publishing and PyPI Trusted Publishing with the pypi environment.
# 3. Run this workflow manually with the exact Studio version.
name: Release / ReMe Studio
run-name: Publish ReMe Studio ${{ inputs.version }} (${{ inputs.npm_tag }})
on:
workflow_dispatch:
inputs:
version:
description: Version from the Studio Python and npm manifests
required: true
type: string
npm_tag:
description: npm distribution tag
required: true
default: latest
type: choice
options:
- next
- latest
permissions:
contents: read
concurrency:
group: publish-reme-studio
cancel-in-progress: false
jobs:
build:
runs-on: ubuntu-latest
env:
RELEASE_VERSION: ${{ inputs.version }}
NPM_TAG: ${{ inputs.npm_tag }}
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: "22.22.3"
cache: npm
cache-dependency-path: reme_studio/package-lock.json
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.11"
- name: Validate Studio package names and version
run: |
python - <<'PY'
import json
import os
import tomllib
from pathlib import Path
studio = Path("reme_studio")
python_manifest = tomllib.loads((studio / "pyproject.toml").read_text(encoding="utf-8"))["project"]
npm_manifest = json.loads((studio / "package.json").read_text(encoding="utf-8"))
expected = os.environ["RELEASE_VERSION"].removeprefix("v")
if python_manifest["name"] != "reme_studio":
raise SystemExit(f"Unexpected Python package name: {python_manifest['name']}")
if npm_manifest["name"] != "@agentscope-ai/reme_studio":
raise SystemExit(f"Unexpected npm package name: {npm_manifest['name']}")
if python_manifest["version"] != expected or npm_manifest["version"] != expected:
raise SystemExit(
f"Studio manifests are {python_manifest['version']} and {npm_manifest['version']}; "
f"workflow input is {expected}",
)
prerelease = "-" in expected
if prerelease != (os.environ["NPM_TAG"] == "next"):
raise SystemExit("Prereleases must use next; stable releases must use latest")
PY
- name: Install dependencies and run checks
working-directory: reme_studio
run: |
npm ci
npm run format:check
npm run lint
npm test
- name: Build Studio distributions
run: |
python -m pip install build twine
mkdir -p dist/studio-python dist/studio-npm
npm pack ./reme_studio --pack-destination dist/studio-npm
python scripts/package_studio.py
python -m build reme_studio --outdir dist/studio-python
python -m twine check dist/studio-python/*
- name: Verify Studio distributions and isolated installation
run: |
STUDIO_WHEEL="$(pwd)/$(ls dist/studio-python/reme_studio-*.whl)"
tar -tzf dist/studio-npm/*.tgz | grep '^package/dist-static/index.html$'
python -m venv "${RUNNER_TEMP}/reme-studio-package-smoke"
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" -m pip install "${STUDIO_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-studio-package-smoke/bin/python" - <<'PY'
from reme_studio import static_dir
assert (static_dir() / "index.html").is_file()
PY
- uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: reme-studio-${{ inputs.version }}
path: |
dist/studio-python/*
dist/studio-npm/*
if-no-files-found: error
publish-python:
needs: build
runs-on: ubuntu-latest
environment: pypi
permissions:
contents: read
id-token: write
steps:
- uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: reme-studio-${{ inputs.version }}
path: dist
- name: Publish ReMe Studio to PyPI
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # release/v1
with:
packages-dir: dist/studio-python
skip-existing: true
publish-npm:
needs: build
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: "24"
registry-url: https://registry.npmjs.org
- uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: reme-studio-${{ inputs.version }}
path: dist
- name: Publish ReMe Studio to npm
env:
NPM_TAG: ${{ inputs.npm_tag }}
run: npm publish dist/studio-npm/*.tgz --access public --tag "${NPM_TAG}" --provenance

View file

@ -1,167 +0,0 @@
# Release checklist:
# 1. Update typescript/package.json and package-lock.json to the release version and merge them.
# 2. Configure npm Trusted Publishing for agentscope-ai/ReMe and this workflow file.
# 3. Run this workflow manually with the exact package version (an optional v prefix is accepted).
# 4. Configure ClawHub Trusted Publishing or CLAWHUB_TOKEN before enabling ClawHub publication.
# 5. Use the `next` tag for prereleases and `latest` only for stable releases.
name: Release / TypeScript integrations
run-name: Publish @agentscope-ai/reme ${{ inputs.version }} (${{ inputs.npm_tag }})
on:
workflow_dispatch:
inputs:
version:
description: Version from typescript/package.json (for example, 0.1.0)
required: true
type: string
npm_tag:
description: npm distribution tag
required: true
default: latest
type: choice
options:
- next
- latest
publish_clawhub:
description: Also publish the verified tarball to ClawHub
required: true
default: false
type: boolean
permissions:
contents: read
concurrency:
group: publish-agentscope-ai-reme
cancel-in-progress: false
jobs:
build:
runs-on: ubuntu-latest
outputs:
version: ${{ steps.validate.outputs.version }}
env:
RELEASE_VERSION: ${{ inputs.version }}
NPM_TAG: ${{ inputs.npm_tag }}
steps:
- uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Set up Node
uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: '22.22.3'
- name: Validate package name and release version
id: validate
working-directory: typescript
run: |
node --input-type=module <<'JS'
import { appendFileSync, readFileSync } from 'node:fs';
const manifest = JSON.parse(readFileSync('package.json', 'utf8'));
const expected = process.env.RELEASE_VERSION.replace(/^v/, '');
if (manifest.name !== '@agentscope-ai/reme') {
throw new Error(`Unexpected package name: ${manifest.name}`);
}
if (manifest.version !== expected) {
throw new Error(`package.json is ${manifest.version}, workflow input is ${expected}`);
}
const prerelease = manifest.version.includes('-');
const npmTag = process.env.NPM_TAG;
if (prerelease !== (npmTag === 'next')) {
throw new Error(prerelease
? 'Prerelease versions must use the next npm tag'
: 'Stable versions must use the latest npm tag');
}
console.log(`Preparing ${manifest.name}@${manifest.version}`);
appendFileSync(process.env.GITHUB_OUTPUT, `version=${manifest.version}\n`);
JS
- name: Install dependencies
working-directory: typescript
run: npm ci
- name: Type-check and test
working-directory: typescript
run: |
npm run format:check
npm run lint
npm run typecheck
npm test
npm run test:package
npx --yes clawhub@0.23.3 package validate . --json
- name: Pack npm tarball
working-directory: typescript
run: |
mkdir -p "${RUNNER_TEMP}/reme-typescript-package"
npm pack --pack-destination "${RUNNER_TEMP}/reme-typescript-package"
- name: Upload npm tarball
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: agentscope-ai-reme-${{ inputs.version }}
path: ${{ runner.temp }}/reme-typescript-package/*.tgz
if-no-files-found: error
publish:
needs: build
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
steps:
- name: Set up Node for npm
uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6
with:
node-version: '24'
registry-url: https://registry.npmjs.org
- name: Download npm tarball
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
with:
name: agentscope-ai-reme-${{ inputs.version }}
path: dist/typescript
- name: Reject an existing package version
env:
PACKAGE_VERSION: ${{ inputs.version }}
run: |
PACKAGE_VERSION="${PACKAGE_VERSION#v}"
if npm view "@agentscope-ai/reme@${PACKAGE_VERSION}" version >/dev/null 2>&1; then
echo "@agentscope-ai/reme@${PACKAGE_VERSION} already exists" >&2
exit 1
fi
- name: Publish to npm
env:
NPM_TAG: ${{ inputs.npm_tag }}
run: npm publish dist/typescript/*.tgz --access public --tag "${NPM_TAG}" --provenance
publish-clawhub:
if: ${{ inputs.publish_clawhub }}
needs: build
permissions:
actions: read
contents: read
id-token: write
uses: openclaw/clawhub/.github/workflows/package-publish.yml@87ca030c30f3cfb78ab15c8e66b5ff1469c8f9c8 # v0.23.3
with:
owner: agentscope-ai
family: code-plugin
version: ${{ needs.build.outputs.version }}
tags: ${{ inputs.npm_tag }}
source_repo: ${{ github.repository }}
source_commit: ${{ github.sha }}
source_ref: ${{ github.ref }}
source_path: typescript
package_artifact_name: agentscope-ai-reme-${{ inputs.version }}
wait_for_publication: true
secrets:
clawhub_token: ${{ secrets.CLAWHUB_TOKEN }}

View file

@ -1,46 +0,0 @@
name: Security / CodeQL
on:
push:
branches: [main]
pull_request:
branches: [main]
schedule:
- cron: '0 1 * * 1'
workflow_dispatch:
permissions:
actions: read
contents: read
packages: read
security-events: write
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
analyze:
name: Analyze ${{ matrix.language }}
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
language: [python, javascript-typescript]
steps:
- name: Checkout repository
uses: actions/checkout@d23441a48e516b6c34aea4fa41551a30e30af803 # v6
with:
persist-credentials: false
- name: Initialize CodeQL
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
languages: ${{ matrix.language }}
build-mode: none
- name: Perform CodeQL analysis
uses: github/codeql-action/analyze@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4
with:
category: /language:${{ matrix.language }}

102
.gitignore vendored
View file

@ -1,82 +1,30 @@
# OS / editor
.vscode
.env*
.DS_Store
.idea/
.vscode/
.qoder/
*.code-workspace
# Local environment
.env
.env.*
!.env.example
!example.env
.venv/
.idea
venv/
env/
private*/
# Python caches / test artifacts
__pycache__/
*.py[cod]
*$py.class
.ipynb_checkpoints/
.pytest_cache/
.ruff_cache/
.mypy_cache/
.coverage
coverage.xml
htmlcov/
# Packaging / build outputs
build/
dist/
node_modules/
*.egg-info/
typescript/reports/
# Logs / temporary files
.ipynb_checkpoints
.__pycache__
__pycache__
*.log
nohup.out
nohup*.out
tmp*
temp*
private*
dist/
nohup*
cache
log/
logs/
runs/
tmp*/
temp*/
.trash/
# ReMe runtime data
.reme/
reme_workspace/
reme_workspace_auto_fin_real_test*/
vault/
*.db
*.sqlite
*.sqlite3
# Documentation build outputs
docs/_build/
site/
evaluation/
# The pi-Bench suite ships its own trace-history render config, which must
# stay in git even though it lives under an evaluation/ directory.
!benchmark/pibench/config/bench/evaluation/
!benchmark/pibench/config/bench/evaluation/**
datasets/
# Claude Code skills (local only)
.claude/skills/
# Benchmark memory workspaces (created on demand by run.py via mkdir)
benchmark/*/workspaces/
# Benchmark datasets (LongMemEval via download.py, BEAM via git clone)
benchmark/*/dataset/
# Benchmark outputs (created on demand by run.py via mkdir)
benchmark/*/results/
# integration tests outputs
tests/integration/logs/
daily/
runs
logs
rag_nodes_index.jsonl
alfworld_data
step_experiences/*
build/*
*.egg-info/*
cookbook/appworld/data/*
cookbook/appworld/experiments/*
cookbook/appworld/exp_result/*
file_vector_store/*
cookbook/appworld/file_vector_store/*
/.venv/

View file

@ -1,84 +0,0 @@
exclude: ^skills/
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v6.0.0
hooks:
- id: check-ast
- id: check-yaml
- id: check-xml
- id: check-toml
- id: check-json
- id: detect-private-key
- id: trailing-whitespace
- repo: https://github.com/asottile/add-trailing-comma
rev: v4.0.0
hooks:
- id: add-trailing-comma
- repo: https://github.com/psf/black
rev: 26.5.1
hooks:
- id: black
args: [--line-length=120, --target-version=py311]
- repo: https://github.com/PyCQA/flake8
rev: 7.3.0
hooks:
- id: flake8
args: [
"--extend-ignore=E203",
"--max-line-length=120"
]
- repo: https://github.com/pylint-dev/pylint
rev: v4.0.6
hooks:
- id: pylint
exclude:
(?x)(
^docs
| pb2\.py$
| grpc\.py$
| \.demo$
| \.md$
| \.html$
)
args: [
--disable=W0511,
--disable=W0718,
--disable=W0122,
--disable=W1203,
--disable=C0103,
--disable=R0913,
--disable=R0917,
--disable=E0401,
--disable=E1101,
--disable=E1111,
--disable=C0415,
--disable=W0603,
--disable=R1705,
--disable=R0914,
--disable=E0601,
--disable=W0602,
--disable=W0604,
--disable=R0801,
--disable=R0902,
--disable=R0903,
--disable=R0904,
--disable=C0123,
--disable=W0231,
--disable=W1113,
--disable=W0221,
--disable=R0401,
--disable=W0632,
--disable=W0123,
--disable=C3001,
--disable=R1702,
--disable=R0912,
--max-statements=120,
--max-line-length=120,
--max-module-lines=1500,
]
- repo: https://github.com/regebro/pyroma
rev: "5.0.1"
hooks:
- id: pyroma
args: [--min=10, .]

214
AGENTS.md
View file

@ -1,214 +0,0 @@
# AGENTS.md
This file guides coding agents working in the ReMe repository. Keep changes small, testable, and consistent with the
contracts expressed by the current code.
## Project Principles
ReMe is a local-first, file-native memory system for agents.
- User-owned workspace files are the durable source of truth.
- Indexes, catalogs, graphs, caches, and generated metadata must remain rebuildable.
- Prefer transparent formats and predictable behavior over hidden state.
- Preserve user control over workspace paths, configuration, and service boundaries.
- Keep concepts focused on project intent; let code and schemas describe implementation.
When convenience conflicts with these principles, favor data ownership, recoverability, and explicit behavior.
## Sources of Truth
Use this order when documentation and implementation disagree:
1. Current code and public Pydantic schemas.
2. Tests that describe supported behavior.
3. CLI behavior and the built-in configuration.
4. README files and other development documentation.
Do not duplicate large implementation descriptions in documentation. Express the stable contract and link to the
relevant module where useful. When behavior changes intentionally, update the implementation, schemas, tests, defaults,
and concise documentation together.
## Repository Map
- `reme/reme.py`: CLI entry point; dispatches `start`, `find_reme`, and client calls.
- `reme/application.py`: application assembly, dependency ordering, job execution, and lifecycle.
- `reme/config/config_parser.py`: YAML/JSON loading, environment expansion, dot-notation parsing, and deep config
merging.
- `reme/config/default.yaml`: default service, jobs, steps, and components. Other files in
`reme/config/` are named configuration variants.
- `reme/schema/application_config.py`: typed application, component, and job configuration.
- `reme/schema/`: request, response, streaming, memory, graph, and file contracts.
- `reme/components/application_context.py`: application-wide wiring and in-memory shared state.
- `reme/components/runtime_context.py`: request-scoped data, response, streaming queue, and stop event.
- `reme/components/base_component.py`: component lifecycle, dependency binding, and workspace helpers.
- `reme/components/component_registry.py`: the frozen built-in registry template and application-local registry factory.
- `reme/components/job/`: base, stream, background, and cron job implementations.
- `reme/components/service/`: local CLI, HTTP, and MCP service backends.
- `reme/components/`: agent wrappers, model adapters, stores, catalogs, graphs, indexes, clients, tokenizers, and
outbound proxies.
- `reme/steps/`: registered job steps grouped by common, file I/O, index, evolve, cookbook, benchmark, and transfer
concerns.
- `reme/utils/`: shared utilities, including service discovery, logging, web-static resolution, session I/O, token
accounting, and wikilink handling.
- `tests/unit/`: primary fast, isolated validation suite.
- `tests/integration/`: service/model tests that may need credentials or external processes.
- `reme_studio/`: ReMe Studio frontend source plus the independently published `reme_studio` Python package and
`@agentscope-ai/reme_studio` npm static distribution.
- `typescript/`: the independently published `@agentscope-ai/reme` package, including the shared TypeScript client and
DeepSeek Harness and OpenClaw adapters.
- `plugins/`: installable ReMe extensions, such as Auto Fin.
- `integrations/`: adapters that connect ReMe to external agent hosts, such as Claude Code, DSH, and Hermes Agent.
- `skills/`: standalone skills; `reme_memory` calls ReMe, while other skills may use separate tools or direct-file
conventions.
- `benchmark/` and `cookbook/`: runnable evaluations and example workflows.
- `docs/`: README-linked supporting pages and figures.
## Development Setup
ReMe requires Python 3.11 or newer. Install the editable development environment with:
```bash
pip install -e reme_studio -e ".[dev,core]"
```
Before changing behavior, inspect the adjacent implementation, schema, built-in config, and focused tests. Follow
existing async and typing patterns unless the task explicitly requires a new contract.
## Configuration and CLI Contracts
- CLI syntax is `reme ACTION key=value ...`; leading `-` or `--` on arguments is accepted.
- Nested overrides use dot notation. Values support null, booleans, numbers, JSON collections, and quoted JSON strings;
leading-zero numeric-looking values remain strings.
- `config=<name-or-path>` loads a discovered config name or a `.yaml`, `.yml`, or `.json` file. With no explicit config
path, `default` is loaded when available.
- Config files expand `${VAR}` and `${VAR:-default}` recursively. An undefined variable without a default is an error.
- CLI/config overrides are deep-merged over the loaded file. Do not silently change this merge behavior or stable
configuration keys.
- `ApplicationConfig` normalizes `workspace_dir` to an expanded absolute path. `session_dir`
must remain workspace-relative; standard transcripts live under `{session_dir}/dialog`.
- `reme start` runs the configured service. `reme start job=<name> ...` switches to the one-shot CLI service and runs
the job through the normal application lifecycle.
- Other actions use a client selected from the running service configuration when discoverable, otherwise from local
config. Client-selection arguments must not leak into the job payload.
## Registration and Application Lifecycle
Component and Step discovery is import-driven:
- Implementations declare a non-`BASE` `component_type` and register with `@R.register("backend")`
or `R.register(Class, "backend")`.
- Component packages must be imported through `reme/components/__init__.py`.
- Step packages/modules must be reachable through their package `__init__.py` chain and ultimately
`reme/steps/__init__.py`.
- Adding an implementation without its registration import leaves it undiscoverable at runtime. Treat implementation,
registration, import side effect, defaults, and tests as one change.
`Application` validates config through `ApplicationContext`, creates workspace directories, instantiates the service,
configured components, and jobs, and then manages lifecycle as follows:
- Components start in topological dependency order. Missing required dependencies and cycles fail explicitly; optional
dependencies may resolve to `None`.
- Jobs start after components in this order: base jobs, stream jobs, background jobs, then cron jobs.
- Shutdown closes everything in reverse start order and then shuts down the optional thread pool.
- If startup fails, already-started resources are closed.
- `BaseComponent.start()` and `close()` are lock-protected and idempotent. Dependencies created by a standalone
`default_factory` are owned and closed by the parent component.
Keep async clients, tasks, executors, and services under this lifecycle. Do not introduce an untracked long-lived
resource.
## Jobs, Steps, and State
`BaseJob` resolves configured Step classes during job startup and constructs fresh Step instances for every invocation.
Job-level kwargs are merged into each `RuntimeContext`, with call-time kwargs taking precedence. Sequential Steps in one
invocation share the same `RuntimeContext` and `Response`.
Treat Step instances as invocation-scoped:
- Constructor fields and `self.kwargs` hold Step configuration and resolved dependencies. They may be cached or adjusted
during that one invocation, but must not be relied on across Job calls.
- `self.context.data` holds request inputs and intermediate values shared by sequential Steps.
- `self.context.response.answer`, `success`, and `metadata` are request-scoped output. Because the same response travels
through the Step chain, later Steps may consume metadata produced earlier, but it is not application-lifetime or
durable storage.
- `self.app_context.metadata` holds in-memory state shared across Job/Step invocations for the life of one
`Application`, such as counters, tool-context state, session maps, or locks.
- Workspace files or a dedicated Component/store hold durable state that must survive restart.
Use narrow, namespaced keys in `app_context.metadata` and protect shared mutable values against concurrent access. The
search/draft helpers intentionally mirror tool-context state into
`self.kwargs` only when no `ApplicationContext` exists for standalone use and unit tests; do not generalize that
compatibility fallback into persistent runtime state. If shared state becomes a stable service contract or needs
dedicated lifecycle, locking, or persistence, promote it to a typed context field or Component.
Additional Step contracts:
- `Ref` dependencies resolve in this order: Step kwargs, current `RuntimeContext`, then the named application component.
The value is cached only on the current Step instance and cleared before each call.
- `input_mapping` and `output_mapping` copy keys within `RuntimeContext.data`; missing sources are ignored.
- Dispatched Steps receive the current `RuntimeContext`, so their data and response are shared.
- Base jobs convert uncaught Step errors into `Response(success=False)`; stream jobs emit an error chunk and always a
terminal `DONE`; background jobs let errors reach their supervisor.
- Background jobs are never service-exposed. MCP also skips stream jobs. Respect `enable_serve`
and any configured service job allowlist.
## Workspace and File Safety
- Application startup creates the workspace plus configured metadata, session, memory-session, resource, daily, and
digest directories.
- File-operation paths are resolved against the workspace and must stay inside it. Home-relative paths are unsupported,
traversal escapes are rejected, and `_allowed_paths` restrictions fail closed when invalid.
- Preserve per-path locking, encoding detection, byte limits, truncation behavior, and optimistic
`expected_mtime` checks when modifying file operations.
- Do not bypass the existing file steps or stores in a way that weakens workspace containment.
- Never write test state into the repository's `.reme/`; use `tmp_path` or another isolated workspace.
- Do not delete or rewrite user memory to repair an index or make a test pass. Rebuild derived state from source files
instead.
## Validation
Use the narrowest useful check while iterating, then broaden it according to risk.
Focused test:
```bash
pytest tests/unit/path/to/test_file.py -v
```
Main unit suite:
```bash
pytest tests/unit -v --tb=long -s --log-cli-level=WARNING
```
Repository formatting and lint checks:
```bash
pre-commit run --all-files
```
Black and Flake8 use a 120-character line limit and Python 3.11 formatting; Pylint is also run by pre-commit. If
`reme_studio/` changes, use its Node 22.13+ scripts and run the proportionate checks from that directory, such as
`npm run format:check`, `npm run lint`, or `npm test`.
Integration tests may contact real model providers, services, or agent subprocesses and can require credentials. Do not
run credentialed or externally mutating tests automatically; run them only when the task requires them and the necessary
environment has been supplied or authorized. Mock network, model, and subprocess boundaries in unit tests.
## Change Guardrails
- Preserve unrelated user changes in a dirty working tree.
- Make the smallest coherent change and avoid unrelated cleanup or broad refactors.
- Do not edit generated output when the source can be changed instead. The publish workflow builds
`reme_studio/dist-static` and stages it under `reme_studio/src/reme_studio/static`; change `reme_studio/` source for
frontend work.
- Do not silently change CLI flags, configuration keys, workspace layouts, serialized schemas, endpoint shapes,
streaming termination, or service interfaces. Preserve compatibility where practical and document intentional
migrations.
- Do not introduce dependencies without a concrete repository-level need.
- Do not commit `.env` files, credentials, runtime memory, logs, indexes, caches, benchmark outputs, or generated
Studio distributions.
- State which validations passed and which relevant checks were not run in the final handoff.
If a requirement is ambiguous, infer intent from nearby code, schemas, defaults, and tests. Ask the user only when the
remaining choice would materially alter a public contract, user data, or an external system.

View file

@ -1 +0,0 @@
AGENTS.md

684
README.md
View file

@ -1,401 +1,439 @@
English | [**中文**](./README_ZH.md)
<p align="center">
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
<img src="docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
</p>
<p align="center">
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/python-3.11+-blue" alt="Python Version"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/pypi/v/reme-ai.svg?logo=pypi" alt="PyPI Version"></a>
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/python-3.12+-blue" alt="Python Version"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/pypi-v0.1-blue?logo=pypi" alt="PyPI Version"></a>
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
<a href="https://reme.agentscope.io"><img src="https://img.shields.io/badge/docs-ReMe-blue" alt="Documentation"></a>
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
<a href="https://deepwiki.com/agentscope-ai/ReMe"><img src="https://img.shields.io/badge/DeepWiki-Ask_Devin-navy.svg" alt="DeepWiki"></a>
<a href="https://github.com/modelscope/ReMe"><img src="https://img.shields.io/github/stars/modelscope/ReMe?style=social" alt="GitHub Stars"></a>
</p>
<p align="center">
<a href="https://trendshift.io/repositories/20528" target="_blank"><img src="https://trendshift.io/api/badge/repositories/20528" alt="agentscope-ai%2FReMe | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
<strong>ReMe (formerly MemoryScope): Memory Management Framework for Agents</strong><br>
<em>Remember Me, Refine Me.</em>
</p>
<p align="center">
<strong>A local-first, self-evolving personal knowledge base for AI agents.</strong><br>
</p>
---
ReMe provides AI agents with a unified memory system—enabling the ability to extract, reuse, and share memories across
users, tasks, and agents.
> Previous versions: [0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
```
Personal Memory + Task Memory = Agent Memory
```
## ✨ Why ReMe?
Personal memory helps "**understand user preferences**", while task memory helps agents "**perform better**".
🧠 ReMe turns conversations and resources into readable, editable, searchable, and interconnected Markdown memory. Agents
such as QwenPaw and DeepSeek Harness can share the same workspace to retrieve, maintain, and evolve knowledge, while
users retain control of the durable files.
- **Memory as File, File as Memory**: ReMe stores durable memory as ordinary Markdown with frontmatter and wikilinks.
Users and agents can inspect, edit, move, sync, and back it up with familiar tools, while indexes and generated
metadata remain rebuildable.
- **Self-evolving knowledge base**: ReMe progressively turns conversations and resources into daily notes and long-term
knowledge, preserving sources while refining facts, preferences, procedures, and relationships over time.
- **Recall is precise and context-aware.** BM25, optional embeddings, and wikilink expansion retrieve relevant
line-level passages and their relationships without loading the entire knowledge base into the agent context.
- **One memory workspace works across agents.** Personal assistants, coding agents, and other agent runtimes can share
the same local workspace through native integrations, SKILL.md, CLI, HTTP, MCP, or Python APIs.
<p align="center">
<img src="docs/figure/design-philosophy.svg" alt="ReMe Design Philosophy" width="92%">
</p>
---
## 📰 Latest Updates
- [2026.08] - Published [`@agentscope-ai/reme`](https://www.npmjs.com/package/@agentscope-ai/reme), providing native
ReMe memory integrations for DeepSeek Harness and OpenClaw plus a shared TypeScript HTTP client.
- [2026.08] - Published the [ReMe blog](https://agentscope-ai.github.io/ReMe/?doc=en-reme-blog), an end-to-end introduction to its local-first memory
architecture, self-evolving workflows, hybrid search, proactive discovery, and benchmark results.
- [2026.08] - [Experience-driven enhancement method](https://reme.agentscope.io/?doc=toolmemory-en) of agent tool-use execution built
on ReMe is available on [arXiv:2608.03403](https://arxiv.org/abs/2608.03403).
- [2026.07] - Introduced optional plugins: [Daily Paper](https://reme.agentscope.io/?doc=daily-paper-en) for paper discovery and
analysis, and [Auto Fin](https://reme.agentscope.io/?doc=auto-fin-en) for researching the latest 24 hours of topic-related CLS news
with local-memory search and validated historical wikilinks.
- [2026.07] - Our
paper [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
has been accepted to Findings of ACL 2026.
- **[2025-09]** 🎉 ReMe v0.1.8 has been officially released, adding support for asynchronous operations. It has also been
integrated into the memory service of agentscope-runtime.
- **[2025-09]** 🎉 ReMe v0.1 officially released, integrating task memory and personal memory. If you want to use the
original memoryscope project, you can find it
in [MemoryScope](https://github.com/modelscope/Reme/tree/memoryscope_branch).
- **[2025-09]** 🧪 We validated the effectiveness of task memory extraction and reuse in agents in appworld, bfcl(v3),
and frozenlake environments. For more information,
check [appworld exp](docs/cookbook/appworld/quickstart.md), [bfcl exp](docs/cookbook/bfcl/quickstart.md),
and [frozenlake exp](docs/cookbook/frozenlake/quickstart.md).
- **[2025-08]** 🚀 MCP protocol support is now available -> [MCP Quick Start](docs/mcp_quick_start.md).
- **[2025-06]** 🚀 Multiple backend vector storage support (Elasticsearch &
ChromaDB) -> [Vector DB quick start](docs/vector_store_api_guide.md).
- **[2024-09]** 🧠 [MemoryScope](https://github.com/modelscope/Reme/tree/memoryscope_branch) v0.1 released,
personalized and time-aware memory storage and usage.
---
## ✨ Architecture Design
<p align="center">
<img src="docs/figure/reme_structure.jpg" alt="ReMe Logo" width="100%">
</p>
ReMe integrates two complementary memory capabilities:
#### 🧠 **Task Memory/Experience**
Procedural knowledge reused across agents
- **Success Pattern Recognition**: Identify effective strategies and understand their underlying principles
- **Failure Analysis Learning**: Learn from mistakes and avoid repeating the same issues
- **Comparative Patterns**: Different sampling trajectories provide more valuable memories through comparison
- **Validation Patterns**: Confirm the effectiveness of extracted memories through validation modules
Learn more about how to use task memory from [task memory](docs/task_memory/task_memory.md)
#### 👤 **Personal Memory**
Contextualized memory for specific users
- **Individual Preferences**: User habits, preferences, and interaction styles
- **Contextual Adaptation**: Intelligent memory management based on time and context
- **Progressive Learning**: Gradually build deep understanding through long-term interaction
- **Time Awareness**: Time sensitivity in both retrieval and integration
Learn more about how to use personal memory from [personal memory](docs/personal_memory/personal_memory.md)
---
## 🛠️ Installation
### Install from PyPI (Recommended)
```bash
pip install reme-ai
```
### Install from Source
```bash
git clone https://github.com/modelscope/ReMe.git
cd ReMe
pip install .
```
### Environment Configuration
Copy `example.env` to .env and modify the corresponding parameters:
```bash
FLOW_APP_NAME=ReMe
FLOW_LLM_API_KEY=sk-xxxx
FLOW_LLM_BASE_URL=https://xxxx/v1
FLOW_EMBEDDING_API_KEY=sk-xxxx
FLOW_EMBEDDING_BASE_URL=https://xxxx/v1
```
---
## 🚀 Quick Start
### Installation
ReMe requires Python 3.11+.
Install from pip:
### HTTP Service Startup
```bash
pip install "reme-ai[core]"
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen3-30b-a3b-thinking-2507 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
Install from source:
### MCP Server Support
```bash
git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
pip install -e reme_studio -e ".[core]"
cd reme_studio
npm ci
npm run build:static
cd ..
reme \
backend=mcp \
mcp.transport=stdio \
llm.default.model_name=qwen3-30b-a3b-thinking-2507 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
The static build requires Node.js 22.13 or newer and makes Studio available from the source tree.
### Core API Usage
### Start the Service
#### Task Memory Management
```python
import requests
# Experience Summarizer: Learn from execution trajectories
response = requests.post("http://localhost:8002/summary_task_memory", json={
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [{"role": "user", "content": "Help me create a project plan"}], "score": 1.0}
]
})
# Retriever: Get relevant memories
response = requests.post("http://localhost:8002/retrieve_task_memory", json={
"workspace_id": "task_workspace",
"query": "How to efficiently manage project progress?",
"top_k": 1
})
```
<details>
<summary>curl version</summary>
```bash
reme start
# Experience Summarizer: Learn from execution trajectories
curl -X POST http://localhost:8002/summary_task_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [{"role": "user", "content": "Help me create a project plan"}], "score": 1.0}
]
}'
# Retriever: Get relevant memories
curl -X POST http://localhost:8002/retrieve_task_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"query": "How to efficiently manage project progress?",
"top_k": 1
}'
```
The default service address is `127.0.0.1:2333`. If the port is occupied, specify another port:
</details>
<details>
<summary>Node.js version</summary>
```javascript
// Experience Summarizer: Learn from execution trajectories
fetch("http://localhost:8002/summary_task_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
trajectories: [
{messages: [{role: "user", content: "Help me create a project plan"}], score: 1.0}
]
})
})
.then(response => response.json())
.then(data => console.log(data));
// Retriever: Get relevant memories
fetch("http://localhost:8002/retrieve_task_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
query: "How to efficiently manage project progress?",
top_k: 1
})
})
.then(response => response.json())
.then(data => console.log(data));
```
</details>
#### Personal Memory Management
```python
# Memory Integration: Learn from user interactions
response = requests.post("http://localhost:8002/summary_personal_memory", json={
"workspace_id": "task_workspace",
"trajectories": [
{"messages":
[
{"role": "user", "content": "I like to drink coffee while working in the morning"},
{"role": "assistant",
"content": "I understand, you prefer to start your workday with coffee to stay energized"}
]
}
]
})
# Memory Retrieval: Get personal memory fragments
response = requests.post("http://localhost:8002/retrieve_personal_memory", json={
"workspace_id": "task_workspace",
"query": "What are the user's work habits?",
"top_k": 5
})
```
<details>
<summary>curl version</summary>
```bash
reme start service.port=8181
# reme start workspace_dir=/tmp/reme-demo service.port=8181
# Memory Integration: Learn from user interactions
curl -X POST http://localhost:8002/summary_personal_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [
{"role": "user", "content": "I like to drink coffee while working in the morning"},
{"role": "assistant", "content": "I understand, you prefer to start your workday with coffee to stay energized"}
]}
]
}'
# Memory Retrieval: Get personal memory fragments
curl -X POST http://localhost:8002/retrieve_personal_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"query": "What are the user's work habits?",
"top_k": 5
}'
```
```bash
reme version
reme health_check
reme help
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
</details>
<details>
<summary>Node.js version</summary>
```javascript
// Memory Integration: Learn from user interactions
fetch("http://localhost:8002/summary_personal_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
trajectories: [
{messages: [
{role: "user", content: "I like to drink coffee while working in the morning"},
{role: "assistant", content: "I understand, you prefer to start your workday with coffee to stay energized"}
]}
]
})
})
.then(response => response.json())
.then(data => console.log(data));
// Memory Retrieval: Get personal memory fragments
fetch("http://localhost:8002/retrieve_personal_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
query: "What are the user's work habits?",
top_k: 5
})
})
.then(response => response.json())
.then(data => console.log(data));
```
### 5-Minute Memory Demo
</details>
With the service running, write a memory node, let ReMe index it, then retrieve it:
```bash
reme write \
path=digest/wiki/quick-start-demo \
name="Quick Start Demo" \
description="A first ReMe memory node" \
content="# Quick Start Demo
ReMe stores agent memory as readable Markdown.
Related: [[digest/wiki/memory-as-file.md]]"
reme search query="agent memory markdown" limit=5
reme read path=digest/wiki/quick-start-demo start_line=1 end_line=20
```
The generated file is ordinary Markdown with frontmatter:
```markdown
---
name: Quick Start Demo
description: A first ReMe memory node
---
# Quick Start Demo
## 📦 Ready-to-Use Libraries
ReMe stores agent memory as readable Markdown.
ReMe provides pre-built memory libraries that agents can immediately use with verified best practices:
Related: [[digest/wiki/memory-as-file.md]]
### Available Libraries
- **`appworld.jsonl`**: Memory library for Appworld agent interactions, covering complex task planning and execution
patterns
- **`bfcl_v3.jsonl`**: Working memory library for BFCL tool calls
### Quick Usage
```python
# Load pre-built memories
response = requests.post("http://localhost:8002/vector_store", json={
"workspace_id": "appworld",
"action": "load",
"path": "./docs/library/"
})
# Query relevant memories
response = requests.post("http://localhost:8002/retrieve_task_memory", json={
"workspace_id": "appworld",
"query": "How to navigate to settings and update user profile?",
"top_k": 1
})
```
### ReMe Studio (Optional)
## 🧪 Experiments
The `core` installation includes Studio. After starting ReMe, open <http://127.0.0.1:2333/> to browse, edit, and search
the workspace. To add Studio to a base installation, use `pip install "reme-ai[web]"`. See the
[ReMe Studio guide](https://reme.agentscope.io/?doc=studio-en) for source builds, configuration, and development.
### 🌍 [Appworld Experiment](docs/cookbook/appworld/quickstart.md)
### Optional Model Configuration
We tested ReMe on Appworld using qwen3-8b:
Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are
disabled by default, so the default setup does not start an embedding model or require an embedding API key.
| Method | pass@1 | pass@2 | pass@4 |
|--------------|-------------------|-------------------|-------------------|
| without ReMe | 0.083 | 0.140 | 0.228 |
| with ReMe | 0.109 **(+2.6%)** | 0.175 **(+3.5%)** | 0.281 **(+5.3%)** |
```bash
cat > .env <<'EOF'
# Optional: used only after embedding components are explicitly enabled in the config.
# EMBEDDING_API_KEY=sk-xxx
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
Pass@K measures the probability that at least one of the K generated samples successfully completes the task (
score=1).
The current experiment uses an internal AppWorld environment, which may have slight differences.
# Required for auto_memory, auto_resource, and auto_dream.
LLM_API_KEY=sk-xxx
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
EOF
```
You can find more details on reproducing the experiment in [quickstart.md](docs/cookbook/appworld/quickstart.md).
Basic file operations, BM25 search, wikilink traversal, and reading proactive topics can run without LLM credentials.
### 🧊 [Frozenlake Experiment](docs/cookbook/frozenlake/quickstart.md)
> [!NOTE]
> To enable embedding-based semantic retrieval, uncomment `components.as_embedding` and
> `components.embedding_store` in [`reme/config/default.yaml`](reme/config/default.yaml), then change
> `components.file_store.default.embedding_store` from `""` to `default`. See the
> [memory search guide](docs/en/memory_search.md) for details.
| without ReMe | with ReMe |
|:--------------------------------------------------------------------------------------------:|:--------------------------------------------------------------------------------------------:|
| <p align="center"><img src="docs/figure/frozenlake_failure.gif" alt="GIF 1" width="30%"></p> | <p align="center"><img src="docs/figure/frozenlake_success.gif" alt="GIF 2" width="30%"></p> |
## 🤝 Use ReMe with Your Agent
We tested on 100 random frozenlake maps using qwen3-8b:
ReMe can run as a local memory service accessed through the CLI, HTTP API, or MCP server, or it can be embedded in the
host process through its Python API. Host integrations can add memory guidance, recall, and capture to the agent
lifecycle according to the capabilities of each runtime.
| Method | pass rate |
|--------------|------------------|
| without ReMe | 0.66 |
| with ReMe | 0.72 **(+6.0%)** |
| Agent | Recommended path | Available after integration |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| **DeepSeek Harness** | Install [`@agentscope-ai/reme`](typescript/README.md#deepseek-harness) with `dsh plugin --profile web add @agentscope-ai/reme`. | Long-term memory guidance, the `reme_search` tool, and automatic capture of completed main-agent turns. |
| **OpenClaw** | Install [`@agentscope-ai/reme`](typescript/README.md#openclaw) with `openclaw plugins install @agentscope-ai/reme`. | Native memory tools, recall before user-triggered runs, and automatic turn capture. |
| **QwenPaw** | Embed ReMe in-process through its Python API. | Reuse the host lifecycle and model config while keeping memory local and file-based. |
| **Claude Code** | Start the streamable HTTP MCP service and install [the ReMe plugin](integrations/claude_code/reme). | MCP recall tools, the `reme-memory` skill, and a Stop hook that records sessions automatically. |
| **Hermes** | Start the HTTP service and install [the ReMe provider](integrations/hermes_agent). | Recall before model calls and asynchronous `auto_memory` after each completed turn. |
| **Codex and other CLI agents** | Install or copy the [ReMe Memory skill](skills/reme_memory/SKILL.md). | Search, read, and write memory through the CLI; automatic capture requires host lifecycle integration. |
You can find more details on reproducing the experiment in [quickstart.md](docs/cookbook/frozenlake/quickstart.md).
<p align="center"><b>Integration demos</b></p>
### 🔧 [BFCL-V3 Experiment](docs/cookbook/bfcl/quickstart.md)
<table>
<tr>
<td align="center"></td>
<td width="45%" align="center"><b>Auto Memory</b></td>
<td width="45%" align="center"><b>Auto Dream</b></td>
</tr>
<tr>
<td align="center"><b>QwenPaw</b></td>
<td width="45%">
<img src="docs/figure/qwenpaw-auto-memory.gif" alt="QwenPaw Auto Memory demo" width="100%">
</td>
<td width="45%">
<img src="docs/figure/qwenpaw-auto-dream.gif" alt="QwenPaw Auto Dream demo" width="100%">
</td>
</tr>
<tr>
<td align="center"><b>Claude Code</b></td>
<td width="45%">
<img src="docs/figure/cc-auto-memory.gif" alt="Claude Code Auto Memory demo" width="100%">
</td>
<td width="45%">
<img src="docs/figure/cc-auto-dream.gif" alt="Claude Code Auto Dream demo" width="100%">
</td>
</tr>
</table>
We tested ReMe on BFCL-V3 multi-turn-base (randomly split 50train/150val) using qwen3-8b:
## 🧠 How ReMe Works
| Method | pass@1 | pass@2 | pass@4 |
|--------------|---------------------|---------------------|---------------------|
| without ReMe | 0.2472 | 0.2733 | 0.2922 |
| with ReMe | 0.3061 **(+5.89%)** | 0.3500 **(+7.67%)** | 0.3888 **(+9.66%)** |
> Memory as File, File as Memory.
## 📚 Resources
ReMe treats **memory as files**, progressively processing filtered conversation source records and external resources
from `session/` and `resource/` into `daily/`, then `digest/`. The default workspace is `.reme/` under the current
directory; `workspace_dir=...` selects a different user-owned location.
- **[Quick Start](./cookbook/simple_demo)**: Get started quickly with practical examples
- **[Vector Storage Setup](docs/vector_store_api_guide.md)**: Configure local/vector databases and usage
- **[MCP Guide](docs/mcp_quick_start.md)**: Create MCP services
- **[personal memory](docs/personal_memory)** & **[task memory](docs/task_memory)** : Operators used in personal memory and task memory, You can modify the config to customize the pipelines.
- **[Example Collection](./cookbook)**: Real use cases and best practices
### Workspace Layout
---
```text
<workspace_dir>/
├── metadata/ # Rebuildable indexes, graphs, catalogs, and caches
├── session/ # Conversation source records and agent sessions
│ ├── dialog/
│ │ └── <session_id>.jsonl # Source messages saved by auto_memory
│ └── claude_code/
│ └── <session_id>.jsonl # ReMe copy used by auto_memory_cc
├── mem_session/ # Generated agent-wrapper sessions/config, not user memory
│ ├── agentscope/
│ ├── claude_config/
│ └── codex/
├── resource/ # External raw materials
│ ├── <resource>.<ext> # Root-level files enter today's daily layer
│ └── YYYY-MM-DD/
│ └── <resource>.<ext>
├── daily/ # Lightly processed memory: daily facts, conversation summaries, resource readings
│ ├── YYYY-MM-DD.md
│ └── YYYY-MM-DD/
│ ├── <generated_name>.md # Topic-named conversation or resource card
│ └── interests.yaml
└── digest/ # Long-term memory: personal facts, procedural experience, knowledge nodes
├── personal/
│ └── {topic/event}.md
├── procedure/
│ └── {topic/event}.md
└── wiki/
└── {topic/event}.md
```
## 🤝 Contribution
<p align="center">
<img src="docs/figure/reme-overview.svg" alt="ReMe file-based memory system overview" width="92%">
</p>
We believe the best memory systems come from collective wisdom. Contributions welcome 👉[Guide](docs/contribution.md):
### Memory Lifecycle
### Code Contributions
ReMe follows a capture → index → consolidate → recall loop. Workspace files remain the durable source of truth;
everything under `metadata/` is rebuildable.
- New operation and tool development
- Backend implementation and optimization
- API enhancements and new endpoints
| Capability | Entry point | What it does | Output |
| ------------------------------------------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
| [`auto_memory`](docs/en/auto_memory.md) | Agent hook or `reme auto_memory` | Distills useful conversation facts while preserving a filtered conversation source record. | `session/dialog/*.jsonl`, `daily/<date>/<generated-name>.md` |
| [`auto_resource`](docs/en/auto_resource.md) | Resource watcher or `reme auto_resource` | Turns files under `resource/` into source-linked, content-named daily cards. | `daily/<date>/<resource-card>.md` |
| [`auto_index`](docs/en/memory_search.md) | Background watcher or `reme reindex` | Live-indexes Markdown in `daily/` and `digest/`; a full rebuild also scans `resource/` and JSONL. | Searchable chunks, BM25, wikilink graph, and optional vectors |
| [`auto_dream`](docs/en/auto_dream.md) | `dream_cron` or `reme auto_dream` | By default, extracts up to five reusable units from changed files in the latest two-day window, then creates, corroborates, refines, or corrects digest nodes. | `digest/**`, `daily/<date>/interests.yaml` |
| [`proactive`](docs/en/proactive.md) | `reme proactive` before an agent decides to act | Reads topics generated by `auto_dream`; the host agent decides whether and how to mention them. | Structured topics from `daily/<date>/interests.yaml` |
### Documentation Improvements
<table>
<tr>
<td align="center" width="50%">
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
</td>
<td align="center" width="50%">
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
</td>
<td align="center" width="50%">
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
</td>
</tr>
</table>
- Usage examples and tutorials
- Best practice guides
Search returns matching chunks with line ranges and bounded wikilink neighbors. Optional vector results are fused with
BM25 through reciprocal rank fusion (RRF).
> [!IMPORTANT]
>
> `proactive` only reads and exposes interest topics produced by Auto Dream. It does not independently browse the web,
> send notifications, or rewrite the knowledge base; the host agent decides whether and how to act on a topic.
## 📊 Benchmarks
ReMe evaluates multi-session and long-context memory with agentic search-and-read workflows. The figures below are the
published reference runs in this repository; model, prompt, dataset, and judging details are documented with each
benchmark.
| Benchmark | Setting | Sample size | Agentic score | Focus |
| --------------------------------------------------------------------------- | ------------ | -----------------------: | ------------: | ------------------------------------------------------------------ |
| **[LongMemEval cleaned-s](https://reme.agentscope.io/?doc=longmemeval-en)** | **Overall** | **500 questions** | **89.4%** | Cross-session retrieval, knowledge updates, and temporal reasoning |
| [BEAM](https://reme.agentscope.io/?doc=beam-en) | 100K context | 20 cases / 400 questions | 66.1% | Ten types of long-context memory tasks |
| [BEAM](https://reme.agentscope.io/?doc=beam-en) | 1M context | 35 cases / 700 questions | 65.0% | Ultra-long conversation settings |
ReMe also achieved a **0.580 PROC score across five user personas** in the repository's
[π-Bench evaluation](https://reme.agentscope.io/?doc=pibench-en), 2.4% above NanoBot under the same test-model configuration. PROC
measures proactive handling of hidden intent, clarification, cross-session preferences and conventions, task
dependencies, and underspecified requests.
## 🧩 Extensions and Plugins
Plugins are optional Python distributions that contribute Component, Step, or Job backends and configuration. They are
installed separately and enabled explicitly by configuration. Daily Paper and Auto Fin are independently packaged
plugins; see the source distributions and their documentation for [Daily Paper](plugins/daily_paper/README.md) and
[Auto Fin](plugins/auto-fin/README.md).
| Plugin | Capability |
| ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| [Daily Paper](https://reme.agentscope.io/?doc=daily-paper-en) | Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief. |
| [Auto Fin](https://reme.agentscope.io/?doc=auto-fin-en) | Fetch topic-related CLS news, search ReMe history, and generate wikilink-backed Markdown reports. |
See [Plugin Management](docs/en/plugin_management.md) to install, inspect, validate, enable, and uninstall ReMe plugins.
## 📚 Documentation
These guides cover the main user workflows and the runtime contracts implemented by the current code.
| Guide | What you will learn |
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| [Quick Start](docs/en/quick_start.md) | Install ReMe, start the service, and run the first file and memory operations. |
| [Memory as File](docs/en/memory_as_file.md) | Understand workspace layers, frontmatter, wikilinks, chunks, and the file-as-source-of-truth model. |
| [Auto Memory](docs/en/auto_memory.md) | Preserve source conversations and distill reusable daily memory cards. |
| [Auto Resource](docs/en/auto_resource.md) | Import supported text resources and turn them into source-linked daily cards. |
| [Auto Dream](docs/en/auto_dream.md) and [Auto Link](docs/en/auto_link.md) | Consolidate daily notes into evolving digest nodes and readable wikilink relationships. |
| [Memory Search](docs/en/memory_search.md) | Use BM25, optional vectors, RRF fusion, line-range recall, and progressive link expansion. |
| [Proactive](docs/en/proactive.md) | Read interest topics safely and integrate them into a host agent's decision flow. |
| [Application Scenarios](docs/en/reme_scene.md) | Follow concrete financial research, coding-memory, and personal knowledge-base examples. |
| [Framework](docs/en/framework.md) | Understand Application, Job, Step, Component, service, configuration, and lifecycle boundaries. |
| [TypeScript integrations](typescript/README.md) | Configure the shared client and native DeepSeek Harness and OpenClaw adapters. |
| [ReMe Blog](https://agentscope-ai.github.io/ReMe/?doc=en-reme-blog) | Read the product story, design rationale, examples, and benchmark summary. |
## 🛠️ Common Commands
Run `reme help` for the full job list. Common workspace and maintenance commands are:
| Command | Purpose |
| ----------------------------------------- | --------------------------------------------------------------------------------- |
| `reme status` | Show stateful data-component memory estimates and process RSS. |
| [`reme search`](docs/en/memory_search.md) | Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled. |
| `reme read` / `reme write` / `reme edit` | Inspect and maintain Markdown memory files. |
| `reme traverse` / `reme graph_snapshot` | Explore wikilink neighborhoods or the category-rooted digest graph. |
| `reme chat` | Stream a read-only, workspace-aware agent conversation. Requires LLM credentials. |
| `reme reindex` | Rebuild search and wikilink indexes from existing files. |
## 🤝 Community and Contributing
- **Issues, requests, and help**: Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) first. If there is no
related discussion, open one with the background, expected behavior, and impact scope.
- **Code contributions**: Before making changes, read the repository's
[contribution guide](docs/en/contributing.md). Source, schemas, and tests are the authoritative architecture and
extension guide.
- **Documentation contributions**: Update the canonical files under `docs/en/`, `docs/zh/`, or the relevant package
directory in this repository. The documentation site is generated from these files.
- **Commit convention**: Conventional Commits are recommended, for example `feat(search): add link expansion option` or
`docs(zh): update quick start`.
- **Pre-submit checks**: Before submitting a PR, try to run `pre-commit run --all-files` and `pytest`. If tests that
depend on LLMs, embeddings, or external services cannot run, explain that in the PR.
- **Documentation**: Visit [reme.agentscope.io](https://reme.agentscope.io).
### Contributors
Thanks to everyone who has contributed to ReMe:
<a href="https://github.com/agentscope-ai/ReMe/graphs/contributors">
<img src="https://contrib.rocks/image?repo=agentscope-ai/ReMe" alt="Contributors" />
</a>
---
## 📄 Citation
```bibtex
@software{ReMe2026,
title = {Remember me, Refine me: Memory Management Kit for Agents},
author = {ReMe Team},
url = {https://reme.agentscope.io},
year = {2026}
@software{ReMe2025,
title = {ReMe: Memory Management Framework for Agents},
author = {Li Yu, Jiaji Deng, Zouying Cao},
url = {https://github.com/modelscope/ReMe},
year = {2025}
}
```
---
## ⚖️ License
This project is open source under the Apache License 2.0. See [LICENSE](./LICENSE) for details.
This project is licensed under the Apache License 2.0 - see the [LICENSE](./LICENSE) file for details.
---
## Star History
[![Star History Chart](https://api.star-history.com/svg?repos=modelscope/ReMe&type=Date)](https://www.star-history.com/#modelscope/ReMe&Date)

View file

@ -1,388 +1,412 @@
中文 | [**English**](./README.md)
<p align="center">
<img src="https://raw.githubusercontent.com/agentscope-ai/ReMe/main/docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
<img src="docs/figure/reme_logo.png" alt="ReMe Logo" width="50%">
</p>
<p align="center">
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/python-3.11+-blue" alt="Python Version"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/pypi/v/reme-ai.svg?logo=pypi" alt="PyPI Version"></a>
<a href="https://pepy.tech/project/reme-ai/"><img src="https://img.shields.io/pypi/dm/reme-ai" alt="PyPI Downloads"></a>
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/commit-activity/m/agentscope-ai/ReMe?style=flat-square" alt="GitHub commit activity"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/python-3.12+-blue" alt="Python Version"></a>
<a href="https://pypi.org/project/reme-ai/"><img src="https://img.shields.io/badge/pypi-v0.1-blue?logo=pypi" alt="PyPI Version"></a>
<a href="./LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-black" alt="License"></a>
<a href="https://reme.agentscope.io"><img src="https://img.shields.io/badge/docs-ReMe-blue" alt="文档"></a>
<a href="./README.md"><img src="https://img.shields.io/badge/English-Click-yellow" alt="English"></a>
<a href="./README_ZH.md"><img src="https://img.shields.io/badge/简体中文-点击查看-orange" alt="简体中文"></a>
<a href="https://github.com/agentscope-ai/ReMe"><img src="https://img.shields.io/github/stars/agentscope-ai/ReMe?style=social" alt="GitHub Stars"></a>
<a href="https://deepwiki.com/agentscope-ai/ReMe"><img src="https://img.shields.io/badge/DeepWiki-Ask_Devin-navy.svg" alt="DeepWiki"></a>
<a href="https://github.com/modelscope/ReMe"><img src="https://img.shields.io/github/stars/modelscope/ReMe?style=social" alt="GitHub Stars"></a>
</p>
<p align="center">
<a href="https://trendshift.io/repositories/20528" target="_blank"><img src="https://trendshift.io/api/badge/repositories/20528" alt="agentscope-ai%2FReMe | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>
<strong>ReMe (formerly MemoryScope)为Agent设计的记忆管理框架</strong><br>
<em>Remember Me, Refine Me.</em>
</p>
<p align="center">
<strong>面向 AI Agent 的 local-first 自进化个人知识库。</strong><br>
</p>
---
ReMe为AI智能体提供了统一的记忆与经验系统——在跨用户、跨任务、跨智能体下抽取、复用和分享记忆的能力。
> 历史版本:[0.3.x](https://github.com/agentscope-ai/ReMe/tree/reme_v3) ·
> [0.2.x](https://github.com/agentscope-ai/ReMe/tree/v0.2.0.6) ·
> [MemoryScope](https://github.com/agentscope-ai/ReMe/tree/memoryscope_branch)
```
个性化记忆 (Personal Memory) + 任务经验 (Task Memory)= agent记忆
```
## ✨ 为什么选择 ReMe
个性化记忆能够"**理解用户偏好**"任务记忆让agent"**做得更好**"
🧠 ReMe 将对话和资料持续沉淀为可读、可编辑、可检索、相互链接的 Markdown 记忆。QwenPaw、DeepSeek Harness 等 Agent
可以共享同一个 workspace共同检索、维护和演化知识而持久文件始终由用户掌控。
- **Memory as File, File as Memory**ReMe 使用带 frontmatter 和 wikilink 的普通 Markdown 保存持久记忆。用户和 Agent
都可以使用熟悉的工具查看、编辑、移动、同步和备份;索引及生成的元数据均可重建。
- **自进化知识库**ReMe 将对话和资料逐步加工为 daily note 与长期知识,在保留来源的同时,持续提炼事实、偏好、
流程经验及其关系。
- **精准召回所需上下文。** ReMe 结合 BM25、可选 embedding 和 wikilink 展开,召回带行号的相关片段及其关系,无需把整个知识库塞入
Agent 上下文。
- **一个 workspace可供不同 Agent 共同使用。** 个人助理、coding agent 和其他 Agent runtime 可以通过原生集成、SKILL.md、CLI、
HTTP、MCP 或 Python API 共享同一个本地记忆空间。
<p align="center">
<img src="docs/figure/design-philosophy.svg" alt="ReMe 设计理念" width="92%">
</p>
---
## 📰 最新动态
- [2026.08] - 发布 [`@agentscope-ai/reme`](https://www.npmjs.com/package/@agentscope-ai/reme),提供统一 TypeScript HTTP
client以及 DeepSeek Harness 和 OpenClaw 的原生 ReMe 记忆集成。
- [2026.08] - 发布 [ReMe 博客](https://agentscope-ai.github.io/ReMe/?doc=zh-reme-blog),系统介绍本地优先的记忆架构、自进化工作流、混合检索、
主动发现与评测结果。
- [2026.08] - 基于 ReMe 的智能体工具使用
[经验驱动增强方法](https://reme.agentscope.io/?doc=toolmemory-zh)已发布,见
[arXiv:2608.03403](https://arxiv.org/abs/2608.03403)。
- [2026.07] - 新增可选插件:[每日论文](https://reme.agentscope.io/?doc=daily-paper-zh)用于论文发现与解析,
[Auto Fin](https://reme.agentscope.io/?doc=auto-fin-zh)用于研究最近 24 小时的主题相关财联社新闻,通过本地记忆搜索回顾历史材料并构建
wikilink。
- [2026.07] -
我们的论文 [Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution](https://aclanthology.org/2026.findings-acl.829/)
已被 Findings of ACL 2026 接收。
- **[2025-09]** 🎉 ReMe v0.1
正式发布整合任务记忆与个人记忆。如果想使用原始的memoryscope项目你可以在[MemoryScope](https://github.com/modelscope/Reme/tree/memoryscope_branch)
中找到。
- **[2025-09]** 🧪 我们在appworld, bfcl(v3)
以及frozenlake环境验证了任务记忆抽取与复用在Agent中的效果更多信息请查看 [appworld exp](docs/cookbook/appworld/quickstart.md), [bfcl exp](docs/cookbook/bfcl/quickstart.md)
和 [frozenlake exp](docs/cookbook/frozenlake/quickstart.md)。
- **[2025-08]** 🚀 MCP协议支持已上线-> [MCP指南](docs/mcp_quick_start.md)。
- **[2025-06]** 🚀 多后端向量存储支持 (Elasticsearch & ChromaDB) -> [向量数据库指南](docs/vector_store_api_guide.md)。
- **[2024-09]** 🧠 [MemoryScope](https://github.com/modelscope/Reme/tree/memoryscope_branch) v0.1 发布,个性化和时间感知的记忆存储与使用。
---
## ✨ 功能设计
<p align="center">
<img src="docs/figure/reme_structure.jpg" alt="ReMe Logo" width="100%">
</p>
ReMe整合两种互补的记忆能力
#### 🧠 **任务经验 (Task Memory/Experience)**
跨智能体复用的程序性知识
- **成功模式识别**:识别有效策略并理解其根本原理
- **失败分析学习**:从错误中学习,避免重复同样的问题
- **对比模式**:不同采样轨迹通过对比得到更有价值的经验
- **验证模式**:经过验证模块确认抽取记忆的有效性
你可以从[task memory](docs/task_memory/task_memory.md)了解更多如何使用task memory的方法
#### 👤 **个人记忆 (Personal Memory)**
特定用户的情境化记忆
- **个体偏好**:用户的习惯、偏好和交互风格
- **情境适应**:基于时间和上下文的智能记忆管理
- **渐进学习**:通过长期交互逐步建立深度理解
- **时间感知**:检索和整合时都具备时间敏感性
你可以从[personal memory](docs/personal_memory/personal_memory.md)了解更多如何使用personal memory的方法
---
## 🛠️ 安装
### 从PyPI安装推荐
```bash
pip install reme-ai
```
### 从源码安装
```bash
git clone https://github.com/modelscope/ReMe.git
cd ReMe
pip install .
```
### 环境配置
复制 `example.env` 为 .env并修改其中对应参数
```bash
FLOW_APP_NAME=ReMe
FLOW_LLM_API_KEY=sk-xxxx
FLOW_LLM_BASE_URL=https://xxxx/v1
FLOW_EMBEDDING_API_KEY=sk-xxxx
FLOW_EMBEDDING_BASE_URL=https://xxxx/v1
```
---
## 🚀 快速开始
### 安装
ReMe 要求 Python 3.11+。
从 pip 安装:
### HTTP服务启动
```bash
pip install "reme-ai[core]"
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen3-30b-a3b-thinking-2507 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
从源码安装:
### MCP服务器支持
```bash
git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
pip install -e reme_studio -e ".[core]"
cd reme_studio
npm ci
npm run build:static
cd ..
reme \
backend=mcp \
mcp.transport=stdio \
llm.default.model_name=qwen3-30b-a3b-thinking-2507 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
静态构建要求 Node.js 22.13 或更高版本,并让源码安装可以直接使用 Studio。
### 核心API使用
### 启动服务
#### 任务记忆管理
```python
import requests
```bash
reme start
# 经验总结器:从执行轨迹学习
response = requests.post("http://localhost:8002/summary_task_memory", json={
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [{"role": "user", "content": "帮我制定项目计划"}], "score": 1.0}
]
})
# 经验检索器:获取相关经验
response = requests.post("http://localhost:8002/retrieve_task_memory", json={
"workspace_id": "task_workspace",
"query": "如何高效管理项目进度?",
"top_k": 1
})
```
默认服务地址是 `127.0.0.1:2333`。如果端口被占用,可以指定其他端口:
<details>
<summary>curl 版本</summary>
```bash
reme start service.port=8181
# reme start workspace_dir=/tmp/reme-demo service.port=8181
# 经验总结器:从执行轨迹学习
curl -X POST http://localhost:8002/summary_task_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [{"role": "user", "content": "帮我制定项目计划"}], "score": 1.0}
]
}'
# 经验检索器:获取相关经验
curl -X POST http://localhost:8002/retrieve_task_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"query": "如何高效管理项目进度?",
"top_k": 1
}'
```
</details>
<details>
<summary>Node.js 版本</summary>
```javascript
// 经验总结器:从执行轨迹学习
fetch("http://localhost:8002/summary_task_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
trajectories: [
{messages: [{role: "user", content: "帮我制定项目计划"}], score: 1.0}
]
})
})
.then(response => response.json())
.then(data => console.log(data));
// 经验检索器:获取相关经验
fetch("http://localhost:8002/retrieve_task_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
query: "如何高效管理项目进度?",
top_k: 1
})
})
.then(response => response.json())
.then(data => console.log(data));
```
</details>
#### 个人记忆管理
```python
# 记忆整合:从用户交互中学习
response = requests.post("http://localhost:8002/summary_personal_memory", json={
"workspace_id": "task_workspace",
"trajectories": [
{"messages":
[
{"role": "user", "content": "我喜欢早上喝咖啡工作"},
{"role": "assistant", "content": "了解,您习惯早上用咖啡提神来开始工作"}
]
}
]
})
# 记忆检索:获取个人记忆片段
response = requests.post("http://localhost:8002/retrieve_personal_memory", json={
"workspace_id": "task_workspace",
"query": "用户的工作习惯是什么?",
"top_k": 5
})
```
<details>
<summary>curl 版本</summary>
```bash
reme version
reme health_check
reme help
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
# 记忆整合:从用户交互中学习
curl -X POST http://localhost:8002/summary_personal_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"trajectories": [
{"messages": [
{"role": "user", "content": "我喜欢早上喝咖啡工作"},
{"role": "assistant", "content": "了解,您习惯早上用咖啡提神来开始工作"}
]}
]
}'
# 记忆检索:获取个人记忆片段
curl -X POST http://localhost:8002/retrieve_personal_memory \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "task_workspace",
"query": "用户的工作习惯是什么?",
"top_k": 5
}'
```
</details>
### 5 分钟记忆 Demo
<details>
<summary>Node.js 版本</summary>
服务运行后,可以写入一个记忆节点,让 ReMe 索引并检索它:
```javascript
// 记忆整合:从用户交互中学习
fetch("http://localhost:8002/summary_personal_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
trajectories: [
{messages: [
{role: "user", content: "我喜欢早上喝咖啡工作"},
{role: "assistant", content: "了解,您习惯早上用咖啡提神来开始工作"}
]}
]
})
})
.then(response => response.json())
.then(data => console.log(data));
```bash
reme write \
path=digest/wiki/quick-start-demo \
name="Quick Start Demo" \
description="第一个 ReMe 记忆节点" \
content="# Quick Start Demo
ReMe 会把 Agent 记忆保存为可读的 Markdown。
相关链接:[[digest/wiki/memory-as-file.md]]"
reme search query="agent memory markdown" limit=5
reme read path=digest/wiki/quick-start-demo start_line=1 end_line=20
// 记忆检索:获取个人记忆片段
fetch("http://localhost:8002/retrieve_personal_memory", {
method: "POST",
headers: {
"Content-Type": "application/json",
},
body: JSON.stringify({
workspace_id: "task_workspace",
query: "用户的工作习惯是什么?",
top_k: 5
})
})
.then(response => response.json())
.then(data => console.log(data));
```
</details>
生成的文件是普通 Markdown并带有 frontmatter
```markdown
---
name: Quick Start Demo
description: 第一个 ReMe 记忆节点
---
# Quick Start Demo
## 📦 即用型经验库
ReMe 会把 Agent 记忆保存为可读的 Markdown。
ReMe提供预构建的经验库智能体可以立即使用经过验证的最佳实践
相关链接:[[digest/wiki/memory-as-file.md]]
### 可用经验库
- **`appworld.jsonl`**Appworld智能体交互的记忆库涵盖复杂任务规划和执行模式
- **`bfcl_v3.jsonl`**BFCL工具调用的工作记忆库
### 快速使用
```python
# 加载预构建经验
response = requests.post("http://localhost:8002/vector_store", json={
"workspace_id": "appworld",
"action": "load",
"path": "./docs/library/"
})
# 查询相关经验
response = requests.post("http://localhost:8002/retrieve_task_memory", json={
"workspace_id": "appworld",
"query": "如何导航到设置并更新用户资料?",
"top_k": 1
})
```
### ReMe Studio可选
## 🧪 实验
上面的 `core` 安装已包含 Studio。启动 ReMe 后,打开 <http://127.0.0.1:2333/> 即可浏览、编辑和搜索 workspace。
如需为基础安装单独添加 Studio可使用 `pip install "reme-ai[web]"`。源码构建、配置和开发说明见
[ReMe Studio 指南](https://reme.agentscope.io/?doc=studio-zh)。
### 🌍 [Appworld 实验](docs/cookbook/appworld/quickstart.md)
### 可选模型配置
我们在 Appworld 上使用 qwen3-8b 测试 ReMe
如果需要 LLM 驱动的记忆演化或 embedding 检索可以配置环境变量。embedding 默认关闭,因此默认配置不会启动 embedding 模型,也不需要
embedding API key。
| 方法 | pass@1 | pass@2 | pass@4 |
|--------------|-------------------|-------------------|-------------------|
| without ReMe | 0.083 | 0.140 | 0.228 |
| with ReMe | 0.109 **(+2.6%)** | 0.175 **(+3.5%)** | 0.281 **(+5.3%)** |
```bash
cat > .env <<'EOF'
# 可选:仅在配置中显式启用 embedding 组件后使用。
# EMBEDDING_API_KEY=sk-xxx
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
Pass@K 衡量的是在生成的 K 个样本中至少有一个成功完成任务score=1的概率。
当前实验使用的是一个内部的 AppWorld 环境,可能存在轻微差异。
# 必须auto_memory、auto_resource 和 auto_dream 需要 LLM。
LLM_API_KEY=sk-xxx
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
EOF
```
你可以在 [quickstart.md](docs/cookbook/appworld/quickstart.md) 中找到复现实验的更多细节。
基础文件读写、BM25 检索、wikilink 遍历和 proactive topics 读取可以先不配置 LLM 凭证。
> [!NOTE]
> 如需启用基于 embedding 的语义检索,请取消 [`reme/config/default.yaml`](reme/config/default.yaml) 中
> `components.as_embedding``components.embedding_store` 的注释,并将
> `components.file_store.default.embedding_store``""` 改为 `default`。完整说明见
> [记忆检索文档](docs/zh/memory_search.md)。
### 🧊 [Frozenlake 实验](docs/cookbook/frozenlake/quickstart.md)
## 🤝 将 ReMe 接入你的 Agent
| 不使用ReMe | 使用ReMe |
|:--------------------------------------------------------------------------------------------:|:--------------------------------------------------------------------------------------------:|
| <p align="center"><img src="docs/figure/frozenlake_failure.gif" alt="GIF 1" width="30%"></p> | <p align="center"><img src="docs/figure/frozenlake_success.gif" alt="GIF 2" width="30%"></p> |
ReMe 既可以作为本地记忆服务,通过 CLI、HTTP API 或 MCP server 接入,也可以通过 Python API 嵌入宿主进程。宿主集成可根据不同
runtime 的能力,将记忆指引、召回和捕获接入 Agent 生命周期。
我们在 100 个随机 frozenlake 地图上使用 qwen3-8b 进行测试:
| Agent | 推荐接入方式 | 接入后能力 |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| **DeepSeek Harness** | 使用 `dsh plugin --profile web add @agentscope-ai/reme` 安装 [`@agentscope-ai/reme`](typescript/README_ZH.md#deepseek-harness)。 | 长期记忆指引、`reme_search` 工具,以及自动捕获已完成的主 Agent 对话。 |
| **OpenClaw** | 使用 `openclaw plugins install @agentscope-ai/reme` 安装 [`@agentscope-ai/reme`](typescript/README_ZH.md#openclaw)。 | 原生记忆工具、用户触发运行前召回和自动对话捕获。 |
| **QwenPaw** | 通过 Python API 在进程内嵌入 ReMe。 | 复用宿主生命周期和模型配置,同时保持记忆本地、文件化。 |
| **Claude Code** | 启动 streamable HTTP MCP service并安装 [ReMe 插件](integrations/claude_code/reme)。 | MCP 召回工具、`reme-memory` skill以及自动记录会话的 Stop hook。 |
| **Hermes** | 启动 HTTP service并安装 [ReMe provider](integrations/hermes_agent)。 | 模型调用前召回,每轮对话完成后异步执行 `auto_memory`。 |
| **Codex 及其他 CLI Agent** | 安装或复制 [ReMe Memory skill](skills/reme_memory/SKILL.md)。 | 通过 CLI 搜索、读取和写入记忆;自动捕获需要显式接入宿主生命周期。 |
| 方法 | pass rate |
|--------------|------------------|
| without ReMe | 0.66 |
| with ReMe | 0.72 **(+6.0%)** |
<p align="center"><b>集成演示</b></p>
你可以在 [quickstart.md](docs/cookbook/frozenlake/quickstart.md) 中找到复现实验的更多细节。
<table>
<tr>
<td align="center"></td>
<td width="45%" align="center"><b>Auto Memory</b></td>
<td width="45%" align="center"><b>Auto Dream</b></td>
</tr>
<tr>
<td align="center"><b>QwenPaw</b></td>
<td width="45%">
<img src="docs/figure/qwenpaw-auto-memory.gif" alt="QwenPaw Auto Memory 演示" width="100%">
</td>
<td width="45%">
<img src="docs/figure/qwenpaw-auto-dream.gif" alt="QwenPaw Auto Dream 演示" width="100%">
</td>
</tr>
<tr>
<td align="center"><b>Claude Code</b></td>
<td width="45%">
<img src="docs/figure/cc-auto-memory.gif" alt="Claude Code Auto Memory 演示" width="100%">
</td>
<td width="45%">
<img src="docs/figure/cc-auto-dream.gif" alt="Claude Code Auto Dream 演示" width="100%">
</td>
</tr>
</table>
### 🔧 [BFCL-V3 实验](docs/cookbook/bfcl/quickstart.md)
## 🧠 ReMe 如何工作
我们在 BFCL-V3 multi-turn-base (随机划分50train/150val) 上使用 qwen3-8b 测试 ReMe
> Memory as File, File as Memory.
| 方法 | pass@1 | pass@2 | pass@4 |
|--------------|---------------------|---------------------|---------------------|
| without ReMe | 0.2472 | 0.2733 | 0.2922 |
| with ReMe | 0.3061 **(+5.89%)** | 0.3500 **(+7.67%)** | 0.3888 **(+9.66%)** |
ReMe 将 **记忆视为文件**,让过滤后的对话来源记录和外部资料从 `session/``resource/` 渐进加工到 `daily/`,再沉淀为
`digest/`。默认 workspace 是当前目录下的 `.reme/`;可通过 `workspace_dir=...` 选择其他由用户控制的位置。
## 📚 相关资源
### Workspace 结构
- **[快速开始](./cookbook/simple_demo)**:通过实际示例快速上手
- **[向量存储设置](docs/vector_store_api_guide.md)**:配置本地/向量数据库以及使用
- **[mcp指南](docs/mcp_quick_start.md)**创建mcp服务
- **[个性化记忆](docs/personal_memory)** 与 [任务记忆](docs/task_memory): 个性化记忆与任务记忆中分别使用的算子及其含义你可以修改config以自定义链路
- **[示例集合](./cookbook)**:实际用例和最佳实践
```text
<workspace_dir>/
├── metadata/ # 可重建的索引、图谱、catalog 和缓存
├── session/ # 对话来源记录和 Agent session
│ ├── dialog/
│ │ └── <session_id>.jsonl # auto_memory 保存的来源消息
│ └── claude_code/
│ └── <session_id>.jsonl # auto_memory_cc 使用的 ReMe 副本
├── mem_session/ # Agent wrapper 生成的 session/配置,不是用户记忆
│ ├── agentscope/
│ ├── claude_config/
│ └── codex/
├── resource/ # 外部原始材料
│ ├── <resource>.<ext> # 根目录文件进入当天 daily 层
│ └── YYYY-MM-DD/
│ └── <resource>.<ext>
├── daily/ # 浅加工记忆:当天事实、对话摘要、资源解读
│ ├── YYYY-MM-DD.md
│ └── YYYY-MM-DD/
│ ├── <generated_name>.md # 按主题命名的对话或资源卡片
│ └── interests.yaml
└── digest/ # 长期记忆:个人事实、流程经验、知识节点
├── personal/
│ └── {topic/event}.md
├── procedure/
│ └── {topic/event}.md
└── wiki/
└── {topic/event}.md
```
---
<p align="center">
<img src="docs/figure/reme-overview.svg" alt="ReMe 文件化记忆系统总览" width="92%">
</p>
## 🤝 贡献
### 记忆生命周期
我们相信最好的记忆系统来自集体智慧。欢迎贡献👉[指南](docs/contribution.md)
ReMe 遵循 capture → index → consolidate → recall 的循环。workspace 文件是持久化的事实来源,`metadata/` 中的内容均可重建。
### 代码贡献
- 新操作和工具开发
- 后端实现和优化
- API增强和新端点
| 能力 | 入口 | 作用 | 输出 |
| ------------------------------------------- | ----------------------------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
| [`auto_memory`](docs/zh/auto_memory.md) | Agent hook 或 `reme auto_memory` | 提炼有长期价值的对话事实,同时保留过滤后的对话来源记录。 | `session/dialog/*.jsonl``daily/<date>/<generated-name>.md` |
| [`auto_resource`](docs/zh/auto_resource.md) | 资源监听或 `reme auto_resource` | 将 `resource/` 下的文件转为带来源链接、按内容命名的 daily 卡片。 | `daily/<date>/<resource-card>.md` |
| [`auto_index`](docs/zh/memory_search.md) | 后台监听或 `reme reindex` | 实时索引 `daily/``digest/` 中的 Markdown全量重建还会扫描 `resource/` 和 JSONL。 | 可检索的 chunks、BM25、wikilink 图谱和可选向量 |
| [`auto_dream`](docs/zh/auto_dream.md) | `dream_cron``reme auto_dream` | 默认从最近两天内变化的文件中最多提取 5 个可复用 unit再创建、印证、补充或修正 digest 节点。 | `digest/**``daily/<date>/interests.yaml` |
| [`proactive`](docs/zh/proactive.md) | Agent 决定主动行动前调用 `reme proactive` | 读取 `auto_dream` 生成的 topics是否以及如何提醒用户由宿主 Agent 决定。 | 来自 `daily/<date>/interests.yaml` 的结构化 topics |
### 文档改进
- 使用示例和教程
- 最佳实践指南
<table>
<tr>
<td align="center" width="50%">
<img src="docs/figure/memory-as-file.svg" alt="Memory as File" width="92%">
</td>
<td align="center" width="50%">
<img src="docs/figure/auto-memory-resource.svg" alt="Auto Memory and Resource" width="92%">
</td>
</tr>
<tr>
<td align="center" width="50%">
<img src="docs/figure/auto-dream-and-proactive.svg" alt="Auto Dream and Proactive" width="92%">
</td>
<td align="center" width="50%">
<img src="docs/figure/auto-index-and-memory-search.svg" alt="Auto Index and Memory Search" width="92%">
</td>
</tr>
</table>
搜索返回带行号范围的相关 chunks 和数量受限的 wikilink 邻居;可选向量结果通过 RRF 与 BM25 融合。
> [!IMPORTANT]
>
> `proactive` 只读取并暴露 Auto Dream 生成的兴趣主题,不会自行联网、发送通知或改写知识库;是否以及如何使用主题,由宿主 Agent
> 决定。
## 📊 评测结果
ReMe 通过 Agent 多轮搜索与读取的方式评测多会话和超长上下文中的记忆能力。下表为仓库中已公开的参考实验结果模型、prompt、数据集和评判细节见各评测文档。
| 基准 | 设置 | 样本量 | Agentic 得分 | 主要检验内容 |
| --------------------------------------------------------------------------- | ----------- | ----------------: | -----------: | ------------------------------ |
| **[LongMemEval cleaned-s](https://reme.agentscope.io/?doc=longmemeval-zh)** | **整体** | **500 题** | **89.4%** | 跨会话检索、知识更新与时间推理 |
| [BEAM](https://reme.agentscope.io/?doc=beam-zh) | 100K 上下文 | 20 cases / 400 题 | 66.1% | 十类长上下文记忆任务 |
| [BEAM](https://reme.agentscope.io/?doc=beam-zh) | 1M 上下文 | 35 cases / 700 题 | 65.0% | 超长对话设置 |
在仓库的 [π-Bench 评测](https://reme.agentscope.io/?doc=pibench-zh)中ReMe Agent 在 5 种用户角色上的平均 **PROC 得分为 0.580**
,比相同测试模型配置的 NanoBot 高 2.4%。PROC 用于评估隐藏意图完成、针对性澄清、跨会话偏好和规范复用、跨任务依赖推断以及欠规格请求推进等主动性能力。
## 🧩 扩展与插件
插件是可选的独立 Python distribution可以贡献 Component、Step、Job backend 和配置,并通过配置显式启用。每日论文与 Auto Fin
均已独立打包,源码 distribution 及说明分别见[每日论文](plugins/daily_paper/README_ZH.md)和
[Auto Fin](plugins/auto-fin/README_ZH.md)。
| 插件 | 能力 |
| ---------------------------------------------------------- | ------------------------------------------------------------------------------ |
| [每日论文](https://reme.agentscope.io/?doc=daily-paper-zh) | 发现并排序论文,使用 Agent 解读 PDF生成文件化论文笔记和五分钟简报。 |
| [Auto Fin](https://reme.agentscope.io/?doc=auto-fin-zh) | 拉取主题相关财联社新闻,搜索 ReMe 历史材料并生成带 wikilink 的 Markdown 报告。 |
安装、查看、校验、启用和卸载 ReMe 插件的方法见[插件管理](docs/zh/plugin_management.md)。
## 📚 文档
下列文档覆盖主要使用流程,并以当前代码的运行时契约为准。
| 文档 | 主要内容 |
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------- |
| [快速开始](docs/zh/quick_start.md) | 安装 ReMe、启动服务并执行首次文件和记忆操作。 |
| [Memory as File](docs/zh/memory_as_file.md) | 理解 workspace 分层、frontmatter、wikilink、chunk 和文件事实来源模型。 |
| [Auto Memory](docs/zh/auto_memory.md) | 保留过滤后的对话来源记录,并提炼可复用的 daily 记忆卡片。 |
| [Auto Resource](docs/zh/auto_resource.md) | 导入支持的文本资料,转换为可追溯来源的 daily 卡片。 |
| [Auto Dream](docs/zh/auto_dream.md) 与 [Auto Link](docs/zh/auto_link.md) | 将 daily 记忆整理为持续演化的 digest 节点和可读 wikilink 关系。 |
| [记忆检索](docs/zh/memory_search.md) | 使用 BM25、可选向量、RRF 融合、行号范围召回和渐进式链接扩展。 |
| [Proactive](docs/zh/proactive.md) | 安全读取兴趣主题,并将其接入宿主 Agent 的决策流程。 |
| [应用场景](docs/zh/reme_scene.md) | 查看金融研究、研发记忆和个人知识库的完整使用示例。 |
| [框架说明](docs/zh/framework.md) | 理解 Application、Job、Step、Component、service、配置和生命周期边界。 |
| [TypeScript 集成](typescript/README_ZH.md) | 配置统一 client以及 DeepSeek Harness 和 OpenClaw 原生适配器。 |
| [ReMe 博客](https://agentscope-ai.github.io/ReMe/?doc=zh-reme-blog) | 了解完整产品故事、设计动机、使用示例和评测摘要。 |
## 🛠️ 常用命令
运行 `reme help` 可查看完整 job 列表。常用 workspace 与维护命令如下:
| 命令 | 作用 |
| ----------------------------------------- | ------------------------------------------------------------- |
| `reme status` | 查看有状态数据组件的内存估算及进程 RSS。 |
| [`reme search`](docs/zh/memory_search.md) | 默认使用 BM25 和 wikilink 检索,启用后增加向量检索。 |
| `reme read` / `reme write` / `reme edit` | 检查和维护 Markdown 记忆文件。 |
| `reme traverse` / `reme graph_snapshot` | 浏览 wikilink 邻域或按类别组织的 digest 图。 |
| `reme chat` | 与可感知 workspace 的只读 Agent 进行流式对话;需要 LLM 凭证。 |
| `reme reindex` | 基于已有文件重建检索和 wikilink 索引。 |
## 🤝 社区与贡献
- **问题反馈、需求与帮助**:请先查看 [Open Issues](https://github.com/agentscope-ai/ReMe/issues);如无相关讨论,可新建 Issue
说明背景、目标行为和影响范围。
- **代码贡献**:改动前建议阅读仓库内的[贡献指南](docs/zh/contributing.md)。架构与扩展方式以源码、schema 和测试为准。
- **文档贡献**:请直接更新本仓库 `docs/en/``docs/zh/` 或对应 package 目录中的规范源文件;文档站点会从这些文件生成。
- **提交规范**:建议使用 Conventional Commits例如 `feat(search): add link expansion option`
`docs(zh): update quick start`
- **提交前检查**:提交 PR 前请尽量运行 `pre-commit run --all-files``pytest`;如有依赖 LLM、embedding 或外部服务的测试无法运行,请在
PR 中说明。
- **项目文档**:访问 [reme.agentscope.io](https://reme.agentscope.io)。
### 贡献者
感谢所有为 ReMe 做出贡献的朋友们:
<a href="https://github.com/agentscope-ai/ReMe/graphs/contributors">
<img src="https://contrib.rocks/image?repo=agentscope-ai/ReMe" alt="贡献者" />
</a>
---
## 📄 引用
```bibtex
@software{ReMe2026,
title = {Remember me, Refine me: Memory Management Kit for Agents},
author = {ReMe Team},
url = {https://reme.agentscope.io},
year = {2026}
@software{ReMe2025,
title = {ReMe: Memory Management Framework for Agents},
author = {Li Yu, Jiaji Deng, Zouying Cao},
url = {https://github.com/modelscope/ReMe},
year = {2025}
}
```
---
## ⚖️ 许可证
本项目基于 Apache License 2.0 开源,详情参见 [LICENSE](./LICENSE) 文件。
本项目采用Apache License 2.0许可证 - 详情请参阅[LICENSE](./LICENSE)文件。
---
## Star 历史
[![Star History Chart](https://api.star-history.com/svg?repos=modelscope/ReMe&type=Date)](https://www.star-history.com/#modelscope/ReMe&Date)

View file

@ -1,124 +0,0 @@
[中文版 / Chinese version](./README_ZH.md)
# BEAM Benchmark
BEAM is a benchmark for **memory capability over long-context chat cases**. Each
case contains a very long chat history split into batches; ReMe converts each
batch into a session, ingests them in chronological order, then answers probing
questions via an agentic (ReAct) mode. Answers are scored with BEAM's
rubric-based `answer_judge` job, which produces both a graded score and a binary
verdict, and per-type averages are reported.
BEAM ships dataset variants by chat size — `100K` / `500K` / `1M` / `10M` — so
memory systems can be stressed at different context lengths. Question types
include abstention, contradiction resolution, event ordering, information
extraction, instruction following, knowledge update, multi-session reasoning,
preference following, summarization, and temporal reasoning.
> For the shared setup (dependencies, credentials, log conventions) see the
> [top-level benchmark README](../README.md).
## 1. Get the Dataset
BEAM is a public repository, cloned into `benchmark/beam/dataset/`:
```bash
mkdir -p benchmark/beam/dataset
cd benchmark/beam/dataset
git clone https://github.com/mohammadtavakoli78/BEAM.git
```
After cloning, `benchmark/beam/dataset/BEAM/` should contain `chats/`, `src/`,
`topics/` and other subdirectories.
## 2. Run
From the repository root:
```bash
python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q # quiet
python benchmark/beam/run.py --eval_only # reuse existing workspaces, query + judge only
```
## 3. Pipeline
1. For each case, load `chat.json` and convert each batch into a ReMe session.
2. Ingest sessions in chronological order into an isolated workspace, then `digest_update`.
3. Answer each probing question via agentic (ReAct) mode.
4. Score answers with BEAM's rubric-based `answer_judge` job and print per-type averages.
## 4. Key config — `benchmark/beam/config.yaml`
| Key | Meaning |
| --- | --- |
| `dataset.beam_root` | BEAM dataset root (`benchmark/beam/dataset/BEAM`). |
| `dataset.chat_size` | Variant to run: `100K` / `500K` / `1M` / `10M`. |
| `dataset.case_ids` | Specific cases (e.g. `["1","2"]`); empty = all cases. |
| `dataset.start_index` / `num_items` | Case pagination (`num_items` `0` = all). |
| `dataset.workspace_root` | Per-case workspace root (`benchmark/beam/workspaces/beam`). |
| `evaluation.num_workers` | `0` = auto, `1` = sequential, `>1` = parallel. |
| `reme.config` | ReMe config used (`beam.yaml`). |
| `output.dir` | Results directory (`benchmark/beam/results`). |
## 5. Outputs
Results are JSON files written to `output.dir` as
`results_<chat_size>_<timestamp>.json`, with a per-type score summary also
printed to the console. Logging conventions are shared across benchmarks — see
the [top-level README](../README.md#outputs--logs).
## 6. Reference Results
> The results below use the longmemeval-version prompt.
### 100K
agentscope==2.0.4.post1, conda reme env, 20 workers, eval-only (reusing prebuilt memory)
(2026-08-05, 20 cases / 400 Qs, total 46.0 min)
| Type | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|---|
| abstention | 0.550 | 0.550 | 96,031 | 1,070 | 97,101 | 4.58 |
| contradiction_resolution | 0.438 | 0.412 | 32,263 | 872 | 33,135 | 2.48 |
| event_ordering | 0.501 | 0.423 | 140,195 | 5,163 | 145,358 | 4.70 |
| information_extraction | 0.873 | 0.832 | 50,245 | 883 | 51,128 | 3.15 |
| instruction_following | 0.750 | 0.725 | 37,986 | 848 | 38,834 | 2.67 |
| knowledge_update | 0.688 | 0.675 | 31,198 | 651 | 31,849 | 2.27 |
| multi_session_reasoning | 0.626 | 0.584 | 85,038 | 4,563 | 89,601 | 4.28 |
| preference_following | 0.925 | 0.912 | 34,281 | 989 | 35,270 | 2.50 |
| summarization | 0.623 | 0.461 | 89,657 | 2,056 | 91,713 | 4.12 |
| temporal_reasoning | 0.637 | 0.625 | 34,563 | 1,049 | 35,612 | 2.52 |
| **OVERALL** | **0.661** | **0.620** | **63,146** | **1,814** | **64,960** | **3.33** |
Memory Construction average token consumption (default agent, full build over 20 cases):
| Agent | input tok/case | output tok/case | total tok/case |
|---|---|---|---|
| default | 2,172,316 | 136,697 | 2,309,013 |
### 1M
agentscope==2.0.4.post1, conda reme env, 20 workers, full memory build
(2026-08-05, 35 cases / 700 Qs, total 459.2 min)
| Type | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|---|
| abstention | 0.429 | 0.429 | 118,707 | 1,178 | 119,886 | 4.20 |
| contradiction_resolution | 0.391 | 0.364 | 49,787 | 810 | 50,597 | 2.50 |
| event_ordering | 0.558 | 0.456 | 201,514 | 3,889 | 205,403 | 4.79 |
| information_extraction | 0.809 | 0.772 | 78,950 | 894 | 79,844 | 3.00 |
| instruction_following | 0.852 | 0.832 | 55,757 | 924 | 56,681 | 2.81 |
| knowledge_update | 0.779 | 0.771 | 45,981 | 665 | 46,646 | 2.37 |
| multi_session_reasoning | 0.658 | 0.612 | 138,133 | 2,873 | 141,006 | 4.40 |
| preference_following | 0.798 | 0.777 | 51,796 | 920 | 52,716 | 2.53 |
| summarization | 0.693 | 0.537 | 158,794 | 2,905 | 161,700 | 4.44 |
| temporal_reasoning | 0.536 | 0.536 | 100,176 | 3,148 | 103,324 | 3.90 |
| **OVERALL** | **0.650** | **0.609** | **99,959** | **1,821** | **101,780** | **3.49** |
Memory Construction average token consumption (default agent, full build over 35 cases):
| Agent | input tok/case | output tok/case | total tok/case |
|---|---|---|---|
| default | 31,943,817 | 1,417,061 | 33,360,878 |

View file

@ -1,119 +0,0 @@
# BEAM 评测
[English version](./README.md)
BEAM 是一个面向**长上下文对话场景**的记忆能力评测基准。每个 case 包含一段被切分为多个
batch 的超长对话ReMe 将每个 batch 转换为一个会话,按时间顺序摄入后,以 agenticReAct
模式回答探测问题。答案由 BEAM 基于 rubric 的 `answer_judge` 任务打分,同时给出分级分数与二元
判定,并输出各类型平均分。
BEAM 按对话规模提供多种数据变体 —— `100K` / `500K` / `1M` / `10M`,可在不同上下文长度下
压测记忆系统。题型包括 abstention拒答、contradiction resolution矛盾消解、event
ordering事件排序、information extraction信息抽取、instruction following指令遵循
knowledge update知识更新、multi-session reasoning多会话推理、preference following
偏好遵循、summarization摘要与 temporal reasoning时间推理
> 公共设置(依赖、凭据、日志约定)见[总评测说明](../README_ZH.md)。
## 1. 获取数据集
BEAM 是公开仓库clone 到 `benchmark/beam/dataset/` 下:
```bash
mkdir -p benchmark/beam/dataset
cd benchmark/beam/dataset
git clone https://github.com/mohammadtavakoli78/BEAM.git
```
clone 完成后,`benchmark/beam/dataset/BEAM/` 目录下应包含 `chats/``src/``topics/` 等子目录。
## 2. 运行
在仓库根目录执行:
```bash
python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q # 安静模式
python benchmark/beam/run.py --eval_only # 复用已有工作区,仅执行查询 + 评判
```
## 3. 流程
1. 为每个 case 加载 `chat.json`,将每个 batch 转换为一个 ReMe 会话。
2. 按时间顺序将会话摄入独立工作区,随后执行 `digest_update`
3. 以 agenticReAct模式回答每个探测问题。
4. 通过 BEAM 基于 rubric 的 `answer_judge` 任务打分,并输出各类型平均分。
## 4. 关键配置 —— `benchmark/beam/config.yaml`
| 配置项 | 含义 |
| --- | --- |
| `dataset.beam_root` | BEAM 数据集根目录(`benchmark/beam/dataset/BEAM`)。 |
| `dataset.chat_size` | 运行的变体:`100K` / `500K` / `1M` / `10M`。 |
| `dataset.case_ids` | 指定 case`["1","2"]`),空表示全部。 |
| `dataset.start_index` / `num_items` | case 分页(`num_items``0` 表示全部)。 |
| `dataset.workspace_root` | case 工作区根目录(`benchmark/beam/workspaces/beam`)。 |
| `evaluation.num_workers` | `0` = 自动,`1` = 串行,`>1` = 并行。 |
| `reme.config` | 使用的 ReMe 配置(`beam.yaml`)。 |
| `output.dir` | 结果目录(`benchmark/beam/results`)。 |
## 5. 输出
结果以 JSON 文件写入 `output.dir`,文件名为 `results_<chat_size>_<timestamp>.json`
同时控制台会打印含各类型分数的汇总。日志约定在各基准间通用,见
[总说明](../README_ZH.md#输出与日志)。
## 6. 参考结果
> 以下结果使用 longmemeval 版本的 prompt。
### 100K
agentscope==2.0.4.post1conda reme 环境20 并发eval-only复用已构建 memory
2026-08-0520 cases / 400 Qs总耗时 46.0 min
| 题型 | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|---|
| abstention | 0.550 | 0.550 | 96,031 | 1,070 | 97,101 | 4.58 |
| contradiction_resolution | 0.438 | 0.412 | 32,263 | 872 | 33,135 | 2.48 |
| event_ordering | 0.501 | 0.423 | 140,195 | 5,163 | 145,358 | 4.70 |
| information_extraction | 0.873 | 0.832 | 50,245 | 883 | 51,128 | 3.15 |
| instruction_following | 0.750 | 0.725 | 37,986 | 848 | 38,834 | 2.67 |
| knowledge_update | 0.688 | 0.675 | 31,198 | 651 | 31,849 | 2.27 |
| multi_session_reasoning | 0.626 | 0.584 | 85,038 | 4,563 | 89,601 | 4.28 |
| preference_following | 0.925 | 0.912 | 34,281 | 989 | 35,270 | 2.50 |
| summarization | 0.623 | 0.461 | 89,657 | 2,056 | 91,713 | 4.12 |
| temporal_reasoning | 0.637 | 0.625 | 34,563 | 1,049 | 35,612 | 2.52 |
| **OVERALL** | **0.661** | **0.620** | **63,146** | **1,814** | **64,960** | **3.33** |
Memory Construction 平均 token 消耗default agent20 cases 全量构建):
| Agent | input tok/case | output tok/case | total tok/case |
|---|---|---|---|
| default | 2,172,316 | 136,697 | 2,309,013 |
### 1M
agentscope==2.0.4.post1conda reme 环境20 并发,全量构建 memory
2026-08-0535 cases / 700 Qs总耗时 459.2 min
| 题型 | Agentic | Binary | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|---|
| abstention | 0.429 | 0.429 | 118,707 | 1,178 | 119,886 | 4.20 |
| contradiction_resolution | 0.391 | 0.364 | 49,787 | 810 | 50,597 | 2.50 |
| event_ordering | 0.558 | 0.456 | 201,514 | 3,889 | 205,403 | 4.79 |
| information_extraction | 0.809 | 0.772 | 78,950 | 894 | 79,844 | 3.00 |
| instruction_following | 0.852 | 0.832 | 55,757 | 924 | 56,681 | 2.81 |
| knowledge_update | 0.779 | 0.771 | 45,981 | 665 | 46,646 | 2.37 |
| multi_session_reasoning | 0.658 | 0.612 | 138,133 | 2,873 | 141,006 | 4.40 |
| preference_following | 0.798 | 0.777 | 51,796 | 920 | 52,716 | 2.53 |
| summarization | 0.693 | 0.537 | 158,794 | 2,905 | 161,700 | 4.44 |
| temporal_reasoning | 0.536 | 0.536 | 100,176 | 3,148 | 103,324 | 3.90 |
| **OVERALL** | **0.650** | **0.609** | **99,959** | **1,821** | **101,780** | **3.49** |
Memory Construction 平均 token 消耗default agent35 cases 全量构建):
| Agent | input tok/case | output tok/case | total tok/case |
|---|---|---|---|
| default | 31,943,817 | 1,417,061 | 33,360,878 |

View file

@ -1,24 +0,0 @@
# BEAM evaluation configuration
# This file controls what/how to evaluate.
dataset:
beam_root: "benchmark/beam/dataset/BEAM" # BEAM dataset root
chat_size: "1M" # 100K | 500K | 1M | 10M (dataset variant)
case_ids: [] # empty = all cases; or ["1", "2", "3"]
start_index: 0 # first case index (for pagination)
num_items: 0 # 0 = all cases; >0 = limit
workspace_root: "benchmark/beam/workspaces/beam" # workspace root for case workspaces
evaluation:
num_workers: 20 # 0 = auto; 1 = sequential; >1 = parallel (per-case)
compress_session: false # true = compress session chunks in search_v2 (query-aware); false = no compression
reme:
config: "beam.yaml" # reme config (in reme/config/)
output:
dir: "benchmark/beam/results"
log_dir: "logs" # log directory (relative to project root)
log_prefix: "beam" # benchmark name used in log filenames
log_to_console: true
log_to_file: true

View file

@ -1,76 +0,0 @@
#!/bin/bash
# 杀死指定进程及其所有子进程
# Usage: bash kill.sh <PID>
if [ -z "$1" ]; then
echo "Usage: bash kill.sh <PID>"
echo " 杀死指定进程及其所有子进程"
exit 1
fi
PID=$1
# 检查进程是否存在
if ! kill -0 "$PID" 2>/dev/null; then
echo "进程 $PID 不存在"
exit 1
fi
# 递归收集所有子进程(包括子进程的子进程)
collect_children() {
local parent=$1
local children
children=$(ps -o pid= --ppid "$parent" 2>/dev/null | tr -d ' ')
for child in $children; do
collect_children "$child"
done
echo "$parent"
}
# 收集进程树(子进程在前,父进程在后,保证先杀子再杀父)
PROCESS_TREE=$(collect_children "$PID")
TOTAL=$(echo "$PROCESS_TREE" | wc -l | tr -d ' ')
echo "进程树(共 $TOTAL 个进程):"
while read -r p; do
cmd=$(ps -o args= -p "$p" 2>/dev/null | head -c 80)
printf " PID=%-8s %s\n" "$p" "$cmd"
done <<< "$PROCESS_TREE"
# 先 SIGTERM 优雅终止
echo ""
echo "发送 SIGTERM..."
while read -r p; do
kill "$p" 2>/dev/null
done <<< "$PROCESS_TREE"
# 等待最多 5 秒
for i in $(seq 1 5); do
alive=false
while read -r p; do
if kill -0 "$p" 2>/dev/null; then
alive=true
fi
done <<< "$PROCESS_TREE"
if [ "$alive" = false ]; then
break
fi
sleep 1
done
# 检查是否还有残留,强制 SIGKILL
remaining=false
while read -r p; do
if kill -0 "$p" 2>/dev/null; then
remaining=true
fi
done <<< "$PROCESS_TREE"
if [ "$remaining" = true ]; then
echo "部分进程未响应,发送 SIGKILL..."
while read -r p; do
kill -9 "$p" 2>/dev/null
done <<< "$PROCESS_TREE"
fi
echo "已终止进程树(根 PID=$PID,共 $TOTAL 个进程)"

View file

@ -1,891 +0,0 @@
"""BEAM evaluation runner for ReMe.
Evaluates ReMe's memory capability using the BEAM dataset.
Each case gets an isolated workspace; chat.json batches are ingested as
sessions in chronological order; finally probing questions are answered
via an agentic (ReAct) approach, then
judged by BEAM's rubric-based LLM-as-judge.
Usage:
python benchmark/beam/run.py
python benchmark/beam/run.py --config benchmark/beam/config.yaml
python benchmark/beam/run.py -q # quiet: only eval-level logs
python benchmark/beam/run.py --log-level WARNING # reduce eval runner logs
python benchmark/beam/run.py --reme-log-level WARNING # reduce reme internal logs
python benchmark/beam/run.py --eval_only # query+judge only, reuse existing workspace
"""
import json
import logging
import os
import re
import shutil
import time
import threading
from datetime import datetime
from pathlib import Path
import yaml
from dotenv import load_dotenv
# Load .env from project root
_PROJECT_ROOT = Path(__file__).parent.parent.parent
load_dotenv(_PROJECT_ROOT / ".env")
# Workspace root — read from config.yaml (dataset.workspace_root)
_WORKSPACE_ROOT_DEFAULT = "benchmark/beam/workspaces/beam"
# ---------------------------------------------------------------------------
# Logging
# ---------------------------------------------------------------------------
_DEFAULT_LOG_FORMAT = "%(asctime)s | %(levelname)s | %(message)s"
logging.basicConfig(level=logging.INFO, format=_DEFAULT_LOG_FORMAT)
logger = logging.getLogger("beam")
# Noisy library loggers silenced by default
_NOISY_LOGGERS = [
"httpx",
"httpcore",
"openai",
"uvicorn",
"multipart",
"asyncio",
"watchfiles",
"filelock",
]
def setup_logging(
log_level: str,
reme_log_level: str,
log_dir: str | None = None,
):
"""Configure logging for the eval runner and reme internals.
Args:
log_level: Level for the eval runner logger (DEBUG/INFO/WARNING/ERROR).
reme_log_level: Level for reme's internal loguru logger.
log_dir: Per-run log directory (absolute path). None = no file logging.
"""
numeric = getattr(logging, log_level.upper(), logging.INFO)
# Eval runner logger
logging.getLogger().setLevel(numeric)
logger.setLevel(numeric)
# Suppress noisy library loggers when above DEBUG
if numeric > logging.DEBUG:
for name in _NOISY_LOGGERS:
lib_logger = logging.getLogger(name)
lib_logger.setLevel(max(numeric, logging.WARNING))
# Add file handler for eval runner if log_dir is specified
if log_dir:
os.makedirs(log_dir, exist_ok=True)
log_filepath = os.path.join(log_dir, "runner.log")
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
file_handler.setLevel(numeric)
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
logging.getLogger().addHandler(file_handler)
logger.info(f"Eval runner log file: {log_filepath}")
# Reme internal logger (loguru) — will be applied per-worker via _configure_worker
os.environ["REME_LOG_LEVEL"] = reme_log_level.upper()
if log_dir:
os.environ["REME_LOG_DIR"] = log_dir
def _configure_worker(
log_level: str,
reme_log_level: str,
log_dir: str | None = None,
):
"""Set up logging inside a multiprocessing worker process.
Must be called at the top of each worker because child processes inherit
parent state but loguru sinks are NOT shared across fork/spawn.
"""
numeric = getattr(logging, log_level.upper(), logging.INFO)
logging.basicConfig(level=numeric, format=_DEFAULT_LOG_FORMAT, force=True)
logging.getLogger("beam").setLevel(numeric)
if numeric > logging.DEBUG:
for name in _NOISY_LOGGERS:
logging.getLogger(name).setLevel(max(numeric, logging.WARNING))
# Add file handler for eval runner in worker process
if log_dir:
os.makedirs(log_dir, exist_ok=True)
pid = os.getpid()
log_filepath = os.path.join(log_dir, f"worker-{pid}.log")
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
file_handler.setLevel(numeric)
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
logging.getLogger().addHandler(file_handler)
# Re-initialize loguru for reme internals at the desired level
from reme.utils import get_logger
reme_log_dir = log_dir or "logs"
get_logger(log_dir=reme_log_dir, level=reme_log_level.upper(), force_init=True)
# ---------------------------------------------------------------------------
# Config loading
# ---------------------------------------------------------------------------
def load_eval_config(config_path: str | None = None) -> dict:
"""Load evaluation config yaml with env-var expansion."""
if config_path is None:
config_path = str(Path(__file__).parent / "config.yaml")
with open(config_path, encoding="utf-8") as f:
raw = f.read()
# Expand ${VAR} and ${VAR:-default}
def _expand(m):
expr = m.group(1)
if ":-" in expr:
key, default = expr.split(":-", 1)
return os.environ.get(key, default)
return os.environ.get(expr, "")
raw = re.sub(r"\$\{([^}]+)\}", _expand, raw)
return yaml.safe_load(raw)
# ---------------------------------------------------------------------------
# BEAM data loading
# ---------------------------------------------------------------------------
def parse_beam_time_anchor(time_str: str) -> datetime:
"""Parse BEAM time_anchor format: 'March-15-2024' -> datetime."""
for fmt in ("%B-%d-%Y", "%b-%d-%Y"):
try:
return datetime.strptime(time_str, fmt)
except ValueError:
continue
raise ValueError(f"Cannot parse time_anchor: {time_str!r}")
def load_beam_chat(chat_path: Path, chat_size: str, case_id: str) -> list[dict]:
"""Load BEAM chat.json and convert to ReMe session format.
Each batch becomes one session with all its turns flattened.
Each turn resolves its own time_anchor independently; turns without
an explicit time_anchor inherit from the most recent preceding turn.
Returns list of sessions, each with:
- session_id: str
- date: str (YYYY-MM-DD) derived from the *first* turn's time
- messages: list[dict] with name, role, content, created_at
"""
with open(chat_path, encoding="utf-8") as f:
batches = json.load(f)
sessions = []
for batch in batches:
batch_num = batch["batch_number"]
# Resolve batch-level fallback (used when no turn has a time_anchor)
batch_anchor = batch.get("time_anchor")
if not batch_anchor:
batch_anchor = "January-1-2024"
# Flatten all turns, resolving time_anchor per turn
messages = []
prev_dt = None # carries forward from previous turn
first_dt = None # for session-level date
for turn in batch["turns"]:
# Find this turn's own time_anchor from its messages
turn_anchor = None
for msg in turn:
if msg.get("time_anchor"):
turn_anchor = msg["time_anchor"]
break
if turn_anchor:
dt = parse_beam_time_anchor(turn_anchor)
elif prev_dt is not None:
dt = prev_dt # inherit from previous turn
else:
dt = parse_beam_time_anchor(batch_anchor)
if first_dt is None:
first_dt = dt
prev_dt = dt
for msg in turn:
role = msg["role"]
messages.append(
{
"name": role,
"role": role,
"content": msg["content"],
"created_at": dt.strftime("%Y-%m-%dT%H:%M:%S"),
},
)
sessions.append(
{
"session_id": f"beam_{chat_size}_{case_id}_batch{batch_num}",
"date": first_dt.strftime("%Y-%m-%d"),
"messages": messages,
},
)
return sessions
def get_available_cases(beam_root: Path, chat_size: str) -> list[str]:
"""Return sorted list of case IDs for a given chat size."""
chats_dir = beam_root / "chats" / chat_size
if not chats_dir.exists():
return []
return sorted(
[d.name for d in chats_dir.iterdir() if d.is_dir()],
key=int,
)
# ---------------------------------------------------------------------------
# Answer generation
# ---------------------------------------------------------------------------
async def answer_question_agentic(app, question: str, compress_session: bool = False) -> tuple[str, dict]:
"""Answer a probing question using ReMe's agentic_answer job.
Returns (answer, metadata)
"""
from reme.utils.evaluation_interface import track_agent_token_usage, track_job_counts
with (
track_job_counts(["search"], app.context) as tool_counts,
track_agent_token_usage(
["bench"],
app.context,
) as token_usages,
):
query_resp = await app.run_job(
"agentic_answer",
query=question,
compress_session=compress_session,
)
answer = (query_resp.answer or "").strip()
return answer, {
"mode": "agentic",
"tool_counts": tool_counts,
"token_usage": token_usages["bench"],
}
# ---------------------------------------------------------------------------
# BEAM rubric-based LLM-as-Judge
# ---------------------------------------------------------------------------
async def judge_answer(
app,
question: str,
llm_response: str,
rubric: list[str],
question_type: str = "",
) -> dict:
"""Judge an answer via the answer_judge job (beam_rubric_judge_step)."""
judge_resp = await app.run_job(
"answer_judge",
llm_response=llm_response,
rubric=rubric,
probing_question=question,
question_type=question_type,
)
result = {
"llm_judge_score": (judge_resp.metadata or {}).get("llm_judge_score", 0.0),
"llm_judge_responses": (judge_resp.metadata or {}).get("llm_judge_responses", []),
}
# Include event_ordering extra metrics if present
eo = (judge_resp.metadata or {}).get("event_ordering")
if eo:
result["event_ordering"] = eo
return result
# ---------------------------------------------------------------------------
# Main evaluation pipeline
# ---------------------------------------------------------------------------
async def evaluate_case(eval_config: dict, case_id: str, eval_only: bool = False) -> dict:
"""Evaluate a single BEAM case end-to-end.
Args:
eval_config: The evaluation configuration dict.
case_id: The case directory name (e.g. "1").
eval_only: If True, skip ingestion and only run query+judge
using the existing workspace.
Returns:
A results dict with all questions, answers, and judgments.
"""
from reme import Application
from reme.config import resolve_app_config
dataset_cfg = eval_config["dataset"]
chat_size = dataset_cfg["chat_size"]
compress_session = bool(eval_config["evaluation"].get("compress_session", False))
beam_root = _PROJECT_ROOT / dataset_cfg.get("beam_root", "benchmark/beam/dataset/BEAM")
chat_path = beam_root / "chats" / chat_size / case_id / "chat.json"
probing_questions_path = beam_root / "chats" / chat_size / case_id / "probing_questions" / "probing_questions.json"
if not chat_path.exists():
raise FileNotFoundError(f"Chat file not found: {chat_path}")
if not probing_questions_path.exists():
raise FileNotFoundError(f"Probing questions not found: {probing_questions_path}")
logger.info(
"[Case %s] size=%s%s",
case_id,
chat_size,
" [eval_only]" if eval_only else "",
)
# Workspace setup
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
case_dir = workspace_root / f"{chat_size}_{case_id}"
workspace_dir = str(case_dir / ".reme")
if eval_only:
if not case_dir.exists() or not Path(workspace_dir).exists():
raise FileNotFoundError(
f"[Case {case_id}] eval_only: workspace not found at {case_dir}. "
f"Run without --eval_only first to build the workspace.",
)
else:
if case_dir.exists():
shutil.rmtree(case_dir)
logger.info(f"[Case {case_id}] Cleaned existing workspace: {case_dir}")
else:
logger.info(f"[Case {case_id}] Workspace not found, creating: {case_dir}")
case_dir.mkdir(parents=True, exist_ok=True)
# Pre-initialize ReMe's loguru logger with the correct log_dir
output_cfg = eval_config.get("output", {})
if output_cfg.get("log_to_file", False):
reme_log_dir = os.environ.get("REME_LOG_DIR")
if reme_log_dir:
from reme.utils import get_logger
get_logger(
log_dir=reme_log_dir,
level=os.environ.get("REME_LOG_LEVEL", "INFO"),
log_to_console=output_cfg.get("log_to_console", True),
log_to_file=True,
force_init=True,
)
cfg = resolve_app_config(
config=eval_config["reme"]["config"],
workspace_dir=workspace_dir,
log_to_console=output_cfg.get("log_to_console", True),
log_to_file=output_cfg.get("log_to_file", False),
enable_logo=False,
)
app = Application(**cfg)
await app.start()
from reme.utils.evaluation_interface import check_agent_token_usage # noqa: E402
_MEM_AGENT_NAMES = ("default", "bench")
sessions_ingested = 0
memory_token_usage: dict[str, dict[str, int | None]] = {}
try:
if not eval_only:
# ── Phase 1: Ingest sessions (with token tracking) ─────────
sessions = load_beam_chat(chat_path, chat_size, case_id)
logger.info(f"[Case {case_id}] Loaded {len(sessions)} sessions from chat.json")
# Snapshot token counters before memory construction
mem_token_start = {name: check_agent_token_usage(name, app.context) for name in _MEM_AGENT_NAMES}
for i, session in enumerate(sessions):
logger.info(
f"[Case {case_id}] Ingesting session {i+1}/{len(sessions)}: "
f"id={session['session_id']} date={session['date']} "
f"msgs={len(session['messages'])}",
)
resp = await app.run_job(
"auto_memory",
messages=session["messages"],
session_id=session["session_id"],
date=session["date"],
)
if not resp.success:
logger.warning(f"[Case {case_id}] auto_memory failed: {resp.answer}")
else:
logger.info(
f"[Case {case_id}] auto_memory success: " f"{resp.answer[:100] if resp.answer else ''}",
)
await app.run_job("index_update")
sessions_ingested += 1
# Final digest update
logger.info(f"[Case {case_id}] Running digest_update...")
await app.run_job("digest_update")
logger.info(f"[Case {case_id}] Ingestion complete.")
# Compute memory construction token deltas
for name in _MEM_AGENT_NAMES:
end_usage = check_agent_token_usage(name, app.context)
delta: dict[str, int | None] = {}
for metric in _TOKEN_USAGE_METRICS:
current = end_usage[metric]
start = mem_token_start[name][metric]
delta[metric] = None if current is None else current - (start or 0)
memory_token_usage[name] = delta
logger.info(f"[Case {case_id}] Memory construction token usage: {memory_token_usage}")
# ── Phase 2: Answer + Judge probing questions ───────────────
with open(probing_questions_path, encoding="utf-8") as f:
probing_questions = json.load(f)
total_questions = sum(len(v) for v in probing_questions.values())
logger.info(f"[Case {case_id}] Total probing questions: {total_questions}")
all_question_results = []
q_idx = 0
for q_type in probing_questions:
logger.info(
f"[Case {case_id}] Question type: {q_type} " f"({len(probing_questions[q_type])} questions)",
)
for i, q in enumerate(probing_questions[q_type]):
q_idx += 1
question = q["question"]
rubric = q.get("rubric", [])
logger.info(
f"[Case {case_id}] [{q_idx}/{total_questions}] " f"{q_type} Q{i+1}: {question[:100]}...",
)
q_result = {
"question_type": q_type,
"question_index": i,
"question": question,
"rubric": rubric,
}
# Agentic answer
try:
agentic_answer, agentic_meta = await answer_question_agentic(
app,
question,
compress_session=compress_session,
)
except Exception as e:
logger.error(f"[Case {case_id}] Agentic answer failed: {e}")
agentic_answer = f"(error: {e})"
agentic_meta = {"error": str(e)}
if not agentic_answer:
agentic_answer = "(no answer generated)"
logger.info(f"[Case {case_id}] Agentic answer: {agentic_answer[:200]}...")
logger.info(
f"[Case {case_id}] Agentic tool calls: {agentic_meta.get('tool_counts', {})}",
)
logger.info(f"[Case {case_id}] Bench token usage: {agentic_meta.get('token_usage', {})}")
# Judge agentic answer
logger.info(f"[Case {case_id}] Judging agentic ({q_type})...")
agentic_judgment = await judge_answer(
app,
question,
agentic_answer,
rubric,
question_type=q_type,
)
logger.info(
f"[Case {case_id}] Agentic score: " f"{agentic_judgment['llm_judge_score']:.3f}",
)
q_result["agentic_response"] = agentic_answer
q_result["agentic_judgment"] = agentic_judgment
q_result["agentic_metadata"] = agentic_meta
all_question_results.append(q_result)
finally:
await app.close()
return {
"case_id": case_id,
"chat_size": chat_size,
"sessions_ingested": sessions_ingested,
"total_questions": len(all_question_results),
"questions": all_question_results,
"memory_token_usage": memory_token_usage,
}
# ---------------------------------------------------------------------------
# Worker: runs a single case in its own process with its own event loop
# ---------------------------------------------------------------------------
def _evaluate_case_worker(task_input: tuple) -> dict:
"""Worker function for multiprocessing. Each process gets its own event loop."""
eval_config, case_id, log_level, reme_log_level, eval_only, log_dir = task_input
import asyncio # pylint: disable=import-outside-toplevel
_configure_worker(log_level, reme_log_level, log_dir=log_dir)
# Suppress httpx GC noise
logging.getLogger("asyncio").setLevel(logging.CRITICAL)
return asyncio.run(evaluate_case(eval_config, case_id, eval_only=eval_only))
def _indexed_worker(indexed_input: tuple) -> tuple:
"""Module-level wrapper for imap_unordered with index tracking."""
idx, task_input = indexed_input
return idx, _evaluate_case_worker(task_input)
def _resolve_num_workers(configured: int) -> int:
"""Resolve num_workers: 0=auto (cpu_count-2, min 1), 1=sequential, >1=parallel."""
if configured == 0:
return max(1, (os.cpu_count() or 4) - 2)
return max(1, configured)
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main( # pylint: disable=too-many-statements
config_path: str | None = None,
log_level: str = "INFO",
reme_log_level: str = "INFO",
eval_only: bool = False,
):
"""Run the BEAM evaluation pipeline.
Args:
config_path: Path to the YAML config file.
log_level: Log level for the eval runner.
reme_log_level: Log level for reme internal logs.
eval_only: If True, skip ingestion and only run query+judge using
existing workspaces.
"""
from multiprocessing import Pool # pylint: disable=import-outside-toplevel
# Load config BEFORE logging setup so log_dir is available
eval_config = load_eval_config(config_path)
# Resolve per-run log directory from config
output_cfg = eval_config.get("output", {})
log_dir_abs = None
if output_cfg.get("log_to_file", False):
log_dir_raw = output_cfg.get("log_dir", "logs")
log_prefix = output_cfg.get("log_prefix", "beam")
run_ts = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
log_dir_abs = str(_PROJECT_ROOT / log_dir_raw / f"{log_prefix}_{run_ts}")
setup_logging(log_level, reme_log_level, log_dir=log_dir_abs)
dataset_cfg = eval_config["dataset"]
chat_size = dataset_cfg["chat_size"]
beam_root = _PROJECT_ROOT / dataset_cfg.get("beam_root", "benchmark/beam/dataset/BEAM")
# Determine which cases to run
case_ids = dataset_cfg.get("case_ids") or []
if not case_ids:
case_ids = get_available_cases(beam_root, chat_size)
# Pagination
start = dataset_cfg.get("start_index", 0)
num_items = dataset_cfg.get("num_items", 0)
if num_items > 0:
case_ids = case_ids[start : start + num_items]
elif start > 0:
case_ids = case_ids[start:]
if not case_ids:
logger.error(f"No cases found for chat_size={chat_size}")
return
logger.info(
"Evaluating %d case(s) for chat_size=%s: %s%s",
len(case_ids),
chat_size,
case_ids,
" [eval_only: query+judge only]" if eval_only else "",
)
# Resolve parallelism
num_workers = _resolve_num_workers(eval_config["evaluation"].get("num_workers", 1))
logger.info(f"Using {num_workers} worker(s)")
# Create output directory
output_dir = _PROJECT_ROOT / output_cfg.get("dir", "benchmark/beam/results")
output_dir.mkdir(parents=True, exist_ok=True)
# Create workspace root directory
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
workspace_root.mkdir(parents=True, exist_ok=True)
# Pre-check: verify all workspaces exist in eval_only mode
if eval_only:
missing_cases = []
for case_id in case_ids:
case_dir = workspace_root / f"{chat_size}_{case_id}"
if not case_dir.exists() or not (case_dir / ".reme").exists():
missing_cases.append(case_id)
if missing_cases:
preview = missing_cases[:10]
suffix = "..." if len(missing_cases) > 10 else ""
raise FileNotFoundError(
f"eval_only: {len(missing_cases)} workspace(s) not found under {workspace_root}. "
f"Missing cases: {preview}{suffix}. "
f"Run without --eval_only first to build the workspaces.",
)
# Build task args
task_args = [(eval_config, case_id, log_level, reme_log_level, eval_only, log_dir_abs) for case_id in case_ids]
# Progress tracking
total_items = len(task_args)
completed_count = [0]
start_time = time.time()
progress_lock = threading.Lock()
def _print_progress(prefix: str = "PROGRESS"):
elapsed = time.time() - start_time
elapsed_min = elapsed / 60
done = completed_count[0]
pct = 100.0 * done / total_items if total_items else 0
eta_str = "N/A"
if done > 0:
eta_sec = elapsed / done * (total_items - done)
eta_str = f"{eta_sec/60:.1f}min"
print(
f"[{prefix}] {datetime.now().strftime('%Y-%m-%d %H:%M:%S')} | "
f"{done}/{total_items} ({pct:.1f}%) completed | "
f"elapsed={elapsed_min:.1f}min | ETA={eta_str}",
flush=True,
)
def _progress_timer():
"""Background thread: print progress every 10 minutes."""
while not _timer_stop.is_set():
_timer_stop.wait(600)
if not _timer_stop.is_set():
with progress_lock:
_print_progress()
_timer_stop = threading.Event()
timer_thread = threading.Thread(target=_progress_timer, daemon=True)
timer_thread.start()
# Run evaluation
if num_workers == 1:
results = []
for task_input in task_args:
result = _evaluate_case_worker(task_input)
results.append(result)
with progress_lock:
completed_count[0] += 1
else:
results = [None] * total_items
indexed_args = list(enumerate(task_args))
with Pool(processes=num_workers) as pool:
for idx, result in pool.imap_unordered(_indexed_worker, indexed_args):
results[idx] = result
with progress_lock:
completed_count[0] += 1
# Stop progress timer
_timer_stop.set()
timer_thread.join(timeout=2)
# Save results
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
output_file = output_dir / f"results_{chat_size}_{timestamp}.json"
with open(output_file, "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
logger.info(f"Results saved to {output_file}")
# Final progress
_print_progress("FINAL")
# Print concise summary
print("\n" + "=" * 70)
print(f" BEAM EVALUATION RESULTS | size={chat_size} cases={len(results)}")
print("=" * 70)
# Per-type stats (agentic only)
type_scores: dict[str, list[float]] = {}
type_binary_scores: dict[str, list[float]] = {}
all_scores: list[float] = []
all_binary_scores: list[float] = []
all_tool_call_totals: list[int] = []
all_token_usages: list[dict[str, int | None]] = []
all_memory_token_usages: list[dict[str, dict[str, int | None]]] = []
for case_result in results:
if "error" in case_result:
continue
mem_usage = case_result.get("memory_token_usage", {})
if mem_usage:
all_memory_token_usages.append(mem_usage)
for q in case_result.get("questions", []):
judgment = q.get("agentic_judgment", {})
score = judgment.get("llm_judge_score", 0.0)
# Binary: convert each rubric item score to 0/1, then average
judge_responses = judgment.get("llm_judge_responses", [])
if judge_responses:
binary_scores_per_item = [1.0 if r.get("score", 0) >= 1.0 else 0.0 for r in judge_responses]
binary_score = sum(binary_scores_per_item) / len(binary_scores_per_item)
else:
binary_score = 1.0 if score > 0.99 else 0.0
qtype = q["question_type"]
if qtype not in type_scores:
type_scores[qtype] = []
type_binary_scores[qtype] = []
type_scores[qtype].append(score)
type_binary_scores[qtype].append(binary_score)
all_scores.append(score)
all_binary_scores.append(binary_score)
metadata = q.get("agentic_metadata", {})
all_tool_call_totals.append(sum(metadata.get("tool_counts", {}).values()))
all_token_usages.append(metadata.get("token_usage", {}))
# Memory construction token usage summary
if all_memory_token_usages:
print("\n ── Memory Construction Token Usage ──")
for agent_name in ("default", "bench"):
for metric in _TOKEN_USAGE_METRICS:
values = [
usage[agent_name][metric]
for usage in all_memory_token_usages
if usage.get(agent_name, {}).get(metric) is not None
]
if values:
total = sum(values)
mean, std = _mean_and_std(values)
print(
f" {agent_name}/{metric}: total={total} mean={mean:.2f} std={std:.2f} ({len(values)} cases)",
)
else:
print(f" {agent_name}/{metric}: unavailable")
print()
print("\n ── AGENTIC ──")
if all_scores:
for qtype in sorted(type_scores.keys()):
scores = type_scores[qtype]
avg = sum(scores) / len(scores) if scores else 0
bin_scores = type_binary_scores[qtype]
bin_avg = sum(bin_scores) / len(bin_scores) if bin_scores else 0
print(f" {qtype:<40s}: {avg:.3f} binary={bin_avg:.3f} ({len(scores)} Qs)")
overall = sum(all_scores) / len(all_scores) if all_scores else 0
binary_overall = sum(all_binary_scores) / len(all_binary_scores) if all_binary_scores else 0
print(f" {'-'*38}")
print(f" {'OVERALL':<40s}: {overall:.3f} binary={binary_overall:.3f} ({len(all_scores)} Qs)")
tool_call_mean, tool_call_std = _mean_and_std(all_tool_call_totals)
print(f" Tool calls/query: mean={tool_call_mean:.2f} std={tool_call_std:.2f}")
print(" Bench reported tokens/query:")
for metric in _TOKEN_USAGE_METRICS:
values = [usage[metric] for usage in all_token_usages if usage.get(metric) is not None]
if values:
mean, std = _mean_and_std(values)
print(f" {metric}: mean={mean:.2f} std={std:.2f}")
else:
print(f" {metric}: unavailable")
else:
print(" (no results)")
# Per-case summary
print("\n ── Per-Case Summary ──")
for case_result in results:
case_id = case_result["case_id"]
if "error" in case_result:
print(f" Case {case_id}: ERROR — {case_result['error']}")
continue
n_qs = case_result.get("total_questions", 0)
n_sessions = case_result.get("sessions_ingested", 0)
mem_usage = case_result.get("memory_token_usage", {})
parts = [f"Case {case_id}: {n_sessions} sessions, {n_qs} questions"]
# Append memory construction total tokens if available
for agent_name in ("default", "bench"):
agent_usage = mem_usage.get(agent_name, {})
total = agent_usage.get("total_tokens")
if total is not None:
parts.append(f"mem_{agent_name}_tokens={total}")
questions = case_result.get("questions", [])
scores = [q.get("agentic_judgment", {}).get("llm_judge_score", 0.0) for q in questions]
if scores:
avg = sum(scores) / len(scores)
# Binary: 0/1 per rubric item, average per question, then across questions
bin_scores = []
for q in questions:
judge_responses = q.get("agentic_judgment", {}).get("llm_judge_responses", [])
if judge_responses:
item_bins = [1.0 if r.get("score", 0) >= 1.0 else 0.0 for r in judge_responses]
bin_scores.append(sum(item_bins) / len(item_bins))
else:
s = q.get("agentic_judgment", {}).get("llm_judge_score", 0.0)
bin_scores.append(1.0 if s > 0.99 else 0.0)
bin_avg = sum(bin_scores) / len(bin_scores)
parts.append(f"agentic={avg:.3f} binary={bin_avg:.3f}")
print(f" {' | '.join(parts)}")
print("=" * 70)
total_elapsed = time.time() - start_time
print(f"\n Total time: {total_elapsed/60:.1f} min")
print("\n" + "=" * 70)
print(" [DONE] BEAM EVALUATION COMPLETED SUCCESSFULLY")
print("=" * 70 + "\n")
_TOKEN_USAGE_METRICS = (
"input_tokens",
"output_tokens",
"total_tokens",
)
def _mean_and_std(values: list[int]) -> tuple[float, float]:
"""Return population mean and standard deviation for one per-question metric."""
if not values:
return 0.0, 0.0
mean = sum(values) / len(values)
return mean, (sum((value - mean) ** 2 for value in values) / len(values)) ** 0.5
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser(description="BEAM evaluation runner")
parser.add_argument("--config", type=str, default=None, help="Path to config.yaml")
parser.add_argument(
"--log-level",
type=str,
default="INFO",
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
help="Log level for the eval runner (default: INFO)",
)
parser.add_argument(
"--reme-log-level",
type=str,
default="INFO",
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
help="Log level for reme internal logs — loguru (default: INFO)",
)
parser.add_argument(
"-q",
"--quiet",
action="store_true",
help="Shortcut for --log-level WARNING --reme-log-level WARNING",
)
parser.add_argument(
"--eval_only",
action="store_true",
help="Skip ingestion. Reuse existing workspaces and only run query+judge.",
)
args = parser.parse_args()
if args.quiet:
args.log_level = "WARNING"
args.reme_log_level = "WARNING"
main(args.config, args.log_level, args.reme_log_level, eval_only=args.eval_only)

View file

@ -1,96 +0,0 @@
[中文版 / Chinese version](./README_ZH.md)
# LongMemEval Benchmark
LongMemEval is a benchmark for **long-term memory over multi-session chat
histories**. Each item provides a chronologically ordered set of chat sessions
between a user and an assistant, followed by a probing question whose answer is
only recoverable by reasoning over the user-owned memory. ReMe ingests the
sessions into an isolated per-item workspace, answers the question via an
agentic (ReAct) mode, and scores the answer with an LLM-as-judge.
Question types include single-session (user / assistant / preference),
multi-session reasoning, knowledge update, and temporal reasoning.
> For the shared setup (dependencies, credentials, log conventions) see the
> [top-level benchmark README](../README.md).
## 1. Get the Dataset
ReMe uses only the **cleaned-S** split, hosted on HuggingFace:
[agentscope-ai/ReMe_longmemeval_clean_s_v2](https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2).
The download script fetches it via the hf-mirror.com mirror; to use a different
mirror, modify `BASE_URL` in [`download.py`](./download.py).
```bash
cd benchmark/longmemeval
python download.py # saves dataset/longmemeval_s_reme_cleaned.json; skips if already present
```
Ground truth is embedded in the data file.
## 2. Run
From the repository root:
```bash
python benchmark/longmemeval/run.py
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
python benchmark/longmemeval/run.py --eval_only # reuse existing workspaces, query + judge only
```
## 3. Pipeline
1. Load the dataset (ground truth is embedded in the data file).
2. For each item, create an isolated workspace and ingest sessions in chronological order.
3. Trigger `auto_dream` when consecutive sessions cross the configured hour (default 23:00).
4. Answer each question via agentic (ReAct) mode.
5. Judge the answer (binary yes/no) with the `answer_judge` job and print per-type accuracy.
## 4. Key config — `benchmark/longmemeval/config.yaml`
| Key | Meaning |
| --- | --- |
| `dataset.path` | Dataset file to evaluate (e.g. `longmemeval_s_reme_cleaned.json`); ground truth is included. |
| `dataset.start_index` / `num_items` | Slice of items to evaluate. |
| `dataset.question_types` | Filter by question type; empty = all. |
| `dataset.workspace_root` | Per-item workspace root (`benchmark/longmemeval/workspaces/longmemeval-s`). |
| `evaluation.num_workers` | `0` = auto (cpu-2), `1` = sequential, `>1` = parallel. |
| `evaluation.filter_future_sessions` | Only ingest sessions with timestamp ≤ `question_date`. |
| `reme.config` | ReMe config used (`lme.yaml`). |
| `reme.dream_trigger_hour` / `dream_scan_days` / `dream_max_units` | Dream triggering behavior. |
| `output.dir` | Results directory (`benchmark/longmemeval/results`). |
## 5. Outputs
Results are JSON files written to `output.dir` as `results_<timestamp>.json`,
with a per-type accuracy summary also printed to the console. Logging
conventions are shared across benchmarks — see the
[top-level README](../README.md#outputs--logs).
## 6. Reference Results
### cleaned-s
**Basic settings**
1. Modified auto-memory prompt, auto-dream disabled.
2. All sessions in reme-memory are strictly earlier than the question time.
**Results**
agentscope==2.0.4.post1, conda reme env, 32 workers, eval-only (reusing prebuilt memory)
(2026-08-06, 500 items, total 10.0 min)
| Type | Agentic | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|
| knowledge-update | 0.910 | 31,581 | 589 | 32,169 | 2.90 |
| multi-session | 0.842 | 52,837 | 1,474 | 54,311 | 4.21 |
| single-session-assistant | 1.000 | 15,596 | 279 | 15,875 | 1.89 |
| single-session-preference | 0.633 | 36,802 | 818 | 37,620 | 3.60 |
| single-session-user | 0.986 | 27,433 | 359 | 27,792 | 2.60 |
| temporal-reasoning | 0.902 | 62,674 | 985 | 63,659 | 4.97 |
| **OVERALL** | **0.894** | **43,448** | **876** | **44,324** | **3.69** |

View file

@ -1,90 +0,0 @@
# LongMemEval 评测
[English version](./README.md)
LongMemEval 是一个面向**多轮多会话历史的长期记忆能力**的评测基准。每个条目提供一组按时间
顺序排列的用户与助手之间的会话以及一个只能通过推理用户自有记忆才能回答的探测问题。ReMe
将会话摄入按条目隔离的工作区,以 agenticReAct模式回答问题最后由 LLM-as-judge 打分。
题型包括单会话user / assistant / preference、多会话推理、知识更新与时间推理等。
> 公共设置(依赖、凭据、日志约定)见[总评测说明](../README_ZH.md)。
## 1. 获取数据集
ReMe 仅使用 **cleaned-S** 版本,数据托管在 HuggingFace
[agentscope-ai/ReMe_longmemeval_clean_s_v2](https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2)。
下载脚本经 hf-mirror.com 镜像源获取,如需更换源请修改 [`download.py`](./download.py) 中的
`BASE_URL`
```bash
cd benchmark/longmemeval
python download.py # 保存为 dataset/longmemeval_s_reme_cleaned.json已存在则自动跳过
```
ground truth 已内嵌在数据文件中。
## 2. 运行
在仓库根目录执行:
```bash
python benchmark/longmemeval/run.py
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
python benchmark/longmemeval/run.py -q # 安静模式:仅评测级日志
python benchmark/longmemeval/run.py --log-level WARNING # 降低评测 runner 日志
python benchmark/longmemeval/run.py --reme-log-level WARNING # 降低 reme 内部日志
python benchmark/longmemeval/run.py --eval_only # 复用已有工作区,仅执行查询 + 评判
```
## 3. 流程
1. 加载数据集ground truth 已内嵌在数据文件中)。
2. 为每个条目创建独立工作区,按时间顺序摄入会话。
3. 当相邻会话跨越配置的时刻(默认 23:00时触发 `auto_dream`
4. 以 agenticReAct模式回答每个问题。
5. 通过 `answer_judge` 任务对答案做二元yes/no评判并输出各类型准确率。
## 4. 关键配置 —— `benchmark/longmemeval/config.yaml`
| 配置项 | 含义 |
| --- | --- |
| `dataset.path` | 待评测的数据集文件(如 `longmemeval_s_reme_cleaned.json`),已包含 ground truth。 |
| `dataset.start_index` / `num_items` | 评测条目的切片范围。 |
| `dataset.question_types` | 按问题类型过滤,空表示全部。 |
| `dataset.workspace_root` | 条目工作区根目录(`benchmark/longmemeval/workspaces/longmemeval-s`)。 |
| `evaluation.num_workers` | `0` = 自动cpu-2`1` = 串行,`>1` = 并行。 |
| `evaluation.filter_future_sessions` | 仅摄入时间戳 ≤ `question_date` 的会话。 |
| `reme.config` | 使用的 ReMe 配置(`lme.yaml`)。 |
| `reme.dream_trigger_hour` / `dream_scan_days` / `dream_max_units` | dream 触发行为。 |
| `output.dir` | 结果目录(`benchmark/longmemeval/results`)。 |
## 5. 输出
结果以 JSON 文件写入 `output.dir`,文件名为 `results_<timestamp>.json`
同时控制台会打印含各类型准确率的汇总。日志约定在各基准间通用,见
[总说明](../README_ZH.md#输出与日志)。
## 6. 参考结果
### cleaned-s
**基础设置**
1. 使用修改后的 auto-memory prompt关闭 auto-dream 机制
2. reme-memory 中的全部 session 的时间一定早于 question 的时间
**结果**
agentscope==2.0.4.post1, conda reme env, 32 workers, eval-only复用预构建记忆
2026-08-06500 题,总计 10.0 min
| 类型 | Agentic | input tok/q | output tok/q | total tok/q | tool calls/q |
|---|---|---|---|---|---|
| knowledge-update | 0.910 | 31,581 | 589 | 32,169 | 2.90 |
| multi-session | 0.842 | 52,837 | 1,474 | 54,311 | 4.21 |
| single-session-assistant | 1.000 | 15,596 | 279 | 15,875 | 1.89 |
| single-session-preference | 0.633 | 36,802 | 818 | 37,620 | 3.60 |
| single-session-user | 0.986 | 27,433 | 359 | 27,792 | 2.60 |
| temporal-reasoning | 0.902 | 62,674 | 985 | 63,659 | 4.97 |
| **OVERALL** | **0.894** | **43,448** | **876** | **44,324** | **3.69** |

View file

@ -1,33 +0,0 @@
# LongMemEval evaluation configuration
# This file controls what/how to evaluate.
dataset:
path: "benchmark/longmemeval/dataset/longmemeval_s_reme_cleaned.json"
start_index: 0 # first item index
num_items: 500 # how many items to evaluate (starting from start_index)
max_sessions: 0 # 0 = all sessions; >0 = limit sessions per item for testing
question_types: [] # filter by question_type; empty list = no filtering (all types)
workspace_root: "benchmark/longmemeval/workspaces/longmemeval-s" # workspace root for item workspaces
evaluation:
# LLM-as-judge uses the 'judge' as_llm component defined in lme.yaml
# Model and credentials are configured there (reading from .env)
# Judgment is always binary (yes/no) — defined in lme/llm_judge.yaml
num_workers: 32 # 0 = auto (cpu_count - 2, min 1); 1 = sequential; >1 = parallel
filter_future_sessions: true # true = only ingest sessions with timestamp <= question_date
compress_session: false # true = compress session chunks in search_v2 (query-aware); false = no compression
reme:
config: "lme.yaml" # reme config to use (in reme/config/)
# Dream trigger: when gap between consecutive sessions crosses this hour (23:00)
dream_trigger_hour: 23
# Dream scan_days for each trigger
dream_scan_days: 2
dream_max_units: 5
output:
dir: "benchmark/longmemeval/results"
log_dir: "logs" # log directory (relative to project root)
log_prefix: "longmemeval" # benchmark name used in log filenames
log_to_console: true
log_to_file: true

View file

@ -1,67 +0,0 @@
"""Download the LongMemEval cleaned-S dataset used by ReMe.
Source: https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2
(downloaded via the hf-mirror.com mirror for reliability).
The file ``longmemeval_s_reme_cleaned.json`` is saved under ``dataset/`` next to this
script using the same name as on the remote (``benchmark/longmemeval/config.yaml``
points to it).
Usage:
python download.py # download cleaned-S (skip if it already exists)
"""
import os
import sys
import urllib.request
BASE_URL = "https://hf-mirror.com/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2/resolve/main"
TARGET_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "dataset")
# Files to download (saved with the same name as on the remote).
FILES = [
"longmemeval_s_reme_cleaned.json",
]
def download_file(filename: str):
"""Download a single file from the mirror to the target directory."""
url = f"{BASE_URL}/{filename}"
dest = os.path.join(TARGET_DIR, filename)
if os.path.exists(dest):
size = os.path.getsize(dest)
print(f" [skip] {filename} already exists ({size / 1024 / 1024:.1f} MB)")
return
print(f" [downloading] {filename} ...")
try:
urllib.request.urlretrieve(url, dest, reporthook=_progress)
size = os.path.getsize(dest)
print(f"\n [done] {filename} ({size / 1024 / 1024:.1f} MB)")
except Exception as e:
print(f"\n [error] {filename}: {e}")
if os.path.exists(dest):
os.remove(dest)
sys.exit(1)
def _progress(block_num, block_size, total_size):
downloaded = block_num * block_size
if total_size > 0:
pct = min(100, downloaded * 100 / total_size)
mb = downloaded / 1024 / 1024
total_mb = total_size / 1024 / 1024
sys.stdout.write(f"\r {mb:.1f}/{total_mb:.1f} MB ({pct:.1f}%)")
else:
mb = downloaded / 1024 / 1024
sys.stdout.write(f"\r {mb:.1f} MB downloaded")
sys.stdout.flush()
if __name__ == "__main__":
os.makedirs(TARGET_DIR, exist_ok=True)
print(f"Downloading LongMemEval cleaned-S dataset to: {TARGET_DIR}\n")
for fname in FILES:
download_file(fname)
print("\nAll files downloaded successfully!")

View file

@ -1,76 +0,0 @@
#!/bin/bash
# 杀死指定进程及其所有子进程
# Usage: bash kill.sh <PID>
if [ -z "$1" ]; then
echo "Usage: bash kill.sh <PID>"
echo " 杀死指定进程及其所有子进程"
exit 1
fi
PID=$1
# 检查进程是否存在
if ! kill -0 "$PID" 2>/dev/null; then
echo "进程 $PID 不存在"
exit 1
fi
# 递归收集所有子进程(包括子进程的子进程)
collect_children() {
local parent=$1
local children
children=$(ps -o pid= --ppid "$parent" 2>/dev/null | tr -d ' ')
for child in $children; do
collect_children "$child"
done
echo "$parent"
}
# 收集进程树(子进程在前,父进程在后,保证先杀子再杀父)
PROCESS_TREE=$(collect_children "$PID")
TOTAL=$(echo "$PROCESS_TREE" | wc -l | tr -d ' ')
echo "进程树(共 $TOTAL 个进程):"
while read -r p; do
cmd=$(ps -o args= -p "$p" 2>/dev/null | head -c 80)
printf " PID=%-8s %s\n" "$p" "$cmd"
done <<< "$PROCESS_TREE"
# 先 SIGTERM 优雅终止
echo ""
echo "发送 SIGTERM..."
while read -r p; do
kill "$p" 2>/dev/null
done <<< "$PROCESS_TREE"
# 等待最多 5 秒
for i in $(seq 1 5); do
alive=false
while read -r p; do
if kill -0 "$p" 2>/dev/null; then
alive=true
fi
done <<< "$PROCESS_TREE"
if [ "$alive" = false ]; then
break
fi
sleep 1
done
# 检查是否还有残留,强制 SIGKILL
remaining=false
while read -r p; do
if kill -0 "$p" 2>/dev/null; then
remaining=true
fi
done <<< "$PROCESS_TREE"
if [ "$remaining" = true ]; then
echo "部分进程未响应,发送 SIGKILL..."
while read -r p; do
kill -9 "$p" 2>/dev/null
done <<< "$PROCESS_TREE"
fi
echo "已终止进程树(根 PID=$PID,共 $TOTAL 个进程)"

View file

@ -1,816 +0,0 @@
"""LongMemEval evaluation runner for ReMe.
Evaluates ReMe's long-term memory capability using the LongMemEval dataset.
Each item gets an isolated workspace; sessions are ingested in chronological order;
dream is triggered when sessions cross midnight (23:00); finally questions are
answered via an agentic (ReAct) approach and judged by an LLM.
Usage:
python benchmark/longmemeval/run.py
python benchmark/longmemeval/run.py --config benchmark/longmemeval/config.yaml
python benchmark/longmemeval/run.py -q # quiet: only eval-level logs
python benchmark/longmemeval/run.py --log-level WARNING # reduce eval runner logs
python benchmark/longmemeval/run.py --reme-log-level WARNING # reduce reme internal logs
python benchmark/longmemeval/run.py --eval_only # query+judge only, reuse existing workspace
"""
import json
import logging
import os
import re
import shutil
import time
import threading
from datetime import datetime
from pathlib import Path
import yaml
from dotenv import load_dotenv
# Load .env from project root
_PROJECT_ROOT = Path(__file__).parent.parent.parent
load_dotenv(_PROJECT_ROOT / ".env")
# Workspace root for evaluation items — read from config.yaml (dataset.workspace_root)
_WORKSPACE_ROOT_DEFAULT = "benchmark/longmemeval/workspaces/longmemeval-s"
# ---------------------------------------------------------------------------
# Logging
# ---------------------------------------------------------------------------
_DEFAULT_LOG_FORMAT = "%(asctime)s | %(levelname)s | %(message)s"
logging.basicConfig(level=logging.INFO, format=_DEFAULT_LOG_FORMAT)
logger = logging.getLogger("longmemeval")
# Noisy library loggers silenced by default
_NOISY_LOGGERS = [
"httpx",
"httpcore",
"openai",
"uvicorn",
"multipart",
"asyncio",
"watchfiles",
"filelock",
]
def setup_logging(
log_level: str,
reme_log_level: str,
log_dir: str | None = None,
):
"""Configure logging for the eval runner and reme internals.
Args:
log_level: Level for the eval runner logger (DEBUG/INFO/WARNING/ERROR).
reme_log_level: Level for reme's internal loguru logger.
log_dir: Per-run log directory (absolute path). None = no file logging.
"""
numeric = getattr(logging, log_level.upper(), logging.INFO)
# Eval runner logger
logging.getLogger().setLevel(numeric)
logger.setLevel(numeric)
# Suppress noisy library loggers when above DEBUG
if numeric > logging.DEBUG:
for name in _NOISY_LOGGERS:
lib_logger = logging.getLogger(name)
lib_logger.setLevel(max(numeric, logging.WARNING))
# Add file handler for eval runner if log_dir is specified
if log_dir:
os.makedirs(log_dir, exist_ok=True)
log_filepath = os.path.join(log_dir, "runner.log")
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
file_handler.setLevel(numeric)
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
logging.getLogger().addHandler(file_handler)
logger.info(f"Eval runner log file: {log_filepath}")
# Reme internal logger (loguru) — will be applied per-worker via _configure_worker
os.environ["REME_LOG_LEVEL"] = reme_log_level.upper()
if log_dir:
os.environ["REME_LOG_DIR"] = log_dir
def _configure_worker(
log_level: str,
reme_log_level: str,
log_dir: str | None = None,
):
"""Set up logging inside a multiprocessing worker process.
Must be called at the top of each worker because child processes inherit
parent state but loguru sinks are NOT shared across fork/spawn.
"""
numeric = getattr(logging, log_level.upper(), logging.INFO)
logging.basicConfig(level=numeric, format=_DEFAULT_LOG_FORMAT, force=True)
logging.getLogger("longmemeval").setLevel(numeric)
if numeric > logging.DEBUG:
for name in _NOISY_LOGGERS:
logging.getLogger(name).setLevel(max(numeric, logging.WARNING))
# Add file handler for eval runner in worker process
if log_dir:
os.makedirs(log_dir, exist_ok=True)
pid = os.getpid()
log_filepath = os.path.join(log_dir, f"worker-{pid}.log")
file_handler = logging.FileHandler(log_filepath, encoding="utf-8")
file_handler.setLevel(numeric)
file_handler.setFormatter(logging.Formatter(_DEFAULT_LOG_FORMAT))
logging.getLogger().addHandler(file_handler)
# Re-initialize loguru for reme internals at the desired level
from reme.utils import get_logger
reme_log_dir = log_dir or "logs"
get_logger(log_dir=reme_log_dir, level=reme_log_level.upper(), force_init=True)
# ---------------------------------------------------------------------------
# Config loading
# ---------------------------------------------------------------------------
def load_eval_config(config_path: str | None = None) -> dict:
"""Load evaluation config yaml with env-var expansion."""
if config_path is None:
config_path = str(Path(__file__).parent / "config.yaml")
with open(config_path, encoding="utf-8") as f:
raw = f.read()
# Expand ${VAR} and ${VAR:-default}
def _expand(m):
expr = m.group(1)
if ":-" in expr:
key, default = expr.split(":-", 1)
return os.environ.get(key, default)
return os.environ.get(expr, "")
raw = re.sub(r"\$\{([^}]+)\}", _expand, raw)
return yaml.safe_load(raw)
# ---------------------------------------------------------------------------
# Date utilities
# ---------------------------------------------------------------------------
def parse_haystack_date(date_str: str) -> datetime:
"""Parse LongMemEval date format: '2023/05/20 (Sat) 02:21' -> datetime."""
m = re.match(r"(\d{4}/\d{2}/\d{2})\s+\(\w+\)\s+(\d{2}:\d{2})", date_str)
if not m:
raise ValueError(f"Cannot parse haystack date: {date_str!r}")
return datetime.strptime(f"{m.group(1)} {m.group(2)}", "%Y/%m/%d %H:%M")
def to_iso(dt: datetime) -> str:
"""Convert datetime to ISO-8601 string precise to seconds."""
return dt.strftime("%Y-%m-%dT%H:%M:%S")
def should_trigger_dream(prev_dt: datetime, curr_dt: datetime, _trigger_hour: int = 23) -> bool:
"""Check if the time gap between two sessions crosses trigger_hour (e.g. 23:00)."""
if prev_dt.date() == curr_dt.date():
return False
# There's at least one midnight crossing; check if trigger_hour is between them
# Simple heuristic: if dates differ, dream should run for the previous day
return True
def sessions_sorted_by_time(item: dict) -> list[tuple[int, datetime, str, list[dict]]]:
"""Return (original_index, parsed_datetime, session_id, messages) sorted by time."""
entries = []
for i, (date_str, sid, msgs) in enumerate(
zip(item["haystack_dates"], item["haystack_session_ids"], item["haystack_sessions"]),
):
dt = parse_haystack_date(date_str)
entries.append((i, dt, sid, msgs))
# Sort by time (ascending)
entries.sort(key=lambda x: x[1])
return entries
# ---------------------------------------------------------------------------
# Message formatting
# ---------------------------------------------------------------------------
def format_messages_for_reme(messages: list[dict], session_dt: datetime) -> list[dict]:
"""Convert LongMemEval messages to ReMe auto_memory format.
Adds: name, created_at (ISO seconds). All messages in a session share the
same created_at (the session timestamp).
"""
formatted = []
for msg in messages:
role = msg["role"]
formatted.append(
{
"name": role,
"role": role,
"content": msg["content"],
"created_at": to_iso(session_dt),
},
)
return formatted
# ---------------------------------------------------------------------------
# LLM-as-Judge (delegated to answer_judge_step via app.run_job)
# ---------------------------------------------------------------------------
async def judge_response_via_job(
app,
question: str,
ground_truth: str,
response: str,
question_type: str,
) -> dict:
"""Use the answer_judge_step to evaluate a response against the golden answer."""
judge_resp = await app.run_job(
"answer_judge",
query=question,
agent_answer=response,
golden_answer=ground_truth,
question_type=question_type,
)
verdict = (judge_resp.answer or "").strip().lower()
raw_answer = (judge_resp.metadata or {}).get("raw_answer_judgement", "")
return {
"verdict": verdict,
"reason": raw_answer if verdict not in ("yes", "no") else "",
"metric": "binary",
"question_type": question_type,
}
# ---------------------------------------------------------------------------
# Main evaluation pipeline
# ---------------------------------------------------------------------------
async def evaluate_item(item: dict, eval_config: dict, item_index: int, eval_only: bool = False) -> dict:
"""Evaluate a single LongMemEval item end-to-end.
Args:
item: The dataset item containing question, answer, sessions, etc.
eval_config: The evaluation configuration dict.
item_index: The index of this item in the dataset.
eval_only: If True, skip ingestion (phases 1-3) and only run query+judge
using the existing workspace. Useful for re-evaluating different query
configurations without re-ingesting sessions.
"""
from reme import Application
from reme.config import resolve_app_config
from reme.utils.evaluation_interface import track_agent_token_usage, track_job_counts
reme_cfg = eval_config["reme"]
dream_trigger_hour = reme_cfg.get("dream_trigger_hour", 23)
dream_scan_days = reme_cfg.get("dream_scan_days", 2)
dream_max_units = reme_cfg.get("dream_max_units", 5)
# Sort sessions by time
sorted_sessions = sessions_sorted_by_time(item)
# Filter out sessions that occur after question_date (if enabled)
filter_future = eval_config["evaluation"].get("filter_future_sessions", True)
if filter_future and item.get("question_date"):
question_dt = parse_haystack_date(item["question_date"])
total_before_filter = len(sorted_sessions)
sorted_sessions = [(i, dt, sid, msgs) for i, dt, sid, msgs in sorted_sessions if dt <= question_dt]
if len(sorted_sessions) < total_before_filter:
logger.info(
f"[Item {item_index}] Filtered sessions: {total_before_filter} -> {len(sorted_sessions)} "
f"(removed {total_before_filter - len(sorted_sessions)} future sessions "
f"after question_date={item['question_date']})",
)
logger.info(
"[Item %s] question_id=%s type=%s sessions=%d%s",
item_index,
item["question_id"],
item["question_type"],
len(sorted_sessions),
" [eval_only]" if eval_only else "",
)
# Use fixed workspace directory (clean it for fresh evaluation)
workspace_root = _PROJECT_ROOT / eval_config["dataset"].get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
item_dir = workspace_root / f"item_{item_index}"
workspace_dir = str(item_dir / ".reme")
if eval_only:
if not item_dir.exists() or not Path(workspace_dir).exists():
raise FileNotFoundError(
f"[Item {item_index}] eval_only: workspace not found at {item_dir}. "
f"Run without --eval_only first to build the workspace.",
)
else:
if item_dir.exists():
shutil.rmtree(item_dir)
logger.info(f"[Item {item_index}] Cleaned existing workspace: {item_dir}")
else:
logger.info(f"[Item {item_index}] Workspace not found, creating: {item_dir}")
item_dir.mkdir(parents=True, exist_ok=True)
# Pre-initialize ReMe's loguru logger with the correct log_dir
# (singleton — Application.__init__ will reuse this instance)
output_cfg = eval_config.get("output", {})
if output_cfg.get("log_to_file", False):
reme_log_dir = os.environ.get("REME_LOG_DIR")
if reme_log_dir:
from reme.utils import get_logger
get_logger(
log_dir=reme_log_dir,
level=os.environ.get("REME_LOG_LEVEL", "INFO"),
log_to_console=output_cfg.get("log_to_console", True),
log_to_file=True,
force_init=True,
)
cfg = resolve_app_config(
config=reme_cfg["config"],
workspace_dir=workspace_dir,
log_to_console=output_cfg.get("log_to_console", True),
log_to_file=output_cfg.get("log_to_file", False),
enable_logo=False,
)
app = Application(**cfg)
await app.start()
try:
dream_dates_triggered = set()
dream_available = True # Set to False if auto_dream job is not found
if not eval_only:
# ── Phase 1: Ingest sessions ──────────────────────────────
prev_dt = None
for idx, (_, session_dt, session_id, messages) in enumerate(sorted_sessions):
# Check if dream should be triggered before this session
if (
dream_available
and prev_dt is not None
and should_trigger_dream(prev_dt, session_dt, dream_trigger_hour)
):
dream_date = prev_dt.strftime("%Y-%m-%d")
if dream_date not in dream_dates_triggered:
logger.info(f"[Item {item_index}] Triggering dream for date={dream_date}")
try:
dream_resp = await app.run_job(
"auto_dream",
date=dream_date,
scan_days=dream_scan_days,
max_units=dream_max_units,
)
logger.info(
f"[Item {item_index}] Dream done: success={dream_resp.success} "
f"answer={dream_resp.answer[:100] if dream_resp.answer else ''}",
)
except Exception as e:
if "not found" in str(e).lower():
dream_available = False
logger.warning(f"[Item {item_index}] auto_dream job not found, skipping all dreams")
else:
logger.warning(f"[Item {item_index}] Dream failed for {dream_date}: {e}")
dream_dates_triggered.add(dream_date)
# Index update after dream to pick up new digest nodes
await app.run_job("index_update")
# Format and ingest the session
formatted_msgs = format_messages_for_reme(messages, session_dt)
date_str = session_dt.strftime("%Y-%m-%d")
logger.info(
f"[Item {item_index}] Ingesting session {idx+1}/{len(sorted_sessions)} "
f"id={session_id} date={date_str} msgs={len(formatted_msgs)}",
)
resp = await app.run_job(
"auto_memory",
messages=formatted_msgs,
session_id=session_id,
date=date_str,
)
if not resp.success:
logger.warning(
f"[Item {item_index}] auto_memory failed for session {session_id}: {resp.answer}",
)
# Manual index update after each session
await app.run_job("index_update")
prev_dt = session_dt
# ── Phase 2: Final dream for the last day ─────────────────
if dream_available and prev_dt is not None:
last_dream_date = prev_dt.strftime("%Y-%m-%d")
if last_dream_date not in dream_dates_triggered:
logger.info(f"[Item {item_index}] Final dream for date={last_dream_date}")
try:
await app.run_job(
"auto_dream",
date=last_dream_date,
scan_days=dream_scan_days,
max_units=dream_max_units,
)
except Exception as e:
if "not found" in str(e).lower():
dream_available = False
logger.warning(f"[Item {item_index}] auto_dream job not found, skipping all dreams")
else:
logger.warning(f"[Item {item_index}] Final dream failed: {e}")
dream_dates_triggered.add(last_dream_date)
# Index update after final dream
await app.run_job("index_update")
# ── Phase 3: Digest update ────────────────────────────────
await app.run_job("digest_update")
# ── Phase 4: Ask question via agentic_answer job (ReAct agent) ──
question = item["question"]
compress_session = bool(eval_config["evaluation"].get("compress_session", False))
question_date_raw = item.get("question_date", "")
question_dt = parse_haystack_date(question_date_raw) if question_date_raw else None
query_time = to_iso(question_dt) if question_dt else ""
logger.info(
f"[Item {item_index}] Asking (agentic): {question[:80]}... query_time={query_time}",
)
with (
track_job_counts(["search"], app.context) as tool_counts,
track_agent_token_usage(
["bench"],
app.context,
) as token_usages,
):
query_resp = await app.run_job(
"agentic_answer",
query=question,
query_time=query_time,
compress_session=compress_session,
)
agentic_tool_counts = tool_counts
agentic_token_usage = token_usages["bench"]
agentic_response = (query_resp.answer or "").strip()
if not agentic_response:
agentic_response = "(no answer generated)"
logger.info(f"[Item {item_index}] Agentic response: {agentic_response[:200]}...")
logger.info(f"[Item {item_index}] Agentic tool calls: {agentic_tool_counts}")
logger.info(f"[Item {item_index}] Bench token usage: {agentic_token_usage}")
# ── Phase 5: Judge agentic response (via answer_judge_step) ──────────
logger.info(f"[Item {item_index}] Judging agentic (binary, type={item['question_type']})...")
agentic_judgment = await judge_response_via_job(
app=app,
question=question,
ground_truth=item["answer"],
response=agentic_response,
question_type=item["question_type"],
)
logger.info(f"[Item {item_index}] agentic binary result: {agentic_judgment}")
finally:
await app.close()
return {
"question_id": item["question_id"],
"question_type": item["question_type"],
"question": question,
"ground_truth": item["answer"],
"agentic_response": agentic_response,
"agentic_judgment": agentic_judgment,
"agentic_tool_counts": agentic_tool_counts,
"agentic_token_usage": agentic_token_usage,
"sessions_ingested": len(sorted_sessions),
"dreams_triggered": len(dream_dates_triggered),
}
# ---------------------------------------------------------------------------
# Worker: runs a single item in its own process with its own event loop
# ---------------------------------------------------------------------------
def _evaluate_item_worker(task_input: tuple) -> dict:
"""Worker function for multiprocessing. Each process gets its own event loop."""
item, eval_config, item_index, log_level, reme_log_level, eval_only, log_dir = task_input
import asyncio # pylint: disable=import-outside-toplevel
_configure_worker(log_level, reme_log_level, log_dir=log_dir)
# Permanently suppress "Task exception was never retrieved" /
# "Event loop is closed" noise from httpx AsyncClient GC cleanup.
# These fire AFTER asyncio.run() closes the loop, during Python's
# garbage collection of httpx connection-pool tasks — harmless.
logging.getLogger("asyncio").setLevel(logging.CRITICAL)
return asyncio.run(evaluate_item(item, eval_config, item_index, eval_only=eval_only))
def _indexed_worker(indexed_input: tuple) -> tuple:
"""Module-level wrapper for imap_unordered with index tracking."""
idx, task_input = indexed_input
return idx, _evaluate_item_worker(task_input)
def _resolve_num_workers(configured: int) -> int:
"""Resolve num_workers: 0=auto (cpu_count-2, min 1), 1=sequential, >1=parallel."""
if configured == 0:
return max(1, (os.cpu_count() or 4) - 2)
return max(1, configured)
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main(
config_path: str | None = None,
log_level: str = "INFO",
reme_log_level: str = "INFO",
eval_only: bool = False,
):
"""Run the LongMemEval evaluation pipeline.
Args:
config_path: Path to the YAML config file.
log_level: Log level for the eval runner.
reme_log_level: Log level for reme internal logs.
eval_only: If True, skip ingestion and only run query+judge using
existing workspaces.
"""
from multiprocessing import Pool # pylint: disable=import-outside-toplevel
# Load config BEFORE logging setup so log_dir is available
eval_config = load_eval_config(config_path)
# Resolve per-run log directory from config
output_cfg = eval_config.get("output", {})
log_dir_abs = None
if output_cfg.get("log_to_file", False):
log_dir_raw = output_cfg.get("log_dir", "logs")
log_prefix = output_cfg.get("log_prefix", "longmemeval")
run_ts = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
log_dir_abs = str(_PROJECT_ROOT / log_dir_raw / f"{log_prefix}_{run_ts}")
setup_logging(log_level, reme_log_level, log_dir=log_dir_abs)
dataset_cfg = eval_config["dataset"]
# Load dataset
dataset_path = _PROJECT_ROOT / dataset_cfg["path"]
logger.info(f"Loading dataset from {dataset_path}")
with open(dataset_path, encoding="utf-8") as f:
data = json.load(f)
start = dataset_cfg.get("start_index", 0)
num_items = dataset_cfg.get("num_items", 0)
if num_items > 0:
raw_items = data[start : start + num_items]
else:
raw_items = data[start:]
# Build item list
items_with_idx = [(start + i, item) for i, item in enumerate(raw_items)]
# Filter by question_type if specified
question_types = dataset_cfg.get("question_types") or []
if question_types:
before_filter = len(items_with_idx)
items_with_idx = [(idx, item) for idx, item in items_with_idx if item.get("question_type") in question_types]
logger.info(
f"Filtered by question_types={question_types}: {before_filter} -> {len(items_with_idx)} items",
)
# Filter by question_id if specified
question_ids = dataset_cfg.get("question_ids") or []
if question_ids:
qid_set = set(question_ids)
before_filter = len(items_with_idx)
items_with_idx = [(idx, item) for idx, item in items_with_idx if item.get("question_id") in qid_set]
logger.info(
f"Filtered by question_ids ({len(qid_set)} ids): {before_filter} -> {len(items_with_idx)} items",
)
logger.info(
"Evaluating %d item(s) starting from index %d%s",
len(items_with_idx),
start,
" [eval_only: query+judge only]" if eval_only else "",
)
# Resolve parallelism
num_workers = _resolve_num_workers(eval_config["evaluation"].get("num_workers", 1))
logger.info(f"Using {num_workers} worker(s)")
# Create output directory
output_dir = _PROJECT_ROOT / output_cfg.get("dir", "benchmark/longmemeval/results")
output_dir.mkdir(parents=True, exist_ok=True)
# Create workspace root directory
workspace_root = _PROJECT_ROOT / dataset_cfg.get("workspace_root", _WORKSPACE_ROOT_DEFAULT)
workspace_root.mkdir(parents=True, exist_ok=True)
# Pre-check: verify all workspaces exist in eval_only mode
if eval_only:
missing_items = []
for orig_idx, _ in items_with_idx:
item_dir = workspace_root / f"item_{orig_idx}"
if not item_dir.exists() or not (item_dir / ".reme").exists():
missing_items.append(orig_idx)
if missing_items:
preview = missing_items[:10]
suffix = "..." if len(missing_items) > 10 else ""
raise FileNotFoundError(
f"eval_only: {len(missing_items)} workspace(s) not found under {workspace_root}. "
f"Missing item indices: {preview}{suffix}. "
f"Run without --eval_only first to build the workspaces.",
)
# Build task args — include log levels, eval_only flag, and log paths (use original index for workspace lookup)
task_args = [
(item, eval_config, orig_idx, log_level, reme_log_level, eval_only, log_dir_abs)
for orig_idx, item in items_with_idx
]
# Progress tracking (force print regardless of log level, every 10 minutes)
total_items = len(task_args)
completed_count = [0] # use list for mutability in closure
start_time = time.time()
progress_lock = threading.Lock()
def _print_progress(prefix: str = "PROGRESS"):
elapsed = time.time() - start_time
elapsed_min = elapsed / 60
done = completed_count[0]
pct = 100.0 * done / total_items if total_items else 0
eta_str = "N/A"
if done > 0:
eta_sec = elapsed / done * (total_items - done)
eta_str = f"{eta_sec/60:.1f}min"
print(
f"[{prefix}] {datetime.now().strftime('%Y-%m-%d %H:%M:%S')} | "
f"{done}/{total_items} ({pct:.1f}%) completed | "
f"elapsed={elapsed_min:.1f}min | ETA={eta_str}",
flush=True,
)
def _progress_timer():
"""Background thread: print progress every 10 minutes."""
while not _timer_stop.is_set():
_timer_stop.wait(600) # 10 minutes
if not _timer_stop.is_set():
with progress_lock:
_print_progress()
_timer_stop = threading.Event()
timer_thread = threading.Thread(target=_progress_timer, daemon=True)
timer_thread.start()
# Run evaluation
if num_workers == 1:
# Sequential mode
results = []
for task_input in task_args:
result = _evaluate_item_worker(task_input)
results.append(result)
with progress_lock:
completed_count[0] += 1
else:
# Parallel mode — use imap_unordered for progress tracking
results = [None] * total_items
indexed_args = list(enumerate(task_args))
with Pool(processes=num_workers) as pool:
for idx, result in pool.imap_unordered(_indexed_worker, indexed_args):
results[idx] = result
with progress_lock:
completed_count[0] += 1
# Stop progress timer
_timer_stop.set()
timer_thread.join(timeout=2)
# Save results
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
output_file = output_dir / f"results_{timestamp}.json"
with open(output_file, "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
logger.info(f"Results saved to {output_file}")
# Final progress
_print_progress("FINAL")
_print_summary(results, start_time)
# ---------------------------------------------------------------------------
# Summary printing
# ---------------------------------------------------------------------------
def _print_summary(results: list[dict], start_time: float) -> None:
"""Print per-item verdicts and per-type accuracy."""
print("\n" + "=" * 60)
print("EVALUATION RESULTS")
print("=" * 60)
def _accumulate(judgment_key):
correct = 0
stats: dict = {} # {question_type: {correct: int, total: int}}
for r in results:
qtype = r["question_type"]
verdict = r.get(judgment_key, {}).get("verdict", "N/A")
if qtype not in stats:
stats[qtype] = {"correct": 0, "total": 0}
stats[qtype]["total"] += 1
if verdict == "yes":
correct += 1
stats[qtype]["correct"] += 1
return correct, stats
agentic_correct, agentic_type_stats = _accumulate("agentic_judgment")
total = len(results)
# Per-item verdict rows
for r in results:
a_verdict = r.get("agentic_judgment", {}).get("verdict", "N/A")
print(f" [{r['question_id']}] type={r['question_type']} agentic={a_verdict}")
print("\n" + "-" * 60)
print(f" Items: {total}")
# Agentic stats
print("\n ── Agentic (ReAct) ──")
print(f" Overall accuracy: {agentic_correct}/{total} ({100*agentic_correct/total:.1f}%)")
tool_call_totals = [sum(r.get("agentic_tool_counts", {}).values()) for r in results]
tool_call_mean, tool_call_std = _mean_and_std(tool_call_totals)
print(f" Tool calls/query: mean={tool_call_mean:.2f} std={tool_call_std:.2f}")
token_usages = [r.get("agentic_token_usage", {}) for r in results]
print(" Bench reported tokens/query:")
for metric in _TOKEN_USAGE_METRICS:
values = [usage[metric] for usage in token_usages if usage.get(metric) is not None]
if values:
mean, std = _mean_and_std(values)
print(f" {metric}: mean={mean:.2f} std={std:.2f}")
else:
print(f" {metric}: unavailable")
print(" Per-type accuracy:")
for qtype, stats in sorted(agentic_type_stats.items()):
acc = 100 * stats["correct"] / stats["total"] if stats["total"] else 0
print(f" {qtype}: {stats['correct']}/{stats['total']} ({acc:.1f}%)")
print("=" * 60)
total_elapsed = time.time() - start_time
print(f"\n Total time: {total_elapsed/60:.1f} min")
print("\n" + "=" * 60)
print(" [DONE] EVALUATION COMPLETED SUCCESSFULLY")
print("=" * 60 + "\n")
_TOKEN_USAGE_METRICS = (
"input_tokens",
"output_tokens",
"total_tokens",
)
def _mean_and_std(values: list[int]) -> tuple[float, float]:
"""Return population mean and standard deviation for one per-query metric."""
if not values:
return 0.0, 0.0
mean = sum(values) / len(values)
return mean, (sum((value - mean) ** 2 for value in values) / len(values)) ** 0.5
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser(description="LongMemEval evaluation runner")
parser.add_argument("--config", type=str, default=None, help="Path to config.yaml")
parser.add_argument(
"--log-level",
type=str,
default="INFO",
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
help="Log level for the eval runner (default: INFO)",
)
parser.add_argument(
"--reme-log-level",
type=str,
default="INFO",
choices=["DEBUG", "INFO", "WARNING", "ERROR"],
help="Log level for reme internal logs — loguru (default: INFO)",
)
parser.add_argument(
"-q",
"--quiet",
action="store_true",
help="Shortcut for --log-level WARNING --reme-log-level WARNING",
)
parser.add_argument(
"--eval_only",
action="store_true",
help="Skip ingestion (phases 1-3). Reuse existing workspaces and only run query+judge.",
)
args = parser.parse_args()
if args.quiet:
args.log_level = "WARNING"
args.reme_log_level = "WARNING"
main(args.config, args.log_level, args.reme_log_level, eval_only=args.eval_only)

View file

@ -1,14 +0,0 @@
# 含真实 API key绝不入库
env.sh
# 运行时产物(含对话内容,勿入库)
logs/
outputs/
reme_workspace/
nanobot_workspace/
# 数据符号链接(指向外部 π-Bench 仓库)
data
__pycache__/
*.pyc

View file

@ -1,327 +0,0 @@
[中文版 / Chinese version](./README_ZH.md)
# π-Bench Evaluation Suite
A glue layer that connects the **ReMe agent (with persistent memory)** to
**π-Bench** (Proactive Personal Assistant Benchmark). This directory contains
only the minimal code and configuration needed for the integration: the
π-Bench framework (`src/`), evaluation data (`data/`), the AppWorld tool
environment, and ReMe itself are all **external third-party dependencies**,
referenced in place via symlink and environment variables and never bundled
with this suite.
- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench (arXiv: 2605.14678)
- ReMe: the root of the ReMe repository this suite lives in (recommended
location: `ReMe/benchmark/pibench/`)
## 1. Architecture
```
π-Bench runner (src.main --mode run)
│ user_agent (simulated-user LLM) walks data/{persona}/episode.yaml
│ task by task, chatting with the agent over multiple turns and judging
│ hidden intents (PROC) during the run phase
test server (π-Bench scripts/test_server.py, HTTP long-polling)
▲ /send │ /poll
│ ▼
bridge_reme.py ──────────────► ReMe Application (embedded as a library)
│ ├─ agent_wrapper: agent under test (AgentScope)
│ ├─ jobs: search / auto_memory / daily_write
│ └─ workspace: reme_workspace/{persona}/
│ (isolated persistent memory per persona)
└──── MCP ────► AppWorld MCP ────► AppWorld APIs (tool/app environment)
π-Bench runner (src.main --mode eval)
judger (judge LLM) reads the traces and scores each checklist item (COMP)
```
Key points:
- The bridge runs on **ReMe's own venv python** and uses ReMe as a library
(`resolve_app_config` + `Application`); **no ReMe source modification** is
required.
- Every incoming user message automatically triggers a ReMe memory `search`
and injects the matched memories (tuning knobs in §8); on task end (reset)
the session is distilled into daily notes by `auto_memory`.
- Tool calls executed by the agent (AppWorld MCP + ReMe job tools) are
captured per turn into the trace as `tool_steps`, so π-Bench
`tools_evaluation_path` scripts can score tool behavior (§7).
- π-Bench's `data/`, `src/` and AppWorld are not part of this suite; install
π-Bench first (§3.1).
## 2. Directory layout
```
pibench/
├── README.md / README_ZH.md # this document (English / Chinese)
├── env.sh.example # environment template (copy to env.sh, fill TODOs)
├── bridge_reme.py # ReMe ↔ test server bridge (memory inject/save,
│ # profile injection, tool-trace capture)
├── run_persona.sh # full pipeline for ONE persona (5 services + run + eval)
├── run_all.sh # batch over 5 personas (fresh/resume, default parallel=2)
├── resume.py # checkpoint resume: completion detection + surgical
│ # cleanup of interrupted tasks' residual memory
├── fix_trace_logs.py # run outputs → ~/.nanobot/trace_logs conversion,
│ # merging tool sidecars into turn files (pre-eval)
├── .gitignore # excludes env.sh and all runtime artifacts
└── config/
├── models/reme.yaml # runner model config (model_id=reme)
└── bench/evaluation/trace_history.yaml # trace render policy (shipped with
# the suite; passed via --history-config-path)
```
Generated at runtime (all git-ignored): `data` (symlink), `logs/`, `outputs/`,
`reme_workspace/`, `nanobot_workspace/`.
## 3. Prerequisites (third-party, install first)
### 3.1 π-Bench repository (with AppWorld)
```bash
git clone https://github.com/Simplified-Reasoning/Pi-Bench.git <pi-bench-dir>
cd <pi-bench-dir>
python3.11 -m venv .venv # scripts expect exactly this venv name
source .venv/bin/activate
pip install -e . # pibench runner (src.main)
bash scripts/setup_appworld.sh # install AppWorld and download its data (large)
```
Post-install sanity checks:
```bash
ls data/ # should contain researcher marketer pharmacist law_trainee Financier
.venv/bin/python -c "import src" && echo OK
.venv/bin/appworld --help >/dev/null && echo OK
```
### 3.2 ReMe repository
```bash
cd <reme-dir> # ReMe repository root (contains the reme/ package)
python3.11 -m venv .venv # scripts expect exactly this venv name
source .venv/bin/activate
pip install -e . # or ReMe's own install flow; `import reme` must work
```
Sanity check: `.venv/bin/python -c "import reme; print('ok')"`
## 4. Install this suite (step by step)
1. **Place the suite** (recommended inside the ReMe repo so `REME_DIR` is
inferred automatically):
```bash
cp -r pibench <reme-dir>/benchmark/pibench
cd <reme-dir>/benchmark/pibench
```
If placed elsewhere, set `REME_DIR` explicitly in env.sh later.
2. **Create the environment file and fill in the custom parameters**:
```bash
cp env.sh.example env.sh
```
Open `env.sh`; required items (marked TODO):
| Variable | Description |
|---|---|
| `PI_BENCH_ROOT` | π-Bench repo root (contains `src/` `data/` `.venv` `third_party/appworld`) |
| `USER_API_KEY` | API key of the simulated-user LLM (run phase, hidden-intent judging) |
| `JUDGER_API_KEY` | API key of the judger LLM (eval phase, checklist scoring) |
| `BRAVE_SEARCH_API_KEY` | optional; for the agent's web_search tool, `dummy` when unused |
Optional tuning: `REME_MODEL_NAME` (base model of the agent under test),
`REME_DIR`, `REME_LLM_BASE_URL` (default: DashScope OpenAI-compatible
endpoint).
3. **Link the evaluation data** (referenced in place, never copied):
```bash
ln -s "$PI_BENCH_ROOT/data" data
```
4. **(Optional) adjust model config** `config/models/reme.yaml`:
- `user_agent.model` / `judger.model`: model names for the simulated user
and the judger (literal values; π-Bench only expands `${ENV}` in
base_url/api_key).
- `run.turn_timeout`, `max_tool_iterations`, etc. as needed.
5. **Smoke check** (does not start the evaluation):
```bash
bash -n run_all.sh && bash -n run_persona.sh
source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
```
## 5. Run the evaluation
> ⚠️ For long runs use `screen`, **not nohup** (nohup loses the permission
> context in sandboxed/restricted environments and breaks child processes).
```bash
# Full official run: wipe ALL personas' memory/outputs/traces first (default
# fresh mode, parallel=2)
mkdir -p logs # on a fresh deployment logs/ does not exist yet
screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
# Checkpoint continuation (after an interruption; no wipe, completed tasks skipped)
bash run_all.sh --resume
# Other usages
bash run_all.sh --parallel 1 # sequential
bash run_all.sh --resume --skip-eval # run phase only
bash run_persona.sh researcher # single persona (default --resume semantics)
bash run_persona.sh researcher --fresh
```
Time reference: 5 personas × 20 tasks, parallel=2, fresh full run ≈ 1214 hours.
`run_all.sh` exits non-zero when any persona fails, so upstream automation
cannot mistake a partially failed suite run for a success.
## 6. Port allocation (parallel personas never collide)
| persona | AppWorld API | AppWorld MCP | Test Server | ReMe internal service |
|-------------|------|-------|------|-------|
| marketer | 9001 | 10001 | 9998 | 18766 |
| law_trainee | 9002 | 10002 | 9997 | 18767 |
| pharmacist | 9003 | 10003 | 9996 | 18768 |
| researcher | 9004 | 10004 | 9995 | 18765 |
| Financier | 9005 | 10005 | 9994 | 18769 |
## 7. Outputs and scores
- **Results**: `outputs/reme/{persona}/{task}/eval/results/*_result.json`
- `overall_average_score`: checklist completeness (COMP; the judger scores
each criterion YES/NO, weighted across dependency groups)
- `overall_proactiveness_average_score`: proactiveness (PROC; the
user_agent judges hidden-intent coverage during the run phase; each task
file also carries the global average)
- **Traces**: `~/.nanobot/trace_logs/reme/{persona}/{task}/...` (the scoring
input of the eval phase)
- **Logs**: `logs/` (`suite_<persona>.log` per persona; `bridge_*`,
`runner_run/eval_*`, `appworld_*`, `test_server_*` per service)
- **Memory store**: `reme_workspace/{persona}/` (daily/digest notes, raw
session dialogs, BM25 index, etc.; persistent across runs, wiped only in
fresh mode)
Score summary:
```bash
grep -h "overall_average_score\|overall_proactiveness" \
outputs/reme/*/*/eval/results/*_result.json | head
```
### Tool-trace capture (tools_evaluation support)
Some tasks define `objectives.tools_evaluation_path`: Python scripts that
score tool behavior (e.g. "the temporary Todoist board was created and
removed"). They need the executed tool calls in the trace. The pipeline:
1. During `reply()`, the bridge reads the persisted AgentScope session state
after each turn and extracts the new `tool_call` / `tool_result` blocks
(tool name, arguments, result).
2. Records are appended to
`outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`, tagged with the
turn number; AgentScope MCP names (`mcp__AppWorld__<tool>`) are normalized
to the π-Bench convention (`mcp_appworld_<tool>`).
3. `fix_trace_logs.py` pairs each `{ts}-messages.jsonl` run with the
temporally closest tools sidecar and merges the records into the generated
`turn_N.json` files under the `tool_steps` key — one of the two
tool-history formats understood by π-Bench's `collect_tool_history()`.
4. The eval phase then feeds `tool_steps` to both the tools_evaluation
scripts and the rendered `<tool_trace_extracts>` seen by the judger.
## 8. Memory mechanism (core design of this suite)
- **Persona isolation**: each persona has its own workspace
(`reme_workspace/{persona}/`); the bridge takes an exclusive
`.bridge.lock` on it at startup, so two bridges can never share one memory
store, and one persona's memory search can never reach another's memories.
- **Writes**: on task end (runner sends reset), the session is distilled by
the `auto_memory` job into daily notes and indexed by the background
watcher (BM25). Saves are non-blocking background tasks; the first message
of a new session waits for in-flight writes before searching.
- **Reads**: on every incoming user message the bridge runs one `search` and
injects matched memories (`[Relevant memories from previous sessions]`
prefix); without matches the message passes through unchanged. Retrieval
tuning (bridge CLI flags, adjustable in run_persona.sh):
- `--search-limit 3`: at most 3 memory chunks injected per message;
- `--search-min-score 2.0`: weak BM25 hits are filtered out;
- `tool_context_id` rotates per task: chunks already injected within the
same task are not re-injected (ReMe's seen-chunk dedup, 24h TTL); normal
recall resumes after task boundaries.
- **No self-leakage**: the in-progress session is not in the store yet
(saves happen on reset), so a task can never retrieve its own unfinished
content.
- The agent also holds `search`/`daily_write` tools and can retrieve/record
proactively.
- **System prompt**: `bridge_reme.py:build_system_prompt()` embeds the
HIDDEN-NEEDS protocol (proactiveness-oriented) and injects the persona
profile from `data/{persona}/profile.yaml` into every turn's system prompt.
## 9. Checkpoint resume and memory-cleanup semantics
- **Completion detection** (resume.py): scans
`outputs/reme/{persona}/**/history/*-log.jsonl` and
`outputs/reme/{persona}/run/*-log.jsonl` for
`Task finished task_id=X status=Y`. The status with the **newest event
timestamp** wins per task (record `timestamp`, falling back to
`timestamp_iso`, then to the timestamp embedded in the log file name) —
file category and read order alone can never override a newer record, so an
old run-level SUCCESS cannot mask a newer per-task ERROR. `SUCCESS /
MAX_TURNS / TIMEOUT` count as completed; `ERROR` and never-started tasks
are re-run (passed to the runner as repeated `--task-id` flags in episode
order).
- **Answer-leak prevention**: an interrupted task may already have been
distilled into daily notes during graceful shutdown; re-running it with
that memory injected would inflate scores. Before resuming,
`resume.py cleanup` therefore removes residual memory **only for tasks
about to be re-run** (daily/digest notes, session/dialog, mem_session;
matched via `session_id = pibench_{task}_*`). Completed tasks' memories are
never touched. Daily index files are refreshed **only for the dates that
lost notes**, by full workspace-relative wikilink path — and when the ReMe
package is importable, the refresh reuses ReMe's own daily-index rebuild
logic (`refresh_day_index`), so same-named notes on other dates are never
modified.
- **fresh vs resume are mutually exclusive**: a full memory wipe belongs to
fresh mode only (`run_all.sh` default, executed before any service starts);
resume never wipes.
## 10. Customization entry points
| Goal | Location |
|---|---|
| Base model of the agent under test | `REME_MODEL_NAME` in `env.sh` |
| user_agent / judger models | `config/models/reme.yaml` |
| Agent system prompt | `bridge_reme.py` `build_system_prompt()` |
| Memory retrieval limit/threshold | `--search-limit/--search-min-score` on the bridge command in `run_persona.sh` |
| ReMe internal parameters | **Do not modify ReMe source**; write a dedicated config modeled on `reme/config/beam.yaml` and override via `resolve_app_config(config=...)` (see bridge `_init_reme_app`) |
| Turn timeout / tool iteration cap | `config/models/reme.yaml` `run.turn_timeout`, `model.max_tool_iterations` |
## 11. Troubleshooting
- **Port already in use**: the scripts auto-kill residual processes on the
four port groups above; if another suite (e.g. a different π-Bench
experiment) holds them, stop it first or change the port table in
run_persona.sh.
- **Bridge exits immediately with workspace locked**: another bridge already
holds the same workspace; make sure each persona uses its own
`--workspace-dir` (the scripts allocate one per persona).
- **Runner reports `${USER_API_KEY} ... empty`**: env.sh is unfilled or not
sourced; run_persona.sh sources env.sh automatically — when running the
runner manually, `source env.sh` first.
- **`Cannot import 'reme'`**: the bridge must run with
`${REME_DIR}/.venv/bin/python` (run_persona.sh already does); otherwise
check that `REME_DIR` points at the ReMe repository root.
- **AppWorld fails to start**: run `bash scripts/setup_appworld.sh` in the
π-Bench repo first (downloads data); inspect
`logs/appworld_*_<persona>.log`.
- **trace_history.yaml not found**: the runner needs
`config/bench/evaluation/trace_history.yaml`; this suite ships the file and
passes it explicitly via `--history-config-path`, and run_persona.sh fails
fast with a clear error if it is missing. Always launch run_persona.sh /
run_all.sh from the suite directory.
## 12. Privacy and security
- The suite code and config templates contain **no real API keys, user names
or absolute paths**; real keys live only in your local `env.sh`
(git-ignored).
- `logs/`, `outputs/`, `reme_workspace/` and `nanobot_workspace/` contain
full conversations and model outputs; never commit or share them.
- The `data` symlink points at the official π-Bench evaluation data; respect
its data license terms.

View file

@ -1,284 +0,0 @@
# π-Bench 评测说明
[English version](./README.md)
**ReMe agent带持久记忆** 接入 **π-Bench**Proactive Personal Assistant
Benchmark的胶水层评测套件。只含对接所需的最小代码与配置π-Bench 框架
`src/`)、评测数据(`data/`、AppWorld 工具环境、ReMe 本体均为**外部第三方
依赖**,通过符号链接与环境变量原位引用,不随本套件分发。
- π-Bench: https://github.com/Simplified-Reasoning/Pi-Bench arXiv: 2605.14678
- ReMe: 你所在 ReMe 仓库的根目录(本套件推荐放在 `ReMe/benchmark/pibench/`
## 1. 架构总览
```
π-Bench runner (src.main --mode run)
│ user_agent模拟用户 LLM按 data/{persona}/episode.yaml 顺序
│ 逐任务、多轮地与 agent 对话,并在 run 阶段判定隐藏意图(PROC)
test server (π-Bench scripts/test_server.py, HTTP 长轮询)
▲ /send │ /poll
│ ▼
bridge_reme.py ──────────────► ReMe Application以库方式内嵌启动
│ ├─ agent_wrapper: 被测 agentAgentScope
│ ├─ jobs: search / auto_memory / daily_write
│ └─ workspace: reme_workspace/{persona}/
│ (每 persona 独立持久记忆库,互不可见)
└──── MCP ────► AppWorld MCP ────► AppWorld API工具/应用环境)
π-Bench runner (src.main --mode eval)
judger裁判 LLM读取 trace按 checklist 逐条 YES/NO 打分(COMP)
```
要点:
- bridge 用 **ReMe 自己的 venv python** 运行,把 ReMe 当库用(`resolve_app_config`
+ `Application`**ReMe 源码零改动**。
- 每条用户消息都会自动触发一次 ReMe memory `search` 并把命中记忆注入当前消息
(参数见 §8任务结束reset时会话被 `auto_memory` 提炼为 daily 笔记落盘。
- agent 执行的每一轮工具调用AppWorld MCP + ReMe job 工具)都会被采集并以
`tool_steps` 形式写入 trace供 π-Bench 的 `tools_evaluation_path` 脚本
对工具行为评分§7
- π-Bench 的 `data/``src/`、AppWorld 均不属于本套件,需先装好 π-Bench§3.1)。
## 2. 目录结构
```
pibench/
├── README.md / README_ZH.md # 本文档(英文 / 中文)
├── env.sh.example # 环境配置模板(复制为 env.sh 后填写 TODO 项)
├── bridge_reme.py # ReMe ↔ test server 桥接(记忆注入/保存、
│ # profile 注入、工具调用轨迹采集)
├── run_persona.sh # 单 persona 全流程5 个服务 + run + eval
├── run_all.sh # 5 个 persona 批跑fresh/resume默认 2 并行)
├── resume.py # 断点续跑:完成判定 + 中断任务残留记忆的外科清理
├── fix_trace_logs.py # run 输出 → ~/.nanobot/trace_logs 转换,
│ # 并把工具轨迹合并进 turn 文件eval 前置)
├── .gitignore # 排除 env.sh 与全部运行产物
└── config/
├── models/reme.yaml # runner 模型配置model_id=reme
└── bench/evaluation/trace_history.yaml # trace 渲染策略(随套件提供,
# 经 --history-config-path 显式传入)
```
运行时自动生成(均被 .gitignore 排除):`data`(符号链接)、`logs/`
`outputs/``reme_workspace/``nanobot_workspace/`
## 3. 前置依赖(第三方,先装好)
### 3.1 π-Bench 仓库(含 AppWorld
```bash
git clone https://github.com/Simplified-Reasoning/Pi-Bench.git <pi-bench-dir>
cd <pi-bench-dir>
python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
source .venv/bin/activate
pip install -e . # pibench runnersrc.main
bash scripts/setup_appworld.sh # 安装 AppWorld 并下载其数据(体积较大,需网络)
```
装完自检:
```bash
ls data/ # 应含 researcher marketer pharmacist law_trainee Financier
.venv/bin/python -c "import src" && echo OK
.venv/bin/appworld --help >/dev/null && echo OK
```
### 3.2 ReMe 仓库
```bash
cd <reme-dir> # ReMe 仓库根目录(含 reme/ 包)
python3.11 -m venv .venv # 脚本约定使用 .venv 这个目录名
source .venv/bin/activate
pip install -e . # 或按 ReMe 自身安装方式,保证 `import reme` 可用
```
自检:`.venv/bin/python -c "import reme; print('ok')"`
## 4. 安装本套件(逐步)
1. **放置套件**(推荐放进 ReMe 仓库,`REME_DIR` 可自动推断):
```bash
cp -r pibench <reme-dir>/benchmark/pibench
cd <reme-dir>/benchmark/pibench
```
若放在其他位置,稍后在 env.sh 中显式设置 `REME_DIR`
2. **创建环境文件并填写自定义参数**
```bash
cp env.sh.example env.sh
```
打开 `env.sh`,必填项(标 TODO 的):
| 变量 | 说明 |
|---|---|
| `PI_BENCH_ROOT` | π-Bench 仓库根目录(含 `src/` `data/` `.venv` `third_party/appworld` |
| `USER_API_KEY` | 模拟用户 LLM 的 API keyrun 阶段判定隐藏意图) |
| `JUDGER_API_KEY` | 裁判 LLM 的 API keyeval 阶段 checklist 打分) |
| `BRAVE_SEARCH_API_KEY` | 可选agent 的 web_search 工具用,不用填 `dummy` |
可选调整:`REME_MODEL_NAME`(被测 agent 基模)、`REME_DIR`
`REME_LLM_BASE_URL`(默认 DashScope OpenAI 兼容端点)。
3. **链接评测数据**(π-Bench 数据原位引用,不复制):
```bash
ln -s "$PI_BENCH_ROOT/data" data
```
4. **(可选)调整模型配置** `config/models/reme.yaml`
- `user_agent.model` / `judger.model`:模拟用户与裁判的模型名(字面量,
π-Bench 仅对 base_url/api_key 做 `${ENV}` 展开)。
- `run.turn_timeout``max_tool_iterations` 等按需。
5. **冒烟自检**(不启动评测):
```bash
bash -n run_all.sh && bash -n run_persona.sh
source env.sh && "$REME_DIR/.venv/bin/python" -c "import reme; print('reme ok')"
```
## 5. 运行评测
> ⚠️ 长时间运行请放进 `screen`**不要用 nohup**nohup 在沙箱/受限环境下
> 会丢失权限上下文导致子进程异常)。
```bash
# 完整正式评测:先清空全部 persona 的记忆/输出/trace再从头跑默认 fresh2 并行)
mkdir -p logs # 全新部署时 logs/ 尚不存在,先建再重定向
screen -dmS pibench_suite bash -c "cd $(pwd) && bash run_all.sh > logs/run_all_master.log 2>&1"
# 断点续跑(中断后继续;不清记忆,跳过已完成任务)
bash run_all.sh --resume
# 其他用法
bash run_all.sh --parallel 1 # 串行
bash run_all.sh --resume --skip-eval # 只跑 run 阶段
bash run_persona.sh researcher # 单 persona默认 --resume 语义)
bash run_persona.sh researcher --fresh
```
耗时参考5 persona × 20 任务、2 并行fresh 全量约 1214 小时。
任一 persona 失败时 `run_all.sh` 以非零状态退出,上层自动化不会把部分失败
的评测误判为成功。
## 6. 端口分配(多 persona 并行互不冲突)
| persona | AppWorld API | AppWorld MCP | Test Server | ReMe 内部服务 |
|-------------|------|-------|------|-------|
| marketer | 9001 | 10001 | 9998 | 18766 |
| law_trainee | 9002 | 10002 | 9997 | 18767 |
| pharmacist | 9003 | 10003 | 9996 | 18768 |
| researcher | 9004 | 10004 | 9995 | 18765 |
| Financier | 9005 | 10005 | 9994 | 18769 |
## 7. 输出与分数
- **结果**`outputs/reme/{persona}/{task}/eval/results/*_result.json`
- `overall_average_score`checklist 完整度COMPjudger 逐条 YES/NO 按依赖组加权)
- `overall_proactiveness_average_score`主动性PROCrun 阶段 user_agent
判定隐藏意图覆盖率;每个任务文件同时携带全局均值)
- **trace**`~/.nanobot/trace_logs/reme/{persona}/{task}/...`eval 的判分输入)
- **日志**`logs/``suite_<persona>.log` 为每 persona 总日志,`bridge_*`
`runner_run/eval_*``appworld_*``test_server_*` 分服务)
- **记忆库**`reme_workspace/{persona}/`daily/digest 笔记、session 原始对话、
BM25 索引等跨运行持久fresh 才清空)
查看汇总:
```bash
grep -h "overall_average_score\|overall_proactiveness" \
outputs/reme/*/*/eval/results/*_result.json | head
```
### 工具轨迹采集tools_evaluation 支持)
部分任务定义了 `objectives.tools_evaluation_path`:用 Python 脚本对工具行为
打分(例如"临时 Todoist 看板已创建并被删除")。这些脚本需要 trace 里有真实
的工具调用记录。采集链路:
1. 每轮 `reply()` 之后bridge 读取 AgentScope 落盘的会话状态,提取本轮新增
`tool_call` / `tool_result` 块(工具名、参数、结果)。
2. 记录按 turn 编号追加写入
`outputs/reme/{persona}/{task}/history/{ts}-tools.jsonl`AgentScope 的
MCP 工具名(`mcp__AppWorld__<tool>`)会规范化为 π-Bench 约定
`mcp_appworld_<tool>`)。
3. `fix_trace_logs.py` 将每个 `{ts}-messages.jsonl` 运行与时间上最接近的
tools 旁路文件配对,把记录合并进生成的 `turn_N.json``tool_steps`
字段——这是 π-Bench `collect_tool_history()` 支持的两种工具轨迹格式之一。
4. eval 阶段 `tool_steps` 既提供给 tools_evaluation 脚本,也会被渲染为
judger 可见的 `<tool_trace_extracts>`
## 8. 记忆机制(本套件的核心设计)
- **persona 隔离**:每个 persona 独立 workspace`reme_workspace/{persona}/`
bridge 启动时对 workspace 加 `.bridge.lock` 排他锁,两个 bridge 不可能共用
同一记忆库;一个 persona 的 memory search 永远接触不到其他 persona 的记忆。
- **写入**任务结束runner 发送 reset会话经 `auto_memory` job 提炼为
daily 笔记落盘,后台 watcher 建 BM25 索引。保存为非阻塞后台任务,
新会话首条消息会先等待在途写入完成再检索。
- **读取**bridge 每收到一条用户消息自动 `search` 一次并注入命中记忆
`[Relevant memories from previous sessions]` 前缀),无命中则原样透传。
检索参数bridge 命令行,可在 run_persona.sh 中调整):
- `--search-limit 3`:每条消息最多注入 3 个记忆块;
- `--search-min-score 2.0`:过滤弱 BM25 命中;
- `tool_context_id` 按任务轮换:同一任务内已注入的记忆块不重复注入
ReMe 自带 seen-chunk 去重24h TTL任务边界后恢复正常召回。
- **无自泄漏**进行中的会话尚未入库save 发生在 reset任务不会检索到
自己未完成的内容。
- agent 同时持有 `search`/`daily_write` 工具,可主动检索/记录。
- **system prompt**`bridge_reme.py:build_system_prompt()` 内置
HIDDEN-NEEDS 协议(面向 proactiveness并把 `data/{persona}/profile.yaml`
的 persona profile 注入每轮 system prompt。
## 9. 断点续跑与记忆清理语义
- **完成判定**resume.py扫描 `outputs/reme/{persona}/**/history/*-log.jsonl`
`outputs/reme/{persona}/run/*-log.jsonl` 中的
`Task finished task_id=X status=Y`。每个任务以**事件时间最新**的记录为准
(优先取记录的 `timestamp`,回退 `timestamp_iso`,再回退日志文件名中的
时间戳)——文件类别与读取顺序本身不能覆盖更新的记录,因此旧的 run 级
SUCCESS 不会掩盖更新的 per-task ERROR。`SUCCESS/MAX_TURNS/TIMEOUT` 记为
完成,`ERROR`/未开始的任务重跑(按 episode 顺序以 `--task-id` 传给 runner
- **防答案泄漏**:被中断的任务可能已在优雅退出时提炼成 daily 笔记,直接重跑会
把答案注入、抬高分数。因此 resume 启动前 `resume.py cleanup` **只删除待重跑
任务**的残留记忆daily/digest 笔记、session/dialog、mem_session
`session_id = pibench_{task}_*` 匹配已完成任务的记忆一律不动。daily
索引**只刷新实际发生删除的日期**,按完整的 workspace 相对 wikilink 路径
匹配;当 ReMe 包可导入时,刷新直接复用 ReMe 自带的 daily 索引重建逻辑
`refresh_day_index`),不会误改其他日期下的同名笔记条目。
- **fresh vs resume 互斥**:全量清记忆只属于 fresh 模式(`run_all.sh` 默认,
在任何服务启动前执行resume 永不清全量。
## 10. 自定义与调优入口
| 目标 | 位置 |
|---|---|
| 被测 agent 基模 | `env.sh``REME_MODEL_NAME` |
| user_agent / judger 模型 | `config/models/reme.yaml` |
| agent system prompt | `bridge_reme.py` `build_system_prompt()` |
| 记忆检索条数/阈值 | `run_persona.sh` bridge 启动命令的 `--search-limit/--search-min-score` |
| ReMe 内部参数 | **不要改 ReMe 源码**;仿照 `reme/config/beam.yaml` 写专有配置,经 `resolve_app_config(config=...)` 覆盖(见 bridge `_init_reme_app` |
| 轮超时/工具迭代上限 | `config/models/reme.yaml` `run.turn_timeout``model.max_tool_iterations` |
## 11. 故障排查
- **端口被占用**:脚本会自动 kill 上述 4 组端口上的残留进程;若与其他套件
(如别的 π-Bench 实验)冲突,请先停掉对方或改 run_persona.sh 的端口表。
- **bridge 启动即退出,提示 workspace locked**:另一个 bridge 正占用同一
workspace确认每个 persona 用各自的 `--workspace-dir`(脚本已按 persona 分配)。
- **runner 报 `${USER_API_KEY} ... empty`**env.sh 未填写或未生效;
run_persona.sh 会自动 source env.sh手动运行 runner 时请先 `source env.sh`
- **`Cannot import 'reme'`**bridge 必须用 `${REME_DIR}/.venv/bin/python` 运行
run_persona.sh 已如此),或检查 `REME_DIR` 是否指向 ReMe 仓库根目录。
- **AppWorld 启动失败**:先在 π-Bench 仓库执行 `bash scripts/setup_appworld.sh`
下载数据;查看 `logs/appworld_*_<persona>.log`
- **trace_history.yaml 找不到**runner 需要
`config/bench/evaluation/trace_history.yaml`;本套件已随附该文件并通过
`--history-config-path` 显式传入run_persona.sh 启动前会做存在性检查,
缺失时立即报出清晰错误。请始终从套件目录启动 run_persona.sh / run_all.sh。
## 12. 隐私与安全
- 套件代码与配置模板中**不含任何真实 API key、用户名或绝对路径**
真实 key 只存在于你本地的 `env.sh`(已被 .gitignore 排除)。
- `logs/``outputs/``reme_workspace/``nanobot_workspace/` 含完整对话内容
与模型输出,请勿提交仓库或外传。
- `data` 符号链接指向 π-Bench 官方评测数据,请遵守其数据许可条款。

File diff suppressed because it is too large Load diff

View file

@ -1,53 +0,0 @@
version: 1
format:
root_tag: trace
turn_tag: turn
message_tag: message
file_tag: file
tool_call_tag_prefix: tool_call
tool_result_tag_prefix: tool_result
text_policy:
default:
truncate_chars: 1200
mask_newlines: false
field_overrides:
files_read:
truncate_chars: 40000
assistant_content:
truncate_chars: 40000
tool_result_content:
truncate_chars: 40000
fields:
turn:
include_session_key: false
files:
enabled: true
messages:
enabled: true
include_message_role_attr: true
include_message_index_attr: false
include_system: false
include_user: true
include_assistant_thinking_content: false
include_assistant_thinking_reasoning: false
include_assistant_content: true
include_assistant_reasoning: false
include_assistant_tool_calls: false
require_matching_tool_call: true
tool_calls:
include_tool_call_id: false
tools:
web_fetch:
enabled: true
include_tool_call_keys: [url]
include_tool_result: false
web_search:
enabled: true
include_tool_call_keys: [query]
include_tool_result: false

View file

@ -1,40 +0,0 @@
# ReMe model configuration for Pi-Bench
# Uses ReMe's AgentScope agent with Dashscope as the LLM backend
model:
model: reme
base_url: "http://localhost:8088"
api_key: "dummy"
provider: custom
max_tokens: 16384
max_tool_iterations: 120
memory_window: 100
user_agent:
model: qwen3.8-max
base_url: "${USER_BASE_URL}"
api_key: "${USER_API_KEY}"
temperature: 0.0
request_timeout: 360.0
judger:
model: qwen3.8-max
base_url: "${JUDGER_BASE_URL}"
api_key: "${JUDGER_API_KEY}"
temperature: 0.0
request_timeout: 360.0
tools:
brave_search_api_key: "${BRAVE_SEARCH_API_KEY}"
web_search_max_results: 10
nanobot:
trace_logs_dir: "~/.nanobot/trace_logs"
workspace_dir: "~/.nanobot/workspace"
copy_task_assets_to_workspace: true
run:
output_dir: outputs
log_level: INFO
user_mode: llm
turn_timeout: 2400.0

View file

@ -1,57 +0,0 @@
#!/bin/bash
# ═══════════════════════════════════════════════════════════════════════
# pibench evaluation suite - environment configuration template
# Usage: cp env.sh.example env.sh, then fill in the TODO items below.
# ⚠️ env.sh contains real API keys; never commit or share it
# (already excluded via .gitignore).
# ═══════════════════════════════════════════════════════════════════════
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# ─── TODO: π-Bench repository root ────────────────────────────────────
# Must contain src/, data/, scripts/test_server.py, third_party/appworld
# and .venv (see README setup).
export PI_BENCH_ROOT=""
# ─── ReMe repository ──────────────────────────────────────────────────
# Defaults to two levels above this directory (the layout this suite uses
# when placed at ReMe/benchmark/pibench); point it at the actual ReMe
# repository root if the suite lives elsewhere.
export REME_DIR="${REME_DIR:-$(cd "${SUITE_DIR}/../.." && pwd)}"
# ─── Base model of the agent under test (LLM used by the ReMe agent) ──
export REME_MODEL_NAME="${REME_MODEL_NAME:-qwen3.6-plus}"
# ─── LLM service endpoint (default: DashScope OpenAI-compatible; any
# OpenAI-compatible endpoint works) ────────────────────────────────
DASHSCOPE_BASE_URL="https://dashscope.aliyuncs.com/compatible-mode/v1"
export REME_LLM_BASE_URL="${REME_LLM_BASE_URL:-${DASHSCOPE_BASE_URL}}"
# ─── TODO: API keys ───────────────────────────────────────────────────
# USER_API_KEY : drives the simulated user LLM (run phase; judges whether
# hidden intents are satisfied and asks follow-ups)
# JUDGER_API_KEY: drives the judger LLM (eval phase; scores the checklist)
# The two may be identical; one strong model is recommended for both.
export USER_BASE_URL="${DASHSCOPE_BASE_URL}"
export USER_API_KEY="TODO-fill-in-user-agent-api-key"
export JUDGER_BASE_URL="${DASHSCOPE_BASE_URL}"
export JUDGER_API_KEY="TODO-fill-in-judger-api-key"
# The ReMe agent's key reuses USER_API_KEY by default (no need to repeat
# it when both use the same service and key).
export REME_LLM_API_KEY="${REME_LLM_API_KEY:-${USER_API_KEY}}"
# Brave Search (optional; used by the agent's web_search tool - use
# "dummy" when not needed).
export BRAVE_SEARCH_API_KEY="TODO-optional-brave-search-key-or-dummy"
# ─── Persistent memory workspaces (one subdirectory per persona,
# created automatically) ───────────────────────────────────────────
export REME_WORKSPACE_ROOT="${REME_WORKSPACE_ROOT:-${SUITE_DIR}/reme_workspace}"
# ─── Variables consumed by ReMe's default.yaml model config expansion;
# do not remove ────────────────────────────────────────────────────
export LLM_MODEL_NAME="${REME_MODEL_NAME}"
export LLM_BASE_URL="${REME_LLM_BASE_URL}"
export LLM_API_KEY="${REME_LLM_API_KEY}"

View file

@ -1,198 +0,0 @@
#!/usr/bin/env python3
"""Convert reme_eval run outputs into eval-compatible trace logs.
outputs/{model_id}/{user_id}/{task_id}/history/{ts}-messages.jsonl
-> ~/.nanobot/trace_logs/{model_id}/{user_id}/{task_id}/{ts}/turn_N.json
The bridge additionally writes {ts}-tools.jsonl sidecar files next to the
message histories: one JSON object per executed tool call with fields
{turn, name, arguments, result}. Each messages run is paired with the
temporally closest sidecar, and the records are merged into the generated
turn files under the "tool_steps" key, which is one of the tool-history
formats π-Bench's collect_tool_history() understands. Without this step,
tools_evaluation scripts would see no tool evidence at all.
Usage: python fix_trace_logs.py [user_id ...] (no args = all users)
"""
import json
import re
import sys
from datetime import datetime
from pathlib import Path
SUITE_DIR = Path(__file__).resolve().parent
OUTPUTS_DIR = SUITE_DIR / "outputs"
TRACE_LOGS_DIR = Path.home() / ".nanobot" / "trace_logs"
MESSAGES_FILE_RE = re.compile(r"^(\d{8}_\d{6})-messages\.jsonl$")
TOOLS_FILE_RE = re.compile(r"^(\d{8}_\d{6})-tools\.jsonl$")
TIME_FORMAT = "%Y%m%d_%H%M%S"
# A tool sidecar belongs to the messages run that started at most this many
# seconds earlier (the bridge stamps the sidecar when the task's first user
# message arrives, shortly after the runner opened the messages file).
MAX_PAIR_DELTA_SECONDS = 6 * 3600
def _to_epoch(timestamp: str) -> float:
"""Parse a YYYYMMDD_HHMMSS timestamp into epoch seconds."""
try:
return datetime.strptime(timestamp, TIME_FORMAT).timestamp()
except ValueError:
return 0.0
def load_tool_records(tools_file: Path) -> dict:
"""Group sidecar tool records by turn number."""
by_turn: dict = {}
try:
with open(tools_file, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
continue
if not isinstance(record, dict) or not record.get("name"):
continue
turn = int(record.get("turn") or 0)
by_turn.setdefault(turn, []).append(
{
"name": record["name"],
"arguments": record.get("arguments", {}),
"result": record.get("result", ""),
},
)
except OSError as exc:
print(f" WARNING: cannot read tool sidecar {tools_file}: {exc}")
return by_turn
def pair_tool_sidecars(message_runs: list, tool_runs: list) -> dict:
"""Pair each messages run with the temporally closest unused tool sidecar.
Fresh runs produce exactly one messages file and one sidecar per task;
re-runs append matching pairs, so sorted greedy nearest-timestamp
matching is stable. Sidecars farther away than MAX_PAIR_DELTA_SECONDS
(e.g. leftovers of a crashed bridge) stay unpaired.
"""
pairing: dict = {}
unused = list(tool_runs)
for msg_ts, _ in message_runs:
best_delta = None
best_item = None
for tool_ts, tool_path in unused:
delta = abs(_to_epoch(tool_ts) - _to_epoch(msg_ts))
if best_delta is None or delta < best_delta:
best_delta = delta
best_item = (tool_ts, tool_path)
if best_delta is not None and best_item is not None and best_delta <= MAX_PAIR_DELTA_SECONDS:
pairing[msg_ts] = best_item[1]
unused.remove(best_item)
return pairing
def build_turns(messages: list) -> list:
"""Split the flat message list into per-turn [user, assistant] groups."""
turns = []
i = 0
while i < len(messages):
turn_msgs = []
if messages[i]["role"] == "user":
turn_msgs.append({"role": "user", "content": messages[i]["message"]})
i += 1
if i < len(messages) and messages[i]["role"] == "assistant":
turn_msgs.append({"role": "assistant", "content": messages[i]["message"]})
i += 1
if not turn_msgs:
i += 1 # defensive: never spin on unexpected roles
continue
turns.append(turn_msgs)
return turns
def convert_task(model_id: str, user_id: str, task_dir: Path) -> None:
"""Convert one task's history dir into trace turn files with tool_steps."""
history_dir = task_dir / "history"
if not history_dir.is_dir():
return
message_runs = []
tool_runs = []
for msg_file in history_dir.glob("*-messages.jsonl"):
match = MESSAGES_FILE_RE.match(msg_file.name)
if match:
message_runs.append((match.group(1), msg_file))
for tools_file in history_dir.glob("*-tools.jsonl"):
match = TOOLS_FILE_RE.match(tools_file.name)
if match:
tool_runs.append((match.group(1), tools_file))
if not message_runs:
return
message_runs.sort(key=lambda item: item[0])
tool_runs.sort(key=lambda item: item[0])
pairing = pair_tool_sidecars(message_runs, tool_runs)
print(f"\n{model_id}/{user_id}/{task_dir.name}")
for timestamp, msg_file in message_runs:
trace_dir = TRACE_LOGS_DIR / model_id / user_id / task_dir.name / timestamp
trace_dir.mkdir(parents=True, exist_ok=True)
messages = []
with open(msg_file, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
msg = json.loads(line)
if msg.get("role") == "user" and msg.get("message") == "/new":
continue
messages.append(msg)
tools_file = pairing.get(timestamp)
tools_by_turn = load_tool_records(tools_file) if tools_file else {}
if tools_file is not None:
print(f" {timestamp}: paired tool sidecar {tools_file.name}")
turns = build_turns(messages)
for turn_idx, turn_msgs in enumerate(turns, start=1):
turn_data = {"messages": turn_msgs}
tool_steps = tools_by_turn.get(turn_idx)
if tool_steps:
turn_data["tool_steps"] = tool_steps
turn_file = trace_dir / f"turn_{turn_idx}.json"
with open(turn_file, "w", encoding="utf-8") as f:
json.dump(turn_data, f, indent=2, ensure_ascii=False)
tool_total = sum(len(steps) for steps in tools_by_turn.values())
print(f" {timestamp}: {len(turns)} turns, {tool_total} tool step(s) -> {trace_dir}")
def convert_outputs(user_filter=None):
"""Convert message history JSONL files into per-turn trace JSON files."""
if not OUTPUTS_DIR.exists():
print(f"outputs dir not found: {OUTPUTS_DIR}")
return
for model_dir in sorted(OUTPUTS_DIR.iterdir()):
if not model_dir.is_dir():
continue
model_id = model_dir.name
for user_dir in sorted(model_dir.iterdir()):
if not user_dir.is_dir():
continue
user_id = user_dir.name
if user_filter and user_id not in user_filter:
continue
for task_dir in sorted(user_dir.iterdir()):
if task_dir.is_dir():
convert_task(model_id, user_id, task_dir)
if __name__ == "__main__":
convert_outputs(set(sys.argv[1:]) or None)
print("\ndone")

View file

@ -1,332 +0,0 @@
#!/usr/bin/env python3
"""Checkpoint-resume support for the reme_eval suite.
Completion source of truth:
- outputs/reme/<persona>/<task_id>/history/*-log.jsonl (per-task logs,
flushed incrementally, survive mid-run kills)
- outputs/reme/<persona>/run/*-log.jsonl (run-level logs,
may be truncated if the process was killed before flush)
lines: "Task finished task_id=<id> status=<STATUS>"
A task counts as COMPLETED when its latest terminal status is one of
SUCCESS / MAX_TURNS / TIMEOUT. ERROR or never-started tasks stay pending.
"Latest" is decided by EVENT TIME, not by file category or read order:
each record's "timestamp" (epoch seconds, or "timestamp_iso" as fallback)
is compared across per-task and run-level logs alike, with the timestamp
embedded in the log file name as a last-resort fallback. This keeps an
old run-level SUCCESS from overriding a newer per-task ERROR when the
re-run died before the new run-level log captured the task.
Commands:
remaining <persona> [--json]
Print task_ids still to run, in data/<persona>/episode.yaml order
(one per line; --json prints {"completed": [...], "remaining": [...]}).
cleanup <persona> [--dry-run]
Surgically remove residual memory artifacts of tasks that are about
to be RE-RUN (i.e. pending tasks that left partial state because a
previous run was interrupted). This prevents answer leakage: an
interrupted task's conversation may already have been distilled into
daily notes during graceful shutdown, and re-running the task with
that memory injected would inflate scores.
Removed artifacts (only for pending tasks with residual state):
- daily/<date>/<note>.md whose frontmatter session_id matches
pibench_<task_id>_*, plus a refresh of ONLY the daily index of
the affected date(s) (daily/<date>.md), matched by the full
workspace-relative note path, never by bare file name
- digest notes with matching session_id
- session/dialog/pibench_<task_id>_*.jsonl
- mem_session/**.jsonl files containing pibench_<task_id>_
When the ReMe package is importable, the daily index refresh reuses
ReMe's own rebuild logic (reme.steps.file_io._daily_index.
refresh_day_index); otherwise index lines are dropped by exact
wikilink path match. Either way, indexes of other dates are never
touched. The ReMe watcher (init_changes_step) detects the deleted
daily notes on next bridge startup and removes them from the BM25
index itself.
Completed tasks' memories are NEVER touched by this command.
Design note (resume vs memory-wipe conflict):
A full memory wipe is a suite-level action of fresh mode (run_all.sh
without --resume) and happens before any service starts. Resume mode
never wipes; it only performs the surgical cleanup above. The two modes
are mutually exclusive, so a resumed run can never lose the cross-session
memory accumulated by completed tasks.
"""
import asyncio
import json
import os
import re
import sys
from datetime import datetime
from pathlib import Path
import yaml
try: # Reuse ReMe's daily-index rebuild when running inside the ReMe venv.
from reme.steps.file_io._daily_index import refresh_day_index
except ImportError: # pragma: no cover - depends on runtime venv
refresh_day_index = None
SUITE_DIR = Path(__file__).resolve().parent
DATA_DIR = Path(os.environ.get("REME_EVAL_DATA_DIR", SUITE_DIR / "data")).resolve()
OUTPUTS_DIR = Path(os.environ.get("REME_EVAL_OUTPUTS_DIR", SUITE_DIR / "outputs")) / "reme"
WORKSPACE_ROOT = Path(
os.environ.get("REME_WORKSPACE_ROOT", SUITE_DIR / "reme_workspace"),
).resolve()
COMPLETED_STATUSES = {"SUCCESS", "MAX_TURNS", "TIMEOUT"}
TASK_FINISHED_RE = re.compile(r"Task finished task_id=(\S+) status=(\S+)")
SESSION_ID_RE = re.compile(r"^session_id:\s*(\S+)", re.MULTILINE)
NOTE_COUNT_RE = re.compile(r"(description:\s*)\d+(\s*note\(s\) today)")
LOG_FILE_TS_RE = re.compile(r"^(\d{8}_\d{6})-log\.jsonl$")
TIME_FORMAT = "%Y%m%d_%H%M%S"
def log(msg: str) -> None:
"""Print a status message to stderr."""
print(msg, file=sys.stderr)
def episode_task_order(persona: str) -> list[str]:
"""Return the ordered task ids from the persona's episode.yaml."""
episode_path = DATA_DIR / persona / "episode.yaml"
with open(episode_path, "r", encoding="utf-8") as f:
episode = yaml.safe_load(f)
return [task["task_id"] for task in episode.get("tasks", [])]
def _event_time(record: dict, file_ts: str) -> float:
"""Best-effort event time (epoch seconds) of one log record.
Prefers the record's own timestamp fields; falls back to the timestamp
embedded in the log file name so that even stripped records keep a
meaningful order. Returns 0.0 when nothing is parseable.
"""
timestamp = record.get("timestamp")
if isinstance(timestamp, (int, float)) and not isinstance(timestamp, bool):
return float(timestamp)
iso = record.get("timestamp_iso")
if isinstance(iso, str):
try:
return datetime.fromisoformat(iso).timestamp()
except ValueError:
pass
if file_ts:
try:
return datetime.strptime(file_ts, TIME_FORMAT).timestamp()
except ValueError:
pass
return 0.0
def latest_task_statuses(persona: str) -> dict[str, str]:
"""Scan per-task and run-level logs; the newest EVENT TIME wins per task.
Every "Task finished" record across both log categories is keyed by
(event_time, file timestamp, file order, line number); the record with
the highest key decides the task's status. File category and read order
alone can never override a newer record from the other category.
"""
persona_dir = OUTPUTS_DIR / persona
if not persona_dir.is_dir():
return {}
log_files = sorted(persona_dir.glob("*/history/*-log.jsonl"))
log_files += sorted(persona_dir.glob("run/*-log.jsonl"))
best: dict[str, tuple[tuple, str]] = {}
for file_order, log_file in enumerate(log_files):
ts_match = LOG_FILE_TS_RE.match(log_file.name)
file_ts = ts_match.group(1) if ts_match else ""
try:
with open(log_file, "r", encoding="utf-8") as f:
for line_no, line in enumerate(f):
if "Task finished" not in line:
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
continue
match = TASK_FINISHED_RE.search(str(record.get("message", "")))
if not match:
continue
task_id, status = match.group(1), match.group(2)
sort_key = (_event_time(record, file_ts), file_ts, file_order, line_no)
current = best.get(task_id)
if current is None or sort_key > current[0]:
best[task_id] = (sort_key, status)
except OSError:
continue
return {task_id: status for task_id, (_, status) in best.items()}
def split_tasks(persona: str) -> tuple[list[str], list[str]]:
"""Split the episode task order into completed and remaining tasks."""
order = episode_task_order(persona)
statuses = latest_task_statuses(persona)
completed = [t for t in order if statuses.get(t) in COMPLETED_STATUSES]
remaining = [t for t in order if t not in set(completed)]
return completed, remaining
def _daily_note_session_id(note_path: Path) -> str:
try:
text = note_path.read_text(encoding="utf-8")
except OSError:
return ""
match = SESSION_ID_RE.search(text)
return match.group(1) if match else ""
class _WorkspaceFileStoreShim:
"""Structural stand-in for ReMe's file store; only workspace_path is read."""
def __init__(self, workspace_path: Path):
self.workspace_path = workspace_path
def _refresh_daily_indexes(
workspace: Path,
removed_by_date: dict[str, set[str]],
removed: list[str],
) -> None:
"""Rebuild the daily index of each affected date via ReMe's own logic."""
for date in sorted(removed_by_date):
result = asyncio.run(
refresh_day_index(_WorkspaceFileStoreShim(workspace), date, "daily"),
)
if result.get("error"):
log(f"[resume] WARNING: daily index refresh failed for {date}: {result['error']}")
continue
removed.append(f"daily/{date}.md (refreshed, {len(removed_by_date[date])} note(s) removed)")
def _strip_index_lines(
workspace: Path,
removed_by_date: dict[str, set[str]],
removed: list[str],
dry_run: bool,
) -> None:
"""Fallback index edit: drop lines that reference removed notes by full
workspace-relative wikilink path, and fix the note count. Only the index
files of affected dates are touched."""
for date in sorted(removed_by_date):
index_path = workspace / "daily" / f"{date}.md"
if not index_path.is_file():
continue
wikilinks = [f"[[{rel_path}]]" for rel_path in sorted(removed_by_date[date])]
lines = index_path.read_text(encoding="utf-8").splitlines()
kept = [line for line in lines if not any(link in line for link in wikilinks)]
if len(kept) == len(lines):
continue
note_count = sum(1 for line in kept if line.startswith("- [[daily/"))
kept = [NOTE_COUNT_RE.sub(rf"\g<1>{note_count}\2", line) for line in kept]
removed.append(f"{index_path.relative_to(workspace)} (rewritten)")
if not dry_run:
index_path.write_text("\n".join(kept) + "\n", encoding="utf-8")
def cleanup_partial_memory(persona: str, remaining: list[str], dry_run: bool = False) -> list[str]:
"""Remove partial memory artifacts of remaining tasks so they can be re-run cleanly."""
workspace = WORKSPACE_ROOT / persona
removed: list[str] = []
if not workspace.is_dir() or not remaining:
return removed
prefixes = tuple(f"pibench_{task_id}_" for task_id in remaining)
def act(path: Path, label: str) -> None:
removed.append(label)
if not dry_run:
path.unlink()
# 1) daily / digest notes distilled from interrupted sessions. For daily
# notes, remember the full workspace-relative path grouped by date so only
# the affected daily indexes are refreshed below.
removed_by_date: dict[str, set[str]] = {}
for section in ("daily", "digest"):
section_root = workspace / section
if not section_root.is_dir():
continue
for note_path in section_root.rglob("*.md"):
if note_path.parent == section_root:
continue # index files handled below
session_id = _daily_note_session_id(note_path)
if session_id.startswith(prefixes):
rel_path = note_path.relative_to(workspace).as_posix()
act(note_path, rel_path)
if section == "daily":
removed_by_date.setdefault(note_path.parent.name, set()).add(rel_path)
# 2) daily index files: refresh only the dates that lost notes, matching
# notes by their full wikilink path instead of their bare file name.
if removed_by_date:
if dry_run:
for date in sorted(removed_by_date):
removed.append(f"daily/{date}.md (would refresh index)")
elif refresh_day_index is not None:
_refresh_daily_indexes(workspace, removed_by_date, removed)
else:
_strip_index_lines(workspace, removed_by_date, removed, dry_run)
# 3) raw dialog logs of interrupted sessions
dialog_dir = workspace / "session" / "dialog"
if dialog_dir.is_dir():
for task_id in remaining:
for dialog_path in dialog_dir.glob(f"pibench_{task_id}_*.jsonl"):
act(dialog_path, str(dialog_path.relative_to(workspace)))
# 4) agent-scope session states that contain interrupted-task sessions
mem_session_dir = workspace / "mem_session"
if mem_session_dir.is_dir():
for session_path in mem_session_dir.rglob("*.jsonl"):
try:
content = session_path.read_text(encoding="utf-8", errors="ignore")
except OSError:
continue
if any(prefix in content for prefix in prefixes):
act(session_path, str(session_path.relative_to(workspace)))
return removed
def main() -> int:
"""CLI entrypoint: run 'remaining' or 'cleanup' action for a persona."""
args = sys.argv[1:]
if len(args) < 2 or args[0] not in {"remaining", "cleanup"}:
print(__doc__, file=sys.stderr)
return 2
command, persona = args[0], args[1]
completed, remaining = split_tasks(persona)
if command == "remaining":
if "--json" in args:
print(json.dumps({"completed": completed, "remaining": remaining}))
else:
for task_id in remaining:
print(task_id)
log(
f"[resume] {persona}: completed={len(completed)} "
f"({', '.join(completed) if completed else '-'}) remaining={len(remaining)}",
)
return 0
dry_run = "--dry-run" in args
removed = cleanup_partial_memory(persona, remaining, dry_run=dry_run)
if removed:
verb = "would remove" if dry_run else "removed"
log(f"[resume] {persona}: {verb} {len(removed)} partial-memory artifact(s):")
for item in removed:
log(f" - {item}")
else:
log(f"[resume] {persona}: no partial-memory artifacts to clean")
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -1,119 +0,0 @@
#!/bin/bash
# Run all 5 personas with the ReMe agent, PARALLEL at a time (default 2).
# Each persona's tasks follow data/{persona}/episode.yaml order.
#
# Usage:
# bash run_all.sh # FRESH official run: wipes ALL personas'
# # ReMe memory/outputs/trace logs first,
# # then runs everything from scratch.
# bash run_all.sh --resume # Checkpoint continuation: no wipe; every
# # persona skips already-completed tasks.
# bash run_all.sh --parallel 1 # sequential (original behavior)
# bash run_all.sh --skip-eval # run phase only
#
# Memory-wipe vs resume conflict resolution:
# The full ReMe memory wipe happens ONLY here, ONLY in fresh mode (the
# default), and ONLY before any service/bridge starts. --resume never
# wipes; run_persona.sh then additionally performs a surgical cleanup of
# residual memory belonging to interrupted (to-be-re-run) tasks, so a
# resumed run keeps all completed-task memory but never inherits a partial
# task's own answer. The two modes are mutually exclusive.
set -uo pipefail
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PERSONAS=(researcher marketer law_trainee pharmacist Financier)
TRACE_ROOT="${HOME}/.nanobot/trace_logs"
PARALLEL=2
MODE="fresh"
PASS_ARGS=()
while [[ $# -gt 0 ]]; do
case $1 in
--parallel)
PARALLEL="${2:-}"; shift 2 || true
case "$PARALLEL" in (""|*[!0-9]*) echo "--parallel needs a positive integer"; exit 2 ;; esac
[ "$PARALLEL" -lt 1 ] && PARALLEL=1
[ "$PARALLEL" -gt ${#PERSONAS[@]} ] && PARALLEL=${#PERSONAS[@]}
;;
--resume)
if [ "$MODE" = "fresh_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
MODE="resume"; shift ;;
--fresh)
if [ "$MODE" = "resume" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
MODE="fresh_set"; shift ;;
--skip-eval) PASS_ARGS+=(--skip-eval); shift ;;
*) echo "Unknown option: $1"; exit 1 ;;
esac
done
[ "$MODE" = "fresh_set" ] && MODE="fresh"
START_TS=$(date +%Y%m%d_%H%M%S)
SUMMARY_LOG="${SUITE_DIR}/logs/run_all_${START_TS}.summary"
mkdir -p "${SUITE_DIR}/logs"
echo "############################################################"
echo "# reme_eval suite | mode=${MODE} parallel=${PARALLEL} | ${START_TS}"
echo "############################################################"
# ─── Fresh mode: suite-level wipe BEFORE anything starts ──────────────
if [ "$MODE" = "fresh" ]; then
echo "[fresh] wiping ALL personas' memory workspaces, outputs and trace logs..."
for persona in "${PERSONAS[@]}"; do
rm -rf "${SUITE_DIR}/reme_workspace/${persona}"
rm -rf "${SUITE_DIR}/outputs/reme/${persona}"
rm -rf "${TRACE_ROOT}/reme/${persona}"
rm -rf "${SUITE_DIR}/nanobot_workspace/${persona}"
done
echo "[fresh] wipe done."
else
echo "[resume] no memory wipe; personas resume after their last completed task."
fi
# ─── Run personas in batches of PARALLEL ──────────────────────────────
STATUS_LIST=()
ANY_FAILED=0
OVERALL_START=$(date +%s)
TOTAL=${#PERSONAS[@]}
for ((i = 0; i < TOTAL; i += PARALLEL)); do
BATCH=("${PERSONAS[@]:i:PARALLEL}")
BATCH_PIDS=()
BATCH_NAMES=()
echo ""
echo "============================================================"
echo "# BATCH $(( i / PARALLEL + 1 )): ${BATCH[*]} started $(date '+%F %T')"
echo "============================================================"
for persona in "${BATCH[@]}"; do
bash "${SUITE_DIR}/run_persona.sh" "${persona}" --resume ${PASS_ARGS[@]+"${PASS_ARGS[@]}"} \
> "${SUITE_DIR}/logs/suite_${persona}.log" 2>&1 &
BATCH_PIDS+=($!)
BATCH_NAMES+=("$persona")
done
for j in $(seq 0 $(( ${#BATCH[@]} - 1 ))); do
pid=${BATCH_PIDS[$j]}
persona=${BATCH_NAMES[$j]}
if wait "$pid"; then
STATUS_LIST+=("${persona}: OK")
else
rc=$?
ANY_FAILED=1
STATUS_LIST+=("${persona}: FAILED rc=${rc}")
echo "[run_all] ${persona} FAILED (rc=${rc}); see logs/suite_${persona}.log"
fi
done
done
total=$(( $(date +%s) - OVERALL_START ))
echo ""
echo "================ FINAL SUMMARY (${total}s total) ================" | tee -a "${SUMMARY_LOG}"
for line in "${STATUS_LIST[@]}"; do
echo " ${line}" | tee -a "${SUMMARY_LOG}"
done
echo "Summary: ${SUMMARY_LOG}"
if [ "${ANY_FAILED}" -ne 0 ]; then
FAILED_COUNT=$(printf '%s\n' "${STATUS_LIST[@]}" | grep -c "FAILED")
echo "[run_all] ${FAILED_COUNT} persona(s) FAILED; suite run is marked as failed." | tee -a "${SUMMARY_LOG}"
exit 1
fi
exit 0

View file

@ -1,301 +0,0 @@
#!/bin/bash
# Run the full pi-bench evaluation for ONE persona with the ReMe agent.
# Tasks follow data/{persona}/episode.yaml order (runner-native).
#
# Usage: bash run_persona.sh <persona> [--fresh|--resume] [--skip-eval]
#
# Modes (default: --resume):
# --resume Checkpoint continuation. Never wipes memory. Tasks already
# finished (SUCCESS/MAX_TURNS/TIMEOUT in the task history logs)
# are skipped via repeated --task-id flags. Before starting, any
# residual memory of tasks that are about to be RE-RUN (partial
# sessions from an interrupted run) is surgically removed by
# resume.py cleanup, so re-runs don't inherit leaked answers.
# --fresh Wipes THIS persona's ReMe memory, outputs and trace logs first,
# then runs all tasks from scratch.
# The two flags are mutually exclusive. A full multi-persona memory wipe is a
# suite-level action of `run_all.sh` (fresh mode), never done here implicitly.
set -uo pipefail
SUITE_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
TRACE_ROOT="${HOME}/.nanobot/trace_logs"
# ─── External dependencies (pi-bench / ReMe are NOT bundled; see README) ──
if [ ! -f "${SUITE_DIR}/env.sh" ]; then
echo "env.sh not found. Run: cp env.sh.example env.sh (then fill in the TODO items)"
exit 1
fi
source "${SUITE_DIR}/env.sh"
PIBENCH_DIR="${PI_BENCH_ROOT:-}"
if [ -z "${PIBENCH_DIR}" ] || [ ! -f "${PIBENCH_DIR}/src/main.py" ]; then
echo "PI_BENCH_ROOT is unset or invalid (src/main.py not found). Set it in env.sh."
exit 1
fi
if [ ! -x "${PIBENCH_DIR}/.venv/bin/python" ] || [ ! -x "${PIBENCH_DIR}/.venv/bin/appworld" ]; then
echo "pi-bench venv incomplete: ${PIBENCH_DIR}/.venv must provide python + appworld (see README setup)."
exit 1
fi
if [ ! -x "${REME_DIR}/.venv/bin/python" ]; then
echo "ReMe venv not found: ${REME_DIR}/.venv/bin/python (check REME_DIR in env.sh)"
exit 1
fi
if [ ! -e "${SUITE_DIR}/data" ]; then
echo 'Benchmark data not linked. Run: ln -s "$PI_BENCH_ROOT/data" data'
exit 1
fi
# ─── Pre-flight: files the runner needs before any service starts ─────
MODEL_CONFIG="${SUITE_DIR}/config/models/reme.yaml"
HISTORY_CONFIG="${SUITE_DIR}/config/bench/evaluation/trace_history.yaml"
if [ ! -f "${MODEL_CONFIG}" ]; then
echo "Model config not found: ${MODEL_CONFIG} (see README directory layout)."
exit 1
fi
if [ ! -f "${HISTORY_CONFIG}" ]; then
echo "Trace history config not found: ${HISTORY_CONFIG}"
echo "pi-bench requires config/bench/evaluation/trace_history.yaml; see README."
exit 1
fi
APPWORLD_DIR="${PIBENCH_DIR}/third_party/appworld"
PI_PYTHON="${PIBENCH_DIR}/.venv/bin/python"
APPWORLD_BIN="${PIBENCH_DIR}/.venv/bin/appworld"
# resume.py runs on the ReMe venv so it can reuse ReMe's daily-index rebuild.
REME_PYTHON="${REME_DIR}/.venv/bin/python"
PERSONA="${1:-}"
if [ -z "$PERSONA" ]; then
echo "Usage: $0 <persona> [--fresh|--resume] [--skip-eval]"
exit 1
fi
shift
MODE="resume"
SKIP_EVAL=false
while [[ $# -gt 0 ]]; do
case $1 in
--fresh)
if [ "$MODE" = "resume_set" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
MODE="fresh"; shift ;;
--resume)
if [ "$MODE" = "fresh" ]; then echo "--fresh and --resume are mutually exclusive"; exit 2; fi
MODE="resume_set"; shift ;;
--skip-eval) SKIP_EVAL=true; shift ;;
*) echo "Unknown option: $1"; exit 1 ;;
esac
done
[ "$MODE" = "resume_set" ] && MODE="resume"
# ─── Per-persona ports (pi-bench AGENTS.md convention) ────────────────
# REME_PORT: ReMe's internal HTTP service; must be unique per concurrent bridge.
case "$PERSONA" in
marketer) API_PORT=9001; MCP_PORT=10001; TEST_PORT=9998; REME_PORT=18766 ;;
law_trainee) API_PORT=9002; MCP_PORT=10002; TEST_PORT=9997; REME_PORT=18767 ;;
pharmacist) API_PORT=9003; MCP_PORT=10003; TEST_PORT=9996; REME_PORT=18768 ;;
researcher) API_PORT=9004; MCP_PORT=10004; TEST_PORT=9995; REME_PORT=18765 ;;
Financier) API_PORT=9005; MCP_PORT=10005; TEST_PORT=9994; REME_PORT=18769 ;;
*) echo "Unknown persona: $PERSONA"; exit 1 ;;
esac
API_URL="http://127.0.0.1:${API_PORT}"
MCP_URL="http://127.0.0.1:${MCP_PORT}/mcp"
TEST_URL="http://127.0.0.1:${TEST_PORT}"
LOG_DIR="${SUITE_DIR}/logs"
mkdir -p "${LOG_DIR}"
# ─── Environment (env.sh already sourced at the top) ──────────────────
WORKSPACE_DIR="${REME_WORKSPACE_ROOT}/${PERSONA}"
NANOBOT_WORKSPACE_DIR="${SUITE_DIR}/nanobot_workspace/${PERSONA}"
mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
echo "========================================="
echo "ReMe x Pi-Bench | persona=${PERSONA} | mode=${MODE}"
echo " api=${API_PORT} mcp=${MCP_PORT} test=${TEST_PORT} reme=${REME_PORT}"
echo " model=${REME_MODEL_NAME}"
echo " memory workspace=${WORKSPACE_DIR} (persistent)"
echo "========================================="
# ─── Fresh mode: wipe this persona's state ────────────────────────────
if [ "$MODE" = "fresh" ]; then
echo "[fresh] wiping persona state: memory workspace, outputs, trace logs"
rm -rf "${WORKSPACE_DIR}"
rm -rf "${SUITE_DIR}/outputs/reme/${PERSONA}"
rm -rf "${TRACE_ROOT}/reme/${PERSONA}"
rm -rf "${NANOBOT_WORKSPACE_DIR}"
mkdir -p "${WORKSPACE_DIR}" "${NANOBOT_WORKSPACE_DIR}"
fi
# ─── Resume: determine remaining tasks + clean partial memories ───────
TASK_ARGS=()
RUN_PHASE_NEEDED=true
if [ "$MODE" = "resume" ]; then
REMAINING_JSON="$("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" --json)"
if [ -z "$REMAINING_JSON" ]; then
echo "Failed to compute remaining tasks"; exit 1
fi
echo "[resume] ${REMAINING_JSON}"
REMAINING_TASKS=()
while IFS= read -r tid_line; do
[ -n "$tid_line" ] && REMAINING_TASKS+=("$tid_line")
done < <("${REME_PYTHON}" "${SUITE_DIR}/resume.py" remaining "${PERSONA}" 2>/dev/null)
if [ ${#REMAINING_TASKS[@]} -eq 0 ]; then
RUN_PHASE_NEEDED=false
echo "[resume] all tasks already completed; skipping run phase"
else
# Remove residual memory of interrupted (to-be-re-run) tasks so
# re-runs don't get their own partial answers injected.
"${REME_PYTHON}" "${SUITE_DIR}/resume.py" cleanup "${PERSONA}"
for tid in "${REMAINING_TASKS[@]}"; do
TASK_ARGS+=(--task-id "$tid")
done
echo "[resume] running ${#REMAINING_TASKS[@]} remaining task(s): ${REMAINING_TASKS[*]}"
fi
fi
# ─── Port cleanup from previous runs ──────────────────────────────────
for port in ${API_PORT} ${MCP_PORT} ${TEST_PORT} ${REME_PORT}; do
pids=$(lsof -ti :${port} 2>/dev/null || true)
if [ -n "$pids" ]; then
echo "Killing stale processes on port ${port}: ${pids}"
kill -9 $pids 2>/dev/null || true
fi
done
sleep 2
PIDS=()
cleanup() {
echo "[${PERSONA}] cleaning up services..."
for pid in "${PIDS[@]:-}"; do
kill "$pid" 2>/dev/null || true
done
wait 2>/dev/null || true
}
trap cleanup EXIT INT TERM
wait_for_service() {
local url="$1" name="$2" port="$3" timeout="${4:-180}"
echo -n " waiting for ${name}..."
local start=$(date +%s)
while true; do
if curl -sf --max-time 5 "${url}" > /dev/null 2>&1; then
echo " ready"; return 0
fi
if [ -n "$port" ] && lsof -ti :${port} > /dev/null 2>&1; then
local elapsed=$(( $(date +%s) - start ))
if [ "$elapsed" -ge 10 ]; then echo " ready (port)"; return 0; fi
fi
if [ $(( $(date +%s) - start )) -ge "$timeout" ]; then
echo " TIMEOUT"; return 1
fi
sleep 2
done
}
# ─── [1/5] AppWorld API ────────────────────────────────────────────────
echo "[1/5] AppWorld API (:${API_PORT})"
(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve apis --root . \
--port ${API_PORT}) > "${LOG_DIR}/appworld_api_${PERSONA}.log" 2>&1 &
PIDS+=($!)
if ! wait_for_service "${API_URL}/docs" "AppWorld API" "${API_PORT}" 180; then
tail -20 "${LOG_DIR}/appworld_api_${PERSONA}.log"; exit 1
fi
# ─── [2/5] AppWorld MCP ────────────────────────────────────────────────
echo "[2/5] AppWorld MCP (:${MCP_PORT})"
TOOLS_CONFIG="${SUITE_DIR}/data/${PERSONA}/tools.yaml"
(cd "${APPWORLD_DIR}" && exec "${APPWORLD_BIN}" serve mcp http --root . \
--remote-apis-url "${API_URL}" --port ${MCP_PORT} \
--tools-config-file "${TOOLS_CONFIG}") > "${LOG_DIR}/appworld_mcp_${PERSONA}.log" 2>&1 &
PIDS+=($!)
if ! wait_for_service "${MCP_URL}" "AppWorld MCP" "${MCP_PORT}" 180; then
tail -20 "${LOG_DIR}/appworld_mcp_${PERSONA}.log"; exit 1
fi
# ─── [3/5] Test Server ─────────────────────────────────────────────────
echo "[3/5] Test Server (:${TEST_PORT})"
PORT=${TEST_PORT} "${PI_PYTHON}" "${PIBENCH_DIR}/scripts/test_server.py" \
> "${LOG_DIR}/test_server_${PERSONA}.log" 2>&1 &
PIDS+=($!)
if ! wait_for_service "${TEST_URL}/sent?after=-1" "Test Server" "${TEST_PORT}" 30; then
tail -20 "${LOG_DIR}/test_server_${PERSONA}.log"; exit 1
fi
# ─── [4/5] ReMe Bridge (ReMe venv) ─────────────────────────────────────
echo "[4/5] ReMe Bridge (reme service port ${REME_PORT})"
"${REME_DIR}/.venv/bin/python" "${SUITE_DIR}/bridge_reme.py" \
--test-server-url "${TEST_URL}" \
--appworld-mcp-url "${MCP_URL}" \
--reme-dir "${REME_DIR}" \
--data-root "${SUITE_DIR}/data" \
--user-id "${PERSONA}" \
--workspace-dir "${WORKSPACE_DIR}" \
--reme-port "${REME_PORT}" \
--model-name "${REME_MODEL_NAME}" \
--model-base-url "${REME_LLM_BASE_URL}" \
--model-api-key "${REME_LLM_API_KEY}" \
> "${LOG_DIR}/bridge_${PERSONA}.log" 2>&1 &
BRIDGE_PID=$!
PIDS+=(${BRIDGE_PID})
sleep 5
if ! kill -0 "${BRIDGE_PID}" 2>/dev/null; then
echo "Bridge failed to start:"; tail -30 "${LOG_DIR}/bridge_${PERSONA}.log"; exit 1
fi
for i in $(seq 1 12); do
if grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null; then
echo " bridge initialized"; break
fi
sleep 5
done
grep -q "Bridge started:" "${LOG_DIR}/bridge_${PERSONA}.log" 2>/dev/null || {
echo "WARNING: bridge may not be ready:"; tail -20 "${LOG_DIR}/bridge_${PERSONA}.log"; }
# ─── [5/5] Runner (run phase) ──────────────────────────────────────────
if [ "$RUN_PHASE_NEEDED" = true ]; then
echo "[5/5] Runner: run phase (episode order from data/${PERSONA}/episode.yaml)"
cd "${SUITE_DIR}"
BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
"${PI_PYTHON}" -m src.main \
--model-config "${MODEL_CONFIG}" \
--history-config-path "${HISTORY_CONFIG}" \
--mode run --user-id "${PERSONA}" \
--workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
${TASK_ARGS[@]+"${TASK_ARGS[@]}"} \
2>&1 | tee "${LOG_DIR}/runner_run_${PERSONA}.log"
RUN_EXIT=${PIPESTATUS[0]}
if [ ${RUN_EXIT} -ne 0 ]; then
echo "Run phase failed (exit ${RUN_EXIT}). Logs: ${LOG_DIR}/"
exit ${RUN_EXIT}
fi
else
echo "[5/5] Runner: run phase skipped (all tasks completed)"
fi
if [ "$SKIP_EVAL" = true ]; then
echo "Skipping eval (--skip-eval)"
exit 0
fi
# ─── Trace conversion + eval phase (always over all available traces) ──
echo "Converting trace logs..."
"${PI_PYTHON}" "${SUITE_DIR}/fix_trace_logs.py" "${PERSONA}"
echo "Runner: eval phase"
cd "${SUITE_DIR}"
BENCH_TEST_SERVER_URL="${TEST_URL}" PYTHONPATH="${PIBENCH_DIR}" \
"${PI_PYTHON}" -m src.main \
--model-config "${MODEL_CONFIG}" \
--history-config-path "${HISTORY_CONFIG}" \
--mode eval --user-id "${PERSONA}" \
--workspace-dir "${NANOBOT_WORKSPACE_DIR}" \
2>&1 | tee "${LOG_DIR}/runner_eval_${PERSONA}.log"
EVAL_EXIT=${PIPESTATUS[0]}
echo ""
echo "========================================="
echo "persona=${PERSONA} finished (eval exit=${EVAL_EXIT})"
echo " results : ${SUITE_DIR}/outputs/reme/${PERSONA}/"
echo " memory : ${WORKSPACE_DIR}/"
echo " logs : ${LOG_DIR}/"
echo "========================================="
exit ${EVAL_EXIT}

View file

@ -1,98 +0,0 @@
## Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
**Language**: English (default) / [中文](./README_ZH.md)
> Paper: [arXiv:2608.03403](https://arxiv.org/abs/2608.03403)
> Code: [https://github.com/WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)
<p align="center">
<img src="gitcha.png" alt="ExpG challenges and overview" width="85%">
</p>
### Overview
This folder archives **ExpG**, a tool-use enhancement built on [Agentscope ReMe](https://github.com/agentscope-ai/ReMe). ExpG mines, distills, and reuses experience from historical tool calls to provide **capability boundaries** and **best-practice guidance**, which helps agents:
- Select and invoke tools more robustly under dynamic or noisy environments;
- Let smaller models with guidance outperform larger, memoryless baselines;
- Improve consistently across tool selection, tool calling, and response generation.
**How ReMe is used:** Start the Tool Memory service; historical tool calls are written and evaluated via `add_tool_call_result`, distilled into tool-level guidance via `summary_tool_memory`, then retrieved and injected into later reasoning via `retrieve_tool_memory`. ReMe provides the vector store and service APIs; the acquisition / distillation / reuse strategy is implemented by ExpG. Full implementation and experiments are in [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG).
---
### ExpG Mechanism
ExpG treats tool invocations as learnable experience and runs a three-stage pipeline:
1. **Experience Acquisition**
- Analyze invocation quality from historical trajectories (success/failure, cost, latency, etc.);
- Build structured experience units per tool, recording context, parameter patterns, and outcomes.
2. **Experience Distillation**
- Filter noisy or unhelpful experiences and keep representative patterns;
- Aggregate by equivalence classes to cover common and rare failure modes;
- Summarize with an LLM into generalizable textual guidance.
3. **Experience Reuse**
- Retrieve relevant experience / guidance for future tasks;
- Inject guidance into tool selection, argument generation, and response synthesis;
- Improve stability under dynamic environments and imperfect feedback.
---
### Main Results
Performance comparison (%) across MetaTool, API-Bank, and BFCL-V3. **Bold** indicates the best results within each model.
| Model | Method | MetaTool Pass@1 | MetaTool Avg@3 | MetaTool Pass@3 | API-Bank Pass@1 | API-Bank Avg@3 | API-Bank Pass@3 | BFCL-V3 Pass@1 | BFCL-V3 Avg@3 | BFCL-V3 Pass@3 | Total Pass@1 | Total Avg@3 | Total Pass@3 |
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| GPT-5 nano | No Method | 72.62 | 72.76 | 78.49 | 82.96 | 83.46 | 86.97 | 53.80 | 53.00 | 60.95 | 70.82 | 70.62 | 76.63 |
| GPT-5 nano | Few-shot | 74.12 | 75.11 | 82.32 | 83.71 | 83.96 | **87.22** | 56.18 | 55.24 | 61.39 | 72.36 | 72.65 | 79.28 |
| GPT-5 nano | DRAFT | 73.94 | 73.04 | 78.97 | 84.21 | 83.46 | **87.22** | 57.27 | 57.27 | 62.26 | 72.52 | 71.58 | 77.23 |
| GPT-5 nano | Mem0 | 74.96 | 76.13 | 82.92 | 84.96 | 85.21 | **87.22** | 60.95 | 61.61 | 65.08 | 73.98 | 74.67 | 80.35 |
| GPT-5 nano | **ExpG** | **81.67** | **82.07** | **84.60** | **86.72** | **86.55** | **87.22** | **64.43** | **63.99** | **66.38** | **79.32** | **79.22** | **81.69** |
| DeepSeek-V3 | No Method | 83.10 | 82.94 | 84.66 | 84.71 | 84.38 | 85.46 | 58.79 | 59.65 | 65.94 | 78.92 | 78.66 | 81.37 |
| DeepSeek-V3 | Few-shot | 82.74 | 83.90 | 86.28 | 85.21 | 84.63 | 86.22 | 60.52 | 60.30 | 67.90 | 79.08 | 79.45 | 82.92 |
| DeepSeek-V3 | DRAFT | 80.23 | 80.79 | 82.44 | 84.96 | 85.63 | 86.47 | 62.26 | 61.61 | 68.55 | 77.70 | 77.80 | 80.54 |
| DeepSeek-V3 | Mem0 | 83.88 | 84.56 | 86.40 | 85.46 | 85.55 | 86.47 | 65.08 | 65.15 | 68.33 | 80.70 | 80.91 | 83.12 |
| DeepSeek-V3 | **ExpG** | **85.26** | **85.38** | **86.52** | **87.72** | **87.39** | **87.97** | **69.41** | **69.92** | **72.02** | **82.76** | **82.61** | **84.11** |
| Qwen3-8B | No Method | 76.51 | 76.97 | 77.71 | 83.96 | 83.88 | 84.21 | 58.79 | 58.28 | 60.30 | 74.46 | 74.41 | 75.56 |
| Qwen3-8B | Few-shot | 79.93 | 79.83 | 82.92 | 83.71 | 82.62 | 84.96 | 60.09 | 59.29 | 61.39 | 76.91 | 76.27 | 79.32 |
| Qwen3-8B | DRAFT | 78.19 | 77.33 | 77.89 | 85.71 | 84.96 | 85.46 | 60.74 | 60.30 | 62.91 | 76.20 | 75.18 | 76.35 |
| Qwen3-8B | Mem0 | 75.07 | 75.47 | 82.38 | 86.22 | 86.05 | 86.47 | 63.34 | 64.93 | 66.16 | 74.69 | 74.98 | 80.07 |
| Qwen3-8B | **ExpG** | **83.52** | **84.88** | **85.08** | **86.47** | **87.89** | **87.97** | **67.46** | **66.96** | **68.33** | **81.06** | **81.82** | **82.48** |
| Qwen3-32B | No Method | 80.05 | 79.43 | 80.17 | 84.71 | 84.88 | 85.21 | 65.15 | 65.08 | 66.16 | 78.05 | 77.55 | 78.41 |
| Qwen3-32B | **ExpG** | **84.68** | **85.02** | **86.28** | **86.97** | **87.30** | **87.72** | **70.72** | **71.01** | **73.32** | **82.48** | **82.56** | **84.14** |
| Qwen3-235B | No Method | 78.25 | 79.23 | 80.29 | 85.46 | 85.46 | 85.71 | 71.37 | 71.15 | 73.54 | 78.13 | 78.49 | 79.91 |
| Qwen3-235B | **ExpG** | **86.34** | **86.70** | **86.94** | **87.47** | **86.97** | **88.22** | **79.61** | **78.52** | **80.04** | **85.29** | **84.98** | **85.69** |
---
### Reference Code
| Path | Role |
| --- | --- |
| [`tool_memory.py`](./tool_memory.py) | HTTP client for official ReMe Tool Memory APIs (`add_tool_call_result` / `summary_tool_memory` / `retrieve_tool_memory`) |
| [`parse_tool_call_result_prompt.yaml`](./parse_tool_call_result_prompt.yaml) | Prompt for multi-aspect evaluation of each tool call |
| [`summary_tool_memory_prompt.yaml`](./summary_tool_memory_prompt.yaml) | Prompt for summarizing tool call history into guidance |
| [`tool_memory_flows.yaml`](./tool_memory_flows.yaml) | Tool Memory flow / op config excerpt |
These are reference snippets. For the full runnable codebase, see [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG).
---
### Citation
```bibtex
@misc{wang2026expg,
title = {Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance},
author = {Can Wang and Haoran Chen and Li Yu and Ding Hao and Bohai Zhao and Zhaoyang Liu and Zhiying Tu},
year = {2026},
eprint = {2608.03403},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.03403},
howpublished = {\url{https://github.com/WangCan1178/ExpG}}
}
```

View file

@ -1,98 +0,0 @@
## Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
**语言**:中文 / [English](./README.md)
> 论文:[arXiv:2608.03403](https://arxiv.org/abs/2608.03403)
> 代码:[https://github.com/WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)
<p align="center">
<img src="gitcha.png" alt="ExpG 挑战与概览" width="85%">
</p>
### 简介
本目录归档基于 [Agentscope ReMe](https://github.com/agentscope-ai/ReMe) 的工具使用增强工作 **ExpG**:在 ReMe 记忆框架之上,从历史工具调用中挖掘、提炼并复用经验,为智能体提供工具的 **能力边界****最佳实践指导**,从而:
- 在动态或有噪环境下更鲁棒地选择和调用工具;
- 让较小模型在带有经验指导时超越更大、但无记忆的基线;
- 在工具选择、工具调用和响应生成等多个阶段带来一致收益。
**如何使用 ReMe** 启动 Tool Memory 服务后,历史工具调用经 `add_tool_call_result` 写入并评估,经 `summary_tool_memory` 蒸馏成工具级指导,再经 `retrieve_tool_memory` 取回并注入后续推理。向量存储与服务接口由 ReMe 提供,经验获取 / 蒸馏 / 复用策略由 ExpG 实现。完整实现与实验见 [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)。
---
### ExpG 机制概览
ExpG 将工具调用视为可学习经验,并通过三阶段流水线完成经验的获取、提炼与复用:
1. **经验获取Experience Acquisition**
- 从历史工具调用轨迹中分析调用质量(成功/失败、代价、时间等);
- 针对不同工具构建结构化的经验单元,记录调用上下文、参数模式和结果。
2. **经验蒸馏Experience Distillation**
- 过滤无效 / 噪声经验,保留具有代表性的调用模式;
- 基于“等价类”视角对经验进行聚合,覆盖常见模式与稀有失败模式;
- 使用 LLM 对经验进行总结形成可泛化的文本化指导guidance
3. **经验复用Experience Reuse**
- 在未来任务中,根据当前工具调用上下文检索相关经验 / 指导;
- 将经验引导融入到工具选择、参数生成和响应整理等环节;
- 使得代理在面对动态环境和不完美反馈时仍能保持稳定表现。
---
### 主实验结果
MetaTool、API-Bank、BFCL-V3 上的性能对比(%)。**加粗**为各模型组内最优。
| Model | Method | MetaTool Pass@1 | MetaTool Avg@3 | MetaTool Pass@3 | API-Bank Pass@1 | API-Bank Avg@3 | API-Bank Pass@3 | BFCL-V3 Pass@1 | BFCL-V3 Avg@3 | BFCL-V3 Pass@3 | Total Pass@1 | Total Avg@3 | Total Pass@3 |
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| GPT-5 nano | No Method | 72.62 | 72.76 | 78.49 | 82.96 | 83.46 | 86.97 | 53.80 | 53.00 | 60.95 | 70.82 | 70.62 | 76.63 |
| GPT-5 nano | Few-shot | 74.12 | 75.11 | 82.32 | 83.71 | 83.96 | **87.22** | 56.18 | 55.24 | 61.39 | 72.36 | 72.65 | 79.28 |
| GPT-5 nano | DRAFT | 73.94 | 73.04 | 78.97 | 84.21 | 83.46 | **87.22** | 57.27 | 57.27 | 62.26 | 72.52 | 71.58 | 77.23 |
| GPT-5 nano | Mem0 | 74.96 | 76.13 | 82.92 | 84.96 | 85.21 | **87.22** | 60.95 | 61.61 | 65.08 | 73.98 | 74.67 | 80.35 |
| GPT-5 nano | **ExpG** | **81.67** | **82.07** | **84.60** | **86.72** | **86.55** | **87.22** | **64.43** | **63.99** | **66.38** | **79.32** | **79.22** | **81.69** |
| DeepSeek-V3 | No Method | 83.10 | 82.94 | 84.66 | 84.71 | 84.38 | 85.46 | 58.79 | 59.65 | 65.94 | 78.92 | 78.66 | 81.37 |
| DeepSeek-V3 | Few-shot | 82.74 | 83.90 | 86.28 | 85.21 | 84.63 | 86.22 | 60.52 | 60.30 | 67.90 | 79.08 | 79.45 | 82.92 |
| DeepSeek-V3 | DRAFT | 80.23 | 80.79 | 82.44 | 84.96 | 85.63 | 86.47 | 62.26 | 61.61 | 68.55 | 77.70 | 77.80 | 80.54 |
| DeepSeek-V3 | Mem0 | 83.88 | 84.56 | 86.40 | 85.46 | 85.55 | 86.47 | 65.08 | 65.15 | 68.33 | 80.70 | 80.91 | 83.12 |
| DeepSeek-V3 | **ExpG** | **85.26** | **85.38** | **86.52** | **87.72** | **87.39** | **87.97** | **69.41** | **69.92** | **72.02** | **82.76** | **82.61** | **84.11** |
| Qwen3-8B | No Method | 76.51 | 76.97 | 77.71 | 83.96 | 83.88 | 84.21 | 58.79 | 58.28 | 60.30 | 74.46 | 74.41 | 75.56 |
| Qwen3-8B | Few-shot | 79.93 | 79.83 | 82.92 | 83.71 | 82.62 | 84.96 | 60.09 | 59.29 | 61.39 | 76.91 | 76.27 | 79.32 |
| Qwen3-8B | DRAFT | 78.19 | 77.33 | 77.89 | 85.71 | 84.96 | 85.46 | 60.74 | 60.30 | 62.91 | 76.20 | 75.18 | 76.35 |
| Qwen3-8B | Mem0 | 75.07 | 75.47 | 82.38 | 86.22 | 86.05 | 86.47 | 63.34 | 64.93 | 66.16 | 74.69 | 74.98 | 80.07 |
| Qwen3-8B | **ExpG** | **83.52** | **84.88** | **85.08** | **86.47** | **87.89** | **87.97** | **67.46** | **66.96** | **68.33** | **81.06** | **81.82** | **82.48** |
| Qwen3-32B | No Method | 80.05 | 79.43 | 80.17 | 84.71 | 84.88 | 85.21 | 65.15 | 65.08 | 66.16 | 78.05 | 77.55 | 78.41 |
| Qwen3-32B | **ExpG** | **84.68** | **85.02** | **86.28** | **86.97** | **87.30** | **87.72** | **70.72** | **71.01** | **73.32** | **82.48** | **82.56** | **84.14** |
| Qwen3-235B | No Method | 78.25 | 79.23 | 80.29 | 85.46 | 85.46 | 85.71 | 71.37 | 71.15 | 73.54 | 78.13 | 78.49 | 79.91 |
| Qwen3-235B | **ExpG** | **86.34** | **86.70** | **86.94** | **87.47** | **86.97** | **88.22** | **79.61** | **78.52** | **80.04** | **85.29** | **84.98** | **85.69** |
---
### 参考代码
| 路径 | 作用 |
| --- | --- |
| [`tool_memory.py`](./tool_memory.py) | 官方风格 ReMe Tool Memory HTTP 客户端(`add_tool_call_result` / `summary_tool_memory` / `retrieve_tool_memory` |
| [`parse_tool_call_result_prompt.yaml`](./parse_tool_call_result_prompt.yaml) | 单次工具调用多维评估用的 prompt |
| [`summary_tool_memory_prompt.yaml`](./summary_tool_memory_prompt.yaml) | 将工具调用历史总结为 guidance 的 prompt |
| [`tool_memory_flows.yaml`](./tool_memory_flows.yaml) | Tool Memory 相关的 flow / op 配置摘录 |
以上为参考片段。完整可运行代码见 [WangCan1178/ExpG](https://github.com/WangCan1178/ExpG)。
---
### 引用
```bibtex
@misc{wang2026expg,
title = {Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance},
author = {Can Wang and Haoran Chen and Li Yu and Ding Hao and Bohai Zhao and Zhaoyang Liu and Zhiying Tu},
year = {2026},
eprint = {2608.03403},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.03403},
howpublished = {\url{https://github.com/WangCan1178/ExpG}}
}
```

Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.9 MiB

View file

@ -1,49 +0,0 @@
prompt: |
You are an expert in evaluating tool invocation process. The tool is invoked by an AI agent.
Tool invocation Information:
- Tool Name: {tool_name}
- Success Flag: {success_flag}
- Time Cost: {time_cost}s
- Token Cost: {token_cost} tokens
- Agent Context: {context}
- Input Parameters: {input_params}
- Tool Response: {response}
- Tool Schema: {schema}
Evaluation Method:
Start from a default score list of scores = [0, 0, 0, 0, 0, 0, 0, 0, 0, 0].
For each item below that is satisfied, assign 1 point to the corresponding index.
The final scores should be a list of 10 integers, each being either 0 or 1.
1. Use Quality (total 2 points. If context is provided, use it as an aid when evaluating):
- Index 1: Should the tool be invoked now? Consider whether all necessary information for the tool's invocation is ready, and whether the tool execution environment is correct. If it is a multi-round conversation, also consider the dependency relationships of the tool chain.
- Index 2: If should, is the chosen tool appropriate?
2. Input Quality (total 4 points. When evaluating, consider both the context and the tool schema):
- Index 3: Are all required parameters provided?
- Index 4: Are the input parameters valid and supported by the tool?
- Index 5: Are the input parameters in the correct format for their respective fields?
- Index 6: Does the value (content) of input parameter correctly reflect and match the given context?
3. Response Quality (total 4 points):
- Index 7: Does the response provide meaningful and useful information? Or are there any error messages or information that can be used as guidance for agent invoking tool better?
- Index 8: Does the response match the tool's intended purpose/function?
- Index 9: Does the response value correct (content appropriate) given the input parameters?
- Index 10: Does the response help accomplish the task within the given context?
Important:
1. Sometimes there is not enough information in the context or schema to make a complete evaluation. In such cases, make your best judgment based on the available information.
2. Some tools (commonly system tools such as mkdir, touch, echo, etc.) modify the external environment. Since these results cannot be obtained, they return "None" as the response. At this point, all the scores in the quality of the response should be obtained and should not be seen as a problem for the tool.
3. Evaluation independently from the success flag. The success_flag indicates whether the tool executed without technical errors. The evaluation should evaluate the quality of the tool invocation. A tool can execute successfully (Success Flag=1) but still produce low-quality or irrelevant responses, leading to a low evaluation score.
4. Sometimes an agent will execute multiple steps and invoke multiple tools to complete a task, but you only need to evaluate the use of one tool for one of the steps, not whether the final task is completed or not.
Answer Format:
Please provide your answer in the following JSON format:
```json
{
"scores": [0,0,0,0,0,0,0,0,0,0],
"explanation": "A brief evaluation (2-3 sentences) explaining the quality of the tool invocation, based on your evaluation. Low-quality aspects need to be reified, especially the causes of tool invocation errors."
}
```

View file

@ -1,32 +0,0 @@
prompt: |
You are an expert in analyzing tool usage patterns and generating practical usage guidance for agents.
Tool Information:
- Tool Name: {tool_name}
- Tool Schema: {tool_schema}
Recent Tool Invocation Experiences:
{experiences}
Important:
1. Assume the tool (tool schema) can't be changed, your task is to guide agent to use it better.
2. Your answer must be based on the information given, don't make it up. If not enough data, state "Not enough data to determine Core Function/Success Patterns/Common Issues/Best Practices."
3. Your answer will be used to guide the use of the tool in the future, so do not include content related to recent tool invocation experience such as "case #3" or "Call #2", but some values can be used as examples.
4. Pay attention to information not mentioned in the tool schema, such as the response upon successful tool invocation. It's also welcome to uncover insights, such as how tools can be used more effectively, and possible dependencies between tools. But if they aren't, don't make them up.
5. Finally, to avoid deriving incorrect guidance from individual invocation, check whether, if the agent follows the proposed guidance, it can perform better on all recent invocation histories. If not, revise the guidance until it can. Specifically:
- Don't write guidance in an absolute tone without a very deterministic message (meaning that all invocation histories are satisfied, otherwise it will result in failure).
- Sometimes there may be inconsistencies. Consider whether this is due to the context in which the tool is being used.
Your Task:
Based on the tool invocation history, generate a concise and logical tool usage guidance following this structure:
1. Core Function: What this tool does and when to use it.
2. Success Patterns: Parameter patterns and usage scenarios that work well.
3. Common Issues: Main pitfalls to avoid and why they fail.
4. Best Practices: 2-3 actionable recommendations.
Answer Format:
Provide a structured, concise guidance (max 200 words). Focus on actionable insights derived from actual usage data. Avoid generic advice and think step by step.
```txt
Your concise, data-driven tool usage guidance
```

View file

@ -1,234 +0,0 @@
"""Official-style ReMe Tool Memory HTTP helpers.
Aligned with ReMe Tool Memory HTTP APIs (see ReMe cookbook
``use_tool_memory_demo.py`` and docs under ``docs/tool_memory/``):
- ``add_tool_call_result``
- ``summary_tool_memory``
- ``retrieve_tool_memory``
Response memories are read from ``metadata.memory_list[].content``.
This module does not use ExpG-only fields such as ``no_persist``,
``source_task``, or ``add_to``.
"""
from __future__ import annotations
import logging
from typing import Any, Dict, List, Optional
import httpx
logger = logging.getLogger(__name__)
DEFAULT_BASE_URL = "http://localhost:8002"
class ToolMemoryFetcher:
"""HTTP client for ReMe Tool Memory endpoints."""
def __init__(
self,
workspace_id: str,
base_url: str = DEFAULT_BASE_URL,
timeout: float = 60.0,
) -> None:
self.workspace_id = workspace_id
self.base_url = base_url.rstrip("/")
self.timeout = timeout
def _url(self, endpoint: str) -> str:
return f"{self.base_url}/{endpoint.lstrip('/')}"
@staticmethod
def _join_tool_names(tool_names: List[str] | str) -> str:
if isinstance(tool_names, str):
return tool_names
return ",".join(tool_names)
@staticmethod
def _memory_list(payload: Dict[str, Any]) -> List[Dict[str, Any]]:
metadata = payload.get("metadata") or {}
if not isinstance(metadata, dict):
return []
memory_list = metadata.get("memory_list") or []
return memory_list if isinstance(memory_list, list) else []
@classmethod
def _content_by_tool(cls, payload: Dict[str, Any]) -> Dict[str, str]:
result: Dict[str, str] = {}
for memory in cls._memory_list(payload):
if not isinstance(memory, dict):
continue
tool_name = str(memory.get("when_to_use") or "").strip()
content = memory.get("content") or ""
if tool_name:
result[tool_name] = str(content)
return result
async def add_tool_call_result_async(
self,
tool_call_results: List[Dict[str, Any]],
) -> Dict[str, Any]:
"""Call ``add_tool_call_result``."""
async with httpx.AsyncClient() as client:
response = await client.post(
self._url("add_tool_call_result"),
json={
"workspace_id": self.workspace_id,
"tool_call_results": tool_call_results,
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
async def summary_tool_memory_async(
self,
tool_names: List[str] | str,
) -> Dict[str, Any]:
"""Call ``summary_tool_memory``."""
async with httpx.AsyncClient() as client:
response = await client.post(
self._url("summary_tool_memory"),
json={
"workspace_id": self.workspace_id,
"tool_names": self._join_tool_names(tool_names),
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
async def retrieve_tool_memory_async(
self,
tool_names: List[str] | str,
) -> Dict[str, Any]:
"""Call ``retrieve_tool_memory``."""
async with httpx.AsyncClient() as client:
response = await client.post(
self._url("retrieve_tool_memory"),
json={
"workspace_id": self.workspace_id,
"tool_names": self._join_tool_names(tool_names),
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
async def collect_memory_async(
self,
tool_names: List[str],
) -> Dict[str, str]:
"""Summarize then retrieve guidance for tools.
Returns:
Mapping from tool name to memory ``content`` string.
"""
if not tool_names:
return {}
names = self._join_tool_names(tool_names)
try:
summary = await self.summary_tool_memory_async(names)
if not summary.get("success"):
logger.warning("summary_tool_memory failed for %s", names)
except Exception as exc: # noqa: BLE001
logger.warning("summary_tool_memory error for %s: %s", names, exc)
try:
retrieved = await self.retrieve_tool_memory_async(names)
except Exception as exc: # noqa: BLE001
logger.warning("retrieve_tool_memory error for %s: %s", names, exc)
return {}
if not retrieved.get("success"):
logger.warning("retrieve_tool_memory failed for %s", names)
return {}
return self._content_by_tool(retrieved)
def add_tool_call_result(
self,
tool_call_results: List[Dict[str, Any]],
) -> Dict[str, Any]:
"""Sync wrapper for ``add_tool_call_result``."""
with httpx.Client() as client:
response = client.post(
self._url("add_tool_call_result"),
json={
"workspace_id": self.workspace_id,
"tool_call_results": tool_call_results,
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
def summary_tool_memory(self, tool_names: List[str] | str) -> Dict[str, Any]:
"""Sync wrapper for ``summary_tool_memory``."""
with httpx.Client() as client:
response = client.post(
self._url("summary_tool_memory"),
json={
"workspace_id": self.workspace_id,
"tool_names": self._join_tool_names(tool_names),
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
def retrieve_tool_memory(self, tool_names: List[str] | str) -> Dict[str, Any]:
"""Sync wrapper for ``retrieve_tool_memory``."""
with httpx.Client() as client:
response = client.post(
self._url("retrieve_tool_memory"),
json={
"workspace_id": self.workspace_id,
"tool_names": self._join_tool_names(tool_names),
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()
def collect_memory(self, tool_names: List[str]) -> Dict[str, str]:
"""Sync wrapper for summarize + retrieve.
Prefer ``collect_memory_async`` inside an existing event loop.
"""
if not tool_names:
return {}
names = self._join_tool_names(tool_names)
try:
summary = self.summary_tool_memory(names)
if not summary.get("success"):
logger.warning("summary_tool_memory failed for %s", names)
except Exception as exc: # noqa: BLE001
logger.warning("summary_tool_memory error for %s: %s", names, exc)
try:
retrieved = self.retrieve_tool_memory(names)
except Exception as exc: # noqa: BLE001
logger.warning("retrieve_tool_memory error for %s: %s", names, exc)
return {}
if not retrieved.get("success"):
logger.warning("retrieve_tool_memory failed for %s", names)
return {}
return self._content_by_tool(retrieved)
def get_memory_content(
self,
tool_names: List[str] | str,
) -> Optional[str]:
"""Retrieve and join memory contents for the given tools."""
payload = self.retrieve_tool_memory(tool_names)
if not payload.get("success"):
return None
contents = [content for content in self._content_by_tool(payload).values() if content]
return "\n\n".join(contents) if contents else None

View file

@ -1,45 +0,0 @@
# Tool Memory flow / op config excerpt used by ExpG.
# Full runnable code: https://github.com/WangCan1178/ExpG
flow:
retrieve_tool_memory:
flow_content: retrieve_tool_memory_op
description: "Retrieves tool memories from the vector database based on tool names to provide tool usage patterns and best practices"
input_schema:
tool_names:
type: string
description: "Comma-separated tool names (e.g., 'tool_name1,tool_name2')"
required: true
add_tool_call_result:
flow_content: parse_tool_call_result_op >> update_vector_store_op
description: "Evaluates and adds tool call results to the tool memory database, creating new memory or updating existing memory for the specified tool"
input_schema:
tool_call_results:
type: array
description: "List of tool call result objects, each containing: tool_name, input, output, success, time_cost, token_cost, create_time"
required: true
summary_tool_memory:
flow_content: summary_tool_memory_op >> update_vector_store_op
description: "Analyzes tool call history and generates comprehensive usage patterns, best practices, and recommendations for the specified tools"
input_schema:
tool_names:
type: string
description: "Comma-separated tool names to summarize (e.g., 'tool_name1,tool_name2')"
required: true
op:
parse_tool_call_result_op:
backend: parse_tool_call_result_op
llm: default
params:
max_history_tool_call_cnt: 100
evaluation_sleep_interval: 1.0
summary_tool_memory_op:
backend: summary_tool_memory_op
llm: default
params:
data_from: '2025-09-10 10:56:58'
summary_sleep_interval: 1.0

View file

@ -0,0 +1,225 @@
import os
from typing import List
from tqdm import tqdm
os.environ["APPWORLD_ROOT"] = "."
from dotenv import load_dotenv
load_dotenv("../../../.env")
import re
import time
import json
import ray
import requests
from appworld import AppWorld, load_task_ids
from jinja2 import Template
from loguru import logger
from openai import OpenAI
from prompt import PROMPT_TEMPLATE, PROMPT_TEMPLATE_WITH_EXPERIENCE
@ray.remote
class AppworldReactAgent:
"""A minimal ReAct Agent for AppWorld tasks."""
def __init__(self,
index: int,
task_ids: List[str],
experiment_name: str,
model_name: str = "qwen3-8b",
temperature: float = 0.9,
max_interactions: int = 30,
max_response_size: int = 2048,
num_runs: int = 1,
use_task_memory: bool = False,
make_task_memory: bool = False,
api_url: str = "http://0.0.0.0:8002/",
workspace_id: str="appworld_v1"):
self.index: int = index
self.task_ids: List[str] = task_ids
self.experiment_name: str = experiment_name
self.model_name: str = model_name
self.temperature: float = temperature
self.max_interactions: int = max_interactions
self.max_response_size: int = max_response_size
self.num_runs: int = num_runs
self.use_task_memory: bool = use_task_memory
self.make_task_memory: bool = make_task_memory
self.api_url = api_url
self.workspace_id = workspace_id
self.llm_client = OpenAI()
def call_llm(self, messages: list) -> str:
for i in range(100):
try:
response = self.llm_client.chat.completions.create(
model=self.model_name,
messages=messages,
temperature=self.temperature,
extra_body={"enable_thinking": False},
seed=0)
return response.choices[0].message.content
except Exception as e:
logger.exception(f"encounter error with {e.args}")
time.sleep(1 + i * 10)
return "call llm error"
def prompt_messages(self,world: AppWorld) -> list[dict]:
if self.use_task_memory:
task_memory = self.get_task_memory(world.task.instruction)
logger.info(f"loaded task_memory: {task_memory}")
dictionary = {"supervisor": world.task.supervisor, "instruction": world.task.instruction, "experience": task_memory}
else:
dictionary = {"supervisor": world.task.supervisor, "instruction": world.task.instruction ,"experience": ""}
print(dictionary)
prompt = Template(PROMPT_TEMPLATE_WITH_EXPERIENCE.lstrip()).render(dictionary)
messages: list[dict] = []
# last_start = 0
# for match in re.finditer("(USER|ASSISTANT|SYSTEM):\n", prompt):
# last_end = match.span()[0]
# if len(messages) == 0:
# if last_end != 0:
# raise ValueError(
# f"Start of the prompt has no assigned role: {prompt[:last_end]}"
# )
# else:
# messages[-1]["content"] = prompt[last_start:last_end]
# role_type = match.group(1).lower()
# messages.append({"role": role_type, "content": None})
# last_start = match.span()[1]
# messages[-1]["content"] = prompt[last_start:]
messages.append({"role":"user", "content":prompt})
return messages
@staticmethod
def get_reward(world) -> float:
tracker = world.evaluate()
num_passes = len(tracker.passes)
num_failures = len(tracker.failures)
return num_passes / (num_passes + num_failures)
def execute(self):
result = []
for task_index, task_id in enumerate(tqdm(self.task_ids, desc=f"ray_index={self.index}")):
# Run each task num_runs times
for run_id in range(self.num_runs):
with AppWorld(task_id=task_id, experiment_name=f"{self.experiment_name}_run_{run_id}") as world:
history = self.prompt_messages(world=world)
before_score = self.get_reward(world)
for i in range(self.max_interactions):
code = self.call_llm(history)
history.append({"role": "assistant", "content": code})
output = world.execute(code)
if len(output) > self.max_response_size:
# logger.warning(f"output exceed max size={len(output)}")
output = output[:self.max_response_size]
history.append({"role": "user", "content": output})
if world.task_completed():
break
after_score = self.get_reward(world)
uplift_score = after_score - before_score
t_result = {
"task_id": world.task_id,
"run_id": run_id, # Add run_id field
"experiment_name": self.experiment_name,
"task_completed": world.task_completed(),
"before_score": before_score,
"after_score": after_score,
"uplift_score": uplift_score,
"task_history": history,
}
result.append(t_result)
if self.make_task_memory:
memory_list = self.make_task_memory(result)
logger.info(f"Created {len(memory_list) if memory_list else 0} task memories")
return result
def handle_api_response(self, response: requests.Response):
"""Handle API response with proper error checking"""
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return None
return response.json()
def get_task_memory(self, query: str):
"""Retrieve relevant task memories based on a query"""
response = requests.post(
url=f"{self.api_url}retrieve_task_memory",
json={
"workspace_id": self.workspace_id,
"query": query,
}
)
result = self.handle_api_response(response)
if not result:
return ""
# Extract and return the answer
answer = result.get("answer", "")
print(f"Retrieved task memory: {answer}")
return answer
def make_task_memory(self, result):
"""Generate a summary of conversation messages and create task memories"""
if not result:
print("No results to summarize")
return
# Prepare trajectories from results
trajectories = []
for r in result:
if "task_history" in r:
trajectories.append({
"messages": r["task_history"],
"score": float(r.get("uplift_score", 0.0))
})
if not trajectories:
print("No trajectories to summarize")
return
response = requests.post(
url=f"{self.api_url}summary_task_memory",
json={
"workspace_id": self.workspace_id,
"trajectories": trajectories
}
)
result = self.handle_api_response(response)
if not result:
return
# Extract memory list from response
memory_list = result.get("metadata", {}).get("memory_list", [])
print(f"Task memory list created: {len(memory_list)} memories")
return memory_list
def main():
dataset_name = "train"
task_ids = load_task_ids(dataset_name)
agent = AppworldReactAgent(index=0, task_ids=task_ids[0:1], experiment_name=dataset_name, num_runs=4)
result = agent.execute()
logger.info(f"result={json.dumps(result)}")
if __name__ == "__main__":
main()

312
cookbook/appworld/prompt.py Normal file
View file

@ -0,0 +1,312 @@
# This is a basic prompt template containing all the necessary onboarding information to solve AppWorld tasks. It explains the role of the agent and the supervisor, how to explore the API documentation, how to operate the interactive coding environment and call APIs via a simple task, and provides key instructions and disclaimers.
# You can adapt it as needed by your agent. You can also choose to bypass API docs app and build your own API retrieval, e.g., for FullCodeRefl, IPFunCall, etc, we asked an LLM to predict relevant APIs separately and put its documentation directly in the prompt.
PROMPT_TEMPLATE = """
USER:
I am your supervisor and you are a super intelligent AI Assistant whose job is to achieve my day-to-day tasks completely autonomously.
To do this, you will need to interact with app/s (e.g., spotify, venmo, etc) using their associated APIs on my behalf. For this you will undertake a *multi-step conversation* using a python REPL environment. That is, you will write the python code and the environment will execute it and show you the result, based on which, you will write python code for the next step and so on, until you've achieved the goal. This environment will let you interact with app/s using their associated APIs on my behalf.
Here are three key APIs that you need to know to get more information
# To get a list of apps that are available to you.
print(apis.api_docs.show_app_descriptions())
# To get the list of apis under any app listed above, e.g. supervisor
print(apis.api_docs.show_api_descriptions(app_name='supervisor'))
# To get the specification of a particular api, e.g. supervisor app's show_account_passwords
print(apis.api_docs.show_api_doc(app_name='supervisor', api_name='show_account_passwords'))
Each code execution will produce an output that you can use in subsequent calls. Using these APIs, you can now generate code, that the environment will execute, to solve the task.
For example, consider the task:
My name is: {{ supervisor.first_name }} {{ supervisor.last_name }}. My personal email is {{ supervisor.email }} and phone number is {{ supervisor.phone_number }}.
Task:
What is the password for my Spotify account?
ASSISTANT:
# Okay. Lets first find which apps are available to get the password by looking at the app descriptions.
print(apis.api_docs.show_app_descriptions())
USER:
[
{
"name": "api_docs",
"description": "An app to search and explore API documentation."
},
{
"name": "supervisor",
"description": "An app to access supervisor's personal information, account credentials, addresses, payment cards, and manage the assigned task."
},
...
{
"name": "spotify",
"description": "A music streaming app to stream songs and manage song, album and playlist libraries."
},
{
"name": "venmo",
"description": "A social payment app to send, receive and request money to and from others."
},
...
]
ASSISTANT:
# Looks like the supervisor app could help me with that. Lets see what apis are available under this app.
print(apis.api_docs.show_api_descriptions(app_name='supervisor'))
USER:
[
...
"show_account_passwords : Show your supervisor's account passwords."
...
]
ASSISTANT:
# I can use `show_account_passwords` to get the passwords. Let me see its detailed specification to understand its arguments and output structure.
print(apis.api_docs.show_api_doc(app_name='supervisor', api_name='show_account_passwords'))
USER:
{
'app_name': 'supervisor',
'api_name': 'show_account_passwords',
'path': '/account_passwords',
'method': 'GET',
'description': "Show your supervisor's app account passwords.",
'parameters': [],
'response_schemas': {
'success': [{'account_name': 'string', 'password': 'string'}],
'failure': {'message': 'string'}
}
}
ASSISTANT:
# Okay, it requires no arguments. So I can just call it directly.
print(apis.supervisor.show_account_passwords())
USER:
[
{
"account_name": "spotify",
"password": "dummy_spotify_pass"
},
{
"account_name": "file_system",
"password": "dummy_fs_pass"
},
...
]
ASSISTANT:
# So the Spotify password is an entry in the `passwords` list with the account_name=spotify.
spotify_password = [account_password["account_name"] == "spotify" for account_password in passwords][0]["password"]
print(spotify_password)
USER:
dummy_spotify_pass
ASSISTANT:
# When the task is completed, I need to call apis.supervisor.complete_task(). If there is an answer, I need to pass it as an argument `answer`. I will pass the spotify_password as an answer.
apis.supervisor.complete_task(answer=spotify_password)
USER:
Marked the active task complete.
----------------------------------------------
USER:
**Key instructions and disclaimers**:
1. The email addresses, access tokens and variables (e.g. spotify_password) in the example above were only for demonstration. Obtain the correct information by calling relevant APIs yourself.
2. Only generate valid code blocks, i.e., do not put them in ```...``` or add any extra formatting. Any thoughts should be put as code comments.
3. You can use the variables from the previous code blocks in the subsequent code blocks.
4. Write small chunks of code and only one chunk of code in every step. Make sure everything is working correctly before making any irreversible change.
5. The provided Python environment has access to its standard library. But modules and functions that have a risk of affecting the underlying OS, file system or process are disabled. You will get an error if do call them.
6. Any reference to a file system in the task instructions means the file system *app*, operable via given APIs, and not the actual file system the code is running on. So do not write code making calls to os-level modules and functions.
7. To interact with apps, only use the provided APIs, and not the corresponding Python packages. E.g., do NOT use `spotipy` for Spotify. Remember, the environment only has the standard library.
8. The provided API documentation has both the input arguments and the output JSON schemas. All calls to APIs and parsing its outputs must be as per this documentation.
9. For APIs that return results in "pages", make sure to consider all pages.
10. To obtain current date or time, use Python functions like `datetime.now()` or obtain it from the phone app. Do not rely on your existing knowledge of what the current date or time is.
11. For all temporal requests, use proper time boundaries, e.g., if I ask for something that happened yesterday, make sure to consider the time between 00:00:00 and 23:59:59. All requests are concerning a single, default (no) time zone.
12. Any reference to my friends, family or any other person or relation refers to the people in my phone's contacts list.
13. All my personal information, and information about my app account credentials, physical addresses and owned payment cards are stored in the "supervisor" app. You can access them via the APIs provided by the supervisor app.
14. Once you have completed the task, call `apis.supervisor.complete_task()`. If the task asks for some information, return it as the answer argument, i.e. call `apis.supervisor.complete_task(answer=<answer>)`. For tasks that do not require an answer, just skip the answer argument or pass it as None.
15. The answers, when given, should be just entity or number, not full sentences, e.g., `answer=10` for "How many songs are in the Spotify queue?". When an answer is a number, it should be in numbers, not in words, e.g., "10" and not "ten".
16. You can also pass `status="fail"` in the complete_task API if you are sure you cannot solve it and want to exit.
17. You must make all decisions completely autonomously and not ask for any clarifications or confirmations from me or anyone else.
USER:
Using these APIs, now generate code to solve the actual task:
My name is: {{ supervisor.first_name }} {{ supervisor.last_name }}. My personal email is {{ supervisor.email }} and phone number is {{ supervisor.phone_number }}.
Task:
{{ instruction }}
"""
PROMPT_TEMPLATE_WITH_EXPERIENCE = """
USER:
I am your supervisor and you are a super intelligent AI Assistant whose job is to achieve my day-to-day tasks completely autonomously.
To do this, you will need to interact with app/s (e.g., spotify, venmo, etc) using their associated APIs on my behalf. For this you will undertake a *multi-step conversation* using a python REPL environment. That is, you will write the python code and the environment will execute it and show you the result, based on which, you will write python code for the next step and so on, until you've achieved the goal. This environment will let you interact with app/s using their associated APIs on my behalf.
Here are three key APIs that you need to know to get more information
# To get a list of apps that are available to you.
print(apis.api_docs.show_app_descriptions())
# To get the list of apis under any app listed above, e.g. supervisor
print(apis.api_docs.show_api_descriptions(app_name='supervisor'))
# To get the specification of a particular api, e.g. supervisor app's show_account_passwords
print(apis.api_docs.show_api_doc(app_name='supervisor', api_name='show_account_passwords'))
Each code execution will produce an output that you can use in subsequent calls. Using these APIs, you can now generate code, that the environment will execute, to solve the task.
For example, consider the task:
My name is: {{ supervisor.first_name }} {{ supervisor.last_name }}. My personal email is {{ supervisor.email }} and phone number is {{ supervisor.phone_number }}.
Task:
What is the password for my Spotify account?
ASSISTANT:
# Okay. Lets first find which apps are available to get the password by looking at the app descriptions.
print(apis.api_docs.show_app_descriptions())
USER:
[
{
"name": "api_docs",
"description": "An app to search and explore API documentation."
},
{
"name": "supervisor",
"description": "An app to access supervisor's personal information, account credentials, addresses, payment cards, and manage the assigned task."
},
...
{
"name": "spotify",
"description": "A music streaming app to stream songs and manage song, album and playlist libraries."
},
{
"name": "venmo",
"description": "A social payment app to send, receive and request money to and from others."
},
...
]
ASSISTANT:
# Looks like the supervisor app could help me with that. Lets see what apis are available under this app.
print(apis.api_docs.show_api_descriptions(app_name='supervisor'))
USER:
[
...
"show_account_passwords : Show your supervisor's account passwords."
...
]
ASSISTANT:
# I can use `show_account_passwords` to get the passwords. Let me see its detailed specification to understand its arguments and output structure.
print(apis.api_docs.show_api_doc(app_name='supervisor', api_name='show_account_passwords'))
USER:
{
'app_name': 'supervisor',
'api_name': 'show_account_passwords',
'path': '/account_passwords',
'method': 'GET',
'description': "Show your supervisor's app account passwords.",
'parameters': [],
'response_schemas': {
'success': [{'account_name': 'string', 'password': 'string'}],
'failure': {'message': 'string'}
}
}
ASSISTANT:
# Okay, it requires no arguments. So I can just call it directly.
print(apis.supervisor.show_account_passwords())
USER:
[
{
"account_name": "spotify",
"password": "dummy_spotify_pass"
},
{
"account_name": "file_system",
"password": "dummy_fs_pass"
},
...
]
ASSISTANT:
# So the Spotify password is an entry in the `passwords` list with the account_name=spotify.
spotify_password = [account_password["account_name"] == "spotify" for account_password in passwords][0]["password"]
print(spotify_password)
USER:
dummy_spotify_pass
ASSISTANT:
# When the task is completed, I need to call apis.supervisor.complete_task(). If there is an answer, I need to pass it as an argument `answer`. I will pass the spotify_password as an answer.
apis.supervisor.complete_task(answer=spotify_password)
USER:
Marked the active task complete.
----------------------------------------------
USER:
**Key instructions and disclaimers**:
1. The email addresses, access tokens and variables (e.g. spotify_password) in the example above were only for demonstration. Obtain the correct information by calling relevant APIs yourself.
2. Only generate valid code blocks, i.e., do not put them in ```...``` or add any extra formatting. Any thoughts should be put as code comments.
3. You can use the variables from the previous code blocks in the subsequent code blocks.
4. Write small chunks of code and only one chunk of code in every step. Make sure everything is working correctly before making any irreversible change.
5. The provided Python environment has access to its standard library. But modules and functions that have a risk of affecting the underlying OS, file system or process are disabled. You will get an error if do call them.
6. Any reference to a file system in the task instructions means the file system *app*, operable via given APIs, and not the actual file system the code is running on. So do not write code making calls to os-level modules and functions.
7. To interact with apps, only use the provided APIs, and not the corresponding Python packages. E.g., do NOT use `spotipy` for Spotify. Remember, the environment only has the standard library.
8. The provided API documentation has both the input arguments and the output JSON schemas. All calls to APIs and parsing its outputs must be as per this documentation.
9. For APIs that return results in "pages", make sure to consider all pages.
10. To obtain current date or time, use Python functions like `datetime.now()` or obtain it from the phone app. Do not rely on your existing knowledge of what the current date or time is.
11. For all temporal requests, use proper time boundaries, e.g., if I ask for something that happened yesterday, make sure to consider the time between 00:00:00 and 23:59:59. All requests are concerning a single, default (no) time zone.
12. Any reference to my friends, family or any other person or relation refers to the people in my phone's contacts list.
13. All my personal information, and information about my app account credentials, physical addresses and owned payment cards are stored in the "supervisor" app. You can access them via the APIs provided by the supervisor app.
14. Once you have completed the task, call `apis.supervisor.complete_task()`. If the task asks for some information, return it as the answer argument, i.e. call `apis.supervisor.complete_task(answer=<answer>)`. For tasks that do not require an answer, just skip the answer argument or pass it as None.
15. The answers, when given, should be just entity or number, not full sentences, e.g., `answer=10` for "How many songs are in the Spotify queue?". When an answer is a number, it should be in numbers, not in words, e.g., "10" and not "ten".
16. You can also pass `status="fail"` in the complete_task API if you are sure you cannot solve it and want to exit.
17. You must make all decisions completely autonomously and not ask for any clarifications or confirmations from me or anyone else.
18. Some Related Experience to help you to complete the task:
{{experience}}
USER:
Using these APIs, now generate code to solve the actual task:
My name is: {{ supervisor.first_name }} {{ supervisor.last_name }}. My personal email is {{ supervisor.email }} and phone number is {{ supervisor.phone_number }}.
Task:
{{ instruction }}
"""

View file

@ -0,0 +1,5 @@
jinja2
loguru
openai
ray
pandas

View file

@ -0,0 +1,178 @@
import os
import time
import requests
import ray
from ray import logger
os.environ["APPWORLD_ROOT"] = "."
from dotenv import load_dotenv
load_dotenv("../../.env")
import json
from pathlib import Path
from appworld import load_task_ids
from appworld_react_agent import AppworldReactAgent
def handle_api_response(response: requests.Response):
"""Handle API response with proper error checking"""
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return None
return response.json()
def delete_workspace(workspace_id: str, api_url: str = "http://0.0.0.0:8002/"):
"""Delete the current workspace from the vector store"""
response = requests.post(
url=f"{api_url}vector_store",
json={
"workspace_id": workspace_id,
"action": "delete",
}
)
result = handle_api_response(response)
if result:
print(f"Workspace '{workspace_id}' deleted successfully")
def dump_memory(workspace_id: str, path: str = "./", api_url: str = "http://0.0.0.0:8002/"):
"""Dump the vector store memories to disk"""
response = requests.post(
url=f"{api_url}vector_store",
json={
"workspace_id": workspace_id,
"action": "dump",
"path": path,
}
)
result = handle_api_response(response)
if result:
print(f"Memory dumped to {path}")
def load_memory(workspace_id: str, path: str = "docs/library", api_url: str = "http://0.0.0.0:8002/"):
"""Load memories from disk into the vector store"""
response = requests.post(
url=f"{api_url}vector_store",
json={
"workspace_id": workspace_id,
"action": "load",
"path": path,
}
)
result = handle_api_response(response)
if result:
print(f"Memory loaded from {path}")
def run_agent(dataset_name: str, experiment_suffix: str, max_workers: int, num_runs: int = 1, use_task_memory: bool = False, make_task_memory: bool = False, workspace_id: str="appworld_v1", api_url: str = "http://0.0.0.0:8002/") :
experiment_name = dataset_name + "_" + experiment_suffix
path: Path = Path(f"./exp_result")
path.mkdir(parents=True, exist_ok=True)
task_ids = load_task_ids(dataset_name)
result: list = []
def dump_file():
with open(path / f"{experiment_name}.jsonl", "a") as f:
for x in result:
f.write(json.dumps(x) + "\n")
if max_workers > 1:
future_list: list = []
for i in range(max_workers):
# Assign tasks to each worker, ensuring each task runs num_runs times
worker_task_ids = task_ids[i::max_workers]
actor = AppworldReactAgent.remote(index=i,
task_ids=worker_task_ids,
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=use_task_memory,
make_task_memory=make_task_memory,
workspace_id=workspace_id,
api_url=api_url)
future = actor.execute.remote()
future_list.append(future)
time.sleep(1)
logger.info("submit complete")
for i, future in enumerate(future_list):
t_result = ray.get(future)
if t_result:
if isinstance(t_result, list):
result.extend(t_result)
else:
result.append(t_result)
logger.info(f"worker {i + 1}/{max_workers} complete")
dump_file()
else:
for index, task_id in enumerate(task_ids):
agent = AppworldReactAgent(index=index,
task_ids=[task_id],
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=use_task_memory,
make_task_memory=make_task_memory,
workspace_id=workspace_id,
api_url=api_url)
task_results = agent.execute()
if isinstance(task_results, list):
result.extend(task_results)
else:
result.append(task_results)
dump_file()
def main():
max_workers = 8
num_runs = 1 # Run each task once
workspace_id = "appworld"
api_url = "http://0.0.0.0:8002/"
if max_workers > 1:
ray.init(num_cpus=8)
# Clean up workspace before starting
logger.info("Deleting workspace...")
delete_workspace(workspace_id=workspace_id, api_url=api_url)
# First run to build task memories
logger.info("Start load experiments to build task memories")
load_memory(workspace_id=workspace_id, api_url=api_url)
# run_agent(dataset_name="dev", experiment_suffix="build-memory",
# max_workers=max_workers, num_runs=1,
# use_task_memory=False, make_task_memory=True,
# workspace_id=workspace_id, api_url=api_url)
for i in range(num_runs):
# Run experiments with task memory
logger.info("Start running experiments with task memory")
run_agent(dataset_name="dev", experiment_suffix=f"with-memory",
max_workers=max_workers, num_runs=1,
use_task_memory=True, make_task_memory=False,
workspace_id=workspace_id, api_url=api_url)
# Run experiments without task memory
logger.info("Start running experiments without task memory")
run_agent(dataset_name="dev", experiment_suffix=f"no-memory",
max_workers=max_workers, num_runs=1,
use_task_memory=False, make_task_memory=False,
workspace_id=workspace_id, api_url=api_url)
if __name__ == "__main__":
main()

View file

@ -0,0 +1,160 @@
import json
from pathlib import Path
from collections import defaultdict
import pandas as pd
from loguru import logger
def calculate_best_at_k(scores: list, k: int) -> float:
"""
Calculate best@k
Divide scores into groups of size k, take the maximum value in each group,
then average these maximum values
Args:
scores: List of after_score values for all runs of a task
k: Group size
Returns:
best@k value
"""
if len(scores) % k != 0:
raise ValueError(f"Length of scores ({len(scores)}) must be divisible by k ({k})")
group_maxs = []
for i in range(0, len(scores), k):
group = scores[i:i + k]
group_maxs.append(max(group))
return sum(group_maxs) / len(group_maxs)
def calculate_pass_at_k(scores: list, k: int) -> float:
if len(scores) % k != 0:
raise ValueError(f"Length of scores ({len(scores)}) must be divisible by k ({k})")
group_maxs = []
for i in range(0, len(scores), k):
group = scores[i:i + k]
is_pass = 1.0 if max(group) >=1.0 else 0.0
group_maxs.append(is_pass)
return sum(group_maxs) / len(group_maxs)
def get_possible_k_values(total_runs: int) -> list:
"""
Get all possible k values (factors of total_runs)
Args:
total_runs: Total number of runs
Returns:
List of k values in descending order
"""
k_values = []
for k in range(1, total_runs + 1):
if total_runs % k == 0:
k_values.append(k)
return sorted(k_values, reverse=True) # Sort from large to small
def run_exp_statistic():
path: Path = Path(f"./exp_result")
# Store results for all experiments
all_results = {}
for file in [f for f in path.glob("*.jsonl") if not f.stem[-1].isdigit()]:
# Group results by task_id
task_results = defaultdict(list)
with open(file, "r") as f:
for line in f:
if not line.strip():
continue
data = json.loads(line)
if isinstance(data, list):
for part_data in data:
task_id = part_data["task_id"]
after_score = part_data["after_score"]
task_results[task_id].append(after_score)
else:
task_id = data["task_id"]
after_score = data["after_score"]
task_results[task_id].append(after_score)
if not task_results:
logger.warning(f"No valid data found in file {file}")
continue
# Check if each task has consistent number of runs
run_counts = [len(scores) for scores in task_results.values()]
if len(set(run_counts)) > 1:
logger.warning(f"Inconsistent number of runs for different tasks in file {file}: {set(run_counts)}")
continue
num_runs = run_counts[0]
logger.info(f"File {file}: {len(task_results)} tasks, {num_runs} runs per task")
# Get all possible k values
k_values = get_possible_k_values(num_runs)
logger.info(f"Calculable best@k values: {k_values}")
# Calculate various best@k values
file_results = {"file": file.name}
for k in k_values:
best_at_k_scores = []
pass_at_k_scores = []
for task_id, scores in task_results.items():
try:
best_k_score = calculate_best_at_k(scores, k)
pass_at_k_score = calculate_pass_at_k(scores, k)
pass_at_k_scores.append(pass_at_k_score)
best_at_k_scores.append(best_k_score)
except ValueError as e:
logger.error(f"Error calculating best@{k} for task {task_id}: {e}")
continue
if best_at_k_scores:
avg_best_at_k = sum(best_at_k_scores) / len(best_at_k_scores)
file_results[f"best@{k}"] = avg_best_at_k
logger.info(f"file={file.name} best@{k}={avg_best_at_k:.4f}")
if pass_at_k_scores:
avg_pass_at_k = sum(pass_at_k_scores) / len(pass_at_k_scores)
file_results[f"pass@{k}"] = avg_pass_at_k
logger.info(f"file={file.name} pass@{k}={avg_pass_at_k:.4f}")
all_results[file.name] = file_results
# Create and display table
if all_results:
df = pd.DataFrame(list(all_results.values()))
df = df.set_index('file')
# Sort columns by the number in column name (best@8, best@4, best@2, best@1)
pass_columns = [col for col in df.columns if col.startswith('pass@')]
# best_columns = [col for col in df.columns]
pass_columns.sort(key=lambda x: x, reverse=False)
df = df[pass_columns]
print("\n" + "=" * 80)
print("Experiment Results Summary Table")
print("=" * 80)
print(df.round(4))
print("=" * 80)
# Save table to CSV
output_path = path / "experiment_summary.csv"
df.to_csv(output_path)
logger.info(f"Results table saved to: {output_path}")
else:
logger.warning("No valid experiment results found")
if __name__ == "__main__":
run_exp_statistic()

View file

611
cookbook/bfcl/bfcl_agent.py Normal file
View file

@ -0,0 +1,611 @@
import os
os.environ["BFCL_DATA_PATH"] = "data/multiturn_data_base_val.jsonl"
os.environ["BFCL_ANSWER_PATH"] = "data/possible_answer"
from dotenv import load_dotenv
load_dotenv("../../.env")
import re
import time
import json
import ray
import warnings
import tempfile
import requests
import datetime
from tqdm import tqdm
from pathlib import Path
from loguru import logger
from openai import OpenAI
from typing import Dict, List, Any
from bfcl_utils import (
load_test_case,
handle_user_turn,
handle_tool_calls,
extract_tool_schema,
extract_single_turn_response,
extract_multi_turn_responses,
capture_and_print_score_files,
create_error_response
)
from bfcl_eval.model_handler.api_inference.qwen import QwenAPIHandler
from bfcl_eval.eval_checker.multi_turn_eval.multi_turn_utils import (
is_empty_execute_response,
)
from bfcl_eval.eval_checker.eval_runner import (
multi_turn_runner,
ast_file_runner,
)
from bfcl_eval.eval_checker.eval_runner_helper import record_cost_latency
from bfcl_eval.utils import (
is_multi_turn,
is_relevance_or_irrelevance,
find_file_with_suffix,
load_file,
)
@ray.remote
class BFCLAgent:
"""A minimal ReAct Agent for BFCL-v3(multi-turn) tasks."""
def __init__(self,
index: int,
task_ids: List[str],
experiment_name: str,
data_path: str = os.getenv("BFCL_DATA_PATH"),
answer_path: Path = Path(os.getenv("BFCL_ANSWER_PATH")),
model_name: str = "qwen3-8b",
temperature: float = 0.9,
max_interactions: int = 30,
max_response_size: int = 2000,
num_runs: int = 1,
enable_thinking: bool = False,
use_memory: bool = False,
use_memory_addition: bool = False,
use_memory_deletion: bool = False,
delete_freq: int = 10,
freq_threshold: int = 5,
utility_threshold: float = 0.5,
memory_base_url: str = "http://0.0.0.0:8001/",
memory_workspace_id: str = "bfcl_8b_0725"):
self.index: int = index
self.task_ids: List[str] = task_ids
self.categories: List[str] = [task_id.rsplit("_", 1)[0] if "_" in task_id else task_id for task_id in task_ids]
self.experiment_name: str = experiment_name
self.data_path: str = data_path
self.answer_path: Path = answer_path
self.model_name: str = model_name
self.temperature: float = temperature
self.max_interactions: int = max_interactions
self.max_response_size: int = max_response_size
self.num_runs: int = num_runs
self.enable_thinking: bool = enable_thinking
self.use_memory: bool = use_memory
self.use_memory_addition: bool = use_memory_addition if use_memory else False
self.use_memory_deletion: bool = use_memory_deletion if use_memory else False
self.delete_freq: int = delete_freq
self.freq_threshold: int = freq_threshold
self.utility_threshold: float = utility_threshold
self.memory_base_url: str = memory_base_url
self.memory_workspace_id: str = memory_workspace_id
self.history: List[List[List[dict]]] = [[] for _ in range(num_runs)]
self.retrieved_memory_list: List[List[List[Any]]] = [[] for _ in range(num_runs)]
self.test_entry: List[List[Dict[str, Any]]] = [[] for _ in range(num_runs)]
self.original_test_entry: List[List[Dict[str, Any]]] = [[] for _ in range(num_runs)]
self.tool_schema: List[List[List[dict]]] = [[] for _ in range(num_runs)]
self.current_turn = [[0 for _ in range(len(task_ids))] for _ in range(num_runs)]
for run_id in range(num_runs):
for task_index in range(len(task_ids)):
self.init_state(run_id, task_index)
def init_state(self, run_id, i) -> Dict[str, Any]:
self.test_entry[run_id].append(load_test_case(self.data_path, self.task_ids[i]))
self.original_test_entry[run_id].append(self.test_entry[run_id][i].get("extra", {}))
self.tool_schema[run_id].append(extract_tool_schema(self.test_entry[run_id][i].get("tools", [{}])))
msg = self.test_entry[run_id][i].get("messages", [])[0]
if self.use_memory:
query = msg["content"]
response = self.get_memory(query)
if len(response["metadata"]["memory_list"]):
self.retrieved_memory_list[run_id].append(response["metadata"]["memory_list"])
exp: str = response["answer"]
# print(f"memory_merged={exp}")
self.history[run_id].append([self.get_query_with_memory(query, exp)])
else:
self.retrieved_memory_list[run_id].append([])
self.history[run_id].append([msg])
else:
self.history[run_id].append([msg])
self.current_turn[run_id][i] = 1
def get_query_with_memory(self, query: str, memory: str):
return {
"role": "user",
"content": "Task:\n" + query + "\n\nSome Related Experience to help you to complete the task:\n" + memory
}
def get_traj_from_task_history(self, task_id: str, task_history: list, reward: float):
return {
"task_id": task_id,
"messages": task_history,
"score": reward
}
def get_memory(self, query: str):
response = requests.post(url=self.memory_base_url + "retrieve_task_memory", json={
"workspace_id": self.memory_workspace_id,
"query": query,
"top_k": 5
})
if response.status_code != 200:
logger.info(response.text)
return ""
response = response.json()
logger.info(f"query: {query}, response: {response}")
return response
def add_memory(self, trajectories):
response = requests.post(url=self.memory_base_url + "summary_task_memory", json={
"workspace_id": self.memory_workspace_id,
"trajectories": trajectories,
})
response.raise_for_status()
response = response.json()
logger.info(f"add new memorys: {response["metadata"]["memory_list"]}")
def update_memory_information(self, memory_list, update_utility: bool=False):
response = requests.post(url=self.memory_base_url + "record_task_memory", json={
"workspace_id": self.memory_workspace_id,
"memory_dicts": memory_list,
"update_utility": update_utility
})
response.raise_for_status()
logger.info(response.json())
def delete_memory(self):
response = requests.post(url=self.memory_base_url + "delete_task_memory", json={
"workspace_id": self.memory_workspace_id,
"freq_threshold": self.freq_threshold,
"utility_threshold": self.utility_threshold
})
response.raise_for_status()
def call_llm(self, messages: list, tool_schemas: list[dict]) -> str:
for i in range(100):
try:
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
# Change this function to modify the base llm
response = client.chat.completions.create(
model=self.model_name,
messages=messages,
tools=tool_schemas,
temperature=self.temperature,
seed=0,
extra_body={"enable_thinking": self.enable_thinking},
stream=self.enable_thinking,
parallel_tool_calls=True,
)
if not self.enable_thinking:
out_msg = response.choices[0].message
return out_msg.model_dump(exclude_unset=True, exclude_none=True)
else:
reasoning_content = "" # Complete reasoning process
answer_content = "" # Define complete response
tool_info = [] # Store tool invocation information
is_answering = False # Determine whether the reasoning process has finished and response has started
for chunk in response:
if not chunk.choices:
# Handle usage information
continue
else:
delta = chunk.choices[0].delta
# Handle AI's thought process (chain reasoning)
if hasattr(delta, 'reasoning_content') and delta.reasoning_content is not None:
reasoning_content += delta.reasoning_content
# Handle final response content
else:
if not is_answering: # Print title when entering the response phase for the first time
is_answering = True
if delta.content is not None:
answer_content += delta.content
# Handle tool invocation information (support parallel tool calls)
if delta.tool_calls is not None:
for tool_call in delta.tool_calls:
index = tool_call.index # Tool call index, used for parallel calls
# Dynamically expand tool information storage list
while len(tool_info) <= index:
tool_info.append({"id": "", "type": "function", "index": index, "function": { "name": "", "arguments": "" }})
# Collect tool call ID (used for subsequent function calls)
if tool_call.id:
tool_info[index]['id'] += tool_call.id
# Collect function name (used for subsequent routing to specific functions)
if tool_call.function and tool_call.function.name:
tool_info[index]['function']['name'] += tool_call.function.name
# Collect function parameters (in JSON string format, need subsequent parsing)
if tool_call.function and tool_call.function.arguments:
tool_info[index]['function']['arguments'] += tool_call.function.arguments
msg = {
"role": "assistant",
"content": answer_content,
"reasoning_content": reasoning_content,
}
if tool_info:
msg["tool_calls"] = tool_info
return msg
except Exception as e:
logger.exception(f"encounter error with {e.args}")
time.sleep(1 + i * 10)
return "call llm error"
def env_step(self, run_id: int, index: int, messages: str) -> str:
"""
Process one step in the conversation.
Both single turn and multi turn are supported.
Args:
messages: List of conversation messages, with the last one being assistant response
test_entry: Test entry containing initial_config, involved_classes, question etc.
**kwargs: Additional arguments for compatibility
Returns:
Dict containing next message and tools if applicable
"""
try:
if not messages:
return handle_user_turn(self.original_test_entry[run_id][index], self.current_turn[run_id][index])
if messages[-1]["role"] != "assistant":
return create_error_response(
"Last message must be from assistant"
)
if "tool_calls" in messages[-1] and len(messages[-1]["tool_calls"]) > 0:
try:
tool_calls = messages[-1]["tool_calls"]
decoded_calls = self._convert_tool_calls_to_execution_format(
tool_calls
)
# decoded_calls:[function(param=xxx)]
print(f"decoded_calls: {decoded_calls}")
if is_empty_execute_response(decoded_calls):
warnings.warn(
f"is_empty_execute_response: {is_empty_execute_response(decoded_calls)}"
)
return handle_user_turn(self.original_test_entry[run_id][index], self.current_turn[run_id][index])
return handle_tool_calls(
tool_calls, decoded_calls, self.original_test_entry[run_id][index], self.current_turn[run_id][index]
)
except Exception as e:
warnings.warn(f"Errors during tool invocation: {str(e)}")
return handle_user_turn(self.original_test_entry[run_id][index], self.current_turn[run_id][index])
else:
return handle_user_turn(self.original_test_entry[run_id][index], self.current_turn[run_id][index])
except Exception as e:
return create_error_response(f"Failed to process request: {str(e)}")
def _convert_tool_calls_to_execution_format(
self, tool_calls: List[Dict[str, Any]]
) -> List[str]:
"""
Convert OpenAI format tool calls to execution format.
Args:
tool_calls: List of tool calls in OpenAI format
Returns:
List of function calls in string format
"""
execution_list = []
for tool_call in tool_calls:
function = tool_call.get("function", {})
function_name = function.get("name", "")
try:
arguments = function.get("arguments", "{}")
if isinstance(arguments, str):
args_dict = json.loads(arguments)
else:
args_dict = arguments
args_str = ", ".join([f"{k}={repr(v)}" for k, v in args_dict.items()])
execution_list.append(f"{function_name}({args_str})")
except Exception as e:
execution_list.append(f"{function_name}()")
return execution_list
def get_reward(self, run_id, index) -> float:
try:
if not self.history[run_id][index] or not self.original_test_entry[run_id][index]:
return 0.0
model_name = "env_handler"
handler = QwenAPIHandler(
model_name, temperature=1.0
) # FIXME: magic number
model_result_data = self._convert_conversation_to_eval_format(run_id, index)
prompt_data = [self.original_test_entry[run_id][index]]
state = {"leaderboard_table": {}}
record_cost_latency(
state["leaderboard_table"], model_name, [model_result_data]
)
if is_relevance_or_irrelevance(self.categories[index]):
accuracy, _ = self._eval_relevance_test(
handler, model_result_data, prompt_data, model_name, self.category
)
else:
# Find the corresponding possible answer file
possible_answer_file = find_file_with_suffix(
self.answer_path, self.categories[index]
)
possible_answer = load_file(possible_answer_file, sort_by_id=True)
possible_answer = [
item for item in possible_answer if item["id"] == self.task_ids[index]
]
if is_multi_turn(self.categories[index]):
accuracy, _ = self._eval_multi_turn_test(
handler,
model_result_data,
prompt_data,
possible_answer,
model_name,
self.categories[index],
)
else:
accuracy, _ = self._eval_single_turn_test(
handler,
model_result_data,
prompt_data,
possible_answer,
model_name,
self.categories[index],
)
print(f"model_result_data: {model_result_data}")
print(f"possible_answer: {possible_answer}") if possible_answer else None
return accuracy
except Exception as e:
import traceback
traceback.print_exc()
return 0
def _convert_conversation_to_eval_format(self, run_id, index) -> Dict[str, Any]:
"""
Convert conversation history to evaluation format.
Args:
conversation_result: Result from run_conversation
original_test_entry: Original test entry data
Returns:
Data in format expected by multi_turn_runner or other runners
"""
if is_multi_turn(self.categories[index]):
turns_data = extract_multi_turn_responses(self.history[run_id][index])
else:
turns_data = extract_single_turn_response(self.history[run_id][index])
model_result_data = {
"id": self.task_ids[index],
"result": turns_data,
"latency": 0,
"input_token_count": 0,
"output_token_count": 0,
}
return model_result_data
def _eval_multi_turn_test(
self,
handler,
model_result_data,
prompt_data,
possible_answer,
model_name,
test_category,
):
"""
Evaluate multi-turn test.
Args:
handler: Model handler instance
model_result_data: Model result data
prompt_data: Prompt data
possible_answer: Possible answer data
model_name: Name of the model
test_category: Category of the test
Returns:
Tuple of (accuracy, total_count)
"""
with tempfile.TemporaryDirectory() as temp_dir:
score_dir = Path(temp_dir)
accuracy, total_count = multi_turn_runner(
handler=handler,
model_result=[model_result_data],
prompt=prompt_data,
possible_answer=possible_answer,
model_name=model_name,
test_category=test_category,
score_dir=score_dir,
)
capture_and_print_score_files(
score_dir, model_name, test_category, "multi_turn"
)
return accuracy, total_count
def _eval_single_turn_test(
self,
handler,
model_result_data,
prompt_data,
possible_answer,
model_name,
test_category,
):
"""
Evaluate single-turn AST test.
Args:
handler: Model handler instance
model_result_data: Model result data
prompt_data: Prompt data
possible_answer: Possible answer data
model_name: Name of the model
test_category: Category of the test
Returns:
Tuple of (accuracy, total_count)
"""
language = "Python"
if "java" in test_category.lower():
language = "Java"
elif "js" in test_category.lower() or "javascript" in test_category.lower():
language = "JavaScript"
with tempfile.TemporaryDirectory() as temp_dir:
score_dir = Path(temp_dir)
accuracy, total_count = ast_file_runner(
handler=handler,
model_result=[model_result_data],
prompt=prompt_data,
possible_answer=possible_answer,
language=language,
test_category=test_category,
model_name=model_name,
score_dir=score_dir,
)
capture_and_print_score_files(
score_dir, model_name, test_category, "single_turn"
)
return accuracy, total_count
def execute(self):
result = []
counter = 0
for task_index, task_id in enumerate(tqdm(self.task_ids, desc=f"ray_index={self.index}")):
for run_id in range(self.num_runs):
try:
start_time = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
for i in range(self.max_interactions):
llm_output = self.call_llm(self.history[run_id][task_index], self.tool_schema[run_id][task_index])
self.history[run_id][task_index].append(llm_output)
env_output = self.env_step(run_id, task_index, self.history[run_id][task_index])
# Possible env_output returns after environment interaction:
# 1. Triggers a query with available tools list: {"messages": [{"role": "user", "content": user_query}], "tools": tools}
# 2. Returns tool invocation result: {"messages": [{"role": "tool", "content": {<execution_results>}, 'tool_call_id': 'chatcmpl-tool-xxx'}]}
# <execution_results>: when success, returns result dicts, e.g., {"travel_cost_list": [1140.0]}, when error, returns error message, e.g., {"error": "cd: temporary: No such directory. You cannot use path to change directory."}
# 3. Conversation completion: {"messages": [{"role": "env", "content": "[CONVERSATION_COMPLETED]"}]}
# 4. Program error: {"messages": [{"role": "env", "content": f"[ERROR] {error_message}"}]}
# tool_list update
if "tools" in env_output:
self.tool_schema[run_id][task_index] = extract_tool_schema(env_output["tools"])
new_tool_calls=[]
new_tool_call_ids=[]
next_user_msg = ""
for idx, msg in enumerate(env_output.get("messages", [])):
if msg["role"] == "tool" and len(msg["content"])>0:
new_tool_calls.append(msg.get("content", ""))
new_tool_call_ids.append(msg.get("tool_call_id", ""))
elif msg["role"] == "user":
next_user_msg = msg.get("content", "")
self.current_turn[run_id][task_index] += 1
else: # for env role messages
next_user_msg = msg.get("content", "")
if new_tool_calls:
for idx, call in enumerate(new_tool_calls):
self.history[run_id][task_index].append({"role": "tool", "content": str(call), "tool_call_id": new_tool_call_ids[idx]})
else:
self.history[run_id][task_index].append({"role": "user", "content": next_user_msg})
logger.info(f"index={self.index} task_id={task_id} iteration={i}")
if self.task_completed(run_id, task_index):
break
reward = self.get_reward(run_id, task_index)
if self.use_memory:
if reward == 1 and self.use_memory_addition: # selectively add memories when succeed
new_traj_list = [self.get_traj_from_task_history(task_id, self.history[run_id][task_index], reward)]
self.add_memory(new_traj_list)
# update the freq & utility attributes of retrieved memories
update_utility: bool = (reward == 1)
self.update_memory_information(self.retrieved_memory_list[run_id][task_index], update_utility)
counter += 1
if self.use_memory_deletion and counter % self.delete_freq == 0:
self.delete_memory()
t_result = {
"run_id": run_id,
"task_id": self.task_ids[task_index],
"experiment_name": self.experiment_name,
"task_completed": self.task_completed(run_id, task_index),
"reward": reward,
"task_history": self.history[run_id][task_index],
"task_start_time": start_time,
}
result.append(t_result)
except Exception as e:
logger.exception(f"encounter error with {e.args}")
result.append({})
return result
def task_completed(self, run_id, index):
"""
Check if task is completed.
Returns:
True if task is completed, False otherwise
"""
return self.history[run_id][index][-1]["content"] == "[CONVERSATION_COMPLETED]"
def main():
with open(os.getenv("BFCL_DATA_PATH"), "r", encoding="utf-8") as f:
task_ids = [json.loads(l)["id"] for l in f]
dataset_name = "dev"
agent = BFCLAgent(
index=0,
task_id=task_ids[0],
experiment_name=f"zouying_{dataset_name}",
)
result = agent.execute()
logger.info(f"result={json.dumps(result)}")
if __name__ == "__main__":
main()

385
cookbook/bfcl/bfcl_utils.py Normal file
View file

@ -0,0 +1,385 @@
import json
import tempfile
from pathlib import Path
from typing import Dict, List, Any
from bfcl_eval.constants.type_mappings import GORILLA_TO_OPENAPI
from bfcl_eval.constants.default_prompts import (
DEFAULT_USER_PROMPT_FOR_ADDITIONAL_FUNCTION_FC,
)
from bfcl_eval.model_handler.model_style import ModelStyle
from bfcl_eval.model_handler.utils import (
convert_to_function_call,
convert_to_tool,
default_decode_ast_prompting,
default_decode_execute_prompting,
format_execution_results_prompting,
func_doc_language_specific_pre_processing,
retry_with_backoff,
system_prompt_pre_processing_chat_model,
)
from bfcl_eval.eval_checker.multi_turn_eval.multi_turn_utils import (
execute_multi_turn_func_call,
)
def load_test_case(data_path: str, test_id: str | None) -> Dict[str, Any]:
if not Path(data_path).exists():
raise FileNotFoundError(f"BFCL data file '{data_path}' not found")
if test_id is None:
raise ValueError("task_id is required")
with open(data_path, "r", encoding="utf-8") as f:
if str(test_id).isdigit():
idx = int(test_id)
for line_no, line in enumerate(f):
if line_no == idx:
return json.loads(line)
raise ValueError(f"Test case index {idx} not found in {data_path}")
else:
for line in f:
data = json.loads(line)
if data.get("id") == test_id:
return data
raise ValueError(f"Test case id '{test_id}' not found in {data_path}")
def handle_user_turn(
test_entry: Dict[str, Any], current_turn: int
) -> Dict[str, Any]:
"""
Handle user turn by returning appropriate content from test_entry["question"].
For non-first turns, processes user query and tools.
Args:
test_entry: Test entry containing conversation data
current_turn: Current turn number
Returns:
Response containing next user message and tools
"""
try:
current_turn_message = []
tools = compile_tools(test_entry)
questions = test_entry.get("question", [])
holdout_function = test_entry.get("holdout_function", {})
if str(current_turn) in holdout_function:
test_entry["function"].extend(holdout_function[str(current_turn)])
tools = compile_tools(test_entry)
assert (
len(questions[current_turn]) == 0
), "Holdout turn should not have user message."
current_turn_message = [
{
"role": "user",
"content": DEFAULT_USER_PROMPT_FOR_ADDITIONAL_FUNCTION_FC,
}
]
return create_user_response(current_turn_message, tools)
if current_turn >= len(questions):
return create_completion_response()
current_turn_message = questions[current_turn]
return create_user_response(current_turn_message, tools)
except Exception as e:
return create_error_response(f"Failed to process user message: {str(e)}")
def handle_tool_calls(
tool_calls: List[Dict[str, Any]],
decoded_calls: list[str],
test_entry: Dict[str, Any],
current_turn: int,
) -> Dict[str, Any]:
"""
Handle tool calls from assistant.
Args:
tool_calls: List of tool calls in OpenAI format
decoded_calls: List of decoded function calls
test_entry: Test entry containing environment data
current_turn: Current turn number
Returns:
Response containing tool execution results
"""
execution_results, _ = execute_multi_turn_func_call(
func_call_list=decoded_calls,
initial_config=test_entry["initial_config"],
involved_classes=test_entry["involved_classes"],
model_name="env_handler",
test_entry_id=test_entry["id"],
long_context=(
"long_context" in test_entry["id"] or "composite" in test_entry["id"]
),
is_evaL_run=False,
)
# print('execution_results in handler_tool_calls:', execution_results)
return create_tool_response(tool_calls, execution_results)
def compile_tools(test_entry: dict) -> list:
"""
Compile functions into tools format.
Args:
test_entry: Test entry containing functions
Returns:
List of tools in OpenAI format
"""
functions: list = test_entry["function"]
test_category: str = test_entry["id"].rsplit("_", 1)[0]
functions = func_doc_language_specific_pre_processing(functions, test_category)
tools = convert_to_tool(functions, GORILLA_TO_OPENAPI, ModelStyle.OpenAI_Completions)
return tools
def create_tool_response(
tool_calls: List[Dict[str, Any]], execution_results: List[str]
) -> Dict[str, Any]:
"""
Create response for tool calls.
Args:
tool_calls: List of tool calls
execution_results: List of execution results
Returns:
Response containing tool execution results
"""
tool_messages = []
for i, (tool_call, result) in enumerate(zip(tool_calls, execution_results)):
tool_messages.append(
{
"role": "tool",
"content": result,
"tool_call_id": tool_call.get("id", f"call_{i}"),
}
)
return {"messages": tool_messages}
def create_user_response(
question_turn: List[Dict[str, Any]], tools: List[Dict[str, Any]]
) -> Dict[str, Any]:
"""
Create response containing user message.
Args:
question_turn: List of messages for current turn
tools: List of available tools
Returns:
Response containing user message and tools
"""
user_content = ""
for msg in question_turn:
if msg["role"] == "user":
user_content = msg["content"]
break
return {"messages": [{"role": "user", "content": user_content}], "tools": tools}
def create_completion_response() -> Dict[str, Any]:
"""
Create response indicating conversation completion.
Returns:
Response with completion message
"""
return {"messages": [{"role": "env", "content": "[CONVERSATION_COMPLETED]"}]}
def create_error_response(error_message: str) -> Dict[str, Any]:
"""
Create response for error conditions.
Args:
error_message: Error message to include
Returns:
Response containing error message
"""
return {"messages": [{"role": "env", "content": f"[ERROR] {error_message}"}]}
def decode_execute(result):
"""
Decode execute results for compatibility with evaluation framework.
Args:
result: Result to decode
Returns:
List of decoded function calls
"""
return default_decode_execute_prompting(result)
def extract_single_turn_response(messages: List[Dict[str, Any]]) -> str:
"""
Extract single-turn response from conversation messages.
Args:
messages: List of conversation messages
Returns:
String representation of the response
"""
for message in reversed(messages):
if message["role"] == "assistant":
if "tool_calls" in message and message["tool_calls"]:
formatted_calls = []
for tool_call in message["tool_calls"]:
formatted_call = format_single_tool_call_for_eval(
tool_call
)
if formatted_call:
formatted_calls.append(formatted_call)
return "\n".join(formatted_calls) if formatted_calls else ""
elif message.get("content"):
return message["content"]
return ""
def extract_multi_turn_responses(
messages: List[Dict[str, Any]]
) -> List[List[str]]:
"""
Extract multi-turn responses from conversation messages.
Args:
messages: List of conversation messages
Returns:
List of turns, each turn is a list of function call strings
"""
turns_data = []
current_turn_responses = []
i = 0
while i < len(messages):
message = messages[i]
if message["role"] == "user":
if current_turn_responses:
turns_data.append(current_turn_responses)
current_turn_responses = []
i += 1
while i < len(messages) and messages[i]["role"] == "assistant":
assistant_msg = messages[i]
if "tool_calls" in assistant_msg and assistant_msg["tool_calls"]:
for tool_call in assistant_msg["tool_calls"]:
formatted_call = format_single_tool_call_for_eval(
tool_call
)
if formatted_call:
current_turn_responses.append(formatted_call)
i += 1
while i < len(messages) and messages[i]["role"] == "tool":
i += 1
else:
i += 1
if current_turn_responses:
turns_data.append(current_turn_responses)
return turns_data
def format_single_tool_call_for_eval(tool_call: Dict[str, Any]) -> str:
"""
Format a single tool call into string representation for evaluation.
Args:
tool_call: Single tool call in OpenAI format
Returns:
Formatted string representation
"""
function = tool_call.get("function", {})
function_name = function.get("name", "")
try:
arguments = function.get("arguments", "{}")
if isinstance(arguments, str):
args_dict = json.loads(arguments)
else:
args_dict = arguments
args_str = ", ".join([f"{k}={repr(v)}" for k, v in args_dict.items()])
return f"{function_name}({args_str})"
except Exception as e:
return f"{function_name}()"
def capture_and_print_score_files(
score_dir: Path, model_name: str, test_category: str, eval_type: str
):
"""
Capture and print contents of score files written to score_dir.
Args:
score_dir: Directory containing score files
model_name: Name of the model
test_category: Category of the test
eval_type: Type of evaluation (relevance/multi_turn/single_turn)
"""
try:
print(f"\n=== {eval_type.upper()} Evaluation Result Files ===")
print(f"Model: {model_name}")
print(f"Test Category: {test_category}")
print(f"Evaluation Type: {eval_type}")
for file_path in score_dir.rglob("*"):
if file_path.is_file():
relative_path = file_path.relative_to(score_dir)
print(f"\n--- File: {relative_path} ---")
try:
with open(file_path, "r", encoding="utf-8") as f:
content = f.read()
if (
file_path.suffix == ".json"
or content.strip().startswith("{")
or content.strip().startswith("[")
):
try:
import json
lines = content.strip().split("\n")
formatted_lines = []
for line in lines:
if line.strip():
parsed = json.loads(line)
formatted_lines.append(
json.dumps(
parsed, ensure_ascii=False, indent=2
)
)
content = "\n".join(formatted_lines)
except json.JSONDecodeError:
pass
print(content)
except UnicodeDecodeError:
print(f"[Binary file, size: {file_path.stat().st_size} bytes]")
except Exception as e:
print(f"[Error reading file: {str(e)}]")
print(f"=== {eval_type.upper()} Evaluation Result Files End ===\n")
except Exception as e:
print(f"Error capturing evaluation result files: {str(e)}")
def extract_tool_schema(tools):
for i in range(len(tools)):
tools[i]['function'].pop("response")
return tools

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,253 @@
import json
import requests
import argparse
from pathlib import Path
from typing import List, Dict, Any
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor, as_completed
def load_task_case(data_path: str, task_id: str | None) -> Dict[str, Any]:
"""按 ID加载单条 JSONL 训练用例。找不到就抛错。"""
if not Path(data_path).exists():
raise FileNotFoundError(f"BFCL data file '{data_path}' not found")
if task_id is None:
raise ValueError("task_id is required")
with open(data_path, "r", encoding="utf-8") as f:
if str(task_id).isdigit():
idx = int(task_id)
for line_no, line in enumerate(f):
if line_no == idx:
return json.loads(line)
raise ValueError(f"Task case index {idx} not found in {data_path}")
else:
for line in f:
data = json.loads(line)
if data.get("id") == task_id:
return data
raise ValueError(f"Task case id '{task_id}' not found in {data_path}")
def get_tool_prompt(tools):
tool_prompt = "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>"
for tool in tools:
tool_prompt += "\n" + json.dumps(tool)
tool_prompt += "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call>"
return tool_prompt
def group_trajectories_by_task_id(jsonl_entries: List[Dict[str, Any]]) -> List[List[Any]]:
"""
根据task_id字段对trajectories进行分组
Args:
jsonl_entries: JSONL条目列表
Returns:
List[List[Any]]: 按task_id分组的trajectory列表
"""
# 按task_id分组
grouped = defaultdict(list)
for entry in jsonl_entries:
task_id = entry.get("task_id", "")
taks_case = load_task_case("data/multiturn_data_base.jsonl", task_id)
tools = taks_case.get("tools", [{}])
from bfcl_utils import extract_tool_schema
tool_schema = extract_tool_schema(tools)
entry["task_history"][0]["content"] += get_tool_prompt(tool_schema)
grouped[task_id].append(entry)
# 对每组只保留最大和最小reward的两个
filtered_groups = []
for key, trajectories in grouped.items():
if len(trajectories) == 1:
# 只有一个trajectory直接保留
filtered_groups.append(trajectories)
elif len(trajectories) == 2:
# 有两个trajectory直接保留
filtered_groups.append(trajectories)
else:
# 多个trajectory选择最大和最小reward的
trajectories.sort(key=lambda t: t["reward"])
min_reward_traj = trajectories[0] # 最小reward
max_reward_traj = trajectories[-1] # 最大reward
filtered_groups.append([min_reward_traj, max_reward_traj])
return filtered_groups
def post_to_summarizer(trajectories: List[Any], service_url: str, workspace_id: str) -> Dict[str, Any]:
"""
将trajectories发送到summarizer服务
Args:
trajectories: trajectory列表
service_url: 服务URL
workspace_id: 工作空间ID
Returns:
响应结果
"""
trajectory_dicts = [{
"task_id": traj["task_id"],
"messages": traj["task_history"],
"score": traj["reward"]
} for traj in trajectories]
request_data = {
"traj_list": trajectory_dicts,
"workspace_id": workspace_id
}
try:
response = requests.post(f"{service_url}/summarizer", json=request_data)
response.raise_for_status()
return response.json()
except Exception as e:
return {"error": str(e), "trajectories_count": len(trajectories)}
def process_trajectories_with_threads(grouped_trajectories: List[List[Any]],
service_url: str,
workspace_id: str,
n_threads: int = 4) -> List[Dict[str, Any]]:
"""
使用多线程处理trajectories组
Args:
grouped_trajectories: 按task_id分组的trajectory列表
service_url: summarizer服务URL
workspace_id: 工作空间ID
n_threads: 线程数
Returns:
所有结果列表
"""
results = []
with ThreadPoolExecutor(max_workers=n_threads) as executor:
# 提交所有任务
future_to_group = {
executor.submit(post_to_summarizer, group, service_url, workspace_id): i
for i, group in enumerate(grouped_trajectories)
}
# 收集结果
for future in as_completed(future_to_group):
group_index = future_to_group[future]
try:
result = future.result()
result["group_index"] = group_index
result["group_size"] = len(grouped_trajectories[group_index])
results.append(result)
print(f"✅ Group {group_index} processed: {result.get('experience_list', 0) if 'experience_list' in result else 'error'}")
except Exception as e:
error_result = {
"group_index": group_index,
"group_size": len(grouped_trajectories[group_index]),
"error": str(e)
}
results.append(error_result)
print(f"❌ Group {group_index} failed: {e}")
return results
def main():
"""
主函数支持命令行参数
"""
parser = argparse.ArgumentParser(description='Convert JSONL to experiences using experience maker service')
parser.add_argument('--jsonl_file', type=str, required=True, help='Path to the JSONL file')
parser.add_argument('--service_url', type=str, default='http://localhost:8001', help='Experience maker service URL')
parser.add_argument('--workspace_id', type=str, required=True, help='Workspace ID for the experience')
parser.add_argument('--output_file', type=str, help='Output file to save results (optional)')
parser.add_argument('--n_threads', type=int, default=4, help='Number of threads for processing')
args = parser.parse_args()
print(f"Processing JSONL file: {args.jsonl_file}")
print(f"Service URL: {args.service_url}")
print(f"Workspace ID: {args.workspace_id}")
print(f"Threads: {args.n_threads}")
# 读取JSONL文件
try:
with open(args.jsonl_file, "r") as f:
data = [json.loads(line) for line in f]
print(f"Loaded {len(data)} entries from JSONL file")
except Exception as e:
print(f"Error reading JSONL file: {e}")
return
# 分组处理
grouped_trajectories = group_trajectories_by_task_id(data)
print(f"Total groups: {len(grouped_trajectories)}")
# 多线程处理
results = process_trajectories_with_threads(
grouped_trajectories,
args.service_url,
args.workspace_id,
n_threads=args.n_threads
)
print(f"Processed {len(results)} groups")
# 统计结果
success_count = sum(1 for r in results if 'error' not in r)
error_count = len(results) - success_count
total_experiences = sum(len(r.get('experiences', [])) for r in results if 'experiences' in r)
print(f"✅ Success: {success_count}")
print(f"❌ Errors: {error_count}")
print(f"📊 Total experiences created: {total_experiences}")
# 保存结果到文件
if args.output_file:
try:
summary = {
"workspace_id": args.workspace_id,
"jsonl_file": args.jsonl_file,
"total_groups": len(grouped_trajectories),
"success_count": success_count,
"error_count": error_count,
"total_experiences": total_experiences,
"results": results
}
with open(args.output_file, 'w') as f:
json.dump(summary, f, indent=2)
print(f"Results saved to: {args.output_file}")
except Exception as e:
print(f"Error saving results: {e}")
# 保持原有的使用示例(向后兼容)
if __name__ == "__main__":
# 检查是否有命令行参数
import sys
if len(sys.argv) > 1:
# 使用新的命令行接口
main()
else:
# 保持原有的行为(向后兼容)
print("Running in compatibility mode...")
with open("exp_result/qwen-max-2025-01-25/no_think/bfcl-multi-turn-base-train50_wo-exp.jsonl", "r") as f:
data = [json.loads(line) for line in f]
# 分组
grouped_trajectories = group_trajectories_by_task_id(data)
print(f"Total groups: {len(grouped_trajectories)}")
results = process_trajectories_with_threads(
grouped_trajectories,
"http://localhost:8001",
"bfcl_train50_qwen_max_2025_01_25_extract_compare_validate",
n_threads=4
)
print(f"Processed {len(results)} groups")

View file

@ -0,0 +1,223 @@
import json
import requests
import argparse
from pathlib import Path
from typing import List, Dict, Any
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor, as_completed
def load_task_case(data_path: str, task_id: str | None) -> Dict[str, Any]:
"""
load training cases by id
"""
if not Path(data_path).exists():
raise FileNotFoundError(f"BFCL data file '{data_path}' not found")
if task_id is None:
raise ValueError("task_id is required")
with open(data_path, "r", encoding="utf-8") as f:
if str(task_id).isdigit():
idx = int(task_id)
for line_no, line in enumerate(f):
if line_no == idx:
return json.loads(line)
raise ValueError(f"Task case index {idx} not found in {data_path}")
else:
for line in f:
data = json.loads(line)
if data.get("id") == task_id:
return data
raise ValueError(f"Task case id '{task_id}' not found in {data_path}")
def get_tool_prompt(tools):
tool_prompt = "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>"
for tool in tools:
tool_prompt += "\n" + json.dumps(tool)
tool_prompt += "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call>"
return tool_prompt
def group_trajectories_by_task_id(jsonl_entries: List[Dict[str, Any]]) -> List[List[Any]]:
"""
group trajectories by task_id
Args:
jsonl_entries: JSONL entry list
Returns:
List[List[Any]]: trajectory list grouped by task_id
"""
grouped = defaultdict(list)
for entry in jsonl_entries:
task_id = entry.get("task_id", "")
taks_case = load_task_case("data/multiturn_data_base.jsonl", task_id)
tools = taks_case.get("tools", [{}])
from bfcl_utils import extract_tool_schema
tool_schema = extract_tool_schema(tools)
entry["task_history"][0]["content"] += get_tool_prompt(tool_schema)
grouped[task_id].append(entry)
# retain only the two with the highest and lowest rewards
filtered_groups = []
for key, trajectories in grouped.items():
if len(trajectories) == 1:
# when only one trajectory, retain it
filtered_groups.append(trajectories)
elif len(trajectories) == 2:
# when there are two trajectories, retain them
filtered_groups.append(trajectories)
else:
# when there are more than two trajectories, choose the two with the highest and lowest rewards
trajectories.sort(key=lambda t: t["reward"])
min_reward_traj = trajectories[0] # highest reward
max_reward_traj = trajectories[-1] # lowest reward
filtered_groups.append([min_reward_traj, max_reward_traj])
return filtered_groups
def post_to_summarizer(trajectories: List[Any], service_url: str, workspace_id: str) -> Dict[str, Any]:
trajectory_dicts = [{
"task_id": traj["task_id"],
"messages": traj["task_history"],
"score": traj["reward"]
} for traj in trajectories]
request_data = {
"trajectories": trajectory_dicts,
"workspace_id": workspace_id
}
try:
response = requests.post(f"{service_url}/summary_task_memory", json=request_data)
response.raise_for_status()
return response.json()
except Exception as e:
return {"error": str(e), "trajectories_count": len(trajectories)}
def process_trajectories_with_threads(grouped_trajectories: List[List[Any]],
service_url: str,
workspace_id: str,
n_threads: int = 4) -> List[Dict[str, Any]]:
"""
use threads to process trajectories
Args:
grouped_trajectories: group trajectory list by task_id
service_url: memory summarizer service URL
workspace_id: workspace ID
n_threads: number of threads
Returns:
all results
"""
results = []
with ThreadPoolExecutor(max_workers=n_threads) as executor:
future_to_group = {
executor.submit(post_to_summarizer, group, service_url, workspace_id): i
for i, group in enumerate(grouped_trajectories)
}
for future in as_completed(future_to_group):
group_index = future_to_group[future]
try:
result = future.result()
result["group_index"] = group_index
result["group_size"] = len(grouped_trajectories[group_index])
results.append(result)
print(f"✅ Group {group_index} processed: {result["metadata"].get('memory_list', 0) if 'memory_list' in result["metadata"] else 'error'}")
except Exception as e:
error_result = {
"group_index": group_index,
"group_size": len(grouped_trajectories[group_index]),
"error": str(e)
}
results.append(error_result)
print(f"❌ Group {group_index} failed: {e}")
return results
def main():
parser = argparse.ArgumentParser(description='Convert JSONL to memories using ReMe service')
parser.add_argument('--jsonl_file', type=str, required=True, help='Path to the JSONL file')
parser.add_argument('--service_url', type=str, default='http://localhost:8001', help='ReMe service URL')
parser.add_argument('--workspace_id', type=str, required=True, help='Workspace ID for the task memory pool')
parser.add_argument('--output_file', type=str, help='Output file to save results (optional)')
parser.add_argument('--n_threads', type=int, default=4, help='Number of threads for processing')
args = parser.parse_args()
print(f"Processing JSONL file: {args.jsonl_file}")
print(f"Service URL: {args.service_url}")
print(f"Workspace ID: {args.workspace_id}")
print(f"Threads: {args.n_threads}")
with open(args.jsonl_file, "r") as f:
data = [json.loads(line) for line in f]
print(f"Loaded {len(data)} entries from JSONL file")
grouped_trajectories = group_trajectories_by_task_id(data)
print(f"Total groups: {len(grouped_trajectories)}")
results = process_trajectories_with_threads(
grouped_trajectories,
args.service_url,
args.workspace_id,
n_threads=args.n_threads
)
print(f"Processed {len(results)} groups")
success_count = sum(1 for r in results if 'error' not in r)
error_count = len(results) - success_count
total_memories = sum(len(r["metadata"].get('memory_list', [])) for r in results if 'memory_list' in r["metadata"])
print(f"✅ Success: {success_count}")
print(f"❌ Errors: {error_count}")
print(f"📊 Total task memories created: {total_memories}")
if args.output_file:
try:
summary = {
"workspace_id": args.workspace_id,
"jsonl_file": args.jsonl_file,
"total_groups": len(grouped_trajectories),
"success_count": success_count,
"error_count": error_count,
"total_task_memories": total_memories,
"results": results
}
with open(args.output_file, 'w') as f:
json.dump(summary, f, indent=2)
print(f"Results saved to: {args.output_file}")
except Exception as e:
print(f"Error saving results: {e}")
if __name__ == "__main__":
import sys
if len(sys.argv) > 1:
main()
else:
print("Running in compatibility mode...")
with open("exp_result/qwen3-14b/no_think/bfcl-multi-turn-base_wo-exp.jsonl", "r") as f:
data = [json.loads(line) for line in f]
grouped_trajectories = group_trajectories_by_task_id(data)
print(f"Total groups: {len(grouped_trajectories)}")
results = process_trajectories_with_threads(
grouped_trajectories,
"http://localhost:8001",
"bfcl_test",
n_threads=4
)
print(f"Processed {len(results)} groups")

View file

@ -0,0 +1,27 @@
import json
with open("../../file_vector_store/bfcl_test.jsonl", 'r') as f:
bfcl = [json.loads(line) for line in f]
new_bfcl = []
for exp in bfcl:
new_exp = {}
new_exp["workspace_id"] = exp["workspace_id"]
new_exp["memory_id"] = exp["unique_id"]
new_exp["memory_type"] = exp["metadata"]["memory_type"]
new_exp["when_to_use"] = exp["content"]
new_exp["content"] = exp["metadata"]["content"]
new_exp["score"] = exp["metadata"]["score"]
new_exp["time_created"] = exp["metadata"]["time_created"]
new_exp["time_modified"] = exp["metadata"]["time_modified"]
new_exp["author"] = exp["metadata"]["author"]
new_exp["metadata"]= exp["metadata"]["metadata"]
new_bfcl.append(new_exp)
with open('../../library/bfcl_test.jsonl', 'w', encoding='utf-8') as f:
f.writelines(json.dumps(item, ensure_ascii=False) + '\n' for item in new_bfcl)

View file

@ -0,0 +1,5 @@
jinja2
loguru
openai
ray
pandas

117
cookbook/bfcl/run_bfcl.py Normal file
View file

@ -0,0 +1,117 @@
import os
import time
import ray
# from ray import logger
from loguru import logger
from dotenv import load_dotenv
load_dotenv("../../.env")
import json
from pathlib import Path
from bfcl_agent import BFCLAgent
def run_agent(dataset_name: str,
experiment_suffix: str,
max_workers: int,
num_runs: int = 4,
model_name: str = "qwen3-8b",
data_path: str = "data/multiturn_data_base_val.jsonl",
answer_path: Path = Path("data/possible_answer"),
use_memory: bool = False,
use_memory_addition: bool = True,
use_memory_deletion: bool = False,
delete_freq: int = 10,
freq_threshold: int = 5,
utility_threshold: float = 0.5,
enable_thinking: bool = False,
memory_base_url: str = "http://0.0.0.0:8001/",
memory_workspace_id: str = "bfcl_test"):
experiment_name = dataset_name + "_" + experiment_suffix
path: Path = Path(f"./exp_result/{model_name}/with_think" if enable_thinking else f"./exp_result/{model_name}/no_think")
path.mkdir(parents=True, exist_ok=True)
with open(data_path, "r", encoding="utf-8") as f:
task_ids = [json.loads(l)["id"] for l in f]
result: list = []
def dump_file():
with open(path / f"{experiment_name}.jsonl", "a") as f:
for x in result:
f.write(json.dumps(x) + "\n")
future_list: list = []
for i in range(max_workers):
actor = BFCLAgent.remote(
index=i,
task_ids=task_ids[i::max_workers],
experiment_name=experiment_name,
data_path=data_path,
answer_path=answer_path,
model_name=model_name,
num_runs=num_runs,
use_memory=use_memory,
use_memory_addition=use_memory_addition,
use_memory_deletion=use_memory_deletion,
delete_freq=delete_freq,
freq_threshold=freq_threshold,
utility_threshold=utility_threshold,
enable_thinking=enable_thinking,
memory_base_url=memory_base_url,
memory_workspace_id=memory_workspace_id
)
future = actor.execute.remote()
future_list.append(future)
time.sleep(1)
logger.info("submit complete")
for i, future in enumerate(future_list):
t_result = ray.get(future)
if t_result:
if isinstance(t_result, list):
result.extend(t_result)
else:
result.append(t_result)
logger.info(f"{i + 1}/{len(task_ids)} complete")
dump_file()
def main():
max_workers = 4
num_runs = 1
use_memory = False
use_memory_addition = False
use_memory_deletion = False
memory_base_url = "http://0.0.0.0:8001/"
memory_workspace_id = "bfcl_test"
if max_workers > 1:
ray.init(num_cpus=max_workers)
for run_id in range(num_runs):
run_agent(
dataset_name="bfcl-multi-turn-base",
experiment_suffix=f"wo-exp",
model_name="qwen3-8b",
max_workers=max_workers,
num_runs=1,
data_path="data/multiturn_data_base_val.jsonl",
answer_path=Path("data/possible_answer"),
enable_thinking=False,
use_memory=use_memory,
use_memory_addition=use_memory_addition,
use_memory_deletion=use_memory_deletion,
delete_freq=5,
freq_threshold=5,
utility_threshold=0.5,
memory_base_url=memory_base_url,
memory_workspace_id=memory_workspace_id,
)
if __name__ == "__main__":
main()

View file

@ -0,0 +1,159 @@
import json
from pathlib import Path
from collections import defaultdict
import pandas as pd
from loguru import logger
def calculate_best_at_k(scores: list, k: int) -> float:
"""
Calculate best@k
Divide scores into groups of size k, take the maximum value in each group,
then average these maximum values
Args:
scores: List of after_score values for all runs of a task
k: Group size
Returns:
best@k value
"""
if len(scores) % k != 0:
raise ValueError(f"Length of scores ({len(scores)}) must be divisible by k ({k})")
group_maxs = []
for i in range(0, len(scores), k):
group = scores[i:i + k]
group_maxs.append(max(group))
return sum(group_maxs) / len(group_maxs)
def calculate_pass_at_k(scores: list, k: int) -> float:
if len(scores) % k != 0:
raise ValueError(f"Length of scores ({len(scores)}) must be divisible by k ({k})")
group_maxs = []
for i in range(0, len(scores), k):
group = scores[i:i + k]
is_pass = 1.0 if max(group) >=1.0 else 0.0
group_maxs.append(is_pass)
return sum(group_maxs) / len(group_maxs)
def get_possible_k_values(total_runs: int) -> list:
"""
Get all possible k values (factors of total_runs)
Args:
total_runs: Total number of runs
Returns:
List of k values in descending order
"""
k_values = []
for k in range(1, total_runs + 1):
if total_runs % k == 0:
k_values.append(k)
return sorted(k_values, reverse=True) # Sort from large to small
def run_exp_statistic():
path: Path = Path(f"./exp_result/qwen3-8b/no_think")
# Store results for all experiments
all_results = {}
for file in [f for f in path.glob("*.jsonl")]:
# Group results by task_id
task_results = defaultdict(list)
print(file)
with open(file, "r") as f:
for line in f:
if not line.strip():
continue
data = json.loads(line)
if isinstance(data, list):
for part_data in data:
task_id = part_data["task_id"]
after_score = part_data["reward"]
task_results[task_id].append(after_score)
else:
task_id = data["task_id"]
after_score = data["reward"]
task_results[task_id].append(after_score)
if not task_results:
logger.warning(f"No valid data found in file {file}")
continue
# Check if each task has consistent number of runs
run_counts = [len(scores) for scores in task_results.values()]
if len(set(run_counts)) > 1:
logger.warning(f"Inconsistent number of runs for different tasks in file {file}: {set(run_counts)}")
continue
num_runs = run_counts[0]
logger.info(f"File {file}: {len(task_results)} tasks, {num_runs} runs per task")
# Get all possible k values
k_values = get_possible_k_values(num_runs)
logger.info(f"Calculable best@k values: {k_values}")
# Calculate various best@k values
file_results = {"file": file.name}
for k in k_values:
best_at_k_scores = []
pass_at_k_scores = []
for task_id, scores in task_results.items():
try:
best_k_score = calculate_best_at_k(scores, k)
pass_at_k_score = calculate_pass_at_k(scores, k)
pass_at_k_scores.append(pass_at_k_score)
best_at_k_scores.append(best_k_score)
except ValueError as e:
logger.error(f"Error calculating best@{k} for task {task_id}: {e}")
continue
if best_at_k_scores:
avg_best_at_k = sum(best_at_k_scores) / len(best_at_k_scores)
file_results[f"best@{k}"] = avg_best_at_k
logger.info(f"file={file.name} best@{k}={avg_best_at_k:.4f}")
if pass_at_k_scores:
avg_pass_at_k = sum(pass_at_k_scores) / len(pass_at_k_scores)
file_results[f"pass@{k}"] = avg_pass_at_k
logger.info(f"file={file.name} pass@{k}={avg_pass_at_k:.4f}")
all_results[file.name] = file_results
# Create and display table
if all_results:
df = pd.DataFrame(list(all_results.values()))
df = df.set_index('file')
# Sort columns by the number in column name (best@8, best@4, best@2, best@1)
# best_columns = [col for col in df.columns if col.startswith('best@')]
best_columns = [col for col in df.columns]
best_columns.sort(key=lambda x: x, reverse=False)
df = df[best_columns]
print("\n" + "=" * 80)
print("Experiment Results Summary Table")
print("=" * 80)
print(df.round(4))
print("=" * 80)
# Save table to CSV
output_path = path / "experiment_summary.csv"
df.to_csv(output_path)
logger.info(f"Results table saved to: {output_path}")
else:
logger.warning("No valid experiment results found")
if __name__ == "__main__":
run_exp_statistic()

View file

@ -0,0 +1,29 @@
import json
import random
import argparse
def split_jsonl(input_file, train_file, val_file, ratio=0.8):
with open(input_file, 'r', encoding='utf-8') as f:
data = [json.loads(line) for line in f]
random.shuffle(data)
split_idx = int(len(data) * ratio)
train_data = data[:split_idx]
val_data = data[split_idx:]
with open(train_file, 'w', encoding='utf-8') as f:
for item in train_data:
f.write(json.dumps(item, ensure_ascii=False) + '\n')
with open(val_file, 'w', encoding='utf-8') as f:
for item in val_data:
f.write(json.dumps(item, ensure_ascii=False) + '\n')
if __name__ == "__main__":
parser = argparse.ArgumentParser(description='Split JSONL file into train and validation sets.')
parser.add_argument('--input', required=True, help='Path to input JSONL file')
parser.add_argument('--train', required=True, help='Path to output train file')
parser.add_argument('--val', required=True, help='Path to output validation file')
parser.add_argument('--ratio', type=float, default=0.5, help='Train ratio (default: 0.8)')
args = parser.parse_args()
split_jsonl(args.input, args.train, args.val, args.ratio)

View file

View file

@ -0,0 +1,47 @@
frozenlake_sys_prompt_no_slippery: |
You are an AI agent playing FrozenLake game. Your goal is to navigate from Start (S) to Goal (G) while avoiding Holes (H).
Game Rules:
- S: Starting position (safe)
- F: Frozen surface (safe to walk on)
- H: Hole (you fall in and lose)
- G: Goal (you win!)
- []: Your current position
Actions:
- 0: Move LEFT
- 1: Move DOWN
- 2: Move RIGHT
- 3: Move UP
Your task: Analyze the current state and choose the best action (0-3) to reach the Goal while avoiding Holes.
While ensuring a safe arrival at the goal, you should aim to complete the task in as few steps as possible.
Think step by step, and respond with your thoughts and then clearly state your action as a number (0-3) in format {"action":"(0-3)"}.
frozenlake_sys_prompt_slippery: |
You are an AI agent playing FrozenLake game. Your goal is to navigate from Start (S) to Goal (G) while avoiding Holes (H).
Game Rules:
- S: Starting position (safe)
- F: Frozen surface (safe to walk on)
- H: Hole (you fall in and lose)
- G: Goal (you win!)
- []: Your current position
Actions:
- 0: Move LEFT
- 1: Move DOWN
- 2: Move RIGHT
- 3: Move UP
The ice is slippery, so you might not always move in the intended direction!
you will move in intended direction with probability of 1/3 else will move in either perpendicular direction with equal probability of 1/3 in both directions.
For example, if action is left, then:
- P(move left)=1/3
- P(move up)=1/3
- P(move down)=1/3
Your task: Analyze the current state and choose the best action (0-3) to reach the Goal while avoiding Holes.
While ensuring a safe arrival at the goal, you should aim to complete the task in as few steps as possible.
Think step by step, and respond with your thoughts and then clearly state your action as a number (0-3) in format {{"action":"(0-3)"}}.

View file

@ -0,0 +1,373 @@
import os
import re
import time
import json
import ray
import requests
import random
from typing import List, Dict, Any, Optional
from dataclasses import dataclass
import numpy as np
import gymnasium as gym
from gymnasium.envs.toy_text.frozen_lake import generate_random_map
from openai import OpenAI
from loguru import logger
import yaml
from dotenv import load_dotenv
from tqdm import tqdm
load_dotenv("../../.env")
@dataclass
class GameResult:
task_id: str
run_id: int
experiment_name: str
success: bool
steps: int
reward: float
trajectory: List[Dict]
map_config: Dict[str, Any]
@ray.remote
class FrozenLakeReactAgent:
"""A ReAct Agent for FrozenLake game with task memory learning."""
def __init__(self,
index: int,
task_configs: List[Dict],
experiment_name: str,
model_name: str = "qwen3-8b",
temperature: float = 0.7,
max_steps: int = 50,
num_runs: int = 1,
use_task_memory: bool = False,
make_task_memory: bool = False):
self.index = index
self.task_configs = task_configs
self.experiment_name = experiment_name
self.model_name = model_name
self.temperature = temperature
self.max_steps = max_steps
self.num_runs = num_runs
self.use_task_memory = use_task_memory
self.make_task_memory = make_task_memory
self.llm_client = OpenAI()
self.action_map = {0: "LEFT", 1: "DOWN", 2: "RIGHT", 3: "UP"}
# Load prompts
self.prompts = self._load_prompts()
def _load_prompts(self) -> Dict[str, str]:
"""Load prompts from yaml file"""
try:
with open("frozenlake_prompts.yaml", 'r', encoding='utf-8') as f:
return yaml.safe_load(f)
except FileNotFoundError:
logger.warning("Prompt file not found, using default prompts")
raise FileNotFoundError("Prompt file not found. Please check your current path (should be ./cook/frozenlake) and try again.")
def call_llm(self, messages: List[Dict]) -> str:
"""Call LLM with retry logic"""
for i in range(5):
try:
response = self.llm_client.chat.completions.create(
model=self.model_name,
messages=messages,
temperature=self.temperature,
extra_body={"enable_thinking": False},
seed=0
)
return response.choices[0].message.content
except Exception as e:
logger.warning(f"LLM call failed (attempt {i + 1}): {e}")
time.sleep(1 + i * 2)
return "LLM call failed"
def observe_state(self, env, observation: int) -> str:
"""Convert environment observation to text description"""
desc = env.unwrapped.desc
nrow, ncol = desc.shape
# Convert to string grid
grid = [[cell.decode('utf-8') for cell in row] for row in desc]
# Get current position
row, col = observation // ncol, observation % ncol
# Create visual representation
state_text = "Current State:\n"
for i in range(nrow):
for j in range(ncol):
if i == row and j == col:
state_text += f"[{grid[i][j]}]"
else:
state_text += f" {grid[i][j]} "
state_text += "\n"
state_text += "\nLegend: S=Start, F=Frozen, H=Hole, G=Goal, []=Your Position"
return state_text
def build_system_prompt(self, is_slippery: bool) -> str:
"""Build system prompt based on game configuration"""
if is_slippery:
return self.prompts["frozenlake_sys_prompt_slippery"]
else:
return self.prompts["frozenlake_sys_prompt_no_slippery"]
def get_task_memory(self, map_desc: str, is_slippery: bool) -> str:
"""Retrieve relevant task memory from task memory service"""
if not self.use_task_memory:
return ""
try:
query = f"FrozenLake game map: {map_desc}, slippery: {is_slippery}"
base_url = "http://0.0.0.0:8002/"
workspace_id = self.experiment_name
response = requests.post(
url=base_url + "retrieve_task_memory",
json={
"workspace_id": workspace_id,
"query": query,
},
timeout=60
)
if response.status_code == 200:
data = response.json()
return data.get("answer", "")
else:
logger.warning(f"Task memory retrieval failed: {response.status_code}")
return ""
except Exception as e:
logger.warning(f"Failed to get task memory: {e}")
return ""
def action_parser(self, response: str) -> int:
"""Parse action from LLM response"""
# Look for {"action":"X"} pattern
patterns = [
r'["\']action["\']\s*:\s*["\']([0-3])["\']',
r'"action"\s*:\s*"([0-3])"',
r"'action'\s*:\s*'([0-3])'",
r'\baction["\']?\s*[:=]\s*["\']?([0-3])'
]
for pattern in patterns:
match = re.search(pattern, response)
if match:
action = int(match.group(1))
if 0 <= action <= 3:
return action
# Random fallback
action = random.randint(0, 3)
logger.warning(f"Could not parse action from response, using random: {action}")
return action
def run_single_episode(self, task_config: Dict, run_id: int) -> GameResult:
"""Run a single episode of the game"""
map_size = task_config.get("map_size", 4)
is_slippery = task_config.get("is_slippery", True)
map_desc = task_config.get("map_desc", None)
# Create environment
env_kwargs = {
"render_mode": None,
"is_slippery": is_slippery,
}
if map_desc is not None:
env_kwargs["desc"] = map_desc
else:
env_kwargs["desc"] = generate_random_map(size=map_size)
env = gym.make("FrozenLake-v1", **env_kwargs)
# Get map description for task memory
map_str = '\n'.join([''.join([cell.decode('utf-8') for cell in row])
for row in env.unwrapped.desc])
# Build messages
system_prompt = self.build_system_prompt(is_slippery)
task_memory = self.get_task_memory(map_str, is_slippery)
messages = [{"role": "system", "content": system_prompt}]
if task_memory:
memory_content = f"Here are some relevant tips from previous successful games:\n\n{task_memory}\n\nUse these tips to help you succeed."
messages.append({"role": "user", "content": memory_content})
messages.append(
{"role": "assistant", "content": "I'll use these tips to navigate the frozen lake successfully."})
# Initialize game
observation, info = env.reset()
trajectory = []
# Add initial state
initial_state = self.observe_state(env, observation)
messages.append({"role": "user", "content": initial_state})
success = False
total_reward = 0
for step in range(self.max_steps):
# Get action from LLM
response = self.call_llm(messages)
logger.info(response)
action = self.action_parser(response)
messages.append({"role": "assistant", "content": response})
# Take action
next_observation, reward, terminated, truncated, info = env.step(action)
total_reward += reward
done = terminated or truncated
# Record trajectory step
trajectory.append({
"step": step,
"state": observation,
"action": action,
"action_name": self.action_map[action],
"reward": reward,
"next_state": next_observation,
"done": done,
"llm_response": response
})
if done:
if terminated and reward > 0:
success = True
result_msg = f"Success! You reached the goal in {step + 1} steps!"
else:
result_msg = f"Game over! You fell into a hole or ran out of time."
messages.append({"role": "user", "content": result_msg})
break
else:
# Continue game
next_state = self.observe_state(env, next_observation)
step_msg = f"Step {step + 1}: You moved {self.action_map[action]}. Reward: {reward}\n{next_state}"
messages.append({"role": "user", "content": step_msg})
observation = next_observation
env.close()
# Create result
map_id = task_config.get("map_id", f"unknown_{self.index}_{run_id}")
task_id = f"{task_config.get('task_type', 'test')}_map{map_id}_{run_id}"
result = GameResult(
task_id=task_id,
run_id=run_id,
experiment_name=self.experiment_name,
success=success,
steps=len(trajectory),
reward=total_reward,
trajectory=trajectory,
map_config={
"map_desc": map_str,
"map_id": map_id,
"is_slippery": is_slippery,
"map_size": map_size,
"use_task_memory": self.use_task_memory
}
)
return result, messages
def save_task_memory(self, results: List[GameResult], messages_list: List[List[Dict]]):
"""Save successful trajectories as task memory"""
if not self.make_task_memory:
return
trajectories = []
for result, messages in zip(results, messages_list):
if result.success:
# Create trajectory for task memory service
traj = {
"messages": messages,
"score": 1.0, # Success
}
trajectories.append(traj)
else:
traj = {
"messages": messages,
"score": 0.0, # Failure
}
trajectories.append(traj)
if trajectories:
try:
base_url = "http://0.0.0.0:8002/"
workspace_id = self.experiment_name
response = requests.post(
url=base_url + "summary_task_memory",
json={
"workspace_id": workspace_id,
"trajectories": trajectories
},
timeout=300
)
if response.status_code == 200:
logger.info(f"Saved {len(trajectories)} trajectories as task memory")
else:
logger.warning(f"Failed to save task memory: {response.status_code}")
except Exception as e:
logger.error(f"Error saving task memory: {e}")
def execute(self) -> List[Dict]:
"""Execute all tasks"""
all_results = []
all_messages = []
for task_index, task_config in tqdm(enumerate(self.task_configs), desc="Processing tasks:"):
for run_id in range(self.num_runs):
logger.info(f"Ray {self.index}, Task {task_index}, Run {run_id}")
result, messages = self.run_single_episode(task_config, run_id)
all_results.append(result)
all_messages.append(messages)
# Convert result to dict for JSON serialization
result_dict = {
"task_id": result.task_id,
"run_id": result.run_id,
"experiment_name": result.experiment_name,
"task_completed": result.success,
"success": result.success,
"steps": result.steps,
"reward": result.reward,
"map_config": result.map_config,
"trajectory": result.trajectory
}
all_results[-1] = result_dict
# Save task memory if needed
if self.make_task_memory:
# Convert back to GameResult objects for task memory saving
game_results = []
for i, result_dict in enumerate(all_results):
game_result = GameResult(
task_id=result_dict["task_id"],
run_id=result_dict["run_id"],
experiment_name=result_dict["experiment_name"],
success=result_dict["success"],
steps=result_dict["steps"],
reward=result_dict["reward"],
trajectory=result_dict["trajectory"],
map_config=result_dict["map_config"]
)
game_results.append(game_result)
self.save_task_memory(game_results, all_messages)
return all_results

View file

@ -0,0 +1,121 @@
#!/usr/bin/env python3
"""
Map Management Tool - Pre-generate and manage test maps
"""
import json
import numpy as np
from pathlib import Path
from typing import List, Optional, Dict, Any
from loguru import logger
from gymnasium.envs.toy_text.frozen_lake import generate_random_map
class MapManager:
"""Map Manager - pre-generating, storing and loading test maps"""
def __init__(self, data_dir: str = "./map/"):
self.data_dir = Path(data_dir)
self.data_dir.mkdir(parents=True, exist_ok=True)
def generate_test_maps(self, num_maps: int, map_size: int = 4,
base_seed: int = 10000) -> str:
"""
Generate test map collection and save
Args:
num_maps: Number of maps to generate
map_size: Map size
base_seed: Base random seed
Returns:
Path of saved file
"""
logger.info(f"🗺️ Generating {num_maps} test maps (size={map_size})")
maps_data = []
for i in range(num_maps):
seed = base_seed + i
np.random.seed(seed)
map_desc = generate_random_map(size=map_size)
maps_data.append({
"map_id": i,
"seed": seed,
"map_size": map_size,
"map_desc": map_desc # Convert to list for JSON serialization
})
# Save to file
filename = f"test_maps_{num_maps}_{map_size}x{map_size}.jsonl"
filepath = self.data_dir / filename
with open(filepath, "w", encoding="utf-8") as f:
for map_data in maps_data:
f.write(json.dumps(map_data, ensure_ascii=False) + "\n")
logger.info(f"✅ Test maps saved to {filepath}")
return str(filepath)
def load_test_maps(self, filepath: str) -> List[Dict[str, Any]]:
"""
Load test maps
Args:
filepath: Map file path
Returns:
Map data list
"""
if not Path(filepath).exists():
raise FileNotFoundError(f"Map file not found: {filepath}")
maps_data = []
with open(filepath, "r", encoding="utf-8") as f:
for line in f:
if line.strip():
map_data = json.loads(line)
# Convert list back to numpy array
maps_data.append(map_data)
logger.info(f"📖 Loaded {len(maps_data)} test maps from {filepath}")
return maps_data
def get_map_by_index(self, maps_data: List[Dict], index: int) -> Optional[list]:
"""Get map by index"""
if 0 <= index < len(maps_data):
return maps_data[index]["map_desc"]
return None
def get_or_create_test_maps(self, num_maps: int, map_size: int = 4) -> List[Dict[str, Any]]:
"""
Get or create test maps
If file exists and has sufficient quantity, load directly; otherwise regenerate
"""
filename = f"test_maps_{num_maps}_{map_size}x{map_size}.jsonl"
filepath = self.data_dir / filename
if filepath.exists():
try:
maps_data = self.load_test_maps(str(filepath))
if len(maps_data) >= num_maps:
logger.info(f"✅ Using existing test maps: {filepath}")
return maps_data[:num_maps] # Return required number of maps
except Exception as e:
logger.warning(f"⚠️ Failed to load existing maps: {e}, regenerating...")
# File doesn't exist or insufficient quantity, regenerate
self.generate_test_maps(num_maps, map_size)
return self.load_test_maps(str(filepath))
if __name__ == "__main__":
# Usage example
manager = MapManager()
# Generate 100 4x4 test maps
manager.generate_test_maps(num_maps=100, map_size=4)
# Load and view the first map
maps = manager.load_test_maps("./map/test_maps_100_4x4.jsonl")
print(f"First map:\n{maps[0]['map_desc']}")

View file

@ -0,0 +1,370 @@
import json
import pandas as pd
from pathlib import Path
from collections import defaultdict
from typing import Dict, List, Tuple
from loguru import logger
def calculate_best_at_k(scores: List[float], k: int) -> float:
"""
Calculate best@k metric.
Divide scores into groups of size k, take the maximum value in each group,
then average these maximum values.
Args:
scores: List of success scores (0 or 1) for all runs of a task
k: Group size
Returns:
best@k value
"""
if len(scores) % k != 0:
raise ValueError(f"Length of scores ({len(scores)}) must be divisible by k ({k})")
group_maxs = []
for i in range(0, len(scores), k):
group = scores[i:i + k]
group_maxs.append(max(group))
return sum(group_maxs) / len(group_maxs)
def get_possible_k_values(total_runs: int) -> List[int]:
"""Get all possible k values (divisors of total_runs)"""
k_values = []
for k in range(1, total_runs + 1):
if total_runs % k == 0:
k_values.append(k)
return sorted(k_values, reverse=True)
def parse_task_config(task_id: str, map_config: Dict) -> Tuple[str, bool, bool]:
"""
Parse task configuration from task_id and map_config.
Returns:
(condition, is_slippery, use_experience)
"""
is_slippery = map_config.get("is_slippery", True)
use_experience = map_config.get("use_experience", False)
# Create condition string
slip_str = "slippery" if is_slippery else "no_slip"
exp_str = "with_exp" if use_experience else "no_exp"
condition = f"{slip_str}_{exp_str}"
return condition, is_slippery, use_experience
def analyze_frozenlake_results():
"""Analyze FrozenLake experiment results"""
path = Path("./exp_result")
if not path.exists():
logger.error("Experiment results directory not found!")
return
all_results = {}
# Process all result files
for file in path.glob("*test*.jsonl"):
logger.info(f"Processing {file.name}")
# Group results by condition and map
condition_results = defaultdict(lambda: defaultdict(list))
with open(file, "r") as f:
for line in f:
if not line.strip():
continue
try:
data = json.loads(line)
if isinstance(data, list):
for item in data:
process_single_result(item, condition_results)
else:
process_single_result(data, condition_results)
except json.JSONDecodeError as e:
logger.warning(f"Invalid JSON in {file.name}: {e}")
continue
if not condition_results:
logger.warning(f"No valid data found in {file.name}")
continue
# Calculate metrics for this file
file_metrics = calculate_file_metrics(condition_results, file.name)
all_results[file.name] = file_metrics
# Generate comprehensive report
if all_results:
generate_analysis_report(all_results)
else:
logger.warning("No valid results found!")
def process_single_result(data: Dict, condition_results: Dict):
"""Process a single result entry"""
map_config = data.get("map_config", {})
task_id = data.get("task_id", "unknown")
success = data.get("success", False)
# Parse condition
condition, is_slippery, use_experience = parse_task_config(task_id, map_config)
# Extract map identifier - prefer map_id from map_config
map_id = map_config.get("map_id", "unknown")
if map_id == "unknown" and "test_map" in task_id:
# Fallback to parsing from task_id
parts = task_id.split("_")
for part in parts:
if part.startswith("map"):
try:
# Extract number from "mapXX"
map_num = ''.join(filter(str.isdigit, part))
if map_num:
map_id = int(map_num)
break
except:
pass
# Store result
success_score = 1.0 if success else 0.0
condition_results[condition][f"map_{map_id}"].append(success_score)
def calculate_file_metrics(condition_results: Dict, filename: str) -> Dict:
"""Calculate metrics for a single file"""
file_metrics = {"file": filename}
for condition, map_results in condition_results.items():
condition_scores = []
# Collect all scores for this condition
for map_id, scores in map_results.items():
condition_scores.extend(scores)
if not condition_scores:
continue
# Check if all maps have the same number of runs
run_counts = [len(scores) for scores in map_results.values()]
if len(set(run_counts)) > 1:
logger.warning(f"Inconsistent runs for {condition}: {set(run_counts)}")
continue
num_runs = run_counts[0] if run_counts else 0
if num_runs == 0:
continue
# Calculate overall success rate
overall_success = sum(condition_scores) / len(condition_scores)
file_metrics[f"{condition}_success_rate"] = overall_success
# Calculate best@k metrics
k_values = get_possible_k_values(num_runs)
for k in k_values:
try:
# Calculate best@k for each map, then average
map_best_k_scores = []
for map_id, scores in map_results.items():
map_best_k = calculate_best_at_k(scores, k)
map_best_k_scores.append(map_best_k)
avg_best_k = sum(map_best_k_scores) / len(map_best_k_scores)
file_metrics[f"{condition}_best@{k}"] = avg_best_k
except ValueError as e:
logger.warning(f"Error calculating best@{k} for {condition}: {e}")
# Map-level analysis
map_success_rates = {}
for map_id, scores in map_results.items():
map_success_rate = sum(scores) / len(scores)
map_success_rates[map_id] = map_success_rate
file_metrics[f"{condition}_map_details"] = map_success_rates
logger.info(f"{filename} - {condition}: {overall_success:.3f} success rate, "
f"{len(map_results)} maps, {num_runs} runs each")
return file_metrics
def generate_analysis_report(all_results: Dict):
"""Generate comprehensive analysis report"""
logger.info("Generating comprehensive analysis report...")
# 1. Create summary table
summary_data = []
for file_name, metrics in all_results.items():
row = {"file": file_name}
# Extract success rates and best@k metrics
for key, value in metrics.items():
if key != "file" and not key.endswith("_map_details"):
row[key] = value
summary_data.append(row)
if summary_data:
df_summary = pd.DataFrame(summary_data)
df_summary = df_summary.set_index('file')
print("\n" + "=" * 100)
print("FROZENLAKE EXPERIMENT RESULTS SUMMARY")
print("=" * 100)
print(df_summary.round(4))
print("=" * 100)
# Save summary table
output_path = Path("./exp_result") / "frozenlake_summary.csv"
df_summary.to_csv(output_path)
logger.info(f"Summary table saved to: {output_path}")
# 2. Condition comparison
print("\n" + "=" * 80)
print("CONDITION COMPARISON")
print("=" * 80)
condition_comparison = defaultdict(list)
for file_name, metrics in all_results.items():
for key, value in metrics.items():
if "_success_rate" in key:
condition = key.replace("_success_rate", "")
condition_comparison[condition].append(value)
# Calculate average performance per condition
condition_avg = {}
for condition, scores in condition_comparison.items():
if scores:
avg_score = sum(scores) / len(scores)
condition_avg[condition] = avg_score
print(f"{condition:20s}: {avg_score:.4f}{pd.Series(scores).std():.4f})")
# 3. Experience effect analysis
print("\n" + "=" * 80)
print("EXPERIENCE EFFECT ANALYSIS")
print("=" * 80)
experience_analysis = analyze_experience_effect(condition_avg)
for analysis_line in experience_analysis:
print(analysis_line)
# 4. Map difficulty analysis
print("\n" + "=" * 80)
print("MAP DIFFICULTY ANALYSIS")
print("=" * 80)
map_analysis = analyze_map_difficulty(all_results)
for map_id, difficulty in map_analysis.items():
print(f"{map_id:10s}: {difficulty:.4f} average success rate")
# 5. Detailed statistics
print("\n" + "=" * 80)
print("DETAILED STATISTICS")
print("=" * 80)
generate_detailed_stats(all_results)
def analyze_experience_effect(condition_avg: Dict[str, float]) -> List[str]:
"""Analyze the effect of experience on performance"""
analysis = []
# Compare with/without experience for each slippery condition
slippery_no_exp = condition_avg.get("slippery_no_exp", 0)
slippery_with_exp = condition_avg.get("slippery_with_exp", 0)
no_slip_no_exp = condition_avg.get("no_slip_no_exp", 0)
no_slip_with_exp = condition_avg.get("no_slip_with_exp", 0)
if slippery_no_exp > 0 and slippery_with_exp > 0:
improvement_slippery = (slippery_with_exp - slippery_no_exp) / slippery_no_exp * 100
analysis.append(f"Slippery condition - Experience effect: {improvement_slippery:+.1f}%")
analysis.append(f" Without exp: {slippery_no_exp:.4f}")
analysis.append(f" With exp: {slippery_with_exp:.4f}")
if no_slip_no_exp > 0 and no_slip_with_exp > 0:
improvement_no_slip = (no_slip_with_exp - no_slip_no_exp) / no_slip_no_exp * 100
analysis.append(f"No-slip condition - Experience effect: {improvement_no_slip:+.1f}%")
analysis.append(f" Without exp: {no_slip_no_exp:.4f}")
analysis.append(f" With exp: {no_slip_with_exp:.4f}")
# Overall experience effect
exp_conditions = [v for k, v in condition_avg.items() if "with_exp" in k]
no_exp_conditions = [v for k, v in condition_avg.items() if "no_exp" in k]
if exp_conditions and no_exp_conditions:
avg_with_exp = sum(exp_conditions) / len(exp_conditions)
avg_without_exp = sum(no_exp_conditions) / len(no_exp_conditions)
overall_improvement = (avg_with_exp - avg_without_exp) / avg_without_exp * 100
analysis.append(f"Overall experience effect: {overall_improvement:+.1f}%")
return analysis
def analyze_map_difficulty(all_results: Dict) -> Dict[str, float]:
"""Analyze difficulty of different maps"""
map_scores = defaultdict(list)
for file_name, metrics in all_results.items():
for key, value in metrics.items():
if key.endswith("_map_details") and isinstance(value, dict):
for map_id, success_rate in value.items():
map_scores[map_id].append(success_rate)
# Calculate average difficulty per map
map_difficulty = {}
for map_id, scores in map_scores.items():
if scores:
avg_success = sum(scores) / len(scores)
map_difficulty[map_id] = avg_success
# Sort by difficulty (hardest first)
return dict(sorted(map_difficulty.items(), key=lambda x: x[1]))
def generate_detailed_stats(all_results: Dict):
"""Generate detailed statistics"""
total_experiments = len(all_results)
total_conditions = set()
for metrics in all_results.values():
for key in metrics.keys():
if "_success_rate" in key:
condition = key.replace("_success_rate", "")
total_conditions.add(condition)
print(f"Total experiment files: {total_experiments}")
print(f"Total conditions tested: {len(total_conditions)}")
print(f"Conditions: {', '.join(sorted(total_conditions))}")
# Best performing conditions
all_success_rates = []
for metrics in all_results.values():
for key, value in metrics.items():
if "_success_rate" in key and isinstance(value, (int, float)):
all_success_rates.append((key.replace("_success_rate", ""), value))
if all_success_rates:
best_condition = max(all_success_rates, key=lambda x: x[1])
worst_condition = min(all_success_rates, key=lambda x: x[1])
print(f"Best performance: {best_condition[0]} ({best_condition[1]:.4f})")
print(f"Worst performance: {worst_condition[0]} ({worst_condition[1]:.4f})")
def main():
"""Main function for statistics analysis"""
logger.info("🔍 Starting FrozenLake Results Analysis")
analyze_frozenlake_results()
logger.info("📊 Analysis completed!")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,296 @@
import os
import time
import json
import ray
from pathlib import Path
from typing import List, Dict
import numpy as np
from loguru import logger
from gymnasium.envs.toy_text.frozen_lake import generate_random_map
from frozenlake_react_agent import FrozenLakeReactAgent
from map_manager import MapManager
def generate_training_configs(num_maps: int = 20, map_size: int = 4, is_slippery: bool=False) -> List[Dict]:
"""Generate random maps for training/task memory generation"""
configs = []
for i in range(num_maps):
# Generate both slippery and non-slippery versions
random_map = generate_random_map(size=map_size)
config = {
"task_type": "training",
"map_desc": random_map,
"map_size": map_size,
"is_slippery": is_slippery,
"task_id": f"train_{i}_{is_slippery}"
}
configs.append(config)
return configs
def generate_test_configs(num_test_maps: int = 100, is_slippery: bool = False) -> List[Dict]:
"""Generate test configurations using MapManager"""
logger.info(f"📋 Generating test configurations for {num_test_maps} maps")
# Initialize MapManager and get test maps
map_manager = MapManager()
maps_data = map_manager.get_or_create_test_maps(num_maps=num_test_maps, map_size=4)
configs = []
for map_data in maps_data:
map_desc = np.array([list(row) for row in map_data["map_desc"]], dtype='c')
map_id = map_data["map_id"]
for use_memory in [True, False]:
config = {
"task_type": "test",
"map_desc": map_desc,
"map_size": 4,
"is_slippery": is_slippery,
"use_task_memory": use_memory,
"map_id": map_id,
"task_id": f"test_map{map_id}_slip{is_slippery}_mem{use_memory}"
}
configs.append(config)
logger.info(f"✅ Generated {len(configs)} test configurations")
return configs
def train(experiment_name: str, max_workers: int = 2, num_runs: int = 3, num_training_maps= 15, is_slippery: bool= False) -> None:
"""Phase 1: Generate task memory from random maps"""
logger.info("🎯 Starting Training Phase - Generating Task Memory")
logger.info("=" * 60)
training_configs = generate_training_configs(num_maps=num_training_maps, map_size=4, is_slippery=is_slippery)
path = Path("./exp_result")
path.mkdir(parents=True, exist_ok=True)
results = []
def dump_results():
output_file = path / f"{experiment_name}_training.jsonl"
with open(output_file, "w") as f:
for result in results:
f.write(json.dumps(result) + "\n")
logger.info(f"Training results saved to {output_file}")
if max_workers > 1:
# Distributed training
future_list = []
for i in range(max_workers):
worker_configs = training_configs[i::max_workers]
if worker_configs: # Only create worker if it has tasks
agent = FrozenLakeReactAgent.remote(
index=i,
task_configs=worker_configs,
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=False, # No task memory in training phase
make_task_memory=True, # Generate task memory
)
future = agent.execute.remote()
future_list.append(future)
time.sleep(1)
logger.info(f"Started {len(future_list)} training workers")
for i, future in enumerate(future_list):
worker_results = ray.get(future)
if worker_results:
results.extend(worker_results)
logger.info(f"results: {results[0]}")
logger.info(f"Training worker {i + 1}/{len(future_list)} completed")
dump_results()
else:
# Single process training
agent = FrozenLakeReactAgent(
index=0,
task_configs=training_configs,
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=False,
make_task_memory=True
)
results = agent.execute()
dump_results()
# Calculate training statistics
successful_runs = [r for r in results if r["success"]]
total_runs = len(results)
success_rate = len(successful_runs) / total_runs if total_runs > 0 else 0
logger.info(f"Training completed: {len(successful_runs)}/{total_runs} successful ({success_rate:.2%})")
return results
def test(experiment_name: str, max_workers: int = 2, num_runs: int = 5, num_test_maps: int = 100, is_slippery: bool=False) -> None:
"""Phase 2: Test on fixed maps with/without task memory"""
logger.info("🧪 Starting Test Phase - Evaluating Performance")
logger.info(f"📊 Testing on {num_test_maps} maps with {num_runs} runs each")
logger.info("=" * 60)
test_configs = generate_test_configs(num_test_maps=num_test_maps, is_slippery=is_slippery)
path = Path("./exp_result")
path.mkdir(parents=True, exist_ok=True)
# Group configs by task memory usage for separate experiments
memory_configs = [c for c in test_configs if c.get("use_task_memory", False)]
no_memory_configs = [c for c in test_configs if not c.get("use_task_memory", False)]
logger.info(f"📝 Configs without task memory: {len(no_memory_configs)}")
logger.info(f"📝 Configs with task memory: {len(memory_configs)}")
def dump_results(suffix: str):
output_file = path / f"{experiment_name}_test_{suffix}.jsonl"
with open(output_file, "w") as f:
for result in all_results:
f.write(json.dumps(result) + "\n")
logger.info(f"💾 Test results saved to {output_file}")
# Test without task memory first
logger.info("🚫 Testing WITHOUT task memory...")
all_results = []
results_no_memory = run_test_configs(
configs=no_memory_configs,
experiment_name=experiment_name,
max_workers=max_workers,
num_runs=num_runs,
use_task_memory=False
)
all_results.extend(results_no_memory)
dump_results("no_memory")
# Test with task memory
logger.info("✅ Testing WITH task memory...")
all_results = []
results_with_memory = run_test_configs(
configs=memory_configs,
experiment_name=experiment_name,
max_workers=max_workers,
num_runs=num_runs,
use_task_memory=True
)
all_results.extend(results_with_memory)
dump_results("with_memory")
return all_results
def run_test_configs(configs: List[Dict], experiment_name: str, max_workers: int,
num_runs: int, use_task_memory: bool) -> List[Dict]:
"""Run a set of test configurations"""
results = []
if max_workers > 1:
future_list = []
for i in range(max_workers):
worker_configs = configs[i::max_workers]
if worker_configs:
agent = FrozenLakeReactAgent.remote(
index=i,
task_configs=worker_configs,
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=use_task_memory,
make_task_memory=False
)
future = agent.execute.remote()
future_list.append(future)
time.sleep(1)
for i, future in enumerate(future_list):
worker_results = ray.get(future)
if worker_results:
results.extend(worker_results)
logger.info(f"Test worker {i + 1}/{len(future_list)} completed")
else:
agent = FrozenLakeReactAgent(
index=0,
task_configs=configs,
experiment_name=experiment_name,
num_runs=num_runs,
use_task_memory=use_task_memory,
make_task_memory=False
)
results = agent.execute()
return results
def main():
"""Main execution function"""
experiment_name = "frozenlake_no_slippery"
max_workers = 4
training_runs = 4 # Runs per training map
num_training_maps = 50
test_runs = 1 # Runs per test configuration
num_test_maps = 100 # Number of test maps to use
is_slippery = False
# model_name = "qwen-max-latest"
# Initialize Ray if using multiple workers
if max_workers > 1:
ray.init(num_cpus=max_workers)
try:
# Phase 1: Training (Experience Generation)
logger.info("🚀 Starting FrozenLake Experiment")
logger.info(f"🎯 Experiment: {experiment_name}")
logger.info(f"🏃 Workers: {max_workers}")
logger.info(f"📊 Test maps: {num_test_maps}")
logger.info(f"🔄 Test runs per map: {test_runs}")
training_results = train(
experiment_name=experiment_name,
max_workers=max_workers,
num_runs=training_runs,
num_training_maps=num_training_maps,
is_slippery=is_slippery
)
# Wait a bit for task memory service to process
logger.info("⏰ Waiting for task memory service to process data...")
time.sleep(10)
# Phase 2: Testing (Performance Evaluation)
test_results = test(
experiment_name=experiment_name,
max_workers=max_workers,
num_runs=test_runs,
num_test_maps=num_test_maps,
is_slippery=is_slippery
)
# Summary
logger.info("🎉 Experiment completed!")
logger.info(f"📈 Training results: {len(training_results)} episodes")
logger.info(f"📈 Test results: {len(test_results)} episodes")
# Quick statistics
successful_training = sum(1 for r in training_results if r.get("success", False))
training_success_rate = successful_training / len(training_results) if training_results else 0
successful_test = sum(1 for r in test_results if r.get("success", False))
test_success_rate = successful_test / len(test_results) if test_results else 0
logger.info(f"📊 Training success rate: {training_success_rate:.2%}")
logger.info(f"📊 Test success rate: {test_success_rate:.2%}")
finally:
if max_workers > 1:
ray.shutdown()
if __name__ == "__main__":
main()

View file

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,123 @@
{
"answer": "",
"messages": [],
"success": true,
"metadata": {
"memory_list": [
{
"workspace_id": "personal_memory_demo",
"memory_id": "45b3d01c803a41fab029568ec289a82d",
"memory_type": "personal",
"when_to_use": "John Smith, 28, San Francisco, tech company",
"content": "user's name is John Smith, aged 28, works at a tech company in San Francisco",
"score": 0.0,
"time_created": "2025-09-06 23:44:34",
"time_modified": "2025-09-06 23:44:34",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "John Smith, 28, San Francisco, tech company",
"source_message": "My name is John Smith, I'm 28 years old, and I work at a tech company in San Francisco",
"observation_type": "personal_info"
},
"target": "user",
"reflection_subject": ""
},
{
"workspace_id": "personal_memory_demo",
"memory_id": "0918a87e11344b8981ee8588e2729d22",
"memory_type": "personal",
"when_to_use": "software engineer, backend, Python, Go",
"content": "user is a software engineer specializing in backend development using Python and Go",
"score": 0.0,
"time_created": "2025-09-06 23:44:34",
"time_modified": "2025-09-06 23:44:34",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "software engineer, backend, Python, Go",
"source_message": "I'm a software engineer, mainly doing backend development using Python and Go",
"observation_type": "personal_info"
},
"target": "user",
"reflection_subject": ""
},
{
"workspace_id": "personal_memory_demo",
"memory_id": "073ed894a05e43d5badb0fdc04401368",
"memory_type": "personal",
"when_to_use": "basketball, sci-fi movies, Dune Part 2",
"content": "user enjoys playing basketball and watching sci-fi movies, recently watched Dune Part 2",
"score": 0.0,
"time_created": "2025-09-06 23:44:34",
"time_modified": "2025-09-06 23:44:34",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "basketball, sci-fi movies, Dune Part 2",
"source_message": "I enjoy playing basketball and watching sci-fi movies. I recently watched Dune Part 2",
"observation_type": "personal_info"
},
"target": "user",
"reflection_subject": ""
},
{
"workspace_id": "personal_memory_demo",
"memory_id": "1c81c798bd5843debcf1e4b0dd393dc4",
"memory_type": "personal",
"when_to_use": "cat, Shadow, pet",
"content": "user has a 3-year-old cat named Shadow",
"score": 0.0,
"time_created": "2025-09-06 23:44:34",
"time_modified": "2025-09-06 23:44:34",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "cat, Shadow, pet",
"source_message": "I have a cat named Shadow who is 3 years old",
"observation_type": "personal_info"
},
"target": "user",
"reflection_subject": ""
},
{
"workspace_id": "personal_memory_demo",
"memory_id": "1c35f20456d847428653ddb24c03f61f",
"memory_type": "personal",
"when_to_use": "Japanese cuisine, sushi, ramen",
"content": "user is interested in Japanese cuisine, especially sushi and ramen",
"score": 0.0,
"time_created": "2025-09-06 23:44:34",
"time_modified": "2025-09-06 23:44:34",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "Japanese cuisine, sushi, ramen",
"source_message": "I'm really interested in Japanese cuisine, especially sushi and ramen",
"observation_type": "personal_info"
},
"target": "user",
"reflection_subject": ""
},
{
"workspace_id": "personal_memory_demo",
"memory_id": "7bb563c1e03f4c41a52000de7deb007e",
"memory_type": "personal",
"when_to_use": "Japan, Tokyo, Kyoto, travel plan",
"content": "user is planning a trip to Tokyo and Kyoto, Japan in October 2025",
"score": 0.0,
"time_created": "2025-09-06 23:44:31",
"time_modified": "2025-09-06 23:44:31",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"keywords": "Japan, Tokyo, Kyoto, travel plan",
"time_info": "October 2025",
"source_message": "I'm planning a trip to Japan next month, mainly to Tokyo and Kyoto",
"observation_type": "personal_info_with_time"
},
"target": "user",
"reflection_subject": ""
}
],
"deleted_memory_ids": [],
"update_result": {
"deleted_count": 0,
"inserted_count": 6
}
}
}

View file

@ -0,0 +1,87 @@
[
{
"workspace_id": "test_workspace",
"memory_id": "e1103be06ec24ffebf441257d385211b",
"memory_type": "task",
"when_to_use": "When analyzing a complex, multi-faceted company with historical, financial, operational, and competitive dimensions—especially in dynamic industries like tech or EVs.",
"content": "The agent successfully decomposed the broad query 'Analyze the company Tesla' into four distinct subtasks: (1) historical context, (2) business model and revenue streams, (3) financial performance, and (4) innovation and market position. Each subtask was addressed via targeted web searches using specific, focused queries that extracted precise, high-value information. The use of multiple search iterations with varying angles (e.g., 'Tesla innovation technology advancements 2024' vs. 'market position competitors electric vehicles 2024') ensured comprehensive coverage across different domains. This structured, layered approach prevented information overload while ensuring depth in each critical area.",
"score": 0.92,
"time_created": "2025-09-07 15:57:06",
"time_modified": "2025-09-07 15:57:06",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"when_to_use": "When analyzing a complex, multi-faceted company with historical, financial, operational, and competitive dimensions—especially in dynamic industries like tech or EVs.",
"experience": "The agent successfully decomposed the broad query 'Analyze the company Tesla' into four distinct subtasks: (1) historical context, (2) business model and revenue streams, (3) financial performance, and (4) innovation and market position. Each subtask was addressed via targeted web searches using specific, focused queries that extracted precise, high-value information. The use of multiple search iterations with varying angles (e.g., 'Tesla innovation technology advancements 2024' vs. 'market position competitors electric vehicles 2024') ensured comprehensive coverage across different domains. This structured, layered approach prevented information overload while ensuring depth in each critical area.",
"tags": [
"decomposition",
"multi-dimensional analysis",
"targeted search",
"information layering",
"business model",
"financials",
"competitive landscape"
],
"confidence": 0.9,
"step_type": "reasoning",
"tools_used": [
"web_search"
]
}
},
{
"workspace_id": "test_workspace",
"memory_id": "989fceaa659548d6b85740de02a3ea83",
"memory_type": "task",
"when_to_use": "When initial search results are insufficient or fragmented, especially for time-sensitive or evolving topics like AI advancements or quarterly financials.",
"content": "After receiving partial results from the first three searches, the agent proactively initiated two additional web searches to fill critical knowledge gaps: one on recent technological innovations (2024) and another on current market competition. These follow-up queries were highly specific and timed to capture up-to-date developments (e.g., FSD V12.5, Optimus robot production plans). This iterative search strategy allowed the agent to identify real-time trends and emerging strategic moves, which were essential for a forward-looking analysis. The ability to dynamically adjust the research plan based on incomplete early data is a key indicator of adaptive intelligence.",
"score": 0.85,
"time_created": "2025-09-07 15:57:06",
"time_modified": "2025-09-07 15:57:06",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"when_to_use": "When initial search results are insufficient or fragmented, especially for time-sensitive or evolving topics like AI advancements or quarterly financials.",
"experience": "After receiving partial results from the first three searches, the agent proactively initiated two additional web searches to fill critical knowledge gaps: one on recent technological innovations (2024) and another on current market competition. These follow-up queries were highly specific and timed to capture up-to-date developments (e.g., FSD V12.5, Optimus robot production plans). This iterative search strategy allowed the agent to identify real-time trends and emerging strategic moves, which were essential for a forward-looking analysis. The ability to dynamically adjust the research plan based on incomplete early data is a key indicator of adaptive intelligence.",
"tags": [
"iterative research",
"dynamic query refinement",
"real-time updates",
"gap detection",
"adaptive planning",
"AI innovation"
],
"confidence": 0.85,
"step_type": "action",
"tools_used": [
"web_search"
]
}
},
{
"workspace_id": "test_workspace",
"memory_id": "1aa50187c2da41a483256a55aae8e260",
"memory_type": "task",
"when_to_use": "When synthesizing diverse data sources into a coherent, structured narrative for executive-level understanding.",
"content": "The agent did not merely aggregate raw facts but synthesized findings into a well-organized, thematic report that connected history, business model, financials, innovation, and competition. It highlighted critical contradictions (e.g., declining profits despite strong Q4 growth) and strategic shifts (e.g., move toward software/services). By identifying key metrics (like carbon credit income and FSD safety record) as differentiators, it transformed data into insight. This demonstrates the importance of post-data synthesis reasoning—turning fragmented inputs into actionable, narrative-driven conclusions that reflect both factual accuracy and strategic interpretation.",
"score": 0.85,
"time_created": "2025-09-07 15:57:06",
"time_modified": "2025-09-07 15:57:06",
"author": "qwen3-30b-a3b-instruct-2507",
"metadata": {
"when_to_use": "When synthesizing diverse data sources into a coherent, structured narrative for executive-level understanding.",
"experience": "The agent did not merely aggregate raw facts but synthesized findings into a well-organized, thematic report that connected history, business model, financials, innovation, and competition. It highlighted critical contradictions (e.g., declining profits despite strong Q4 growth) and strategic shifts (e.g., move toward software/services). By identifying key metrics (like carbon credit income and FSD safety record) as differentiators, it transformed data into insight. This demonstrates the importance of post-data synthesis reasoning—turning fragmented inputs into actionable, narrative-driven conclusions that reflect both factual accuracy and strategic interpretation.",
"tags": [
"synthesis",
"narrative structuring",
"insight generation",
"strategic interpretation",
"data integration",
"executive summary"
],
"confidence": 0.9,
"step_type": "reasoning",
"tools_used": [
"web_search"
]
}
}
]

View file

@ -0,0 +1,106 @@
[
{
"role": "user",
"content": "You are a helpful assistant.\nThe current time is 2025-09-07 15:55:08.\nPlease proactively choose the most suitable tool or combination of tools based on the user's question, including web_search etc.\nPlease first think about how to break down the problem into subtasks, what tools and parameters should be used for each subtask, and finally provide the tool call name and parameters.\nTry calling the same tool multiple times with different parameters to obtain information from various perspectives.\nPlease determine the response language based on the language of the user's question.\n\nAnalyze Xiaomi Corporation\n",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "",
"time_created": "2025-09-07 15:55:08",
"metadata": {}
},
{
"role": "assistant",
"content": "",
"reasoning_content": "",
"tool_calls": [
{
"index": 0,
"id": "call_cf6bfde51c5f4980bc6d68",
"type": "function",
"name": "web_search",
"arguments": "{\"query\": \"Xiaomi Corporation company overview\"}",
"description": "",
"input_schema": {},
"output_schema": {}
},
{
"index": 1,
"id": "call_d537006a5856429db13013",
"type": "function",
"name": "web_search",
"arguments": "{\"query\": \"Xiaomi Corporation financial performance 2024\"}",
"description": "",
"input_schema": {},
"output_schema": {}
},
{
"index": 2,
"id": "call_a23e2b5e71ed4afd91bfb9",
"type": "function",
"name": "web_search",
"arguments": "{\"query\": \"Xiaomi Corporation product portfolio and innovations\"}",
"description": "",
"input_schema": {},
"output_schema": {}
},
{
"index": 3,
"id": "call_adec4f13bc4c4b859d2b08",
"type": "function",
"name": "web_search",
"arguments": "{\"query\": \"Xiaomi Corporation market position in smartphones and IoT devices\"}",
"description": "",
"input_schema": {},
"output_schema": {}
}
],
"tool_call_id": "",
"time_created": "2025-09-07 15:55:27",
"metadata": {}
},
{
"role": "tool",
"content": "小米公司正式成立於2010年4月, 並於2018年7月9日在香港聯交所主板上市。小米公司是一家以智能手機、智能硬件和 IoT 平台為核心的消費電子及智能製造公司。 創新與品質並重,小米不斷追求極致的用戶體驗和運營效率。小米的使命是,始終堅持做\"感動人心、價格厚道\"的好產品,讓全球每個人都能享受科技帶來的美好生活。 小米是全球領先的智能手機公司之一。\n\nXiaomi Corporation was founded in April 2010 and listed on the Main Board of the Hong Kong Stock Exchange on July 9, 2018 (1810.HK). Xiaomi is a consumer electronics and smart manufacturing company with smartphones and smart hardware connected by an IoT platform at its core. With an equal emphasis on innovation and quality, Xiaomi continuously pursues high-quality user experience and operational efficiency. The company relentlessly builds amazing products with honest prices to let everyone in the world enjoy a better life through innovative technology. Xiaomi is one of the world's leading smartphone companies. As of 2024, Xiaomi ranked among the top 3 in the global smartphone market, in terms of smartphone shipments, according to Canalys. The company has also established the worlds leading consumer AIoT (AI+IoT) platform, with 904.6 million smart devices connected to its platform, excluding smartphones, tablets and laptops, as of December 31, 2024. Xiaomi products are present in more than 100 countries and regions around the world. In August 2024, the company listed as Fortune Global 500 for the 6th consecutive year. Xiaomi is a constituent of the Hang Seng Index, Hang Seng China Enterprises Index, Hang Seng TECH Index and Hang Seng China 50 Index.",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "call_cf6bfde51c5f4980bc6d68",
"time_created": "2025-09-07 15:55:46",
"metadata": {}
},
{
"role": "tool",
"content": "- Xiaomi Corporation reported a historical high in annual revenue for 2024 with JPY365.9 billion, marking a 35% year-on-year increase.\n- The smartphone revenue reached CNY 191.8 billion ($26.5 billion), accounting for 52% of overall sales. Xiaomi sold 169 million units in 2024.\n- In Q2 2024, Xiaomi's ASP (Average Selling Price) was around 1,102 yuan, slightly down year-over-year but still above the 1,100 yuan median.\n- Xiaomi announced its 2024 Annual Results on 18 March 2025, and hosted an investor conference call/audio webcast on the same day at 08:30 pm (Beijing Time).",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "call_d537006a5856429db13013",
"time_created": "2025-09-07 15:55:46",
"metadata": {}
},
{
"role": "tool",
"content": "Xiaomi expands its product portfolio with new smart products such as Smart TVs, Air Purifier, Vacuum Cleaner, Trimmer and more. \nFrom immersive entertainment to fitness, fashion, and smart mobility, Xiaomi has enriched its AIoT portfolio, reinforcing its commitment to innovation. \nProduct portfolio expands through branding of hardware made by partner companies. \nXiaomi Corporation is an Internet company producing smartphones and AIoT (Artificial Intelligence of Things) devices.",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "call_a23e2b5e71ed4afd91bfb9",
"time_created": "2025-09-07 15:55:46",
"metadata": {}
},
{
"role": "tool",
"content": "小米公司是全球领先的智能手机公司之一。根据Canalys数据,2024年在全球范围内手机出货量排名第三。For smartphones priced at RMB4,000 and above, Xiaomi's smartphone units sold garnered a market share of 9.6%, up 2.9 percentage points YoY. In the 5000 - 6000 yuan price segment, Xiaomi's market share was 15.4%, a year - on - year increase of 6.5 percentage points.\n\n小米还建立了世界领先的消费级AIoT (AI+IoT)平台,截至2024年12月31日,AIoT平台已连接的IoT设备数(不包括智能手机、平板及笔记本电脑)已达到9.05亿。",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "call_adec4f13bc4c4b859d2b08",
"time_created": "2025-09-07 15:55:46",
"metadata": {}
},
{
"role": "assistant",
"content": "Based on the provided information, here is a comprehensive analysis of Xiaomi Corporation:\n\n### **Company Overview**\n- **Founded**: April 2010\n- **Listing**: Hong Kong Stock Exchange (1810.HK) on July 9, 2018\n- **Core Business**: Smartphones, smart hardware, and IoT platforms.\n- **Mission**: To create \"emotionally touching, reasonably priced\" products that enhance global lifestyles through technology.\n\n### **Market Position & Performance (2024)**\n- **Global Smartphone Market**: Ranked among the top 3 in smartphone shipments globally, according to Canalys.\n - Sold **169 million units** in 2024.\n - Achieved significant market share growth in premium segments:\n - **RMB 4,000+ segment**: 9.6% market share (+2.9 pp YoY).\n - **5,0006,000 yuan segment**: 15.4% market share (+6.5 pp YoY).\n- **Financial Highlights (2024)**:\n - Record annual revenue: **JPY 365.9 billion** (+35% YoY).\n - Smartphone revenue: **CNY 191.8 billion ($26.5 billion)**, accounting for **52%** of total sales.\n - Q2 2024 Average Selling Price (ASP): ~**1,102 yuan**, maintaining strong pricing power above the 1,100 yuan median.\n\n### **AIoT Ecosystem & Innovation**\n- **World-Leading Consumer AIoT Platform**:\n - As of December 31, 2024, **904.6 million IoT devices** were connected to Xiaomis platform (excluding smartphones, tablets, and laptops).\n- **Product Portfolio Expansion**:\n - Diversified into smart home appliances (Smart TVs, Air Purifiers, Vacuum Cleaners, Trimmers), fitness, fashion, and smart mobility.\n - Enhanced product offerings through partnerships and branded hardware from third-party manufacturers.\n\n### **Global Reach & Recognition**\n- Products available in **over 100 countries and regions**.\n- Recognized as a **Fortune Global 500 company** for the **6th consecutive year** (August 2024).\n- Constituent of major indices: Hang Seng Index, Hang Seng China Enterprises Index, Hang Seng TECH Index, and Hang Seng China 50 Index.\n\n### **Conclusion**\nXiaomi Corporation has solidified its position as a global leader in both smartphones and consumer AIoT ecosystems. With record-breaking financial performance in 2024, strong growth in premium smartphone segments, and an expanding portfolio of innovative smart devices, Xiaomi continues to deliver value through innovation and operational efficiency. Its mission-driven approach—combining high-quality products with affordable pricing—resonates across markets worldwide.",
"reasoning_content": "",
"tool_calls": [],
"tool_call_id": "",
"time_created": "2025-09-07 15:55:55",
"metadata": {}
}
]

View file

@ -0,0 +1,92 @@
import asyncio
import json
import aiohttp
# API base URL
base_url = "http://0.0.0.0:8002"
async def main():
# Create a unique workspace ID
workspace_id = "personal_memory_demo"
async with aiohttp.ClientSession() as session:
# Step 1: Clear existing memories in the workspace
print("Clearing existing memories...")
async with session.post(
f"{base_url}/vector_store",
json={
"action": "delete",
"workspace_id": workspace_id,
},
headers={"Content-Type": "application/json"}
) as response:
result = await response.json()
print(json.dumps(result, ensure_ascii=False))
# Step 2: Create a conversation with rich personal information
print("\nCreating conversation with personal information...")
messages = [
{"role": "user", "content": "My name is John Smith, I'm 28 years old, and I work at a tech company in San Francisco"},
{"role": "assistant", "content": "Nice to meet you, John!"},
{"role": "user", "content": "I'm a software engineer, mainly doing backend development using Python and Go"},
{"role": "assistant", "content": "I see, you're a backend engineer working with Python and Go."},
{"role": "user", "content": "I enjoy playing basketball and watching sci-fi movies. I recently watched Dune Part 2"},
{"role": "assistant", "content": "Basketball and sci-fi movies are great hobbies! Dune Part 2 was indeed amazing."},
{"role": "user", "content": "I have a cat named Shadow who is 3 years old"},
{"role": "assistant", "content": "Shadow sounds adorable! 3-year-old cats are quite playful."},
{"role": "user", "content": "I'm planning a trip to Japan next month, mainly to Tokyo and Kyoto"},
{"role": "assistant", "content": "Your Japan trip sounds exciting! Tokyo and Kyoto are both wonderful destinations with their own unique charm."},
{"role": "user", "content": "I'm really interested in Japanese cuisine, especially sushi and ramen"},
{"role": "assistant", "content": "Japanese cuisine is delicious! Sushi and ramen are very popular choices."},
]
# Step 3: Summarize personal memories from the conversation
print("\nSummarizing personal memories...")
async with session.post(
f"{base_url}/summary_personal_memory",
json={
"trajectories": [
{"messages": messages, "score": 1.0}
],
"workspace_id": workspace_id,
},
headers={"Content-Type": "application/json"}
) as response:
result = await response.json()
result = json.dumps(result, ensure_ascii=False, indent=2)
print(result)
with open("personal_memory.jsonl", "w") as f:
f.write(result)
# Wait for the memories to be processed and stored
print("\nWaiting for memories to be processed...")
await asyncio.sleep(2)
# Step 4: Retrieve personal memories with different queries
queries = [
"What's my name and age?",
"What do I do for work?",
"What are my hobbies?",
"Do I have any pets?",
"What are my travel plans?",
"What foods do I like?"
]
print("\nRetrieving personal memories...")
for query in queries:
print(f"\nQuery: {query}")
async with session.post(
f"{base_url}/retrieve_personal_memory",
json={
"query": query,
"workspace_id": workspace_id,
},
headers={"Content-Type": "application/json"}
) as response:
result = await response.json()
print(json.dumps(result, ensure_ascii=False, indent=2))
if __name__ == "__main__":
asyncio.run(main())

View file

@ -0,0 +1,291 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
Task Memory Demo for MemoryScope
This script demonstrates how to use the task memory capabilities of MemoryScope.
It shows how to run an agent, summarize conversations, retrieve memories, and
manage the memory workspace.
"""
import json
import time
from typing import List, Dict, Any, Optional
import requests
from dotenv import load_dotenv
# Load environment variables from .env file
load_dotenv()
# API configuration
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "test_workspace"
def handle_api_response(response: requests.Response) -> Optional[Dict[str, Any]]:
"""
Handle API response with proper error checking
Args:
response: Response object from requests
Returns:
Response JSON if successful, None otherwise
"""
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return None
return response.json()
def delete_workspace() -> None:
"""
Delete the current workspace from the vector store
Returns:
None
"""
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "delete",
}
)
result = handle_api_response(response)
if result:
print(f"Workspace '{WORKSPACE_ID}' deleted successfully")
def run_agent(query: str, dump_messages: bool = False) -> List[Dict[str, Any]]:
"""
Run the agent with a specific query
Args:
query: The query to send to the agent
dump_messages: Whether to save messages to a file
Returns:
List of message objects from the conversation
"""
response = requests.post(
url=f"{BASE_URL}react",
json={"query": query}
)
result = handle_api_response(response)
if not result:
return []
# Extract and display the answer
answer = result.get("answer", "")
print(f"Agent response: {answer}")
# Get the conversation messages
messages = result.get("messages", [])
# Optionally save messages to file
if dump_messages and messages:
with open("task_messages.jsonl", "w") as f:
f.write(json.dumps(messages, indent=2, ensure_ascii=False))
print(f"Messages saved to messages.jsonl")
return messages
def run_summary(messages: List[Dict[str, Any]], enable_dump_memory: bool = True) -> None:
"""
Generate a summary of conversation messages and create task memories
Args:
messages: List of message objects from a conversation
enable_dump_memory: Whether to save memory list to a file
Returns:
None
"""
if not messages:
print("No messages to summarize")
return
response = requests.post(
# url=f"{BASE_URL}summary_task_memory_simple",
url=f"{BASE_URL}summary_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"trajectories": [
{"messages": messages, "score": 1.0}
]
}
)
result = handle_api_response(response)
if not result:
return
# Extract memory list from response
memory_list = result.get("metadata", {}).get("memory_list", [])
print(f"Memory list: {memory_list}")
# Optionally save memory list to file
if enable_dump_memory and memory_list:
with open("task_memory.jsonl", "w") as f:
f.write(json.dumps(memory_list, indent=2, ensure_ascii=False))
print(f"Memory saved to memory.jsonl")
def run_retrieve(query: str) -> str:
"""
Retrieve relevant task memories based on a query
Args:
query: The query to retrieve relevant memories
Returns:
String containing the retrieved memory answer
"""
response = requests.post(
# url=f"{BASE_URL}retrieve_task_memory_simple",
url=f"{BASE_URL}retrieve_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"query": query,
}
)
result = handle_api_response(response)
if not result:
return ""
# Extract and return the answer
answer = result.get("answer", "")
print(f"Retrieved memory: {answer}")
return answer
def run_agent_with_memory(query_first: str, query_second: str, enable_dump_memory: bool = True) -> List[Dict[str, Any]]:
"""
Run the agent with memory augmentation
This function demonstrates how to use task memory to enhance agent responses:
1. First run the agent with the second query to build memory
2. Then summarize the conversation to create memories
3. Retrieve relevant memories for the first query
4. Run the agent with the first query augmented with retrieved memories
Args:
query_first: The query to run with memory augmentation
query_second: The query to build initial memories
enable_dump_memory: Whether to save memory list to a file
Returns:
List of message objects from the final conversation
"""
# Run agent with second query to build initial memories
print(f"\n--- Building memories with query: '{query_second}' ---")
messages = run_agent(query=query_second)
# Summarize conversation to create memories
print("\n--- Summarizing conversation to create memories ---")
run_summary(messages, enable_dump_memory)
time.sleep(1)
# Retrieve relevant memories for the first query
print(f"\n--- Retrieving memories for query: '{query_first}' ---")
retrieved_memory = run_retrieve(query_first)
# Run agent with first query augmented with retrieved memories
print(f"\n--- Running agent with memory-augmented query ---")
augmented_query = f"{retrieved_memory}\n\nUser Question:\n{query_first}"
print(f"Augmented query: {augmented_query}")
messages = run_agent(query=augmented_query)
return messages
def dump_memory(path: str = "./") -> None:
"""
Dump the vector store memories to disk
Args:
path: Directory path to save the memories
Returns:
None
"""
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "dump",
"path": path,
}
)
result = handle_api_response(response)
if result:
print(f"Memory dumped to {path}")
def load_memory(path: str = "./") -> None:
"""
Load memories from disk into the vector store
Args:
path: Directory path to load the memories from
Returns:
None
"""
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "load",
"path": path,
}
)
result = handle_api_response(response)
if result:
print(f"Memory loaded from {path}")
def main() -> None:
"""
Main function to demonstrate task memory workflow
"""
# Define example queries
query1 = "Analyze Xiaomi Corporation"
query2 = "Analyze the company Tesla."
print("=== Task Memory Demo ===")
# Step 1: Clean up workspace
print("\n1. Deleting workspace...")
delete_workspace()
# Step 2: Run agent with first query and save messages
print("\n2. Running agent with first query...")
run_agent(query=query1, dump_messages=True)
# Step 3: Demonstrate memory-augmented agent
print("\n3. Running memory-augmented agent workflow...")
run_agent_with_memory(query_first=query1, query_second=query2)
# Step 4: Demonstrate memory persistence
print("\n4. Dumping memory to disk...")
dump_memory()
print("\n5. Loading memory from disk...")
load_memory()
print("\n=== Demo Complete ===")
if __name__ == "__main__":
main()

View file

@ -0,0 +1,233 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
Task Memory Demo for MemoryScope using MCP Client
This script demonstrates how to use the task memory capabilities of MemoryScope
through the MCP client interface. It shows how to run an agent, summarize conversations,
retrieve memories, and manage the memory workspace.
"""
import asyncio
import json
from typing import List, Dict, Any
from dotenv import load_dotenv
from fastmcp import Client
# Load environment variables from .env file
load_dotenv()
# API configuration
MCP_URL = "http://0.0.0.0:8002/sse/"
WORKSPACE_ID = "test_workspace"
async def delete_workspace(client: Client) -> None:
"""
Delete the current workspace from the vector store
Args:
client: MCP client instance
Returns:
None
"""
result = await client.call_tool(
"vector_store",
arguments={
"workspace_id": WORKSPACE_ID,
"action": "delete",
}
)
print(f"Workspace '{WORKSPACE_ID}' deleted successfully")
async def run_agent(client: Client, query: str, dump_messages: bool = False) -> List[Dict[str, Any]]:
with open("task_messages.jsonl") as f:
messages = json.loads(f.read())
print(f"messages={messages}")
return messages
async def run_summary(client: Client, messages: List[Dict[str, Any]], enable_dump_memory: bool = True) -> None:
"""
Generate a summary of conversation messages and create task memories
Args:
client: MCP client instance
messages: List of message objects from a conversation
enable_dump_memory: Whether to save memory list to a file
Returns:
None
"""
if not messages:
print("No messages to summarize")
return
result = await client.call_tool(
"summary_task_memory",
arguments={
"workspace_id": WORKSPACE_ID,
"trajectories": [
{"messages": messages, "score": 1.0}
]
}
)
answer = result.content[0].text
# Extract memory list from response
print(f"Memory list: {answer}")
if enable_dump_memory:
with open("mcp_task_memory.jsonl", "w") as f:
f.write(answer)
print(f"Memory saved to mcp_task_memory.jsonl")
async def run_retrieve(client: Client, query: str) -> str:
"""
Retrieve relevant task memories based on a query
Args:
client: MCP client instance
query: The query to retrieve relevant memories
Returns:
String containing the retrieved memory answer
"""
result = await client.call_tool(
"retrieve_task_memory",
arguments={
"workspace_id": WORKSPACE_ID,
"query": query,
}
)
answer = result.content[0].text
print(f"Retrieved memory: {answer}")
return answer
async def run_agent_with_memory(client: Client, query_first: str, query_second: str, enable_dump_memory: bool = True) -> List[Dict[str, Any]]:
"""
Run the agent with memory augmentation
This function demonstrates how to use task memory to enhance agent responses:
1. First run the agent with the second query to build memory
2. Then summarize the conversation to create memories
3. Retrieve relevant memories for the first query
4. Run the agent with the first query augmented with retrieved memories
Args:
client: MCP client instance
query_first: The query to run with memory augmentation
query_second: The query to build initial memories
enable_dump_memory: Whether to save memory list to a file
Returns:
List of message objects from the final conversation
"""
# Run agent with second query to build initial memories
print(f"\n--- Building memories with query: '{query_second}' ---")
messages = await run_agent(client, query=query_second)
# Summarize conversation to create memories
print("\n--- Summarizing conversation to create memories ---")
await run_summary(client, messages, enable_dump_memory)
await asyncio.sleep(1)
# Retrieve relevant memories for the first query
print(f"\n--- Retrieving memories for query: '{query_first}' ---")
retrieved_memory = await run_retrieve(client, query_first)
# Run agent with first query augmented with retrieved memories
print(f"\n--- Running agent with memory-augmented query ---")
augmented_query = f"{retrieved_memory}\n\nUser Question:\n{query_first}"
print(f"Augmented query: {augmented_query}")
messages = await run_agent(client, query=augmented_query)
return messages
async def dump_memory(client: Client, path: str = "./") -> None:
"""
Dump the vector store memories to disk
Args:
client: MCP client instance
path: Directory path to save the memories
Returns:
None
"""
result = await client.call_tool(
"vector_store",
arguments={
"workspace_id": WORKSPACE_ID,
"action": "dump",
"path": path,
}
)
print(f"Memory dumped to {path}")
async def load_memory(client: Client, path: str = "./") -> None:
"""
Load memories from disk into the vector store
Args:
client: MCP client instance
path: Directory path to load the memories from
Returns:
None
"""
result = await client.call_tool(
"vector_store",
arguments={
"workspace_id": WORKSPACE_ID,
"action": "load",
"path": path,
}
)
print(f"Memory loaded from {path}")
async def main() -> None:
"""
Main function to demonstrate task memory workflow
"""
# Define example queries
query1 = "Analyze Xiaomi Corporation"
query2 = "Analyze the company Tesla."
print("=== Task Memory Demo (MCP Client) ===")
async with Client(MCP_URL) as client:
# Step 1: Clean up workspace
print("\n1. Deleting workspace...")
await delete_workspace(client)
# Step 2: Run agent with first query and save messages
print("\n2. Running agent with first query...")
await run_agent(client, query=query1, dump_messages=True)
# Step 3: Demonstrate memory-augmented agent
print("\n3. Running memory-augmented agent workflow...")
await run_agent_with_memory(client, query_first=query1, query_second=query2)
# Step 4: Demonstrate memory persistence
print("\n4. Dumping memory to disk...")
await dump_memory(client)
print("\n5. Loading memory from disk...")
await load_memory(client)
print("\n=== Demo Complete ===")
if __name__ == "__main__":
asyncio.run(main())

36
docs/contribution.md Normal file
View file

@ -0,0 +1,36 @@
# Contribute to ReMe
Our community thrives on the diverse ideas and contributions of its members. Whether you're fixing a bug, adding a new feature, improving the documentation, or adding examples, your help is welcome. Here's how you can contribute:
## Report Bugs and Ask For New Features?
Did you find a bug or have a feature request? Please first check the issue tracker to see if it has already been reported. If not, feel free to open a new issue. Include as much detail as possible:
- A descriptive title
- Clear description of the issue
- Steps to reproduce the problem
- Version of the ReMe you are using
- Any relevant code snippets or error messages
## Contribute to Codebase
### Fork and Clone the Repository
To work on an issue or a new feature, start by forking the ReMe repository and then cloning your fork locally.
```bash
git clone https://github.com/your-username/ReMe.git
cd ReMe
```
### Create a New Branch
Create a new branch for your work. This helps keep proposed changes organized and separate from the `main` branch.
```bash
git checkout -b your-feature-branch-name
```
### Making Changes
With your new branch checked out, you can now make your changes to the code. Remember to keep your changes as focused as possible. If you're addressing multiple issues or features, it's better to create separate branches and pull requests for each.
### Commit Your Changes
Once you've made your changes, it's time to commit them. Write clear and concise commit messages that explain your changes.
```bash
git add -A
git commit -m "A brief description of the changes"
```
### Submit a Pull Request
When you're ready for feedback, submit a pull request to the ReMe `main` branch. In your pull request description, explain the changes you've made and any other relevant context.
We will review your pull request. This process might involve some discussion, additional changes on your part, or both.
### Code Review
Wait for us to review your pull request. We may suggest some changes or improvements. Keep an eye on your GitHub notifications and be responsive to any feedback.

View file

@ -0,0 +1,141 @@
# AppWorld Experiment Quick Start Guide
This guide helps you quickly set up and run AppWorld experiments with ReMe integration.
## Env Setup
### 1. Clone the Repository
```bash
git clone https://github.com/modelscope/ReMe.git
cd ReMe/cookbook/appworld
```
### 2. Appworld Environment Setup
Create a new conda environment with Python 3.12:
```bash
conda create -p ./appworld-env python==3.12
conda activate ./appworld-env
```
Install required Python packages:
```bash
pip install -r requirements.txt
```
Install AppWorld and download the dataset:
```bash
pip install appworld
appworld install
appworld download data
```
**Note**: The AppWorld data will be saved in the current directory.
### 3. Start ReMe Service
Install ReMe (if not already installed)
If you haven't installed the ReMe environment yet, follow these steps:
```bash
# Go back to the project root
cd ../..
# Create ReMe environment
conda create -p ./reme-env python==3.12
conda activate ./reme-env
# Install ReMe
pip install .
```
Launch the ReMe service to enable memory library functionality:
```bash
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen-max-latest \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
add memories for appworld:
```bash
curl -X POST "http://0.0.0.0:8002/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "appworld",
"action": "load",
"path": "./docs/library"
}'
```
Now you have loaded the ReMe memory library to enable memory-based agent!
### 4. Common Issues
**AppWorld data not found**: Ensure `appworld download data` completed successfully
**pydantic version issue**: AppWorld depends on an older version of pydantic, which is why a separate environment is needed. If you encounter issues running the experiments, try `pip install appworld` to override the dependencies.
## Run Experiments
### 1. Test: With Memory vs Without Memory
Run the main experiment script to compare performance with and without memory:
```bash
python run_appworld.py
```
**What this does:**
- Runs AppWorld tasks on the development dataset
- Compares agent performance with ReMe memory (`use_memory=True`) vs without memory
- Uses multiple workers for parallel processing
- Runs each task multiple times for statistical significance
- Results are automatically saved to `./exp_result/` directory
**Configuration options in `run_appworld.py`:**
- `max_workers`: Number of parallel workers (default: 6)
- `num_runs`: Number of times each task is repeated (default: 4)
- `use_memory`: Whether to use ReMe memory library
### 2. View Experiment Results
After running experiments, analyze the statistical results:
```bash
python run_exp_statistic.py
```
**What this script does:**
- Processes all result files in `./exp_result/`
- Calculates best@k metrics for different k values
- Generates a summary table showing performance comparisons
- Saves results to `experiment_summary.csv`
**Metrics explained:**
- `best@k`: Takes groups of k runs per task, finds the maximum score in each group, then averages these maximums
- Higher k values show potential performance, lower k values show consistency
**Output Files**
- `./exp_result/*.jsonl`: Raw experiment results for each configuration
- `./exp_result/experiment_summary.csv`: Statistical summary table
- Console output: Real-time progress and summary statistics
## Understanding Results
The experiment compares:
1. **Baseline**: Agent without memory library
2. **With Memory**: Agent enhanced with ReMe memory library
Key metrics to look for:
- **best@1**: Average performance across all single runs
- **best@k**: Performance when taking the best of k attempts
- Improvement percentage when using memory vs baseline

View file

@ -0,0 +1,120 @@
# BFCL Experiment Quick Start Guide
This guide helps you quickly set up and run BFCL experiments with ReMe integration.
## Env Setup
### 1. BFCL installation
#### clone the repository
```bash
git clone https://github.com/ShishirPatil/gorilla.git
```
#### Change directory to the `berkeley-function-call-leaderboard`
```bash
cd gorilla/berkeley-function-call-leaderboard
```
#### Install the package in editable mode
```bash
conda create -n bfcl-env python==3.12
conda activate bfcl-env
pip install -e .
pip install -r requirements.txt
```
#### Move the dataset to the data folder under bfcl
```bash
cp -r bfcl_eval/data {/path/to/bfcl/data}
```
**Note**: The original BFCL data is designed as a benchmark dataset and does not have a train/validation split, you can use ``split_into_trainval.py`` to split JSONL file into train and validation sets.
### 2. Collect agent trajectories on training data set
Run the main experiment script to collect agent trajectories on training data set without task memory(`use_memory=False`):
```bash
python run_bfcl.py
```
**Note**:
- `max_workers`: Number of parallel workers (default: `4`)
- `num_runs`: Number of times each task is repeated (default: `1`)
- `model_name`: LLM model name (default: `qwen3-8b`)
- `enable_thinking`: Control the model's thinking mode (default: `False`)
- `data_path`: Path to the training dataset (default: `./data/multiturn_data_base_train.jsonl`)
- `answer_path`: Path to the possible answer, which are used to evaluate the model's output function (default: `./data/possible_answer`)
- Results are automatically saved to `./exp_result/{model_name}/{no_think/with_think}` directory
### 3. Start ReMe Service and Init the task memory pool
After collecting trajectories, Launch the ReMe service (make sure you have installed ReMe environment, if not please follow the steps in the [ReMe Installation Guide](https://github.com/modelscope/ReMe/blob/main/doc/README.md) to install):
```bash
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen-max-2025-01-25 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
and then init the task memory pool:
```bash
python init_exp_pool.py
```
**Configuration options in `init_exp_pool.py`:**
- `jsonl_file`: Path to the collloaded trajectories
- `service_url`: ReMe service URL (default: `http://localhost:8002`)
- `workspace_id`: Workspace ID for the task memory pool (default: `bfcl_test`)
- `n_threads`: Number of threads for processing (default: `4`)
- `output_file`: Output file to save results (optional)
Now you have inited the task memory pool using `local` backend (start on `http://localhost:8002`). Then, use `local_file_to_library.py` script to convert the local file to the memory library or run the following `curl` command:
```bash
curl -X POST "http://0.0.0.0:8002/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "bfcl_test",
"action": "dump",
"path": "./library"
}'
```
to dump the memory library (default in `./library/bfcl_test.jsonl`).
Next time, you can import this previously exported task memory data to populate the new started workspace with existing knowledge:
```bash
curl -X POST "http://0.0.0.0:8002/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "bfcl_test",
"action": "load",
"path": "./library"
}'
```
### 4. Run Experiments on Validation Set
Run you can compare agent performance on the validation set with task memory (`use_memory=True`) and without task memory:
```bash
# remember to change the configuration options, e.g., `data_path=./data/multiturn_data_base_val.jsonl`
python run_bfcl.py
```
After running experiments, analyze the statistical results:
```bash
python run_exp_statistic.py
```
**What this script does:**
- Processes all result files in `./exp_result/`
- Calculates best@k metrics for different k values
- Generates a summary table showing performance comparisons
- Saves results to `experiment_summary.csv`

View file

@ -0,0 +1,153 @@
# FrozenLake Experiment Quick Start Guide
This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task.
## Environment Setup
### 1. Clone the Repository
```bash
git clone https://github.com/modelscope/ReMe.git
cd ReMe/cookbook/frozenlake
```
### 2. FrozenLake Environment Setup
Install Gymnasium for FrozenLake environment:
```bash
pip install gymnasium
```
This will install:
- gymnasium - for the FrozenLake environment
- ray - for parallel execution
- openai - for LLM API access
- other dependencies
### 3. Start ReMe Service
If you haven't installed ReMe yet, follow these steps:
```bash
# Go back to the project root
cd ../..
# Create a virtual environment (optional)
conda create -p ./reme-env python==3.10
conda activate ./reme-env
# Install ReMe
pip install .
```
Launch the ReMe service to enable memory library functionality:
```bash
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen-max-2025-01-25 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local
```
Add your api key for agent:
```bash
export OPENAI_API_KEY="xxx"
export OPENAI_BASE_URL="xxx"
```
## Run Experiments
### 1. Quick Test: Performance Evaluation Only (Default)
Run the main experiment script to test agent performance using existing memory:
```bash
cd cookbook/frozenlake
python run_frozenlake.py
```
**What this does:**
- Tests the agent on randomly generated FrozenLake maps
- Uses the default memory library (`frozenlake_no_slippery`)
- Evaluates performance with multiple runs for statistical significance
- Results are automatically saved to `./exp_result/` directory
### 2. Advanced: Training + Testing (Memory Generation)
To create new memories through training and then test performance:
You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function:
```python
def main():
experiment_name = "frozenlake_no_slippery" # Name of the experiment
max_workers = 4 # Number of parallel workers
training_runs = 4 # Runs per training map
num_training_maps = 50 # Number of maps for training
test_runs = 1 # Runs per test configuration
num_test_maps = 100 # Number of test maps
is_slippery = False # Enable slippery mode
```
Key parameters to consider:
- `experiment_name`: Used as the workspace ID for task memory
- `is_slippery`: When True, agent movement becomes stochastic (harder)
- `max_workers`: Increase for faster execution on multi-core systems
### 3. View Experiment Results
After running experiments, analyze the statistical results:
```bash
python run_exp_statistic.py
```
**What this script does:**
- Processes all result files in `./exp_result/`
- Calculates success rates and performance metrics
- Generates a summary table showing performance comparisons
- Analyzes the effect of task memory on performance
- Saves results to `frozenlake_summary.csv`
## Understanding the Implementation
### Key Components
1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`)
- Implements a ReAct agent that interacts with the FrozenLake environment
- Handles task memory retrieval and storage
- Uses LLM (via OpenAI API) for decision making
2. **Experiment Runner** (`run_frozenlake.py`)
- Manages the overall experiment flow
- Handles training and testing phases
- Uses Ray for parallel execution
3. **Map Manager** (`map_manager.py`)
- Generates and manages test maps
- Ensures consistent evaluation across experiments
4. **Statistics Analyzer** (`run_exp_statistic.py`)
- Processes experiment results
- Calculates performance metrics
- Generates comparative analysis
### Output Files
- `./exp_result/*_training.jsonl`: Results from training phase
- `./exp_result/*_test_no_memory.jsonl`: Test results without task memory
- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory
- `./exp_result/frozenlake_summary.csv`: Statistical summary
### Task Memory Mechanism
The task memory system works as follows:
1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service
2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map
3. **Memory Application**: The agent uses retrieved memories to guide its decision-making
The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake.

View file

@ -1,238 +0,0 @@
# Auto Dream
`auto_dream` is ReMe's long-term memory distillation flow from daily to digest. By default it scans the target date and
the previous day, processes only files changed since the previous dream, extracts a small set of high-value memory units
across that window, integrates them into `digest/`, and writes the target day's `interests.yaml` for proactive use.
<p align="center">
<img src="../figure/auto-dream-and-proactive.svg" alt="ReMe Auto Dream and Proactive flow from daily to digest to proactive" width="92%">
</p>
Its daily inputs usually come from [Auto Memory](./auto_memory.md) and [Auto Resource](./auto_resource.md). For the file
semantics of `digest/`, Sources sections, and wikilinks, see [Memory as File](./memory_as_file.md). For the linking
strategy used during Integrate, see [Auto Link](./auto_link.md). To read `interests.yaml`,
use [Proactive](./proactive.md).
## Configuration
The default configuration is in `reme/config/default.yaml`:
```yaml
auto_dream:
backend: base
parameters:
date:
type: string
default: ""
hint:
type: string
default: ""
scan_days:
type: integer
default: 2
max_units:
type: integer
default: 5
topic_count:
type: integer
default: 3
topic_diversity_days:
type: integer
default: 7
steps:
- backend: dream_extract_step
file_catalog: dream
topic_session_id: interests
scan_days: 2
max_units: 5
- backend: dream_integrate_step
- backend: dream_topics_step
topic_count: 3
topic_diversity_days: 7
- backend: dream_finish_step
file_catalog: dream
```
Parameters:
| Parameter | Purpose |
|------------------------|---------------------------------------------------------------------------------------------------------|
| `date` | Date to process in `YYYY-MM-DD` format. When empty, use today in the application's timezone. |
| `hint` | Additional guidance from the caller for the Extract and Integrate stages. |
| `scan_days` | Recent-date window ending at `date`; defaults to 2 and has a minimum of 1. |
| `max_units` | Maximum reusable units extracted in one run; defaults to 5. |
| `topic_count` | Maximum number of topics written to `interests.yaml`. Defaults to 3. |
| `topic_diversity_days` | Number of past days of `interests.yaml` files considered when avoiding duplicate topics. Defaults to 7. |
## Inputs and Outputs
Inputs are daily Markdown files from the most recent `scan_days` ending at the specified date. For example,
`date=2026-06-20` with `scan_days=2` scans:
```text
daily/2026-06-19.md
daily/2026-06-19/**/*.md
daily/2026-06-20.md
daily/2026-06-20/**/*.md
```
Every `daily/<date>/interests.yaml` in the scan window is excluded from extraction so previous proactive output cannot
feed back into the next run. Final topics are written only for the target date.
The main outputs are:
| Output | Description |
|--------------------------------|------------------------------------------------------------------------------|
| `digest/procedure/*.md` | Methods, workflows, runbooks, and executable experience. |
| `digest/personal/*.md` | User-, team-, and project-related preferences, facts, and long-term context. |
| `digest/wiki/*.md` | General knowledge, concepts, observations, and decision precedents. |
| `daily/<date>/interests.yaml` | Topics worth proactive attention from the host agent that day. |
| `metadata/file_catalog/dream*` | Dream-specific catalog used to detect changes in daily inputs. |
## Four Stages
### 1. Extract
`dream_extract_step` performs three tasks:
1. Refresh each `daily/<date>.md` in the scan window.
2. Scan those day indexes and `daily/<date>/**/*.md`, comparing mtimes with `file_catalog: dream`.
3. Send all changed files together to the LLM and globally extract two structured result types: `units` and `topics`.
`units` are long-term memory units ready to be distilled into digest. Each has `name`, `bucket`, `summary`, and `paths`.
A run returns at most `max_units`; extraction merges cross-file evidence for the same abstraction and drops passing
mentions, per-file summaries, and weak candidates without reusable value. `bucket` may only be `procedure`, `personal`,
or `wiki`; unknown values are routed to `wiki`.
`topics` are proactive-interest candidates for the day. They contain `title`, `reason`, `evidence`, `keywords`, and
`paths` and are filtered again in the Topics stage.
If there are no changed files, Extract succeeds with no units; Integrate then has no unit work, Topics preserves any
existing target-day topics, and Finish still performs its normal catalog summary. If files changed but no LLM is
configured, Extract fails because extraction requires an LLM.
### 2. Integrate
`dream_integrate_step` invokes an agent independently for each unit and integrates that unit into one digest node. It
exposes these tools to the agent:
```text
node_search, read, frontmatter_read, write, edit, frontmatter_update
```
This stage carries the core responsibility of `auto_link`. It first uses `node_search` to recall similar or related
nodes at digest-node granularity, decides whether to create or update a node, and finally writes sources and related
digest nodes as wikilinks. See [Auto Link](./auto_link.md) for the recall, deduplication, and edge-writing rules.
Extract is the gate for deciding whether material is worth remembering, so Integrate has no `SKIP` action: each admitted
unit must land in exactly one digest node. Creates and updates must retain provenance and weave related digest links
into contextual sentences; bare wikilinks and standalone relationship fields are not valid output.
There are four integration actions:
| Action | Meaning |
|---------------|--------------------------------------------------------------------------------|
| `CREATE` | No equivalent abstraction exists; create a new digest node. |
| `CORROBORATE` | The same memory appeared again; append a source or strengthen the description. |
| `REFINE` | New material adds boundaries, steps, prerequisites, applicability, or detail. |
| `CORRECT` | New material corrects errors, omissions, or conflicts in the existing node. |
Successfully integrated units are recorded in `integrate_results`. Failed units enter `failed_units`, and their source
paths enter `failed_paths`. The Finish stage does not checkpoint failed paths, ensuring that they can be retried later.
### 3. Topics
`dream_topics_step` turns topic candidates from Extract into the final `daily/<date>/interests.yaml` for the day.
It reads:
```text
daily/<date>/interests.yaml
daily/<each of the previous topic_diversity_days dates>/interests.yaml
```
Existing topics from the same day are preserved, while similar topics from the previous `topic_diversity_days` days are
deduplicated. At most three topics are written by default. With an LLM configured, the LLM selects topics that are more
specific, actionable, and non-repetitive. Without an LLM, the step falls back to local normalization and deduplication.
Example output format. See [Proactive](./proactive.md) for the interface that reads this file:
```yaml
date: 2026-06-20
topic_count: 3
diversity_days: 7
topics:
- title: Quality regression in the memory retrieval pipeline
reason: The user has recently made repeated changes to search, node_search, and dream integration.
evidence: daily/2026-06-20/session.md
keywords:
- memory search
- auto dream
paths:
- daily/2026-06-20/session.md
```
### 4. Finish
`dream_finish_step` completes the run:
1. Write successfully processed changed paths to `file_catalog: dream`.
2. Also write the target `daily/<date>/interests.yaml` and every refreshed day-index page in the scan window to the
catalog.
3. Persist the dream catalog if there were upserts or deletions.
4. Return a summary containing counts for scanned, changed, integrated, topics, checkpoints, and related values.
Failed paths are not checkpointed. The next `auto_dream` run therefore continues to treat them as changed inputs until
integration succeeds.
## Running Auto Dream
CLI:
```bash
reme auto_dream date=2026-06-20
```
With caller guidance:
```bash
reme auto_dream date=2026-06-20 hint="Prioritize engineering decisions and long-term preferences"
```
Override the default scan window and unit cap:
```bash
reme auto_dream date=2026-06-20 scan_days=3 max_units=8
```
The same set of steps can also be placed in a `cron` Job, for example to run every morning:
```yaml
jobs:
daily_auto_dream:
backend: cron
cron: "30 3 * * *"
steps:
- backend: dream_extract_step
file_catalog: dream
- backend: dream_integrate_step
- backend: dream_topics_step
- backend: dream_finish_step
file_catalog: dream
```
## Important Boundaries
`auto_dream` consumes only daily inputs and does not rewrite daily bodies. Daily preserves facts and the original
situation; digest is the abstracted long-term memory layer.
`digest` is not a copy of the source text. Its body should preserve reusable abstractions, while a Sources section
points back with contextual sentences such as `The decision was recorded in [[daily/<date>/decision.md]].` Links follow
the workspace-relative wikilink semantics described in
[Memory as File](./memory_as_file.md).
`auto_dream` does not invent an overview from nothing. Only content that actually appears in daily input and is
extracted as a unit or topic can enter digest or `interests.yaml`.
The complete flow depends on an LLM for Extract and Integrate. Topics can perform local deduplication without an LLM,
but that does not mean the full dream flow can run offline.

View file

@ -1,146 +0,0 @@
# Auto Link
In the current implementation, `auto_link` is not a separately registered Job. It is a capability of the Integrate stage
in
`auto_dream`: when `dream_integrate_step` writes a memory unit to `digest/`, it also recalls digest nodes, makes a
deduplication decision, links sources, and weaves wikilinks to related nodes into the result.
For the complete dream flow, see [Auto Dream](./auto_dream.md). For general wikilink, frontmatter, and
workspace-relative path semantics, see [Memory as File](./memory_as_file.md). For question-answering retrieval, see
[Memory Search](./memory_search.md).
## Where It Runs
The default `auto_dream` flow is:
```yaml
auto_dream:
steps:
- dream_extract_step
- dream_integrate_step # where auto_link actually happens
- dream_topics_step
- dream_finish_step
```
The Integrate stage processes each unit independently. A unit is written to exactly one target digest node, but that
node may link to multiple sources and multiple related digest nodes.
## Goals
`auto_link` addresses graph quality at write time:
| Problem | Handling |
|------------------------------------------------|----------------------------------------------------------------------|
| The same memory already exists | Recall and update the existing node instead of creating a duplicate. |
| New and existing material are related | Write workspace-relative wikilinks into the body. |
| A digest node is disconnected from its sources | Add daily/resource links under a `## Sources` section. |
| A node contains only isolated prose | Add links to related digest nodes on both CREATE and UPDATE. |
## Toolchain
`dream_integrate_step` exposes these tools to the agent:
```text
node_search
read
frontmatter_read
write
edit
frontmatter_update
```
`node_search` is digest-only node retrieval designed for dream integration. It returns node-level signals such as the
digest node's `path` and the `name` and `description` from frontmatter. It does not expand the body and does not perform
the link expansion used by ordinary search.
`read` and `frontmatter_read` are used only for candidates that may be relevant, avoiding expansion of every recalled
result into a large context.
## Linking Flow
### 1. Recall candidate nodes
The agent first calls `node_search` with the unit's triggers, verbs, nouns, synonyms, and possible failure modes. Broad
recall, for example `limit=20-30`, is recommended by default because this step serves both deduplication and link
discovery.
Recalled results are internally classified into three groups:
| Classification | Meaning | Next action |
|--------------------|---------------------------------------------------------------------------------------------------------|---------------------------|
| `same_abstraction` | The trigger or underlying abstraction is the same, with substantial content overlap. | Use as the UPDATE target. |
| `related` | An adjacent process, prerequisite, failure mode, concept, preference, or upstream/downstream knowledge. | Write a body wikilink. |
| `unrelated` | Only superficially similar or unrelated. | Ignore. |
### 2. Choose a write action
Every unit must select one action:
| Action | Linking semantics |
|---------------|-------------------------------------------------------------------------------------------------------------------------------|
| `CREATE` | Write a new `digest/<bucket>/<slug>.md` and add source and related-node links to its body. |
| `CORROBORATE` | The same abstraction appeared again; append its source link and strengthen the description when needed. |
| `REFINE` | New material extends the existing node; insert the additional content in the appropriate section and preserve existing links. |
| `CORRECT` | New material corrects the existing node; use source links to identify the basis for the correction. |
An UPDATE should be additive whenever possible: do not delete existing wikilinks or source entries. This prevents later
graph indexing and retrieval from losing edges.
### 3. Write source edges
Source edges are ordinary wikilinks grouped under a Markdown heading:
```markdown
## Sources
The decision was recorded in [[daily/2026-06-20/session.md]], while the supporting technical evidence comes from
[[resource/2026-06-20/paper.md]].
```
These edges represent the evidence behind a digest node. Plain-text descriptions do not count as source edges because
only wikilinks can be parsed reliably by the file graph. The surrounding sentence must explain what each source
supports; a bare wikilink line is not valid Integrate output. For the complete parsing rules, see
[Memory as File](./memory_as_file.md#wikilink).
### 4. Write relationships between digest nodes
Relationships between digest nodes use complete workspace-relative paths woven into natural prose:
```markdown
This design extends [[digest/wiki/hybrid-search.md]] and uses
[[digest/procedure/rebuild-index.md]]. Follow
[[digest/personal/team-review-preference.md]] during review.
```
## Bucket Differences
`auto_link` adjusts the shape of its output according to the unit bucket:
| Bucket | Writing focus |
|-------------|-----------------------------------------------------------------------------------------------------------------------------|
| `procedure` | Write a runbook with triggers, steps, inputs, and failure modes. Link prerequisites, substeps, and related preferences. |
| `personal` | Write user-, team-, or project-specific facts and preferences. Link related projects, habits, and decision context. |
| `wiki` | Write general knowledge, principles, observations, and decision precedents. Link concepts, methods, and adjacent knowledge. |
Regardless of bucket, preserve source edges and weave recalled related digest nodes into the body whenever possible.
## Relationship to Search
`auto_link` uses `node_search`, not the question-answering `search`.
| Capability | Purpose |
|---------------|-----------------------------------------------------------------------------------------------------------|
| `search` | External question answering; returns chunks and can expand upstream/downstream link context. |
| `node_search` | Dream integration; recalls only digest node-level summaries for deduplication and related-link decisions. |
This boundary matters. The Integrate stage needs to decide whether the same abstraction already exists and which nodes
should be linked; it should not load large numbers of body chunks into context. [Memory Search](./memory_search.md)
handles question-oriented chunk retrieval, RRF fusion, and link expansion.
## Failure and Retry
If integration of a unit fails, `dream_integrate_step` records `failed_units` and `failed_paths`.
`dream_finish_step` does not checkpoint those source paths, so the next `auto_dream` run processes them again.
This makes auto_link writes retryable: a failure does not mark the input as complete or silently discard digest edges
that should have been created.

View file

@ -1,114 +0,0 @@
# Auto Memory
Auto Memory is ReMe's entry point for conversational memory. Within a target date, it uses `session_id` to find or update at
most one daily memory card, whose filename is a concise topic or event name chosen by the Agent. The day's `YYYY-MM-DD.md`
page indexes those cards. It turns "we talked about it" into "it was remembered" while retaining a source conversation record
as evidence.
<p align="center">
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
</p>
For the general file semantics of `daily/`, `session/`, frontmatter, and wikilinks, see
[Memory as File](./memory_as_file.md).
```text
Conversation
├─ step 1: daily/YYYY-MM-DD/<generated_name>.md # one topic-named card per session
├─ step 2: daily/YYYY-MM-DD.md # daily index linking the cards
└─ source: session/dialog/<session_id>.jsonl # source conversation record
```
## What It Records
Auto Memory does not preserve a chat transcript as a running summary. It records information that may remain useful later:
- User preferences: preferred style, collaboration habits, and long-term requirements.
- Key facts: project background, important numbers, explicit conclusions, and constraints.
- Process decisions: what happened, why a choice was made, and which alternatives were rejected.
- Current state: what has been completed, what is blocked, and what comes next.
- Reusable experience: commands, workflows, diagnostic methods, and solutions.
## Write Location
Auto Memory writes distilled memories to `daily/`. Conversations from the same day first become individual cards:
Example directory:
```text
workspace/
daily/
2026-06-20.md
2026-06-20/
login-refactor-decision.md
retrieval-regression.md
```
The two files under the date directory are topic-named cards distilled from different conversations.
`daily/2026-06-20.md` is the index page for that day. Resource files enter the same daily memory layer; see
[Auto Resource](./auto_resource.md).
When a call includes `session_id`, Auto Memory uses it to find the corresponding card through frontmatter, while the Agent
chooses a readable filename through `name`:
```yaml
name: login-refactor-decision
session_id: session-a
source_conversation: "[[session/dialog/session-a.jsonl]]"
```
This keeps different conversations separate without forcing opaque IDs into filenames. An update locates the existing note by
`session_id` or `source_conversation`; if the Agent supplies a better frontmatter `name`, the system can rename the note and
retarget inbound wikilinks. To see what happened on a day, start with `YYYY-MM-DD.md`.
## Preserving the Original Information
The distilled daily note is optimized for readability; a filtered source conversation record is retained for trust and
verification.
While generating memory cards, Auto Memory also saves the source messages:
```text
session/
dialog/
session-a.jsonl
session-b.jsonl
```
Each daily note points to its corresponding conversation record. Saved messages omit tool-result blocks and base64 data
blocks, preventing recalled memory and binary payloads from being mistaken for user-provided evidence later.
## Message Timestamps
Auto Memory preserves each retained message's `created_at` in both the prompt and the source conversation JSONL. When importing historical
conversations or benchmark data, provide the actual occurrence time for every message so the model does not confuse event
time with execution time:
```bash
reme auto_memory \
session_id=locomo-session \
messages='[
{"role":"user","content":"Jon lost his job today.","created_at":"2023-01-19T08:00:00"},
{"role":"assistant","content":"I am sorry to hear that.","created_at":"2023-01-19T08:01:00"}
]'
```
For compatibility with common dataset schemas, `auto_memory` also checks `time_created`, `timestamp`, `createdAt`,
`timeCreated`, and `created_time` when `created_at` is absent. These fields may appear either at the top level of a message
or inside `metadata`.
When a call does not explicitly provide `date`, Auto Memory uses the latest valid `created_at` date in the messages. If no
message contains a valid timestamp, it falls back to the current date. Historical imports may also specify the
target date directly:
```bash
reme auto_memory \
session_id=locomo-session \
date=2023-01-19 \
messages='[{"role":"user","content":"Jon lost his job today."}]'
```
## What Happens Next
Auto Memory only creates memory in the daily layer. To distill this material further into long-term `digest/` nodes, use
[Auto Dream](./auto_dream.md). To search daily and digest content, use [Memory Search](./memory_search.md).

View file

@ -1,105 +0,0 @@
# Auto Resource `Beta`
Auto Resource is ReMe's entry point for interpreting resources and is currently in **Beta**. Resource files first enter
`resource/`, preferably under a date directory, and are then interpreted into daily resource cards. Each card's filename
comes from the LLM-generated frontmatter `name`, and `source_resource` links the card back to its original file.
<p align="center">
<img src="../figure/auto-memory-resource.svg" alt="ReMe Auto Memory and Auto Resource writing daily memory cards" width="92%">
</p>
For the general file semantics of workspace layers, `resource/`, and `daily/`, see
[Memory as File](./memory_as_file.md). For the flow that writes conversations to daily, see
[Auto Memory](./auto_memory.md).
```text
resource/[YYYY-MM-DD/]<resource_file>
├─ step 1: daily/YYYY-MM-DD/<generated_name>.md # interpreted resource card
├─ step 2: source_resource points to the original resource
└─ step 3: daily/YYYY-MM-DD.md # daily index linking the cards
```
## What It Records
Auto Resource does more than copy file content. It extracts information that will make the resource easier to retrieve
and understand later:
- Core content: what the resource is mainly about.
- Structure: its sections, tables, fields, and data organization.
- Key details: important numbers, names, dates, and conclusions.
- Context and purpose: why the resource exists and how it relates to current work.
- Actionable items: tasks, deadlines, and follow-up work.
In short, it turns "a file was archived" into "the resource is usable."
## Original Resource Entry Point
Auto Resource uses `resource/` as the entry point for source material. Date directories are recommended, and their date
determines which daily memory layer receives the interpreted card. A file directly under `resource/` is also supported
and uses today in the application timezone.
Example directory:
```text
workspace/
resource/
quick-note.txt # enters today's daily layer
2026-06-20/
market-report.md
meeting-notes.csv
```
The current Beta version is best suited to text-based resources such as `md`, `txt`, `json`, `jsonl`, `csv`, `yaml`, and
`html`.
## Resource Cards
Each resource file produces one daily resource card. The system initially uses the resource file's stem as a temporary
path. After the agent writes the card, the file is renamed according to its frontmatter `name`:
```text
resource/2026-06-20/market-report.md
daily/2026-06-20/market-report-highlights.md
```
The resource card links to the original file through frontmatter:
```yaml
source_resource: "[[resource/2026-06-20/market-report.md]]"
```
When a resource changes, Auto Resource finds and updates the corresponding card through `source_resource`. When a
resource is deleted, its daily note is also removed. The older `daily/YYYY-MM-DD/<resource_stem>.md` naming convention
remains supported as a fallback.
## Daily Index
Resource cards enter the same daily memory layer as Auto Memory cards. The day's `YYYY-MM-DD.md` page acts as an index
and organizes those resource cards:
```text
daily/
2026-06-20.md
2026-06-20/
market-report-highlights.md
meeting-notes-summary.md
```
To review which resources were processed on a day, start with `YYYY-MM-DD.md`. To inspect what was distilled from one
resource, open its corresponding resource card.
## Preserving the Original Resource
The interpreted daily note is optimized for readability; the original resource is retained for trust and verification.
Auto Resource does not move the original file. It remains at its original path under `resource/`. Text resources can
therefore enter the daily memory flow while their source files stay in their original location.
## What Happens Next
Auto Resource only creates resource interpretations in the daily layer. To distill long-term knowledge from resources
into
`digest/`, use [Auto Dream](./auto_dream.md). The default live index covers daily cards and digest nodes. Run
`reme reindex`
when original resource files must also be directly searchable; see [Memory Search](./memory_search.md).

View file

@ -1,230 +0,0 @@
# Open Source and Contributing
ReMe is open source and hosted on GitHub:
**https://github.com/agentscope-ai/ReMe**
---
## How to Contribute
Thank you for your interest in ReMe. ReMe is a file-first, self-evolving memory system for agents. Contributions are
welcome through issue reports, documentation improvements, additional tests, bug fixes, and new capabilities.
If this is your first time running ReMe locally, start with [Quick Start](./quick_start.md). If your change affects
runtime layers, Jobs, Steps, or components, read [ReMe Framework](./framework.md). If it affects workspace directories,
frontmatter, wikilinks, or chunking, read [Memory as File](./memory_as_file.md).
### 1. Before You Begin
Before investing in an implementation:
- Check [Open Issues](https://github.com/agentscope-ai/ReMe/issues) for an existing issue or discussion.
- If a related issue is still open, comment that you would like to work on it to avoid duplicate effort.
- If no issue exists, create one describing the context, expected behavior, possible implementation, and scope of
impact.
- For larger feature changes, align with maintainers on interfaces, configuration, compatibility, and test strategy
before submitting an implementation.
### 2. Local Development Environment
The core ReMe code is located in:
- `reme/`: Python package source, including configuration, components, services, Jobs, Steps, schemas, and utilities.
- `pyproject.toml`: project metadata, dependencies, optional dependencies, command entry points, and test configuration.
- `tests/`: unit and integration tests.
The project requires Python 3.11 or later. A virtual environment is recommended:
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e reme_studio -e ".[dev,full]"
cd reme_studio
npm ci
npm run build:static
cd ..
pre-commit install
```
### 3. Development Model
Before developing ReMe code, read [ReMe Framework](./framework.md). New or modified core capabilities should follow the
layers and call chain described there:
```text
CLI / Client -> Service -> Application -> Job -> Step -> Component / Workspace
```
In practice:
- Capabilities exposed to users or external systems should normally be orchestrated by a Job, then exposed by a Service
as a CLI-, HTTP-, or MCP-callable interface.
- Reusable infrastructure belongs in `reme/components/`, with dependencies declared through `BaseComponent.bind()`.
- Atomic business operations belong in `reme/steps/` and access the file store, agent wrapper, catalog, LLM, and other
components through `BaseStep.Ref`.
- Request, response, and persistent data structures belong in `reme/schema/` or `reme/enumeration/`. Do not scatter
implicit structures through Step implementations.
- Configuration-driven defaults belong in `reme/config/default.yaml`, and the default configuration must remain runnable
and testable.
When adding a Step or Job, pay particular attention to these conventions:
- Register implementations with `@R.register("<backend_name>")`. Registration names should be stable, clear, and match
the configured `backend`.
- After adding a Step file, make sure its package `__init__.py` imports the module; otherwise, the registry will not
load it.
- A Step should perform one atomic business operation. Cross-step flows belong in Job configuration or a dedicated
orchestration Step.
- A Job composes Steps and selects normal, streaming, background, or scheduled execution. `enable_serve` controls
whether it is externally exposed.
- When a Step needs components, prefer `BaseStep.Ref`. Do not reconstruct global components inside a Step or bypass
`ApplicationContext`.
- File, index, graph, frontmatter, and wikilink behavior must preserve consistent workspace-relative path semantics.
- Add fast tests under `tests/unit/` for new capabilities. Put cross-component, LLM, embedding, or service behavior
under
`tests/integration/` when appropriate.
### 4. Code and Documentation Changes
Choose the appropriate entry point for the type of change:
| Change type | Primary location | Guidance |
|-----------------------------------|-------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------|
| Configuration or startup behavior | `reme/config/`, `reme/application.py`, `reme/reme.py` | Keep the default configuration runnable and avoid breaking existing CLI, HTTP, and MCP entry points. |
| Component capability | `reme/components/` | Reuse `BaseComponent`, the registry, and context objects. |
| Job or Step | `reme/components/job/`, `reme/steps/` | Follow the Job -> Step model in [ReMe Framework](./framework.md), keep request and response schemas clear, and add corresponding tests. |
| Data structure | `reme/schema/`, `reme/enumeration/` | Preserve serialization compatibility and existing frontmatter and wikilink semantics. |
| Utility | `reme/utils/` | Keep function boundaries small and cover edge cases with unit tests. |
| User documentation | `docs/en/`, `README.md` | Update documentation when user-visible behavior changes. |
If a change involves an LLM, embeddings, an external service, file watching, or a background task, also describe its
dependencies, failure behavior, and local validation method.
### 5. Commit Message Format
Use [Conventional Commits](https://www.conventionalcommits.org/) to keep history clear.
Format:
```text
<type>(<scope>): <subject>
```
Common types:
- `feat`: new feature
- `fix`: bug fix
- `docs`: documentation only
- `style`: code-style change with no behavior change
- `refactor`: refactoring that neither fixes a bug nor adds a feature
- `perf`: performance improvement
- `test`: add or update tests
- `chore`: build, tooling, or maintenance work
Examples:
```bash
feat(search): add link expansion option
fix(file-graph): handle pending wikilinks after move
docs(memory): update auto memory guide
test(config): cover default yaml parsing
chore(pre-commit): update lint hooks
```
### 6. Pull Request Titles
PR titles should use the same format:
```text
<type>(<scope>): <description>
```
Requirements:
- Use `feat`, `fix`, `docs`, `test`, `refactor`, `chore`, `perf`, `style`, `build`, or `revert` as the type.
- Use lowercase letters, numbers, hyphens, or underscores for the scope.
- Keep the description short and state the actual effect of the PR.
Examples:
```text
feat(auto-memory): persist source conversation metadata
fix(markdown): keep wikilink aliases during edit
docs(en): add contribution guide
```
### 7. Pre-submit Checks
Before committing or opening a PR, run at least:
```bash
pre-commit run --all-files
pytest
```
For a localized code change, start with a narrower test set:
```bash
pytest tests/unit/test_search_step.py
pytest tests/unit/test_reme_cli.py
```
If `pre-commit` modifies files automatically, commit those changes and rerun the checks until everything passes.
The current pre-commit configuration includes YAML/TOML/JSON validation, private-key detection, trailing-whitespace
checks,
`black`, `flake8`, `pylint`, and `pyroma`. The main formatting rules are:
- `black --line-length=120`
- `flake8 --max-line-length=120`
- `pylint --max-line-length=120`
Some integration tests may require an LLM, embeddings, or external service configuration. If you cannot run them
locally, state why they were skipped and what alternative validation you completed in the PR description.
### 8. Testing Requirements
Add tests according to the risk of the change:
- For a bug fix, first add a regression test that reproduces the issue.
- For a new Step, Job, or component, cover at least the main path and a failure path.
- For changes to shared logic such as indexes, graphs, wikilinks, frontmatter, or file operations, add edge cases.
- For changes to the CLI, services, or configuration parsing, cover the user-visible entry point.
- Documentation-only changes usually do not require new tests, but running `pre-commit run --all-files` is still
recommended.
Place tests according to the existing structure:
- `tests/unit/`: fast tests that require no real external service.
- `tests/integration/`: integration tests spanning components or requiring external configuration.
### 9. Documentation Contributions
When a change affects how users install, configure, invoke, or understand ReMe, update the documentation as well.
Documentation lives under:
```text
docs/
```
Documentation should:
- Use clear titles that directly identify a capability or flow.
- Provide commands that can be copied and run.
- Use real repository paths such as `reme/config/default.yaml`, `reme/steps/`, and `tests/unit/`.
- Describe default behavior according to the current code, `pyproject.toml`, and default configuration.
---
## Getting Help
- Bugs and feature requests: [GitHub Issues](https://github.com/agentscope-ai/ReMe/issues)
- Project home: [GitHub Repository](https://github.com/agentscope-ai/ReMe)
- Documentation site: [https://reme.agentscope.io](https://reme.agentscope.io)
---
Thank you for contributing to ReMe. Your improvements help make long-term memory for agents more readable, controllable,
and maintainable.

View file

@ -1,840 +0,0 @@
# ReMe Framework
## 1. Overview
The ReMe runtime can be understood as follows: **a configuration-driven Application assembles components and Jobs; the
Service exposes service-enabled Jobs to the CLI, HTTP, or MCP; and each Job executes its Steps in sequence**.
<p align="center">
<img src="../figure/framework-structure.svg" alt="ReMe framework structure: CLI, Service, Application, Job, Step, and Component" width="92%">
</p>
To run and use ReMe first, see [Quick Start](./quick_start.md). For workspace file semantics, see
[Memory as File](./memory_as_file.md). User-facing guides for retrieval, automatic memory, and proactive reading are
[Memory Search](./memory_search.md), [Auto Memory](./auto_memory.md), [Auto Resource](./auto_resource.md),
[Auto Dream](./auto_dream.md), and [Proactive](./proactive.md).
### Capability Boundary
ReMe v4 focuses on long-term memory: it distills conversations and resources into `daily/`, organizes them into
`digest/`, and exposes write, retrieval, and proactive-read capabilities through the CLI, HTTP, and MCP.
Single-session context-window management is outside the scope of ReMe v4. This includes compressing the current
conversation, injecting summaries, trimming tool output, or providing an independent `/compact` interface. Those
capabilities belong in the host agent framework. ReMe accepts conversations, resources, and file changes that have
already occurred and persists the information with long-term value.
```mermaid
flowchart LR
CLI["reme CLI<br/>reme/reme.py"] --> Client["Client<br/>http / mcp"]
Client --> Service["Service<br/>HTTP / MCP"]
Service --> App["Application<br/>reme/application.py"]
App --> Jobs["Jobs<br/>base / stream / background / cron"]
Jobs --> Steps["Steps<br/>reme/steps/**"]
Steps --> Ctx["RuntimeContext<br/>data + Response + stream queue"]
Steps --> Components["Components<br/>store / graph / index / llm / agent / catalog"]
Components --> Workspace["Workspace<br/>daily / digest / resource / metadata"]
```
Core layers:
| Layer | Main location | Responsibility |
|-------------|----------------------------|---------------------------------------------------------------------------------------------------|
| CLI | `reme/reme.py` | Parse commands; `start` launches the service; other actions call the service through a client. |
| Service | `reme/components/service/` | Register Jobs as HTTP endpoints or MCP tools. |
| Application | `reme/application.py` | Assemble configured objects, start them in dependency order, close them, and invoke Jobs. |
| Job | `reme/components/job/` | Orchestrate Steps and select normal, streaming, background, or scheduled execution. |
| Step | `reme/steps/` | Atomic business operations such as file I/O, retrieval, indexing, and self-evolution. |
| Component | `reme/components/` | Reusable infrastructure such as file_store, file_graph, keyword_index, and agent_wrapper. |
| Schema | `reme/schema/` | Data structures such as `Request`, `Response`, `FileChunk`, `FileNode`, and configuration models. |
| Config | `reme/config/` | Default YAML configuration and command-line override parsing. |
## 2. Directory Structure
```text
reme/
reme.py # CLI entry point
application.py # Application assembly and lifecycle
plugin.py # installed plugin contract and entry-point loader
config/
default.yaml # default service / jobs / components
config_parser.py # config=, dot notation, and env placeholder parsing
components/
component_registry.py # backend registry and application-local copies
base_component.py # ComponentMixin / BaseComponent / bind dependency declarations
runtime_context.py # context for one Job execution
job/ # BaseJob / StreamJob / BackgroundJob / CronJob
service/ # HTTP / MCP services
client/ # HTTP / MCP clients
file_store/ # file-index coordination layer
file_graph/ # wikilink graph
keyword_index/ # BM25 and other keyword indexes
file_chunker/ # Markdown / JSON / JSONL / generic text chunking
file_catalog/ # change checkpoints
as_llm/, as_embedding/ # model wrappers
agent_wrapper/ # AgentScope / Claude Code / Codex wrappers
steps/
base_step.py # BaseStep, Ref, dispatch_steps
common/ # version, help, health_check, status, chat
benchmark/ # LongMemEval / BEAM evaluation steps
cookbook/ # built-in cookbook support steps
file_io/ # read/write/edit/delete/move/frontmatter/daily
index/ # watch/init/update/search/traverse
evolve/ # auto_memory, auto_resource, auto_dream, proactive
transfer/ # upload/download
plugins/
auto-fin/ # independent example plugin distribution
daily_paper/ # independent paper-research plugin distribution
integrations/
claude_code/ # Claude Code adapter and marketplace
hermes_agent/ # Hermes Agent memory-provider adapter
```
The default workspace directories are defined by `ApplicationConfig`:
```text
<workspace_dir>/
metadata/ # persistent file_store, file_graph, keyword_index, file_catalog, and related state
session/ # source conversations used by memory workflows
mem_session/ # generated Agent wrapper sessions and configuration
resource/ # external resources
daily/ # lightly processed memory
digest/ # long-term digest memory
```
`Application.__init__()` first ensures that these directories exist, then initializes the service, components, and Jobs.
## 3. Startup and Call Chain
### 3.1 CLI
The entry point is `reme/reme.py::main()`:
```mermaid
flowchart LR
A["main()"] --> B["parse_args(*sys.argv[1:])"]
B --> C{action}
C -->|" start "| D["load_env()"]
D --> E["resolve_app_config(**kwargs)"]
E --> F["precheck_start(service)"]
F --> G["ReMe(**config).run_app()"]
C -->|" find_reme "| H["cli_find_reme()"]
C -->|" other actions "| I["call_server(action, **kwargs)"]
I --> J["R.get(ComponentEnum.CLIENT, backend)"]
J --> K["client(action=action, **kwargs)"]
```
Common commands:
```bash
reme start
reme start service.port=8181
reme version
reme search query="memory" limit=5
reme search query="memory" backend=mcp
```
Configuration parsing supports:
| Capability | Source | Description |
|------------------------|-------------------------|--------------------------------------------------------------------------|
| Default configuration | `resolve_app_config()` | Load `reme/config/default.yaml` when `config` is not specified. |
| Explicit configuration | `config=<name-or-path>` | Accept a built-in configuration name or a YAML/JSON file path. |
| Dot notation | `parse_dot_notation()` | For example, `service.port=8181`. |
| Environment variables | `_expand_env_vars()` | Support `${VAR}` and `${VAR:-default}`. |
| Value conversion | `_convert_value()` | Convert bool, int, float, JSON list/dict, and null values automatically. |
### 3.2 Service
`BaseService.run_app()` executes in this order:
Set the optional `service.jobs` list to restrict HTTP or MCP exposure to those job names. If omitted, all jobs with
`enable_serve: true` remain eligible; an empty list exposes none. The whitelist does not override `enable_serve: false`.
When the list is configured, a missing, disabled, unsupported, or invalid selected job fails service startup.
```mermaid
flowchart LR
A["Service.build_service(app)"] --> B["read app.context.jobs"]
B --> C{"enabled and selected by service.jobs?"}
C -->|yes| D["Service.add_job(job)"]
C -->|no| E["skip registration"]
D --> F["Service.start_service(app)"]
E --> F
F --> G["app.start() during lifespan"]
G --> H["Application starts jobs"]
```
HTTP service behavior:
| Job type | HTTP exposure |
|-------------------------------------------|---------------------------------------------------|
| Non-`StreamJob` with `enable_serve: true` | `POST /<job.name>` returning `Response` JSON. |
| `StreamJob` | `POST /<job.name>` returning `text/event-stream`. |
| `enable_serve: false` | No endpoint is registered. |
After registering Job endpoints, the HTTP service can also mount the ReMe Studio single-page application. The default is
`service.web_enabled=true`. Builds are resolved from `service.web_static_dir`, `REME_WEB_STATIC_DIR`, the optional
`reme_studio` package installed by the `web` and `core` extras, and source-tree locations such as
`reme_studio/dist-static`. If no `index.html` is found, only the frontend is skipped and the Job API remains available. The
Studio `GET` fallback does not replace existing `POST /<job.name>` routes.
MCP service behavior:
| Job type | MCP exposure |
|-------------------------------------------|-------------------------------------------------------------------|
| Non-`StreamJob` with `enable_serve: true` | Registered as an MCP tool. |
| `StreamJob` | Currently skipped and not registered. |
| `BackgroundJob` | Forces `enable_serve=False` at construction and is never exposed. |
MCP services can inject server-owned arguments with `injected_job_kwargs`; callers cannot override those arguments. Set
`tool_error_on_failure: true` to expose an unsuccessful ReMe `Response` as an MCP tool error.
## 4. Registry and Dependency Injection
### 4.1 Global Registry R
ReMe uses the process-wide singleton `R = ComponentRegistry()`. Every component, Job, and Step is registered with
`@R.register("name")`.
```python
from ...components import R
@R.register("version_step")
class VersionStep(BaseStep):
...
```
The registry key is:
```text
(component_type, register_name) -> class
```
`component_type` comes from a class attribute:
| Type | Class attribute |
|-----------|-----------------------------------------------------------|
| Step | `BaseStep.component_type = ComponentEnum.STEP` |
| Job | `BaseJob.component_type = ComponentEnum.JOB` |
| Service | `BaseService.component_type = ComponentEnum.SERVICE` |
| FileStore | `BaseFileStore.component_type = ComponentEnum.FILE_STORE` |
The same backend name can therefore exist under different component types. For example, `http` can be both a service
backend and a client backend.
`ComponentEnum` provides the built-in identifiers, but installed plugins may declare a new type with a namespaced
string such as `example.reranker`. Custom identifiers use lowercase letters and numbers separated by `.`, `_`, or `-`.
They are configured under `components` and participate in the same dependency ordering and lifecycle as built-ins.
### 4.2 Built-in and Plugin Registration
Built-in implementations populate the built-in registry through package imports. ReMe freezes that template after
bootstrap, and each `Application` receives a mutable copy. Runtime code resolves backends through the application's
registry rather than changing the process-wide template. ReMe then loads only the installed plugins explicitly named by
`plugins` in the resolved configuration. A plugin exposes its package through the `reme.plugins` Python entry-point
group. The package's `plugin.yaml` has two optional mappings: `backends` maps registration names to
`module:Class` targets, and `application_defaults` contributes a low-priority `ApplicationConfig` fragment. The
entry-point name is the plugin's identity.
Plugins are enabled explicitly through the application config's `plugins` list or a `plugins=[...]` CLI override.
Plugin registration therefore stays local to one application;
duplicate `(component_type, backend)` providers fail during assembly instead of overwriting each other.
The legacy Python `Plugin` descriptor and `reme.configs` entry points remain accepted during migration. Configuration
files can use `extends` to inherit another built-in, legacy plugin, or file-based configuration. See the independently
packaged [Auto Fin](../../plugins/auto-fin/README.md) and [Daily Paper](../../plugins/daily_paper/README.md) plugins.
Plugin packages are managed locally and remain separate from per-application activation:
```bash
reme plugins list
reme plugins install reme-auto-fin
reme plugins install reme-daily-paper
reme plugins show daily-paper
reme plugins validate daily-paper
reme plugins uninstall daily-paper
reme start plugins='["auto-fin","daily-paper"]'
```
These management commands use the current Python interpreter's pip and never run through an HTTP or MCP service.
### 4.3 Component.bind
Dependencies between components are declared with `BaseComponent.bind()`. At startup,
`Application._topological_order()` reads every component's `dependencies` and starts them in topological order.
```mermaid
flowchart LR
A["Component.__init__<br/>self.keyword_index = self.bind(...)"] --> B["Dependency placeholder"]
B --> C["Application._topological_order()"]
C --> D["component.start()"]
D --> E["_resolve_bindings()"]
E --> F["self.keyword_index = app_context.components[type][name]"]
F --> G["component._start()"]
```
Rules for `BaseComponent.bind(name, BaseClass, optional=True)`:
| Scenario | Behavior |
|-----------------------------------------|------------------------------------------------------------|
| `name` is empty | Return `None` and skip the dependency. |
| `app_context` exists | Look up `app_context.components[ctype][name]`. |
| Dependency missing and `optional=True` | Resolve to `None`. |
| Dependency missing and `optional=False` | Fail at startup. |
| Standalone mode | A private component can be created with `default_factory`. |
### 4.4 Step.Ref
Steps do not participate in component topological startup. They are created temporarily for each Job invocation. Steps
access components primarily through `BaseStep.Ref`:
```python
file_store: BaseFileStore = Ref(BaseFileStore, ComponentEnum.FILE_STORE)
agent_wrapper: BaseAgentWrapper = Ref(BaseAgentWrapper, ComponentEnum.AGENT_WRAPPER, optional=True)
```
Resolution priority:
```mermaid
flowchart LR
A["access self.file_store"] --> B{"same-named object in kwargs?"}
B -->|yes| C["use kwargs object"]
B -->|no| D{"same-named object in context.data?"}
D -->|yes| E["use context object"]
D -->|no| F["read name from kwargs['file_store']; default is default"]
F --> G["app_context.components[FILE_STORE][name]"]
```
A Step configuration can therefore specify:
```yaml
steps:
- backend: update_catalog_step
file_catalog: resource
```
Here, `file_catalog: resource` means to resolve the `file_catalog` component named `resource`.
## 5. Application Lifecycle
The Application converts configuration into runtime objects and starts and closes them in order.
```mermaid
flowchart LR
A["Application(**kwargs)"] --> B["ApplicationContext(**kwargs)<br/>parse ApplicationConfig"]
B --> C["_setup_workspace_directories()"]
C --> D["_init_service()"]
D --> E["_init_components()"]
E --> F["_init_jobs()"]
F --> G["run_app()"]
G --> H["service.run_app(app)"]
```
Startup order in `Application._start()`:
```mermaid
flowchart LR
A["create optional thread_pool"] --> B["topologically sort components"]
B --> C["start components"]
C --> D["start BaseJob"]
D --> E["start StreamJob"]
E --> F["start BackgroundJob"]
F --> G["start CronJob"]
```
During shutdown, objects in `_started_components` are closed in reverse order so dependents close before their
dependencies.
## 6. Job Model
A Job is the orchestration unit for an externally callable capability or background task. Jobs are configured under
`jobs:`
in `reme/config/default.yaml`.
### 6.1 BaseJob
`BaseJob` is the most common request-oriented Job:
```mermaid
flowchart LR
Caller["Caller"] --> Job["BaseJob<br/>job(**kwargs)"]
Job --> Ctx["RuntimeContext<br/>merged_kwargs"]
Ctx --> S1["Step 1<br/>await step(context)"]
S1 --> D1["read/write context.data / response"]
D1 --> S2["Step 2<br/>await step(context)"]
S2 --> D2["read/write context.data / response"]
D2 --> Resp["context.response"]
Resp --> Caller
```
Important source behavior:
| Source | Behavior |
|--------------------|----------------------------------------------------------------------------------|
| `_start()` | Parse each Step config from YAML into `(step_cls, params)`. |
| `_build_steps()` | Create new Step instances for every call, avoiding state shared across requests. |
| `__call__()` | Create a `RuntimeContext` and execute Steps sequentially. |
| Exception handling | Catch the exception, set `response.success=False`, and set `answer=str(e)`. |
### 6.2 StreamJob
`StreamJob` extends `BaseJob` but returns streaming chunks:
| Behavior | Description |
|-------------|------------------------------------------------------------|
| Context | Includes `stream_queue`. |
| Step output | Call `context.add_stream_string(text, ChunkEnum.CONTENT)`. |
| Exception | Write `ChunkEnum.ERROR`. |
| Completion | Always send a `DONE` chunk. |
### 6.3 BackgroundJob
`BackgroundJob` runs long-lived loops such as file watchers. Its constructor forces `enable_serve=False`.
```mermaid
flowchart LR
A["Application starts BackgroundJob"] --> B["_start() creates stop_event and task"]
B --> C["_run_with_supervisor()"]
C --> D["await self()"]
D --> E{"exception?"}
E -->|no, returned normally| F["finish"]
E -->|yes and supervisor = True| G["exponential backoff + jitter"]
G --> C
E -->|yes and supervisor = False| H["raise exception"]
I["close()"] --> J["stop_event.set()"]
J --> K["wait close_timeout; cancel on timeout"]
```
The default `BackgroundJob.__call__()` also executes configured Steps in sequence, but it does not swallow exceptions,
which allows the supervisor to restart the task.
### 6.4 CronJob
`CronJob` extends `BackgroundJob` with a `cron` expression:
```yaml
jobs:
nightly_dream:
backend: cron
cron: "0 3 * * *"
steps:
- backend: dream_extract_step
- backend: dream_integrate_step
- backend: dream_topics_step
- backend: dream_finish_step
```
The current implementation uses `croniter` to calculate the next trigger time. The timezone comes from
`app_config.timezone`.
### 6.5 Default Job Types
```mermaid
flowchart LR
Jobs["default.yaml jobs"] --> BG["background<br/>index_update_loop<br/>resource_watch_loop<br/>digest_watch_loop"]
Jobs --> Cron["cron<br/>dream_cron<br/>optimize_index_cron"]
Jobs --> Stream["stream<br/>chat"]
Jobs --> Base["base<br/>version / help / health_check / status / app_config<br/>search / node_search / traverse / graph_snapshot / reindex<br/>read / load / read_image / write / save / edit / delete / move / list / stat / frontmatter_*<br/>daily_list / daily_reindex / daily_write<br/>auto_memory / auto_memory_cc / auto_resource / auto_dream / proactive"]
```
## 7. Step Model
A Step is a concrete business action. Every Step extends `BaseStep` and implements `execute()`.
```mermaid
flowchart LR
A["Job._build_steps()"] --> B["Step.__init__()"]
B --> C["load prompt<br/>class-named YAML + prompt_dict override"]
C --> D["Step.__call__(context, **kwargs)"]
D --> E["clear Ref cache"]
E --> F["RuntimeContext.from_context()"]
F --> G["input_mapping"]
G --> H["execute()"]
H --> I["output_mapping"]
I --> J["return result"]
```
### 7.1 RuntimeContext
`RuntimeContext` is shared by all Steps within one Job invocation:
| Field | Description |
|----------------|----------------------------------------------------------------------------|
| `response` | Final `Response(answer, success, metadata)`. |
| `data` | Free-form dictionary containing input parameters and intermediate results. |
| `stream_queue` | Output queue for streaming Jobs. |
| `stop_event` | Stop signal for background Jobs. |
Common Step code:
```python
assert self.context is not None
query = self.context.get("query", "")
self.context["processed_query"] = query.strip().lower()
self.context.response.answer = "..."
self.context.response.metadata["key"] = "value"
return self.context.response
```
### 7.2 input_mapping / output_mapping
`BaseStep.__call__()` invokes `RuntimeContext.apply_mapping()` before and after execution:
```yaml
steps:
- backend: some_step
input_mapping:
user_query: query
output_mapping:
result: final_result
```
The semantics are to copy `context.data[source]` to `context.data[target]`.
### 7.3 dispatch_steps
Some Steps produce batches of events and dispatch them to other Steps. `BaseStep.dispatch_steps()` resolves and executes
child Steps according to configuration.
Example from the default configuration:
```yaml
index_update_loop:
backend: background
watch_dirs: [ daily_dir, digest_dir ]
watch_suffixes: [ md ]
steps:
- backend: init_changes_step
monitor_type: file_store
monitor_name: default
dispatch_steps: [ update_index_step ]
- backend: watch_changes_step
dispatch_steps: [ update_index_step ]
```
Flow:
```mermaid
flowchart LR
Init["init_changes_step"] --> Batch["changes batch"]
Watch["watch_changes_step"] --> Batch
Batch --> Dispatch["dispatch_steps(...)"]
Dispatch --> Update["update_index_step"]
Update --> Store["file_store"]
```
## 8. Components in the Default Configuration
Current default components in `reme/config/default.yaml`:
| ComponentEnum | Name | Backend | Description |
|-------------------|---------------------------------|--------------------------------------------------|--------------------------------------------------------------------------------|
| `service` | singleton | `http` | Default HTTP service. |
| `tokenizer` | `default` | `regex` | BM25 tokenizer. |
| `as_embedding` | `default` | Not configured by default; example uses `openai` | Provides the embedding model wrapper after uncommenting the example config. |
| `embedding_store` | `default` | Not configured by default; example uses `local` | Depends on `as_embedding: default` after uncommenting the example config. |
| `as_llm` | `default` | `${LLM_BACKEND:-openai}` | LLM model wrapper. |
| `agent_wrapper` | `default` | `agentscope` | AgentScope wrapper. |
| `agent_wrapper` | `claude_code` | `claude_code` | Claude Code wrapper. |
| `agent_wrapper` | `codex/codex_oauth` | `codex` | Codex wrappers for API-key and OAuth authentication. |
| `file_graph` | `default` | `local` | Wikilink graph. |
| `file_catalog` | `default/resource/digest/dream` | `local` | File-change checkpoints. |
| `file_chunker` | `markdown` | `markdown` | Markdown AST chunking. |
| `file_chunker` | `json/jsonl/default` | `json/jsonl/default` | JSON, JSONL, and generic text chunkers; generic text supports `txt` and `log`. |
| `keyword_index` | `default` | `bm25` | BM25 keyword index. |
| `file_store` | `default` | `local` | Combines file_graph and keyword_index; defaults to `embedding_store: ""`. |
Note that the `search` Step configuration contains `vector_weight`, but `file_store.default.embedding_store` is empty by
default. Vector retrieval is available only when the runtime configuration enables an embedding store.
## 9. Adding a Step
### 9.1 Minimal Step
Suppose you want to add a Step that converts input text to uppercase.
Create a file such as `reme/steps/common/uppercase.py`:
```python
from ..base_step import BaseStep
from ...components import R
@R.register("uppercase_step")
class UppercaseStep(BaseStep):
async def execute(self):
assert self.context is not None
text = self.context.get("text", "")
result = str(text).upper()
self.context["uppercase_text"] = result
self.context.response.answer = result
self.context.response.metadata["length"] = len(result)
return self.context.response
```
### 9.2 Registering the Step
Make sure `reme/steps/common/__init__.py` imports the new module. Add:
```python
from . import uppercase
```
The reason is that `@R.register("uppercase_step")` only executes after the module is imported.
### 9.3 Accessing Components
If a Step needs an existing component, prefer the Refs provided by `BaseStep`:
```python
class MySearchStep(BaseStep):
async def execute(self):
assert self.context is not None
results = await self.file_store.keyword_search(
self.context.get("query", ""),
limit=5,
)
...
```
Common attributes available directly:
| Attribute | Component resolved by default |
|----------------------|-------------------------------------|
| `self.as_llm` | `.model` from `as_llm: default`. |
| `self.agent_wrapper` | `agent_wrapper: default`; optional. |
| `self.file_catalog` | `file_catalog: default`; optional. |
| `self.file_store` | `file_store: default`. |
To select a non-default component from Job configuration:
```yaml
steps:
- backend: my_step
file_catalog: dream
```
### 9.4 Step Design Guidance
| Guidance | Reason |
|---------------------------------------------------------------------------------|-------------------------------------------------------------------------------|
| Read input from `context` and write intermediate results to `context`. | A multi-Step Job passes data through the same context. |
| Write the final result to `context.response`. | Services and clients consume the standard `Response`. |
| Do not store request-scoped state on a Step instance. | A Step is rebuilt for every Job call, and stateless Steps are easier to test. |
| A background loop that supports interruption should check `context.stop_event`. | `BackgroundJob.close()` relies on the stop event for graceful shutdown. |
| Call `add_stream_string()` only from a StreamJob. | A normal Job has no stream queue. |
### 9.5 Unit Test Example
A Step can be instantiated directly and passed a `RuntimeContext`:
```python
import pytest
from reme.components.runtime_context import RuntimeContext
from reme.steps.common.uppercase import UppercaseStep
@pytest.mark.asyncio
async def test_uppercase_step():
ctx = RuntimeContext(text="hello")
resp = await UppercaseStep()(ctx)
assert resp.answer == "HELLO"
assert ctx["uppercase_text"] == "HELLO"
```
## 10. Adding a Job
A Job usually requires no new Python class; configure existing Steps instead. Add a new Job backend only when a new
execution model is required.
### 10.1 Adding a Normal Request Job
Add the Job under `jobs:` in a YAML configuration:
```yaml
jobs:
uppercase:
backend: base
description: "Convert text to uppercase."
parameters:
type: object
properties:
text:
type: string
description: "input text"
required:
- text
steps:
- backend: uppercase_step
```
Start and call it:
```bash
reme start
reme uppercase text="hello"
```
Call chain:
```mermaid
flowchart LR
CLI["CLI<br/>reme uppercase text=hello"] --> HTTP["HTTP Client"]
HTTP --> Req["POST /uppercase"]
Req --> S["HttpService"]
S --> J["uppercase BaseJob<br/>job(text='hello')"]
J --> Step["uppercase_step<br/>await step(context)"]
Step --> Resp["context.response.answer = HELLO"]
Resp --> JSON["Response JSON"]
JSON --> CLIOut["CLI prints answer"]
```
### 10.2 Adding a Multi-Step Job
A Job can chain multiple Steps:
```yaml
jobs:
demo_echo:
backend: base
description: "Normalize query, then echo it."
parameters:
type: object
properties:
query:
type: string
default: ""
min_score:
type: number
default: 0.5
steps:
- backend: demo_echo_step1
- backend: demo_echo_step2
```
The first Step writes:
```text
context["processed_query"]
context["adjusted_min_score"]
```
The second Step reads those fields and writes the final `response`.
### 10.3 Adding a Stream Job
Use `backend: stream` in configuration:
```yaml
jobs:
stream_uppercase:
backend: stream
description: "Stream uppercase text."
parameters:
type: object
properties:
text:
type: string
required:
- text
steps:
- backend: uppercase_prepare_step
- backend: uppercase_stream_step
```
Example streaming Step:
```python
from ..base_step import BaseStep
from ...components import R
from ...enumeration import ChunkEnum
@R.register("uppercase_stream_step")
class UppercaseStreamStep(BaseStep):
async def execute(self):
assert self.context is not None
for ch in self.context.get("uppercase_text", ""):
await self.context.add_stream_string(ch, ChunkEnum.CONTENT)
return self.context.response
```
### 10.4 Adding a Background Job
Use `backend: background` in configuration:
```yaml
jobs:
my_watch_loop:
backend: background
watch_dirs: [ daily_dir ]
watch_suffixes: [ md ]
steps:
- backend: init_changes_step
monitor_type: file_store
monitor_name: default
dispatch_steps: [ update_index_step ]
- backend: watch_changes_step
dispatch_steps: [ update_index_step ]
```
Characteristics of a background Job:
| Characteristic | Description |
|---------------------------------|--------------------------------------------------------------------|
| Not externally exposed | `BackgroundJob.__init__()` forces `enable_serve=False`. |
| Has a supervisor | Restarts with exponential backoff after an exception by default. |
| Has a stop event | Notifies the loop to exit during close. |
| Suitable for watching/consuming | File watching, queue consumption, and periodic long-running loops. |
### 10.5 Adding a Cron Job
Use `backend: cron` in configuration:
```yaml
jobs:
daily_auto_dream:
backend: cron
cron: "30 3 * * *"
steps:
- backend: dream_extract_step
file_catalog: dream
- backend: dream_integrate_step
- backend: dream_topics_step
- backend: dream_finish_step
file_catalog: dream
```
An invalid `cron` expression fails at startup.
### 10.6 When a New Job Backend Is Needed
Most use cases require only a new Step plus a YAML Job. Consider adding `reme/components/job/*.py` only in these cases:
| Requirement | New Job class? |
|---------------------------------------------------------------------|--------------------------------|
| Add a business command | No; use `backend: base`. |
| Chain existing steps | No; use `steps:`. |
| Need SSE/streaming output | No; use `backend: stream`. |
| Need a background loop | No; use `backend: background`. |
| Need cron scheduling | No; use `backend: cron`. |
| Need entirely new scheduling, concurrency, or transaction semantics | Yes; add a Job backend. |
Minimal shape of a new Job backend:
```python
from .base_job import BaseJob
from ..component_registry import R
@R.register("my_job_backend")
class MyJob(BaseJob):
async def __call__(self, **kwargs):
# custom scheduling logic
return await super().__call__(**kwargs)
```
Also ensure the module is imported by `reme/components/job/__init__.py`.

Some files were not shown because too many files have changed in this diff Show more