Find a file
xyf2020 23d4c96c15
refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422)
* chore(benchmark): isolate dataset/workspaces/results per benchmark

- Move shared benchmark/{datasets,memory_workspaces,results} into per-benchmark subdirs benchmark/<name>/{dataset,workspaces,results}
- Update beam/longmemeval config.yaml and run.py path defaults
- Relocate longmemeval download.py to benchmark/longmemeval/ (downloads into dataset/ subdir); inline dataset download docs into README
- Update .gitignore: benchmark/*/{dataset,workspaces,results}/
- Move result-{beam,longmemeval}.md to benchmark/results_md/ and drop result- prefix; update README links
- Fix stale path refs in llm_judge.py and logs/demo_search_format.py

* feat(benchmark): add read tool to agentic answer and update BEAM results

- Add 'read' to job_tools in BaseAgenticAnswerStep for file reading capability
- Document read tool usage in lme/agentic_answer.yaml system prompt
- Update result-beam.md with latest evaluation scores (OVERALL: 0.623/0.580)

* feat(auto_memory): add source line-number markers for note traceability

- Add _format_history hook in AutoMemoryStep with line-number annotation
- Override in BeamAutoMemoryStep to prefix each turn with [Ln] for citation
- Add session_file variable to prompt templates for source marker paths
- Simplify repeated extraction rules by referencing system prompt
- Enhance agentic_answer search strategy (multi-search, read tool hint)
- Add warning log on ReadStep failure

* feat(beam): enhance auto_memory with source markers and pilot ingest tooling

* refactor(beam): rename max_chunk_words to max_segment_words, drop one-off pilot scripts

* feat: add CompressorStep and search_v2 dual-mode session compression

- Add CompressorStep (reme/steps/evolve/compressor.py) for direct LLM
  text compression with optional query-guided relevance filtering
- Extend search_v2_step to support query-aware and query-independent
  session transcript compression via _compress injected kwargs
- Refactor _source_format.py: split into render_chunk_entries +
  join_chunk_entries; session chunks now render line-aligned with
  L<n>: prefixes for verbatim/compressed parity
- Add JOB_TOOLS and INJECTED_JOB_KWARGS to BaseAgenticAnswerStep for
  per-subclass tool and parameter injection
- LmeAgenticAnswerStep injects _search._compress payload to enable
  query-aware compression during benchmark evaluation
- Record compression ablation results in result-longmemeval.md
- Add unit tests for CompressorStep and search compression paths

* refactor(compress): relax session compression to lenient format-preserving strategy and update LME results

* refactor(benchmark): make session compression config-driven via compress_session flag

Move session-transcript compression from LME hard-coded injection to a
runtime context flag set by evaluation.compress_session in each
benchmark config. Compression is off by default for both BEAM and LME,
and BaseAgenticAnswerStep now conditionally injects the _search compress
payload only when the flag is truthy.

* feat(lme/auto_memory): add source attribution markers with line numbers

Add _format_history to annotate each turn with [Ln] line numbers and
expose {session_file} in prompts so the agent can emit bare wikilink-style
source markers like [[session/dialog/s1.jsonl#L1-L2,L5-L6]] at the end
of factual entries. Consolidate the per-prompt body/format rules into
references to the system prompt to avoid drift, and add frontmatter-
protection guidance for the edit tool.

* feat: improve agentic answer prompt and update beam 100K results

- Strengthen abstention rule: prohibit extrapolation from related but
  non-direct evidence
- Add multi-angle search after preliminary answer to check for
  conflicting/supplementary/updated information
- Add max-iteration fallback to 'Information not found'
- Update beam.md with 100K results (agentscope 2.0.4.post1, from scratch)
  including per-type token consumption and memory construction stats
- config.yaml: 100K dataset, 20 workers for BEAM evaluation
- run.py: add memory construction token usage tracking (default agent)
- Overall: 0.635 → 0.654 (+0.019), contradiction_resolution: 0.338 → 0.478
  (+0.140), abstention: 0.500 → 0.525 (+0.025)

* feat(read): add session-aware formatting for read tool and update BEAM eval

- Add truncate_session_output in _file_io.py to render jsonl session
  lines as [speaker @ time] content before byte-budget truncation
- Add read_step_format_session flag to ReadStep, honoring injected
  job kwargs (precedence) and YAML fallback
- Inject read_step_format_session=True into BaseAgenticAnswerStep
  so agentic answer reads render session transcripts human-readably
- Refine BEAM agentic_answer prompt: continue multi-angle search
  after preliminary answer, forbid fabrication/extrapolation
- Update BEAM config to 1M variant and add sequential 100K-eval /
  1M-build shell script
- Refresh benchmark/results_md/beam.md with latest results

* chore(config): disable expand_links in beam and lme search_v2 configs

* refactor(beam): drop one-off sequential 100K-eval-then-1M-build script

* fix(benchmark): add compressor job to beam config and fix BEAM clone instructions

- Add compressor job and compressor as_llm component to reme/config/beam.yaml
  (aligned with lme.yaml) so that compress_session: true works for BEAM
- Add graceful degradation guard in search_v2._compress_session_entries:
  when the compressor job is missing from the active config, log a warning
  and skip compression instead of raising 'Job compressor not found'.
  Skipped when there is no app_context so unit tests mocking run_job still
  drive compression behavior.
- Fix BEAM download instructions in README.md/README_ZH.md: add mkdir -p
  before cd benchmark/beam/dataset (the directory is gitignored and absent
  in a fresh clone)

* fix(steps): guard compressor exceptions and fix ReadStep boolean override

1. search_v2: catch per-entry exceptions from run_job('compressor') inside
   compress() so asyncio.gather never propagates a compressor failure (e.g.
   temporary LLM outage). The failing entry keeps its original body while
   remaining entries are still compressed, preserving already-retrieved
   search results.

2. read: replace 'context_value or yaml_value' with an existence check so
   that a runtime-injected False can explicitly disable a YAML-true
   read_step_format_session flag.

Add focused unit tests for both paths.

* fix(search_v2): use existence check for strict_date_filter boolean override

Replace 'context_value or yaml_value' with an existence-based check so
that a runtime-injected False can explicitly disable a YAML-true
strict_date_filter flag, consistent with the read_step_format_session fix.

* refactor(search): simplify strict_date_filter fallback to truthiness-or

* style(test): rename unused param to satisfy pylint W0613

* refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt

- Move shared benchmark/README, README_ZH, kill.sh, and results_md/*.md into
  per-benchmark subdirs (benchmark/beam/, benchmark/longmemeval/) so each
  benchmark owns its own docs, scripts, and result snapshots.
- Simplify lme/agentic_answer.yaml system prompt: drop verbose memory-system
  description, keep search strategy, draft tool, and answer rules concise.

* docs(benchmark): update LME README_ZH results to latest eval run

---------

Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
2026-08-06 15:13:52 +08:00
.github/workflows feat: simplify wikilink semantics and support line anchors (#412) 2026-08-05 11:47:50 +08:00
benchmark refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422) 2026-08-06 15:13:52 +08:00
cookbook feat: add Auto Fin cookbook and managed outbound proxy support (#392) 2026-07-25 18:09:39 +08:00
docs feat: simplify wikilink semantics and support line anchors (#412) 2026-08-05 11:47:50 +08:00
plugins Revert "feat(plugin): add ReMe integration for Codex (#372)" (#400) 2026-07-29 18:14:05 +08:00
reme refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt (#422) 2026-08-06 15:13:52 +08:00
scripts feat: add workspace web APIs and star growth report (#416) 2026-08-05 18:03:37 +08:00
skills feat: add Auto Fin cookbook and managed outbound proxy support (#392) 2026-07-25 18:09:39 +08:00
tests Revert "feat(backend): improve workspace support for web clients (#417)" (#419) 2026-08-05 23:17:08 +08:00
.gitignore feat(benchmark): enhance session memory retrieval and isolate benchmark assets (#409) 2026-08-05 19:23:42 +08:00
.pre-commit-config.yaml feat: add daily paper cookbook and DingTalk agent integration (#385) 2026-07-22 19:17:01 +08:00
AGENTS.md refactor(agent): unify agent subprocess env, sessions, skills, and MCP/service jobs (#382) 2026-07-20 23:52:14 +08:00
CLAUDE.md docs: restructure documentation and update content organization (#339) 2026-07-13 16:50:46 +08:00
example.env feat: add Auto Fin cookbook and managed outbound proxy support (#392) 2026-07-25 18:09:39 +08:00
LICENSE feat(reme_ai): implement memory retrieval and merging functionality 2025-08-25 16:10:53 +08:00
pyproject.toml feat(file_store): add ZvecLocalFileStore backend (#410) 2026-08-05 22:08:30 +08:00
README.md docs: link ExpG news entry to toolmemory README (#415) 2026-08-05 17:00:19 +08:00
README_ZH.md docs: link ExpG news entry to toolmemory README (#415) 2026-08-05 17:00:19 +08:00

ReMe Logo

Python Version PyPI Version PyPI Downloads GitHub commit activity License Documentation English 简体中文 GitHub Stars DeepWiki

agentscope-ai%2FReMe | Trendshift

An agent memory layer that turns conversations and resources into readable, editable, searchable Markdown memory.

Previous versions: 0.3.x · 0.2.x · MemoryScope

🧠 ReMe is a local-first memory layer for AI agents. It turns conversations and resources into file-based long-term memory, then continuously indexes, links, and consolidates that memory for future recall.

✨ Core Ideas

  • Memory as File: Markdown files with frontmatter and wikilinks serve as memory nodes that both users and agents can read and write directly.
  • Self-evolving knowledge base: Auto Memory, Auto Resource, and Auto Dream progressively transform conversations and resources into long-term memories, while automatically building wikilink relationships.
  • Progressive hybrid search: ReMe combines wikilinks, BM25, and embeddings for hybrid retrieval across keyword matching, semantic recall, and relationship expansion.
  • Agent-friendly integration: SKILL.md + CLI integration makes it easy for different agents to read, write, maintain, and reuse memory.

ReMe Design Philosophy

🔭 Use Cases

  • Personal assistants: Give personal assistants such as QwenPaw, OpenClaw, and Hermes a user-editable long-term memory layer.
  • Coding agents: Preserve coding style, project background, repository decisions, and workflow experience across sessions when integrating with coding agents such as Claude Code.
  • LLM Wiki: Turn conversations, notes, and resources into a searchable, traceable, and linked Markdown knowledge base that both users and agents can maintain.
  • Self-evolving agents: Support agents that learn from experience by saving successful paths, failed attempts, reusable procedures, and periodic reflections as memory.

📰 News

🚀 Quick Start

Installation

ReMe requires Python 3.11+.

Install from pip:

pip install "reme-ai[core]"

Install from source:

git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
pip install -e ".[core]"

Environment Variables

Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are disabled by default, so the default setup does not start an embedding model or require an embedding API key.

cat > .env <<'EOF'
# Optional: used only after embedding components are explicitly enabled in the config.
# EMBEDDING_API_KEY=sk-xxx
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1

# Required for auto_memory, auto_resource, and auto_dream.
LLM_API_KEY=sk-xxx
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
EOF

Basic file operations, BM25 search, wikilink traversal, and reading proactive topics can run without LLM credentials.

Note

To enable embedding-based semantic retrieval, uncomment components.as_embedding and components.embedding_store in reme/config/default.yaml, then change components.file_store.default.embedding_store from "" to default. See the memory search guide for details.

Start the Service

reme start

The default service address is 127.0.0.1:2333. If the port is occupied, specify another port:

reme start service.port=8181
# reme start workspace_dir=/tmp/reme-demo service.port=8181

After startup, check the service status. If you use a custom port, replace 2333 in the URL below with that port.

reme version
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'

5-Minute Memory Demo

With the service running, write a memory node, let ReMe index it, then retrieve it:

reme write \
  path=digest/wiki/quick-start-demo \
  name="Quick Start Demo" \
  description="A first ReMe memory node" \
  content="# Quick Start Demo

ReMe stores agent memory as readable Markdown.

Related: [[digest/wiki/memory-as-file.md]]"

reme search query="agent memory markdown" limit=5
reme read path=digest/wiki/quick-start-demo start_line=1 end_line=20

The generated file is ordinary Markdown with frontmatter:

---
name: Quick Start Demo
description: A first ReMe memory node
---

# Quick Start Demo

ReMe stores agent memory as readable Markdown.

Related: [[digest/wiki/memory-as-file.md]]

🧑‍🍳 Cookbooks

Cookbooks are optional, end-to-end workflows assembled from ReMe jobs and steps. They are not enabled by the default configuration; select the cookbook's standalone configuration when starting ReMe. Each new cookbook will be added as another row in this table.

Cookbook Capability
Daily Paper Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief.
Auto Fin Match CLS events to liquid ETFs, study historical reactions, and generate file-native research reports.

📁 Memory System

Memory as File, File as Memory.

ReMe treats memory as files, progressively processing raw conversations and external resources from session/ and resource/ into daily/, then consolidating them into reusable long-term memory nodes under digest/.

Directory Structure

<workspace_dir>/
├── metadata/       # Persistent system state such as indexes, graphs, and catalogs
├── session/        # Raw conversations and agent sessions
│   ├── dialog/
│   │   └── <session_id>.jsonl
│   ├── agentscope/
│   └── claude_code/
├── resource/            # External raw materials
│   └── YYYY-MM-DD/
│       └── <resource>.<ext>
├── daily/               # Lightly processed memory: daily facts, conversation summaries, resource readings
│   ├── YYYY-MM-DD.md
│   └── YYYY-MM-DD/
│       ├── <session_event>.md
│       ├── <resource_stem>.md
│       └── interests.yaml
└── digest/              # Long-term memory: personal facts, procedural experience, knowledge nodes
    ├── personal/
    │   └── {topic/event}.md
    ├── procedure/
    │   └── {topic/event}.md
    └── wiki/
        └── {topic/event}.md

ReMe file-based memory system overview

🧭 Memory Design Philosophy

Capture raw dialogs and resources, refine them into long-term preferences, reusable experience, and valuable knowledge, while keeping the result editable by humans and agents.

Automatic Memory Flow

ReMe follows a capture → index → consolidate → recall loop. Conversations and resources first become daily memory cards; background jobs keep files searchable; auto_dream distills stable knowledge into digest/; agents recall memory through search, wikilinks, or proactive topics.

Capability Entry point What it does Output
auto_memory Agent hook or reme auto_memory Distills useful conversation facts while preserving the raw session. session/dialog/*.jsonl, daily/<date>/<session>.md
auto_resource Resource watcher or reme auto_resource Turns files under resource/<date>/ into source-linked daily cards. daily/<date>/<resource-card>.md
auto_index Background watcher or reme reindex Maintains chunks, the BM25 index, the wikilink graph, and the optional embedding index. Searchable daily/, digest/, and resource/ content
auto_dream dream_cron or reme auto_dream Consolidates changed daily cards into long-term personal, procedure, and wiki memory. digest/**, daily/<date>/interests.yaml
proactive reme proactive before an agent decides to act Reads topics generated by auto_dream; the host agent decides whether and how to mention them. Structured topics from daily/<date>/interests.yaml
Memory as File Auto Memory and Resource
Auto Dream and Proactive Auto Index and Memory Search

🤝 Agent-friendly Integration

ReMe runs as a local memory service and offers multiple integration paths: CLI, HTTP API, MCP server, and SDK. Different agents can choose the path that fits their runtime while sharing the same local memory workspace.

Agents Recommended path What works out of the box
QwenPaw Embed ReMe via the Python SDK. Reuse the app's own lifecycle and model config while keeping memory local and file-based.
Claude Code Start ReMe as an MCP service and install plugins/reme. MCP recall tools, a reme-memory skill, and a Stop hook that records sessions automatically.
Other CLI-capable agents (OpenClaw/Hermes/Codex) Copy or install skills/reme_memory/SKILL.md. Search/read/write memory and call auto_memory, auto_dream, and proactive via the CLI.

Integration demos

Auto Memory Auto Dream
QwenPaw QwenPaw Auto Memory demo QwenPaw Auto Dream demo
Claude Code Claude Code Auto Memory demo Claude Code Auto Dream demo

🛠️ ReMe Operations

ReMe operates the workspace through a unified job interface exposed by the CLI. Agents usually only need retrieval, reading, writing, editing, and automatic memory commands. Lower-level indexing, frontmatter, and file operation commands are mainly for maintenance, debugging, or advanced integration. Run reme help for the full job list.

Command Purpose
reme start Start the local ReMe service.
reme version / reme health_check Check package and component status.
reme status Show stateful data-component memory estimates and process RSS.
reme search Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled.
reme read / reme write / reme edit Inspect and maintain Markdown memory files.
reme auto_memory Turn conversation messages into daily memory cards. Requires LLM credentials.
reme auto_resource Interpret files under resource/ into daily resource cards. Requires LLM credentials.
reme auto_dream / reme proactive Consolidate daily memory into long-term digest and surface topics worth attention.
reme reindex Rebuild search and wikilink indexes from existing files.

🤝 Community and Support

  • Issues and requests: Check Open Issues first. If there is no related discussion, open a new issue with background, expected behavior, and impact scope.
  • Code contributions: Before making changes, read the contribution guide. Source, schemas, and tests are the authoritative architecture and extension guide.
  • Documentation contributions: Submit user-facing documentation changes to the unified documentation repository under reme/<version>/{en,zh}/.
  • Commit convention: Conventional Commits are recommended, for example feat(search): add link expansion option or docs(zh): update quick start.
  • Pre-submit checks: Before submitting a PR, try to run pre-commit run --all-files and pytest. If tests that depend on LLMs, embeddings, or external services cannot run, explain that in the PR.
  • Get help: Use GitHub Issues for bugs and feature requests. Project documentation is available at https://docs.agentscope.io/reme.

Contributors

Thanks to everyone who has contributed to ReMe:

Contributors

📄 Citation

@software{ReMe2026,
  title = {Remember me, Refine me: Memory Management Kit for Agents},
  author = {ReMe Team},
  url = {https://reme.agentscope.io},
  year = {2026}
}

⚖️ License

This project is open source under the Apache License 2.0. See LICENSE for details.