* chore(benchmark): isolate dataset/workspaces/results per benchmark
- Move shared benchmark/{datasets,memory_workspaces,results} into per-benchmark subdirs benchmark/<name>/{dataset,workspaces,results}
- Update beam/longmemeval config.yaml and run.py path defaults
- Relocate longmemeval download.py to benchmark/longmemeval/ (downloads into dataset/ subdir); inline dataset download docs into README
- Update .gitignore: benchmark/*/{dataset,workspaces,results}/
- Move result-{beam,longmemeval}.md to benchmark/results_md/ and drop result- prefix; update README links
- Fix stale path refs in llm_judge.py and logs/demo_search_format.py
* feat(benchmark): add read tool to agentic answer and update BEAM results
- Add 'read' to job_tools in BaseAgenticAnswerStep for file reading capability
- Document read tool usage in lme/agentic_answer.yaml system prompt
- Update result-beam.md with latest evaluation scores (OVERALL: 0.623/0.580)
* feat(auto_memory): add source line-number markers for note traceability
- Add _format_history hook in AutoMemoryStep with line-number annotation
- Override in BeamAutoMemoryStep to prefix each turn with [Ln] for citation
- Add session_file variable to prompt templates for source marker paths
- Simplify repeated extraction rules by referencing system prompt
- Enhance agentic_answer search strategy (multi-search, read tool hint)
- Add warning log on ReadStep failure
* feat(beam): enhance auto_memory with source markers and pilot ingest tooling
* refactor(beam): rename max_chunk_words to max_segment_words, drop one-off pilot scripts
* feat: add CompressorStep and search_v2 dual-mode session compression
- Add CompressorStep (reme/steps/evolve/compressor.py) for direct LLM
text compression with optional query-guided relevance filtering
- Extend search_v2_step to support query-aware and query-independent
session transcript compression via _compress injected kwargs
- Refactor _source_format.py: split into render_chunk_entries +
join_chunk_entries; session chunks now render line-aligned with
L<n>: prefixes for verbatim/compressed parity
- Add JOB_TOOLS and INJECTED_JOB_KWARGS to BaseAgenticAnswerStep for
per-subclass tool and parameter injection
- LmeAgenticAnswerStep injects _search._compress payload to enable
query-aware compression during benchmark evaluation
- Record compression ablation results in result-longmemeval.md
- Add unit tests for CompressorStep and search compression paths
* refactor(compress): relax session compression to lenient format-preserving strategy and update LME results
* refactor(benchmark): make session compression config-driven via compress_session flag
Move session-transcript compression from LME hard-coded injection to a
runtime context flag set by evaluation.compress_session in each
benchmark config. Compression is off by default for both BEAM and LME,
and BaseAgenticAnswerStep now conditionally injects the _search compress
payload only when the flag is truthy.
* feat(lme/auto_memory): add source attribution markers with line numbers
Add _format_history to annotate each turn with [Ln] line numbers and
expose {session_file} in prompts so the agent can emit bare wikilink-style
source markers like [[session/dialog/s1.jsonl#L1-L2,L5-L6]] at the end
of factual entries. Consolidate the per-prompt body/format rules into
references to the system prompt to avoid drift, and add frontmatter-
protection guidance for the edit tool.
* feat: improve agentic answer prompt and update beam 100K results
- Strengthen abstention rule: prohibit extrapolation from related but
non-direct evidence
- Add multi-angle search after preliminary answer to check for
conflicting/supplementary/updated information
- Add max-iteration fallback to 'Information not found'
- Update beam.md with 100K results (agentscope 2.0.4.post1, from scratch)
including per-type token consumption and memory construction stats
- config.yaml: 100K dataset, 20 workers for BEAM evaluation
- run.py: add memory construction token usage tracking (default agent)
- Overall: 0.635 → 0.654 (+0.019), contradiction_resolution: 0.338 → 0.478
(+0.140), abstention: 0.500 → 0.525 (+0.025)
* feat(read): add session-aware formatting for read tool and update BEAM eval
- Add truncate_session_output in _file_io.py to render jsonl session
lines as [speaker @ time] content before byte-budget truncation
- Add read_step_format_session flag to ReadStep, honoring injected
job kwargs (precedence) and YAML fallback
- Inject read_step_format_session=True into BaseAgenticAnswerStep
so agentic answer reads render session transcripts human-readably
- Refine BEAM agentic_answer prompt: continue multi-angle search
after preliminary answer, forbid fabrication/extrapolation
- Update BEAM config to 1M variant and add sequential 100K-eval /
1M-build shell script
- Refresh benchmark/results_md/beam.md with latest results
* chore(config): disable expand_links in beam and lme search_v2 configs
* refactor(beam): drop one-off sequential 100K-eval-then-1M-build script
* fix(benchmark): add compressor job to beam config and fix BEAM clone instructions
- Add compressor job and compressor as_llm component to reme/config/beam.yaml
(aligned with lme.yaml) so that compress_session: true works for BEAM
- Add graceful degradation guard in search_v2._compress_session_entries:
when the compressor job is missing from the active config, log a warning
and skip compression instead of raising 'Job compressor not found'.
Skipped when there is no app_context so unit tests mocking run_job still
drive compression behavior.
- Fix BEAM download instructions in README.md/README_ZH.md: add mkdir -p
before cd benchmark/beam/dataset (the directory is gitignored and absent
in a fresh clone)
* fix(steps): guard compressor exceptions and fix ReadStep boolean override
1. search_v2: catch per-entry exceptions from run_job('compressor') inside
compress() so asyncio.gather never propagates a compressor failure (e.g.
temporary LLM outage). The failing entry keeps its original body while
remaining entries are still compressed, preserving already-retrieved
search results.
2. read: replace 'context_value or yaml_value' with an existence check so
that a runtime-injected False can explicitly disable a YAML-true
read_step_format_session flag.
Add focused unit tests for both paths.
* fix(search_v2): use existence check for strict_date_filter boolean override
Replace 'context_value or yaml_value' with an existence-based check so
that a runtime-injected False can explicitly disable a YAML-true
strict_date_filter flag, consistent with the read_step_format_session fix.
* refactor(search): simplify strict_date_filter fallback to truthiness-or
* style(test): rename unused param to satisfy pylint W0613
* refactor(benchmark): isolate per-benchmark assets and simplify LME agentic prompt
- Move shared benchmark/README, README_ZH, kill.sh, and results_md/*.md into
per-benchmark subdirs (benchmark/beam/, benchmark/longmemeval/) so each
benchmark owns its own docs, scripts, and result snapshots.
- Simplify lme/agentic_answer.yaml system prompt: drop verbose memory-system
description, keep search strategy, draft tool, and answer rules concise.
* docs(benchmark): update LME README_ZH results to latest eval run
---------
Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
|
||
|---|---|---|
| .github/workflows | ||
| benchmark | ||
| cookbook | ||
| docs | ||
| plugins | ||
| reme | ||
| scripts | ||
| skills | ||
| tests | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| example.env | ||
| LICENSE | ||
| pyproject.toml | ||
| README.md | ||
| README_ZH.md | ||
An agent memory layer that turns conversations and resources into readable, editable, searchable Markdown memory.
Previous versions: 0.3.x · 0.2.x · MemoryScope
🧠 ReMe is a local-first memory layer for AI agents. It turns conversations and resources into file-based long-term memory, then continuously indexes, links, and consolidates that memory for future recall.
✨ Core Ideas
- Memory as File: Markdown files with frontmatter and wikilinks serve as memory nodes that both users and agents can read and write directly.
- Self-evolving knowledge base: Auto Memory, Auto Resource, and Auto Dream progressively transform conversations and resources into long-term memories, while automatically building wikilink relationships.
- Progressive hybrid search: ReMe combines wikilinks, BM25, and embeddings for hybrid retrieval across keyword matching, semantic recall, and relationship expansion.
- Agent-friendly integration: SKILL.md + CLI integration makes it easy for different agents to read, write, maintain, and reuse memory.
🔭 Use Cases
- Personal assistants: Give personal assistants such as QwenPaw, OpenClaw, and Hermes a user-editable long-term memory layer.
- Coding agents: Preserve coding style, project background, repository decisions, and workflow experience across sessions when integrating with coding agents such as Claude Code.
- LLM Wiki: Turn conversations, notes, and resources into a searchable, traceable, and linked Markdown knowledge base that both users and agents can maintain.
- Self-evolving agents: Support agents that learn from experience by saving successful paths, failed attempts, reusable procedures, and periodic reflections as memory.
📰 News
- [2026.08] - Experience-driven enhancement method of agent tool-use execution built on ReMe is available on arXiv:2608.03403.
- [2026.07] - Introduced optional Cookbooks: Daily Paper for paper discovery and analysis, and Auto Fin for file-native ETF event research based on CLS news and historical market reactions.
- [2026.07] - Our paper Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution has been accepted to Findings of ACL 2026.
🚀 Quick Start
Installation
ReMe requires Python 3.11+.
Install from pip:
pip install "reme-ai[core]"
Install from source:
git clone https://github.com/agentscope-ai/ReMe.git
cd ReMe
pip install -e ".[core]"
Environment Variables
Configure environment variables when you want LLM-powered memory evolution or embedding retrieval. Embeddings are disabled by default, so the default setup does not start an embedding model or require an embedding API key.
cat > .env <<'EOF'
# Optional: used only after embedding components are explicitly enabled in the config.
# EMBEDDING_API_KEY=sk-xxx
# EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
# Required for auto_memory, auto_resource, and auto_dream.
LLM_API_KEY=sk-xxx
LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
EOF
Basic file operations, BM25 search, wikilink traversal, and reading proactive topics can run without LLM credentials.
Note
To enable embedding-based semantic retrieval, uncomment
components.as_embeddingandcomponents.embedding_storeinreme/config/default.yaml, then changecomponents.file_store.default.embedding_storefrom""todefault. See the memory search guide for details.
Start the Service
reme start
The default service address is 127.0.0.1:2333. If the port is occupied, specify another port:
reme start service.port=8181
# reme start workspace_dir=/tmp/reme-demo service.port=8181
After startup, check the service status. If you use a custom port, replace 2333 in the URL below with that port.
reme version
curl -s http://127.0.0.1:2333/version -H 'Content-Type: application/json' -d '{}'
5-Minute Memory Demo
With the service running, write a memory node, let ReMe index it, then retrieve it:
reme write \
path=digest/wiki/quick-start-demo \
name="Quick Start Demo" \
description="A first ReMe memory node" \
content="# Quick Start Demo
ReMe stores agent memory as readable Markdown.
Related: [[digest/wiki/memory-as-file.md]]"
reme search query="agent memory markdown" limit=5
reme read path=digest/wiki/quick-start-demo start_line=1 end_line=20
The generated file is ordinary Markdown with frontmatter:
---
name: Quick Start Demo
description: A first ReMe memory node
---
# Quick Start Demo
ReMe stores agent memory as readable Markdown.
Related: [[digest/wiki/memory-as-file.md]]
🧑🍳 Cookbooks
Cookbooks are optional, end-to-end workflows assembled from ReMe jobs and steps. They are not enabled by the default configuration; select the cookbook's standalone configuration when starting ReMe. Each new cookbook will be added as another row in this table.
| Cookbook | Capability |
|---|---|
| Daily Paper | Discover and rank papers, analyze PDFs with an agent, and generate file-native notes and a five-minute brief. |
| Auto Fin | Match CLS events to liquid ETFs, study historical reactions, and generate file-native research reports. |
📁 Memory System
Memory as File, File as Memory.
ReMe treats memory as files, progressively processing raw conversations and external resources from session/ and
resource/ into daily/, then consolidating them into reusable long-term memory nodes under digest/.
Directory Structure
<workspace_dir>/
├── metadata/ # Persistent system state such as indexes, graphs, and catalogs
├── session/ # Raw conversations and agent sessions
│ ├── dialog/
│ │ └── <session_id>.jsonl
│ ├── agentscope/
│ └── claude_code/
├── resource/ # External raw materials
│ └── YYYY-MM-DD/
│ └── <resource>.<ext>
├── daily/ # Lightly processed memory: daily facts, conversation summaries, resource readings
│ ├── YYYY-MM-DD.md
│ └── YYYY-MM-DD/
│ ├── <session_event>.md
│ ├── <resource_stem>.md
│ └── interests.yaml
└── digest/ # Long-term memory: personal facts, procedural experience, knowledge nodes
├── personal/
│ └── {topic/event}.md
├── procedure/
│ └── {topic/event}.md
└── wiki/
└── {topic/event}.md
🧭 Memory Design Philosophy
Capture raw dialogs and resources, refine them into long-term preferences, reusable experience, and valuable knowledge, while keeping the result editable by humans and agents.
Automatic Memory Flow
ReMe follows a capture → index → consolidate → recall loop. Conversations and resources first become daily memory cards;
background jobs keep files searchable; auto_dream distills stable knowledge into digest/; agents recall memory
through search, wikilinks, or proactive topics.
| Capability | Entry point | What it does | Output |
|---|---|---|---|
auto_memory |
Agent hook or reme auto_memory |
Distills useful conversation facts while preserving the raw session. | session/dialog/*.jsonl, daily/<date>/<session>.md |
auto_resource |
Resource watcher or reme auto_resource |
Turns files under resource/<date>/ into source-linked daily cards. |
daily/<date>/<resource-card>.md |
auto_index |
Background watcher or reme reindex |
Maintains chunks, the BM25 index, the wikilink graph, and the optional embedding index. | Searchable daily/, digest/, and resource/ content |
auto_dream |
dream_cron or reme auto_dream |
Consolidates changed daily cards into long-term personal, procedure, and wiki memory. | digest/**, daily/<date>/interests.yaml |
proactive |
reme proactive before an agent decides to act |
Reads topics generated by auto_dream; the host agent decides whether and how to mention them. |
Structured topics from daily/<date>/interests.yaml |
|
|
|
|
|
|
🤝 Agent-friendly Integration
ReMe runs as a local memory service and offers multiple integration paths: CLI, HTTP API, MCP server, and SDK. Different agents can choose the path that fits their runtime while sharing the same local memory workspace.
| Agents | Recommended path | What works out of the box |
|---|---|---|
| QwenPaw | Embed ReMe via the Python SDK. | Reuse the app's own lifecycle and model config while keeping memory local and file-based. |
| Claude Code | Start ReMe as an MCP service and install plugins/reme. | MCP recall tools, a reme-memory skill, and a Stop hook that records sessions automatically. |
| Other CLI-capable agents (OpenClaw/Hermes/Codex) | Copy or install skills/reme_memory/SKILL.md. | Search/read/write memory and call auto_memory, auto_dream, and proactive via the CLI. |
Integration demos
| Auto Memory | Auto Dream | |
| QwenPaw |
|
|
| Claude Code |
|
|
🛠️ ReMe Operations
ReMe operates the workspace through a unified job interface exposed by the CLI. Agents usually only need retrieval,
reading, writing, editing, and automatic memory commands. Lower-level indexing, frontmatter, and file operation commands
are mainly for maintenance, debugging, or advanced integration. Run reme help for the full job list.
| Command | Purpose |
|---|---|
reme start |
Start the local ReMe service. |
reme version / reme health_check |
Check package and component status. |
reme status |
Show stateful data-component memory estimates and process RSS. |
reme search |
Retrieve memory with BM25 and wikilinks by default, plus vectors when enabled. |
reme read / reme write / reme edit |
Inspect and maintain Markdown memory files. |
reme auto_memory |
Turn conversation messages into daily memory cards. Requires LLM credentials. |
reme auto_resource |
Interpret files under resource/ into daily resource cards. Requires LLM credentials. |
reme auto_dream / reme proactive |
Consolidate daily memory into long-term digest and surface topics worth attention. |
reme reindex |
Rebuild search and wikilink indexes from existing files. |
🤝 Community and Support
- Issues and requests: Check Open Issues first. If there is no related discussion, open a new issue with background, expected behavior, and impact scope.
- Code contributions: Before making changes, read the contribution guide. Source, schemas, and tests are the authoritative architecture and extension guide.
- Documentation contributions: Submit user-facing documentation changes to the
unified documentation repository under
reme/<version>/{en,zh}/. - Commit convention: Conventional Commits are recommended, for example
feat(search): add link expansion optionordocs(zh): update quick start. - Pre-submit checks: Before submitting a PR, try to run
pre-commit run --all-filesandpytest. If tests that depend on LLMs, embeddings, or external services cannot run, explain that in the PR. - Get help: Use GitHub Issues for bugs and feature requests. Project documentation is available at https://docs.agentscope.io/reme.
Contributors
Thanks to everyone who has contributed to ReMe:
📄 Citation
@software{ReMe2026,
title = {Remember me, Refine me: Memory Management Kit for Agents},
author = {ReMe Team},
url = {https://reme.agentscope.io},
year = {2026}
}
⚖️ License
This project is open source under the Apache License 2.0. See LICENSE for details.