mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-09-30 01:52:29 +00:00
Some checks failed
Pre-commit / run (ubuntu-latest) (push) Has been cancelled
* feat(reme): 添加配置选项以启用或禁用个人资料功能 - 在 ReMe 初始化方法中添加 enable_profile 参数,默认值为 True - 根据 enable_profile 设置决定是否创建 profile 目录和设置 profile_dir - 在 PersonalSummarizer 中根据 enable_profile 条件性地添加个人资料相关工具 - 在 PersonalRetriever 中根据 enable_profile 条件性地添加 ReadAllProfiles 工具 - 修改 profile_path 属性以在禁用个人资料时返回 None - 修改 get_profile_handler 方法以在禁用个人资料时返回 None - 为 enable_profile 参数添加文档说明其用于云向量存储场景 * refactor(benchmark): 重构LongMemEval基准测试中的ReMe实例管理 - 移除未使用的shutil导入 - 将固定的ReMe实例改为每个问题创建独立实例以实现隔离 - 更新LLM配置名称从qwen3-max-think到qwen-max-t - 修改模型调用逻辑使用正确的model_name参数 - 添加qwen-flash和GPT-4o-mini等新模型配置 - 统一使用"User"作为用户名,通过集合名实现隔离 - 调整并发处理数从4降至1,批处理大小从10增至30 - 每个问题类型采样数从2增至4 - 添加异步上下文管理确保资源正确释放 * reformat 2 files * refactor(benchmark): 重构长记忆评估中的模型配置 - 将原有的 eval_model_name 替换为专门的 retrieve_model_name 用于检索操作 - 添加对 qwen-max 模型配置的支持 - 更新参数解析器以支持新的检索模型参数 - 修改最大并发数默认值从 1 提升到 4 - 调整样本数量默认值从 4 减少到 1 - 统一模型参数命名规范,区分摘要、检索和评估模型 - 优化内存处理器初始化逻辑,支持独立的检索模型配置 * fix(benchmark): 移除数据路径默认值并设为必填参数 - 将LongMemEval评估脚本中的data_path参数改为必需参数 - 将HaluMem评估脚本中的data_path参数改为必需参数 - 删除了硬编码的默认文件路径配置 - 强制用户显式指定数据集文件路径以避免路径错误 * Update __init__.py * Update __init__.py * fix(benchmark): 修复ReMe评估中的模型配置和空值处理问题 - 移除了retrieve_memory调用中不需要的llm_config_name参数 - 修复了长字符串打印的换行格式问题 - 添加了eval_result为空时的初始化处理 - 在accuracy评估中加入了eval_model_name参数传递 * style(benchmark): 格式化模型名称打印输出 - 移除了多行字符串中的换行符和多余空格 - 将模型名称信息合并为单行连续显示 - 保持了原有的打印格式和信息完整性 * docs(readme): 更新文档添加实验结果表格 - 在英文版 README 中添加 🧪 Experiments 章节 - 添加 LoCoMo 和 HaluMem 两个基准测试的结果表格 - 在中文版 README_ZH 中添加 🧪 实验 章节 - 添加 LoCoMo 和 HaluMem 测试集的实验配置说明 - 添加完整的实验数据对比表格和评估协议说明 * docs(readme): 更新文档中的内存系统链接 - 为基于文件的记忆系统添加锚点链接 - 为基于向量库的记忆系统添加锚点链接 - 修复英文文档中的链接格式 - 修复中文文档中的链接格式和空行问题 * docs(readme): update experimental results section in documentation - Remove outdated experimental data placeholder "Coming soon..." - Add complete evaluation results for LoCoMo and HaluMem benchmarks - Include detailed performance metrics tables for all memory methods - Update experimental settings description with ReMe backbone details - Align evaluation protocol information with LLM-as-a-Judge approach - Maintain consistent formatting between English and Chinese documentation * docs(benchmark): add quick start guides for halumem and longmemeval experiments - Created HaluMem experiment quick start guide with ReMe integration setup - Added detailed steps for installing ReMe environment using conda - Included repository cloning instructions for HaluMem benchmark - Provided complete command examples for running HaluMem experiments - Created LongMeMEval quick start guide with data download procedures - Added wget commands for downloading cleaned dataset files - Included evaluation script instructions for computing experiment statistics - Documented parameter configurations for different model types and batch sizes * docs(longmemeval): update quickstart guide documentation - Changed project name from Halumem to Longmemeval in title - Updated description to reference Longmemeval experiments instead of Halumem - Maintained existing ReMe integration instructions unchanged * chore(logger): add test comment to logger configuration - Added test comment in logger utility function - Removed duplicate log handling by keeping the remove() call * chore(logger): add test comment to logger configuration - Added test comment in logger utility function - Removed duplicate log handling by keeping the remove() call * feat(core): add file logging capability to application - Added log_to_file parameter to Application class constructor - Integrated log_to_file option in logger initialization - Updated ServiceContext to support file logging configuration - Modified init_logger function to conditionally enable file logging - Added log_to_file field to ServiceConfig schema - Updated ReMe class to include file logging option - Wrapped file logging setup in conditional check to prevent unnecessary operations * docs(benchmark): update HaluMem quickstart guide with dataset download instructions - Replace repository cloning with direct dataset download using curl - Add commands to download HaluMem-Medium.jsonl and HaluMem-Long.jsonl files - Include both official Hugging Face and mirror download sources - Update data path reference from nested directory to local data folder - Add dataset page link and mirror usage instructions for mainland China access * feat(memory): add profile retrieval tool and refactor profile management - Introduce RetrieveProfile tool for fetching specific user profiles - Refactor ProfileHandler to support both filesystem and vector backends - Add async methods to ProfileHandler with synchronous fallbacks - Update PersonalRetriever to support two-stage profile and memory retrieval - Enhance PersonalSummarizer with improved tool partitioning logic - Add profile_backend, profile_store_name, and profile_max_capacity configuration options - Replace direct ProfileHandler imports with get_profile_handler method - Implement profile search functionality with dedicated prompts and workflows - Add FileProfileBackend and VectorProfileBackend implementations - Update base memory tool with new profile configuration parameters * feat(profile): add custom profile collection name support - Add profile_collection_name parameter to Application constructor - Allow custom database collection name for vector profiles instead of default suffix - Update profile vector store configuration logic to use custom collection name - Modify _ensure_profile_vector_store_config to handle custom collection names - Update docstring with detailed parameter descriptions for profile configuration options * test(history): add single history id acceptance test for multiple mode - Add test case to verify multiple-mode history lookup accepts a single history_id string - Create FakeVectorStore stub with minimal implementation for ReadHistory tests - Return requested history node from vector store mock - Initialize ReadHistory tool with multiple mode enabled - Add pylint disable comment for protected access to vector store property * refactor(memory): update profile handler and vector tools with improved formatting and error handling - Add module docstring to profiles/__init__.py - Add pylint disable comments for no-name-in-module and missing-function-docstring - Format long error message in ProfileHandler.sync_run method for better readability - Reformat parameters in ProfileHandler.aadd method to separate lines - Update model_copy call in reme.py to span multiple lines for better readability - Format aadd_batch call in update_profile.py to span multiple lines * feat(profiles): add profile management system with file and vector storage backends - Add FileProfileBackend for filesystem-based profile persistence - Add VectorProfileBackend for vector store-based profile management - Create abstract BaseProfileBackend interface for profile operations - Implement ProfileVectorHandler for vector-backed profile storage - Add RetrieveProfile tool for semantic profile retrieval - Update eval_reme.py to use user_message_s2 for retriever prompt - Modify eval_reme.yaml to use {profiles} instead of {user_profile} - Implement complete CRUD operations for profile management - Add batch operations for efficient profile handling - Include search functionality with semantic matching capabilities - Add capacity limits and automatic cleanup for profile storage * docs(profiles): add comprehensive docstrings for profile backend and handler methods - Added documentation for get_all_sync, get_by_sync, delete_sync, delete_all_sync methods - Documented add_sync and add_batch_sync functionality with deduping behavior - Added docstrings for update_sync and search_sync operations - Updated ProfileHandler.format_node method with proper documentation - Refactored private _format_node to public format_node method - Added comprehensive documentation for profile vector handler operations - Documented _vector_profile_matches, _get_by_profile_id, _get_by_profile_key helper methods - Added docstrings for retrieve_profile functionality and formatting methods
180 lines
8.5 KiB
YAML
180 lines
8.5 KiB
YAML
TEMPLATE_MEMOS: |
|
|
Memories for user {user_id}:
|
|
{memories}
|
|
|
|
PROMPT_MEMZERO_JSON: |
|
|
# CONTEXT:
|
|
{context}
|
|
|
|
# CONTEXT PRIORITY:
|
|
When the context contains information from multiple sources, follow this strict priority order:
|
|
1. **Historical Dialogue** (highest priority) - Direct conversation content
|
|
2. **Extracted Memories** (medium priority) - Summarized memory points
|
|
3. **User Profile** (lowest priority) - General user information
|
|
|
|
# Question:
|
|
{question}
|
|
|
|
# INSTRUCTIONS:
|
|
1. Carefully analyze all provided memories (facts and entities)
|
|
2. Pay special attention to the timestamps (event_time) to determine when events occurred
|
|
3. If the question asks about a specific event or fact, look for direct evidence in the memories
|
|
4. If the memories contain contradictory information, prioritize the most recent memory
|
|
5. Always convert relative time references to specific dates, months, or years
|
|
6. Be as specific as possible when talking about people, places, and events
|
|
7. Timestamps in memories represent the time the event was mentioned in a message, not the actual time the event occurred
|
|
|
|
|
|
# OUTPUT FORMAT:
|
|
Please provide your response in the following JSON format:
|
|
|
|
```json
|
|
{{
|
|
"reasoning": "reasoning content",
|
|
"answer": "Provide a detailed answer"
|
|
}}
|
|
```
|
|
|
|
SYSTEM_PROMPT: |
|
|
You are an expert grader that determines if answers to questions match a gold standard answer
|
|
|
|
USER_PROMPT: |
|
|
Your task is to label an answer to a question as 'CORRECT' or 'WRONG'. You will be given the following data:
|
|
(1) a question (posed by one user to another user),
|
|
(2) a 'gold' (ground truth) answer,
|
|
(3) a generated answer
|
|
which you will score as CORRECT/WRONG.
|
|
|
|
The point of the question is to ask about something one user should know about the other user based on their prior conversations.
|
|
The gold answer will usually be a concise and short answer that includes the referenced topic, for example:
|
|
Question: Do you remember what I got the last time I went to Hawaii?
|
|
Gold answer: A shell necklace
|
|
The generated answer might be much longer, but you should be generous with your grading - as long as it touches on the same topic as the gold answer, it should be counted as CORRECT.
|
|
|
|
For time related questions, the gold answer will be a specific date, month, year, etc. The generated answer might be much longer or use relative time references (like "last Tuesday" or "next month"), but you should be generous with your grading - as long as it refers to the same date or time period as the gold answer, it should be counted as CORRECT. Even if the format differs (e.g., "May 7th" vs "7 May"), consider it CORRECT if it's the same date.
|
|
|
|
Now it's time for the real question:
|
|
Question: {question}
|
|
Gold answer: {golden_answer}
|
|
Generated answer: {generated_answer}
|
|
|
|
First, provide a short (one sentence) explanation of your reasoning, then finish with CORRECT or WRONG.
|
|
Do NOT include both CORRECT and WRONG in your response, or it will break the evaluation script.
|
|
|
|
Just return the label CORRECT or WRONG in a json format with the key as "label".
|
|
|
|
user_message_summary_1: |
|
|
You are a Memory Agent responsible for managing {memory_type} memories about {memory_target}.
|
|
|
|
## Latest Conversation
|
|
Format: round<index> [<timestamp>] <role/name>: <content>
|
|
{context}
|
|
|
|
## Task
|
|
### Step 1: Create Memory Draft
|
|
Use `add_draft_and_retrieve_similar_memory` to create a memory draft list based on the latest conversation.
|
|
- For each memory draft, fill in the required parameters:
|
|
* `message_time`: timestamp from the conversation (e.g., '2020-01-01 00:00:00')
|
|
* `memory_content`: concise memory content extracted from the conversation
|
|
- Use actual names from the conversation (e.g., "Bob likes apples") instead of generic references (e.g., "user likes apples")
|
|
- Extract all important information comprehensively—do not miss critical details, but avoid any fabrications or unfounded assumptions
|
|
- The tool will retrieve similar historical memories via vector search to help you in Step 2
|
|
|
|
### Step 2: Add Memories
|
|
Review each memory draft from Step 1 and compare it with the retrieved historical memories, then use `add_memory` to manage all memories in one call:
|
|
|
|
- For each new memory, fill in the required parameters:
|
|
* `message_time`: timestamp from the conversation (e.g., '2020-01-01 00:00:00')
|
|
* `memory_content`: memory content
|
|
- Add memories when:
|
|
* The draft contains new information not present in historical memories
|
|
|
|
|
|
**General Guidelines:**
|
|
- **Skip** drafts if their content is already fully covered by historical memories (avoid redundancy)
|
|
- You can add memories in a single `add_memory` tool call
|
|
|
|
user_message_summary_2: |
|
|
You are a Profile Agent responsible for managing profiles about {memory_target}.
|
|
|
|
## Latest Conversation
|
|
Format: round<index> [<timestamp>] <role/name>: <content>
|
|
{context}
|
|
|
|
## Current Profiles
|
|
{profiles}
|
|
|
|
## Task
|
|
Analyze the Latest Conversation and use `update_profiles` to manage profiles (both updates and additions in one call):
|
|
|
|
**For profiles_to_update** (updating existing profiles):
|
|
- For each profile to update, fill in the required parameters:
|
|
* `profile_id`: ID of the profile to update (from Current Profiles)
|
|
* `message_time`: timestamp from the conversation (e.g., '2020-01-01 00:00:00')
|
|
* `profile_key`: profile key or category (e.g., 'name', 'age', 'occupation')
|
|
* `profile_value`: updated profile value, please be concise. (e.g., 'John Smith')
|
|
|
|
**For profiles_to_add** (adding new profiles):
|
|
- For each new profile, fill in the required parameters:
|
|
* `message_time`: timestamp from the conversation (e.g., '2020-01-01 00:00:00')
|
|
* `profile_key`: profile key or category (e.g., 'name', 'age', 'occupation')
|
|
* `profile_value`: profile value (e.g., 'John Smith')
|
|
- Add profiles when:
|
|
* The information represents a new distinct profile not present in Current Profiles
|
|
* The profile key doesn't exist in Current Profiles
|
|
* The information cannot be merged into existing profiles
|
|
|
|
**General Guidelines:**
|
|
- Extract all important information comprehensively—do not miss critical details, but avoid any fabrications or unfounded assumptions
|
|
- You can update and add profiles in a single tool call
|
|
|
|
user_message_retrieve: |
|
|
You are a Memory Retrieval Agent specialized in retrieving {memory_type} memories about {memory_target}.
|
|
|
|
## User Profile
|
|
{profiles}
|
|
|
|
## User Question
|
|
{context}
|
|
|
|
## Multi-Phase Retrieval Strategy
|
|
Follow these phases sequentially to gather comprehensive information:
|
|
|
|
### Phase 1: Semantic Search (No Time Filter)
|
|
**Tool**: `retrieve_memory` (without time constraints)
|
|
**Objective**: Cast a wide net to find potentially relevant memories
|
|
**Approach**:
|
|
- Execute 3-5 diverse search queries using different formulations:
|
|
* Original question verbatim
|
|
* Rephrased variations (different wording, synonyms)
|
|
* Entity-focused queries (extract and search specific names, places, events)
|
|
* Keyword-based searches (core concepts, topics)
|
|
* Related context queries (broader themes)
|
|
- Review all results before proceeding to next phase
|
|
|
|
### Phase 2: Deep Dive into History
|
|
**Tool**: `read_history`
|
|
**When to use**: After exhausting retrieval attempts OR when specific conversation context is needed
|
|
**Important Constraints**:
|
|
- Each history is very long and resource-intensive to read
|
|
- **Maximum limit: Read no more than 3 histories total**
|
|
- Only use this phase when absolutely necessary for answering the question
|
|
**Approach**:
|
|
- Extract `history_id` from retrieved memory references
|
|
- Prioritize the most relevant or recent histories
|
|
- Can read multiple histories at once by passing multiple history_ids
|
|
- Be selective: choose only the top 1-3 most promising histories
|
|
- Use this to understand the full conversation surrounding a memory
|
|
|
|
## Response Guidelines
|
|
- Base your answer EXCLUSIVELY on user profile, retrieved memories, and history data
|
|
- Never infer, assume, or hallucinate information
|
|
- Always cite sources with timestamps: `[timestamp] Memory content`
|
|
- Present conflicting information transparently with respective timestamps
|
|
- If you find sufficient information to answer the user's question, you may output directly without exhausting all search phases
|
|
- Exhaust all search strategies before concluding information doesn't exist
|
|
|
|
### Output any tangentially related findings, Format:
|
|
[timestamp] [memory/profile/history] [relevant content1]
|
|
[timestamp] [memory/profile/history] [relevant content2]
|
|
|