diff --git a/docs/task_memory/task_memory.md b/docs/task_memory/task_memory.md index 5e232225..a3e07295 100644 --- a/docs/task_memory/task_memory.md +++ b/docs/task_memory/task_memory.md @@ -11,9 +11,39 @@ Task Memory represents knowledge extracted from previous task executions, includ Each task memory contains: - `when_to_use`: Conditions that indicate when this memory is relevant -- `content`: The actual knowledge or memory to be applied +- `content`: The actual knowledge or experience to be applied +- `score`: Quality score of the memory (assigned during validation) - Metadata about the memory's source and utility +## Task Memory的数据结构 + +### TaskMemory + +```python +class TaskMemory(BaseMemory): + memory_type: str = "task" + workspace_id: str # 工作空间ID + memory_id: str # 记忆的唯一ID + when_to_use: str # 何时使用此记忆的条件 + content: str # 具体的经验内容 + score: float # 质量评分(validation阶段设置) + time_created: str # 创建时间 + time_modified: str # 最后修改时间 + author: str # 创建者(通常是LLM模型名) + metadata: dict # 其他元数据 +``` + +### Trajectory + +```python +class Trajectory(BaseModel): + messages: List[Message] # 对话消息列表 + score: float # 轨迹评分(用于判断成功/失败) + metadata: dict # 可选的元数据(如segments、query等) +``` + +Task Memory通过分析Trajectory来提取经验。轨迹的score决定了它被分类为成功还是失败案例。 + ## Configuration Logic Task Memory in ReMe is configured through two main flows: @@ -196,7 +226,119 @@ response = requests.post( ReMe also provides additional task memory operations: -- `record_task_memory`: Update frequency and utility attributes of retrieved memories -- `delete_task_memory`: Delete memories based on utility/frequency thresholds +### Record Task Memory + +`record_task_memory` flow用于更新检索到的任务记忆的使用频率和效用属性: + +```yaml +record_task_memory: + flow_content: update_memory_freq_op >> update_memory_utility_op >> update_vector_store_op + description: "Update the freq & utility attributes of retrieved task memories" + input_schema: + workspace_id: + type: string + required: true + memory_dicts: + type: array + description: "A list of retrieved task memory" + required: true + update_utility: + type: boolean + description: "Whether to update the utility attribute" + required: true +``` + +使用示例: + +```python +# 记录检索到的记忆被使用 +response = requests.post( + url=f"{BASE_URL}record_task_memory", + json={ + "workspace_id": WORKSPACE_ID, + "memory_dicts": retrieved_memories, # 从retrieve_task_memory获取的记忆列表 + "update_utility": True # 是否更新效用值 + } +) +``` + +### Delete Task Memory + +`delete_task_memory` flow用于删除低效用的任务记忆: + +```yaml +delete_task_memory: + flow_content: delete_memory_op >> update_vector_store_op + description: "Delete task memories when utility/freq < utility_threshold and freq >= freq_threshold" + input_schema: + workspace_id: + type: string + required: true + freq_threshold: + type: integer + description: "The retrieved frequency threshold" + required: true + utility_threshold: + type: number + description: "The utility/freq threshold" + required: true +``` + +使用示例: + +```python +# 删除使用频率高但效用低的记忆 +response = requests.post( + url=f"{BASE_URL}delete_task_memory", + json={ + "workspace_id": WORKSPACE_ID, + "freq_threshold": 10, # 至少被检索过10次 + "utility_threshold": 0.3 # 但效用/频率比 < 0.3 + } +) +``` + +这个机制确保记忆库的质量:频繁被检索但实际没什么帮助的记忆会被清理。 + +## 与Tool Memory的对比 + +| 特性 | Task Memory | Tool Memory | +|------|-------------|-------------| +| **目的** | 记录任务解决的经验 | 记录工具调用的使用模式 | +| **输入** | 对话轨迹(Trajectory) | 工具调用记录(ToolCallResult) | +| **索引方式** | when_to_use条件(语义搜索) | 工具名称(精确匹配) | +| **提取方式** | 从完整轨迹中分析提取 | 从单次调用评估和统计 | +| **内容** | 任务解决经验和策略 | 工具使用指南和最佳实践 | +| **更新方式** | 从成功/失败/对比中提取 | 每次调用后追加记录 | +| **检索流程** | 复杂:构建查询→召回→重排→重写 | 简单:精确匹配→返回 | +| **去重机制** | 语义去重(deduplication) | 按工具名唯一,限制历史条数 | +| **验证机制** | LLM验证质量(validation) | 评估单次调用质量 | + +## 最佳实践 + +1. **轨迹评分标准**: + - 明确定义成功和失败的评分标准 + - `success_threshold`通常设为1.0(完美成功) + - 部分成功的轨迹(如0.5-0.9)可以提供对比经验 + +2. **轨迹质量**: + - 确保轨迹包含完整的问题解决过程 + - 包含足够的上下文信息在metadata中 + - 避免过短或过长的轨迹 + +3. **定期维护**: + - 使用`record_task_memory`跟踪记忆使用情况 + - 定期使用`delete_task_memory`清理低效记忆 + - 保持记忆库的质量和相关性 + +4. **合理使用简化版本**: + - 简单场景使用`summary_task_memory_simple` + - 复杂场景使用完整的`summary_task_memory` flow + - 根据实际需求选择是否启用轨迹分割 + +5. **检索优化**: + - 调整`top_k`参数(默认5)控制返回的记忆数量 + - 启用`enable_llm_rerank`提升相关性 + - 启用`enable_llm_rewrite`使记忆更适配当前任务 For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project. diff --git a/docs/task_memory/task_retrieve_ops.md b/docs/task_memory/task_retrieve_ops.md index a12e8698..6c61c88f 100644 --- a/docs/task_memory/task_retrieve_ops.md +++ b/docs/task_memory/task_retrieve_ops.md @@ -8,16 +8,61 @@ Constructs a query for memory retrieval either from a direct query input or by a ### Functionality -- If a direct `query` is provided in the context, it uses that query +- If a direct `query` is provided in the context, it uses that query directly - If `messages` are provided in the context, it can: - - Use an LLM to generate a query based on the conversation context - - Or create a simple query from recent messages without using an LLM + - **LLM-based**: Use an LLM to generate a focused query based on the conversation context + - **Simple**: Create a simple query from the last 3 messages without using an LLM +- Raises an error if neither `query` nor `messages` is provided + +### Processing Flow + +1. **Check for Direct Query**: + - If `context.query` exists, use it directly + - This is the most common case when retrieving for a specific task + +2. **Build from Messages** (if no direct query): + - Extract last 3 messages from `context.messages` + - If `enable_llm_build=true`: + - Merge all messages into execution_process text + - Use LLM to analyze and generate a focused query + - If `enable_llm_build=false`: + - Truncate each message content to 200 chars + - Format as "- role: content" for each message + - Concatenate into a simple query + +3. **Set Context**: Store the built query in `context.query` for downstream ops ### Parameters - `op.build_query_op.params.enable_llm_build` (boolean, default: `true`): - When `true`, uses an LLM to generate a query from conversation messages - When `false`, creates a simple query by concatenating recent messages + - LLM-based queries are more focused but require an extra LLM call + - Simple queries are faster but may be less targeted + +### Usage Example + +```python +import requests + +# Method 1: Direct query (recommended) +response = requests.post( + url="http://0.0.0.0:8002/retrieve_task_memory", + json={ + "workspace_id": "my_workspace", + "query": "How to handle file upload errors in Python?" + } +) + +# Method 2: From messages (when integrated in conversation) +# This is handled internally by the flow when messages are provided +``` + +### When to Use Which Mode + +- **Direct Query**: Most common, use when you know what to search for +- **LLM-based Build**: When you have a conversation and want the LLM to extract the key question +- **Simple Build**: When you want fast retrieval without extra LLM cost ## RerankMemoryOp @@ -53,12 +98,78 @@ Rewrites and formats the retrieved memories to make them more relevant and actio - Formats retrieved memories into a structured format - Can use an LLM to rewrite memories to better fit the current context (optional) - Generates a cohesive context message from multiple memories +- Extracts current conversation context to inform the rewriting process + +### Processing Flow + +1. **Get Inputs**: + - `memory_list`: Retrieved and reranked memories from previous ops + - `query`: The search query + - `messages`: Current conversation messages (optional) + +2. **Format Memories**: + - Create numbered memory entries with "When to use" and "Content" + - Example format: + ``` + Memory 1: + When to use: [condition] + Content: [experience] + ``` + +3. **LLM Rewrite** (if `enable_llm_rewrite=true`): + - Extract current context from last 3 messages + - Provide original formatted memories and current context to LLM + - LLM adapts the memories to be more relevant to the current query + - Extracts `rewritten_context` from LLM response (JSON format) + +4. **Return**: Set `context.response.answer` to the final formatted/rewritten memories ### Parameters - `op.rewrite_memory_op.params.enable_llm_rewrite` (boolean, default: `true`): - When `true`, uses an LLM to rewrite the memories to make them more relevant and actionable - When `false`, simply formats the memories without LLM-based rewriting + - LLM rewriting makes memories more contextual but adds latency + +### Format Examples + +**Without LLM Rewrite** (enable_llm_rewrite=false): +``` +Memory 1 : + When to use: When handling file upload errors in web applications + Content: Use try-except blocks to catch specific exceptions like PermissionError... + +Memory 2 : + When to use: When validating file types before processing + Content: Check file extensions and MIME types to ensure security... +``` + +**With LLM Rewrite** (enable_llm_rewrite=true): +``` +Based on your question about file upload errors in Python web applications: + +1. Error Handling Strategy: + - Wrap file operations in try-except blocks + - Catch specific exceptions: PermissionError, IOError, OSError + - [Adapted from Memory 1] + +2. Input Validation: + - Always validate file types before processing + - Check both extension and MIME type + - [Adapted from Memory 2] +``` + +### When to Enable LLM Rewrite + +- **Enable** when: + - Memories need to be adapted to specific context + - User query is complex or nuanced + - Better integration with conversation flow is needed + +- **Disable** when: + - Fast retrieval is prioritized + - Memories are already clear and relevant + - Reducing LLM calls and costs is important ## MergeMemoryOp diff --git a/docs/task_memory/task_summary_ops.md b/docs/task_memory/task_summary_ops.md index 138f5812..f1364baf 100644 --- a/docs/task_memory/task_summary_ops.md +++ b/docs/task_memory/task_summary_ops.md @@ -138,13 +138,66 @@ A simplified version of memory extraction that processes entire trajectories in ### Functionality - Classifies trajectories as success or failure based on score threshold -- Extracts memories directly from complete trajectories +- Extracts memories directly from complete trajectories without segmentation +- Uses a single LLM prompt to extract multiple memories from each trajectory +- Parses JSON response containing `when_to_use` and `memory` pairs - Useful for simpler use cases where detailed segmentation is not required +### Processing Flow + +1. **Receive Trajectories**: Get a list of Trajectory objects with messages and scores +2. **For Each Trajectory**: + - Merge all messages into execution_process text + - Determine execution_result ("success" or "fail") based on score threshold + - Format prompt with execution_process and execution_result + - Call LLM to extract memories in JSON format +3. **Parse Response**: Extract JSON array of {when_to_use, memory} objects +4. **Create TaskMemory**: Create TaskMemory objects for each extracted memory +5. **Return**: Set memory_list in context.response.metadata + ### Parameters - `op.simple_summary_op.params.success_score_threshold` (float, default: `0.9`): - The threshold score that determines if a trajectory is considered successful + - Trajectories with `score >= success_score_threshold` are marked as "success" + - Trajectories with `score < success_score_threshold` are marked as "fail" + +### Usage Example + +```python +import requests + +# Prepare trajectory with messages +trajectory = { + "messages": [ + {"role": "user", "content": "How do I sort a list in Python?"}, + {"role": "assistant", "content": "You can use the sorted() function..."}, + {"role": "user", "content": "Thanks! What about reverse sorting?"}, + {"role": "assistant", "content": "Use sorted(list, reverse=True)..."} + ], + "score": 1.0 # Successful trajectory +} + +# Summarize using simple summary +response = requests.post( + url="http://0.0.0.0:8002/summary_task_memory_simple", + json={ + "workspace_id": "my_workspace", + "trajectories": [trajectory] + } +) +``` + +### Comparison with Full Summary Pipeline + +| Feature | SimpleSummaryOp | Full Summary Pipeline | +|---------|-----------------|----------------------| +| **Complexity** | Single operation | Multi-stage pipeline | +| **Segmentation** | No segmentation | Optional segmentation | +| **Extraction** | Single prompt per trajectory | Separate prompts for success/failure/comparative | +| **Validation** | No validation | LLM validation step | +| **Deduplication** | No deduplication | Optional deduplication | +| **Use Case** | Quick prototyping | Production environments | ## SimpleComparativeSummaryOp