docs(task_memory): enhance documentation with data structures and usage examples

- Add TaskMemory and Trajectory data structure definitions
- Include detailed processing flows for record and delete operations
- Add comparison table between Task Memory and Tool Memory
- Provide best practices for trajectory scoring and maintenance
- Enhance build_query_op documentation with processing flow
- Add format examples for rewrite_memory_op with and without LLM
- Include usage examples for simple_summary_op and comparative summary
- Add comparison table between simple and full summary pipelines
- Improve parameter descriptions and when-to-use guidance
- Add Chinese translations for key concepts and examples
This commit is contained in:
jinli.yl 2025-10-16 16:36:40 +08:00
parent b07213791f
commit 82405c4a7c
3 changed files with 313 additions and 7 deletions

View file

@ -11,9 +11,39 @@ Task Memory represents knowledge extracted from previous task executions, includ
Each task memory contains:
- `when_to_use`: Conditions that indicate when this memory is relevant
- `content`: The actual knowledge or memory to be applied
- `content`: The actual knowledge or experience to be applied
- `score`: Quality score of the memory (assigned during validation)
- Metadata about the memory's source and utility
## Task Memory的数据结构
### TaskMemory
```python
class TaskMemory(BaseMemory):
memory_type: str = "task"
workspace_id: str # 工作空间ID
memory_id: str # 记忆的唯一ID
when_to_use: str # 何时使用此记忆的条件
content: str # 具体的经验内容
score: float # 质量评分(validation阶段设置)
time_created: str # 创建时间
time_modified: str # 最后修改时间
author: str # 创建者(通常是LLM模型名)
metadata: dict # 其他元数据
```
### Trajectory
```python
class Trajectory(BaseModel):
messages: List[Message] # 对话消息列表
score: float # 轨迹评分(用于判断成功/失败)
metadata: dict # 可选的元数据(如segments、query等)
```
Task Memory通过分析Trajectory来提取经验。轨迹的score决定了它被分类为成功还是失败案例。
## Configuration Logic
Task Memory in ReMe is configured through two main flows:
@ -196,7 +226,119 @@ response = requests.post(
ReMe also provides additional task memory operations:
- `record_task_memory`: Update frequency and utility attributes of retrieved memories
- `delete_task_memory`: Delete memories based on utility/frequency thresholds
### Record Task Memory
`record_task_memory` flow用于更新检索到的任务记忆的使用频率和效用属性:
```yaml
record_task_memory:
flow_content: update_memory_freq_op >> update_memory_utility_op >> update_vector_store_op
description: "Update the freq & utility attributes of retrieved task memories"
input_schema:
workspace_id:
type: string
required: true
memory_dicts:
type: array
description: "A list of retrieved task memory"
required: true
update_utility:
type: boolean
description: "Whether to update the utility attribute"
required: true
```
使用示例:
```python
# 记录检索到的记忆被使用
response = requests.post(
url=f"{BASE_URL}record_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"memory_dicts": retrieved_memories, # 从retrieve_task_memory获取的记忆列表
"update_utility": True # 是否更新效用值
}
)
```
### Delete Task Memory
`delete_task_memory` flow用于删除低效用的任务记忆:
```yaml
delete_task_memory:
flow_content: delete_memory_op >> update_vector_store_op
description: "Delete task memories when utility/freq < utility_threshold and freq >= freq_threshold"
input_schema:
workspace_id:
type: string
required: true
freq_threshold:
type: integer
description: "The retrieved frequency threshold"
required: true
utility_threshold:
type: number
description: "The utility/freq threshold"
required: true
```
使用示例:
```python
# 删除使用频率高但效用低的记忆
response = requests.post(
url=f"{BASE_URL}delete_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"freq_threshold": 10, # 至少被检索过10次
"utility_threshold": 0.3 # 但效用/频率比 < 0.3
}
)
```
这个机制确保记忆库的质量:频繁被检索但实际没什么帮助的记忆会被清理。
## 与Tool Memory的对比
| 特性 | Task Memory | Tool Memory |
|------|-------------|-------------|
| **目的** | 记录任务解决的经验 | 记录工具调用的使用模式 |
| **输入** | 对话轨迹(Trajectory) | 工具调用记录(ToolCallResult) |
| **索引方式** | when_to_use条件(语义搜索) | 工具名称(精确匹配) |
| **提取方式** | 从完整轨迹中分析提取 | 从单次调用评估和统计 |
| **内容** | 任务解决经验和策略 | 工具使用指南和最佳实践 |
| **更新方式** | 从成功/失败/对比中提取 | 每次调用后追加记录 |
| **检索流程** | 复杂:构建查询→召回→重排→重写 | 简单:精确匹配→返回 |
| **去重机制** | 语义去重(deduplication) | 按工具名唯一,限制历史条数 |
| **验证机制** | LLM验证质量(validation) | 评估单次调用质量 |
## 最佳实践
1. **轨迹评分标准**:
- 明确定义成功和失败的评分标准
- `success_threshold`通常设为1.0(完美成功)
- 部分成功的轨迹(如0.5-0.9)可以提供对比经验
2. **轨迹质量**:
- 确保轨迹包含完整的问题解决过程
- 包含足够的上下文信息在metadata中
- 避免过短或过长的轨迹
3. **定期维护**:
- 使用`record_task_memory`跟踪记忆使用情况
- 定期使用`delete_task_memory`清理低效记忆
- 保持记忆库的质量和相关性
4. **合理使用简化版本**:
- 简单场景使用`summary_task_memory_simple`
- 复杂场景使用完整的`summary_task_memory` flow
- 根据实际需求选择是否启用轨迹分割
5. **检索优化**:
- 调整`top_k`参数(默认5)控制返回的记忆数量
- 启用`enable_llm_rerank`提升相关性
- 启用`enable_llm_rewrite`使记忆更适配当前任务
For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project.

View file

@ -8,16 +8,61 @@ Constructs a query for memory retrieval either from a direct query input or by a
### Functionality
- If a direct `query` is provided in the context, it uses that query
- If a direct `query` is provided in the context, it uses that query directly
- If `messages` are provided in the context, it can:
- Use an LLM to generate a query based on the conversation context
- Or create a simple query from recent messages without using an LLM
- **LLM-based**: Use an LLM to generate a focused query based on the conversation context
- **Simple**: Create a simple query from the last 3 messages without using an LLM
- Raises an error if neither `query` nor `messages` is provided
### Processing Flow
1. **Check for Direct Query**:
- If `context.query` exists, use it directly
- This is the most common case when retrieving for a specific task
2. **Build from Messages** (if no direct query):
- Extract last 3 messages from `context.messages`
- If `enable_llm_build=true`:
- Merge all messages into execution_process text
- Use LLM to analyze and generate a focused query
- If `enable_llm_build=false`:
- Truncate each message content to 200 chars
- Format as "- role: content" for each message
- Concatenate into a simple query
3. **Set Context**: Store the built query in `context.query` for downstream ops
### Parameters
- `op.build_query_op.params.enable_llm_build` (boolean, default: `true`):
- When `true`, uses an LLM to generate a query from conversation messages
- When `false`, creates a simple query by concatenating recent messages
- LLM-based queries are more focused but require an extra LLM call
- Simple queries are faster but may be less targeted
### Usage Example
```python
import requests
# Method 1: Direct query (recommended)
response = requests.post(
url="http://0.0.0.0:8002/retrieve_task_memory",
json={
"workspace_id": "my_workspace",
"query": "How to handle file upload errors in Python?"
}
)
# Method 2: From messages (when integrated in conversation)
# This is handled internally by the flow when messages are provided
```
### When to Use Which Mode
- **Direct Query**: Most common, use when you know what to search for
- **LLM-based Build**: When you have a conversation and want the LLM to extract the key question
- **Simple Build**: When you want fast retrieval without extra LLM cost
## RerankMemoryOp
@ -53,12 +98,78 @@ Rewrites and formats the retrieved memories to make them more relevant and actio
- Formats retrieved memories into a structured format
- Can use an LLM to rewrite memories to better fit the current context (optional)
- Generates a cohesive context message from multiple memories
- Extracts current conversation context to inform the rewriting process
### Processing Flow
1. **Get Inputs**:
- `memory_list`: Retrieved and reranked memories from previous ops
- `query`: The search query
- `messages`: Current conversation messages (optional)
2. **Format Memories**:
- Create numbered memory entries with "When to use" and "Content"
- Example format:
```
Memory 1:
When to use: [condition]
Content: [experience]
```
3. **LLM Rewrite** (if `enable_llm_rewrite=true`):
- Extract current context from last 3 messages
- Provide original formatted memories and current context to LLM
- LLM adapts the memories to be more relevant to the current query
- Extracts `rewritten_context` from LLM response (JSON format)
4. **Return**: Set `context.response.answer` to the final formatted/rewritten memories
### Parameters
- `op.rewrite_memory_op.params.enable_llm_rewrite` (boolean, default: `true`):
- When `true`, uses an LLM to rewrite the memories to make them more relevant and actionable
- When `false`, simply formats the memories without LLM-based rewriting
- LLM rewriting makes memories more contextual but adds latency
### Format Examples
**Without LLM Rewrite** (enable_llm_rewrite=false):
```
Memory 1 :
When to use: When handling file upload errors in web applications
Content: Use try-except blocks to catch specific exceptions like PermissionError...
Memory 2 :
When to use: When validating file types before processing
Content: Check file extensions and MIME types to ensure security...
```
**With LLM Rewrite** (enable_llm_rewrite=true):
```
Based on your question about file upload errors in Python web applications:
1. Error Handling Strategy:
- Wrap file operations in try-except blocks
- Catch specific exceptions: PermissionError, IOError, OSError
- [Adapted from Memory 1]
2. Input Validation:
- Always validate file types before processing
- Check both extension and MIME type
- [Adapted from Memory 2]
```
### When to Enable LLM Rewrite
- **Enable** when:
- Memories need to be adapted to specific context
- User query is complex or nuanced
- Better integration with conversation flow is needed
- **Disable** when:
- Fast retrieval is prioritized
- Memories are already clear and relevant
- Reducing LLM calls and costs is important
## MergeMemoryOp

View file

@ -138,13 +138,66 @@ A simplified version of memory extraction that processes entire trajectories in
### Functionality
- Classifies trajectories as success or failure based on score threshold
- Extracts memories directly from complete trajectories
- Extracts memories directly from complete trajectories without segmentation
- Uses a single LLM prompt to extract multiple memories from each trajectory
- Parses JSON response containing `when_to_use` and `memory` pairs
- Useful for simpler use cases where detailed segmentation is not required
### Processing Flow
1. **Receive Trajectories**: Get a list of Trajectory objects with messages and scores
2. **For Each Trajectory**:
- Merge all messages into execution_process text
- Determine execution_result ("success" or "fail") based on score threshold
- Format prompt with execution_process and execution_result
- Call LLM to extract memories in JSON format
3. **Parse Response**: Extract JSON array of {when_to_use, memory} objects
4. **Create TaskMemory**: Create TaskMemory objects for each extracted memory
5. **Return**: Set memory_list in context.response.metadata
### Parameters
- `op.simple_summary_op.params.success_score_threshold` (float, default: `0.9`):
- The threshold score that determines if a trajectory is considered successful
- Trajectories with `score >= success_score_threshold` are marked as "success"
- Trajectories with `score < success_score_threshold` are marked as "fail"
### Usage Example
```python
import requests
# Prepare trajectory with messages
trajectory = {
"messages": [
{"role": "user", "content": "How do I sort a list in Python?"},
{"role": "assistant", "content": "You can use the sorted() function..."},
{"role": "user", "content": "Thanks! What about reverse sorting?"},
{"role": "assistant", "content": "Use sorted(list, reverse=True)..."}
],
"score": 1.0 # Successful trajectory
}
# Summarize using simple summary
response = requests.post(
url="http://0.0.0.0:8002/summary_task_memory_simple",
json={
"workspace_id": "my_workspace",
"trajectories": [trajectory]
}
)
```
### Comparison with Full Summary Pipeline
| Feature | SimpleSummaryOp | Full Summary Pipeline |
|---------|-----------------|----------------------|
| **Complexity** | Single operation | Multi-stage pipeline |
| **Segmentation** | No segmentation | Optional segmentation |
| **Extraction** | Single prompt per trajectory | Separate prompts for success/failure/comparative |
| **Validation** | No validation | LLM validation step |
| **Deduplication** | No deduplication | Optional deduplication |
| **Use Case** | Quick prototyping | Production environments |
## SimpleComparativeSummaryOp