docs(task_memory): simplify documentation and remove redundant content

- Removed detailed data structure definitions for TaskMemory and Trajectory
- Eliminated extensive examples and usage instructions for record_task_memory and delete_task_memory
- Removed comparison table with Tool Memory and best practices section
- Simplified build_query_op and rewrite_memory_op documentation by removing processing flows
- Removed detailed examples and parameter descriptions for simple_summary_op- Added new tool_memory.md documentation with complete tool memory implementation
- Added new tool_retrieve_ops.md documentation with detailed retrieval operations- Created comprehensive tool memory data structures and API documentation
- Added usage examples and integration workflows for tool memory operations
- Documented configuration parameters and best practices for tool memory management
This commit is contained in:
jinli.yl 2025-10-16 16:39:40 +08:00
parent 82405c4a7c
commit 40538f974a
6 changed files with 1580 additions and 316 deletions

View file

@ -11,39 +11,9 @@ Task Memory represents knowledge extracted from previous task executions, includ
Each task memory contains:
- `when_to_use`: Conditions that indicate when this memory is relevant
- `content`: The actual knowledge or experience to be applied
- `score`: Quality score of the memory (assigned during validation)
- `content`: The actual knowledge or memory to be applied
- Metadata about the memory's source and utility
## Task Memory的数据结构
### TaskMemory
```python
class TaskMemory(BaseMemory):
memory_type: str = "task"
workspace_id: str # 工作空间ID
memory_id: str # 记忆的唯一ID
when_to_use: str # 何时使用此记忆的条件
content: str # 具体的经验内容
score: float # 质量评分(validation阶段设置)
time_created: str # 创建时间
time_modified: str # 最后修改时间
author: str # 创建者(通常是LLM模型名)
metadata: dict # 其他元数据
```
### Trajectory
```python
class Trajectory(BaseModel):
messages: List[Message] # 对话消息列表
score: float # 轨迹评分(用于判断成功/失败)
metadata: dict # 可选的元数据(如segments、query等)
```
Task Memory通过分析Trajectory来提取经验。轨迹的score决定了它被分类为成功还是失败案例。
## Configuration Logic
Task Memory in ReMe is configured through two main flows:
@ -226,119 +196,7 @@ response = requests.post(
ReMe also provides additional task memory operations:
### Record Task Memory
- `record_task_memory`: Update frequency and utility attributes of retrieved memories
- `delete_task_memory`: Delete memories based on utility/frequency thresholds
`record_task_memory` flow用于更新检索到的任务记忆的使用频率和效用属性:
```yaml
record_task_memory:
flow_content: update_memory_freq_op >> update_memory_utility_op >> update_vector_store_op
description: "Update the freq & utility attributes of retrieved task memories"
input_schema:
workspace_id:
type: string
required: true
memory_dicts:
type: array
description: "A list of retrieved task memory"
required: true
update_utility:
type: boolean
description: "Whether to update the utility attribute"
required: true
```
使用示例:
```python
# 记录检索到的记忆被使用
response = requests.post(
url=f"{BASE_URL}record_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"memory_dicts": retrieved_memories, # 从retrieve_task_memory获取的记忆列表
"update_utility": True # 是否更新效用值
}
)
```
### Delete Task Memory
`delete_task_memory` flow用于删除低效用的任务记忆:
```yaml
delete_task_memory:
flow_content: delete_memory_op >> update_vector_store_op
description: "Delete task memories when utility/freq < utility_threshold and freq >= freq_threshold"
input_schema:
workspace_id:
type: string
required: true
freq_threshold:
type: integer
description: "The retrieved frequency threshold"
required: true
utility_threshold:
type: number
description: "The utility/freq threshold"
required: true
```
使用示例:
```python
# 删除使用频率高但效用低的记忆
response = requests.post(
url=f"{BASE_URL}delete_task_memory",
json={
"workspace_id": WORKSPACE_ID,
"freq_threshold": 10, # 至少被检索过10次
"utility_threshold": 0.3 # 但效用/频率比 < 0.3
}
)
```
这个机制确保记忆库的质量:频繁被检索但实际没什么帮助的记忆会被清理。
## 与Tool Memory的对比
| 特性 | Task Memory | Tool Memory |
|------|-------------|-------------|
| **目的** | 记录任务解决的经验 | 记录工具调用的使用模式 |
| **输入** | 对话轨迹(Trajectory) | 工具调用记录(ToolCallResult) |
| **索引方式** | when_to_use条件(语义搜索) | 工具名称(精确匹配) |
| **提取方式** | 从完整轨迹中分析提取 | 从单次调用评估和统计 |
| **内容** | 任务解决经验和策略 | 工具使用指南和最佳实践 |
| **更新方式** | 从成功/失败/对比中提取 | 每次调用后追加记录 |
| **检索流程** | 复杂:构建查询→召回→重排→重写 | 简单:精确匹配→返回 |
| **去重机制** | 语义去重(deduplication) | 按工具名唯一,限制历史条数 |
| **验证机制** | LLM验证质量(validation) | 评估单次调用质量 |
## 最佳实践
1. **轨迹评分标准**:
- 明确定义成功和失败的评分标准
- `success_threshold`通常设为1.0(完美成功)
- 部分成功的轨迹(如0.5-0.9)可以提供对比经验
2. **轨迹质量**:
- 确保轨迹包含完整的问题解决过程
- 包含足够的上下文信息在metadata中
- 避免过短或过长的轨迹
3. **定期维护**:
- 使用`record_task_memory`跟踪记忆使用情况
- 定期使用`delete_task_memory`清理低效记忆
- 保持记忆库的质量和相关性
4. **合理使用简化版本**:
- 简单场景使用`summary_task_memory_simple`
- 复杂场景使用完整的`summary_task_memory` flow
- 根据实际需求选择是否启用轨迹分割
5. **检索优化**:
- 调整`top_k`参数(默认5)控制返回的记忆数量
- 启用`enable_llm_rerank`提升相关性
- 启用`enable_llm_rewrite`使记忆更适配当前任务
For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project.
For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project.

View file

@ -8,61 +8,16 @@ Constructs a query for memory retrieval either from a direct query input or by a
### Functionality
- If a direct `query` is provided in the context, it uses that query directly
- If a direct `query` is provided in the context, it uses that query
- If `messages` are provided in the context, it can:
- **LLM-based**: Use an LLM to generate a focused query based on the conversation context
- **Simple**: Create a simple query from the last 3 messages without using an LLM
- Raises an error if neither `query` nor `messages` is provided
### Processing Flow
1. **Check for Direct Query**:
- If `context.query` exists, use it directly
- This is the most common case when retrieving for a specific task
2. **Build from Messages** (if no direct query):
- Extract last 3 messages from `context.messages`
- If `enable_llm_build=true`:
- Merge all messages into execution_process text
- Use LLM to analyze and generate a focused query
- If `enable_llm_build=false`:
- Truncate each message content to 200 chars
- Format as "- role: content" for each message
- Concatenate into a simple query
3. **Set Context**: Store the built query in `context.query` for downstream ops
- Use an LLM to generate a query based on the conversation context
- Or create a simple query from recent messages without using an LLM
### Parameters
- `op.build_query_op.params.enable_llm_build` (boolean, default: `true`):
- When `true`, uses an LLM to generate a query from conversation messages
- When `false`, creates a simple query by concatenating recent messages
- LLM-based queries are more focused but require an extra LLM call
- Simple queries are faster but may be less targeted
### Usage Example
```python
import requests
# Method 1: Direct query (recommended)
response = requests.post(
url="http://0.0.0.0:8002/retrieve_task_memory",
json={
"workspace_id": "my_workspace",
"query": "How to handle file upload errors in Python?"
}
)
# Method 2: From messages (when integrated in conversation)
# This is handled internally by the flow when messages are provided
```
### When to Use Which Mode
- **Direct Query**: Most common, use when you know what to search for
- **LLM-based Build**: When you have a conversation and want the LLM to extract the key question
- **Simple Build**: When you want fast retrieval without extra LLM cost
## RerankMemoryOp
@ -98,78 +53,12 @@ Rewrites and formats the retrieved memories to make them more relevant and actio
- Formats retrieved memories into a structured format
- Can use an LLM to rewrite memories to better fit the current context (optional)
- Generates a cohesive context message from multiple memories
- Extracts current conversation context to inform the rewriting process
### Processing Flow
1. **Get Inputs**:
- `memory_list`: Retrieved and reranked memories from previous ops
- `query`: The search query
- `messages`: Current conversation messages (optional)
2. **Format Memories**:
- Create numbered memory entries with "When to use" and "Content"
- Example format:
```
Memory 1:
When to use: [condition]
Content: [experience]
```
3. **LLM Rewrite** (if `enable_llm_rewrite=true`):
- Extract current context from last 3 messages
- Provide original formatted memories and current context to LLM
- LLM adapts the memories to be more relevant to the current query
- Extracts `rewritten_context` from LLM response (JSON format)
4. **Return**: Set `context.response.answer` to the final formatted/rewritten memories
### Parameters
- `op.rewrite_memory_op.params.enable_llm_rewrite` (boolean, default: `true`):
- When `true`, uses an LLM to rewrite the memories to make them more relevant and actionable
- When `false`, simply formats the memories without LLM-based rewriting
- LLM rewriting makes memories more contextual but adds latency
### Format Examples
**Without LLM Rewrite** (enable_llm_rewrite=false):
```
Memory 1 :
When to use: When handling file upload errors in web applications
Content: Use try-except blocks to catch specific exceptions like PermissionError...
Memory 2 :
When to use: When validating file types before processing
Content: Check file extensions and MIME types to ensure security...
```
**With LLM Rewrite** (enable_llm_rewrite=true):
```
Based on your question about file upload errors in Python web applications:
1. Error Handling Strategy:
- Wrap file operations in try-except blocks
- Catch specific exceptions: PermissionError, IOError, OSError
- [Adapted from Memory 1]
2. Input Validation:
- Always validate file types before processing
- Check both extension and MIME type
- [Adapted from Memory 2]
```
### When to Enable LLM Rewrite
- **Enable** when:
- Memories need to be adapted to specific context
- User query is complex or nuanced
- Better integration with conversation flow is needed
- **Disable** when:
- Fast retrieval is prioritized
- Memories are already clear and relevant
- Reducing LLM calls and costs is important
## MergeMemoryOp
@ -181,4 +70,4 @@ An alternative to RewriteMemoryOp that merges multiple memories into a single re
- Collects the content from all memories in the memory list
- Formats them into a single response with a standard structure
- Adds a prompt to consider the helpful parts when answering the question
- Adds a prompt to consider the helpful parts when answering the question

View file

@ -138,66 +138,13 @@ A simplified version of memory extraction that processes entire trajectories in
### Functionality
- Classifies trajectories as success or failure based on score threshold
- Extracts memories directly from complete trajectories without segmentation
- Uses a single LLM prompt to extract multiple memories from each trajectory
- Parses JSON response containing `when_to_use` and `memory` pairs
- Extracts memories directly from complete trajectories
- Useful for simpler use cases where detailed segmentation is not required
### Processing Flow
1. **Receive Trajectories**: Get a list of Trajectory objects with messages and scores
2. **For Each Trajectory**:
- Merge all messages into execution_process text
- Determine execution_result ("success" or "fail") based on score threshold
- Format prompt with execution_process and execution_result
- Call LLM to extract memories in JSON format
3. **Parse Response**: Extract JSON array of {when_to_use, memory} objects
4. **Create TaskMemory**: Create TaskMemory objects for each extracted memory
5. **Return**: Set memory_list in context.response.metadata
### Parameters
- `op.simple_summary_op.params.success_score_threshold` (float, default: `0.9`):
- The threshold score that determines if a trajectory is considered successful
- Trajectories with `score >= success_score_threshold` are marked as "success"
- Trajectories with `score < success_score_threshold` are marked as "fail"
### Usage Example
```python
import requests
# Prepare trajectory with messages
trajectory = {
"messages": [
{"role": "user", "content": "How do I sort a list in Python?"},
{"role": "assistant", "content": "You can use the sorted() function..."},
{"role": "user", "content": "Thanks! What about reverse sorting?"},
{"role": "assistant", "content": "Use sorted(list, reverse=True)..."}
],
"score": 1.0 # Successful trajectory
}
# Summarize using simple summary
response = requests.post(
url="http://0.0.0.0:8002/summary_task_memory_simple",
json={
"workspace_id": "my_workspace",
"trajectories": [trajectory]
}
)
```
### Comparison with Full Summary Pipeline
| Feature | SimpleSummaryOp | Full Summary Pipeline |
|---------|-----------------|----------------------|
| **Complexity** | Single operation | Multi-stage pipeline |
| **Segmentation** | No segmentation | Optional segmentation |
| **Extraction** | Single prompt per trajectory | Separate prompts for success/failure/comparative |
| **Validation** | No validation | LLM validation step |
| **Deduplication** | No deduplication | Optional deduplication |
| **Use Case** | Quick prototyping | Production environments |
## SimpleComparativeSummaryOp
@ -213,4 +160,4 @@ A simplified version of comparative memory extraction.
### Parameters
No specific parameters beyond the LLM configuration.
No specific parameters beyond the LLM configuration.

View file

@ -0,0 +1,557 @@
# Tool Memory in ReMe
Tool Memory is a specialized component of ReMe that captures and learns from tool usage patterns, enabling AI agents to improve their tool invocation strategies over time. This document explains how tool memory works and how to use it in your applications.
## What is Tool Memory?
Tool Memory represents knowledge extracted from historical tool invocations, including:
- Usage patterns and best practices for specific tools
- Common parameter configurations that lead to success or failure
- Performance characteristics (time cost, token cost, success rate)
- Actionable recommendations based on real usage data
Each tool memory contains:
- `when_to_use`: The tool name (used as the unique identifier)
- `content`: Synthesized usage guidelines and best practices
- `tool_call_results`: Historical invocation records with evaluations
- Statistical metrics about tool performance
## Tool Memory Data Structure
### ToolMemory
```python
class ToolMemory(BaseMemory):
memory_type: str = "tool"
workspace_id: str # Workspace identifier
memory_id: str # Unique memory ID
when_to_use: str # Tool name (serves as identifier)
content: str # Synthesized usage guidelines
score: float # Overall quality score
time_created: str # Creation timestamp
time_modified: str # Last modification timestamp
author: str # Creator (typically LLM model name)
tool_call_results: List[ToolCallResult] # Historical invocation records
metadata: dict # Additional metadata
```
### ToolCallResult
```python
class ToolCallResult(BaseModel):
create_time: str # Invocation timestamp
tool_name: str # Name of the tool
input: dict | str # Input parameters
output: str # Tool output
token_cost: int # Token consumption
success: bool # Whether invocation succeeded
time_cost: float # Time consumed (seconds)
summary: str # Brief summary of the result
evaluation: str # Detailed evaluation
score: float # Evaluation score (0.0, 0.5, or 1.0)
metadata: dict # Additional metadata
```
Tool Memory learns from each tool invocation by evaluating the result and accumulating insights over time.
## Configuration Logic
Tool Memory in ReMe is configured through three main flows:
### 1. Add Tool Call Result
The `add_tool_call_result` flow processes individual tool invocations and adds them to memory:
```yaml
add_tool_call_result:
flow_content: parse_tool_call_result_op >> update_vector_store_op
description: "Evaluates and adds tool call results to the tool memory database"
```
```mermaid
graph LR
A[Tool Call Results] --> B[parse_tool_call_result_op]
B --> C[Evaluate Each Call]
C --> D[Generate Summary & Score]
D --> E[Update/Create ToolMemory]
E --> F[update_vector_store_op]
F --> G[Store in Vector DB]
```
This flow:
1. Receives tool call results with input, output, and metadata
2. Evaluates each call using LLM (generates summary, evaluation, score)
3. Appends evaluated results to the tool's memory
4. Updates the vector store with the modified memory
### 2. Retrieve Tool Memory
The `retrieve_tool_memory` flow fetches usage guidelines for specific tools:
```yaml
retrieve_tool_memory:
flow_content: retrieve_tool_memory_op
description: "Retrieves tool memories from the vector database based on tool names"
```
```mermaid
graph LR
A[Tool Names] --> B[retrieve_tool_memory_op]
B --> C[Search by Tool Name]
C --> D[Exact Match Check]
D --> E[Return ToolMemory]
E --> F[Usage Guidelines + History]
```
This flow:
1. Takes comma-separated tool names as input
2. Searches the vector store for exact matches
3. Returns tool memories with usage guidelines and call history
### 3. Summary Tool Memory
The `summary_tool_memory` flow analyzes historical data and generates comprehensive usage guidelines:
```yaml
summary_tool_memory:
flow_content: summary_tool_memory_op >> update_vector_store_op
description: "Analyzes tool call history and generates comprehensive usage patterns"
```
```mermaid
graph LR
A[Tool Names] --> B[summary_tool_memory_op]
B --> C[Retrieve Tool Memory]
C --> D[Analyze Recent Calls]
D --> E[Calculate Statistics]
E --> F[Generate Guidelines]
F --> G[update_vector_store_op]
G --> H[Update Vector DB]
```
This flow:
1. Retrieves existing tool memories by tool name
2. Analyzes recent N tool calls (default: 20)
3. Calculates statistical metrics (success rate, avg score, costs)
4. Uses LLM to synthesize actionable usage guidelines
5. Updates the tool memory content with new insights
## Complete Interaction Flow
The following diagram illustrates how the three operations interact with the vector store and agent tool calls:
```mermaid
graph LR
Agent[Agent] -->|1. Before tool call| Retrieve[retrieve_tool_memory]
Retrieve -->|Read| VectorStore[(Vector Store)]
VectorStore -->|Usage Guidelines| Agent
Agent -->|2. Execute| Tool[Tool Call]
Tool -->|Result| Agent
Agent -->|3. After tool call| Add[add_tool_call_result]
Add -->|Evaluate & Write| VectorStore
Summary[summary_tool_memory] -->|4. Periodic| VectorStore
VectorStore -->|Read History| Summary
Summary -->|Update Guidelines| VectorStore
style Agent fill:#e1f5ff
style Retrieve fill:#fff4e1
style Add fill:#ffe1f5
style Summary fill:#e1ffe1
style VectorStore fill:#f0f0f0
```
### Workflow Steps
**1. retrieve_tool_memory** (Before tool execution)
- Agent queries usage guidelines from Vector Store
- Returns best practices and historical patterns
**2. Tool Call Execution**
- Agent executes tool with recommended parameters
- Gets result (success/failure, output, costs)
**3. add_tool_call_result** (After tool execution)
- Evaluates the tool call result
- Stores evaluated result to Vector Store
**4. summary_tool_memory** (Periodic)
- Analyzes accumulated call history
- Generates comprehensive usage guidelines
- Updates Vector Store with new insights
## Basic Usage
Here's how to use Tool Memory in your application:
### Step 1: Set Up Your Environment
```python
import requests
# API configuration
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "your_workspace_id"
```
### Step 2: Record Tool Call Results
After your agent executes a tool, record the invocation:
```python
# Example: Record a web search tool call
tool_call_results = [
{
"create_time": "2025-10-15 14:30:00",
"tool_name": "web_search",
"input": {
"query": "Python asyncio tutorial",
"max_results": 10,
"language": "en"
},
"output": "Found 10 relevant results including official docs and tutorials",
"token_cost": 150,
"success": True,
"time_cost": 2.3
}
]
# Add the tool call result
response = requests.post(
url=f"{BASE_URL}add_tool_call_result",
json={
"workspace_id": WORKSPACE_ID,
"tool_name": "web_search",
"tool_call_results": tool_call_results
}
)
```
### Step 3: Retrieve Tool Usage Guidelines
Before using a tool, retrieve its usage guidelines:
```python
# Retrieve memory for specific tools
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search,file_reader" # Comma-separated
}
)
memory_list = response.json().get("metadata", {}).get("memory_list", [])
for memory in memory_list:
print(f"Tool: {memory['when_to_use']}")
print(f"Guidelines: {memory['content']}")
print(f"Total Calls: {len(memory['tool_call_results'])}")
```
### Step 4: Generate Comprehensive Usage Guidelines
After accumulating sufficient call history, generate synthesized guidelines:
```python
# Summarize tool usage patterns
response = requests.post(
url=f"{BASE_URL}summary_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search"
}
)
if response.json().get("success"):
print("Successfully generated usage guidelines")
```
## Complete Example Workflow
Here's a complete example demonstrating the tool memory lifecycle:
```python
import requests
from datetime import datetime
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "demo_workspace"
def record_tool_usage(tool_name, input_params, output, success, time_cost, token_cost):
"""Record a single tool invocation"""
tool_call_result = {
"create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"),
"tool_name": tool_name,
"input": input_params,
"output": output,
"token_cost": token_cost,
"success": success,
"time_cost": time_cost
}
response = requests.post(
url=f"{BASE_URL}add_tool_call_result",
json={
"workspace_id": WORKSPACE_ID,
"tool_name": tool_name,
"tool_call_results": [tool_call_result]
}
)
return response.json()
def get_tool_guidelines(tool_name):
"""Retrieve usage guidelines for a tool"""
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
result = response.json()
if result.get("success"):
memory_list = result.get("metadata", {}).get("memory_list", [])
if memory_list:
return memory_list[0]
return None
def generate_guidelines(tool_name):
"""Generate comprehensive usage guidelines"""
response = requests.post(
url=f"{BASE_URL}summary_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
return response.json()
# Example usage
if __name__ == "__main__":
tool_name = "web_search"
# 1. Record multiple tool invocations
print("Recording tool invocations...")
for i in range(5):
record_tool_usage(
tool_name=tool_name,
input_params={"query": f"test query {i}", "max_results": 10},
output=f"Found results for query {i}",
success=True,
time_cost=2.0 + i * 0.5,
token_cost=100 + i * 20
)
# 2. Generate usage guidelines
print("\nGenerating usage guidelines...")
generate_guidelines(tool_name)
# 3. Retrieve and display guidelines
print("\nRetrieving guidelines...")
memory = get_tool_guidelines(tool_name)
if memory:
print(f"\nTool: {memory['when_to_use']}")
print(f"Guidelines:\n{memory['content']}")
print(f"\nTotal invocations: {len(memory['tool_call_results'])}")
```
## Use Cases
### Use Case 1: Learning Optimal Parameters
```mermaid
sequenceDiagram
participant Agent
participant ToolMemory
participant Tool
Agent->>ToolMemory: retrieve_tool_memory("api_caller")
ToolMemory-->>Agent: Guidelines: Use timeout=30s, retry=3
Agent->>Tool: Call with recommended params
Tool-->>Agent: Success (time: 2.5s)
Agent->>ToolMemory: add_tool_call_result(success=True)
ToolMemory->>ToolMemory: Update statistics
```
**Scenario**: An agent needs to call an external API. Tool Memory has learned from 50+ previous calls that:
- Setting `timeout=30s` achieves 95% success rate
- Using `retry=3` handles transient failures effectively
- Requests with `max_results > 100` often timeout
The agent retrieves these guidelines before making the call, leading to higher success rates.
### Use Case 2: Avoiding Common Pitfalls
```mermaid
sequenceDiagram
participant Agent
participant ToolMemory
participant FileReader
Agent->>ToolMemory: retrieve_tool_memory("file_reader")
ToolMemory-->>Agent: Warning: Large files (>10MB) cause timeouts
Agent->>Agent: Check file size first
Agent->>FileReader: Read file with streaming mode
FileReader-->>Agent: Success
Agent->>ToolMemory: add_tool_call_result(success=True)
```
**Scenario**: Tool Memory has recorded that the `file_reader` tool fails when:
- File paths contain special characters without escaping
- Files larger than 10MB are read without streaming mode
- Binary files are opened in text mode
The agent retrieves these warnings and adjusts its approach accordingly.
### Use Case 3: Performance Optimization
```python
# Before using Tool Memory
average_time_cost = 5.2s
success_rate = 75%
# After learning from Tool Memory
# - Use batch processing for multiple queries
# - Set appropriate timeout values
# - Cache frequently accessed data
average_time_cost = 2.8s # 46% improvement
success_rate = 92% # 17% improvement
```
**Scenario**: Tool Memory analyzes 100+ invocations of a `database_query` tool and discovers:
- Batch queries are 3x faster than individual queries
- Connection pooling reduces overhead by 40%
- Queries during peak hours (2-4 PM) have higher failure rates
The synthesized guidelines help the agent optimize its database interactions.
## Managing Tool Memories
### Delete a Workspace
```python
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "delete"
}
)
```
### Dump Memories to Disk
```python
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "dump",
"path": "./"
}
)
```
### Load Memories from Disk
```python
response = requests.post(
url=f"{BASE_URL}vector_store",
json={
"workspace_id": WORKSPACE_ID,
"action": "load",
"path": "./"
}
)
```
## Configuration Parameters
### ParseToolCallResultOp Parameters
Configure in `default.yaml`:
```yaml
op:
parse_tool_call_result_op:
backend: parse_tool_call_result_op
llm: default
params:
max_history_tool_call_cnt: 100 # Max historical calls to retain
evaluation_sleep_interval: 1.0 # Delay between evaluations (seconds)
```
- `max_history_tool_call_cnt`: Limits the number of historical tool call results stored per tool. Older results are removed when this limit is exceeded.
- `evaluation_sleep_interval`: Controls the delay between concurrent evaluations to avoid rate limiting.
### SummaryToolMemoryOp Parameters
```yaml
op:
summary_tool_memory_op:
backend: summary_tool_memory_op
llm: default
params:
recent_call_count: 20 # Number of recent calls to analyze
summary_sleep_interval: 1.0 # Delay between summaries (seconds)
```
- `recent_call_count`: Number of most recent tool calls to analyze when generating guidelines.
- `summary_sleep_interval`: Controls the delay between concurrent summarizations.
## Best Practices
1. **Regular Recording**:
- Record every tool invocation, including failures
- Include detailed input parameters and output
- Capture performance metrics (time_cost, token_cost)
2. **Periodic Summarization**:
- Generate guidelines after accumulating 20-50 tool calls
- Re-summarize when usage patterns change significantly
- Update guidelines when new tool versions are deployed
3. **Retrieval Strategy**:
- Always retrieve guidelines before using unfamiliar tools
- Cache retrieved guidelines for the duration of a task
- Re-retrieve after tool memory updates
4. **Quality Maintenance**:
- Monitor success rates and average scores
- Investigate tools with declining performance
- Clean up outdated memories when tools are deprecated
5. **Parameter Tuning**:
- Adjust `max_history_tool_call_cnt` based on tool usage frequency
- Increase `recent_call_count` for tools with diverse usage patterns
- Reduce `evaluation_sleep_interval` if rate limiting is not a concern
## Integration with Agent Workflows
```mermaid
graph TB
A[Agent Receives Task] --> B{Tool Required?}
B -->|Yes| C[Retrieve Tool Memory]
C --> D[Apply Guidelines]
D --> E[Execute Tool]
E --> F[Record Result]
F --> G{Sufficient History?}
G -->|Yes| H[Generate Summary]
G -->|No| I[Continue]
H --> I
B -->|No| I[Process Task]
I --> J[Task Complete]
```
Tool Memory seamlessly integrates into agent workflows:
1. Before tool execution: Retrieve usage guidelines
2. During execution: Apply recommended parameters
3. After execution: Record results with evaluation
4. Periodically: Generate updated guidelines
For more detailed examples, see the implementation in `reme_ai/summary/tool/` directory of the ReMe project.

View file

@ -0,0 +1,548 @@
# Tool Memory Retrieval Operations
## RetrieveToolMemoryOp
### Purpose
Retrieves tool memories from the vector database based on tool names, providing usage patterns, best practices, and historical call data.
### Functionality
- Accepts comma-separated tool names as input
- Searches the vector store for exact tool name matches
- Validates that retrieved memories are of type "tool"
- Returns complete tool memories including usage guidelines and call history
- Provides detailed logging for debugging and monitoring
### Processing Flow
```mermaid
graph TB
A[Receive Tool Names] --> B[Validate Input]
B --> C[Split by Comma]
C --> D[Trim Whitespace]
D --> E[For Each Tool Name]
E --> F[Search Vector Store]
F --> G{Results Found?}
G -->|Yes| H[Get Top Result]
G -->|No| I[Log Warning: Not Found]
H --> J{Type = tool?}
J -->|Yes| K{Name Matches?}
J -->|No| L[Log Warning: Wrong Type]
K -->|Yes| M[Add to Results]
K -->|No| N[Log Warning: Name Mismatch]
I --> O[Continue Next Tool]
L --> O
N --> O
M --> O
O --> P{More Tools?}
P -->|Yes| E
P -->|No| Q{Any Matches?}
Q -->|Yes| R[Return Memory List]
Q -->|No| S[Return Empty]
```
1. **Input Validation**:
- Check if `tool_names` parameter is provided
- Return error if empty
- Log workspace_id and tool count
2. **Tool Name Processing**:
- Split input by comma delimiter
- Strip whitespace from each tool name
- Filter out empty strings
- Log the list of tools to retrieve
3. **Vector Store Search**:
- For each tool name, search with `top_k=1`
- Use tool name as the query (exact match preferred)
- Retrieve the top matching result
4. **Result Validation**:
- Verify result is of type `ToolMemory`
- Check that `when_to_use` field exactly matches tool name
- Log match details (memory_id, total_calls)
- Warn if no match or mismatch found
5. **Response Preparation**:
- Collect all matched tool memories
- Set success status based on matches found
- Return memory list in metadata
### Parameters
This operation has no configurable parameters. It uses the default vector store configuration.
### Input Schema
```yaml
input_schema:
tool_names:
type: string
description: "Comma-separated tool names (e.g., 'tool_name1,tool_name2')"
required: true
```
### Output Format
The operation sets the following in `context.response`:
```python
{
"success": True,
"answer": "Successfully retrieved 2 tool memories",
"metadata": {
"memory_list": [
{
"workspace_id": "demo_workspace",
"memory_id": "abc123def456",
"memory_type": "tool",
"when_to_use": "web_search",
"content": "Core Function: The web_search tool retrieves...",
"score": 0.85,
"time_created": "2025-10-15 10:00:00",
"time_modified": "2025-10-15 14:30:00",
"author": "qwen3-30b-a3b-instruct-2507",
"tool_call_results": [
{
"create_time": "2025-10-15 14:30:00",
"tool_name": "web_search",
"input": {...},
"output": "...",
"summary": "...",
"evaluation": "...",
"score": 1.0,
"success": True,
"time_cost": 2.3,
"token_cost": 150
}
],
"metadata": {}
}
]
}
}
```
### Usage Example
#### Basic Retrieval
```python
import requests
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "demo_workspace"
# Retrieve memory for a single tool
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search"
}
)
result = response.json()
if result.get("success"):
memory_list = result.get("metadata", {}).get("memory_list", [])
for memory in memory_list:
print(f"Tool: {memory['when_to_use']}")
print(f"Guidelines:\n{memory['content']}")
print(f"Total Calls: {len(memory['tool_call_results'])}")
```
#### Multiple Tools Retrieval
```python
# Retrieve memories for multiple tools at once
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search,file_reader,api_caller,database_query"
}
)
result = response.json()
if result.get("success"):
memory_list = result.get("metadata", {}).get("memory_list", [])
print(f"Retrieved {len(memory_list)} tool memories")
for memory in memory_list:
print(f"\n{'='*60}")
print(f"Tool: {memory['when_to_use']}")
print(f"{'='*60}")
print(memory['content'])
```
#### Extracting Specific Information
```python
def get_tool_statistics(tool_name):
"""Get statistical information for a specific tool"""
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
result = response.json()
if not result.get("success"):
return None
memory_list = result.get("metadata", {}).get("memory_list", [])
if not memory_list:
return None
memory = memory_list[0]
tool_calls = memory['tool_call_results']
# Calculate statistics
total_calls = len(tool_calls)
successful_calls = sum(1 for call in tool_calls if call['success'])
avg_score = sum(call['score'] for call in tool_calls) / total_calls if total_calls > 0 else 0
avg_time = sum(call['time_cost'] for call in tool_calls) / total_calls if total_calls > 0 else 0
avg_tokens = sum(call['token_cost'] for call in tool_calls) / total_calls if total_calls > 0 else 0
return {
"tool_name": tool_name,
"total_calls": total_calls,
"success_rate": successful_calls / total_calls if total_calls > 0 else 0,
"avg_score": avg_score,
"avg_time_cost": avg_time,
"avg_token_cost": avg_tokens,
"guidelines": memory['content']
}
# Usage
stats = get_tool_statistics("web_search")
if stats:
print(f"Tool: {stats['tool_name']}")
print(f"Total Calls: {stats['total_calls']}")
print(f"Success Rate: {stats['success_rate']:.1%}")
print(f"Avg Score: {stats['avg_score']:.2f}")
print(f"Avg Time: {stats['avg_time_cost']:.2f}s")
print(f"Avg Tokens: {stats['avg_token_cost']:.0f}")
```
### Integration with Agent Workflows
#### Pre-Execution Retrieval
```python
def execute_tool_with_guidelines(tool_name, input_params):
"""Execute a tool after retrieving its usage guidelines"""
# 1. Retrieve tool memory
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
# 2. Extract guidelines
guidelines = ""
if response.json().get("success"):
memory_list = response.json().get("metadata", {}).get("memory_list", [])
if memory_list:
guidelines = memory_list[0]['content']
print(f"Guidelines for {tool_name}:")
print(guidelines)
# 3. Adjust parameters based on guidelines
# (This would be done by an LLM or rule-based system)
adjusted_params = adjust_parameters(input_params, guidelines)
# 4. Execute tool
result = execute_tool(tool_name, adjusted_params)
# 5. Record result
record_tool_call(tool_name, adjusted_params, result)
return result
```
#### Batch Retrieval for Agent Initialization
```python
def initialize_agent_with_tool_memories(available_tools):
"""Load all tool memories at agent initialization"""
# Retrieve all tool memories at once
tool_names = ",".join(available_tools)
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_names
}
)
# Build tool memory cache
tool_memory_cache = {}
if response.json().get("success"):
memory_list = response.json().get("metadata", {}).get("memory_list", [])
for memory in memory_list:
tool_memory_cache[memory['when_to_use']] = {
"guidelines": memory['content'],
"total_calls": len(memory['tool_call_results']),
"last_modified": memory['time_modified']
}
return tool_memory_cache
# Usage
available_tools = ["web_search", "file_reader", "api_caller"]
tool_cache = initialize_agent_with_tool_memories(available_tools)
# Agent can now quickly access guidelines
if "web_search" in tool_cache:
print(tool_cache["web_search"]["guidelines"])
```
### Error Handling
#### Tool Not Found
```python
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "nonexistent_tool"
}
)
result = response.json()
if not result.get("success"):
print(f"Error: {result.get('answer')}")
# Output: "No matching tool memories found"
```
#### Empty Tool Names
```python
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": ""
}
)
result = response.json()
if not result.get("success"):
print(f"Error: {result.get('answer')}")
# Output: "tool_names is required"
```
#### Partial Matches
```python
# Request 3 tools, but only 2 exist
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search,file_reader,nonexistent_tool"
}
)
result = response.json()
if result.get("success"):
memory_list = result.get("metadata", {}).get("memory_list", [])
print(f"Found {len(memory_list)} out of 3 requested tools")
# Output: "Found 2 out of 3 requested tools"
```
### Retrieval Workflow
```mermaid
sequenceDiagram
participant Agent
participant RetrieveOp
participant VectorStore
Agent->>RetrieveOp: retrieve_tool_memory(tool_names)
RetrieveOp->>RetrieveOp: Split and validate names
loop For each tool name
RetrieveOp->>VectorStore: search(tool_name, top_k=1)
VectorStore-->>RetrieveOp: Return top result
RetrieveOp->>RetrieveOp: Validate type and name
alt Valid match
RetrieveOp->>RetrieveOp: Add to results
else No match
RetrieveOp->>RetrieveOp: Log warning
end
end
RetrieveOp-->>Agent: Return memory_list
Agent->>Agent: Apply guidelines
```
### Use Cases
#### Use Case 1: Pre-Execution Guidance
**Scenario**: Before executing a tool, the agent retrieves usage guidelines to optimize parameters.
```python
# Agent needs to search the web
tool_name = "web_search"
query = "machine learning basics"
# Retrieve guidelines
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
# Guidelines suggest: max_results=10-20, use filter_type="technical_docs"
# Agent adjusts parameters accordingly
optimized_params = {
"query": query,
"max_results": 15, # Within recommended range
"filter_type": "technical_docs", # As suggested
"language": "en"
}
# Execute with optimized parameters
result = execute_tool(tool_name, optimized_params)
```
#### Use Case 2: Performance Monitoring
**Scenario**: Monitor tool performance trends over time.
```python
def monitor_tool_performance(tool_name):
"""Monitor tool performance and detect degradation"""
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_name
}
)
if not response.json().get("success"):
return None
memory = response.json()["metadata"]["memory_list"][0]
calls = memory["tool_call_results"]
# Analyze recent vs historical performance
recent_calls = calls[-20:] # Last 20 calls
historical_calls = calls[:-20] if len(calls) > 20 else []
recent_success_rate = sum(1 for c in recent_calls if c['success']) / len(recent_calls)
historical_success_rate = (sum(1 for c in historical_calls if c['success']) / len(historical_calls)
if historical_calls else recent_success_rate)
# Detect degradation
if recent_success_rate < historical_success_rate - 0.1:
print(f"Warning: {tool_name} performance degraded!")
print(f"Recent: {recent_success_rate:.1%}, Historical: {historical_success_rate:.1%}")
return "degraded"
return "healthy"
# Usage
status = monitor_tool_performance("web_search")
```
#### Use Case 3: Tool Selection
**Scenario**: Choose the best tool for a task based on historical performance.
```python
def select_best_tool(task_description, candidate_tools):
"""Select the best tool based on historical performance"""
# Retrieve memories for all candidate tools
tool_names = ",".join(candidate_tools)
response = requests.post(
url=f"{BASE_URL}retrieve_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": tool_names
}
)
if not response.json().get("success"):
return candidate_tools[0] # Default to first
memory_list = response.json()["metadata"]["memory_list"]
# Score each tool
tool_scores = {}
for memory in memory_list:
calls = memory["tool_call_results"]
if not calls:
continue
# Calculate composite score
success_rate = sum(1 for c in calls if c['success']) / len(calls)
avg_score = sum(c['score'] for c in calls) / len(calls)
avg_time = sum(c['time_cost'] for c in calls) / len(calls)
# Composite: prioritize success and quality, penalize slow tools
composite = (success_rate * 0.4 + avg_score * 0.4 + (1 / (1 + avg_time)) * 0.2)
tool_scores[memory['when_to_use']] = composite
# Select best tool
if tool_scores:
best_tool = max(tool_scores, key=tool_scores.get)
print(f"Selected {best_tool} with score {tool_scores[best_tool]:.2f}")
return best_tool
return candidate_tools[0]
# Usage
best = select_best_tool(
"Search for technical documentation",
["web_search", "doc_search", "api_search"]
)
```
## Best Practices
1. **Retrieval Timing**:
- Retrieve guidelines before first use of a tool
- Cache retrieved memories for the duration of a task
- Re-retrieve after tool memory updates (post-summarization)
2. **Batch Retrieval**:
- Retrieve multiple tool memories in a single request
- Use comma-separated tool names for efficiency
- Initialize agent with all available tool memories
3. **Error Handling**:
- Always check `success` status in response
- Handle cases where tool memory doesn't exist
- Provide fallback behavior for missing guidelines
4. **Guidelines Application**:
- Parse guidelines to extract parameter recommendations
- Use LLM to interpret guidelines in context
- Combine guidelines with task-specific requirements
5. **Performance Optimization**:
- Cache retrieved memories to avoid repeated calls
- Invalidate cache after summarization updates
- Monitor retrieval latency and optimize if needed
6. **Monitoring**:
- Log retrieval requests for debugging
- Track which tools are frequently retrieved
- Identify tools without memory (candidates for recording)

View file

@ -0,0 +1,465 @@
# Tool Summary Operations
## ParseToolCallResultOp
### Purpose
Evaluates individual tool invocations and adds them to the tool memory database with comprehensive assessments.
### Functionality
- Receives tool call results with input parameters, output, and metadata
- Uses LLM to evaluate each tool call based on success and parameter alignment
- Generates summary, evaluation, and score (0.0, 0.5, or 1.0) for each call
- Appends evaluated results to existing tool memory or creates new memory
- Maintains a sliding window of recent tool calls (configurable limit)
### Processing Flow
```mermaid
graph TB
A[Receive Tool Call Results] --> B[Validate Input]
B --> C[Search for Existing Memory]
C --> D{Memory Exists?}
D -->|Yes| E[Load Existing Memory]
D -->|No| F[Create New Memory]
E --> G[Concurrent Evaluation]
F --> G
G --> H[Evaluate Call #1]
G --> I[Evaluate Call #2]
G --> J[Evaluate Call #N]
H --> K[Generate Summary & Score]
I --> K
J --> K
K --> L[Append to Memory]
L --> M[Trim to Max History]
M --> N[Update Modified Time]
N --> O[Prepare for Vector Store]
```
1. **Input Validation**:
- Verify `tool_name` is provided
- Convert dict objects to `ToolCallResult` instances
- Check if `tool_call_results` list is not empty
2. **Memory Lookup**:
- Search vector store for existing tool memory by tool name
- Verify exact match (memory type = "tool" and when_to_use = tool_name)
- Create new `ToolMemory` if no match found
3. **Concurrent Evaluation**:
- Submit all tool call results for parallel evaluation
- Each evaluation uses LLM to analyze the call
- Generate structured evaluation with summary, assessment, and score
4. **Memory Update**:
- Append evaluated results to tool memory
- Trim to `max_history_tool_call_cnt` if limit exceeded
- Update modification timestamp
- Prepare for vector store update
### Parameters
Configure in `default.yaml`:
```yaml
op:
parse_tool_call_result_op:
backend: parse_tool_call_result_op
llm: default
params:
max_history_tool_call_cnt: 100
evaluation_sleep_interval: 1.0
```
- `max_history_tool_call_cnt` (integer, default: `100`):
- Maximum number of historical tool call results to retain per tool
- When exceeded, oldest results are removed (FIFO)
- Balances memory size with historical context
- Recommended: 50-200 depending on tool usage frequency
- `evaluation_sleep_interval` (float, default: `1.0`):
- Delay in seconds between concurrent evaluations
- Prevents rate limiting when evaluating multiple calls
- Set to 0 for maximum speed (if no rate limits)
- Increase if encountering API throttling
### Evaluation Criteria
The LLM evaluates each tool call based on two dimensions:
1. **Success Evaluation**:
- Check if the success flag indicates successful execution
- Verify output contains no error messages
- Assess if time and token costs are reasonable
- Identify any error indicators in the output
2. **Parameter Alignment Evaluation**:
- Evaluate if output matches expected behavior given input
- Consider if input parameters are appropriate for the tool
- Check for parameter mismatches or unexpected behaviors
- Verify output is consistent with tool's intended purpose
### Scoring Guidelines
- **1.0 (Success)**: Tool executed successfully with good parameter alignment
- Example: Query returned relevant results, parameters were appropriate
- **0.5 (Partial Success)**: Tool executed but with issues
- Example: Query succeeded but parameters were suboptimal (e.g., too generic)
- Example: Results returned but with warnings about invalid parameters
- **0.0 (Failure)**: Tool execution failed or severe parameter misalignment
- Example: Timeout due to excessive max_results parameter
- Example: Error due to invalid parameter format
### Output Format
The operation sets the following in `context.response.metadata`:
```python
{
"deleted_memory_ids": ["memory_id_if_updating"],
"memory_list": [
{
"memory_id": "abc123",
"when_to_use": "web_search",
"tool_call_results": [
{
"create_time": "2025-10-15 14:30:00",
"tool_name": "web_search",
"input": {...},
"output": "...",
"summary": "Successfully retrieved 10 relevant results",
"evaluation": "Good parameter alignment...",
"score": 1.0,
"success": True,
"time_cost": 2.3,
"token_cost": 150
}
]
}
]
}
```
### Usage Example
```python
import requests
from datetime import datetime
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "demo_workspace"
# Prepare tool call results
tool_call_results = [
{
"create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"),
"tool_name": "web_search",
"input": {
"query": "Python asyncio tutorial",
"max_results": 10,
"language": "en",
"filter_type": "technical_docs"
},
"output": "Found 10 relevant results including official documentation and tutorials",
"token_cost": 150,
"success": True,
"time_cost": 2.3
},
{
"create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"),
"tool_name": "web_search",
"input": {
"query": "test", # Too generic
"max_results": 100, # Too many
"language": "unknown" # Invalid
},
"output": "Warning: language 'unknown' not supported. Query too generic, limited results.",
"token_cost": 80,
"success": True,
"time_cost": 3.5
}
]
# Add tool call results
response = requests.post(
url=f"{BASE_URL}add_tool_call_result",
json={
"workspace_id": WORKSPACE_ID,
"tool_name": "web_search",
"tool_call_results": tool_call_results
}
)
result = response.json()
print(f"Success: {result.get('success')}")
print(f"Answer: {result.get('answer')}")
# Check evaluated results
memory_list = result.get("metadata", {}).get("memory_list", [])
if memory_list:
tool_memory = memory_list[0]
for call_result in tool_memory["tool_call_results"]:
print(f"\nCall Summary: {call_result['summary']}")
print(f"Evaluation: {call_result['evaluation']}")
print(f"Score: {call_result['score']}")
```
## SummaryToolMemoryOp
### Purpose
Analyzes accumulated tool call history and generates comprehensive usage patterns, best practices, and recommendations.
### Functionality
- Retrieves existing tool memories from the vector store
- Analyzes the most recent N tool calls (configurable)
- Calculates statistical metrics (success rate, average scores, costs)
- Uses LLM to synthesize actionable usage guidelines
- Updates tool memory content with generated insights
### Processing Flow
```mermaid
graph TB
A[Receive Tool Names] --> B[Split by Comma]
B --> C[For Each Tool Name]
C --> D[Search Vector Store]
D --> E{Exact Match?}
E -->|Yes| F[Load Tool Memory]
E -->|No| G[Log Warning]
F --> H[Extract Recent N Calls]
H --> I[Calculate Statistics]
I --> J[Format Call Summaries]
J --> K[Format Statistics]
K --> L[Concurrent Summarization]
L --> M[LLM Analysis #1]
L --> N[LLM Analysis #2]
L --> O[LLM Analysis #N]
M --> P[Generate Guidelines]
N --> P
O --> P
P --> Q[Update Memory Content]
Q --> R[Update Modified Time]
R --> S[Prepare for Vector Store]
```
1. **Tool Name Processing**:
- Split comma-separated tool names
- Trim whitespace from each name
- Log the list of tools to process
2. **Memory Retrieval**:
- Search vector store for each tool name
- Verify exact match (memory type and when_to_use)
- Skip tools without existing memory
3. **Data Preparation**:
- Extract the most recent N tool call results
- Calculate statistical metrics:
- Total calls vs recent calls analyzed
- Success rate (overall and recent)
- Average score (overall and recent)
- Average time cost
- Average token cost
- Format call summaries as markdown
- Format statistics as markdown
4. **Concurrent Summarization**:
- Submit all tools for parallel summarization
- Each summarization uses LLM to analyze patterns
- Generate structured usage guidelines
5. **Memory Update**:
- Update tool memory content with new guidelines
- Update modification timestamp
- Prepare for vector store update
### Parameters
Configure in `default.yaml`:
```yaml
op:
summary_tool_memory_op:
backend: summary_tool_memory_op
llm: default
params:
recent_call_count: 20
summary_sleep_interval: 1.0
```
- `recent_call_count` (integer, default: `20`):
- Number of most recent tool calls to analyze
- Focuses on recent usage patterns
- Recommended: 10-50 depending on tool usage frequency
- Higher values provide more context but may dilute recent patterns
- `summary_sleep_interval` (float, default: `1.0`):
- Delay in seconds between concurrent summarizations
- Prevents rate limiting when summarizing multiple tools
- Set to 0 for maximum speed (if no rate limits)
- Increase if encountering API throttling
### Statistical Metrics
The operation calculates the following metrics:
```python
{
"total_calls": 100, # Total number of calls in history
"recent_calls": 20, # Number of recent calls analyzed
"success_rate": 0.85, # Overall success rate (85%)
"recent_success_rate": 0.90, # Recent success rate (90%)
"avg_score": 0.78, # Average evaluation score
"recent_avg_score": 0.82, # Recent average score
"avg_time_cost": 2.45, # Average time in seconds
"avg_token_cost": 125.3 # Average token consumption
}
```
### Generated Guidelines Structure
The LLM generates guidelines following this structure:
1. **Core Function**: What the tool does and when to use it
2. **Success Patterns**: Parameter patterns and scenarios that work well
3. **Common Issues**: Main pitfalls to avoid and why they fail
4. **Best Practices**: 2-3 actionable recommendations
### Usage Example
```python
import requests
BASE_URL = "http://0.0.0.0:8002/"
WORKSPACE_ID = "demo_workspace"
# After accumulating tool call history, generate guidelines
response = requests.post(
url=f"{BASE_URL}summary_tool_memory",
json={
"workspace_id": WORKSPACE_ID,
"tool_names": "web_search,file_reader,api_caller" # Multiple tools
}
)
result = response.json()
print(f"Success: {result.get('success')}")
print(f"Answer: {result.get('answer')}")
# Display generated guidelines
memory_list = result.get("metadata", {}).get("memory_list", [])
for memory in memory_list:
print(f"\n{'='*60}")
print(f"Tool: {memory['when_to_use']}")
print(f"{'='*60}")
print(memory['content'])
# Display statistics
stats = memory.get('metadata', {}).get('statistics', {})
print(f"\nStatistics:")
print(f" Total Calls: {stats.get('total_calls', 0)}")
print(f" Success Rate: {stats.get('success_rate', 0):.1%}")
print(f" Avg Score: {stats.get('avg_score', 0):.2f}")
```
### Example Generated Guidelines
```
Core Function:
The web_search tool retrieves information from the internet based on query parameters.
Use it when you need up-to-date information, documentation, or external data.
Success Patterns:
- Specific queries (e.g., "Python asyncio tutorial") achieve 95% success rate
- Setting max_results=10-20 balances quality and performance
- Using language="en" and filter_type="technical_docs" improves relevance
- Average successful call: 2.3s, 150 tokens
Common Issues:
- Generic queries (e.g., "test") return poor results (score: 0.5)
- max_results > 50 often leads to timeouts (avg: 8.2s vs 2.3s)
- Invalid language codes default to English but add latency
- Missing filter_type returns mixed-quality results
Best Practices:
1. Use specific, descriptive queries with clear intent
2. Set max_results=10-20 for optimal balance
3. Always specify language and filter_type for technical searches
```
### When to Run Summarization
- **Initial Setup**: After accumulating 20-30 tool calls
- **Regular Updates**: Every 50-100 new calls
- **Pattern Changes**: When success rate changes significantly
- **Tool Updates**: After tool version changes or parameter updates
- **Performance Issues**: When investigating declining performance
### Integration with Retrieval
```mermaid
sequenceDiagram
participant Agent
participant AddOp as add_tool_call_result
participant SummaryOp as summary_tool_memory
participant RetrieveOp as retrieve_tool_memory
loop Every Tool Call
Agent->>AddOp: Record tool call result
AddOp->>AddOp: Evaluate and store
end
Note over Agent,SummaryOp: After 20+ calls
Agent->>SummaryOp: Generate guidelines
SummaryOp->>SummaryOp: Analyze patterns
SummaryOp->>SummaryOp: Update content
Note over Agent,RetrieveOp: Before next use
Agent->>RetrieveOp: Get tool memory
RetrieveOp-->>Agent: Return guidelines + history
Agent->>Agent: Apply recommendations
```
The summarization operation works in conjunction with retrieval:
1. `add_tool_call_result` continuously records invocations
2. `summary_tool_memory` periodically generates guidelines
3. `retrieve_tool_memory` provides guidelines before tool use
4. Agent applies recommendations to improve success rates
## Best Practices
1. **Recording Strategy**:
- Record every tool invocation, including failures
- Include detailed input parameters and complete output
- Capture accurate performance metrics
- Add relevant metadata for context
2. **Evaluation Quality**:
- Ensure LLM has sufficient context for evaluation
- Monitor score distribution (should not be all 1.0 or 0.0)
- Review evaluations periodically for quality
- Adjust evaluation prompts if needed
3. **Summarization Timing**:
- Wait for 20-30 calls before first summarization
- Re-summarize after significant new data (50+ calls)
- Update when usage patterns change
- Regenerate after tool updates
4. **Parameter Tuning**:
- Adjust `max_history_tool_call_cnt` based on tool usage frequency
- Increase `recent_call_count` for tools with diverse patterns
- Balance `sleep_interval` between speed and rate limits
- Monitor vector store size and adjust retention limits
5. **Quality Maintenance**:
- Review generated guidelines for accuracy
- Validate statistical metrics match expectations
- Clean up deprecated tools from vector store
- Archive historical data before major changes