From 40538f974a63cbb65f7a30b97edbf4439c544374 Mon Sep 17 00:00:00 2001 From: "jinli.yl" Date: Thu, 16 Oct 2025 16:39:40 +0800 Subject: [PATCH] docs(task_memory): simplify documentation and remove redundant content - Removed detailed data structure definitions for TaskMemory and Trajectory - Eliminated extensive examples and usage instructions for record_task_memory and delete_task_memory - Removed comparison table with Tool Memory and best practices section - Simplified build_query_op and rewrite_memory_op documentation by removing processing flows - Removed detailed examples and parameter descriptions for simple_summary_op- Added new tool_memory.md documentation with complete tool memory implementation - Added new tool_retrieve_ops.md documentation with detailed retrieval operations- Created comprehensive tool memory data structures and API documentation - Added usage examples and integration workflows for tool memory operations - Documented configuration parameters and best practices for tool memory management --- docs/task_memory/task_memory.md | 150 +------ docs/task_memory/task_retrieve_ops.md | 119 +----- docs/task_memory/task_summary_ops.md | 57 +-- docs/tool_memory/tool_memory.md | 557 ++++++++++++++++++++++++++ docs/tool_memory/tool_retrieve_ops.md | 548 +++++++++++++++++++++++++ docs/tool_memory/tool_summary_ops.md | 465 +++++++++++++++++++++ 6 files changed, 1580 insertions(+), 316 deletions(-) create mode 100644 docs/tool_memory/tool_memory.md create mode 100644 docs/tool_memory/tool_retrieve_ops.md create mode 100644 docs/tool_memory/tool_summary_ops.md diff --git a/docs/task_memory/task_memory.md b/docs/task_memory/task_memory.md index a3e07295..ae66d9f3 100644 --- a/docs/task_memory/task_memory.md +++ b/docs/task_memory/task_memory.md @@ -11,39 +11,9 @@ Task Memory represents knowledge extracted from previous task executions, includ Each task memory contains: - `when_to_use`: Conditions that indicate when this memory is relevant -- `content`: The actual knowledge or experience to be applied -- `score`: Quality score of the memory (assigned during validation) +- `content`: The actual knowledge or memory to be applied - Metadata about the memory's source and utility -## Task Memory的数据结构 - -### TaskMemory - -```python -class TaskMemory(BaseMemory): - memory_type: str = "task" - workspace_id: str # 工作空间ID - memory_id: str # 记忆的唯一ID - when_to_use: str # 何时使用此记忆的条件 - content: str # 具体的经验内容 - score: float # 质量评分(validation阶段设置) - time_created: str # 创建时间 - time_modified: str # 最后修改时间 - author: str # 创建者(通常是LLM模型名) - metadata: dict # 其他元数据 -``` - -### Trajectory - -```python -class Trajectory(BaseModel): - messages: List[Message] # 对话消息列表 - score: float # 轨迹评分(用于判断成功/失败) - metadata: dict # 可选的元数据(如segments、query等) -``` - -Task Memory通过分析Trajectory来提取经验。轨迹的score决定了它被分类为成功还是失败案例。 - ## Configuration Logic Task Memory in ReMe is configured through two main flows: @@ -226,119 +196,7 @@ response = requests.post( ReMe also provides additional task memory operations: -### Record Task Memory +- `record_task_memory`: Update frequency and utility attributes of retrieved memories +- `delete_task_memory`: Delete memories based on utility/frequency thresholds -`record_task_memory` flow用于更新检索到的任务记忆的使用频率和效用属性: - -```yaml -record_task_memory: - flow_content: update_memory_freq_op >> update_memory_utility_op >> update_vector_store_op - description: "Update the freq & utility attributes of retrieved task memories" - input_schema: - workspace_id: - type: string - required: true - memory_dicts: - type: array - description: "A list of retrieved task memory" - required: true - update_utility: - type: boolean - description: "Whether to update the utility attribute" - required: true -``` - -使用示例: - -```python -# 记录检索到的记忆被使用 -response = requests.post( - url=f"{BASE_URL}record_task_memory", - json={ - "workspace_id": WORKSPACE_ID, - "memory_dicts": retrieved_memories, # 从retrieve_task_memory获取的记忆列表 - "update_utility": True # 是否更新效用值 - } -) -``` - -### Delete Task Memory - -`delete_task_memory` flow用于删除低效用的任务记忆: - -```yaml -delete_task_memory: - flow_content: delete_memory_op >> update_vector_store_op - description: "Delete task memories when utility/freq < utility_threshold and freq >= freq_threshold" - input_schema: - workspace_id: - type: string - required: true - freq_threshold: - type: integer - description: "The retrieved frequency threshold" - required: true - utility_threshold: - type: number - description: "The utility/freq threshold" - required: true -``` - -使用示例: - -```python -# 删除使用频率高但效用低的记忆 -response = requests.post( - url=f"{BASE_URL}delete_task_memory", - json={ - "workspace_id": WORKSPACE_ID, - "freq_threshold": 10, # 至少被检索过10次 - "utility_threshold": 0.3 # 但效用/频率比 < 0.3 - } -) -``` - -这个机制确保记忆库的质量:频繁被检索但实际没什么帮助的记忆会被清理。 - -## 与Tool Memory的对比 - -| 特性 | Task Memory | Tool Memory | -|------|-------------|-------------| -| **目的** | 记录任务解决的经验 | 记录工具调用的使用模式 | -| **输入** | 对话轨迹(Trajectory) | 工具调用记录(ToolCallResult) | -| **索引方式** | when_to_use条件(语义搜索) | 工具名称(精确匹配) | -| **提取方式** | 从完整轨迹中分析提取 | 从单次调用评估和统计 | -| **内容** | 任务解决经验和策略 | 工具使用指南和最佳实践 | -| **更新方式** | 从成功/失败/对比中提取 | 每次调用后追加记录 | -| **检索流程** | 复杂:构建查询→召回→重排→重写 | 简单:精确匹配→返回 | -| **去重机制** | 语义去重(deduplication) | 按工具名唯一,限制历史条数 | -| **验证机制** | LLM验证质量(validation) | 评估单次调用质量 | - -## 最佳实践 - -1. **轨迹评分标准**: - - 明确定义成功和失败的评分标准 - - `success_threshold`通常设为1.0(完美成功) - - 部分成功的轨迹(如0.5-0.9)可以提供对比经验 - -2. **轨迹质量**: - - 确保轨迹包含完整的问题解决过程 - - 包含足够的上下文信息在metadata中 - - 避免过短或过长的轨迹 - -3. **定期维护**: - - 使用`record_task_memory`跟踪记忆使用情况 - - 定期使用`delete_task_memory`清理低效记忆 - - 保持记忆库的质量和相关性 - -4. **合理使用简化版本**: - - 简单场景使用`summary_task_memory_simple` - - 复杂场景使用完整的`summary_task_memory` flow - - 根据实际需求选择是否启用轨迹分割 - -5. **检索优化**: - - 调整`top_k`参数(默认5)控制返回的记忆数量 - - 启用`enable_llm_rerank`提升相关性 - - 启用`enable_llm_rewrite`使记忆更适配当前任务 - -For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project. +For more detailed examples, see the `use_task_memory_demo.py` file in the cookbook directory of the ReMe project. \ No newline at end of file diff --git a/docs/task_memory/task_retrieve_ops.md b/docs/task_memory/task_retrieve_ops.md index 6c61c88f..f4ec6eb5 100644 --- a/docs/task_memory/task_retrieve_ops.md +++ b/docs/task_memory/task_retrieve_ops.md @@ -8,61 +8,16 @@ Constructs a query for memory retrieval either from a direct query input or by a ### Functionality -- If a direct `query` is provided in the context, it uses that query directly +- If a direct `query` is provided in the context, it uses that query - If `messages` are provided in the context, it can: - - **LLM-based**: Use an LLM to generate a focused query based on the conversation context - - **Simple**: Create a simple query from the last 3 messages without using an LLM -- Raises an error if neither `query` nor `messages` is provided - -### Processing Flow - -1. **Check for Direct Query**: - - If `context.query` exists, use it directly - - This is the most common case when retrieving for a specific task - -2. **Build from Messages** (if no direct query): - - Extract last 3 messages from `context.messages` - - If `enable_llm_build=true`: - - Merge all messages into execution_process text - - Use LLM to analyze and generate a focused query - - If `enable_llm_build=false`: - - Truncate each message content to 200 chars - - Format as "- role: content" for each message - - Concatenate into a simple query - -3. **Set Context**: Store the built query in `context.query` for downstream ops + - Use an LLM to generate a query based on the conversation context + - Or create a simple query from recent messages without using an LLM ### Parameters - `op.build_query_op.params.enable_llm_build` (boolean, default: `true`): - When `true`, uses an LLM to generate a query from conversation messages - When `false`, creates a simple query by concatenating recent messages - - LLM-based queries are more focused but require an extra LLM call - - Simple queries are faster but may be less targeted - -### Usage Example - -```python -import requests - -# Method 1: Direct query (recommended) -response = requests.post( - url="http://0.0.0.0:8002/retrieve_task_memory", - json={ - "workspace_id": "my_workspace", - "query": "How to handle file upload errors in Python?" - } -) - -# Method 2: From messages (when integrated in conversation) -# This is handled internally by the flow when messages are provided -``` - -### When to Use Which Mode - -- **Direct Query**: Most common, use when you know what to search for -- **LLM-based Build**: When you have a conversation and want the LLM to extract the key question -- **Simple Build**: When you want fast retrieval without extra LLM cost ## RerankMemoryOp @@ -98,78 +53,12 @@ Rewrites and formats the retrieved memories to make them more relevant and actio - Formats retrieved memories into a structured format - Can use an LLM to rewrite memories to better fit the current context (optional) - Generates a cohesive context message from multiple memories -- Extracts current conversation context to inform the rewriting process - -### Processing Flow - -1. **Get Inputs**: - - `memory_list`: Retrieved and reranked memories from previous ops - - `query`: The search query - - `messages`: Current conversation messages (optional) - -2. **Format Memories**: - - Create numbered memory entries with "When to use" and "Content" - - Example format: - ``` - Memory 1: - When to use: [condition] - Content: [experience] - ``` - -3. **LLM Rewrite** (if `enable_llm_rewrite=true`): - - Extract current context from last 3 messages - - Provide original formatted memories and current context to LLM - - LLM adapts the memories to be more relevant to the current query - - Extracts `rewritten_context` from LLM response (JSON format) - -4. **Return**: Set `context.response.answer` to the final formatted/rewritten memories ### Parameters - `op.rewrite_memory_op.params.enable_llm_rewrite` (boolean, default: `true`): - When `true`, uses an LLM to rewrite the memories to make them more relevant and actionable - When `false`, simply formats the memories without LLM-based rewriting - - LLM rewriting makes memories more contextual but adds latency - -### Format Examples - -**Without LLM Rewrite** (enable_llm_rewrite=false): -``` -Memory 1 : - When to use: When handling file upload errors in web applications - Content: Use try-except blocks to catch specific exceptions like PermissionError... - -Memory 2 : - When to use: When validating file types before processing - Content: Check file extensions and MIME types to ensure security... -``` - -**With LLM Rewrite** (enable_llm_rewrite=true): -``` -Based on your question about file upload errors in Python web applications: - -1. Error Handling Strategy: - - Wrap file operations in try-except blocks - - Catch specific exceptions: PermissionError, IOError, OSError - - [Adapted from Memory 1] - -2. Input Validation: - - Always validate file types before processing - - Check both extension and MIME type - - [Adapted from Memory 2] -``` - -### When to Enable LLM Rewrite - -- **Enable** when: - - Memories need to be adapted to specific context - - User query is complex or nuanced - - Better integration with conversation flow is needed - -- **Disable** when: - - Fast retrieval is prioritized - - Memories are already clear and relevant - - Reducing LLM calls and costs is important ## MergeMemoryOp @@ -181,4 +70,4 @@ An alternative to RewriteMemoryOp that merges multiple memories into a single re - Collects the content from all memories in the memory list - Formats them into a single response with a standard structure -- Adds a prompt to consider the helpful parts when answering the question +- Adds a prompt to consider the helpful parts when answering the question \ No newline at end of file diff --git a/docs/task_memory/task_summary_ops.md b/docs/task_memory/task_summary_ops.md index f1364baf..db6dd1af 100644 --- a/docs/task_memory/task_summary_ops.md +++ b/docs/task_memory/task_summary_ops.md @@ -138,66 +138,13 @@ A simplified version of memory extraction that processes entire trajectories in ### Functionality - Classifies trajectories as success or failure based on score threshold -- Extracts memories directly from complete trajectories without segmentation -- Uses a single LLM prompt to extract multiple memories from each trajectory -- Parses JSON response containing `when_to_use` and `memory` pairs +- Extracts memories directly from complete trajectories - Useful for simpler use cases where detailed segmentation is not required -### Processing Flow - -1. **Receive Trajectories**: Get a list of Trajectory objects with messages and scores -2. **For Each Trajectory**: - - Merge all messages into execution_process text - - Determine execution_result ("success" or "fail") based on score threshold - - Format prompt with execution_process and execution_result - - Call LLM to extract memories in JSON format -3. **Parse Response**: Extract JSON array of {when_to_use, memory} objects -4. **Create TaskMemory**: Create TaskMemory objects for each extracted memory -5. **Return**: Set memory_list in context.response.metadata - ### Parameters - `op.simple_summary_op.params.success_score_threshold` (float, default: `0.9`): - The threshold score that determines if a trajectory is considered successful - - Trajectories with `score >= success_score_threshold` are marked as "success" - - Trajectories with `score < success_score_threshold` are marked as "fail" - -### Usage Example - -```python -import requests - -# Prepare trajectory with messages -trajectory = { - "messages": [ - {"role": "user", "content": "How do I sort a list in Python?"}, - {"role": "assistant", "content": "You can use the sorted() function..."}, - {"role": "user", "content": "Thanks! What about reverse sorting?"}, - {"role": "assistant", "content": "Use sorted(list, reverse=True)..."} - ], - "score": 1.0 # Successful trajectory -} - -# Summarize using simple summary -response = requests.post( - url="http://0.0.0.0:8002/summary_task_memory_simple", - json={ - "workspace_id": "my_workspace", - "trajectories": [trajectory] - } -) -``` - -### Comparison with Full Summary Pipeline - -| Feature | SimpleSummaryOp | Full Summary Pipeline | -|---------|-----------------|----------------------| -| **Complexity** | Single operation | Multi-stage pipeline | -| **Segmentation** | No segmentation | Optional segmentation | -| **Extraction** | Single prompt per trajectory | Separate prompts for success/failure/comparative | -| **Validation** | No validation | LLM validation step | -| **Deduplication** | No deduplication | Optional deduplication | -| **Use Case** | Quick prototyping | Production environments | ## SimpleComparativeSummaryOp @@ -213,4 +160,4 @@ A simplified version of comparative memory extraction. ### Parameters -No specific parameters beyond the LLM configuration. +No specific parameters beyond the LLM configuration. \ No newline at end of file diff --git a/docs/tool_memory/tool_memory.md b/docs/tool_memory/tool_memory.md new file mode 100644 index 00000000..024a6409 --- /dev/null +++ b/docs/tool_memory/tool_memory.md @@ -0,0 +1,557 @@ +# Tool Memory in ReMe + +Tool Memory is a specialized component of ReMe that captures and learns from tool usage patterns, enabling AI agents to improve their tool invocation strategies over time. This document explains how tool memory works and how to use it in your applications. + +## What is Tool Memory? + +Tool Memory represents knowledge extracted from historical tool invocations, including: +- Usage patterns and best practices for specific tools +- Common parameter configurations that lead to success or failure +- Performance characteristics (time cost, token cost, success rate) +- Actionable recommendations based on real usage data + +Each tool memory contains: +- `when_to_use`: The tool name (used as the unique identifier) +- `content`: Synthesized usage guidelines and best practices +- `tool_call_results`: Historical invocation records with evaluations +- Statistical metrics about tool performance + +## Tool Memory Data Structure + +### ToolMemory + +```python +class ToolMemory(BaseMemory): + memory_type: str = "tool" + workspace_id: str # Workspace identifier + memory_id: str # Unique memory ID + when_to_use: str # Tool name (serves as identifier) + content: str # Synthesized usage guidelines + score: float # Overall quality score + time_created: str # Creation timestamp + time_modified: str # Last modification timestamp + author: str # Creator (typically LLM model name) + tool_call_results: List[ToolCallResult] # Historical invocation records + metadata: dict # Additional metadata +``` + +### ToolCallResult + +```python +class ToolCallResult(BaseModel): + create_time: str # Invocation timestamp + tool_name: str # Name of the tool + input: dict | str # Input parameters + output: str # Tool output + token_cost: int # Token consumption + success: bool # Whether invocation succeeded + time_cost: float # Time consumed (seconds) + summary: str # Brief summary of the result + evaluation: str # Detailed evaluation + score: float # Evaluation score (0.0, 0.5, or 1.0) + metadata: dict # Additional metadata +``` + +Tool Memory learns from each tool invocation by evaluating the result and accumulating insights over time. + +## Configuration Logic + +Tool Memory in ReMe is configured through three main flows: + +### 1. Add Tool Call Result + +The `add_tool_call_result` flow processes individual tool invocations and adds them to memory: + +```yaml +add_tool_call_result: + flow_content: parse_tool_call_result_op >> update_vector_store_op + description: "Evaluates and adds tool call results to the tool memory database" +``` + +```mermaid +graph LR + A[Tool Call Results] --> B[parse_tool_call_result_op] + B --> C[Evaluate Each Call] + C --> D[Generate Summary & Score] + D --> E[Update/Create ToolMemory] + E --> F[update_vector_store_op] + F --> G[Store in Vector DB] +``` + +This flow: +1. Receives tool call results with input, output, and metadata +2. Evaluates each call using LLM (generates summary, evaluation, score) +3. Appends evaluated results to the tool's memory +4. Updates the vector store with the modified memory + +### 2. Retrieve Tool Memory + +The `retrieve_tool_memory` flow fetches usage guidelines for specific tools: + +```yaml +retrieve_tool_memory: + flow_content: retrieve_tool_memory_op + description: "Retrieves tool memories from the vector database based on tool names" +``` + +```mermaid +graph LR + A[Tool Names] --> B[retrieve_tool_memory_op] + B --> C[Search by Tool Name] + C --> D[Exact Match Check] + D --> E[Return ToolMemory] + E --> F[Usage Guidelines + History] +``` + +This flow: +1. Takes comma-separated tool names as input +2. Searches the vector store for exact matches +3. Returns tool memories with usage guidelines and call history + +### 3. Summary Tool Memory + +The `summary_tool_memory` flow analyzes historical data and generates comprehensive usage guidelines: + +```yaml +summary_tool_memory: + flow_content: summary_tool_memory_op >> update_vector_store_op + description: "Analyzes tool call history and generates comprehensive usage patterns" +``` + +```mermaid +graph LR + A[Tool Names] --> B[summary_tool_memory_op] + B --> C[Retrieve Tool Memory] + C --> D[Analyze Recent Calls] + D --> E[Calculate Statistics] + E --> F[Generate Guidelines] + F --> G[update_vector_store_op] + G --> H[Update Vector DB] +``` + +This flow: +1. Retrieves existing tool memories by tool name +2. Analyzes recent N tool calls (default: 20) +3. Calculates statistical metrics (success rate, avg score, costs) +4. Uses LLM to synthesize actionable usage guidelines +5. Updates the tool memory content with new insights + +## Complete Interaction Flow + +The following diagram illustrates how the three operations interact with the vector store and agent tool calls: + +```mermaid +graph LR + Agent[Agent] -->|1. Before tool call| Retrieve[retrieve_tool_memory] + Retrieve -->|Read| VectorStore[(Vector Store)] + VectorStore -->|Usage Guidelines| Agent + + Agent -->|2. Execute| Tool[Tool Call] + Tool -->|Result| Agent + + Agent -->|3. After tool call| Add[add_tool_call_result] + Add -->|Evaluate & Write| VectorStore + + Summary[summary_tool_memory] -->|4. Periodic| VectorStore + VectorStore -->|Read History| Summary + Summary -->|Update Guidelines| VectorStore + + style Agent fill:#e1f5ff + style Retrieve fill:#fff4e1 + style Add fill:#ffe1f5 + style Summary fill:#e1ffe1 + style VectorStore fill:#f0f0f0 +``` + +### Workflow Steps + +**1. retrieve_tool_memory** (Before tool execution) +- Agent queries usage guidelines from Vector Store +- Returns best practices and historical patterns + +**2. Tool Call Execution** +- Agent executes tool with recommended parameters +- Gets result (success/failure, output, costs) + +**3. add_tool_call_result** (After tool execution) +- Evaluates the tool call result +- Stores evaluated result to Vector Store + +**4. summary_tool_memory** (Periodic) +- Analyzes accumulated call history +- Generates comprehensive usage guidelines +- Updates Vector Store with new insights + +## Basic Usage + +Here's how to use Tool Memory in your application: + +### Step 1: Set Up Your Environment + +```python +import requests + +# API configuration +BASE_URL = "http://0.0.0.0:8002/" +WORKSPACE_ID = "your_workspace_id" +``` + +### Step 2: Record Tool Call Results + +After your agent executes a tool, record the invocation: + +```python +# Example: Record a web search tool call +tool_call_results = [ + { + "create_time": "2025-10-15 14:30:00", + "tool_name": "web_search", + "input": { + "query": "Python asyncio tutorial", + "max_results": 10, + "language": "en" + }, + "output": "Found 10 relevant results including official docs and tutorials", + "token_cost": 150, + "success": True, + "time_cost": 2.3 + } +] + +# Add the tool call result +response = requests.post( + url=f"{BASE_URL}add_tool_call_result", + json={ + "workspace_id": WORKSPACE_ID, + "tool_name": "web_search", + "tool_call_results": tool_call_results + } +) +``` + +### Step 3: Retrieve Tool Usage Guidelines + +Before using a tool, retrieve its usage guidelines: + +```python +# Retrieve memory for specific tools +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search,file_reader" # Comma-separated + } +) + +memory_list = response.json().get("metadata", {}).get("memory_list", []) +for memory in memory_list: + print(f"Tool: {memory['when_to_use']}") + print(f"Guidelines: {memory['content']}") + print(f"Total Calls: {len(memory['tool_call_results'])}") +``` + +### Step 4: Generate Comprehensive Usage Guidelines + +After accumulating sufficient call history, generate synthesized guidelines: + +```python +# Summarize tool usage patterns +response = requests.post( + url=f"{BASE_URL}summary_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search" + } +) + +if response.json().get("success"): + print("Successfully generated usage guidelines") +``` + +## Complete Example Workflow + +Here's a complete example demonstrating the tool memory lifecycle: + +```python +import requests +from datetime import datetime + +BASE_URL = "http://0.0.0.0:8002/" +WORKSPACE_ID = "demo_workspace" + +def record_tool_usage(tool_name, input_params, output, success, time_cost, token_cost): + """Record a single tool invocation""" + tool_call_result = { + "create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"), + "tool_name": tool_name, + "input": input_params, + "output": output, + "token_cost": token_cost, + "success": success, + "time_cost": time_cost + } + + response = requests.post( + url=f"{BASE_URL}add_tool_call_result", + json={ + "workspace_id": WORKSPACE_ID, + "tool_name": tool_name, + "tool_call_results": [tool_call_result] + } + ) + return response.json() + +def get_tool_guidelines(tool_name): + """Retrieve usage guidelines for a tool""" + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } + ) + + result = response.json() + if result.get("success"): + memory_list = result.get("metadata", {}).get("memory_list", []) + if memory_list: + return memory_list[0] + return None + +def generate_guidelines(tool_name): + """Generate comprehensive usage guidelines""" + response = requests.post( + url=f"{BASE_URL}summary_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } + ) + return response.json() + +# Example usage +if __name__ == "__main__": + tool_name = "web_search" + + # 1. Record multiple tool invocations + print("Recording tool invocations...") + for i in range(5): + record_tool_usage( + tool_name=tool_name, + input_params={"query": f"test query {i}", "max_results": 10}, + output=f"Found results for query {i}", + success=True, + time_cost=2.0 + i * 0.5, + token_cost=100 + i * 20 + ) + + # 2. Generate usage guidelines + print("\nGenerating usage guidelines...") + generate_guidelines(tool_name) + + # 3. Retrieve and display guidelines + print("\nRetrieving guidelines...") + memory = get_tool_guidelines(tool_name) + if memory: + print(f"\nTool: {memory['when_to_use']}") + print(f"Guidelines:\n{memory['content']}") + print(f"\nTotal invocations: {len(memory['tool_call_results'])}") +``` + +## Use Cases + +### Use Case 1: Learning Optimal Parameters + +```mermaid +sequenceDiagram + participant Agent + participant ToolMemory + participant Tool + + Agent->>ToolMemory: retrieve_tool_memory("api_caller") + ToolMemory-->>Agent: Guidelines: Use timeout=30s, retry=3 + Agent->>Tool: Call with recommended params + Tool-->>Agent: Success (time: 2.5s) + Agent->>ToolMemory: add_tool_call_result(success=True) + ToolMemory->>ToolMemory: Update statistics +``` + +**Scenario**: An agent needs to call an external API. Tool Memory has learned from 50+ previous calls that: +- Setting `timeout=30s` achieves 95% success rate +- Using `retry=3` handles transient failures effectively +- Requests with `max_results > 100` often timeout + +The agent retrieves these guidelines before making the call, leading to higher success rates. + +### Use Case 2: Avoiding Common Pitfalls + +```mermaid +sequenceDiagram + participant Agent + participant ToolMemory + participant FileReader + + Agent->>ToolMemory: retrieve_tool_memory("file_reader") + ToolMemory-->>Agent: Warning: Large files (>10MB) cause timeouts + Agent->>Agent: Check file size first + Agent->>FileReader: Read file with streaming mode + FileReader-->>Agent: Success + Agent->>ToolMemory: add_tool_call_result(success=True) +``` + +**Scenario**: Tool Memory has recorded that the `file_reader` tool fails when: +- File paths contain special characters without escaping +- Files larger than 10MB are read without streaming mode +- Binary files are opened in text mode + +The agent retrieves these warnings and adjusts its approach accordingly. + +### Use Case 3: Performance Optimization + +```python +# Before using Tool Memory +average_time_cost = 5.2s +success_rate = 75% + +# After learning from Tool Memory +# - Use batch processing for multiple queries +# - Set appropriate timeout values +# - Cache frequently accessed data + +average_time_cost = 2.8s # 46% improvement +success_rate = 92% # 17% improvement +``` + +**Scenario**: Tool Memory analyzes 100+ invocations of a `database_query` tool and discovers: +- Batch queries are 3x faster than individual queries +- Connection pooling reduces overhead by 40% +- Queries during peak hours (2-4 PM) have higher failure rates + +The synthesized guidelines help the agent optimize its database interactions. + +## Managing Tool Memories + +### Delete a Workspace + +```python +response = requests.post( + url=f"{BASE_URL}vector_store", + json={ + "workspace_id": WORKSPACE_ID, + "action": "delete" + } +) +``` + +### Dump Memories to Disk + +```python +response = requests.post( + url=f"{BASE_URL}vector_store", + json={ + "workspace_id": WORKSPACE_ID, + "action": "dump", + "path": "./" + } +) +``` + +### Load Memories from Disk + +```python +response = requests.post( + url=f"{BASE_URL}vector_store", + json={ + "workspace_id": WORKSPACE_ID, + "action": "load", + "path": "./" + } +) +``` + +## Configuration Parameters + +### ParseToolCallResultOp Parameters + +Configure in `default.yaml`: + +```yaml +op: + parse_tool_call_result_op: + backend: parse_tool_call_result_op + llm: default + params: + max_history_tool_call_cnt: 100 # Max historical calls to retain + evaluation_sleep_interval: 1.0 # Delay between evaluations (seconds) +``` + +- `max_history_tool_call_cnt`: Limits the number of historical tool call results stored per tool. Older results are removed when this limit is exceeded. +- `evaluation_sleep_interval`: Controls the delay between concurrent evaluations to avoid rate limiting. + +### SummaryToolMemoryOp Parameters + +```yaml +op: + summary_tool_memory_op: + backend: summary_tool_memory_op + llm: default + params: + recent_call_count: 20 # Number of recent calls to analyze + summary_sleep_interval: 1.0 # Delay between summaries (seconds) +``` + +- `recent_call_count`: Number of most recent tool calls to analyze when generating guidelines. +- `summary_sleep_interval`: Controls the delay between concurrent summarizations. + +## Best Practices + +1. **Regular Recording**: + - Record every tool invocation, including failures + - Include detailed input parameters and output + - Capture performance metrics (time_cost, token_cost) + +2. **Periodic Summarization**: + - Generate guidelines after accumulating 20-50 tool calls + - Re-summarize when usage patterns change significantly + - Update guidelines when new tool versions are deployed + +3. **Retrieval Strategy**: + - Always retrieve guidelines before using unfamiliar tools + - Cache retrieved guidelines for the duration of a task + - Re-retrieve after tool memory updates + +4. **Quality Maintenance**: + - Monitor success rates and average scores + - Investigate tools with declining performance + - Clean up outdated memories when tools are deprecated + +5. **Parameter Tuning**: + - Adjust `max_history_tool_call_cnt` based on tool usage frequency + - Increase `recent_call_count` for tools with diverse usage patterns + - Reduce `evaluation_sleep_interval` if rate limiting is not a concern + +## Integration with Agent Workflows + +```mermaid +graph TB + A[Agent Receives Task] --> B{Tool Required?} + B -->|Yes| C[Retrieve Tool Memory] + C --> D[Apply Guidelines] + D --> E[Execute Tool] + E --> F[Record Result] + F --> G{Sufficient History?} + G -->|Yes| H[Generate Summary] + G -->|No| I[Continue] + H --> I + B -->|No| I[Process Task] + I --> J[Task Complete] +``` + +Tool Memory seamlessly integrates into agent workflows: +1. Before tool execution: Retrieve usage guidelines +2. During execution: Apply recommended parameters +3. After execution: Record results with evaluation +4. Periodically: Generate updated guidelines + +For more detailed examples, see the implementation in `reme_ai/summary/tool/` directory of the ReMe project. + diff --git a/docs/tool_memory/tool_retrieve_ops.md b/docs/tool_memory/tool_retrieve_ops.md new file mode 100644 index 00000000..38ac37e1 --- /dev/null +++ b/docs/tool_memory/tool_retrieve_ops.md @@ -0,0 +1,548 @@ +# Tool Memory Retrieval Operations + +## RetrieveToolMemoryOp + +### Purpose + +Retrieves tool memories from the vector database based on tool names, providing usage patterns, best practices, and historical call data. + +### Functionality + +- Accepts comma-separated tool names as input +- Searches the vector store for exact tool name matches +- Validates that retrieved memories are of type "tool" +- Returns complete tool memories including usage guidelines and call history +- Provides detailed logging for debugging and monitoring + +### Processing Flow + +```mermaid +graph TB + A[Receive Tool Names] --> B[Validate Input] + B --> C[Split by Comma] + C --> D[Trim Whitespace] + D --> E[For Each Tool Name] + E --> F[Search Vector Store] + F --> G{Results Found?} + G -->|Yes| H[Get Top Result] + G -->|No| I[Log Warning: Not Found] + H --> J{Type = tool?} + J -->|Yes| K{Name Matches?} + J -->|No| L[Log Warning: Wrong Type] + K -->|Yes| M[Add to Results] + K -->|No| N[Log Warning: Name Mismatch] + I --> O[Continue Next Tool] + L --> O + N --> O + M --> O + O --> P{More Tools?} + P -->|Yes| E + P -->|No| Q{Any Matches?} + Q -->|Yes| R[Return Memory List] + Q -->|No| S[Return Empty] +``` + +1. **Input Validation**: + - Check if `tool_names` parameter is provided + - Return error if empty + - Log workspace_id and tool count + +2. **Tool Name Processing**: + - Split input by comma delimiter + - Strip whitespace from each tool name + - Filter out empty strings + - Log the list of tools to retrieve + +3. **Vector Store Search**: + - For each tool name, search with `top_k=1` + - Use tool name as the query (exact match preferred) + - Retrieve the top matching result + +4. **Result Validation**: + - Verify result is of type `ToolMemory` + - Check that `when_to_use` field exactly matches tool name + - Log match details (memory_id, total_calls) + - Warn if no match or mismatch found + +5. **Response Preparation**: + - Collect all matched tool memories + - Set success status based on matches found + - Return memory list in metadata + +### Parameters + +This operation has no configurable parameters. It uses the default vector store configuration. + +### Input Schema + +```yaml +input_schema: + tool_names: + type: string + description: "Comma-separated tool names (e.g., 'tool_name1,tool_name2')" + required: true +``` + +### Output Format + +The operation sets the following in `context.response`: + +```python +{ + "success": True, + "answer": "Successfully retrieved 2 tool memories", + "metadata": { + "memory_list": [ + { + "workspace_id": "demo_workspace", + "memory_id": "abc123def456", + "memory_type": "tool", + "when_to_use": "web_search", + "content": "Core Function: The web_search tool retrieves...", + "score": 0.85, + "time_created": "2025-10-15 10:00:00", + "time_modified": "2025-10-15 14:30:00", + "author": "qwen3-30b-a3b-instruct-2507", + "tool_call_results": [ + { + "create_time": "2025-10-15 14:30:00", + "tool_name": "web_search", + "input": {...}, + "output": "...", + "summary": "...", + "evaluation": "...", + "score": 1.0, + "success": True, + "time_cost": 2.3, + "token_cost": 150 + } + ], + "metadata": {} + } + ] + } +} +``` + +### Usage Example + +#### Basic Retrieval + +```python +import requests + +BASE_URL = "http://0.0.0.0:8002/" +WORKSPACE_ID = "demo_workspace" + +# Retrieve memory for a single tool +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search" + } +) + +result = response.json() +if result.get("success"): + memory_list = result.get("metadata", {}).get("memory_list", []) + for memory in memory_list: + print(f"Tool: {memory['when_to_use']}") + print(f"Guidelines:\n{memory['content']}") + print(f"Total Calls: {len(memory['tool_call_results'])}") +``` + +#### Multiple Tools Retrieval + +```python +# Retrieve memories for multiple tools at once +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search,file_reader,api_caller,database_query" + } +) + +result = response.json() +if result.get("success"): + memory_list = result.get("metadata", {}).get("memory_list", []) + print(f"Retrieved {len(memory_list)} tool memories") + + for memory in memory_list: + print(f"\n{'='*60}") + print(f"Tool: {memory['when_to_use']}") + print(f"{'='*60}") + print(memory['content']) +``` + +#### Extracting Specific Information + +```python +def get_tool_statistics(tool_name): + """Get statistical information for a specific tool""" + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } + ) + + result = response.json() + if not result.get("success"): + return None + + memory_list = result.get("metadata", {}).get("memory_list", []) + if not memory_list: + return None + + memory = memory_list[0] + tool_calls = memory['tool_call_results'] + + # Calculate statistics + total_calls = len(tool_calls) + successful_calls = sum(1 for call in tool_calls if call['success']) + avg_score = sum(call['score'] for call in tool_calls) / total_calls if total_calls > 0 else 0 + avg_time = sum(call['time_cost'] for call in tool_calls) / total_calls if total_calls > 0 else 0 + avg_tokens = sum(call['token_cost'] for call in tool_calls) / total_calls if total_calls > 0 else 0 + + return { + "tool_name": tool_name, + "total_calls": total_calls, + "success_rate": successful_calls / total_calls if total_calls > 0 else 0, + "avg_score": avg_score, + "avg_time_cost": avg_time, + "avg_token_cost": avg_tokens, + "guidelines": memory['content'] + } + +# Usage +stats = get_tool_statistics("web_search") +if stats: + print(f"Tool: {stats['tool_name']}") + print(f"Total Calls: {stats['total_calls']}") + print(f"Success Rate: {stats['success_rate']:.1%}") + print(f"Avg Score: {stats['avg_score']:.2f}") + print(f"Avg Time: {stats['avg_time_cost']:.2f}s") + print(f"Avg Tokens: {stats['avg_token_cost']:.0f}") +``` + +### Integration with Agent Workflows + +#### Pre-Execution Retrieval + +```python +def execute_tool_with_guidelines(tool_name, input_params): + """Execute a tool after retrieving its usage guidelines""" + + # 1. Retrieve tool memory + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } + ) + + # 2. Extract guidelines + guidelines = "" + if response.json().get("success"): + memory_list = response.json().get("metadata", {}).get("memory_list", []) + if memory_list: + guidelines = memory_list[0]['content'] + print(f"Guidelines for {tool_name}:") + print(guidelines) + + # 3. Adjust parameters based on guidelines + # (This would be done by an LLM or rule-based system) + adjusted_params = adjust_parameters(input_params, guidelines) + + # 4. Execute tool + result = execute_tool(tool_name, adjusted_params) + + # 5. Record result + record_tool_call(tool_name, adjusted_params, result) + + return result +``` + +#### Batch Retrieval for Agent Initialization + +```python +def initialize_agent_with_tool_memories(available_tools): + """Load all tool memories at agent initialization""" + + # Retrieve all tool memories at once + tool_names = ",".join(available_tools) + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_names + } + ) + + # Build tool memory cache + tool_memory_cache = {} + if response.json().get("success"): + memory_list = response.json().get("metadata", {}).get("memory_list", []) + for memory in memory_list: + tool_memory_cache[memory['when_to_use']] = { + "guidelines": memory['content'], + "total_calls": len(memory['tool_call_results']), + "last_modified": memory['time_modified'] + } + + return tool_memory_cache + +# Usage +available_tools = ["web_search", "file_reader", "api_caller"] +tool_cache = initialize_agent_with_tool_memories(available_tools) + +# Agent can now quickly access guidelines +if "web_search" in tool_cache: + print(tool_cache["web_search"]["guidelines"]) +``` + +### Error Handling + +#### Tool Not Found + +```python +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "nonexistent_tool" + } +) + +result = response.json() +if not result.get("success"): + print(f"Error: {result.get('answer')}") + # Output: "No matching tool memories found" +``` + +#### Empty Tool Names + +```python +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "" + } +) + +result = response.json() +if not result.get("success"): + print(f"Error: {result.get('answer')}") + # Output: "tool_names is required" +``` + +#### Partial Matches + +```python +# Request 3 tools, but only 2 exist +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search,file_reader,nonexistent_tool" + } +) + +result = response.json() +if result.get("success"): + memory_list = result.get("metadata", {}).get("memory_list", []) + print(f"Found {len(memory_list)} out of 3 requested tools") + # Output: "Found 2 out of 3 requested tools" +``` + +### Retrieval Workflow + +```mermaid +sequenceDiagram + participant Agent + participant RetrieveOp + participant VectorStore + + Agent->>RetrieveOp: retrieve_tool_memory(tool_names) + RetrieveOp->>RetrieveOp: Split and validate names + + loop For each tool name + RetrieveOp->>VectorStore: search(tool_name, top_k=1) + VectorStore-->>RetrieveOp: Return top result + RetrieveOp->>RetrieveOp: Validate type and name + alt Valid match + RetrieveOp->>RetrieveOp: Add to results + else No match + RetrieveOp->>RetrieveOp: Log warning + end + end + + RetrieveOp-->>Agent: Return memory_list + Agent->>Agent: Apply guidelines +``` + +### Use Cases + +#### Use Case 1: Pre-Execution Guidance + +**Scenario**: Before executing a tool, the agent retrieves usage guidelines to optimize parameters. + +```python +# Agent needs to search the web +tool_name = "web_search" +query = "machine learning basics" + +# Retrieve guidelines +response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } +) + +# Guidelines suggest: max_results=10-20, use filter_type="technical_docs" +# Agent adjusts parameters accordingly +optimized_params = { + "query": query, + "max_results": 15, # Within recommended range + "filter_type": "technical_docs", # As suggested + "language": "en" +} + +# Execute with optimized parameters +result = execute_tool(tool_name, optimized_params) +``` + +#### Use Case 2: Performance Monitoring + +**Scenario**: Monitor tool performance trends over time. + +```python +def monitor_tool_performance(tool_name): + """Monitor tool performance and detect degradation""" + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_name + } + ) + + if not response.json().get("success"): + return None + + memory = response.json()["metadata"]["memory_list"][0] + calls = memory["tool_call_results"] + + # Analyze recent vs historical performance + recent_calls = calls[-20:] # Last 20 calls + historical_calls = calls[:-20] if len(calls) > 20 else [] + + recent_success_rate = sum(1 for c in recent_calls if c['success']) / len(recent_calls) + historical_success_rate = (sum(1 for c in historical_calls if c['success']) / len(historical_calls) + if historical_calls else recent_success_rate) + + # Detect degradation + if recent_success_rate < historical_success_rate - 0.1: + print(f"Warning: {tool_name} performance degraded!") + print(f"Recent: {recent_success_rate:.1%}, Historical: {historical_success_rate:.1%}") + return "degraded" + + return "healthy" + +# Usage +status = monitor_tool_performance("web_search") +``` + +#### Use Case 3: Tool Selection + +**Scenario**: Choose the best tool for a task based on historical performance. + +```python +def select_best_tool(task_description, candidate_tools): + """Select the best tool based on historical performance""" + + # Retrieve memories for all candidate tools + tool_names = ",".join(candidate_tools) + response = requests.post( + url=f"{BASE_URL}retrieve_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": tool_names + } + ) + + if not response.json().get("success"): + return candidate_tools[0] # Default to first + + memory_list = response.json()["metadata"]["memory_list"] + + # Score each tool + tool_scores = {} + for memory in memory_list: + calls = memory["tool_call_results"] + if not calls: + continue + + # Calculate composite score + success_rate = sum(1 for c in calls if c['success']) / len(calls) + avg_score = sum(c['score'] for c in calls) / len(calls) + avg_time = sum(c['time_cost'] for c in calls) / len(calls) + + # Composite: prioritize success and quality, penalize slow tools + composite = (success_rate * 0.4 + avg_score * 0.4 + (1 / (1 + avg_time)) * 0.2) + tool_scores[memory['when_to_use']] = composite + + # Select best tool + if tool_scores: + best_tool = max(tool_scores, key=tool_scores.get) + print(f"Selected {best_tool} with score {tool_scores[best_tool]:.2f}") + return best_tool + + return candidate_tools[0] + +# Usage +best = select_best_tool( + "Search for technical documentation", + ["web_search", "doc_search", "api_search"] +) +``` + +## Best Practices + +1. **Retrieval Timing**: + - Retrieve guidelines before first use of a tool + - Cache retrieved memories for the duration of a task + - Re-retrieve after tool memory updates (post-summarization) + +2. **Batch Retrieval**: + - Retrieve multiple tool memories in a single request + - Use comma-separated tool names for efficiency + - Initialize agent with all available tool memories + +3. **Error Handling**: + - Always check `success` status in response + - Handle cases where tool memory doesn't exist + - Provide fallback behavior for missing guidelines + +4. **Guidelines Application**: + - Parse guidelines to extract parameter recommendations + - Use LLM to interpret guidelines in context + - Combine guidelines with task-specific requirements + +5. **Performance Optimization**: + - Cache retrieved memories to avoid repeated calls + - Invalidate cache after summarization updates + - Monitor retrieval latency and optimize if needed + +6. **Monitoring**: + - Log retrieval requests for debugging + - Track which tools are frequently retrieved + - Identify tools without memory (candidates for recording) + diff --git a/docs/tool_memory/tool_summary_ops.md b/docs/tool_memory/tool_summary_ops.md new file mode 100644 index 00000000..dd4992e9 --- /dev/null +++ b/docs/tool_memory/tool_summary_ops.md @@ -0,0 +1,465 @@ +# Tool Summary Operations + +## ParseToolCallResultOp + +### Purpose + +Evaluates individual tool invocations and adds them to the tool memory database with comprehensive assessments. + +### Functionality + +- Receives tool call results with input parameters, output, and metadata +- Uses LLM to evaluate each tool call based on success and parameter alignment +- Generates summary, evaluation, and score (0.0, 0.5, or 1.0) for each call +- Appends evaluated results to existing tool memory or creates new memory +- Maintains a sliding window of recent tool calls (configurable limit) + +### Processing Flow + +```mermaid +graph TB + A[Receive Tool Call Results] --> B[Validate Input] + B --> C[Search for Existing Memory] + C --> D{Memory Exists?} + D -->|Yes| E[Load Existing Memory] + D -->|No| F[Create New Memory] + E --> G[Concurrent Evaluation] + F --> G + G --> H[Evaluate Call #1] + G --> I[Evaluate Call #2] + G --> J[Evaluate Call #N] + H --> K[Generate Summary & Score] + I --> K + J --> K + K --> L[Append to Memory] + L --> M[Trim to Max History] + M --> N[Update Modified Time] + N --> O[Prepare for Vector Store] +``` + +1. **Input Validation**: + - Verify `tool_name` is provided + - Convert dict objects to `ToolCallResult` instances + - Check if `tool_call_results` list is not empty + +2. **Memory Lookup**: + - Search vector store for existing tool memory by tool name + - Verify exact match (memory type = "tool" and when_to_use = tool_name) + - Create new `ToolMemory` if no match found + +3. **Concurrent Evaluation**: + - Submit all tool call results for parallel evaluation + - Each evaluation uses LLM to analyze the call + - Generate structured evaluation with summary, assessment, and score + +4. **Memory Update**: + - Append evaluated results to tool memory + - Trim to `max_history_tool_call_cnt` if limit exceeded + - Update modification timestamp + - Prepare for vector store update + +### Parameters + +Configure in `default.yaml`: + +```yaml +op: + parse_tool_call_result_op: + backend: parse_tool_call_result_op + llm: default + params: + max_history_tool_call_cnt: 100 + evaluation_sleep_interval: 1.0 +``` + +- `max_history_tool_call_cnt` (integer, default: `100`): + - Maximum number of historical tool call results to retain per tool + - When exceeded, oldest results are removed (FIFO) + - Balances memory size with historical context + - Recommended: 50-200 depending on tool usage frequency + +- `evaluation_sleep_interval` (float, default: `1.0`): + - Delay in seconds between concurrent evaluations + - Prevents rate limiting when evaluating multiple calls + - Set to 0 for maximum speed (if no rate limits) + - Increase if encountering API throttling + +### Evaluation Criteria + +The LLM evaluates each tool call based on two dimensions: + +1. **Success Evaluation**: + - Check if the success flag indicates successful execution + - Verify output contains no error messages + - Assess if time and token costs are reasonable + - Identify any error indicators in the output + +2. **Parameter Alignment Evaluation**: + - Evaluate if output matches expected behavior given input + - Consider if input parameters are appropriate for the tool + - Check for parameter mismatches or unexpected behaviors + - Verify output is consistent with tool's intended purpose + +### Scoring Guidelines + +- **1.0 (Success)**: Tool executed successfully with good parameter alignment + - Example: Query returned relevant results, parameters were appropriate + +- **0.5 (Partial Success)**: Tool executed but with issues + - Example: Query succeeded but parameters were suboptimal (e.g., too generic) + - Example: Results returned but with warnings about invalid parameters + +- **0.0 (Failure)**: Tool execution failed or severe parameter misalignment + - Example: Timeout due to excessive max_results parameter + - Example: Error due to invalid parameter format + +### Output Format + +The operation sets the following in `context.response.metadata`: + +```python +{ + "deleted_memory_ids": ["memory_id_if_updating"], + "memory_list": [ + { + "memory_id": "abc123", + "when_to_use": "web_search", + "tool_call_results": [ + { + "create_time": "2025-10-15 14:30:00", + "tool_name": "web_search", + "input": {...}, + "output": "...", + "summary": "Successfully retrieved 10 relevant results", + "evaluation": "Good parameter alignment...", + "score": 1.0, + "success": True, + "time_cost": 2.3, + "token_cost": 150 + } + ] + } + ] +} +``` + +### Usage Example + +```python +import requests +from datetime import datetime + +BASE_URL = "http://0.0.0.0:8002/" +WORKSPACE_ID = "demo_workspace" + +# Prepare tool call results +tool_call_results = [ + { + "create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"), + "tool_name": "web_search", + "input": { + "query": "Python asyncio tutorial", + "max_results": 10, + "language": "en", + "filter_type": "technical_docs" + }, + "output": "Found 10 relevant results including official documentation and tutorials", + "token_cost": 150, + "success": True, + "time_cost": 2.3 + }, + { + "create_time": datetime.now().strftime("%Y-%m-%d %H:%M:%S"), + "tool_name": "web_search", + "input": { + "query": "test", # Too generic + "max_results": 100, # Too many + "language": "unknown" # Invalid + }, + "output": "Warning: language 'unknown' not supported. Query too generic, limited results.", + "token_cost": 80, + "success": True, + "time_cost": 3.5 + } +] + +# Add tool call results +response = requests.post( + url=f"{BASE_URL}add_tool_call_result", + json={ + "workspace_id": WORKSPACE_ID, + "tool_name": "web_search", + "tool_call_results": tool_call_results + } +) + +result = response.json() +print(f"Success: {result.get('success')}") +print(f"Answer: {result.get('answer')}") + +# Check evaluated results +memory_list = result.get("metadata", {}).get("memory_list", []) +if memory_list: + tool_memory = memory_list[0] + for call_result in tool_memory["tool_call_results"]: + print(f"\nCall Summary: {call_result['summary']}") + print(f"Evaluation: {call_result['evaluation']}") + print(f"Score: {call_result['score']}") +``` + +## SummaryToolMemoryOp + +### Purpose + +Analyzes accumulated tool call history and generates comprehensive usage patterns, best practices, and recommendations. + +### Functionality + +- Retrieves existing tool memories from the vector store +- Analyzes the most recent N tool calls (configurable) +- Calculates statistical metrics (success rate, average scores, costs) +- Uses LLM to synthesize actionable usage guidelines +- Updates tool memory content with generated insights + +### Processing Flow + +```mermaid +graph TB + A[Receive Tool Names] --> B[Split by Comma] + B --> C[For Each Tool Name] + C --> D[Search Vector Store] + D --> E{Exact Match?} + E -->|Yes| F[Load Tool Memory] + E -->|No| G[Log Warning] + F --> H[Extract Recent N Calls] + H --> I[Calculate Statistics] + I --> J[Format Call Summaries] + J --> K[Format Statistics] + K --> L[Concurrent Summarization] + L --> M[LLM Analysis #1] + L --> N[LLM Analysis #2] + L --> O[LLM Analysis #N] + M --> P[Generate Guidelines] + N --> P + O --> P + P --> Q[Update Memory Content] + Q --> R[Update Modified Time] + R --> S[Prepare for Vector Store] +``` + +1. **Tool Name Processing**: + - Split comma-separated tool names + - Trim whitespace from each name + - Log the list of tools to process + +2. **Memory Retrieval**: + - Search vector store for each tool name + - Verify exact match (memory type and when_to_use) + - Skip tools without existing memory + +3. **Data Preparation**: + - Extract the most recent N tool call results + - Calculate statistical metrics: + - Total calls vs recent calls analyzed + - Success rate (overall and recent) + - Average score (overall and recent) + - Average time cost + - Average token cost + - Format call summaries as markdown + - Format statistics as markdown + +4. **Concurrent Summarization**: + - Submit all tools for parallel summarization + - Each summarization uses LLM to analyze patterns + - Generate structured usage guidelines + +5. **Memory Update**: + - Update tool memory content with new guidelines + - Update modification timestamp + - Prepare for vector store update + +### Parameters + +Configure in `default.yaml`: + +```yaml +op: + summary_tool_memory_op: + backend: summary_tool_memory_op + llm: default + params: + recent_call_count: 20 + summary_sleep_interval: 1.0 +``` + +- `recent_call_count` (integer, default: `20`): + - Number of most recent tool calls to analyze + - Focuses on recent usage patterns + - Recommended: 10-50 depending on tool usage frequency + - Higher values provide more context but may dilute recent patterns + +- `summary_sleep_interval` (float, default: `1.0`): + - Delay in seconds between concurrent summarizations + - Prevents rate limiting when summarizing multiple tools + - Set to 0 for maximum speed (if no rate limits) + - Increase if encountering API throttling + +### Statistical Metrics + +The operation calculates the following metrics: + +```python +{ + "total_calls": 100, # Total number of calls in history + "recent_calls": 20, # Number of recent calls analyzed + "success_rate": 0.85, # Overall success rate (85%) + "recent_success_rate": 0.90, # Recent success rate (90%) + "avg_score": 0.78, # Average evaluation score + "recent_avg_score": 0.82, # Recent average score + "avg_time_cost": 2.45, # Average time in seconds + "avg_token_cost": 125.3 # Average token consumption +} +``` + +### Generated Guidelines Structure + +The LLM generates guidelines following this structure: + +1. **Core Function**: What the tool does and when to use it +2. **Success Patterns**: Parameter patterns and scenarios that work well +3. **Common Issues**: Main pitfalls to avoid and why they fail +4. **Best Practices**: 2-3 actionable recommendations + +### Usage Example + +```python +import requests + +BASE_URL = "http://0.0.0.0:8002/" +WORKSPACE_ID = "demo_workspace" + +# After accumulating tool call history, generate guidelines +response = requests.post( + url=f"{BASE_URL}summary_tool_memory", + json={ + "workspace_id": WORKSPACE_ID, + "tool_names": "web_search,file_reader,api_caller" # Multiple tools + } +) + +result = response.json() +print(f"Success: {result.get('success')}") +print(f"Answer: {result.get('answer')}") + +# Display generated guidelines +memory_list = result.get("metadata", {}).get("memory_list", []) +for memory in memory_list: + print(f"\n{'='*60}") + print(f"Tool: {memory['when_to_use']}") + print(f"{'='*60}") + print(memory['content']) + + # Display statistics + stats = memory.get('metadata', {}).get('statistics', {}) + print(f"\nStatistics:") + print(f" Total Calls: {stats.get('total_calls', 0)}") + print(f" Success Rate: {stats.get('success_rate', 0):.1%}") + print(f" Avg Score: {stats.get('avg_score', 0):.2f}") +``` + +### Example Generated Guidelines + +``` +Core Function: +The web_search tool retrieves information from the internet based on query parameters. +Use it when you need up-to-date information, documentation, or external data. + +Success Patterns: +- Specific queries (e.g., "Python asyncio tutorial") achieve 95% success rate +- Setting max_results=10-20 balances quality and performance +- Using language="en" and filter_type="technical_docs" improves relevance +- Average successful call: 2.3s, 150 tokens + +Common Issues: +- Generic queries (e.g., "test") return poor results (score: 0.5) +- max_results > 50 often leads to timeouts (avg: 8.2s vs 2.3s) +- Invalid language codes default to English but add latency +- Missing filter_type returns mixed-quality results + +Best Practices: +1. Use specific, descriptive queries with clear intent +2. Set max_results=10-20 for optimal balance +3. Always specify language and filter_type for technical searches +``` + +### When to Run Summarization + +- **Initial Setup**: After accumulating 20-30 tool calls +- **Regular Updates**: Every 50-100 new calls +- **Pattern Changes**: When success rate changes significantly +- **Tool Updates**: After tool version changes or parameter updates +- **Performance Issues**: When investigating declining performance + +### Integration with Retrieval + +```mermaid +sequenceDiagram + participant Agent + participant AddOp as add_tool_call_result + participant SummaryOp as summary_tool_memory + participant RetrieveOp as retrieve_tool_memory + + loop Every Tool Call + Agent->>AddOp: Record tool call result + AddOp->>AddOp: Evaluate and store + end + + Note over Agent,SummaryOp: After 20+ calls + Agent->>SummaryOp: Generate guidelines + SummaryOp->>SummaryOp: Analyze patterns + SummaryOp->>SummaryOp: Update content + + Note over Agent,RetrieveOp: Before next use + Agent->>RetrieveOp: Get tool memory + RetrieveOp-->>Agent: Return guidelines + history + Agent->>Agent: Apply recommendations +``` + +The summarization operation works in conjunction with retrieval: +1. `add_tool_call_result` continuously records invocations +2. `summary_tool_memory` periodically generates guidelines +3. `retrieve_tool_memory` provides guidelines before tool use +4. Agent applies recommendations to improve success rates + +## Best Practices + +1. **Recording Strategy**: + - Record every tool invocation, including failures + - Include detailed input parameters and complete output + - Capture accurate performance metrics + - Add relevant metadata for context + +2. **Evaluation Quality**: + - Ensure LLM has sufficient context for evaluation + - Monitor score distribution (should not be all 1.0 or 0.0) + - Review evaluations periodically for quality + - Adjust evaluation prompts if needed + +3. **Summarization Timing**: + - Wait for 20-30 calls before first summarization + - Re-summarize after significant new data (50+ calls) + - Update when usage patterns change + - Regenerate after tool updates + +4. **Parameter Tuning**: + - Adjust `max_history_tool_call_cnt` based on tool usage frequency + - Increase `recent_call_count` for tools with diverse patterns + - Balance `sleep_interval` between speed and rate limits + - Monitor vector store size and adjust retention limits + +5. **Quality Maintenance**: + - Review generated guidelines for accuracy + - Validate statistical metrics match expectations + - Clean up deprecated tools from vector store + - Archive historical data before major changes +