This commit is contained in:
鸣山 2025-07-25 12:36:36 +08:00
commit 0b95a3bd70
16 changed files with 1428 additions and 584 deletions

4
.gitignore vendored
View file

@ -26,4 +26,6 @@ build/*
*.egg-info/*
cookbook/appworld/data/*
cookbook/appworld/experiments/*
cookbook/appworld/exp_result/*
cookbook/appworld/exp_result/*
file_vector_store/*
cookbook/appworld/file_vector_store/*

373
README.md
View file

@ -15,15 +15,17 @@
<strong>A comprehensive framework for AI agent experience generation and reuse</strong><br>
<em>Empowering agents to learn from the past and excel in the future</em>
</p>
---
## 📰 What's New
- **[2025-08]** 🎉 ExperienceMaker v0.1.0 is now available on [PyPI](https://pypi.org/project/experiencemaker/)!
- **[2025-07]** 📚 Complete documentation and quick start guides released
- **[2025-07]** 🚀 Multi-backend vector store support (Elasticsearch & ChromaDB)
- **[2025-06]** 🚀 Multi-backend vector store support (Elasticsearch & ChromaDB)
---
## 📰 What's Next
## 🚀 What's Next
- **Pre-built Experience Libraries**: Domain repositories (Finance/Coding/Education/Research) + community marketplace
- **Rich Experience Formats**: Executable code/tool configs/pipeline templates/workflows
- **Experience Validation**: Quality analysis + cross-task effectiveness + auto-refinement
@ -41,7 +43,7 @@ By automatically extracting, storing, and intelligently reusing experiences from
Traditional AI agents start from scratch with every new task, wasting valuable learning opportunities.
ExperienceMaker changes this paradigm by:
- **🧠 Learning from History**: Automatically extract actionable insights from both successful and failed attempts
- **🔄 Intelligent Reuse**: Apply relevant past experiences to solve new, similar challenges more effectively
- **🔄 Intelligent Reuse**: Apply relevant experiences to solve new, similar challenges more effectively
- **📈 Continuous Improvement**: Build a growing knowledge base that makes agents progressively smarter
### ✨ Core Capabilities
@ -50,12 +52,12 @@ ExperienceMaker changes this paradigm by:
- **Success Pattern Recognition**: Identify what works and understand the underlying principles
- **Failure Analysis**: Learn from mistakes to avoid repeating them in future tasks
- **Comparative Insights**: Understand the critical differences between successful and failed approaches
- **Multi-step Trajectory Processing**: Break down complex tasks into learnable, actionable segments
- **Multistep Trajectory Processing**: Break down complex tasks into learnable, actionable segments
#### 🎯 **Smart Experience Retriever**
- **Semantic Search**: Find relevant experiences using advanced embedding models and semantic understanding
- **Context-Aware Ranking**: Prioritize the most applicable experiences for current task contexts
- **Dynamic Rewriting**: Intelligently adapt past experiences to fit new situations and requirements
- **Dynamic Rewriting**: Intelligently adapt experiences to fit new situations and requirements
- **Multi-modal Support**: Handle various input types including query, messages
#### 🗄️ **Scalable Experience Management**
@ -82,7 +84,12 @@ ExperienceMaker follows a modular, production-ready architecture designed for sc
- **🗄️ Vector Store API**: Database management and workspace operations with full CRUD support
#### ⚙️ **Processing Pipeline**
Our atomic operations can be seamlessly composed into powerful processing pipelines: custom1_op->custom2_op...
Our atomic operations can be seamlessly composed into powerful processing pipelines:
```
custom1_op->custom2_op...
```
#### 🔌 **Extensible Components**
- **LLM Integration**: OpenAI-compatible APIs with flexible model switching and provider support
@ -110,7 +117,7 @@ pip install .
## ⚙️ Environment Setup
Create a `.env` file in your project directory:
Create a `.env` file in your project root directory:
```bash
# Required: LLM API configuration
@ -135,13 +142,16 @@ experiencemaker \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file
```
💡 **Pro Tip**: Check out our [Advanced Guide](./doc/advanced_guide.md) for detailed configuration topics including custom pipelines, operation parameters, and advanced configuration methods.
💡 **Pro Tip**: Check out our [Configuration Guide](./doc/configuration_guide.md) for detailed configuration topics
including custom pipelines, operation parameters, and advanced configuration methods.
The service will start on `http://localhost:8001`
### 🔍 Production Setup with Elasticsearch Backend
```bash
experiencemaker \
http_service.port=8001 \
llm.default.model_name=qwen3-32b \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=elasticsearch
@ -158,79 +168,309 @@ curl -fsSL https://elastic.co/start-local | sh
## 📝 Your First ExperienceMaker Script
Here's how to get started!
- The `load_dotenv()` function loads environment variables from your `.env` file, or you can manually export them.
- The `base_url` points to your ExperienceMaker service.
- The `workspace_id` serves as your experience storage namespace. Experiences in different workspaces remain completely
isolated and cannot access each other.
Note the `workspace_id` serves as your experience storage namespace. Experiences in different workspaces remain completely isolated and cannot access each other.
### 📊 Call Summarizer Examples
Transform conversation trajectories into valuable experiences using batch summarization. Each trajectory contains:
- **Message**: Complete conversation history between user and agent
- **Score**: Performance rating (0-1 scale, where 0=failure, 1=success)
The summarizer analyzes these trajectories to extract actionable insights and patterns for future interactions.
<details open>
<summary><b>Python</b></summary>
```python
import requests
from dotenv import load_dotenv
load_dotenv()
base_url = "http://0.0.0.0:8001/"
workspace_id = "test_workspace"
```
### 📊 Call Summarizer Examples
Batch summarize the trajectory list, where each trajectory consists of a message and a score.
- The message is the conversation history.
- The score represents the rating between 0 and 1, with 0 typically indicating failure and 1 indicating success.
```python
response = requests.post(url=base_url + "summarizer", json={
"workspace_id": workspace_id,
response = requests.post(url="http://0.0.0.0:8001/summarizer", json={
"workspace_id": "test_workspace",
"traj_list": [
{"messages": messages, "score": 1.0}
{"messages": [{"role": "user", "content": "hello world"}], "score": 1.0}
]
})
response = response.json()
experience_list = response["experience_list"]
experience_list = response.json()["experience_list"]
for experience in experience_list:
print(experience)
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/summarizer" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"traj_list": [
{
"messages": [{"role": "user", "content": "hello world"}],
"score": 1.0
}
]
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function callSummarizer() {
try {
const response = await fetch('http://0.0.0.0:8001/summarizer', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
traj_list: [
{
messages: [{ role: "user", content: "hello world" }],
score: 1.0
}
]
})
});
const data = await response.json();
const experienceList = data.experience_list;
experienceList.forEach(experience => {
console.log(experience);
});
} catch (error) {
console.error('Error:', error);
}
}
callSummarizer();
```
</details>
### 🔍 Call Retriever Examples
Retrieve the top_k={top_k} experiences related to {query} in workspace=test_workspace, and finally accept the assembled context.
Alternatively, you can also accept the raw experience_list parameter and assemble the context yourself.
Intelligently search and retrieve the most relevant experiences from your workspace to enhance decision-making. The retriever:
- **Finds** the top-k most similar experiences based on semantic similarity to your query
- **Returns** pre-assembled context ready for immediate use, or raw experience data for custom processing
- **Leverages** your workspace's accumulated knowledge to provide contextually relevant insights
<details open>
<summary><b>Python</b></summary>
```python
response = requests.post(url=base_url + "retriever", json={
"workspace_id": workspace_id,
"query": query,
import requests
response = requests.post(url="http://0.0.0.0:8001/retriever", json={
"workspace_id": "test_workspace",
"query": "what is the meaning of life?",
"top_k": 1,
})
response = response.json()
experience_merged: str = response["experience_merged"]
experience_merged: str = response.json()["experience_merged"]
print(f"experience_merged={experience_merged}")
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/retriever" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"query": "what is the meaning of life?",
"top_k": 1
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function callRetriever() {
try {
const response = await fetch('http://0.0.0.0:8001/retriever', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
query: "what is the meaning of life?",
top_k: 1
})
});
const data = await response.json();
const experienceMerged = data.experience_merged;
console.log(`experience_merged=${experienceMerged}`);
} catch (error) {
console.error('Error:', error);
}
}
callRetriever();
```
</details>
### 💾 Dump Experiences From Vector Store
Dump the experience with workspace_id from the vector store into the {path}/{workspace_id}.jsonl file.
Export and backup your valuable experience data for archival, analysis, or migration purposes. This operation:
- **Extracts** all experiences from the specified workspace in the vector store
- **Saves** them to a structured JSONL file at `{path}/{workspace_id}.jsonl`
- **Preserves** complete experience metadata and embeddings for future restoration
<details open>
<summary><b>Python</b></summary>
```python
response = requests.post(url=base_url + "vector_store", json={
"workspace_id": workspace_id,
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "dump",
"path": "./",
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "dump",
"path": "./"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function dumpExperiences() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "dump",
path: "./"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
dumpExperiences();
```
</details>
### 📥 Load Experiences To Vector Store
Load the {path}/{workspace_id}.jsonl file into the vector store, workspace_id={workspace_id}.
Import and restore previously exported experience data to populate your workspace with existing knowledge. This operation:
- **Reads** experience data from the JSONL file located at `{path}/{workspace_id}.jsonl`
- **Reconstructs** the vector embeddings and indexes them in the specified workspace
- **Enables** immediate access to imported experiences for retrieval and decision-making
<details open>
<summary><b>Python</b></summary>
```python
response = requests.post(url=base_url + "vector_store", json={
"workspace_id": workspace_id,
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "load",
"path": "./",
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "load",
"path": "./"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function loadExperiences() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "load",
path: "./"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
loadExperiences();
```
</details>
💡 **Need More Advanced Operations?** For additional workspace management features(e.g. delete_workspace,
copy_workspace), advanced configuration options, and troubleshooting guidance, check out our
comprehensive [Quick Start Guide](./cookbook/simple_demo/quick_start.md).
🎭 **Want to See It in Action?** We've prepared a [simple react agent](./cookbook/simple_demo/simple_demo.py) that demonstrates how to enhance agent capabilities by integrating summarizer and retriever components, achieving significantly better performance.
@ -259,15 +499,60 @@ Coming Soon! Stay tuned for comprehensive evaluation results.
## 🏪 Ready-made Experience Store
Pre-built experience collections for common domains and use cases are coming soon. This will include ready-to-use experiences for web automation, data processing, API interactions, and more.
ExperienceMaker provides pre-built experience libraries to jumpstart your agent's capabilities.
You can directly load these curated experiences into your workspace and start benefiting from accumulated knowledge
immediately.
### 📦 Available Experience Libraries
- **`appworld_v1.jsonl`**: Comprehensive experiences from Appworld agent interactions, covering complex task planning
and execution patterns
- **`bfcl_v1.jsonl`**: Function calling experiences from Berkeley Function-Calling Leaderboard tasks
### 🚀 Quick Start with Pre-built Experiences
Here's how to load and use the Appworld experience library:
#### Step 1: Load Pre-built Experiences
```python
import requests
# Load Appworld experiences into your workspace
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "appworld_v1",
"action": "load",
"path": "./experience_library/",
})
print(f"loading result result={response.json()}")
```
#### Step 2: Retrieve Relevant Experiences
Now you can query the loaded experiences to get contextual guidance for your tasks:
```python
import requests
# Query for app interaction experiences
response = requests.post(url="http://0.0.0.0:8001/retriever", json={
"workspace_id": "appworld_v1",
"query": "How to navigate to settings and update user profile information?",
"top_k": 1,
})
experience_merged = response.json()["experience_merged"]
print(f"Retrieved experiences: {experience_merged}")
```
---
## 📚 Additional Resources
- **[Quick Start](./cookbook/simple_demo/quick_start.md)**: This guide will help you get started with ExperienceMaker quickly using practical examples.
- **[Vector Store Setup](./doc/vector_store_setup.md)**: Complete production deployment guide
- **[Configuration Guide](./doc/configuration_guide.md)**: Describes all available command-line parameters for ExperienceMaker Service
- **[Advanced Guide](./doc/advanced_guide.md)**: Custom pipelines, operation parameters, and advanced configuration methods
- **[Operations Documentation](./doc/operations_documentation.md)**: Comprehensive operations configuration reference
- **[Example Collection](./cookbook)**: Practical examples and use cases
- **[Future RoadMap](./doc/future_roadmap.md)**: Our vision and upcoming features

View file

@ -0,0 +1,559 @@
# ExperienceMaker Quick Start Guide
This guide will help you get started with ExperienceMaker quickly using practical examples.
## 🚀 What You'll Learn
- How to set up ExperienceMaker service
- Run an agent and generate experiences
- Retrieve and apply experiences to new tasks
- Build experience-enhanced agents
## 📋 Prerequisites
- Python 3.12+
- LLM API access (OpenAI or compatible)
- Embedding model API access
## 🛠️ Installation
### Option 1: Install from PyPI (Recommended)
```bash
pip install experiencemaker
```
### Option 2: Install from Source
```bash
git clone https://github.com/modelscope/ExperienceMaker.git
cd ExperienceMaker
pip install .
```
## ⚙️ Environment Setup
Create a `.env` file in your project directory:
```bash
# Required: LLM API configuration
LLM_API_KEY="sk-xxx"
LLM_BASE_URL="https://xxx.com/v1"
# Required: Embedding model configuration
EMBEDDING_MODEL_API_KEY="sk-xxx"
EMBEDDING_MODEL_BASE_URL="https://xxx.com/v1"
# Optional: Elasticsearch configuration (if using Elasticsearch backend)
```
## 🚀 Start the Service
For testing, use the `local_file` backend:
```bash
experiencemaker \
http_service.port=8001 \
llm.default.model_name=qwen3-32b \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file
```
The service will start on `http://localhost:8001`
### Elasticsearch Backend
```bash
experiencemaker \
http_service.port=8001 \
llm.default.model_name=qwen3-32b \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=elasticsearch
```
**Setup Elasticsearch:**
```bash
export ES_HOSTS="http://localhost:9200"
# Quick setup using Elastic's official script
curl -fsSL https://elastic.co/start-local | sh
```
📖 **Need Help?** Refer to [Vector Store Setup](../../doc/vector_store_setup.md) for comprehensive deployment guidance.
## 📝 Your First ExperienceMaker Script
Here's how to get started!
Note the `workspace_id` serves as your experience storage namespace. Experiences in different workspaces remain
completely isolated and cannot access each other.
### 📊 Call Summarizer Examples
Transform conversation trajectories into valuable experiences using batch summarization. Each trajectory contains:
- **Message**: Complete conversation history between user and agent
- **Score**: Performance rating (0-1 scale, where 0=failure, 1=success)
The summarizer analyzes these trajectories to extract actionable insights and patterns for future interactions.
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/summarizer", json={
"workspace_id": "test_workspace",
"traj_list": [
{"messages": [{"role": "user", "content": "hello world"}], "score": 1.0}
]
})
experience_list = response.json()["experience_list"]
for experience in experience_list:
print(experience)
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/summarizer" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"traj_list": [
{
"messages": [{"role": "user", "content": "hello world"}],
"score": 1.0
}
]
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function callSummarizer() {
try {
const response = await fetch('http://0.0.0.0:8001/summarizer', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
traj_list: [
{
messages: [{ role: "user", content: "hello world" }],
score: 1.0
}
]
})
});
const data = await response.json();
const experienceList = data.experience_list;
experienceList.forEach(experience => {
console.log(experience);
});
} catch (error) {
console.error('Error:', error);
}
}
callSummarizer();
```
</details>
### 🔍 Call Retriever Examples
Intelligently search and retrieve the most relevant experiences from your workspace to enhance decision-making. The retriever:
- **Finds** the top-k most similar experiences based on semantic similarity to your query
- **Returns** pre-assembled context ready for immediate use, or raw experience data for custom processing
- **Leverages** your workspace's accumulated knowledge to provide contextually relevant insights
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/retriever", json={
"workspace_id": "test_workspace",
"query": "what is the meaning of life?",
"top_k": 1,
})
experience_merged: str = response.json()["experience_merged"]
print(f"experience_merged={experience_merged}")
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/retriever" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"query": "what is the meaning of life?",
"top_k": 1
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function callRetriever() {
try {
const response = await fetch('http://0.0.0.0:8001/retriever', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
query: "what is the meaning of life?",
top_k: 1
})
});
const data = await response.json();
const experienceMerged = data.experience_merged;
console.log(`experience_merged=${experienceMerged}`);
} catch (error) {
console.error('Error:', error);
}
}
callRetriever();
```
</details>
### 💾 Dump Experiences From Vector Store
Export and backup your valuable experience data for archival, analysis, or migration purposes. This operation:
- **Extracts** all experiences from the specified workspace in the vector store
- **Saves** them to a structured JSONL file at `{path}/{workspace_id}.jsonl`
- **Preserves** complete experience metadata and embeddings for future restoration
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "dump",
"path": "./",
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "dump",
"path": "./"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function dumpExperiences() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "dump",
path: "./"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
dumpExperiences();
```
</details>
### 📥 Load Experiences To Vector Store
Import and restore previously exported experience data to populate your workspace with existing knowledge. This operation:
- **Reads** experience data from the JSONL file located at `{path}/{workspace_id}.jsonl`
- **Reconstructs** the vector embeddings and indexes them in the specified workspace
- **Enables** immediate access to imported experiences for retrieval and decision-making
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "load",
"path": "./",
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "load",
"path": "./"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function loadExperiences() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "load",
path: "./"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
loadExperiences();
```
</details>
### 🗑️ Delete Workspace
Permanently remove a workspace and all its associated experience data when it's no longer needed. This operation:
- **Removes** all experiences, embeddings, and metadata from the specified workspace
- **Frees up** storage space and computational resources
- **Cannot be undone** - ensure you've backed up important data before deletion
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "delete"
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "delete"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function deleteWorkspace() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "delete"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
deleteWorkspace();
```
</details>
### 📋 Copy Workspace
Duplicate an existing workspace to create a new one with identical experience data, perfect for experimentation or branching. This operation:
- **Clones** all experiences and embeddings from the source workspace
- **Creates** a new independent workspace with the copied data
- **Preserves** original workspace while enabling safe testing and modifications in the copy
<details open>
<summary><b>Python</b></summary>
```python
import requests
response = requests.post(url="http://0.0.0.0:8001/vector_store", json={
"workspace_id": "test_workspace",
"action": "copy",
"src_workspace_id": "src_workspace"
})
print(response.json())
```
</details>
<details>
<summary><b>curl</b></summary>
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "test_workspace",
"action": "copy",
"src_workspace_id": "src_workspace"
}'
```
</details>
<details>
<summary><b>Node.js</b></summary>
```javascript
const fetch = require('node-fetch');
// or: import fetch from 'node-fetch';
async function copyWorkspace() {
try {
const response = await fetch('http://0.0.0.0:8001/vector_store', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
},
body: JSON.stringify({
workspace_id: "test_workspace",
action: "copy",
src_workspace_id: "src_workspace"
})
});
const data = await response.json();
console.log(data);
} catch (error) {
console.error('Error:', error);
}
}
copyWorkspace();
```
</details>
🎭 **Want to See It in Action?** We've prepared a [simple react agent](../../cookbook/simple_demo/simple_demo.py) that
demonstrates how to enhance agent capabilities by integrating summarizer and retriever components, achieving
significantly better performance.
## 🐛 Common Issues
### Service Won't Start
- Check if port 8001 is available
- Verify your API keys in `.env` file
- Ensure Python version is 3.12+
### No Experiences Retrieved
- Make sure you've run the summarizer first
- Check if workspace_id matches between operations
- Verify vector store backend is properly configured
### API Connection Errors
- Confirm LLM_BASE_URL and API keys are correct
- Test API access independently
- Check network connectivity
---
🎯 **You're all set!** You now have a working ExperienceMaker setup that can learn from interactions and improve over time.

View file

@ -64,8 +64,8 @@ def run_retriever(query: str):
def run_agent_with_experience(query_first: str, query_second: str, dump_experience: bool = True):
# messages = run_agent(query=query_second)
# run_summary(messages, dump_experience)
messages = run_agent(query=query_second)
run_summary(messages, dump_experience)
experience_merged = run_retriever(query_first)
messages = run_agent(query=f"{experience_merged}\n\nUser Question:\n{query_first}")
return messages
@ -103,7 +103,7 @@ if __name__ == "__main__":
query1 = "Analyze Xiaomi Corporation"
query2 = "Analyze the company Tesla."
# run_agent(query=query1, dump_messages=True)
# run_agent_with_experience(query_first=query1, query_second=query2)
# dump_experience()
run_agent(query=query1, dump_messages=True)
run_agent_with_experience(query_first=query1, query_second=query2)
dump_experience()
load_experience()

View file

@ -1,310 +0,0 @@
# ExperienceMaker Advanced Configuration Guide
This guide covers advanced configuration topics including custom pipelines, operation parameters, and configuration
methods.
## 🏗️ Configuration Architecture
ExperienceMaker uses a layered configuration system with the following priority order:
1. **Default Configuration** (lowest priority)
2. **YAML Configuration File**
3. **Command Line Arguments** (highest priority)
## 📁 Configuration Structure
```yaml
# Service Configuration
http_service:
host: "0.0.0.0"
port: 8001
timeout_keep_alive: 600
limit_concurrency: 64
# Pipeline Definitions
api:
retriever: recall_experience_op->rerank_experience_op->rewrite_experience_op
summarizer: trajectory_preprocess_op->[success_extraction_op|failure_extraction_op]->experience_validation_op
vector_store: vector_store_action_op
# Operation Configurations
op:
operation_name:
backend: operation_backend
llm: default # Optional: reference to LLM config
embedding_model: default # Optional: reference to embedding config
vector_store: default # Optional: reference to vector store config
params: # Operation-specific parameters
param1: value1
param2: value2
# Resource Configurations
llm:
default:
backend: openai_compatible
model_name: qwen3-32b
params:
temperature: 0.6
embedding_model:
default:
backend: openai_compatible
model_name: text-embedding-v4
params:
dimensions: 1024
vector_store:
default:
backend: local_file
embedding_model: default
```
## 🔧 Pipeline Configuration
### Pipeline Syntax
Pipeline configurations use a special syntax to define operation flows:
- `->`: Sequential execution
- `[]`: Parallel execution group
- `|`: Alternative operations within parallel group
### Examples
```yaml
# Sequential pipeline
api:
retriever: op1->op2->op3
# Parallel execution
api:
summarizer: op1->[op2|op3|op4]->op5
# Complex pipeline with nested parallel operations
api:
retriever: preprocess_op->[recall_op->rerank_op|backup_op]->merge_op
```
## ⚙️ Custom Operation Parameters
### Operation Configuration Structure
```yaml
op:
custom_operation:
backend: custom_backend_name
llm: default # Reference to LLM configuration
vector_store: default # Reference to vector store
params: # Custom parameters for this operation
retrieve_top_k: 15 # Number of top results to retrieve
similarity_threshold: 0.8 # Similarity threshold for filtering
enable_rerank: true # Enable reranking functionality
custom_param: "custom_value" # Any custom parameter
```
### Common Operation Parameters
**Retrieval Operations:**
```yaml
recall_experience_op:
params:
retrieve_top_k: 15
similarity_threshold: 0.5
rerank_experience_op:
params:
enable_llm_rerank: true
enable_score_filter: false
top_k: 5
```
**Extraction Operations:**
```yaml
success_extraction_op:
params:
extraction_mode: "detailed"
include_context: true
experience_validation_op:
params:
validation_threshold: 0.5
strict_mode: false
```
## 🚀 Configuration Methods
### Method 1: Custom Configuration File
**Step 1:** Create your configuration file
```yaml
# my_custom_config.yaml
api:
retriever: custom_recall_op->custom_rerank_op
op:
custom_recall_op:
backend: recall_experience_op
params:
retrieve_top_k: 20
similarity_threshold: 0.7
llm:
default:
model_name: gpt-4
params:
temperature: 0.3
```
**Step 2:** Use the custom configuration
```bash
experiencemaker config_path=/path/to/my_custom_config.yaml
```
### Method 2: Command Line Parameters
Override any configuration parameter using dot notation:
```bash
# Basic parameter override
experiencemaker \
llm.default.model_name=gpt-4 \
embedding_model.default.model_name=text-embedding-3-large
# Operation parameters
experiencemaker \
op.recall_experience_op.params.retrieve_top_k=20 \
op.rerank_experience_op.params.top_k=8
# Service configuration
experiencemaker \
http_service.port=8080 \
thread_pool.max_workers=32
# Pipeline configuration
experiencemaker \
api.retriever="custom_op1->custom_op2"
```
### Method 3: Hybrid Approach
Combine configuration file with command line overrides:
```bash
experiencemaker \
config_path=/path/to/base_config.yaml \
llm.default.model_name=gpt-4 \
op.recall_experience_op.params.retrieve_top_k=25
```
## 🎯 Practical Examples
### Example 1: High-Performance Configuration
```bash
experiencemaker \
http_service.port=8002 \
thread_pool.max_workers=64 \
op.recall_experience_op.params.retrieve_top_k=50 \
op.rerank_experience_op.params.top_k=10 \
llm.default.params.temperature=0.1
```
### Example 2: Development Configuration
```yaml
# dev_config.yaml
http_service:
port: 8003
api:
retriever: recall_experience_op->rerank_experience_op
op:
recall_experience_op:
params:
retrieve_top_k: 5 # Faster for development
rerank_experience_op:
params:
top_k: 3
llm:
default:
model_name: qwen-turbo
params:
temperature: 0.8
```
```bash
experiencemaker config_path=dev_config.yaml
```
### Example 3: Multi-Backend Setup
```yaml
# multi_backend_config.yaml
llm:
fast:
backend: openai_compatible
model_name: qwen-turbo
params:
temperature: 0.9
accurate:
backend: openai_compatible
model_name: gpt-4
params:
temperature: 0.1
op:
quick_extraction_op:
backend: success_extraction_op
llm: fast
detailed_validation_op:
backend: experience_validation_op
llm: accurate
params:
validation_threshold: 0.8
```
## 📋 Configuration Tips
1. **Start Simple**: Begin with the default configuration and override specific parameters
2. **Use Environment Variables**: Set API keys and URLs in `.env` file
3. **Parameter Validation**: Invalid parameters will cause startup errors with detailed messages
4. **Performance Tuning**: Adjust `retrieve_top_k`, `top_k`, and `max_workers` based on your needs
5. **Pipeline Testing**: Use simple pipelines first, then gradually add complexity
## 🔍 Troubleshooting
### Common Issues
**Configuration Not Loading:**
```bash
# Check if config file exists and has correct YAML syntax
experiencemaker config_path=/full/path/to/config.yaml
```
**Parameter Override Not Working:**
```bash
# Use exact parameter path from configuration structure
experiencemaker op.operation_name.params.parameter_name=value
```
**Pipeline Syntax Errors:**
- Check for balanced brackets `[]`
- Ensure operation names exist in `op` section
- Use `|` only within `[]` groups
---
🎯 **Advanced Configuration Mastery!** You can now create sophisticated ExperienceMaker setups tailored to your specific
needs.

View file

@ -1,15 +1,9 @@
# Services Params Documentation
# Configuration Guide
This document describes all available command-line parameters for ExperienceMaker Service.
This document describes all available parameters for ExperienceMaker Service.
The application uses [OmegaConf](https://omegaconf.readthedocs.io/) for configuration management, supporting both YAML
files and command-line overrides.
## Basic Usage
```bash
experiencemaker [parameter1=value1] [parameter2=value2] ...
```
## Configuration Loading Priority
1. Default values from `AppConfig` dataclass
@ -17,7 +11,153 @@ experiencemaker [parameter1=value1] [parameter2=value2] ...
3. Custom YAML file (if `config_path` is specified)
4. Command-line overrides
## Basic Configuration Parameters
## 🏗️ Configuration Architecture
ExperienceMaker uses a layered configuration system with the following priority order:
1. **Default Configuration** (lowest priority)
2. **YAML Configuration File**
3. **Command Line Arguments** (highest priority)
## Basic Bash Usage
```bash
experiencemaker [parameter1=value1] [parameter2=value2] ...
```
## 🧩 YAML Configuration Composition
The YAML configuration file follows a specific composition pattern that enables flexible and modular configuration:
### 1. Resource Declaration
First, you declare the three core resources that form the foundation of the system:
- **`llm`**: Language model configurations
- **`embedding_model`**: Embedding model configurations
- **`vector_store`**: Vector storage configurations
In these sections, `default` (or any custom name) represents a declared configuration object that can be referenced
later:
```yaml
llm:
default: # This is a declared LLM configuration object
backend: openai_compatible
model_name: qwen3-32b
embedding_model:
default: # This is a declared embedding model configuration object
backend: openai_compatible
model_name: text-embedding-v4
vector_store:
default: # This is a declared vector store configuration object
backend: local_file
embedding_model: default
```
### 2. Operation Backend Registration
In the `op` section, each operation declares its `backend` implementation. The backend names are registered through
`@OP_REGISTRY.register()` decorator, typically converting camel-case class names to underscore format:
```yaml
op:
recall_experience_op:
backend: recall_experience_op # Registered via @OP_REGISTRY.register()
```
### 3. Resource References
Operations reference the previously declared resources using their names:
```yaml
op:
recall_experience_op:
backend: recall_experience_op
llm: default # References the declared LLM object
embedding_model: default # References the declared embedding model object
vector_store: default # References the declared vector store object
```
### 4. Pipeline
Pipeline configurations use a special syntax to define operation flows:
- `->`: Sequential execution
- `[]`: Parallel execution group
- `|`: Alternative operations within parallel group
### Examples
```yaml
# Sequential pipeline
api:
retriever: op1->op2->op3
# Parallel execution
summarizer: op1->[op2|op3|op4]->op5
# Complex pipeline with nested parallel operations
vector_store: preprocess_op->[recall_op->rerank_op|backup_op]->merge_op
```
This compositional approach enables:
- **Modularity**: Declare resources once, reference everywhere
- **Flexibility**: Mix and match different backends and configurations
- **Complexity**: Build sophisticated processing chains through pipeline syntax
## 📁 Configuration Structure
```yaml
# Service Configuration
http_service:
host: "0.0.0.0"
port: 8001
timeout_keep_alive: 600
limit_concurrency: 64
# Pipeline Definitions
api:
retriever: recall_experience_op->rerank_experience_op->rewrite_experience_op
summarizer: trajectory_preprocess_op->[success_extraction_op|failure_extraction_op]->experience_validation_op
vector_store: vector_store_action_op
# Operation Configurations
op:
operation_name:
backend: operation_name # Register through `@OP_REGISTRY.register()`, typically by converting camel-cased types into underscored names
llm: default # Optional: reference to LLM config, Register through `@LLM_REGISTRY.register()`
embedding_model: default # Optional: reference to embedding config, Register through `@EMBEDDING_MODEL_REGISTRY.register()`
vector_store: default # Optional: reference to vector store config, Register through `@VECTOR_STORE_REGISTRY.register()`
params: # Operation-specific parameters
param1: value1
param2: value2
# Resource Configurations
llm:
default:
backend: openai_compatible
model_name: qwen3-32b
params:
temperature: 0.6
embedding_model:
default:
backend: openai_compatible
model_name: text-embedding-v4
params:
dimensions: 1024
vector_store:
default:
backend: local_file
embedding_model: default
```
## Detailed Configuration Parameters
| Parameter | Type | Default Value | Description | Example |
|----------------------|--------|-----------------|----------------------------------------------------------------------|-------------------------------------------|
@ -86,38 +226,112 @@ parameters:
| `vector_store.{name}.embedding_model` | string | `""` | Reference to embedding model configuration | `vector_store.default.embedding_model=default` |
| `vector_store.{name}.params.{param}` | any | `{}` | Vector store-specific parameters | `vector_store.default.params.store_dir=file_vector_store` |
## Complete Example
Here's a complete example showing how to configure the entire system:
## 🎯 Practical Examples
### Example 1
```bash
experiencemaker \
http_service.port=8080 \
thread_pool.max_workers=20 \
llm.default.backend=openai_compatible \
llm.default.model_name=qwen3-32b \
llm.default.params.temperature=0.6 \
embedding_model.default.backend=openai_compatible \
embedding_model.default.model_name=text-embedding-v4 \
embedding_model.default.params.dimensions=1024 \
vector_store.default.backend=elasticsearch \
vector_store.default.embedding_model=default \
http_service.port=8002 \
thread_pool.max_workers=64 \
op.recall_experience_op.params.retrieve_top_k=50 \
op.rerank_experience_op.params.top_k=10 \
llm.default.params.temperature=0.1
```
## Configuration File vs Command Line
### Example 2
You can also create a YAML configuration file and override specific parameters:
```yaml
# dev_config.yaml
http_service:
port: 8003
1. Create a custom configuration file (`xxx/my_config.yaml`)
2. Use it with command-line overrides:
api:
retriever: recall_experience_op->rerank_experience_op
op:
recall_experience_op:
params:
retrieve_top_k: 5 # Faster for development
rerank_experience_op:
params:
top_k: 3
llm:
default:
model_name: qwen-turbo
params:
temperature: 0.8
```
```bash
experiencemaker config_path=xxx/my_config.yaml llm.default.model_name=qwen3-32b http_service.port=8080
experiencemaker config_path=dev_config.yaml
```
## Parameter Validation
### Example 3: Multi-Backend Setup
- All parameters are validated according to their types
- Referenced configurations (like `llm`, `embedding_model`, `vector_store`) must exist
- Backend implementations must be registered in their respective registries
- Nested parameters use dot notation for access
```yaml
# multi_backend_config.yaml
llm:
fast:
backend: openai_compatible
model_name: qwen-turbo
params:
temperature: 0.9
accurate:
backend: openai_compatible
model_name: gpt-4
params:
temperature: 0.1
op:
quick_extraction_op:
backend: success_extraction_op
llm: fast
detailed_validation_op:
backend: experience_validation_op
llm: accurate
params:
validation_threshold: 0.8
```
## 📋 Configuration Tips
1. **Start Simple**: Begin with the default configuration and override specific parameters
2. **Use Environment Variables**: Set API keys and URLs in `.env` file
3. **Parameter Validation**: Invalid parameters will cause startup errors with detailed messages
4. **Performance Tuning**: Adjust `retrieve_top_k`, `top_k`, and `max_workers` based on your needs
5. **Pipeline Testing**: Use simple pipelines first, then gradually add complexity
## 🔍 Troubleshooting
### Common Issues
**Configuration Not Loading:**
```bash
# Check if config file exists and has correct YAML syntax
experiencemaker config_path=/full/path/to/config.yaml
```
**Parameter Override Not Working:**
```bash
# Use exact parameter path from configuration structure
experiencemaker op.operation_name.params.parameter_name=value
```
**Pipeline Syntax Errors:**
- Check for balanced brackets `[]`
- Ensure operation names exist in `op` section
- Use `|` only within `[]` groups
---
🎯 **Advanced Configuration Mastery!** You can now create sophisticated ExperienceMaker setups tailored to your specific
needs.

View file

@ -13,7 +13,6 @@ Just as financial analysts develop analytical frameworks, senior engineers estab
- [ ] Coding
- [ ] Education
- [ ] Research
- etc
- [ ] Experience marketplace: community-driven experience sharing and exchange
## P0 - Support for Rich Experience Formats
@ -24,6 +23,14 @@ Expert knowledge extends beyond text to include debugged code, fine-tuned toolch
- [ ] **Tool Integration**: APIs, MCP configurations, and tool setups
- [ ] **Pipeline Templates**: Agent execution pipelines and multi-step tool combinations
## P0 - MCP Integration
Modernize our API architecture by migrating three core APIs to the Model Context Protocol (MCP) standard for improved interoperability and standardization.
- [ ] Summarizer API
- [ ] Retriever API
- [ ] Vector Store API
## P1 - Experience Validation & Optimization
AI-powered analysis of experience usage patterns and effectiveness, with automatic quality optimization and cross-task validation feedback loops.
@ -42,14 +49,26 @@ Transform valuable experience data from daily work into usable insights:
Enable AI to naturally become stronger through everyday work, rather than wasting real-world experience data due to format limitations.
## P2 - Open Source Experience Libraries
Democratize AI experience sharing by making curated experience libraries publicly available on Hugging Face, enabling the broader AI community to benefit from and contribute to professional experience repositories.
- [ ] **Hugging Face Integration**: Upload and maintain experience libraries on Hugging Face Hub
- [ ] **Community Contributions**: Enable community-driven experience library improvements and additions
- [ ] **Standardized Formats**: Establish standard formats for experience sharing across different domains
- [ ] **Version Control**: Implement versioning system for experience library updates and improvements
## Current TODO
- [x] op dev & test ready
- [ ] integrate into beyond-agent @jinli
- [x] op dev & test ready @jiaji
- [ ] cook_book-appworld code & readme @jiaji
- [ ] cook_book-bfcl-v3 op @zouyin delay 0730
- [ ] fix multi-process bug @jinli
- [x] fix multi-process bug @jinli
- [ ] logo optimize @jiaji
- [ ] Ready-made Experience Store @jinli, add appworld/bfcl-v3 default experience store @jiaji
- [x] Ready-made Experience Store @jinli, add appworld/bfcl-v3 default experience store @jiaji
- [ ] op config make up @jiaji
- [ ] config make up, easy to understand @jinli
- [x] config make up, easy to understand @jinli
- [x] refine readme @jinli
- [x] integrate into beyond-agent @jinli
- [ ] rm workspace_id in code @jinli

View file

@ -1,174 +0,0 @@
# ExperienceMaker Quick Start Guide
This guide will help you get started with ExperienceMaker quickly using practical examples.
## 🚀 What You'll Learn
- How to set up ExperienceMaker service
- Run an agent and generate experiences
- Retrieve and apply experiences to new tasks
- Build experience-enhanced agents
## 📋 Prerequisites
- Python 3.12+
- LLM API access (OpenAI or compatible)
- Embedding model API access
## 🛠️ Installation
### Option 1: Install from PyPI (Recommended)
```bash
pip install experiencemaker
```
### Option 2: Install from Source
```bash
git clone https://github.com/modelscope/ExperienceMaker.git
cd ExperienceMaker
pip install .
```
## ⚙️ Environment Setup
Create a `.env` file in your project directory:
```bash
# Required: LLM API configuration
LLM_API_KEY="sk-xxx"
LLM_BASE_URL="https://xxx.com/v1"
# Required: Embedding model configuration
EMBEDDING_MODEL_API_KEY="sk-xxx"
EMBEDDING_MODEL_BASE_URL="https://xxx.com/v1"
# Optional: Elasticsearch configuration (if using Elasticsearch backend)
```
## 🚀 Start the Service
For testing, use the `local_file` backend:
```bash
experiencemaker \
http_service.port=8001 \
llm.default.model_name=qwen3-32b \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file
```
The service will start on `http://localhost:8001`
### Elasticsearch Backend
```bash
experiencemaker \
llm.default.model_name=qwen3-32b \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=elasticsearch
```
**Setup Elasticsearch:**
```bash
export ES_HOSTS="http://localhost:9200"
# Quick setup using Elastic's official script
curl -fsSL https://elastic.co/start-local | sh
```
📖 **Need Help?** Refer to [Vector Store Setup](./doc/vector_store_setup.md) for comprehensive deployment guidance.
## 📝 Your First ExperienceMaker Script
Here's how to get started!
- The `load_dotenv()` function loads environment variables from your `.env` file, or you can manually export them.
- The `base_url` points to your ExperienceMaker service.
- The `workspace_id` serves as your experience storage namespace. Experiences in different workspaces remain completely
isolated and cannot access each other.
```python
import requests
from dotenv import load_dotenv
load_dotenv()
base_url = "http://0.0.0.0:8001/"
workspace_id = "test_workspace"
```
### 📊 Call Summarizer Examples
Batch summarize the trajectory list, where each trajectory consists of a message and a score.
- The message is the conversation history.
- The score represents the rating between 0 and 1, with 0 typically indicating failure and 1 indicating success.
```python
response = requests.post(url=base_url + "summarizer", json={
"workspace_id": workspace_id,
"traj_list": [
{"messages": messages, "score": 1.0}
]
})
response = response.json()
experience_list = response["experience_list"]
for experience in experience_list:
print(experience)
```
### 🔍 Call Retriever Examples
Retrieve the top_k={top_k} experiences related to {query} in workspace=test_workspace, and finally accept the assembled context.
Alternatively, you can also accept the raw experience_list parameter and assemble the context yourself.
```python
response = requests.post(url=base_url + "retriever", json={
"workspace_id": workspace_id,
"query": query,
"top_k": 1,
})
response = response.json()
experience_merged: str = response["experience_merged"]
print(f"experience_merged={experience_merged}")
```
### 💾 Dump Experiences From Vector Store
Dump the experience with workspace_id from the vector store into the {path}/{workspace_id}.jsonl file.
```python
response = requests.post(url=base_url + "vector_store", json={
"workspace_id": workspace_id,
"action": "dump",
"path": "./",
})
print(response.json())
```
### 📥 Load Experiences To Vector Store
Load the {path}/{workspace_id}.jsonl file into the vector store, workspace_id={workspace_id}.
```python
response = requests.post(url=base_url + "vector_store", json={
"workspace_id": workspace_id,
"action": "load",
"path": "./",
})
print(response.json())
```
🎭 **Want to See It in Action?** We've prepared a [simple react agent](./cookbook/simple_demo/simple_demo.py) that demonstrates how to enhance agent capabilities by integrating summarizer and retriever components, achieving significantly better performance.
## 🐛 Common Issues
### Service Won't Start
- Check if port 8001 is available
- Verify your API keys in `.env` file
- Ensure Python version is 3.12+
### No Experiences Retrieved
- Make sure you've run the summarizer first
- Check if workspace_id matches between operations
- Verify vector store backend is properly configured
### API Connection Errors
- Confirm LLM_BASE_URL and API keys are correct
- Test API access independently
- Check network connectivity
---
🎯 **You're all set!** You now have a working ExperienceMaker setup that can learn from interactions and improve over time.

File diff suppressed because one or more lines are too long

View file

@ -65,7 +65,7 @@ class FailureExtractionOp(BaseOp):
experiences = []
for exp_data in experiences_data:
experience = BaseExperience(
experience: BaseExperience = TextExperience(
workspace_id=self.context.request.workspace_id,
when_to_use=exp_data.get("when_to_use", exp_data.get("condition", "")),
content=exp_data.get("experience", ""),

View file

@ -7,7 +7,7 @@ from loguru import logger
from experiencemaker.enumeration.role import Role
from experiencemaker.op import OP_REGISTRY
from experiencemaker.op.base_op import BaseOp
from experiencemaker.schema.experience import BaseExperience, ExperienceMeta
from experiencemaker.schema.experience import BaseExperience, ExperienceMeta, TextExperience
from experiencemaker.schema.message import Message, Trajectory
from experiencemaker.schema.response import SummarizerResponse
from experiencemaker.utils.op_utils import merge_messages_content, parse_json_experience_response, get_trajectory_context
@ -59,12 +59,12 @@ class SuccessExtractionOp(BaseOp):
outcome="successful"
)
def parse_experiences(message: Message) -> List[BaseExperience]:
def parse_experiences(message: Message) -> List[TextExperience]:
experiences_data = parse_json_experience_response(message.content)
experiences = []
for exp_data in experiences_data:
experience = BaseExperience(
experience = TextExperience(
workspace_id=self.context.request.workspace_id,
when_to_use=exp_data.get("when_to_use", exp_data.get("condition", "")),
content=exp_data.get("experience", ""),

View file

@ -1,7 +1,9 @@
import datetime
from abc import ABC
from typing import List
from uuid import uuid4
from loguru import logger
from pydantic import BaseModel, Field
from experiencemaker.schema.vector_node import VectorNode
@ -17,7 +19,7 @@ class ExperienceMeta(BaseModel):
self.modified_time = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
class BaseExperience(BaseModel):
class BaseExperience(BaseModel, ABC):
workspace_id: str = Field(default="")
experience_id: str = Field(default_factory=lambda: uuid4().hex)
@ -101,8 +103,8 @@ def vector_node_to_experience(node: VectorNode) -> BaseExperience:
return KnowledgeExperience.from_vector_node(node)
else:
raise RuntimeError(f"experience type {experience_type} not supported")
logger.warning(f"experience type {experience_type} not supported")
return TextExperience.from_vector_node(node)
if __name__ == "__main__":
e1 = TextExperience(

View file

@ -28,7 +28,9 @@ class BaseVectorStore(BaseModel, ABC):
try:
for line in tqdm(f, desc="load from path"):
if line.strip():
yield VectorNode(**json.loads(line.strip(), **kwargs))
node = VectorNode(**json.loads(line.strip(), **kwargs))
node.workspace_id = workspace_id
yield node
finally:
fcntl.flock(f, fcntl.LOCK_UN)
@ -45,6 +47,7 @@ class BaseVectorStore(BaseModel, ABC):
fcntl.flock(f, fcntl.LOCK_EX)
try:
for node in tqdm(nodes, desc="dump to path"):
node.workspace_id = workspace_id
f.write(json.dumps(node.model_dump(), ensure_ascii=ensure_ascii, **kwargs))
f.write("\n")
count += 1

View file

@ -51,14 +51,15 @@ class EsVectorStore(BaseVectorStore):
def _iter_workspace_nodes(self, workspace_id: str, **kwargs) -> Iterable[VectorNode]:
response = self._client.search(index=workspace_id, body={"query": {"match_all": {}}})
for doc in response['hits']['hits']:
yield self.doc2node(doc)
yield self.doc2node(doc, workspace_id)
def refresh(self, workspace_id: str):
self._client.indices.refresh(index=workspace_id)
@staticmethod
def doc2node(doc) -> VectorNode:
def doc2node(doc, workspace_id: str) -> VectorNode:
node = VectorNode(**doc["_source"])
node.workspace_id = workspace_id
node.unique_id = doc["_id"]
if "_score" in doc:
node.metadata["_score"] = doc["_score"] - 1
@ -105,7 +106,7 @@ class EsVectorStore(BaseVectorStore):
nodes: List[VectorNode] = []
for doc in response['hits']['hits']:
nodes.append(self.doc2node(doc))
nodes.append(self.doc2node(doc, workspace_id))
self.retrieve_filters.clear()
return nodes
@ -124,10 +125,10 @@ class EsVectorStore(BaseVectorStore):
docs = [
{
"_op_type": "index",
"_index": node.workspace_id,
"_index": workspace_id,
"_id": node.unique_id,
"_source": {
"workspace_id": node.workspace_id,
"workspace_id": workspace_id,
"content": node.content,
"metadata": node.metadata,
"vector": node.vector

View file

@ -84,6 +84,7 @@ class FileVectorStore(BaseVectorStore):
workspace_id=workspace_id,
path=self.store_path,
**kwargs)
logger.info(f"update workspace_id={workspace_id} nodes.size={len(nodes)} all.size={len(all_node_dict)} "
f"update_cnt={update_cnt}")