mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-07 03:00:27 +00:00
docs(cookbook): update quickstart guides for AppWorld and FrozenLake experiments
- Update AppWorld quickstart guide to use ReMe instead of ExperienceMaker - Update FrozenLake quickstart guide to use ReMe and improve clarity - Refactor FrozenLake experiment implementation and documentation - Add more detailed explanation of task memory mechanism in FrozenLake
This commit is contained in:
parent
f3911c92fd
commit
d50b55ace0
2 changed files with 107 additions and 107 deletions
|
|
@ -1,14 +1,14 @@
|
||||||
# AppWorld Experiment Quick Start Guide
|
# AppWorld Experiment Quick Start Guide
|
||||||
|
|
||||||
This guide helps you quickly set up and run AppWorld experiments with ExperienceMaker integration.
|
This guide helps you quickly set up and run AppWorld experiments with ReMe integration.
|
||||||
|
|
||||||
## Env Setup
|
## Env Setup
|
||||||
|
|
||||||
### 1. Clone the Repository
|
### 1. Clone the Repository
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/modelscope/ExperienceMaker.git
|
git clone https://github.com/modelscope/ReMe.git
|
||||||
cd ExperienceMaker/cookbook/appworld
|
cd ReMe/cookbook/appworld
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. Appworld Environment Setup
|
### 2. Appworld Environment Setup
|
||||||
|
|
@ -36,43 +36,43 @@ appworld download data
|
||||||
|
|
||||||
**Note**: The AppWorld data will be saved in the current directory.
|
**Note**: The AppWorld data will be saved in the current directory.
|
||||||
|
|
||||||
### 3. Start ExperienceMaker Service
|
### 3. Start ReMe Service
|
||||||
|
|
||||||
Install ExperienceMaker (if not already installed)
|
Install ReMe (if not already installed)
|
||||||
If you haven't installed the ExperienceMaker environment yet, follow these steps:
|
If you haven't installed the ReMe environment yet, follow these steps:
|
||||||
```bash
|
```bash
|
||||||
# Go back to the project root
|
# Go back to the project root
|
||||||
cd ../..
|
cd ../..
|
||||||
|
|
||||||
# Create ExperienceMaker environment
|
# Create ReMe environment
|
||||||
conda create -p ./em-env python==3.12
|
conda create -p ./reme-env python==3.12
|
||||||
conda activate ./em-env
|
conda activate ./reme-env
|
||||||
|
|
||||||
# Install ExperienceMaker
|
# Install ReMe
|
||||||
pip install .
|
pip install .
|
||||||
```
|
```
|
||||||
|
|
||||||
Launch the ExperienceMaker service to enable experience library functionality:
|
Launch the ReMe service to enable memory library functionality:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
experiencemaker \
|
reme \
|
||||||
http_service.port=8001 \
|
http_service.port=8001 \
|
||||||
llm.default.model_name=qwen-max-latest \
|
llm.default.model_name=qwen-max-latest \
|
||||||
embedding_model.default.model_name=text-embedding-v4 \
|
embedding_model.default.model_name=text-embedding-v4 \
|
||||||
vector_store.default.backend=local_file
|
vector_store.default.backend=local_file
|
||||||
```
|
```
|
||||||
|
|
||||||
add experiences for appworld:
|
add memories for appworld:
|
||||||
```bash
|
```bash
|
||||||
curl -X POST "http://0.0.0.0:8001/vector_store" \
|
curl -X POST "http://0.0.0.0:8001/vector_store" \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{
|
-d '{
|
||||||
"workspace_id": "appworld_v1",
|
"workspace_id": "appworld_v1",
|
||||||
"action": "dump",
|
"action": "dump",
|
||||||
"path": "./experience_library"
|
"path": "./memory_library"
|
||||||
}'
|
}'
|
||||||
```
|
```
|
||||||
Now you have loaded the ExperienceMaker experience library to enable experience-based agent!
|
Now you have loaded the ReMe memory library to enable memory-based agent!
|
||||||
|
|
||||||
### 4. Common Issues
|
### 4. Common Issues
|
||||||
|
|
||||||
|
|
@ -84,9 +84,9 @@ Now you have loaded the ExperienceMaker experience library to enable experience-
|
||||||
|
|
||||||
## Run Experiments
|
## Run Experiments
|
||||||
|
|
||||||
### 1. Test: With Experience vs Without Experience
|
### 1. Test: With Memory vs Without Memory
|
||||||
|
|
||||||
Run the main experiment script to compare performance with and without experience:
|
Run the main experiment script to compare performance with and without memory:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python run_appworld.py
|
python run_appworld.py
|
||||||
|
|
@ -94,7 +94,7 @@ python run_appworld.py
|
||||||
|
|
||||||
**What this does:**
|
**What this does:**
|
||||||
- Runs AppWorld tasks on the development dataset
|
- Runs AppWorld tasks on the development dataset
|
||||||
- Compares agent performance with experience (`use_experience=True`) vs without experience
|
- Compares agent performance with ReMe memory (`use_memory=True`) vs without memory
|
||||||
- Uses multiple workers for parallel processing
|
- Uses multiple workers for parallel processing
|
||||||
- Runs each task multiple times for statistical significance
|
- Runs each task multiple times for statistical significance
|
||||||
- Results are automatically saved to `./exp_result/` directory
|
- Results are automatically saved to `./exp_result/` directory
|
||||||
|
|
@ -102,7 +102,7 @@ python run_appworld.py
|
||||||
**Configuration options in `run_appworld.py`:**
|
**Configuration options in `run_appworld.py`:**
|
||||||
- `max_workers`: Number of parallel workers (default: 6)
|
- `max_workers`: Number of parallel workers (default: 6)
|
||||||
- `num_runs`: Number of times each task is repeated (default: 4)
|
- `num_runs`: Number of times each task is repeated (default: 4)
|
||||||
- `use_experience`: Whether to use ExperienceMaker experience library
|
- `use_memory`: Whether to use ReMe memory library
|
||||||
|
|
||||||
### 2. View Experiment Results
|
### 2. View Experiment Results
|
||||||
|
|
||||||
|
|
@ -131,10 +131,10 @@ python run_exp_statistic.py
|
||||||
## Understanding Results
|
## Understanding Results
|
||||||
|
|
||||||
The experiment compares:
|
The experiment compares:
|
||||||
1. **Baseline**: Agent without experience library
|
1. **Baseline**: Agent without memory library
|
||||||
2. **With Experience**: Agent enhanced with ExperienceMaker experience library
|
2. **With Memory**: Agent enhanced with ReMe memory library
|
||||||
|
|
||||||
Key metrics to look for:
|
Key metrics to look for:
|
||||||
- **best@1**: Average performance across all single runs
|
- **best@1**: Average performance across all single runs
|
||||||
- **best@k**: Performance when taking the best of k attempts
|
- **best@k**: Performance when taking the best of k attempts
|
||||||
- Improvement percentage when using experience vs baseline
|
- Improvement percentage when using memory vs baseline
|
||||||
|
|
@ -1,14 +1,14 @@
|
||||||
# FrozenLake Experiment Quick Start Guide
|
# FrozenLake Experiment Quick Start Guide
|
||||||
|
|
||||||
This guide helps you quickly set up and run FrozenLake experiments with ExperienceMaker integration.
|
This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task.
|
||||||
|
|
||||||
## Env Setup
|
## Environment Setup
|
||||||
|
|
||||||
### 1. Clone the Repository
|
### 1. Clone the Repository
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/modelscope/ExperienceMaker.git
|
git clone https://github.com/modelscope/ReMe.git
|
||||||
cd ExperienceMaker/cookbook/frozenlake
|
cd ReMe/cookbook/frozenlake
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. FrozenLake Environment Setup
|
### 2. FrozenLake Environment Setup
|
||||||
|
|
@ -19,49 +19,47 @@ Install Gymnasium for FrozenLake environment:
|
||||||
pip install gymnasium
|
pip install gymnasium
|
||||||
```
|
```
|
||||||
|
|
||||||
### 3. Start ExperienceMaker Service
|
This will install:
|
||||||
|
- gymnasium - for the FrozenLake environment
|
||||||
|
- ray - for parallel execution
|
||||||
|
- openai - for LLM API access
|
||||||
|
- other dependencies
|
||||||
|
|
||||||
Install ExperienceMaker (if not already installed)
|
### 3. Start ReMe Service
|
||||||
If you haven't installed the ExperienceMaker environment yet, follow these steps:
|
|
||||||
|
If you haven't installed ReMe yet, follow these steps:
|
||||||
```bash
|
```bash
|
||||||
# Go back to the project root
|
# Go back to the project root
|
||||||
cd ../..
|
cd ../..
|
||||||
|
|
||||||
# Create ExperienceMaker environment
|
# Create a virtual environment (optional)
|
||||||
conda create -p ./em-env python==3.12
|
conda create -p ./reme-env python==3.10
|
||||||
conda activate ./em-env
|
conda activate ./reme-env
|
||||||
|
|
||||||
# Install ExperienceMaker
|
# Install ReMe
|
||||||
pip install .
|
pip install .
|
||||||
```
|
```
|
||||||
|
|
||||||
Launch the ExperienceMaker service to enable experience library functionality:
|
Launch the ReMe service to enable memory library functionality:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
experiencemaker \
|
reme \
|
||||||
http_service.port=8001 \
|
backend=http \
|
||||||
llm.default.model_name=qwen-max-latest \
|
http.port=8002 \
|
||||||
|
llm.default.model_name=qwen-max-2025-01-25 \
|
||||||
embedding_model.default.model_name=text-embedding-v4 \
|
embedding_model.default.model_name=text-embedding-v4 \
|
||||||
vector_store.default.backend=local_file
|
vector_store.default.backend=local
|
||||||
```
|
```
|
||||||
|
|
||||||
Load default experience library for FrozenLake:
|
Load default memory library for FrozenLake:
|
||||||
```bash
|
```bash
|
||||||
curl -X POST "http://0.0.0.0:8001/vector_store" \
|
|
||||||
-H "Content-Type: application/json" \
|
|
||||||
-d '{
|
|
||||||
"workspace_id": "frozenlake_no_slippery",
|
|
||||||
"action": "dump",
|
|
||||||
"path": "./experience_library"
|
|
||||||
}'
|
|
||||||
```
|
```
|
||||||
Now you have loaded the default FrozenLake experience library to enable experience-based agent!
|
|
||||||
|
|
||||||
## Run Experiments
|
## Run Experiments
|
||||||
|
|
||||||
### 1. Quick Test: Performance Evaluation Only (Default)
|
### 1. Quick Test: Performance Evaluation Only (Default)
|
||||||
|
|
||||||
Run the main experiment script to test agent performance using existing experience:
|
Run the main experiment script to test agent performance using existing memory:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python run_frozenlake.py
|
python run_frozenlake.py
|
||||||
|
|
@ -69,54 +67,33 @@ python run_frozenlake.py
|
||||||
|
|
||||||
**What this does:**
|
**What this does:**
|
||||||
- Tests the agent on randomly generated FrozenLake maps
|
- Tests the agent on randomly generated FrozenLake maps
|
||||||
- Uses the default experience library (`frozenlake_no_slippery`)
|
- Uses the default memory library (`frozenlake_no_slippery`)
|
||||||
- Evaluates performance with multiple runs for statistical significance
|
- Evaluates performance with multiple runs for statistical significance
|
||||||
- Results are automatically saved to `./exp_result/` directory
|
- Results are automatically saved to `./exp_result/` directory
|
||||||
|
|
||||||
### 2. Advanced: Training + Testing (Experience Generation)
|
### 2. Advanced: Training + Testing (Memory Generation)
|
||||||
|
|
||||||
To create new experiences through training and then test performance:
|
To create new memories through training and then test performance:
|
||||||
|
|
||||||
```bash
|
You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function:
|
||||||
python run_frozenlake.py --enable-training
|
|
||||||
|
```python
|
||||||
|
def main():
|
||||||
|
experiment_name = "frozenlake_no_slippery" # Name of the experiment
|
||||||
|
max_workers = 4 # Number of parallel workers
|
||||||
|
training_runs = 4 # Runs per training map
|
||||||
|
num_training_maps = 50 # Number of maps for training
|
||||||
|
test_runs = 1 # Runs per test configuration
|
||||||
|
num_test_maps = 100 # Number of test maps
|
||||||
|
is_slippery = False # Enable slippery mode
|
||||||
```
|
```
|
||||||
|
|
||||||
**What this does:**
|
Key parameters to consider:
|
||||||
- **Stage 1 (Training)**: Generates new experiences by solving training maps
|
- `experiment_name`: Used as the workspace ID for task memory
|
||||||
- **Stage 2 (Testing)**: Evaluates performance using the generated experiences
|
- `is_slippery`: When True, agent movement becomes stochastic (harder)
|
||||||
- Compares baseline performance vs experience-enhanced performance
|
- `max_workers`: Increase for faster execution on multi-core systems
|
||||||
|
|
||||||
### 3. Custom Configuration Examples
|
### 3. View Experiment Results
|
||||||
|
|
||||||
**Basic customization:**
|
|
||||||
```bash
|
|
||||||
python run_frozenlake.py --experiment-name "my_frozenlake_test" --max-workers 8
|
|
||||||
```
|
|
||||||
|
|
||||||
**Enable slippery mode:**
|
|
||||||
```bash
|
|
||||||
python run_frozenlake.py --slippery --experiment-name "frozenlake_slippery"
|
|
||||||
```
|
|
||||||
|
|
||||||
**Full training experiment:**
|
|
||||||
```bash
|
|
||||||
python run_frozenlake.py \
|
|
||||||
--enable-training \
|
|
||||||
--experiment-name "frozenlake_training_experiment" \
|
|
||||||
--max-workers 8 \
|
|
||||||
--training-runs 4 \
|
|
||||||
--num-training-maps 50 \
|
|
||||||
--test-runs 5 \
|
|
||||||
--num-test-maps 100 \
|
|
||||||
--slippery
|
|
||||||
```
|
|
||||||
|
|
||||||
**View all available options:**
|
|
||||||
```bash
|
|
||||||
python run_frozenlake.py --help
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. View Experiment Results
|
|
||||||
|
|
||||||
After running experiments, analyze the statistical results:
|
After running experiments, analyze the statistical results:
|
||||||
|
|
||||||
|
|
@ -128,30 +105,53 @@ python run_exp_statistic.py
|
||||||
- Processes all result files in `./exp_result/`
|
- Processes all result files in `./exp_result/`
|
||||||
- Calculates success rates and performance metrics
|
- Calculates success rates and performance metrics
|
||||||
- Generates a summary table showing performance comparisons
|
- Generates a summary table showing performance comparisons
|
||||||
- Saves results to `experiment_summary.csv`
|
- Analyzes the effect of task memory on performance
|
||||||
|
- Saves results to `frozenlake_summary.csv`
|
||||||
|
|
||||||
## Configuration Parameters
|
## Understanding the Implementation
|
||||||
|
|
||||||
| Parameter | Default Value | Description |
|
### Key Components
|
||||||
|-----------|---------------|-------------|
|
|
||||||
| `--experiment-name` | `frozenlake_no_slippery` | Name of the experiment |
|
1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`)
|
||||||
| `--max-workers` | `4` | Number of parallel workers |
|
- Implements a ReAct agent that interacts with the FrozenLake environment
|
||||||
| `--enable-training` | `False` | Enable training phase (experience generation) |
|
- Handles task memory retrieval and storage
|
||||||
| `--training-runs` | `4` | Number of runs per training map |
|
- Uses LLM (via OpenAI API) for decision making
|
||||||
| `--num-training-maps` | `50` | Number of training maps |
|
|
||||||
| `--test-runs` | `1` | Number of runs per test configuration |
|
2. **Experiment Runner** (`run_frozenlake.py`)
|
||||||
| `--num-test-maps` | `100` | Number of test maps to use |
|
- Manages the overall experiment flow
|
||||||
| `--slippery` | `False` | Enable slippery ice mode |
|
- Handles training and testing phases
|
||||||
|
- Uses Ray for parallel execution
|
||||||
|
|
||||||
|
3. **Map Manager** (`map_manager.py`)
|
||||||
|
- Generates and manages test maps
|
||||||
|
- Ensures consistent evaluation across experiments
|
||||||
|
|
||||||
|
4. **Statistics Analyzer** (`run_exp_statistic.py`)
|
||||||
|
- Processes experiment results
|
||||||
|
- Calculates performance metrics
|
||||||
|
- Generates comparative analysis
|
||||||
|
|
||||||
## Understanding Results
|
## Understanding Results
|
||||||
|
|
||||||
The experiment evaluates agent performance on FrozenLake maps:
|
The experiment evaluates agent performance on FrozenLake maps:
|
||||||
|
|
||||||
- **Success Rate**: Percentage of episodes that reach the goal
|
- **Success Rate**: Percentage of episodes where the agent reaches the goal
|
||||||
- **Default Mode**: Uses existing experience library for quick testing
|
- **With vs. Without Memory**: Compares performance with and without task memory
|
||||||
- **Training Mode**: Generates new experiences then tests performance improvement
|
- **Slippery vs. Non-slippery**: Compares performance in different environment dynamics
|
||||||
|
|
||||||
**Output Files:**
|
### Output Files
|
||||||
- `./exp_result/*.jsonl`: Raw experiment results
|
|
||||||
- `./exp_result/experiment_summary.csv`: Statistical summary
|
- `./exp_result/*_training.jsonl`: Results from training phase
|
||||||
- Console output: Real-time progress and metrics
|
- `./exp_result/*_test_no_memory.jsonl`: Test results without task memory
|
||||||
|
- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory
|
||||||
|
- `./exp_result/frozenlake_summary.csv`: Statistical summary
|
||||||
|
|
||||||
|
### Task Memory Mechanism
|
||||||
|
|
||||||
|
The task memory system works as follows:
|
||||||
|
|
||||||
|
1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service
|
||||||
|
2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map
|
||||||
|
3. **Memory Application**: The agent uses retrieved memories to guide its decision-making
|
||||||
|
|
||||||
|
The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake.
|
||||||
Loading…
Add table
Reference in a new issue