mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-07 03:00:27 +00:00
docs(cookbook): update quickstart guides for AppWorld and FrozenLake experiments
- Update AppWorld quickstart guide to use ReMe instead of ExperienceMaker - Update FrozenLake quickstart guide to use ReMe and improve clarity - Refactor FrozenLake experiment implementation and documentation - Add more detailed explanation of task memory mechanism in FrozenLake
This commit is contained in:
parent
f3911c92fd
commit
d50b55ace0
2 changed files with 107 additions and 107 deletions
|
|
@ -1,14 +1,14 @@
|
|||
# AppWorld Experiment Quick Start Guide
|
||||
|
||||
This guide helps you quickly set up and run AppWorld experiments with ExperienceMaker integration.
|
||||
This guide helps you quickly set up and run AppWorld experiments with ReMe integration.
|
||||
|
||||
## Env Setup
|
||||
|
||||
### 1. Clone the Repository
|
||||
|
||||
```bash
|
||||
git clone https://github.com/modelscope/ExperienceMaker.git
|
||||
cd ExperienceMaker/cookbook/appworld
|
||||
git clone https://github.com/modelscope/ReMe.git
|
||||
cd ReMe/cookbook/appworld
|
||||
```
|
||||
|
||||
### 2. Appworld Environment Setup
|
||||
|
|
@ -36,43 +36,43 @@ appworld download data
|
|||
|
||||
**Note**: The AppWorld data will be saved in the current directory.
|
||||
|
||||
### 3. Start ExperienceMaker Service
|
||||
### 3. Start ReMe Service
|
||||
|
||||
Install ExperienceMaker (if not already installed)
|
||||
If you haven't installed the ExperienceMaker environment yet, follow these steps:
|
||||
Install ReMe (if not already installed)
|
||||
If you haven't installed the ReMe environment yet, follow these steps:
|
||||
```bash
|
||||
# Go back to the project root
|
||||
cd ../..
|
||||
|
||||
# Create ExperienceMaker environment
|
||||
conda create -p ./em-env python==3.12
|
||||
conda activate ./em-env
|
||||
# Create ReMe environment
|
||||
conda create -p ./reme-env python==3.12
|
||||
conda activate ./reme-env
|
||||
|
||||
# Install ExperienceMaker
|
||||
# Install ReMe
|
||||
pip install .
|
||||
```
|
||||
|
||||
Launch the ExperienceMaker service to enable experience library functionality:
|
||||
Launch the ReMe service to enable memory library functionality:
|
||||
|
||||
```bash
|
||||
experiencemaker \
|
||||
reme \
|
||||
http_service.port=8001 \
|
||||
llm.default.model_name=qwen-max-latest \
|
||||
embedding_model.default.model_name=text-embedding-v4 \
|
||||
vector_store.default.backend=local_file
|
||||
```
|
||||
|
||||
add experiences for appworld:
|
||||
add memories for appworld:
|
||||
```bash
|
||||
curl -X POST "http://0.0.0.0:8001/vector_store" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"workspace_id": "appworld_v1",
|
||||
"action": "dump",
|
||||
"path": "./experience_library"
|
||||
"path": "./memory_library"
|
||||
}'
|
||||
```
|
||||
Now you have loaded the ExperienceMaker experience library to enable experience-based agent!
|
||||
Now you have loaded the ReMe memory library to enable memory-based agent!
|
||||
|
||||
### 4. Common Issues
|
||||
|
||||
|
|
@ -84,9 +84,9 @@ Now you have loaded the ExperienceMaker experience library to enable experience-
|
|||
|
||||
## Run Experiments
|
||||
|
||||
### 1. Test: With Experience vs Without Experience
|
||||
### 1. Test: With Memory vs Without Memory
|
||||
|
||||
Run the main experiment script to compare performance with and without experience:
|
||||
Run the main experiment script to compare performance with and without memory:
|
||||
|
||||
```bash
|
||||
python run_appworld.py
|
||||
|
|
@ -94,7 +94,7 @@ python run_appworld.py
|
|||
|
||||
**What this does:**
|
||||
- Runs AppWorld tasks on the development dataset
|
||||
- Compares agent performance with experience (`use_experience=True`) vs without experience
|
||||
- Compares agent performance with ReMe memory (`use_memory=True`) vs without memory
|
||||
- Uses multiple workers for parallel processing
|
||||
- Runs each task multiple times for statistical significance
|
||||
- Results are automatically saved to `./exp_result/` directory
|
||||
|
|
@ -102,7 +102,7 @@ python run_appworld.py
|
|||
**Configuration options in `run_appworld.py`:**
|
||||
- `max_workers`: Number of parallel workers (default: 6)
|
||||
- `num_runs`: Number of times each task is repeated (default: 4)
|
||||
- `use_experience`: Whether to use ExperienceMaker experience library
|
||||
- `use_memory`: Whether to use ReMe memory library
|
||||
|
||||
### 2. View Experiment Results
|
||||
|
||||
|
|
@ -131,10 +131,10 @@ python run_exp_statistic.py
|
|||
## Understanding Results
|
||||
|
||||
The experiment compares:
|
||||
1. **Baseline**: Agent without experience library
|
||||
2. **With Experience**: Agent enhanced with ExperienceMaker experience library
|
||||
1. **Baseline**: Agent without memory library
|
||||
2. **With Memory**: Agent enhanced with ReMe memory library
|
||||
|
||||
Key metrics to look for:
|
||||
- **best@1**: Average performance across all single runs
|
||||
- **best@k**: Performance when taking the best of k attempts
|
||||
- Improvement percentage when using experience vs baseline
|
||||
- Improvement percentage when using memory vs baseline
|
||||
|
|
@ -1,14 +1,14 @@
|
|||
# FrozenLake Experiment Quick Start Guide
|
||||
|
||||
This guide helps you quickly set up and run FrozenLake experiments with ExperienceMaker integration.
|
||||
This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task.
|
||||
|
||||
## Env Setup
|
||||
## Environment Setup
|
||||
|
||||
### 1. Clone the Repository
|
||||
|
||||
```bash
|
||||
git clone https://github.com/modelscope/ExperienceMaker.git
|
||||
cd ExperienceMaker/cookbook/frozenlake
|
||||
git clone https://github.com/modelscope/ReMe.git
|
||||
cd ReMe/cookbook/frozenlake
|
||||
```
|
||||
|
||||
### 2. FrozenLake Environment Setup
|
||||
|
|
@ -19,49 +19,47 @@ Install Gymnasium for FrozenLake environment:
|
|||
pip install gymnasium
|
||||
```
|
||||
|
||||
### 3. Start ExperienceMaker Service
|
||||
This will install:
|
||||
- gymnasium - for the FrozenLake environment
|
||||
- ray - for parallel execution
|
||||
- openai - for LLM API access
|
||||
- other dependencies
|
||||
|
||||
Install ExperienceMaker (if not already installed)
|
||||
If you haven't installed the ExperienceMaker environment yet, follow these steps:
|
||||
### 3. Start ReMe Service
|
||||
|
||||
If you haven't installed ReMe yet, follow these steps:
|
||||
```bash
|
||||
# Go back to the project root
|
||||
cd ../..
|
||||
|
||||
# Create ExperienceMaker environment
|
||||
conda create -p ./em-env python==3.12
|
||||
conda activate ./em-env
|
||||
# Create a virtual environment (optional)
|
||||
conda create -p ./reme-env python==3.10
|
||||
conda activate ./reme-env
|
||||
|
||||
# Install ExperienceMaker
|
||||
# Install ReMe
|
||||
pip install .
|
||||
```
|
||||
|
||||
Launch the ExperienceMaker service to enable experience library functionality:
|
||||
Launch the ReMe service to enable memory library functionality:
|
||||
|
||||
```bash
|
||||
experiencemaker \
|
||||
http_service.port=8001 \
|
||||
llm.default.model_name=qwen-max-latest \
|
||||
reme \
|
||||
backend=http \
|
||||
http.port=8002 \
|
||||
llm.default.model_name=qwen-max-2025-01-25 \
|
||||
embedding_model.default.model_name=text-embedding-v4 \
|
||||
vector_store.default.backend=local_file
|
||||
vector_store.default.backend=local
|
||||
```
|
||||
|
||||
Load default experience library for FrozenLake:
|
||||
Load default memory library for FrozenLake:
|
||||
```bash
|
||||
curl -X POST "http://0.0.0.0:8001/vector_store" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"workspace_id": "frozenlake_no_slippery",
|
||||
"action": "dump",
|
||||
"path": "./experience_library"
|
||||
}'
|
||||
```
|
||||
Now you have loaded the default FrozenLake experience library to enable experience-based agent!
|
||||
|
||||
## Run Experiments
|
||||
|
||||
### 1. Quick Test: Performance Evaluation Only (Default)
|
||||
|
||||
Run the main experiment script to test agent performance using existing experience:
|
||||
Run the main experiment script to test agent performance using existing memory:
|
||||
|
||||
```bash
|
||||
python run_frozenlake.py
|
||||
|
|
@ -69,54 +67,33 @@ python run_frozenlake.py
|
|||
|
||||
**What this does:**
|
||||
- Tests the agent on randomly generated FrozenLake maps
|
||||
- Uses the default experience library (`frozenlake_no_slippery`)
|
||||
- Uses the default memory library (`frozenlake_no_slippery`)
|
||||
- Evaluates performance with multiple runs for statistical significance
|
||||
- Results are automatically saved to `./exp_result/` directory
|
||||
|
||||
### 2. Advanced: Training + Testing (Experience Generation)
|
||||
### 2. Advanced: Training + Testing (Memory Generation)
|
||||
|
||||
To create new experiences through training and then test performance:
|
||||
To create new memories through training and then test performance:
|
||||
|
||||
```bash
|
||||
python run_frozenlake.py --enable-training
|
||||
You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function:
|
||||
|
||||
```python
|
||||
def main():
|
||||
experiment_name = "frozenlake_no_slippery" # Name of the experiment
|
||||
max_workers = 4 # Number of parallel workers
|
||||
training_runs = 4 # Runs per training map
|
||||
num_training_maps = 50 # Number of maps for training
|
||||
test_runs = 1 # Runs per test configuration
|
||||
num_test_maps = 100 # Number of test maps
|
||||
is_slippery = False # Enable slippery mode
|
||||
```
|
||||
|
||||
**What this does:**
|
||||
- **Stage 1 (Training)**: Generates new experiences by solving training maps
|
||||
- **Stage 2 (Testing)**: Evaluates performance using the generated experiences
|
||||
- Compares baseline performance vs experience-enhanced performance
|
||||
Key parameters to consider:
|
||||
- `experiment_name`: Used as the workspace ID for task memory
|
||||
- `is_slippery`: When True, agent movement becomes stochastic (harder)
|
||||
- `max_workers`: Increase for faster execution on multi-core systems
|
||||
|
||||
### 3. Custom Configuration Examples
|
||||
|
||||
**Basic customization:**
|
||||
```bash
|
||||
python run_frozenlake.py --experiment-name "my_frozenlake_test" --max-workers 8
|
||||
```
|
||||
|
||||
**Enable slippery mode:**
|
||||
```bash
|
||||
python run_frozenlake.py --slippery --experiment-name "frozenlake_slippery"
|
||||
```
|
||||
|
||||
**Full training experiment:**
|
||||
```bash
|
||||
python run_frozenlake.py \
|
||||
--enable-training \
|
||||
--experiment-name "frozenlake_training_experiment" \
|
||||
--max-workers 8 \
|
||||
--training-runs 4 \
|
||||
--num-training-maps 50 \
|
||||
--test-runs 5 \
|
||||
--num-test-maps 100 \
|
||||
--slippery
|
||||
```
|
||||
|
||||
**View all available options:**
|
||||
```bash
|
||||
python run_frozenlake.py --help
|
||||
```
|
||||
|
||||
### 4. View Experiment Results
|
||||
### 3. View Experiment Results
|
||||
|
||||
After running experiments, analyze the statistical results:
|
||||
|
||||
|
|
@ -128,30 +105,53 @@ python run_exp_statistic.py
|
|||
- Processes all result files in `./exp_result/`
|
||||
- Calculates success rates and performance metrics
|
||||
- Generates a summary table showing performance comparisons
|
||||
- Saves results to `experiment_summary.csv`
|
||||
- Analyzes the effect of task memory on performance
|
||||
- Saves results to `frozenlake_summary.csv`
|
||||
|
||||
## Configuration Parameters
|
||||
## Understanding the Implementation
|
||||
|
||||
| Parameter | Default Value | Description |
|
||||
|-----------|---------------|-------------|
|
||||
| `--experiment-name` | `frozenlake_no_slippery` | Name of the experiment |
|
||||
| `--max-workers` | `4` | Number of parallel workers |
|
||||
| `--enable-training` | `False` | Enable training phase (experience generation) |
|
||||
| `--training-runs` | `4` | Number of runs per training map |
|
||||
| `--num-training-maps` | `50` | Number of training maps |
|
||||
| `--test-runs` | `1` | Number of runs per test configuration |
|
||||
| `--num-test-maps` | `100` | Number of test maps to use |
|
||||
| `--slippery` | `False` | Enable slippery ice mode |
|
||||
### Key Components
|
||||
|
||||
1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`)
|
||||
- Implements a ReAct agent that interacts with the FrozenLake environment
|
||||
- Handles task memory retrieval and storage
|
||||
- Uses LLM (via OpenAI API) for decision making
|
||||
|
||||
2. **Experiment Runner** (`run_frozenlake.py`)
|
||||
- Manages the overall experiment flow
|
||||
- Handles training and testing phases
|
||||
- Uses Ray for parallel execution
|
||||
|
||||
3. **Map Manager** (`map_manager.py`)
|
||||
- Generates and manages test maps
|
||||
- Ensures consistent evaluation across experiments
|
||||
|
||||
4. **Statistics Analyzer** (`run_exp_statistic.py`)
|
||||
- Processes experiment results
|
||||
- Calculates performance metrics
|
||||
- Generates comparative analysis
|
||||
|
||||
## Understanding Results
|
||||
|
||||
The experiment evaluates agent performance on FrozenLake maps:
|
||||
|
||||
- **Success Rate**: Percentage of episodes that reach the goal
|
||||
- **Default Mode**: Uses existing experience library for quick testing
|
||||
- **Training Mode**: Generates new experiences then tests performance improvement
|
||||
- **Success Rate**: Percentage of episodes where the agent reaches the goal
|
||||
- **With vs. Without Memory**: Compares performance with and without task memory
|
||||
- **Slippery vs. Non-slippery**: Compares performance in different environment dynamics
|
||||
|
||||
**Output Files:**
|
||||
- `./exp_result/*.jsonl`: Raw experiment results
|
||||
- `./exp_result/experiment_summary.csv`: Statistical summary
|
||||
- Console output: Real-time progress and metrics
|
||||
### Output Files
|
||||
|
||||
- `./exp_result/*_training.jsonl`: Results from training phase
|
||||
- `./exp_result/*_test_no_memory.jsonl`: Test results without task memory
|
||||
- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory
|
||||
- `./exp_result/frozenlake_summary.csv`: Statistical summary
|
||||
|
||||
### Task Memory Mechanism
|
||||
|
||||
The task memory system works as follows:
|
||||
|
||||
1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service
|
||||
2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map
|
||||
3. **Memory Application**: The agent uses retrieved memories to guide its decision-making
|
||||
|
||||
The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake.
|
||||
Loading…
Add table
Reference in a new issue