docs(cookbook): update quickstart guides for AppWorld and FrozenLake experiments

- Update AppWorld quickstart guide to use ReMe instead of ExperienceMaker
- Update FrozenLake quickstart guide to use ReMe and improve clarity
- Refactor FrozenLake experiment implementation and documentation
- Add more detailed explanation of task memory mechanism in FrozenLake
This commit is contained in:
jinli.yl 2025-09-01 23:48:10 +08:00
parent f3911c92fd
commit d50b55ace0
2 changed files with 107 additions and 107 deletions

View file

@ -1,14 +1,14 @@
# AppWorld Experiment Quick Start Guide # AppWorld Experiment Quick Start Guide
This guide helps you quickly set up and run AppWorld experiments with ExperienceMaker integration. This guide helps you quickly set up and run AppWorld experiments with ReMe integration.
## Env Setup ## Env Setup
### 1. Clone the Repository ### 1. Clone the Repository
```bash ```bash
git clone https://github.com/modelscope/ExperienceMaker.git git clone https://github.com/modelscope/ReMe.git
cd ExperienceMaker/cookbook/appworld cd ReMe/cookbook/appworld
``` ```
### 2. Appworld Environment Setup ### 2. Appworld Environment Setup
@ -36,43 +36,43 @@ appworld download data
**Note**: The AppWorld data will be saved in the current directory. **Note**: The AppWorld data will be saved in the current directory.
### 3. Start ExperienceMaker Service ### 3. Start ReMe Service
Install ExperienceMaker (if not already installed) Install ReMe (if not already installed)
If you haven't installed the ExperienceMaker environment yet, follow these steps: If you haven't installed the ReMe environment yet, follow these steps:
```bash ```bash
# Go back to the project root # Go back to the project root
cd ../.. cd ../..
# Create ExperienceMaker environment # Create ReMe environment
conda create -p ./em-env python==3.12 conda create -p ./reme-env python==3.12
conda activate ./em-env conda activate ./reme-env
# Install ExperienceMaker # Install ReMe
pip install . pip install .
``` ```
Launch the ExperienceMaker service to enable experience library functionality: Launch the ReMe service to enable memory library functionality:
```bash ```bash
experiencemaker \ reme \
http_service.port=8001 \ http_service.port=8001 \
llm.default.model_name=qwen-max-latest \ llm.default.model_name=qwen-max-latest \
embedding_model.default.model_name=text-embedding-v4 \ embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file vector_store.default.backend=local_file
``` ```
add experiences for appworld: add memories for appworld:
```bash ```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \ curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{ -d '{
"workspace_id": "appworld_v1", "workspace_id": "appworld_v1",
"action": "dump", "action": "dump",
"path": "./experience_library" "path": "./memory_library"
}' }'
``` ```
Now you have loaded the ExperienceMaker experience library to enable experience-based agent! Now you have loaded the ReMe memory library to enable memory-based agent!
### 4. Common Issues ### 4. Common Issues
@ -84,9 +84,9 @@ Now you have loaded the ExperienceMaker experience library to enable experience-
## Run Experiments ## Run Experiments
### 1. Test: With Experience vs Without Experience ### 1. Test: With Memory vs Without Memory
Run the main experiment script to compare performance with and without experience: Run the main experiment script to compare performance with and without memory:
```bash ```bash
python run_appworld.py python run_appworld.py
@ -94,7 +94,7 @@ python run_appworld.py
**What this does:** **What this does:**
- Runs AppWorld tasks on the development dataset - Runs AppWorld tasks on the development dataset
- Compares agent performance with experience (`use_experience=True`) vs without experience - Compares agent performance with ReMe memory (`use_memory=True`) vs without memory
- Uses multiple workers for parallel processing - Uses multiple workers for parallel processing
- Runs each task multiple times for statistical significance - Runs each task multiple times for statistical significance
- Results are automatically saved to `./exp_result/` directory - Results are automatically saved to `./exp_result/` directory
@ -102,7 +102,7 @@ python run_appworld.py
**Configuration options in `run_appworld.py`:** **Configuration options in `run_appworld.py`:**
- `max_workers`: Number of parallel workers (default: 6) - `max_workers`: Number of parallel workers (default: 6)
- `num_runs`: Number of times each task is repeated (default: 4) - `num_runs`: Number of times each task is repeated (default: 4)
- `use_experience`: Whether to use ExperienceMaker experience library - `use_memory`: Whether to use ReMe memory library
### 2. View Experiment Results ### 2. View Experiment Results
@ -131,10 +131,10 @@ python run_exp_statistic.py
## Understanding Results ## Understanding Results
The experiment compares: The experiment compares:
1. **Baseline**: Agent without experience library 1. **Baseline**: Agent without memory library
2. **With Experience**: Agent enhanced with ExperienceMaker experience library 2. **With Memory**: Agent enhanced with ReMe memory library
Key metrics to look for: Key metrics to look for:
- **best@1**: Average performance across all single runs - **best@1**: Average performance across all single runs
- **best@k**: Performance when taking the best of k attempts - **best@k**: Performance when taking the best of k attempts
- Improvement percentage when using experience vs baseline - Improvement percentage when using memory vs baseline

View file

@ -1,14 +1,14 @@
# FrozenLake Experiment Quick Start Guide # FrozenLake Experiment Quick Start Guide
This guide helps you quickly set up and run FrozenLake experiments with ExperienceMaker integration. This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task.
## Env Setup ## Environment Setup
### 1. Clone the Repository ### 1. Clone the Repository
```bash ```bash
git clone https://github.com/modelscope/ExperienceMaker.git git clone https://github.com/modelscope/ReMe.git
cd ExperienceMaker/cookbook/frozenlake cd ReMe/cookbook/frozenlake
``` ```
### 2. FrozenLake Environment Setup ### 2. FrozenLake Environment Setup
@ -19,49 +19,47 @@ Install Gymnasium for FrozenLake environment:
pip install gymnasium pip install gymnasium
``` ```
### 3. Start ExperienceMaker Service This will install:
- gymnasium - for the FrozenLake environment
- ray - for parallel execution
- openai - for LLM API access
- other dependencies
Install ExperienceMaker (if not already installed) ### 3. Start ReMe Service
If you haven't installed the ExperienceMaker environment yet, follow these steps:
If you haven't installed ReMe yet, follow these steps:
```bash ```bash
# Go back to the project root # Go back to the project root
cd ../.. cd ../..
# Create ExperienceMaker environment # Create a virtual environment (optional)
conda create -p ./em-env python==3.12 conda create -p ./reme-env python==3.10
conda activate ./em-env conda activate ./reme-env
# Install ExperienceMaker # Install ReMe
pip install . pip install .
``` ```
Launch the ExperienceMaker service to enable experience library functionality: Launch the ReMe service to enable memory library functionality:
```bash ```bash
experiencemaker \ reme \
http_service.port=8001 \ backend=http \
llm.default.model_name=qwen-max-latest \ http.port=8002 \
llm.default.model_name=qwen-max-2025-01-25 \
embedding_model.default.model_name=text-embedding-v4 \ embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file vector_store.default.backend=local
``` ```
Load default experience library for FrozenLake: Load default memory library for FrozenLake:
```bash ```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "frozenlake_no_slippery",
"action": "dump",
"path": "./experience_library"
}'
``` ```
Now you have loaded the default FrozenLake experience library to enable experience-based agent!
## Run Experiments ## Run Experiments
### 1. Quick Test: Performance Evaluation Only (Default) ### 1. Quick Test: Performance Evaluation Only (Default)
Run the main experiment script to test agent performance using existing experience: Run the main experiment script to test agent performance using existing memory:
```bash ```bash
python run_frozenlake.py python run_frozenlake.py
@ -69,54 +67,33 @@ python run_frozenlake.py
**What this does:** **What this does:**
- Tests the agent on randomly generated FrozenLake maps - Tests the agent on randomly generated FrozenLake maps
- Uses the default experience library (`frozenlake_no_slippery`) - Uses the default memory library (`frozenlake_no_slippery`)
- Evaluates performance with multiple runs for statistical significance - Evaluates performance with multiple runs for statistical significance
- Results are automatically saved to `./exp_result/` directory - Results are automatically saved to `./exp_result/` directory
### 2. Advanced: Training + Testing (Experience Generation) ### 2. Advanced: Training + Testing (Memory Generation)
To create new experiences through training and then test performance: To create new memories through training and then test performance:
```bash You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function:
python run_frozenlake.py --enable-training
```python
def main():
experiment_name = "frozenlake_no_slippery" # Name of the experiment
max_workers = 4 # Number of parallel workers
training_runs = 4 # Runs per training map
num_training_maps = 50 # Number of maps for training
test_runs = 1 # Runs per test configuration
num_test_maps = 100 # Number of test maps
is_slippery = False # Enable slippery mode
``` ```
**What this does:** Key parameters to consider:
- **Stage 1 (Training)**: Generates new experiences by solving training maps - `experiment_name`: Used as the workspace ID for task memory
- **Stage 2 (Testing)**: Evaluates performance using the generated experiences - `is_slippery`: When True, agent movement becomes stochastic (harder)
- Compares baseline performance vs experience-enhanced performance - `max_workers`: Increase for faster execution on multi-core systems
### 3. Custom Configuration Examples ### 3. View Experiment Results
**Basic customization:**
```bash
python run_frozenlake.py --experiment-name "my_frozenlake_test" --max-workers 8
```
**Enable slippery mode:**
```bash
python run_frozenlake.py --slippery --experiment-name "frozenlake_slippery"
```
**Full training experiment:**
```bash
python run_frozenlake.py \
--enable-training \
--experiment-name "frozenlake_training_experiment" \
--max-workers 8 \
--training-runs 4 \
--num-training-maps 50 \
--test-runs 5 \
--num-test-maps 100 \
--slippery
```
**View all available options:**
```bash
python run_frozenlake.py --help
```
### 4. View Experiment Results
After running experiments, analyze the statistical results: After running experiments, analyze the statistical results:
@ -128,30 +105,53 @@ python run_exp_statistic.py
- Processes all result files in `./exp_result/` - Processes all result files in `./exp_result/`
- Calculates success rates and performance metrics - Calculates success rates and performance metrics
- Generates a summary table showing performance comparisons - Generates a summary table showing performance comparisons
- Saves results to `experiment_summary.csv` - Analyzes the effect of task memory on performance
- Saves results to `frozenlake_summary.csv`
## Configuration Parameters ## Understanding the Implementation
| Parameter | Default Value | Description | ### Key Components
|-----------|---------------|-------------|
| `--experiment-name` | `frozenlake_no_slippery` | Name of the experiment | 1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`)
| `--max-workers` | `4` | Number of parallel workers | - Implements a ReAct agent that interacts with the FrozenLake environment
| `--enable-training` | `False` | Enable training phase (experience generation) | - Handles task memory retrieval and storage
| `--training-runs` | `4` | Number of runs per training map | - Uses LLM (via OpenAI API) for decision making
| `--num-training-maps` | `50` | Number of training maps |
| `--test-runs` | `1` | Number of runs per test configuration | 2. **Experiment Runner** (`run_frozenlake.py`)
| `--num-test-maps` | `100` | Number of test maps to use | - Manages the overall experiment flow
| `--slippery` | `False` | Enable slippery ice mode | - Handles training and testing phases
- Uses Ray for parallel execution
3. **Map Manager** (`map_manager.py`)
- Generates and manages test maps
- Ensures consistent evaluation across experiments
4. **Statistics Analyzer** (`run_exp_statistic.py`)
- Processes experiment results
- Calculates performance metrics
- Generates comparative analysis
## Understanding Results ## Understanding Results
The experiment evaluates agent performance on FrozenLake maps: The experiment evaluates agent performance on FrozenLake maps:
- **Success Rate**: Percentage of episodes that reach the goal - **Success Rate**: Percentage of episodes where the agent reaches the goal
- **Default Mode**: Uses existing experience library for quick testing - **With vs. Without Memory**: Compares performance with and without task memory
- **Training Mode**: Generates new experiences then tests performance improvement - **Slippery vs. Non-slippery**: Compares performance in different environment dynamics
**Output Files:** ### Output Files
- `./exp_result/*.jsonl`: Raw experiment results
- `./exp_result/experiment_summary.csv`: Statistical summary - `./exp_result/*_training.jsonl`: Results from training phase
- Console output: Real-time progress and metrics - `./exp_result/*_test_no_memory.jsonl`: Test results without task memory
- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory
- `./exp_result/frozenlake_summary.csv`: Statistical summary
### Task Memory Mechanism
The task memory system works as follows:
1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service
2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map
3. **Memory Application**: The agent uses retrieved memories to guide its decision-making
The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake.