docs(cookbook): update quickstart guides for AppWorld and FrozenLake experiments

- Update AppWorld quickstart guide to use ReMe instead of ExperienceMaker
- Update FrozenLake quickstart guide to use ReMe and improve clarity
- Refactor FrozenLake experiment implementation and documentation
- Add more detailed explanation of task memory mechanism in FrozenLake
This commit is contained in:
jinli.yl 2025-09-01 23:48:10 +08:00
parent f3911c92fd
commit d50b55ace0
2 changed files with 107 additions and 107 deletions

View file

@ -1,14 +1,14 @@
# AppWorld Experiment Quick Start Guide
This guide helps you quickly set up and run AppWorld experiments with ExperienceMaker integration.
This guide helps you quickly set up and run AppWorld experiments with ReMe integration.
## Env Setup
### 1. Clone the Repository
```bash
git clone https://github.com/modelscope/ExperienceMaker.git
cd ExperienceMaker/cookbook/appworld
git clone https://github.com/modelscope/ReMe.git
cd ReMe/cookbook/appworld
```
### 2. Appworld Environment Setup
@ -36,43 +36,43 @@ appworld download data
**Note**: The AppWorld data will be saved in the current directory.
### 3. Start ExperienceMaker Service
### 3. Start ReMe Service
Install ExperienceMaker (if not already installed)
If you haven't installed the ExperienceMaker environment yet, follow these steps:
Install ReMe (if not already installed)
If you haven't installed the ReMe environment yet, follow these steps:
```bash
# Go back to the project root
cd ../..
# Create ExperienceMaker environment
conda create -p ./em-env python==3.12
conda activate ./em-env
# Create ReMe environment
conda create -p ./reme-env python==3.12
conda activate ./reme-env
# Install ExperienceMaker
# Install ReMe
pip install .
```
Launch the ExperienceMaker service to enable experience library functionality:
Launch the ReMe service to enable memory library functionality:
```bash
experiencemaker \
reme \
http_service.port=8001 \
llm.default.model_name=qwen-max-latest \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file
```
add experiences for appworld:
add memories for appworld:
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "appworld_v1",
"action": "dump",
"path": "./experience_library"
"path": "./memory_library"
}'
```
Now you have loaded the ExperienceMaker experience library to enable experience-based agent!
Now you have loaded the ReMe memory library to enable memory-based agent!
### 4. Common Issues
@ -84,9 +84,9 @@ Now you have loaded the ExperienceMaker experience library to enable experience-
## Run Experiments
### 1. Test: With Experience vs Without Experience
### 1. Test: With Memory vs Without Memory
Run the main experiment script to compare performance with and without experience:
Run the main experiment script to compare performance with and without memory:
```bash
python run_appworld.py
@ -94,7 +94,7 @@ python run_appworld.py
**What this does:**
- Runs AppWorld tasks on the development dataset
- Compares agent performance with experience (`use_experience=True`) vs without experience
- Compares agent performance with ReMe memory (`use_memory=True`) vs without memory
- Uses multiple workers for parallel processing
- Runs each task multiple times for statistical significance
- Results are automatically saved to `./exp_result/` directory
@ -102,7 +102,7 @@ python run_appworld.py
**Configuration options in `run_appworld.py`:**
- `max_workers`: Number of parallel workers (default: 6)
- `num_runs`: Number of times each task is repeated (default: 4)
- `use_experience`: Whether to use ExperienceMaker experience library
- `use_memory`: Whether to use ReMe memory library
### 2. View Experiment Results
@ -131,10 +131,10 @@ python run_exp_statistic.py
## Understanding Results
The experiment compares:
1. **Baseline**: Agent without experience library
2. **With Experience**: Agent enhanced with ExperienceMaker experience library
1. **Baseline**: Agent without memory library
2. **With Memory**: Agent enhanced with ReMe memory library
Key metrics to look for:
- **best@1**: Average performance across all single runs
- **best@k**: Performance when taking the best of k attempts
- Improvement percentage when using experience vs baseline
- Improvement percentage when using memory vs baseline

View file

@ -1,14 +1,14 @@
# FrozenLake Experiment Quick Start Guide
This guide helps you quickly set up and run FrozenLake experiments with ExperienceMaker integration.
This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task.
## Env Setup
## Environment Setup
### 1. Clone the Repository
```bash
git clone https://github.com/modelscope/ExperienceMaker.git
cd ExperienceMaker/cookbook/frozenlake
git clone https://github.com/modelscope/ReMe.git
cd ReMe/cookbook/frozenlake
```
### 2. FrozenLake Environment Setup
@ -19,49 +19,47 @@ Install Gymnasium for FrozenLake environment:
pip install gymnasium
```
### 3. Start ExperienceMaker Service
This will install:
- gymnasium - for the FrozenLake environment
- ray - for parallel execution
- openai - for LLM API access
- other dependencies
Install ExperienceMaker (if not already installed)
If you haven't installed the ExperienceMaker environment yet, follow these steps:
### 3. Start ReMe Service
If you haven't installed ReMe yet, follow these steps:
```bash
# Go back to the project root
cd ../..
# Create ExperienceMaker environment
conda create -p ./em-env python==3.12
conda activate ./em-env
# Create a virtual environment (optional)
conda create -p ./reme-env python==3.10
conda activate ./reme-env
# Install ExperienceMaker
# Install ReMe
pip install .
```
Launch the ExperienceMaker service to enable experience library functionality:
Launch the ReMe service to enable memory library functionality:
```bash
experiencemaker \
http_service.port=8001 \
llm.default.model_name=qwen-max-latest \
reme \
backend=http \
http.port=8002 \
llm.default.model_name=qwen-max-2025-01-25 \
embedding_model.default.model_name=text-embedding-v4 \
vector_store.default.backend=local_file
vector_store.default.backend=local
```
Load default experience library for FrozenLake:
Load default memory library for FrozenLake:
```bash
curl -X POST "http://0.0.0.0:8001/vector_store" \
-H "Content-Type: application/json" \
-d '{
"workspace_id": "frozenlake_no_slippery",
"action": "dump",
"path": "./experience_library"
}'
```
Now you have loaded the default FrozenLake experience library to enable experience-based agent!
## Run Experiments
### 1. Quick Test: Performance Evaluation Only (Default)
Run the main experiment script to test agent performance using existing experience:
Run the main experiment script to test agent performance using existing memory:
```bash
python run_frozenlake.py
@ -69,54 +67,33 @@ python run_frozenlake.py
**What this does:**
- Tests the agent on randomly generated FrozenLake maps
- Uses the default experience library (`frozenlake_no_slippery`)
- Uses the default memory library (`frozenlake_no_slippery`)
- Evaluates performance with multiple runs for statistical significance
- Results are automatically saved to `./exp_result/` directory
### 2. Advanced: Training + Testing (Experience Generation)
### 2. Advanced: Training + Testing (Memory Generation)
To create new experiences through training and then test performance:
To create new memories through training and then test performance:
```bash
python run_frozenlake.py --enable-training
You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function:
```python
def main():
experiment_name = "frozenlake_no_slippery" # Name of the experiment
max_workers = 4 # Number of parallel workers
training_runs = 4 # Runs per training map
num_training_maps = 50 # Number of maps for training
test_runs = 1 # Runs per test configuration
num_test_maps = 100 # Number of test maps
is_slippery = False # Enable slippery mode
```
**What this does:**
- **Stage 1 (Training)**: Generates new experiences by solving training maps
- **Stage 2 (Testing)**: Evaluates performance using the generated experiences
- Compares baseline performance vs experience-enhanced performance
Key parameters to consider:
- `experiment_name`: Used as the workspace ID for task memory
- `is_slippery`: When True, agent movement becomes stochastic (harder)
- `max_workers`: Increase for faster execution on multi-core systems
### 3. Custom Configuration Examples
**Basic customization:**
```bash
python run_frozenlake.py --experiment-name "my_frozenlake_test" --max-workers 8
```
**Enable slippery mode:**
```bash
python run_frozenlake.py --slippery --experiment-name "frozenlake_slippery"
```
**Full training experiment:**
```bash
python run_frozenlake.py \
--enable-training \
--experiment-name "frozenlake_training_experiment" \
--max-workers 8 \
--training-runs 4 \
--num-training-maps 50 \
--test-runs 5 \
--num-test-maps 100 \
--slippery
```
**View all available options:**
```bash
python run_frozenlake.py --help
```
### 4. View Experiment Results
### 3. View Experiment Results
After running experiments, analyze the statistical results:
@ -128,30 +105,53 @@ python run_exp_statistic.py
- Processes all result files in `./exp_result/`
- Calculates success rates and performance metrics
- Generates a summary table showing performance comparisons
- Saves results to `experiment_summary.csv`
- Analyzes the effect of task memory on performance
- Saves results to `frozenlake_summary.csv`
## Configuration Parameters
## Understanding the Implementation
| Parameter | Default Value | Description |
|-----------|---------------|-------------|
| `--experiment-name` | `frozenlake_no_slippery` | Name of the experiment |
| `--max-workers` | `4` | Number of parallel workers |
| `--enable-training` | `False` | Enable training phase (experience generation) |
| `--training-runs` | `4` | Number of runs per training map |
| `--num-training-maps` | `50` | Number of training maps |
| `--test-runs` | `1` | Number of runs per test configuration |
| `--num-test-maps` | `100` | Number of test maps to use |
| `--slippery` | `False` | Enable slippery ice mode |
### Key Components
1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`)
- Implements a ReAct agent that interacts with the FrozenLake environment
- Handles task memory retrieval and storage
- Uses LLM (via OpenAI API) for decision making
2. **Experiment Runner** (`run_frozenlake.py`)
- Manages the overall experiment flow
- Handles training and testing phases
- Uses Ray for parallel execution
3. **Map Manager** (`map_manager.py`)
- Generates and manages test maps
- Ensures consistent evaluation across experiments
4. **Statistics Analyzer** (`run_exp_statistic.py`)
- Processes experiment results
- Calculates performance metrics
- Generates comparative analysis
## Understanding Results
The experiment evaluates agent performance on FrozenLake maps:
- **Success Rate**: Percentage of episodes that reach the goal
- **Default Mode**: Uses existing experience library for quick testing
- **Training Mode**: Generates new experiences then tests performance improvement
- **Success Rate**: Percentage of episodes where the agent reaches the goal
- **With vs. Without Memory**: Compares performance with and without task memory
- **Slippery vs. Non-slippery**: Compares performance in different environment dynamics
**Output Files:**
- `./exp_result/*.jsonl`: Raw experiment results
- `./exp_result/experiment_summary.csv`: Statistical summary
- Console output: Real-time progress and metrics
### Output Files
- `./exp_result/*_training.jsonl`: Results from training phase
- `./exp_result/*_test_no_memory.jsonl`: Test results without task memory
- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory
- `./exp_result/frozenlake_summary.csv`: Statistical summary
### Task Memory Mechanism
The task memory system works as follows:
1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service
2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map
3. **Memory Application**: The agent uses retrieved memories to guide its decision-making
The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake.