From d50b55ace0ff06019da3846954491a2ab6f5f653 Mon Sep 17 00:00:00 2001 From: "jinli.yl" Date: Mon, 1 Sep 2025 23:48:10 +0800 Subject: [PATCH] docs(cookbook): update quickstart guides for AppWorld and FrozenLake experiments - Update AppWorld quickstart guide to use ReMe instead of ExperienceMaker - Update FrozenLake quickstart guide to use ReMe and improve clarity - Refactor FrozenLake experiment implementation and documentation - Add more detailed explanation of task memory mechanism in FrozenLake --- cookbook/appworld/quickstart.md | 44 ++++---- cookbook/frozenlake/quickstart.md | 170 +++++++++++++++--------------- 2 files changed, 107 insertions(+), 107 deletions(-) diff --git a/cookbook/appworld/quickstart.md b/cookbook/appworld/quickstart.md index 29a6910b..dd759eb3 100644 --- a/cookbook/appworld/quickstart.md +++ b/cookbook/appworld/quickstart.md @@ -1,14 +1,14 @@ # AppWorld Experiment Quick Start Guide -This guide helps you quickly set up and run AppWorld experiments with ExperienceMaker integration. +This guide helps you quickly set up and run AppWorld experiments with ReMe integration. ## Env Setup ### 1. Clone the Repository ```bash -git clone https://github.com/modelscope/ExperienceMaker.git -cd ExperienceMaker/cookbook/appworld +git clone https://github.com/modelscope/ReMe.git +cd ReMe/cookbook/appworld ``` ### 2. Appworld Environment Setup @@ -36,43 +36,43 @@ appworld download data **Note**: The AppWorld data will be saved in the current directory. -### 3. Start ExperienceMaker Service +### 3. Start ReMe Service -Install ExperienceMaker (if not already installed) -If you haven't installed the ExperienceMaker environment yet, follow these steps: +Install ReMe (if not already installed) +If you haven't installed the ReMe environment yet, follow these steps: ```bash # Go back to the project root cd ../.. -# Create ExperienceMaker environment -conda create -p ./em-env python==3.12 -conda activate ./em-env +# Create ReMe environment +conda create -p ./reme-env python==3.12 +conda activate ./reme-env -# Install ExperienceMaker +# Install ReMe pip install . ``` -Launch the ExperienceMaker service to enable experience library functionality: +Launch the ReMe service to enable memory library functionality: ```bash -experiencemaker \ +reme \ http_service.port=8001 \ llm.default.model_name=qwen-max-latest \ embedding_model.default.model_name=text-embedding-v4 \ vector_store.default.backend=local_file ``` -add experiences for appworld: +add memories for appworld: ```bash curl -X POST "http://0.0.0.0:8001/vector_store" \ -H "Content-Type: application/json" \ -d '{ "workspace_id": "appworld_v1", "action": "dump", - "path": "./experience_library" + "path": "./memory_library" }' ``` -Now you have loaded the ExperienceMaker experience library to enable experience-based agent! +Now you have loaded the ReMe memory library to enable memory-based agent! ### 4. Common Issues @@ -84,9 +84,9 @@ Now you have loaded the ExperienceMaker experience library to enable experience- ## Run Experiments -### 1. Test: With Experience vs Without Experience +### 1. Test: With Memory vs Without Memory -Run the main experiment script to compare performance with and without experience: +Run the main experiment script to compare performance with and without memory: ```bash python run_appworld.py @@ -94,7 +94,7 @@ python run_appworld.py **What this does:** - Runs AppWorld tasks on the development dataset -- Compares agent performance with experience (`use_experience=True`) vs without experience +- Compares agent performance with ReMe memory (`use_memory=True`) vs without memory - Uses multiple workers for parallel processing - Runs each task multiple times for statistical significance - Results are automatically saved to `./exp_result/` directory @@ -102,7 +102,7 @@ python run_appworld.py **Configuration options in `run_appworld.py`:** - `max_workers`: Number of parallel workers (default: 6) - `num_runs`: Number of times each task is repeated (default: 4) -- `use_experience`: Whether to use ExperienceMaker experience library +- `use_memory`: Whether to use ReMe memory library ### 2. View Experiment Results @@ -131,10 +131,10 @@ python run_exp_statistic.py ## Understanding Results The experiment compares: -1. **Baseline**: Agent without experience library -2. **With Experience**: Agent enhanced with ExperienceMaker experience library +1. **Baseline**: Agent without memory library +2. **With Memory**: Agent enhanced with ReMe memory library Key metrics to look for: - **best@1**: Average performance across all single runs - **best@k**: Performance when taking the best of k attempts -- Improvement percentage when using experience vs baseline \ No newline at end of file +- Improvement percentage when using memory vs baseline \ No newline at end of file diff --git a/cookbook/frozenlake/quickstart.md b/cookbook/frozenlake/quickstart.md index 40408b1a..bed7afaa 100644 --- a/cookbook/frozenlake/quickstart.md +++ b/cookbook/frozenlake/quickstart.md @@ -1,14 +1,14 @@ # FrozenLake Experiment Quick Start Guide -This guide helps you quickly set up and run FrozenLake experiments with ExperienceMaker integration. +This guide helps you quickly set up and run FrozenLake experiments with ReMe integration. The FrozenLake experiment demonstrates how task memory can improve an agent's performance in a navigation task. -## Env Setup +## Environment Setup ### 1. Clone the Repository ```bash -git clone https://github.com/modelscope/ExperienceMaker.git -cd ExperienceMaker/cookbook/frozenlake +git clone https://github.com/modelscope/ReMe.git +cd ReMe/cookbook/frozenlake ``` ### 2. FrozenLake Environment Setup @@ -19,49 +19,47 @@ Install Gymnasium for FrozenLake environment: pip install gymnasium ``` -### 3. Start ExperienceMaker Service +This will install: +- gymnasium - for the FrozenLake environment +- ray - for parallel execution +- openai - for LLM API access +- other dependencies -Install ExperienceMaker (if not already installed) -If you haven't installed the ExperienceMaker environment yet, follow these steps: +### 3. Start ReMe Service + +If you haven't installed ReMe yet, follow these steps: ```bash # Go back to the project root cd ../.. -# Create ExperienceMaker environment -conda create -p ./em-env python==3.12 -conda activate ./em-env +# Create a virtual environment (optional) +conda create -p ./reme-env python==3.10 +conda activate ./reme-env -# Install ExperienceMaker +# Install ReMe pip install . ``` -Launch the ExperienceMaker service to enable experience library functionality: +Launch the ReMe service to enable memory library functionality: ```bash -experiencemaker \ - http_service.port=8001 \ - llm.default.model_name=qwen-max-latest \ +reme \ + backend=http \ + http.port=8002 \ + llm.default.model_name=qwen-max-2025-01-25 \ embedding_model.default.model_name=text-embedding-v4 \ - vector_store.default.backend=local_file + vector_store.default.backend=local ``` -Load default experience library for FrozenLake: +Load default memory library for FrozenLake: ```bash -curl -X POST "http://0.0.0.0:8001/vector_store" \ - -H "Content-Type: application/json" \ - -d '{ - "workspace_id": "frozenlake_no_slippery", - "action": "dump", - "path": "./experience_library" - }' ``` -Now you have loaded the default FrozenLake experience library to enable experience-based agent! ## Run Experiments ### 1. Quick Test: Performance Evaluation Only (Default) -Run the main experiment script to test agent performance using existing experience: +Run the main experiment script to test agent performance using existing memory: ```bash python run_frozenlake.py @@ -69,54 +67,33 @@ python run_frozenlake.py **What this does:** - Tests the agent on randomly generated FrozenLake maps -- Uses the default experience library (`frozenlake_no_slippery`) +- Uses the default memory library (`frozenlake_no_slippery`) - Evaluates performance with multiple runs for statistical significance - Results are automatically saved to `./exp_result/` directory -### 2. Advanced: Training + Testing (Experience Generation) +### 2. Advanced: Training + Testing (Memory Generation) -To create new experiences through training and then test performance: +To create new memories through training and then test performance: -```bash -python run_frozenlake.py --enable-training +You can modify the experiment parameters directly in the `run_frozenlake.py` file. The main parameters are in the `main()` function: + +```python +def main(): + experiment_name = "frozenlake_no_slippery" # Name of the experiment + max_workers = 4 # Number of parallel workers + training_runs = 4 # Runs per training map + num_training_maps = 50 # Number of maps for training + test_runs = 1 # Runs per test configuration + num_test_maps = 100 # Number of test maps + is_slippery = False # Enable slippery mode ``` -**What this does:** -- **Stage 1 (Training)**: Generates new experiences by solving training maps -- **Stage 2 (Testing)**: Evaluates performance using the generated experiences -- Compares baseline performance vs experience-enhanced performance +Key parameters to consider: +- `experiment_name`: Used as the workspace ID for task memory +- `is_slippery`: When True, agent movement becomes stochastic (harder) +- `max_workers`: Increase for faster execution on multi-core systems -### 3. Custom Configuration Examples - -**Basic customization:** -```bash -python run_frozenlake.py --experiment-name "my_frozenlake_test" --max-workers 8 -``` - -**Enable slippery mode:** -```bash -python run_frozenlake.py --slippery --experiment-name "frozenlake_slippery" -``` - -**Full training experiment:** -```bash -python run_frozenlake.py \ - --enable-training \ - --experiment-name "frozenlake_training_experiment" \ - --max-workers 8 \ - --training-runs 4 \ - --num-training-maps 50 \ - --test-runs 5 \ - --num-test-maps 100 \ - --slippery -``` - -**View all available options:** -```bash -python run_frozenlake.py --help -``` - -### 4. View Experiment Results +### 3. View Experiment Results After running experiments, analyze the statistical results: @@ -128,30 +105,53 @@ python run_exp_statistic.py - Processes all result files in `./exp_result/` - Calculates success rates and performance metrics - Generates a summary table showing performance comparisons -- Saves results to `experiment_summary.csv` +- Analyzes the effect of task memory on performance +- Saves results to `frozenlake_summary.csv` -## Configuration Parameters +## Understanding the Implementation -| Parameter | Default Value | Description | -|-----------|---------------|-------------| -| `--experiment-name` | `frozenlake_no_slippery` | Name of the experiment | -| `--max-workers` | `4` | Number of parallel workers | -| `--enable-training` | `False` | Enable training phase (experience generation) | -| `--training-runs` | `4` | Number of runs per training map | -| `--num-training-maps` | `50` | Number of training maps | -| `--test-runs` | `1` | Number of runs per test configuration | -| `--num-test-maps` | `100` | Number of test maps to use | -| `--slippery` | `False` | Enable slippery ice mode | +### Key Components + +1. **FrozenLakeReactAgent** (`frozenlake_react_agent.py`) + - Implements a ReAct agent that interacts with the FrozenLake environment + - Handles task memory retrieval and storage + - Uses LLM (via OpenAI API) for decision making + +2. **Experiment Runner** (`run_frozenlake.py`) + - Manages the overall experiment flow + - Handles training and testing phases + - Uses Ray for parallel execution + +3. **Map Manager** (`map_manager.py`) + - Generates and manages test maps + - Ensures consistent evaluation across experiments + +4. **Statistics Analyzer** (`run_exp_statistic.py`) + - Processes experiment results + - Calculates performance metrics + - Generates comparative analysis ## Understanding Results The experiment evaluates agent performance on FrozenLake maps: -- **Success Rate**: Percentage of episodes that reach the goal -- **Default Mode**: Uses existing experience library for quick testing -- **Training Mode**: Generates new experiences then tests performance improvement +- **Success Rate**: Percentage of episodes where the agent reaches the goal +- **With vs. Without Memory**: Compares performance with and without task memory +- **Slippery vs. Non-slippery**: Compares performance in different environment dynamics -**Output Files:** -- `./exp_result/*.jsonl`: Raw experiment results -- `./exp_result/experiment_summary.csv`: Statistical summary -- Console output: Real-time progress and metrics \ No newline at end of file +### Output Files + +- `./exp_result/*_training.jsonl`: Results from training phase +- `./exp_result/*_test_no_memory.jsonl`: Test results without task memory +- `./exp_result/*_test_with_memory.jsonl`: Test results with task memory +- `./exp_result/frozenlake_summary.csv`: Statistical summary + +### Task Memory Mechanism + +The task memory system works as follows: + +1. **Memory Creation**: During training, successful trajectories are sent to the ReMe service +2. **Memory Retrieval**: During testing, the agent queries relevant memories based on the current map +3. **Memory Application**: The agent uses retrieved memories to guide its decision-making + +The experiment demonstrates how task memory can significantly improve performance, especially in challenging environments like the slippery FrozenLake. \ No newline at end of file