ReMe/tests/test_reme_memory_error_handling.py
Zhouwk e0d0e3e568
Some checks are pending
Pre-commit / run (ubuntu-latest) (push) Waiting to run
提供支持向量数据库的profile功能 (#221)
* feat(reme): 添加配置选项以启用或禁用个人资料功能

- 在 ReMe 初始化方法中添加 enable_profile 参数,默认值为 True
- 根据 enable_profile 设置决定是否创建 profile 目录和设置 profile_dir
- 在 PersonalSummarizer 中根据 enable_profile 条件性地添加个人资料相关工具
- 在 PersonalRetriever 中根据 enable_profile 条件性地添加 ReadAllProfiles 工具
- 修改 profile_path 属性以在禁用个人资料时返回 None
- 修改 get_profile_handler 方法以在禁用个人资料时返回 None
- 为 enable_profile 参数添加文档说明其用于云向量存储场景

* refactor(benchmark): 重构LongMemEval基准测试中的ReMe实例管理

- 移除未使用的shutil导入
- 将固定的ReMe实例改为每个问题创建独立实例以实现隔离
- 更新LLM配置名称从qwen3-max-think到qwen-max-t
- 修改模型调用逻辑使用正确的model_name参数
- 添加qwen-flash和GPT-4o-mini等新模型配置
- 统一使用"User"作为用户名,通过集合名实现隔离
- 调整并发处理数从4降至1,批处理大小从10增至30
- 每个问题类型采样数从2增至4
- 添加异步上下文管理确保资源正确释放

* reformat 2 files

* refactor(benchmark): 重构长记忆评估中的模型配置

- 将原有的 eval_model_name 替换为专门的 retrieve_model_name 用于检索操作
- 添加对 qwen-max 模型配置的支持
- 更新参数解析器以支持新的检索模型参数
- 修改最大并发数默认值从 1 提升到 4
- 调整样本数量默认值从 4 减少到 1
- 统一模型参数命名规范,区分摘要、检索和评估模型
- 优化内存处理器初始化逻辑,支持独立的检索模型配置

* fix(benchmark): 移除数据路径默认值并设为必填参数

- 将LongMemEval评估脚本中的data_path参数改为必需参数
- 将HaluMem评估脚本中的data_path参数改为必需参数
- 删除了硬编码的默认文件路径配置
- 强制用户显式指定数据集文件路径以避免路径错误

* Update __init__.py

* Update __init__.py

* fix(benchmark): 修复ReMe评估中的模型配置和空值处理问题

- 移除了retrieve_memory调用中不需要的llm_config_name参数
- 修复了长字符串打印的换行格式问题
- 添加了eval_result为空时的初始化处理
- 在accuracy评估中加入了eval_model_name参数传递

* style(benchmark): 格式化模型名称打印输出

- 移除了多行字符串中的换行符和多余空格
- 将模型名称信息合并为单行连续显示
- 保持了原有的打印格式和信息完整性

* docs(readme): 更新文档添加实验结果表格

- 在英文版 README 中添加 🧪 Experiments 章节
- 添加 LoCoMo 和 HaluMem 两个基准测试的结果表格
- 在中文版 README_ZH 中添加 🧪 实验 章节
- 添加 LoCoMo 和 HaluMem 测试集的实验配置说明
- 添加完整的实验数据对比表格和评估协议说明

* docs(readme): 更新文档中的内存系统链接

- 为基于文件的记忆系统添加锚点链接
- 为基于向量库的记忆系统添加锚点链接
- 修复英文文档中的链接格式
- 修复中文文档中的链接格式和空行问题

* docs(readme): update experimental results section in documentation

- Remove outdated experimental data placeholder "Coming soon..."
- Add complete evaluation results for LoCoMo and HaluMem benchmarks
- Include detailed performance metrics tables for all memory methods
- Update experimental settings description with ReMe backbone details
- Align evaluation protocol information with LLM-as-a-Judge approach
- Maintain consistent formatting between English and Chinese documentation

* docs(benchmark): add quick start guides for halumem and longmemeval experiments

- Created HaluMem experiment quick start guide with ReMe integration setup
- Added detailed steps for installing ReMe environment using conda
- Included repository cloning instructions for HaluMem benchmark
- Provided complete command examples for running HaluMem experiments
- Created LongMeMEval quick start guide with data download procedures
- Added wget commands for downloading cleaned dataset files
- Included evaluation script instructions for computing experiment statistics
- Documented parameter configurations for different model types and batch sizes

* docs(longmemeval): update quickstart guide documentation

- Changed project name from Halumem to Longmemeval in title
- Updated description to reference Longmemeval experiments instead of Halumem
- Maintained existing ReMe integration instructions unchanged

* chore(logger): add test comment to logger configuration

- Added test comment in logger utility function
- Removed duplicate log handling by keeping the remove() call

* chore(logger): add test comment to logger configuration

- Added test comment in logger utility function
- Removed duplicate log handling by keeping the remove() call

* feat(core): add file logging capability to application

- Added log_to_file parameter to Application class constructor
- Integrated log_to_file option in logger initialization
- Updated ServiceContext to support file logging configuration
- Modified init_logger function to conditionally enable file logging
- Added log_to_file field to ServiceConfig schema
- Updated ReMe class to include file logging option
- Wrapped file logging setup in conditional check to prevent unnecessary operations

* docs(benchmark): update HaluMem quickstart guide with dataset download instructions

- Replace repository cloning with direct dataset download using curl
- Add commands to download HaluMem-Medium.jsonl and HaluMem-Long.jsonl files
- Include both official Hugging Face and mirror download sources
- Update data path reference from nested directory to local data folder
- Add dataset page link and mirror usage instructions for mainland China access

* feat(memory): add profile retrieval tool and refactor profile management

- Introduce RetrieveProfile tool for fetching specific user profiles
- Refactor ProfileHandler to support both filesystem and vector backends
- Add async methods to ProfileHandler with synchronous fallbacks
- Update PersonalRetriever to support two-stage profile and memory retrieval
- Enhance PersonalSummarizer with improved tool partitioning logic
- Add profile_backend, profile_store_name, and profile_max_capacity configuration options
- Replace direct ProfileHandler imports with get_profile_handler method
- Implement profile search functionality with dedicated prompts and workflows
- Add FileProfileBackend and VectorProfileBackend implementations
- Update base memory tool with new profile configuration parameters

* feat(profile): add custom profile collection name support

- Add profile_collection_name parameter to Application constructor
- Allow custom database collection name for vector profiles instead of default suffix
- Update profile vector store configuration logic to use custom collection name
- Modify _ensure_profile_vector_store_config to handle custom collection names
- Update docstring with detailed parameter descriptions for profile configuration options

* test(history): add single history id acceptance test for multiple mode

- Add test case to verify multiple-mode history lookup accepts a single history_id string
- Create FakeVectorStore stub with minimal implementation for ReadHistory tests
- Return requested history node from vector store mock
- Initialize ReadHistory tool with multiple mode enabled
- Add pylint disable comment for protected access to vector store property

* refactor(memory): update profile handler and vector tools with improved formatting and error handling

- Add module docstring to profiles/__init__.py
- Add pylint disable comments for no-name-in-module and missing-function-docstring
- Format long error message in ProfileHandler.sync_run method for better readability
- Reformat parameters in ProfileHandler.aadd method to separate lines
- Update model_copy call in reme.py to span multiple lines for better readability
- Format aadd_batch call in update_profile.py to span multiple lines
2026-04-28 15:11:45 +08:00

168 lines
6.3 KiB
Python

"""Tests for ReMe memory error handling and raise_exception propagation."""
from types import SimpleNamespace
import pytest
import reme.reme as reme_module
from reme.core.runtime_context import RuntimeContext
from reme.core.schema import MemoryNode
from reme.memory.vector_tools.history.read_history import ReadHistory
from reme.reme import ReMe
class Recorder:
"""Stub that records constructor args for later inspection."""
instances = []
def __init__(self, *args, **kwargs):
self.args = args
self.kwargs = kwargs
self.call_kwargs = None
Recorder.instances.append(self)
class TopLevelAgent(Recorder):
"""Stub agent that returns a successful structured result."""
async def call(self, **kwargs):
"""Simulate a successful agent call."""
self.call_kwargs = kwargs
return {"answer": "ok", "success": True}
def _make_reme() -> ReMe:
"""Create a ReMe instance with startup bypassed for unit testing."""
reme = ReMe(enable_logo=False, log_to_console=False, enable_profile=False)
reme._started = True # pylint: disable=protected-access
return reme
def _patch_summarize_dependencies(monkeypatch: pytest.MonkeyPatch) -> None:
"""Patch all summarize-path dependencies with stubs."""
Recorder.instances = []
monkeypatch.setattr(reme_module, "AddDraftAndRetrieveSimilarMemory", Recorder)
monkeypatch.setattr(reme_module, "AddMemory", Recorder)
monkeypatch.setattr(reme_module, "AddHistory", Recorder)
monkeypatch.setattr(reme_module, "DelegateTask", Recorder)
monkeypatch.setattr(reme_module, "PersonalSummarizer", Recorder)
monkeypatch.setattr(reme_module, "ProceduralSummarizer", Recorder)
monkeypatch.setattr(reme_module, "ToolSummarizer", Recorder)
monkeypatch.setattr(reme_module, "ReMeSummarizer", TopLevelAgent)
def _patch_retrieve_dependencies(monkeypatch: pytest.MonkeyPatch) -> None:
"""Patch all retrieve-path dependencies with stubs."""
Recorder.instances = []
monkeypatch.setattr(reme_module, "RetrieveMemory", Recorder)
monkeypatch.setattr(reme_module, "ReadHistory", Recorder)
monkeypatch.setattr(reme_module, "DelegateTask", Recorder)
monkeypatch.setattr(reme_module, "PersonalRetriever", Recorder)
monkeypatch.setattr(reme_module, "ProceduralRetriever", Recorder)
monkeypatch.setattr(reme_module, "ToolRetriever", Recorder)
monkeypatch.setattr(reme_module, "ReMeRetriever", TopLevelAgent)
@pytest.mark.asyncio
@pytest.mark.parametrize("raise_exception", [False, True])
async def test_summarize_memory_propagates_raise_exception(
monkeypatch: pytest.MonkeyPatch,
raise_exception: bool,
):
"""Verify raise_exception is forwarded to every sub-agent in summarize."""
_patch_summarize_dependencies(monkeypatch)
reme = _make_reme()
result = await reme.summarize_memory(
messages=[{"role": "user", "content": "hi", "time_created": "2026-03-20 10:00:00"}],
task_name="demo-task",
raise_exception=raise_exception,
)
assert result == "ok"
assert Recorder.instances
assert all(instance.kwargs.get("raise_exception") is raise_exception for instance in Recorder.instances)
@pytest.mark.asyncio
@pytest.mark.parametrize("raise_exception", [False, True])
async def test_retrieve_memory_propagates_raise_exception(
monkeypatch: pytest.MonkeyPatch,
raise_exception: bool,
):
"""Verify raise_exception is forwarded to every sub-agent in retrieve."""
_patch_retrieve_dependencies(monkeypatch)
reme = _make_reme()
result = await reme.retrieve_memory(
query="hello",
task_name="demo-task",
raise_exception=raise_exception,
)
assert result == "ok"
assert Recorder.instances
assert all(instance.kwargs.get("raise_exception") is raise_exception for instance in Recorder.instances)
@pytest.mark.asyncio
async def test_summarize_memory_raises_runtime_error_for_unstructured_result(monkeypatch: pytest.MonkeyPatch):
"""Verify RuntimeError is raised when the top-level summarizer returns a plain string."""
Recorder.instances = []
reme = _make_reme()
monkeypatch.setattr(reme_module, "PersonalSummarizer", Recorder)
monkeypatch.setattr(reme_module, "ProceduralSummarizer", Recorder)
monkeypatch.setattr(reme_module, "ToolSummarizer", Recorder)
monkeypatch.setattr(reme_module, "AddDraftAndRetrieveSimilarMemory", Recorder)
monkeypatch.setattr(reme_module, "AddMemory", Recorder)
monkeypatch.setattr(reme_module, "AddHistory", Recorder)
monkeypatch.setattr(reme_module, "DelegateTask", Recorder)
class FailingTopLevelAgent(Recorder):
"""Stub agent that returns a failure string instead of a dict."""
async def call(self, **kwargs):
"""Simulate a failed agent call returning a plain error string."""
self.call_kwargs = kwargs
return "[ReMeSummarizer] failed: boom"
monkeypatch.setattr(reme_module, "ReMeSummarizer", FailingTopLevelAgent)
with pytest.raises(RuntimeError, match="summarize_memory failed before producing a structured result"):
await reme.summarize_memory(
messages=[{"role": "user", "content": "hi", "time_created": "2026-03-20 10:00:00"}],
task_name="demo-task",
)
@pytest.mark.asyncio
async def test_read_history_accepts_single_history_id_in_multiple_mode():
"""Verify multiple-mode history lookup accepts a single history_id string."""
class FakeVectorStore:
"""Minimal vector store stub for ReadHistory tests."""
async def get(self, vector_ids):
"""Return the requested history node."""
assert vector_ids == ["history_123"]
node = MemoryNode(
memory_id="history_123",
memory_type="history",
memory_target="alice",
content="Alice said hello.",
)
return [node.to_vector_node()]
tool = ReadHistory(enable_multiple=True)
tool._vector_store = FakeVectorStore() # pylint: disable=protected-access
tool.context = RuntimeContext(
history_id="history_123",
retrieved_nodes=[],
service_context=SimpleNamespace(memory_target_type_mapping={"alice": "personal"}),
)
result = await tool.execute()
assert "Historical Dialogue[history_123]" in result