mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-10-07 03:00:27 +00:00
10 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
06fb46fa48
|
feat(tags): add configurable tag indexing and filtered hybrid search (#530)
* Add optional tag generation and normalization to auto memory * Add tag index components and clean up temporary JSONL files * Preserve tag index state when reconciliation fails * Refactor and streamline application implementation * Fix pylint C1803 warnings in tag normalization tests * Document optional tag index configuration * Make tag index failures non-blocking and disable auto-memory tags * Add configurable tag indexing and tag listing * Add tag-filtered hybrid search with exact candidate ranking * Remove obsolete generated files * Rename tag index key to tag_key and reject reserved fields * Extract automatic tagging into a dedicated step * Restrict frontmatter updates to authorized keys * Refine tag filtering and automatic memory tagging * Require underscore-separated tags in auto-tag prompts - Forbid spaces in tags and require underscores (e.g. sam_altman) in both English and Chinese auto_tag prompts, with English examples switched to English entities (OpenAI, gold) - Drop prompt-string assertions superseded by the new tagging rule - Merge construction/runtime tag_key validation tests into one parametrized case * Make max_tags_per_file configurable in auto-tag step * Consolidate tag index tests * Fix search test fixture lint warnings * Align tag contracts and index health behavior * Fall back when tag index is unavailable --------- Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com> |
||
|
|
3f2eb6235f
|
feat(benchmark): extract BEAM and LongMemEval into standalone plugins (#512)
* Simplify project implementation * Centralize benchmark agentic answer base class * Remove bundled ReMe source snapshots * Preserve local benchmark configs and plugin discovery behavior * docs: enrich job parameter descriptions in beam and lme plugin configs * Simplify project structure and remove obsolete code * Move benchmark search step configuration into BEAM and LME plugins * Rename benchmark judge packages to avoid import collisions * Remove explicit plugin package loading in favor of entry-point discovery * Export benchmark plugin Steps from public packages |
||
|
|
fc4a5398a8
|
fix(index): make embedding rebuild explicit and scoped (#508)
* fix(index): make embedding rebuild explicit and scoped * fix(index): harden scoped reindex completion * fix(index): guard embedding space transitions * fix(index): serialize reindex with mutations * docs(index): clarify scoped reindex semantics * docs(index): explain synchronous checkpoint snapshots * fix(index): serialize checkpoint publication |
||
|
|
9218a2d0e3
|
refactor: derive dialog paths from session_dir (#421)
* refactor: derive dialog paths from session directory * fix: normalize configured session paths * fix: align dialog watch paths with writers * fix: reject absolute session directories |
||
|
|
eac8223387
|
feat: add frontend-ready wikilink graph APIs (#414) | ||
|
|
4eb2adf961
|
feat(faiss_file_store): upgrade FAISS to HNSW index with async reindex (#390)
* feat(file_store): upgrade FAISS to HNSW index with async reindex and path constraint - Replace IndexFlatIP with IndexHNSWFlat for better recall/speed tradeoff - Add dynamic efSearch (limit * 5) scaled to query request size - Add async_reindex option: background rebuild with generation-based invalidation - Extract _delete_nodes() in LocalFileStore for subclass reuse - Add unit tests for file store consistency * fix: resolve pylint warnings in faiss store and test file * refactor(file_store): replace generation-based reindex with event-flag worker - Replace _reindex_generation/lock/task with a single long-lived worker coroutine consuming an asyncio.Event flag; repeated submissions coalesce - Use local index reference in vector_search to avoid TOCTOU on self._faiss_index - Pass index explicitly to _set_ef_search for consistency - Track _index_writes to re-arm reindex after concurrent writes - Update tests to match new internal API * fix: resolve pylint too-many-return-statements and implicit-booleaness warnings * feat(file_store): add refine maintenance hook and incremental embedding backfill - Add refine() idle-time maintenance hook to BaseFileStore/LocalFileStore - FaissLocalFileStore: incremental vector add on backfill instead of full rebuild - Dynamic tombstone compaction threshold scaled by index size - Add RefineStoreStep with daily cron job (refine_store_cron) - Enable faiss backend and embedding_store by default in default.yaml - Add unit tests for faiss index maintenance * chore(deps): promote faiss-cpu to core dependencies faiss backend is now the default file_store, so faiss-cpu moves from the optional [core] extra to the base dependencies list. * feat: rename refine_store to optimize_index and add vecdb_path_constraint - Rename refine_store step to optimize_index with cron job scheduling - Add vecdb_path_constraint to file_store components - Update default.yaml with optimize_index_cron and faiss backend comment - Update memory_search docs (en/zh) for FAISS vector management - Update unit tests for index maintenance * feat(faiss): add embedding digest to reject stale sidecar after partial dump Add _chunks_embedding_digest() that computes an order-independent SHA-256 over (chunk_id, float16 embedding) pairs. The digest is written into the idmap sidecar at dump time and verified at load time. A mismatch means the sidecar vectors belong to a different chunk generation than the authoritative JSONL — detectable even when the live-ID set is unchanged (same-ID in-place update crash window). Add test_faiss_rejects_stale_sidecar_after_partial_dump reproducing the crash-between-writes scenario and asserting digest-based rejection. Compress verbose docstrings/comments in existing tests for pylint line budget. --------- Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135> |
||
|
|
630f26b119
|
feat(search): scoped dedup, session-chunk merge, and unified recall formatting (#384)
* feat(search): add tool_context-scoped chunk dedup with TTL
Introduce _ToolContextDedupMixin shared by search/vector_search/bm25_search
to skip already-seen chunks within one agent tool_context. Per-context state
lives in app_context.metadata with configurable TTL (default 24h).
* feat(search): unify chunk answer rendering with merge and explicit empty messages
- Refactor SearchStep/VectorSearchStep/Bm25SearchStep to share format_chunks_answer for consistent source rendering and adjacent session-chunk merging.
- Distinguish empty results: ALL_RETURNED_MESSAGE when dedup removes everything vs NO_RESULTS_MESSAGE when nothing matched.
- Bump JsonlFileChunker default max_chars to 4000.
- Add unit tests for source-format merge and empty-result messages.
* refactor(config): reorganize file_chunker components and move jsonl max_chars into config
- Register explicit markdown/json/jsonl chunkers in beam.yaml and lme.yaml with markdown options (embed_toc, max_ast_sections, frontmatter handling) and jsonl max_chars=4000.
- Restrict default chunker to txt/log extensions.
- Revert JsonlFileChunker code default max_chars back to 2000; the 4000 value now lives in config.
* chore(benchmark): increase longmemeval num_items from 64 to 500
* refactor(search): split SearchStep into simplified and v2 variants, extract counter utility
- Extract global_counter_next from ApplicationContext into reme/utils/counter.py
as a standalone function operating on metadata dict with lazy initialization.
- Split SearchStep into two variants:
- SearchStep (simplified): inline chunk.id dedup, single-branch vector/keyword
optimization based on vector_weight, inline answer formatting.
- SearchV2Step (full): preserves _ToolContextDedupMixin with interval-subset-aware
dedup and format_chunks_answer with session-aware chunk merging.
- Update beam.yaml and lme.yaml to use search_v2_step for benchmark jobs.
- Rename existing search tests to test_search_v2_step_* and add new
test_search_step_* tests covering the simplified variant.
* fix: normalise missing trailing newline in _build_union_chunk to prevent line collision
* refactor: lazy-init counter tree in ApplicationContext metadata
- Remove hardcoded _counter_tree and _counter_tree_lock initialization
from ApplicationContext.metadata; rely on lazy initialization in
reme.utils.counter.global_counter_next on first call
- Set longmemeval num_items back to 500
- Remove obsolete trailing-newline collision tests
---------
Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
|
||
|
|
bf7ca17705
|
feat(benchmark): add LongMemEval golden answer validation (#335)
* feat(benchmark): add golden answer validation and session review for LongMemEval - Introduce GoldenCheckStep to validate LongMemEval golden answers using structured verdicts - Add SessionReviewStep to extract query/answer-relevant evidence from all sessions - Implement concurrent session processing with configurable concurrency limits - Create check_golden job configuration with lme_review and lme_judge agent wrappers - Add Qwen3.7-plus model configuration for enhanced processing capabilities - Include python_execute tool integration for agent-based reasoning and date validation - Generate comprehensive JSON output with session summaries and validation verdicts - Add run_check_golden.py script for batch processing across all LongMemEval samples - Configure proper logging initialization with console and file output options - Update component registry and file I/O modules to support new benchmark features * feat(scripts): add script to summarize LongMemEval check_golden verdicts - Parse check_golden.json files across all LongMemEval samples - Calculate accuracy metrics for golden answers and session IDs - Provide breakdown by question type with percentage calculations - Add command line options for listing bad samples and JSON output - Include progress tracking showing completed vs pending samples - Display confidence scores and date sanity checks statistics * refactor(benchmark): move golden check scripts to longmemeval directory - Moved run_check_golden.py from scripts/ to benchmark/longmemeval/ - Moved stats_check_golden.py from scripts/ to benchmark/longmemeval/ - Updated path resolution to use parents[2] instead of parent.parent - Added new --list-run-failed option to stats script - Added logging directory constant and functions for tracking launched samples - Enhanced stats output with launched count and run failure information - Improved error reporting with run failure details and log file paths * feat(benchmark): add LongMemEval agentic answer workflow with session extraction - Add LmeAgenticAnswerStep, LmeAutoMemoryStep, and LmeExtractSessionStep to __init__.py - Create shared helper render_with_source for displaying search results with session_id - Implement agentic_answer step with vector_search, bm25_search, and extract_session_by_id tools - Add auto_memory step to convert each session into search-friendly daily notes - Create extract_session step to retrieve and analyze raw session content by session_id - Update jinli_lme.yaml with auto_memory, vector_search, bm25_search, and agentic_answer jobs - Configure lme_memory, lme_extract, and lme_agentic_answer agent wrappers - Enhance search steps with include_source option to show session_id metadata - Add proper session_id tracking and collision handling in daily note generation * feat(benchmark): add LongMemEval agentic answer evaluation pipeline - Add session_id tracking to agentic_answer.py result metadata - Introduce run_agentic_answer.py driver for complete pipeline execution - Implement auto_memory, update_index, and agentic_answer job orchestration - Add concurrent execution with configurable limits and staggering - Create aggregation script for collecting tool-call trails and results - Add stats_agentic_answer.py for comprehensive result analysis - Implement resume capability with existing output detection - Generate aggregate.json with per-sample breakdown and tool call summaries * feat(steps): add ClearPathsStep for cleaning workspace outputs before rebuild - Introduce ClearPathsStep to remove stale workspace files/directories - Add support for specifying paths and config_keys as targets to clear - Implement safety checks to prevent deletion of files outside workspace - Add logging for cleared paths and warnings for invalid paths - Configure clear_paths_step in jinli_lme.yaml to clean daily_dir - Add clear_paths_step to clean mem_answer.json before rebuilds * feat(benchmark): add resume functionality to agentic answer runner - Replace --force flag with --resume flag for controlling job execution - By default every job reruns with clean rebuild behavior using config clear steps - Add --resume option to skip samples whose output already exists and continue interrupted batches - Update documentation to reflect new default clean rebuild behavior - Modify job skipping logic to honor resume flag instead of force flag - Update dry-run output to show correct todo jobs based on resume status - Change default example command to use --resume for continuing interrupted runs * feat(benchmark): generate JSONL output for check golden records - Add write_check_golden_list function to create JSONL file - Write all readable check_golden records as JSONL format - Include check_golden_list path in stats output - Display generated JSONL file path in summary report - Maintain UTF-8 encoding with non-ASCII character support * refactor(benchmark): rename answer judge step and integrate LME LLM judge - Rename AnswerJudgeStep to LmeLlmJudgeStep and update imports - Add new llm_judge configuration in jinli_lme.yaml - Update run_agentic_answer.py to include llm_judge in pipeline - Modify LmeLlmJudgeStep to read from query.json and answer.json - Write LLM judgement results back to mem_answer.json - Add command line options for start/end sample range selection - Update aggregate.json generation to include LLM judgement data - Add resume capability for llm_judge job based on judgement presence * refactor(benchmark): rename answer judge step and integrate LME LLM judge - Rename AnswerJudgeStep to LmeLlmJudgeStep and update imports - Add new llm_judge configuration in jinli_lme.yaml - Update run_agentic_answer.py to include llm_judge in pipeline - Modify LmeLlmJudgeStep to read from query.json and answer.json - Write LLM judgement results back to mem_answer.json - Add command line options for start/end sample range selection - Update aggregate.json generation to include LLM judgement data - Add resume capability for llm_judge job based on judgement presence * feat(steps): add wait_for_paths_step to block until workspace files exist - Introduce WaitForPathsStep class that polls for required workspace-relative paths - Add step registration with 'wait_for_paths_step' backend identifier - Implement path validation to ensure targets are within workspace boundaries - Add polling mechanism with configurable intervals via poll_seconds parameter - Include logging functionality with log_every_seconds parameter for status updates - Add metadata tracking of waited paths and duration in response object - Register step in index module and expose in public API - Configure step in jinli_lme.yaml to wait for session_review.json before golden check - Add script rename from run_check_golden.py to run_golden_check.py with enhanced options * feat(benchmark): enhance longmemeval benchmarking with concurrency and progress tracking - Add benchmark extra dependency group with portalocker requirement - Introduce concurrent execution support for golden_check and session_review workflows - Add progress reporting interval option with real-time status updates - Implement global throttling mechanism for session review requests using file locks - Enhance golden check validation with current schema verification - Add active task tracking and graceful shutdown handling - Rename check_golden scripts to golden_check for consistency - Update statistics reporting with correct/incorrect terminology instead of reasonable - Add stale format detection and compatibility handling for verdict fields - Include both_correct rate calculation in accuracy metrics - Add concurrency and staggering options for better resource management * ci(workflow): add Windows smoke test workflow - Create new workflow file .github/workflows/windows-smoke.yml - Configure workflow to trigger on push and pull request events - Set up Python environment with version 3.11 - Install package dependencies using pip - Run version job as smoke test for CLI functionality - Enable concurrency control to prevent duplicate runs - Use matrix strategy for Python version testing * feat(benchmark): add retry mechanism and health check for session review - Added retry configuration options (retry_initial_seconds, retry_max_seconds, retry_max_attempts) to jinli_lme.yaml - Implemented exponential backoff retry logic with configurable parameters in session_review step - Added output_is_healthy function to verify session_review.json integrity and absence of failed reviews - Updated resume functionality to skip only healthy outputs instead of all existing files - Integrated JSON parsing and validation to check for failed reviews in output files - Enhanced error handling and logging for retry attempts and recovery scenarios * feat(benchmark): add LongMemEval session review statistics script - Create stats_session_review.py to summarize session_review.json artifacts - Add command line options for listing failed, missing, and run failed samples - Implement JSON output mode for programmatic consumption - Calculate and display health statistics including total samples, healthy outputs, failed sessions - Provide detailed failure information with session IDs and error messages - Generate re-run commands for samples with failed reviews - Add percentage calculations for better statistical overview - Include support for multiple output formats and detailed logging * feat(benchmark): add LongMemEval output cleanup script and enhance golden check retry logic - Added clean_sample_outputs.py script to remove generated LongMemEval files while preserving source inputs - Implemented configurable retry mechanism in golden_check.py with exponential backoff strategy - Added retry parameters (initial/max seconds and max attempts) to control failure recovery behavior - Integrated asyncio support for asynchronous sleep during retry intervals - Configured default retry settings in jinli_lme.yaml with 5s initial and 300s maximum intervals - Preserved core files (query.json, answer.json, session/) while cleaning generated artifacts * feat(benchmark): add AppleDouble file cleanup to sample output cleaner - Remove AppleDouble files starting with '._' recursively including under session/ - Add is_under helper function to check if path is inside parent directory - Track targets in set to avoid duplicate processing - Include AppleDouble files in cleanup targets when not already covered by existing targets - Maintain dry-run mode as default behavior with --apply flag for actual deletion * refactor(benchmark): update LongMemEval sample output cleaning script - Add time and Iterator imports for enhanced functionality - Add --progress-every argument to control progress reporting frequency - Replace is_under function with iter_sample_targets generator - Implement detailed progress tracking with timing measurements - Add sample-by-sample processing with elapsed time reporting - Include AppleDouble file detection within session directory - Update target counting and deletion statistics display - Add conditional progress updates based on progress-every setting - Improve dry-run mode with would-delete indication * chore(benchmark): increase initial interval for session review step - Changed START_INTERVAL_SECONDS from 1.0 to 3.0 seconds - Adjusted timing parameters for better benchmark stability * refactor(benchmark): implement coordinated retry mechanism for session reviews - Add retry gate condition to coordinate concurrent review attempts - Implement wait_for_healthy_start_slot to handle sequential retries - Create mark_retrying and mark_recovered functions to track retry states - Update reply_with_retry to accept index parameter for coordination - Add has_prior_retry logic to prevent race conditions during recovery - Ensure proper cleanup of retry state on success or failure - Maintain backward compatibility while adding coordination features * chore(benchmark): adjust session review start interval timeout - Changed START_INTERVAL_SECONDS from 3.0 to 5.0 seconds - Increased initial delay for session review benchmark step - Updated timeout configuration for improved stability * refactor(benchmark): update session review concurrency and throttling mechanism - Replace global throttle with per-process concurrency control - Add concurrency parameter with default value of 30 in config - Add start_interval_seconds parameter with default value of 2 seconds - Change default concurrency from 3 to 1 in command line interface - Update documentation to reflect new throttling behavior - Implement semaphore-based concurrency limiting for review tasks - Modify retry mechanism to use local locking instead of global files - Remove portalocker dependency for cross-process throttling * refactor(config): update session review configuration and concurrency settings - Removed deprecated retry configuration parameters from jinli_lme.yaml - Increased MAX_CONCURRENCY from 30 to 60 in session_review.py - Reduced START_INTERVAL_SECONDS from 2.0 to 1.0 in session_review.py - Cleaned up redundant backend specifications in configuration file - Simplified agent wrapper configurations by removing obsolete retry settings * feat(benchmark): enhance LME auto memory step with advanced scheduling and error handling - Add datetime parsing functionality for LongMemEval timestamps with regex pattern - Implement configurable concurrency limits with MAX_CONCURRENCY of 60 - Introduce retry mechanism with exponential backoff for agent interactions - Add session filtering based on date comparison with question_date validation - Create rate limiting with start interval control between requests - Implement sophisticated retry coordination using asyncio conditions - Add comprehensive error tracking for failed and filtered session extracts - Remove deprecated concurrency parameter from jinli_lme.yaml configuration - Add structured output validation in session review step - Include detailed metadata reporting with session statistics and errors * fix(benchmark): adjust default concurrency for auto_memory job - Changed default concurrency from 3 to 1 for auto_memory job to prevent API overload - Updated help text to reflect new default value of 1 for concurrency parameter - Modified documentation to clarify concurrency behavior varies by job type * refactor(search): replace hardcoded candidate multiplier with constant - Introduced _CANDIDATE_MULTIPLIER constant set to 10 - Replaced hardcoded factor of 5 with _CANDIDATE_MULTIPLIER in BM25 search - Replaced hardcoded factor of 5 with _CANDIDATE_MULTIPLIER in vector search - Updated test to verify both search steps use ten times limit for candidates - Imported VectorSearchStep and Bm25SearchStep in test module - Added comprehensive test case for candidate count calculation logic * feat(lme): add data inspection error handling with fallback mechanism - Implemented non-retryable data inspection error markers detection - Added _is_data_inspection_error method to identify inspection failures - Created fallback handling for data inspection errors in auto memory extraction - Added fallback handling for data inspection errors in session review - Extended failed extracts tracking with non-retryable and fallback flags - Separated fallback extracts from regular failed extracts in reporting - Enhanced error logging with specific data inspection failure messages - Updated metrics to track fallback extractions and reviews separately - Maintained existing retry logic for other exception types * feat(benchmark): enhance session review statistics with fallback tracking - Add support for identifying and listing non-retryable fallback reviews - Introduce --list-fallback argument to display fallback review details - Separate retryable failures from non-retryable fallbacks in reporting - Track fallback samples and sessions separately from failed ones - Update console output to show both retryable and non-retryable categories - Include fallback details in JSON output with reasons and session info - Modify failure counting logic to distinguish between retryable and fallback reviews * feat(benchmark): add question_id tracking and enhanced fallback reporting - Add question_id function to extract query.question_id from data - Initialize question_id_by_id dictionary to store question IDs by index - Store question_id for each sample during data processing - Enhance fallback output to include question IDs and session information - Format sample labels with question IDs when available - Display session IDs associated with each fallback case * feat(benchmark): add question_id support and improve bad sample reporting - Add question_id_for function to extract question_id from multiple sources - Add sample_label function to format samples as idx(question_id) when available - Store question_id in data dictionary during processing - Change bad_golden and bad_sessions to store full records instead of just indices - Update list_bad output to show formatted labels with question_id information - Improve error reporting with more detailed sample identification * feat(benchmark): enhance golden check stats with structured output - Add related_session_ids function to extract session IDs from verdict records - Create grouped_records function to group records by question type - Replace flat list output with JSON-formatted grouped records in list_bad option - Replace flat list output with JSON-formatted grouped records in list_bad_sessions option - Maintain Chinese labels while adding structured data presentation - Improve readability of bad verdict record display with hierarchical grouping * feat(benchmark): update data structure for question indexing - Replace sample_label with _idx field for index tracking - Add question_id field to store _question_id values - Maintain backward compatibility with empty string defaults - Preserve existing session_id functionality - Update data mapping to include new fields in grouped results * refactor(benchmark): streamline golden answer verification process - Replace relevance filtering with comprehensive information extraction - Remove is_relevant field and simplify session summary structure - Change relevant_info to extracted_info for clarity - Update golden check logic to work with full extractions instead of filtered summaries - Simplify prompt instructions to focus on complete information extraction - Remove redundant schema validation and structured output requirements - Adjust statistics calculation to match new extraction approach - Update metadata field names to reflect extraction rather than relevance checking * feat(benchmark): add selective file deletion option to clean_sample_outputs - Add --filename argument to delete only specific root-level files - Modify iter_sample_targets function to accept optional filenames filter - Implement validation for root-level filename constraints - Update function calls to pass filenames parameter - Add example usage for selective file deletion in documentation * feat(benchmark): add error count metrics to golden check statistics - Added golden_bad, session_bad, and both_bad calculation fields - Updated console output format to include error counts per question type - Modified table display to show both accuracy rates and error numbers - Enhanced statistical summary with additional error breakdown metrics * test(search): update search step tests with include_source parameter - Added include_source=False parameter to VectorSearchStep initialization - Added include_source=False parameter to Bm25SearchStep initialization - Maintained existing RuntimeContext parameters for both search steps - Updated test calls to match new constructor signature with include_source option |
||
|
|
2612d25959
|
feat(lme): add cli execution and agentic search tooling (#334)
* feat(service): add CLI service for local job execution - Introduce CliService to execute single jobs locally without serving ports - Add prepare_start_config and should_precheck_start functions for CLI job setup - Update reme start command to use CLI service when job argument is provided - Change default service backend from http to cli in jinli_lme config - Modify SearchStep to use constants and rename configuration parameters - Add unit tests for CLI service functionality and configuration handling - Update file extension support to include json format in addition to md and jsonl * feat(search): add BM25 and vector search steps with configuration updates - Add Bm25SearchStep and VectorSearchStep classes with tool context deduplication - Register new search step components in index module - Update configuration to use separate vector_search and bm25_search endpoints - Modify LLM models from qwen3.7-plus/glm-5.1 to glm-5.2 variants - Adjust search parameters and remove hybrid search implementation - Configure embedding store as default in storage settings - Remove auto-memory and file catalog configurations - Update watch directories from multiple paths to session_dir only * feat(agent): add tool result offloading and workspace management - Add tool_results_dir configuration option for offloaded tool results storage - Implement ToolResultOffloadMiddleware to persist large tool results to files - Create WorkspaceBackend to standardize file operations across tools - Add configurable builtin tools selection with sequential execution option - Integrate middleware support for agent wrapper with offloading capability - Update application initialization to create tool results directory - Add safety mechanisms for filesystem operations with sanitized filenames - Enhance agent wrapper with configurable working directory handling - Upgrade agentscope dependency to version 2.0.4 for improved features # Conflicts: # reme/application.py * feat(benchmark): add LongMemEval agentic search and result management - Introduce AgenticAnswerStep for agent-based history search - Add LmePrepareJudgeStep and LmeSaveResultStep for evaluation pipeline - Implement AddDraftStep and ReadAllDraftStep for evidence accumulation - Update configuration with new agent wrapper and search parameters - Add comparison script for analyzing agent run differences - Include documentation for LongMemEval failure analysis - Enhance tool result offloading with skip options - Modify search defaults and indexing behavior * feat(agent): implement tool result offloading with system reminders - Added tool_result_offload_message parameter to agent wrapper reply method - Implemented configurable reminder template for offloaded tool results - Created system reminder messages when tool results are offloaded to files - Added Chinese user message template for agentic answer step - Updated tool result offloading middleware to use custom reminder templates - Enhanced agentic answer instructions to handle long tool results via draft storage * feat(scripts): add LongMemEval results summarization tool - Create summarize_lme_results.py script to analyze result JSON files - Implement command line interface with answer id and dataset root options - Add support for specifying index range with start and end parameters - Include option to show failure details and non-successful completions - Calculate completion statistics and accuracy metrics - Display detailed breakdown of yes/no/other judgements - Handle missing and unreadable result files gracefully - Format output with percentages and comprehensive summary statistics * feat(summarize_lme_results): add question type breakdown to result summary - Import defaultdict from collections module - Add by_type dictionary to track statistics by question type - Count completed, yes, no, and other responses for each question type - Display detailed breakdown table showing accuracy by question type - Include question type column when processing judgements - Print comprehensive summary with question type distribution - Calculate and display accuracy percentage for each question type category * feat(lme): switch to qwen3.7-max model and add shuffle functionality - Changed default LLM model from glm-5.1 to qwen3.7-max in jinli_lme.yaml - Added random module import for shuffle functionality - Implemented --shuffle argument with BooleanOptionalAction for dataset shuffling - Added --seed argument to control random seed for reproducible shuffling - Applied random shuffle to dataset indices when shuffle is enabled - Added console output showing shuffle operation and seed information * fix(cli): set default random seed for shuffle functionality - Changed default seed value from None to 42 for consistent shuffling behavior - Ensures reproducible results when using shuffle option without explicit seed - Maintains backward compatibility while providing deterministic defaults * refactor(benchmark): update agentic answer guidelines for grounding - Updated English instruction to emphasize strict grounding in retrieved context - Modified Chinese instruction to stress evidence-based responses without inference - Removed redundant conciseness requirement in both language versions - Enhanced clarity on proper use of draft saving and retrieval mechanisms - Strengthened emphasis against hallucination of unsupported facts * refactor(benchmark): update agentic search instructions and configuration - Replace separate vector_search and bm25_search with unified search tool - Update agent instructions to use single search tool with multiple strategies - Simplify Chinese instructions for search methodology - Add comprehensive search tool configuration with hybrid vector/BM25 capabilities - Increase model retry attempts from 1 to 3 for better reliability - Remove redundant tool references from job_tools list * feat(search): add configurable search limit with environment variable support - Remove hardcoded limit and min_score parameters from config schema - Increase LLM context size from 200000 to 1000000 - Add REME_SEARCH_LIMIT environment variable support for search configuration - Implement command line argument --search-limit to override default search limit - Add input validation to ensure search limit is positive - Modify subprocess execution to pass environment variables - Update search step to use dynamic default limit from environment or fallback to 5 * refactor(benchmark): remove agentic answer step and related configurations - Removed AgenticAnswerStep class and its registration - Deleted agentic_answer.yaml prompt configuration file - Removed agentic answer related job definitions from jinli_lme.yaml - Cleaned up tool result offloading middleware implementation - Removed tool_results_dir configuration field from application config - Deleted comparison and analysis scripts for agent runs - Removed agentic answer step from LME init module exports - Updated agent wrapper to remove tool result offloading functionality - Removed unused imports and dependencies in agent wrapper module * refactor(benchmark): remove unused LME result processing components - Removed LmePrepareJudgeStep and LmeSaveResultStep classes from benchmark module - Cleaned up imports and exports in lme module initialization - Removed unused middleware configuration from agent wrapper - Deleted obsolete result.py file containing deprecated result processing logic - Simplified agent instantiation by removing middleware parameter - Updated import statements to reflect removed dependencies * refactor(index): remove unused search steps and update imports - Remove Bm25SearchStep and VectorSearchStep from index steps module - Remove unused prepare_start_config and should_precheck_start exports - Move import statements to proper location in reme.py - Update test module to use direct import path for CliService - Remove vector_search and bm25_search configurations from jinli_lme.yaml - Add workspace directory environment variable configuration - Add docstring to getcwd method in agent wrapper - Remove empty middleware list from agent wrapper initialization * feat(index): add BM25 and vector search steps with tool context deduplication - Add Bm25SearchStep for plain BM25 keyword search with tool_context deduplication - Add VectorSearchStep for plain vector search with tool_context deduplication - Implement tool context state management with TTL-based deduplication - Add support for chunk deduplication across tool contexts within TTL window - Update index steps module to include new search step classes - Add test coverage for CLI metadata output functionality - Refactor CLI service to remove unused show_status parameter - Update documentation comments to reflect internal service configuration * feat(steps): add Python code execution capability - Introduce PythonExecuteStep to run Python code in subprocess - Add configuration for python_execute step in jinli_lme.yaml - Register python_execute in available tools list - Implement timeout handling with default 60 second limit - Capture stdout/stderr output and return code metadata - Add comprehensive unit tests for execution scenarios - Support workspace directory context for code execution - Handle timeout errors and runtime exceptions gracefully * refactor(python_execute): replace subprocess with asyncio for Python code execution - Replace subprocess.run with asyncio.create_subprocess_exec for non-blocking execution - Add _PythonResult dataclass to encapsulate execution results and timeout status - Implement proper timeout handling with asyncio.wait_for and process.kill() - Update metadata to include returncode and stderr when timeout occurs - Convert synchronous _run_python method to asynchronous implementation - Maintain backward compatibility while improving execution reliability * refactor(python_execute): replace subprocess with asyncio for Python code execution - Replace subprocess.run with asyncio.create_subprocess_exec for non-blocking execution - Add _PythonResult dataclass to encapsulate execution results and timeout status - Implement proper timeout handling with asyncio.wait_for and process.kill() - Update metadata to include returncode and stderr when timeout occurs - Convert synchronous _run_python method to asynchronous implementation - Maintain backward compatibility while improving execution reliability |
||
|
|
206a53e5ed
|
init: reme version 0.4.0 (#284) |
Renamed from reme4/steps/index/__init__.py (Browse further)