mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-09-07 08:26:06 +00:00
* feat(file_chunker): add dedicated JSON and JSONL file chunkers - Add JsonFileChunker: structure-aware chunking preserving nested key paths, optional list-to-dict conversion, size measured by json.dumps() char count - Add JsonlFileChunker: line-aligned sliding-window chunking with configurable overlap, supports char/byte mode switching - Register both chunkers in default.yaml (json for .json, jsonl for .jsonl) - Add comprehensive unit tests (21 + 20 test cases) * chore(config): update default chunker supported_extensions to txt/log * refactor(json_chunker): optimize _build_tree O(n²) serialization and rewrite tests - Fix O(n²) redundant json.dumps in _build_tree: * Empty containers handled directly as leaves (0 serialization) * Non-empty containers recurse first, then reconstruct+dump once * Only containers that become leaves pay serialization cost - Add _reconstruct_object/_reconstruct_array helpers - Remove dead code: _merge_json method - Apply user changes: min_element_size formula 0.01->0.05, threshold < to <= - Use indent=None for compact output (consistent with _SizeNode estimation) - Remove unused _text_size from JsonlFileChunker Test rewrite: - Replace try/finally boilerplate with make_json fixture - Group tests into TestXxx classes with pytest.mark.parametrize - Add TestOutputValidation: 9 parametrized scenarios verifying: * All chunks are valid JSON * Text length <= chunk_chars (with single-leaf tolerance) * Leaf-value concatenation matches original data (dict + array roots) - Add TestSizeNode: incremental size accuracy tests - Add TestDfsAlgorithm: path wrapping, DFS order, calibration tests - Update test_min_element_size_formula for new 0.05 multiplier - Update test_build_tree_structure for larger min_element_size * chore: apply black formatting to test files
15 lines
421 B
Python
15 lines
421 B
Python
"""File chunker components."""
|
|
|
|
from .base_file_chunker import BaseFileChunker
|
|
from .default_file_chunker import DefaultFileChunker
|
|
from .json_file_chunker import JsonFileChunker
|
|
from .jsonl_file_chunker import JsonlFileChunker
|
|
from .markdown_file_chunker import MarkdownFileChunker
|
|
|
|
__all__ = [
|
|
"BaseFileChunker",
|
|
"DefaultFileChunker",
|
|
"JsonFileChunker",
|
|
"JsonlFileChunker",
|
|
"MarkdownFileChunker",
|
|
]
|