xyf2020
|
82971ac5b0
|
feat(chunker): better json chunker and json chunker (#325)
* feat(file_chunker): add dedicated JSON and JSONL file chunkers
- Add JsonFileChunker: structure-aware chunking preserving nested key paths,
optional list-to-dict conversion, size measured by json.dumps() char count
- Add JsonlFileChunker: line-aligned sliding-window chunking with configurable
overlap, supports char/byte mode switching
- Register both chunkers in default.yaml (json for .json, jsonl for .jsonl)
- Add comprehensive unit tests (21 + 20 test cases)
* chore(config): update default chunker supported_extensions to txt/log
* refactor(json_chunker): optimize _build_tree O(n²) serialization and rewrite tests
- Fix O(n²) redundant json.dumps in _build_tree:
* Empty containers handled directly as leaves (0 serialization)
* Non-empty containers recurse first, then reconstruct+dump once
* Only containers that become leaves pay serialization cost
- Add _reconstruct_object/_reconstruct_array helpers
- Remove dead code: _merge_json method
- Apply user changes: min_element_size formula 0.01->0.05, threshold < to <=
- Use indent=None for compact output (consistent with _SizeNode estimation)
- Remove unused _text_size from JsonlFileChunker
Test rewrite:
- Replace try/finally boilerplate with make_json fixture
- Group tests into TestXxx classes with pytest.mark.parametrize
- Add TestOutputValidation: 9 parametrized scenarios verifying:
* All chunks are valid JSON
* Text length <= chunk_chars (with single-leaf tolerance)
* Leaf-value concatenation matches original data (dict + array roots)
- Add TestSizeNode: incremental size accuracy tests
- Add TestDfsAlgorithm: path wrapping, DFS order, calibration tests
- Update test_min_element_size_formula for new 0.05 multiplier
- Update test_build_tree_structure for larger min_element_size
* chore: apply black formatting to test files
|
2026-07-08 15:18:59 +08:00 |
|