ReMe/reme2/schema
huangsen 51ec09d98f feat(parser): implement AST-based markdown chunking with full document TOC
- Replace legacy line-based chunking with AST tree approach that builds
  a complete document skeleton with content inlined under relevant
  sections
- Add new chunking parameters: chunk_chars (default 2000) and embed_toc
  (default True) to control content size and TOC inclusion
- Implement recursive chunking algorithm that respects structural
  boundaries (code lines, table rows, list items) and prevents splits
  inside blocks
- Introduce part markers [Part X/N] for oversized leaf blocks that
  require splitting
- Add CLI tool for inspecting parsed chunks and edges with options for
  preview and configuration
- Refactor edge extraction to use FileEdge.from_text instead of
  parse_wikilinks for consistency

BREAKING CHANGE: Chunk format changes significantly with full TOC
skeleton wrapping content, affecting embedding models expecting
breadcrumb prefixes.
2026-05-13 14:32:00 +08:00
..
__init__.py feat(parser): implement AST-based markdown chunking with full document TOC 2026-05-13 14:32:00 +08:00
application_config.py feat(components): add token counter and file-based utility components 2026-04-16 20:21:04 +08:00
as_msg_stat.py feat(components): add token counter and file-based utility components 2026-04-16 20:21:04 +08:00
chunk_filter.py ``` 2026-05-08 16:14:42 +08:00
emb_node.py up 2026-05-11 17:54:24 +08:00
file_chunk.py up 2026-05-12 23:49:21 +08:00
file_edge.py feat(parser): implement AST-based markdown chunking with full document TOC 2026-05-13 14:32:00 +08:00
file_node.py up 2026-05-13 12:22:59 +08:00
request.py feat(components): add token counter and file-based utility components 2026-04-16 20:21:04 +08:00
response.py feat(components): add token counter and file-based utility components 2026-04-16 20:21:04 +08:00
stream_chunk.py feat(components): add token counter and file-based utility components 2026-04-16 20:21:04 +08:00