mirror of
https://github.com/agentscope-ai/ReMe.git
synced 2026-08-28 05:25:04 +00:00
Some checks are pending
Pre-commit / run (ubuntu-latest) (push) Waiting to run
Tests ReMe / Unit Tests - py3.13 (push) Waiting to run
Tests ReMe / Unit Tests - py3.11 (push) Waiting to run
Tests ReMe / Unit Tests - py3.12 (push) Waiting to run
Windows Smoke / CLI smoke - py3.11 (push) Waiting to run
* feat: add ssh proxy * feat: add ssh proxy * feat: add ssh proxy * feat: add ssh proxy * feat: add prompt * feat: add agent wrapper * feat: add agent wrapper * feat: add agent wrapper * feat: add tushare skill * feat: add tushare skill * feat: add tushare skill * feat: add none stream * chore(deps): update dependency versions in pyproject.toml - Bump claude-agent-sdk from 0.2.123 to 0.2.126 - Upgrade pre-commit to version 4.6.1 or higher - Upgrade pytest to version 9.1.1 or higher * feat(agent_wrapper): add session compaction support and unify session commands - Introduce compact_session method to BaseAgentWrapper and implement it in AsAgentWrapper, CcAgentWrapper, and CodexAgentWrapper - Add session_command module with SessionCommandResult dataclass and handle_session_command function for /clear and /compact commands - Update __init__.py exports to include session_command handlers - Modify DingTalkWaitStep to handle session commands via handle_session_command function - Remove streaming mode from DingTalkWaitStep and simplify reply handling to final Markdown replies only - Add unit tests for session compaction methods and session command handling across wrappers and DingTalk integration - Clean up and remove obsolete streaming and card rendering code from DingTalk wait step - Adjust daily_cookbook.yaml to remove stream and card_update_interval config entries for DingTalk wait step * feat(auto_fin): add Auto Fin simulated portfolio cookbook workflow - Add comprehensive Auto Fin schema exports for multiple models and enums - Implement base class and helpers for Auto Fin analysis steps - Create file, state, and formatting utilities for Auto Fin with atomic file writes and locking - Define Auto Fin pipeline with four analysis agents: backtest, event, portfolio, and US correlation - Register Auto Fin package in cookbook workflows and schema initialization - Add detailed documentation in markdown describing the system design, workflow, and data contracts * feat(outbound_proxy): add application-scoped outbound HTTP proxy components - Introduce BaseOutboundProxy and OutboundProxyEndpoint as core contracts - Implement FixedHttpOutboundProxy for external HTTP proxy integration - Add SshHttpOutboundProxy providing SSH-backed local HTTP proxy tunnels - Register outbound proxy components in component registry and enumeration - Update components package to include outbound_proxy module - Add dependency on pproxy for SSH HTTP proxy bridging - Include comprehensive unit tests covering proxy lifecycle, validation, environment merging, error handling, readiness, and monitoring mechanisms * refactor(network): replace SSH proxy with explicit HTTP outbound proxy - Remove SSH proxy helper implementation and references in codebase - Add support for explicit HTTP proxy URL in arXiv and HuggingFace clients - Modify clients to use async context manager for consistent resource handling - Update daily paper steps to forward outbound proxy configuration explicitly - Change tests to cover new proxy usage model and remove SSH proxy mocks - Add outbound proxy component configuration in daily_cookbook.yaml - Ensure proxy URL usage disables environment trust in HTTP clients - Fix app context component enum access to be defensive against missing keys * feat(agent_wrapper): add managed proxy support for command environments - Introduce BaseOutboundProxy binding in BaseAgentWrapper for outbound proxy management - Add bash_environment and command_proxy_environment properties to apply proxy settings - Update WorkspaceBackend instantiation in AsAgentWrapper to use bash_environment - Inject managed proxy export commands into Claude Code Bash commands via hooks - Enhance CodexAgentWrapper to include managed proxy in shell environment policy - Modify daily_cookbook.yaml steps to specify outbound_proxy as default where needed - Add comprehensive unit tests verifying managed proxy injection and environment isolation - Ensure subprocess_environment remains unchanged while proxy is applied selectively to commands * refactor(memory): replace search job_tools with memory in daily cookbook config - Change workspace_dir default from .reme to reme_workspace - Replace search job_tools with memory across multiple components and jobs - Update descriptions to reflect long-term memory retrieval instead of search - Modify system prompts to instruct using memory for retrieving notes - Adjust unit tests to verify memory job_tools and job presence instead of search - Ensure consistency in configuration and tests for memory backend usage * refactor(config): rename memory to memory_search in daily cookbook config - Change all occurrences of "memory" to "memory_search" in job_tools and job definitions - Update related system prompts to reflect the new memory_search terminology - Modify unit tests to assert the presence of memory_search instead of memory - Ensure consistency across skills, job tools, and backend configurations in multiple components * feat(auto_fin): add deterministic quantitative research and ranking fusion - Introduce new schema models: EtfScore, RankingMetrics, ExtremeAnalysis, DimensionRanking, and FusionRanking to represent deterministic research outputs - Add ranking data to event, backtest, us_correlation, and portfolio analysis outputs - Implement ranking_section renderer to format Top20 scores and diagnostics in Markdown - Develop AutoFinQuantStep for deterministic ETF ranking using TuShare data, Polars, and a custom extremely randomized tree ensemble - Integrate quantitative rankings into backtest and portfolio analysis steps and reports - Extend auto_fin pipeline with new quant_enabled and quant_required config options - Enforce ranking constraints like unique codes, contiguous ranks, and normalized fusion weights - Update analysis YAMLs with rules limiting data freshness, universe, and ranking usage - Incorporate ranking outputs into all major markdown report bodies in Auto Fin pipeline - Add concurrency-limited asynchronous TuShare client to fetch required market data - Introduce cross-sectional rank correlation and NDCG metrics for ranking quality evaluation * feat(auto_fin): implement stage-wise notification and reporting for analysis pipeline - Refactor notification config in daily_cookbook.yaml to support dispatch steps - Update AutoFinNotificationStep to deduplicate notifications per run stage - Add _notify_stage method in pipeline to send notifications for each analysis stage - Implement persistence and notification for event, backtest, US correlation, and portfolio stages - Modify pipeline flow to persist reports and notify after each stage completion - Adjust metadata to track notifications and errors per stage - Update tests to verify stage-wise notification sending and deduplication - Remove older combined report persistence in favor of modular stage handling * feat(auto_fin): add outbound proxy support for Tushare API usage - Introduce BaseOutboundProxy reference in AutoFinPipelineStep and AutoFinQuantStep - Update TushareResearchClient and trade calendar fetch to accept and use proxy URL - Create _ProxiedTushareApi adapter to route Tushare requests via explicit HTTP proxy - Modify create_tushare_api utility to optionally return proxied API client - Add unit tests covering proxy forwarding and client behavior with managed proxies - Ensure proxy usage respects explicit proxy URL over environment fallback - Integrate outbound proxy into data fetching and quantitative research steps * feat(auto_fin): enforce checkpoint time validation and add state models - Introduce AnalysisState base class and specific states for event, backtest, and US correlation analyses - Replace analysis output types with corresponding state classes in run schemas - Add require_checkpoint_reached method to validate decision_at/data_cutoff against current time - Enforce checkpoint time checks before analysis steps in event, backtest, portfolio, and quant analyses - Refactor quant data loading to include adjustment factors and apply price adjustments without fallback - Update analysis YAML docs to require real-time checkpoint validation and forbid using future data - Improve portfolio run serialization by excluding redundant legacy fields and nested proposed actions - Add helper to extract readable sections from persisted checkpoint documents - Fix event analysis output validation to reject events and sources with future timestamps * feat(auto_fin): auto-select latest reached checkpoint if none specified - Extend checkpoint config to accept empty string for auto selection - Add static method to compute latest checkpoint reached by current time - Modify pipeline step to auto-select checkpoint based on trade calendar and time - Adjust force flag default depending on whether checkpoint is explicit or auto - Log details when checkpoint is auto-selected to improve observability - Add comprehensive tests for auto checkpoint selection logic and edge cases - Remove deprecated default and required constraints from force parameter in config * refactor(auto_fin): unify datetime comparison with compare_datetimes utility - Replace direct datetime comparisons with compare_datetimes function calls - Use cmp_to_key with compare_datetimes for sorting datetime tuples and lists - Update validation logic in backtest, event, analysis, and ledger modules for consistent datetime handling - Add unit tests to verify handling of naive and aware datetime comparisons in event and backtest validations - Ensure marked_at and interval_end timestamps are set and compared consistently using compare_datetimes - Improve correctness of ordering and conditional checks related to timestamps throughout auto_fin steps and ledger code * feat(auto_fin): add datetime comparison helper for mixed timezone data - Implement compare_datetimes function to handle naive and aware datetimes - Ensure naive datetime is interpreted in the known timezone of the counterpart - Facilitate comparisons between legacy and timezone-aware Auto Fin data - Add module docstring explaining purpose of the helpers * docs(auto_fin): enforce unique ETF representative per sub-theme in analysis rules - Update backtest.yaml to recommend or highlight only one ETF per sub-theme for ETF analyses - Modify event.yaml to map only one representative ETF per sub-theme, avoiding duplicate recommendations - Revise portfolio.yaml to restrict holdings/buys to a single ETF per sub-theme, preventing repeated buys of highly overlapping ETFs - Adjust us_correlation.yaml to retain only one representative A-share ETF per sub-theme for mapping or recommendation - Add test to verify presence of new sub-theme uniqueness guidance in step prompts * feat(auto_fin): separate draft model and include deterministic fusion ranking - Introduce _PortfolioProposalDraft pydantic model for agent-authored fields before ranking - Discard any "fusion_ranking" data from draft to prevent conflicts with canonical ranking - Modify AutoFinPortfolioStep to receive draft, enrich with fusion_ranking, and produce final output - Update tests to use _PortfolioProposalDraft and validate deterministic fusion ranking propagation - Add async test verifying fusion ranking is correctly set in portfolio output with no errors * refactor(auto_fin): rewrite and simplify Auto Fin schema and steps - Remove legacy Auto Fin analysis step modules and helpers - Replace complex ranking and portfolio models with simplified current-news models - Update schema to focus on news-case workflow with new domain models - Remove A-share decision checkpoints and backtest details from schema - Simplify recommendation and decision output structures - Clean up deprecated state and utility functions - Update Auto Fin steps initialization to new pipeline steps only - Improve uniqueness validation for themes and ETFs in research plan * feat(auto_fin): implement full local cache and analysis workflow for Auto Fin - Add AutoFinDataStep to prepare and cache daily TuShare data with lookback - Add AutoFinAnalysisStep to analyze cached data and generate Markdown report - Implement detailed time window, ETF filtering, and historical case validation - Introduce YAML prompts for planning and decision-making steps - Update .gitignore to include reme_workspace/ - Clean up config and import structure for auto_fin steps - Remove old pipeline.py and consolidate functionality into new modules - Use polars for efficient CSV reading and data processing - Ensure atomic writes and strict JSON serialization for cache files - Enforce rules on news timing, ETF universe, and historical case usage * fix(auto_fin): restrict news data source to '财联社' in analysis and cache - Update analysis templates to specify current news as from '财联社' only - Modify news fetching functions to filter by source '财联社' - Add validation method to check cached news source correctness - Update news caching logic to exclude non-'财联社' news - Enhance unit tests with multiple sources to ensure filtering works - Confirm news API calls include source filter parameter as '财联社' * refactor(auto_fin): convert I/O methods to asynchronous implementations - Change _news, _dataset, and _theme_data methods to async for improved concurrency - Move JSONL and CSV reading operations to asynchronous wrappers using asyncio.to_thread - Remove synchronous _read_jsonl and _read_csv functions, integrate them as static async class methods - Update cache validation methods to async, awaiting I/O operations accordingly - Adjust usage of dataset and news retrieval in analysis step to await asynchronous methods - Add async unit test to validate JSONL reading with unicode line separators - Preserve existing functionality while enabling non-blocking file and data access * fix(nx_file_graph): defer networkx import and improve dependency handling - Move networkx import inside NxFileGraph constructor for lazy loading - Raise ImportError with original exception context if networkx is missing - Remove module-level fallback assignment of nx to None - Expand test to block loading of multiple optional core dependencies eagerly - Change exception type in test from ModuleNotFoundError to AssertionError - Update test comments to reflect broader optional dependency checks * feat(embedding_store): add quota retry delay mechanism for embedding requests - Introduce quota_retry_delay parameter to configure wait time before retry on quota exhaustion - Implement detection of insufficient quota errors in LocalEmbeddingStore without external SDK - Add retry logic with custom delay when quota is insufficient during embedding requests - Update configuration to set max_retries and quota_retry_delay defaults for embedding store - Add unit tests covering quota exhaustion retry behavior with delay and opt-in control - Ensure existing retry behavior remains unchanged if quota_retry_delay is not set * feat(auto_fin): add detailed logging to analysis and data fetching steps - Add _preview static method for bounded diagnostic output in analysis.py - Log prompt start, completion, errors, and validation details in _reply method - Add info logs for major processing steps in execute method of analysis.py - Add debug and info logs for cache validation, data fetching, and pagination in data.py - Log conditions for skipping reports and cache plans in data.py execute method - Log download summaries and cache writes for news and ETF data - Improve error logging with exception details in cache validation functions - Ensure all logs include context such as record counts, paths, and parameters * refactor(auto_fin): overhaul Auto Fin workflow and schema contracts - Replace old Auto Fin schema models with comprehensive new data classes - Remove legacy Auto Fin analysis step in favor of modular agent-based steps - Introduce AutoFinAgentStep for validating structured agent replies - Simplify data cleaning and JSONL writing utilities for news cache - Remove synchronous and asynchronous dataset methods from analysis step - Redefine Auto Fin analysis configuration for 360-day news retention and multi-step pipeline - Remove embedded analysis prompt templates and replace with agent-driven logic - Update __init__.py exports to match new step implementations and remove deprecated classes - Improve error handling and validation in agent step reply processing - Clean up redundant imports and unused code in analysis and data preparation modules * feat(auto_fin): add detailed logging for analysis and data processing steps - Add timing logs to measure agent prompt processing duration in analysis.py - Log news cache hits and news write paths with record counts in data.py - Include detailed info logs for news download start and completion in data.py - Add start, progress, and completion logs with topic and event counts in history.py - Log start and completion of merge step including path and ETF count in merge.py - Add start and done logs with window and news counts in topic.py * feat(auto_fin): enhance schema and steps with detailed ETF and event modeling - Replace and add multiple AutoFin schema classes to support detailed ETF selection, historical research, market analysis, forecast models, and report output with validation - Implement Shanghai timezone normalization and strict validation in schema models - Remove deprecated AutoFin analysis agent step and consolidate reply handling in base step - Introduce AutoFinStep base class with shared helpers for prompt handling, data fetching, logging, and JSONL file operations - Add AutoFinDataStep to manage daily news data complete with schedule validation, caching, and source validation logic - Update cookbook configuration to customize auto_fin step parameters and simplify outbound proxy settings - Refactor imports and clean unused code for better maintainability * feat(auto_fin): introduce detailed historical event resolution and market similarity analysis - Add AutoFinHistoricalEventReference and AutoFinHistoricalSimilarity models for refined event referencing and similarity judgment - Implement validation to ensure non-empty critical fields and uniqueness of historical news IDs - Develop method to resolve Agent-selected historical event references from workspace files with strict path and existence checks - Enrich historical events with market entry and future returns data after resolution - Redesign market step to calculate similarity-weighted ETF forecasts based on matched historical event similarities - Enforce validation on matched historical events for uniqueness and proper weight summation - Simplify merge step output to final Markdown report without YAML frontmatter and redundant fields - Update user instructions for history search, market, and merge steps to reflect new data structures and responsibilities - Adjust test suite to cover new schema and step behavior changes, including enhanced validation and JSON output formats * feat(auto_fin): add new cron jobs and output analysis jsonl - Add new cron jobs auto_fin_1145_cron and auto_fin_1800_cron with auto_fin_steps - Change auto_fin_0930_cron schedule to run Monday to Sunday - Extend merge step to write analysis data to auto_fin_analysis.jsonl - Update unit tests to verify new cron jobs and their steps configuration * fix(auto_fin): improve atomic file write and refresh daily index - Change temporary file naming to include UUID for uniqueness and hidden prefix - Replace atomic write method from using Path.replace to os.replace with safe unlink - Add import and use os.replace for safer file replace operation - Refresh daily index after writing auto finance markdown and JSONL files - Import and call refresh_day_index in merge step to update file index asynchronously * docs(cookbook): add optional SSH proxy configuration in README files - Introduce optional SSH proxy setup in auto-fin and daily_paper cookbooks - Provide instructions to enable outbound proxy via `daily_cookbook.yaml` and environment variables - Add `REME_PROXY_IP` and `REME_PROXY_ACCOUNT` environment variables descriptions in multiple README files - Update English and Chinese README and README_ZH documents with proxy details - Maintain consistent formatting of environment variable tables across documents * fix(file_io): include schema_version in hidden metadata keys - Added "schema_version" to _INDEX_HIDDEN_METADATA_KEYS in _daily_index.py - Updated _render_notes_block to always include additional keys regardless of schema_version fix(deps): move pproxy dependency to later in pyproject.toml - Removed pproxy from early dependencies list - Added pproxy back near the end of dependency list for better ordering fix(outbound_proxy): require pproxy package for ssh_http proxy - Added importlib.util check for pproxy package presence - Raise RuntimeError if pproxy is not installed when using SSH HTTP outbound proxy - Improved error message suggests installing reme-ai with 'core' extra * docs(readme): update News section with new Cookbook workflows - Clarify introduction of optional Cookbooks with Daily Paper and Auto Fin workflows - Update English README to reflect both paper discovery and file-native ETF event research - Revise Chinese README to include financial news and historical market data research capability - Maintain announcement of paper acceptance at Findings of ACL 2026 * feat(auto_fin): add calculation results to final Markdown output - Implement _calculation_results to summarize forecast for each ETF analyzed - Include program-calculated results in the JSON input for the Markdown report - Update YAML template to incorporate calculation results and adjust recommendation rules - Refine recommendation logic to rely on event impact judgments combined with calculation outputs - Modify tests to verify presence of calculation results and updated report content and format * up prompt * fix(keyword_index): ignore non-indexable chunks during keyword sync - Add is_indexable method to base and BM25 keyword index classes to check text tokenizability - Update local file store to exclude non-indexable chunks from expected document IDs to prevent rebuild - Fix JSONL chunker to correctly handle Unicode line separator U+2028 inside JSON strings without splitting - Add test to ensure non-empty but non-indexable chunk does not trigger keyword index rebuild - Add test to verify U+2028 character does not cause incorrect JSONL record splitting
455 lines
16 KiB
Python
455 lines
16 KiB
Python
"""Tests for JsonlFileChunker."""
|
|
|
|
# pylint: disable=protected-access
|
|
|
|
import asyncio
|
|
import json
|
|
import os
|
|
import tempfile
|
|
|
|
from reme.components.file_chunker import JsonlFileChunker
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Helpers
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _run(coro):
|
|
return asyncio.run(coro)
|
|
|
|
|
|
def _write_jsonl(lines: list[str]) -> str:
|
|
"""Write lines to a temp .jsonl file and return its path."""
|
|
fd, path = tempfile.mkstemp(suffix=".jsonl")
|
|
with os.fdopen(fd, "w", encoding="utf-8") as f:
|
|
for line in lines:
|
|
f.write(line if line.endswith("\n") else line + "\n")
|
|
return path
|
|
|
|
|
|
def _make_records(n: int, width: int = 20) -> list[str]:
|
|
"""Generate *n* JSONL lines, each roughly *width* chars of JSON."""
|
|
return [json.dumps({"id": i, "text": "x" * max(0, width - 20)}) for i in range(n)]
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Basic behaviour
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_empty_file():
|
|
"""Empty file → zero chunks."""
|
|
path = _write_jsonl([])
|
|
try:
|
|
chunker = JsonlFileChunker()
|
|
node, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) == 0
|
|
assert node.links == []
|
|
print("✓ test_empty_file passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_blank_lines_only():
|
|
"""File with only blank lines → zero chunks."""
|
|
fd, path = tempfile.mkstemp(suffix=".jsonl")
|
|
with os.fdopen(fd, "w") as f:
|
|
f.write("\n\n\n")
|
|
try:
|
|
chunker = JsonlFileChunker()
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) == 0
|
|
print("✓ test_blank_lines_only passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_single_line():
|
|
"""One line → one chunk."""
|
|
lines = _make_records(1)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker()
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) == 1
|
|
assert chunks[0].start_line == 1
|
|
assert chunks[0].end_line == 1
|
|
assert lines[0] in chunks[0].text
|
|
print("✓ test_single_line passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_unicode_line_separator_inside_json_string_is_not_a_record_boundary():
|
|
"""U+2028 is valid JSON string content, not a JSONL physical newline."""
|
|
record = json.dumps({"id": 1, "text": "before\u2028\u2028after"}, ensure_ascii=False)
|
|
path = _write_jsonl([record])
|
|
try:
|
|
chunker = JsonlFileChunker(max_lines_per_chunk=1)
|
|
node, chunks = _run(chunker.chunk(path))
|
|
|
|
assert len(chunks) == 1
|
|
assert chunks[0].text == record + "\n"
|
|
assert chunks[0].start_line == chunks[0].end_line == 1
|
|
assert node.chunk_ids == [chunks[0].id]
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_all_lines_fit_in_one_chunk():
|
|
"""Small file fits entirely in one chunk."""
|
|
lines = _make_records(5, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=5000)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) == 1
|
|
assert chunks[0].start_line == 1
|
|
assert chunks[0].end_line == 5
|
|
print("✓ test_all_lines_fit_in_one_chunk passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_max_lines_per_chunk_one():
|
|
"""max_lines_per_chunk=1 emits exactly one complete line per chunk."""
|
|
lines = _make_records(5, width=10)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=5000, max_lines_per_chunk=1)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) == len(lines)
|
|
for line_number, (line, chunk) in enumerate(zip(lines, chunks), start=1):
|
|
assert chunk.start_line == line_number
|
|
assert chunk.end_line == line_number
|
|
assert chunk.text == line + "\n"
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Multi-chunk splitting (no overlap)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_multiple_chunks_no_overlap():
|
|
"""Large file is split into multiple chunks with no overlap."""
|
|
# Each line ≈ 32 chars (with \n), 10 lines ≈ 320 chars.
|
|
lines = _make_records(10, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=100, max_overlap_chars=0)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) > 1, f"Expected >1 chunks, got {len(chunks)}"
|
|
# No overlap: consecutive chunks don't share lines.
|
|
for i in range(len(chunks) - 1):
|
|
assert chunks[i].end_line < chunks[i + 1].start_line
|
|
# All lines are covered.
|
|
assert chunks[0].start_line == 1
|
|
assert chunks[-1].end_line == 10
|
|
print(f" Created {len(chunks)} chunks")
|
|
print("✓ test_multiple_chunks_no_overlap passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_line_aligned_no_intra_line_split():
|
|
"""Every chunk boundary falls on a line boundary."""
|
|
lines = _make_records(20, width=40)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=150, max_overlap_chars=0)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) > 1
|
|
for c in chunks:
|
|
# Each chunk text should be a concatenation of whole lines.
|
|
text_lines = c.text.splitlines()
|
|
assert len(text_lines) == (c.end_line - c.start_line + 1)
|
|
print("✓ test_line_aligned_no_intra_line_split passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Overlap behaviour
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_overlap_lines_shared():
|
|
"""Overlap lines appear at the end of chunk N and the start of chunk N+1."""
|
|
# Each line is ~32 chars. max_chars=100 → ~3 lines per chunk.
|
|
# max_overlap_chars=40 → ~1 line of overlap.
|
|
lines = _make_records(12, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=100, max_overlap_chars=40)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
assert len(chunks) > 1
|
|
# Verify overlap: chunk[i].end_line >= chunk[i+1].start_line
|
|
for i in range(len(chunks) - 1):
|
|
assert chunks[i].end_line >= chunks[i + 1].start_line, (
|
|
f"Chunk {i} end_line={chunks[i].end_line} < " f"chunk {i+1} start_line={chunks[i+1].start_line}"
|
|
)
|
|
# Last chunk still reaches the end of the file.
|
|
assert chunks[-1].end_line == 12
|
|
print(f" Created {len(chunks)} chunks with overlap")
|
|
print("✓ test_overlap_lines_shared passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_zero_overlap():
|
|
"""max_overlap_chars=0 produces strictly non-overlapping chunks."""
|
|
lines = _make_records(10, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=100, max_overlap_chars=0)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
for i in range(len(chunks) - 1):
|
|
assert chunks[i].end_line < chunks[i + 1].start_line
|
|
print("✓ test_zero_overlap passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_overlap_capped_by_max_overlap_chars():
|
|
"""Overlap never exceeds max_overlap_chars in size."""
|
|
lines = _make_records(10, width=50)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
max_overlap = 60 # Should allow at most 1 line of overlap (~51 chars).
|
|
chunker = JsonlFileChunker(max_chars=200, max_overlap_chars=max_overlap)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
if len(chunks) > 1:
|
|
for i in range(len(chunks) - 1):
|
|
overlap_start = chunks[i + 1].start_line
|
|
overlap_end = chunks[i].end_line
|
|
if overlap_start <= overlap_end:
|
|
overlap_lines = list(range(overlap_start, overlap_end + 1))
|
|
overlap_text = "".join(lines[ln - 1] for ln in overlap_lines)
|
|
assert len(overlap_text) <= max_overlap, f"Overlap size {len(overlap_text)} > {max_overlap}"
|
|
print("✓ test_overlap_capped_by_max_overlap_chars passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Oversized single line
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_single_line_exceeds_max():
|
|
"""A line larger than max_chars is emitted alone (no intra-line split)."""
|
|
long_line = json.dumps({"data": "A" * 200})
|
|
lines = ["short", long_line, "short"]
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=64, max_overlap_chars=0)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
# The long line must appear in exactly one chunk by itself (or with
|
|
# no split).
|
|
found = False
|
|
for c in chunks:
|
|
if long_line in c.text:
|
|
found = True
|
|
break
|
|
assert found, "Long line was not found in any chunk"
|
|
assert len(chunks) >= 2 # At least some splitting happened.
|
|
print("✓ test_single_line_exceeds_max passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Byte mode
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_byte_mode():
|
|
"""mode='bytes' counts byte length instead of char length."""
|
|
# CJK characters are 3 bytes each in UTF-8 but 1 char each.
|
|
# Use ensure_ascii=False so raw CJK chars appear in the file.
|
|
cjk_record = json.dumps({"text": "你" * 10}, ensure_ascii=False)
|
|
# The JSON string itself has CJK chars; after writing to file each line
|
|
# has ~13 ASCII chars ({"text": "..."}) + 10 CJK chars.
|
|
# char length ≈ 23, byte length ≈ 13 + 10*3 = 43.
|
|
lines = [cjk_record] * 6
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
# In char mode: each line ≈ 24 chars → 2 lines per chunk at max=50.
|
|
chunker_chars = JsonlFileChunker(max_chars=50, max_overlap_chars=0, mode="chars")
|
|
_, chunks_chars = _run(chunker_chars.chunk(path))
|
|
|
|
# In byte mode: each line ≈ 44 bytes → 1 line per chunk at max=50.
|
|
chunker_bytes = JsonlFileChunker(max_chars=50, max_overlap_chars=0, mode="bytes")
|
|
_, chunks_bytes = _run(chunker_bytes.chunk(path))
|
|
|
|
# Byte mode should produce more chunks since each line is larger in bytes.
|
|
assert len(chunks_bytes) > len(chunks_chars), (
|
|
f"Byte mode ({len(chunks_bytes)}) should produce more chunks "
|
|
f"than char mode ({len(chunks_chars)}) for CJK text"
|
|
)
|
|
print(f" char mode: {len(chunks_chars)} chunks, byte mode: {len(chunks_bytes)} chunks")
|
|
print("✓ test_byte_mode passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_invalid_mode_raises():
|
|
"""An invalid mode string raises ValueError."""
|
|
try:
|
|
JsonlFileChunker(mode="words")
|
|
assert False, "Should have raised ValueError"
|
|
except ValueError:
|
|
pass
|
|
print("✓ test_invalid_mode_raises passed")
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# FileNode / FileChunk properties
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_file_chunk_properties():
|
|
"""Each FileChunk has correct path, line range, text, and hash id."""
|
|
lines = _make_records(5, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=5000)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
c = chunks[0]
|
|
assert c.start_line >= 1
|
|
assert c.end_line >= c.start_line
|
|
assert c.id and len(c.id) > 0
|
|
assert c.text.strip()
|
|
print("✓ test_file_chunk_properties passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_file_node_properties():
|
|
"""FileNode carries path, st_mtime, chunk_ids, and empty links."""
|
|
lines = _make_records(3)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker()
|
|
node, chunks = _run(chunker.chunk(path))
|
|
assert node.st_mtime > 0
|
|
assert node.links == []
|
|
assert node.chunk_ids == [c.id for c in chunks]
|
|
print("✓ test_file_node_properties passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
def test_hash_id_deterministic():
|
|
"""Same file produces the same chunk hash ids on repeated calls."""
|
|
lines = _make_records(5, width=30)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
c1 = JsonlFileChunker(max_chars=100)
|
|
_, chunks1 = _run(c1.chunk(path))
|
|
c2 = JsonlFileChunker(max_chars=100)
|
|
_, chunks2 = _run(c2.chunk(path))
|
|
assert [c.id for c in chunks1] == [c.id for c in chunks2]
|
|
print("✓ test_hash_id_deterministic passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Edge cases
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_max_chars_floor():
|
|
"""max_chars is clamped to at least 64."""
|
|
c = JsonlFileChunker(max_chars=1)
|
|
assert c.max_chars == 64
|
|
print("✓ test_max_chars_floor passed")
|
|
|
|
|
|
def test_max_overlap_floor():
|
|
"""Negative max_overlap_chars is clamped to 0."""
|
|
c = JsonlFileChunker(max_overlap_chars=-10)
|
|
assert c.max_overlap_chars == 0
|
|
print("✓ test_max_overlap_floor passed")
|
|
|
|
|
|
def test_full_coverage():
|
|
"""Every line of the file appears in at least one chunk."""
|
|
lines = _make_records(25, width=40)
|
|
path = _write_jsonl(lines)
|
|
try:
|
|
chunker = JsonlFileChunker(max_chars=150, max_overlap_chars=50)
|
|
_, chunks = _run(chunker.chunk(path))
|
|
covered: set[int] = set()
|
|
for c in chunks:
|
|
covered.update(range(c.start_line, c.end_line + 1))
|
|
assert covered == set(range(1, 26)), f"Missing lines: {set(range(1, 26)) - covered}"
|
|
print("✓ test_full_coverage passed")
|
|
finally:
|
|
os.unlink(path)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Internal algorithm
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_chunk_lines_algorithm_simple():
|
|
"""_chunk_lines produces correct ranges for a simple case."""
|
|
# max_chars floor is 64, so lines must be long enough.
|
|
# Each line = 20 chars + \n = 21 chars. 64 / 21 = 3 lines per chunk.
|
|
chunker = JsonlFileChunker(max_chars=64, max_overlap_chars=0)
|
|
lines = ["a" * 20 + "\n", "b" * 20 + "\n", "c" * 20 + "\n", "d" * 20 + "\n"]
|
|
# 3 lines = 63 chars ≤ 64 → chunk 1: (0,3). 4th line alone → chunk 2: (3,4).
|
|
ranges = chunker._chunk_lines(lines)
|
|
assert ranges == [(0, 3), (3, 4)], f"Got {ranges}"
|
|
print("✓ test_chunk_lines_algorithm_simple passed")
|
|
|
|
|
|
def test_chunk_lines_algorithm_with_overlap():
|
|
"""_chunk_lines with overlap produces overlapping ranges."""
|
|
# Each line = 20 chars + \n = 21 chars. max_chars=64 → 3 lines per chunk.
|
|
# max_overlap_chars=22 → 1 line of overlap (21 chars ≤ 22).
|
|
chunker = JsonlFileChunker(max_chars=64, max_overlap_chars=22)
|
|
lines = ["a" * 20 + "\n"] * 6
|
|
ranges = chunker._chunk_lines(lines)
|
|
assert len(ranges) >= 2, f"Expected ≥2 ranges, got {ranges}"
|
|
# Verify overlap: range[i].end >= range[i+1].start
|
|
for i in range(len(ranges) - 1):
|
|
assert ranges[i][1] >= ranges[i + 1][0], f"No overlap between range {i} {ranges[i]} and {i+1} {ranges[i+1]}"
|
|
# Last range reaches end.
|
|
assert ranges[-1][1] == 6
|
|
print("✓ test_chunk_lines_algorithm_with_overlap passed")
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Runner
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
if __name__ == "__main__":
|
|
test_empty_file()
|
|
test_blank_lines_only()
|
|
test_single_line()
|
|
test_unicode_line_separator_inside_json_string_is_not_a_record_boundary()
|
|
test_all_lines_fit_in_one_chunk()
|
|
test_multiple_chunks_no_overlap()
|
|
test_line_aligned_no_intra_line_split()
|
|
test_overlap_lines_shared()
|
|
test_zero_overlap()
|
|
test_overlap_capped_by_max_overlap_chars()
|
|
test_single_line_exceeds_max()
|
|
test_byte_mode()
|
|
test_invalid_mode_raises()
|
|
test_file_chunk_properties()
|
|
test_file_node_properties()
|
|
test_hash_id_deterministic()
|
|
test_max_chars_floor()
|
|
test_max_overlap_floor()
|
|
test_full_coverage()
|
|
test_chunk_lines_algorithm_simple()
|
|
test_chunk_lines_algorithm_with_overlap()
|
|
print("\n所有测试通过!")
|