ReMe/tests/unit/test_source_format_merge.py
xyf2020 5bc46c88b6
feat(benchmark): enhance session memory retrieval and isolate benchmark assets (#409)
* chore(benchmark): isolate dataset/workspaces/results per benchmark

- Move shared benchmark/{datasets,memory_workspaces,results} into per-benchmark subdirs benchmark/<name>/{dataset,workspaces,results}
- Update beam/longmemeval config.yaml and run.py path defaults
- Relocate longmemeval download.py to benchmark/longmemeval/ (downloads into dataset/ subdir); inline dataset download docs into README
- Update .gitignore: benchmark/*/{dataset,workspaces,results}/
- Move result-{beam,longmemeval}.md to benchmark/results_md/ and drop result- prefix; update README links
- Fix stale path refs in llm_judge.py and logs/demo_search_format.py

* feat(benchmark): add read tool to agentic answer and update BEAM results

- Add 'read' to job_tools in BaseAgenticAnswerStep for file reading capability
- Document read tool usage in lme/agentic_answer.yaml system prompt
- Update result-beam.md with latest evaluation scores (OVERALL: 0.623/0.580)

* feat(auto_memory): add source line-number markers for note traceability

- Add _format_history hook in AutoMemoryStep with line-number annotation
- Override in BeamAutoMemoryStep to prefix each turn with [Ln] for citation
- Add session_file variable to prompt templates for source marker paths
- Simplify repeated extraction rules by referencing system prompt
- Enhance agentic_answer search strategy (multi-search, read tool hint)
- Add warning log on ReadStep failure

* feat(beam): enhance auto_memory with source markers and pilot ingest tooling

* refactor(beam): rename max_chunk_words to max_segment_words, drop one-off pilot scripts

* feat: add CompressorStep and search_v2 dual-mode session compression

- Add CompressorStep (reme/steps/evolve/compressor.py) for direct LLM
  text compression with optional query-guided relevance filtering
- Extend search_v2_step to support query-aware and query-independent
  session transcript compression via _compress injected kwargs
- Refactor _source_format.py: split into render_chunk_entries +
  join_chunk_entries; session chunks now render line-aligned with
  L<n>: prefixes for verbatim/compressed parity
- Add JOB_TOOLS and INJECTED_JOB_KWARGS to BaseAgenticAnswerStep for
  per-subclass tool and parameter injection
- LmeAgenticAnswerStep injects _search._compress payload to enable
  query-aware compression during benchmark evaluation
- Record compression ablation results in result-longmemeval.md
- Add unit tests for CompressorStep and search compression paths

* refactor(compress): relax session compression to lenient format-preserving strategy and update LME results

* refactor(benchmark): make session compression config-driven via compress_session flag

Move session-transcript compression from LME hard-coded injection to a
runtime context flag set by evaluation.compress_session in each
benchmark config. Compression is off by default for both BEAM and LME,
and BaseAgenticAnswerStep now conditionally injects the _search compress
payload only when the flag is truthy.

* feat(lme/auto_memory): add source attribution markers with line numbers

Add _format_history to annotate each turn with [Ln] line numbers and
expose {session_file} in prompts so the agent can emit bare wikilink-style
source markers like [[session/dialog/s1.jsonl#L1-L2,L5-L6]] at the end
of factual entries. Consolidate the per-prompt body/format rules into
references to the system prompt to avoid drift, and add frontmatter-
protection guidance for the edit tool.

* feat: improve agentic answer prompt and update beam 100K results

- Strengthen abstention rule: prohibit extrapolation from related but
  non-direct evidence
- Add multi-angle search after preliminary answer to check for
  conflicting/supplementary/updated information
- Add max-iteration fallback to 'Information not found'
- Update beam.md with 100K results (agentscope 2.0.4.post1, from scratch)
  including per-type token consumption and memory construction stats
- config.yaml: 100K dataset, 20 workers for BEAM evaluation
- run.py: add memory construction token usage tracking (default agent)
- Overall: 0.635 → 0.654 (+0.019), contradiction_resolution: 0.338 → 0.478
  (+0.140), abstention: 0.500 → 0.525 (+0.025)

* feat(read): add session-aware formatting for read tool and update BEAM eval

- Add truncate_session_output in _file_io.py to render jsonl session
  lines as [speaker @ time] content before byte-budget truncation
- Add read_step_format_session flag to ReadStep, honoring injected
  job kwargs (precedence) and YAML fallback
- Inject read_step_format_session=True into BaseAgenticAnswerStep
  so agentic answer reads render session transcripts human-readably
- Refine BEAM agentic_answer prompt: continue multi-angle search
  after preliminary answer, forbid fabrication/extrapolation
- Update BEAM config to 1M variant and add sequential 100K-eval /
  1M-build shell script
- Refresh benchmark/results_md/beam.md with latest results

* chore(config): disable expand_links in beam and lme search_v2 configs

* refactor(beam): drop one-off sequential 100K-eval-then-1M-build script

* fix(benchmark): add compressor job to beam config and fix BEAM clone instructions

- Add compressor job and compressor as_llm component to reme/config/beam.yaml
  (aligned with lme.yaml) so that compress_session: true works for BEAM
- Add graceful degradation guard in search_v2._compress_session_entries:
  when the compressor job is missing from the active config, log a warning
  and skip compression instead of raising 'Job compressor not found'.
  Skipped when there is no app_context so unit tests mocking run_job still
  drive compression behavior.
- Fix BEAM download instructions in README.md/README_ZH.md: add mkdir -p
  before cd benchmark/beam/dataset (the directory is gitignored and absent
  in a fresh clone)

* fix(steps): guard compressor exceptions and fix ReadStep boolean override

1. search_v2: catch per-entry exceptions from run_job('compressor') inside
   compress() so asyncio.gather never propagates a compressor failure (e.g.
   temporary LLM outage). The failing entry keeps its original body while
   remaining entries are still compressed, preserving already-retrieved
   search results.

2. read: replace 'context_value or yaml_value' with an existence check so
   that a runtime-injected False can explicitly disable a YAML-true
   read_step_format_session flag.

Add focused unit tests for both paths.

* fix(search_v2): use existence check for strict_date_filter boolean override

Replace 'context_value or yaml_value' with an existence-based check so
that a runtime-injected False can explicitly disable a YAML-true
strict_date_filter flag, consistent with the read_step_format_session fix.

* refactor(search): simplify strict_date_filter fallback to truthiness-or

* style(test): rename unused param to satisfy pylint W0613

---------

Co-authored-by: sa-buc <jiangniurou.xyf@dail-algo011164204033.ET135>
2026-08-05 19:23:42 +08:00

149 lines
6.3 KiB
Python

"""Unit tests for session-chunk merging via ``merge_session_chunk_intervals``.
Session chunks (``*.jsonl`` under the dialog dir) whose line ranges overlap,
contain one another, or are adjacent are merged into their union by the caller
(``merge_session_chunk_intervals``) before rendering with
``render_chunk_entries`` + ``join_chunk_entries``. Bodies here are plain
(non-``Msg``) text lines, which pass through stripped — letting these tests
assert the union content and line order directly.
"""
from reme.schema import FileChunk
from reme.steps.index._source_format import join_chunk_entries, merge_session_chunk_intervals, render_chunk_entries
_DIALOG_DIR = "session"
def _render(chunks: list[FileChunk], dialog_dir: str, **kwargs) -> str:
"""Merge session chunks, render entries, then join — mirroring the search steps."""
return join_chunk_entries(
render_chunk_entries(merge_session_chunk_intervals(chunks, dialog_dir), dialog_dir, **kwargs),
)
def _chunk(start: int, end: int, text: str, score: float = 1.0, path: str = "session/s1.jsonl") -> FileChunk:
return FileChunk(path=path, start_line=start, end_line=end, text=text, scores={"score": score})
def test_overlapping_session_chunks_merge_into_union_without_duplicates():
"""Overlapping ranges collapse to one passage; the shared line is shown once, in order."""
a = _chunk(1, 3, "m1\nm2\nm3\n")
b = _chunk(3, 5, "m3\nm4\nm5\n")
answer = _render([a, b], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\nm4\nm5"
def test_contained_session_chunk_is_absorbed_by_the_larger_range():
"""When one range fully contains another, only the union (the larger) is shown."""
big = _chunk(1, 5, "m1\nm2\nm3\nm4\nm5\n")
small = _chunk(2, 4, "m2\nm3\nm4\n")
answer = _render([big, small], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\nm4\nm5"
def test_adjacent_session_chunks_merge_end_plus_one_equals_next_start():
"""Gap-free consecutive chunks (prev.end + 1 == next.start) merge into one union."""
a = _chunk(1, 3, "m1\nm2\nm3\n")
b = _chunk(4, 6, "m4\nm5\nm6\n")
answer = _render([a, b], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\nm4\nm5\nm6"
def test_session_chunks_with_a_gap_are_not_merged():
"""A missing line between ranges (start > end + 1) keeps the chunks separate."""
a = _chunk(1, 3, "m1\nm2\nm3\n")
b = _chunk(5, 6, "m5\nm6\n") # line 4 missing -> not adjacent
answer = _render([a, b], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\n\nm5\nm6"
def test_session_chunks_from_different_files_are_not_merged():
"""Overlapping ranges in different session files must stay separate."""
a = _chunk(1, 3, "A1\nA2\nA3\n", path="session/s1.jsonl")
b = _chunk(2, 4, "B2\nB3\nB4\n", path="session/s2.jsonl")
answer = _render([a, b], _DIALOG_DIR, include_source=False)
assert answer == "A1\nA2\nA3\n\nB2\nB3\nB4"
def test_non_session_chunks_are_never_merged():
"""Non-transcript chunks (not ``*.jsonl`` under the dialog dir) pass through untouched."""
a = _chunk(1, 3, "M1\nM2\nM3\n", path="daily/a.md")
b = _chunk(2, 4, "M2\nM3\nM4\n", path="daily/a.md")
answer = _render([a, b], _DIALOG_DIR, include_source=False)
assert answer == "M1\nM2\nM3\n\nM2\nM3\nM4"
def test_merge_preserves_line_order_regardless_of_input_rank_order():
"""A later, higher-ranked chunk does not reorder union content; lines stay chronological."""
later = _chunk(3, 5, "m3\nm4\nm5\n", score=9.0)
earlier = _chunk(1, 3, "m1\nm2\nm3\n", score=1.0)
# Higher-scored later-range chunk is listed first (as a ranker would).
answer = _render([later, earlier], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\nm4\nm5"
def test_merged_header_spans_the_union_range_and_keeps_best_score():
"""With source headers, the merged entry reports the union range and the top score."""
a = _chunk(1, 3, "m1\nm2\nm3\n", score=2.0)
b = _chunk(3, 5, "m3\nm4\nm5\n", score=7.0)
answer = _render([a, b], _DIALOG_DIR, include_source=True)
assert answer.count("==========") == 2 # exactly one header (open + close markers)
assert "session/s1.jsonl:1-5" in answer
assert "score=7.0000" in answer
def test_separate_intervals_in_same_file_stay_separate():
"""Two disjoint interval clusters in one file yield two merged units, ordered by line."""
a = _chunk(1, 2, "m1\nm2\n")
b = _chunk(3, 4, "m3\nm4\n") # adjacent to a -> merges with a into 1-4
c = _chunk(10, 11, "m10\nm11\n") # far away -> separate
answer = _render([a, b, c], _DIALOG_DIR, include_source=False)
assert answer == "m1\nm2\nm3\nm4\n\nm10\nm11"
def test_same_file_units_stay_adjacent_and_sorted_even_when_interleaved_by_rank():
"""Two disjoint units of one session file are grouped together and ordered by
``start_line``, even when a different file is ranked between them and the
lower interval was ranked last."""
s1_high = _chunk(10, 12, "S1x\nS1y\nS1z\n", score=9.0, path="session/s1.jsonl")
s2_mid = _chunk(1, 3, "S2a\nS2b\nS2c\n", score=5.0, path="session/s2.jsonl")
s1_low = _chunk(1, 3, "S1a\nS1b\nS1c\n", score=1.0, path="session/s1.jsonl")
# Rank order (as a ranker would emit, by score desc): s1[10-12], s2[1-3], s1[1-3].
answer = _render([s1_high, s2_mid, s1_low], _DIALOG_DIR, include_source=False)
# s1's two units are adjacent and sorted by start_line (1-3 before 10-12),
# placed at s1's earliest rank (0), so the whole s1 block precedes s2.
assert answer == "S1a\nS1b\nS1c\n\nS1x\nS1y\nS1z\n\nS2a\nS2b\nS2c"
def test_same_file_units_adjacency_with_source_headers():
"""Header view: same-file units are contiguous and ascending; other files follow."""
s1_high = _chunk(10, 12, "S1x\nS1y\nS1z\n", score=9.0, path="session/s1.jsonl")
s2_mid = _chunk(1, 3, "S2a\nS2b\nS2c\n", score=5.0, path="session/s2.jsonl")
s1_low = _chunk(1, 3, "S1a\nS1b\nS1c\n", score=1.0, path="session/s1.jsonl")
answer = _render([s1_high, s2_mid, s1_low], _DIALOG_DIR, include_source=True)
headers = [line for line in answer.splitlines() if line.startswith("==========")]
assert headers[0].split(" [")[0] == "========== session/s1.jsonl:1-3"
assert headers[1].split(" [")[0] == "========== session/s1.jsonl:10-12"
assert headers[2].split(" [")[0] == "========== session/s2.jsonl:1-3"