* Add optional tag generation and normalization to auto memory
* Add tag index components and clean up temporary JSONL files
* Preserve tag index state when reconciliation fails
* Refactor and streamline application implementation
* Fix pylint C1803 warnings in tag normalization tests
* Document optional tag index configuration
* Make tag index failures non-blocking and disable auto-memory tags
* Add configurable tag indexing and tag listing
* Add tag-filtered hybrid search with exact candidate ranking
* Remove obsolete generated files
* Rename tag index key to tag_key and reject reserved fields
* Extract automatic tagging into a dedicated step
* Restrict frontmatter updates to authorized keys
* Refine tag filtering and automatic memory tagging
* Require underscore-separated tags in auto-tag prompts
- Forbid spaces in tags and require underscores (e.g. sam_altman) in both
English and Chinese auto_tag prompts, with English examples switched to
English entities (OpenAI, gold)
- Drop prompt-string assertions superseded by the new tagging rule
- Merge construction/runtime tag_key validation tests into one parametrized case
* Make max_tags_per_file configurable in auto-tag step
* Consolidate tag index tests
* Fix search test fixture lint warnings
* Align tag contracts and index health behavior
* Fall back when tag index is unavailable
---------
Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com>
* refractor(proactive): upgrade proactive feature with disentangled job and steps
* refactor(proactive): apply audit fixes
- rename read-side job 'proactive' -> 'proactive_read' (less confusing vs the refresh pipeline)
- drop dedicated agent_wrapper.proactive; extraction reuses the default wrapper
- simplify schema: remove unused ProactiveExtractOutput/TopicUpdate, drop resource_paths
- extract no longer scans resource/ directly (daily notes already carry resource content)
- update tests and docs accordingly
* feat(proactive): strict extract-output gate and prompt total budget
- parse_extract_reply now requires a contract section (follow_ups/extends/updates
as a list); non-empty replies with misspelled section names trigger the
existing one-shot retry instead of silently checkpointing changed files
- pack_paths gains max_total_chars; extract packs newest daily material first,
keeps the first file on overflow, and records omitted files in a trailer
(default budget 300000 chars, configurable via max_total_chars)
- tests: schema gate unit, schema-error retry e2e, budget unit + e2e
* feat(proactive): add scenario-card plan step and generative agenda step
* feat(proactive): digest-personal profile personalization and leaner LLM contract
- extract/plan/agenda now draw a user profile block from <digest_dir>/personal/*.md
(frontmatter description + body excerpt, per-file budget, profile.md fallback)
- all daily access honours the configured daily_dir (prompt paths parameterized,
config-driven fallbacks) so workspaces using e.g. memory/ work unchanged
- schema trim: drop dead fields errors/material_paths, carry_forward_all -> count
- shrink LLM output contract: new topics emit title/reason/confidence/paths only;
keywords removed end-to-end, evidence derived from paths[0] (updates keep it)
* fix(proactive): skip checkpoint when extract reply stays unusable after retry
Two consecutive unparseable replies now short-circuit the round without
checkpointing, so the same material is retried next round instead of being
silently consumed (closes the residual audit #1 gap: the structural gate
detected schema-wrong output but a double failure still checkpointed).
* fix(proactive): replace running bool with reference-counted job activity tracker for the idle gate
* refactor(proactive): remove job activity tracking and idle gate, restore job tree to upstream
* fix(proactive): address second audit round (readonly reader, mtime checkpoint, wider fallbacks, profile containment, horizon content, expiry boundary)
* refactor(dream): strip interests.yaml ownership from dream, proactive is now the sole writer
* refactor(dream): separate proactive topic generation
* ci: update renamed auto dream smoke test
* fix(proactive): complete refresh migration and docs
---------
Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com>
* feat: refine local-first research workflows
* fix: delegate structured output tool choice
* refactor(auto-fin): fetch and filter rolling CLS news
* fix(auto-fin): keep imports portable across platforms
* feat(auto-fin): expose CLS fetch controls
* fix(auto-fin): propagate configurable news window
* feat(auto_fin): normalize hybrid wikilinks in report body
- Add _normalize_hybrid_wikilinks method to remove redundant Markdown destinations
- Use regex to identify hybrid wikilinks with optional destinations
- Replace redundant destinations with simpler wikilink format for clarity
- Ensure normalization is failure-safe with exception handling and logging
- Update report body normalization process to apply hybrid wikilink fix
- Add unit tests to verify correct normalization and failure safety behavior
* fix(dream): serialize integration with application-wide asyncio lock
- Add application-wide asyncio.Lock to serialize digest writes during integration
- Update _snapshot_digest to capture metadata per bucket
- Validate bucket association when recovering from file changes
- Add tests ensuring recovery only from the correct bucket
- Add tests confirming integration lock is shared across application context
- Enhance strict topic YAML loading validation in dream utils
- Add tests for strict topic loading rejecting invalid or lossy fields
* fix(cookbook): enable configurable job_tools for digest and merge steps
- Update daily_cookbook.yaml to add job_tools: [memory_search, read] in digest steps
- Modify DailyPaperDigestStep to read job_tools from kwargs instead of fixed list
- Modify AutoFinMergeStep to similarly read job_tools from kwargs
- Update tests to pass job_tools explicitly when invoking these steps
- Remove hardcoded _TOOLS constants and replace with dynamic job_tools handling
* fix: retry incomplete dream receipts
* perf(pdf): increase max PDF pages limit from 20 to 35
- Updated configuration max_pdf_pages from 20 to 35 in daily_cookbook.yaml
- Modified code to extract up to 35 pages instead of 20 in analyze.py
- Updated README and README_ZH to document the increased max_pdf_pages
- Adjusted unit test assertions to reflect new max_pdf_pages limit of 35
* fix memory integration and daily paper links
* docs clarify cookbook tool usage
* feat: simplify local links and support line anchors
* fix: align line anchor tests with CI lint
* fix: preserve local links across file moves
* fix: encode markdown paths when rewriting links
* refactor(read): keep explicit line range parameters
* fix: simplify legacy link predicate compatibility
* docs: align local link behavior with implementation
* fix: skip unsupported markdown destination escapes
* fix: normalize workspace link paths across platforms
* fix: bound markdown link scanning
* fix: keep local link processing linear
* docs: clarify permissive markdown link parsing
* fix: handle local link processing failures
* refactor: limit file links to wikilink syntax
* docs: align wikilink contract with implementation
* fix: normalize dream and neighbor paths on Windows
* fix: resolve workspace path for neighbor expansion
* chore(logging): change info logs to debug level for data loading operations
- Changed stopwords loading log from info to debug level
- Changed file catalog nodes loading log from info to debug level
- Changed file graph nodes loading log from info to debug level
* feat(dream): add dream schema definitions and enum for auto-dream functionality
- Add DreamBucketEnum with procedure, personal, and wiki values
- Create comprehensive dream-related Pydantic models including DreamUnit,
DreamTopic, DreamExtractOutput, IntegrateOutcome, TopicSelectionOutput,
ProactiveResult, and DreamState
- Move schema definitions from local step module to shared schema package
- Update dream extraction and integration steps to use new enum-based
bucket validation
- Initialize digest directories for each dream bucket type
- Enhance embedding store health check with workspace directory logging
* refactor(tests): update DreamState import path in test_auto_dream.py
- Move DreamState import from reme.steps.evolve.dream.schema to reme.schema
- Maintain same functionality with updated module reference
- Align import with new schema location in project structure
* fix(core): update version number to 0.4.0.1
- Incremented version from 0.4.0.0 to 0.4.0.1 in __init__.py
* fix(index): remove stopwords path from tokenizer config and add keyword index repair
- Remove stopwords_path from tokenizer config to prevent index forking by install path
- Add _sync_keyword_index_from_chunks method to repair keyword index when persisted state mismatches
- Implement test for keyword index repair from persisted chunks when missing
- Add test to verify tokenizer fingerprint ignores stopwords absolute path
- Update version from 0.4.0.1 to 0.4.0.2
* feat(dream): add scan_days parameter to dream extraction process
- Add scan_days configuration option to default.yaml with default value of 2
- Implement recent_dates utility function to calculate date ranges for scanning
- Modify DreamExtractStep to scan multiple days based on scan_days parameter
- Update dream extraction to process files across multiple dates instead of single day
- Extend DreamState schema to include dates and scan_days fields
- Update DreamTopicsStep to handle multi-day topic processing
- Modify finish step to checkpoint files from all scanned dates
- Add comprehensive tests for multi-day scanning functionality
- Update prompt templates to include scan dates information
- Refactor topics writing logic to target specific date rather than current date
* docs: rename vault_dir to workspace_dir in documentation and examples
* refactor(extract): format long method call across multiple lines
* refactor(extract): format system prompt parameters for better readability