Commit graph

3 commits

Author SHA1 Message Date
jinliyl
cae613d1c4
refactor(auto-fin): one note per topic, merged by a separate digest step (#558)
* fix(auto-fin): parse topic IDs from fenced JSON replies

* fix(auto-fin): limit report agent tool calls in prompt

* refactor(auto-fin): research news by topic before market open

* fix(config): update default model version for claude_code backend

- Change model version from qwen3.8-max to qwen3.7-plus
- Use environment variable LLM_MODEL_NAME to allow override
- Ensure backend configuration reflects updated model setting

* refactor(auto-fin): write one note per topic before the daily digest

The merge step did two jobs at once: it researched every topic and
combined the results into a single report. Split it the way daily-paper
separates analysis from its brief, so each topic earns a durable note of
its own.

- auto_fin_research_step writes one note per topic that had relevant
  news, tagged `kind: auto-fin-topic` and `topic` in frontmatter
- auto_fin_digest_step merges those notes into the day's brief with no
  tools of its own and appends a `## 主题详解` section linking back to
  each note
- a same-day rerun finds a topic's note by its `topic` frontmatter and
  replaces it in place, deleting the old file when the title changed
- base.py now owns the shared Markdown layer: title sanitizing, report
  normalizing, wikilink validation, note lookup, atomic frontmatter
  writes, and change tracking, so both steps share one write path
- the DingTalk step maps `auto_fin_digest_path` to `markdown_path`
  explicitly instead of relying on whichever step ran last
- drop the unused AutoFinTopicOutput schema and read `job_tools` from
  the step config rather than hardcoding `search`

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(dingtalk): surface rejection details when delivery fails

A failed group send only reported HTTPStatusError, so an operator had to
reproduce the request by hand to learn why DingTalk refused it. Include
the status code and the whitelisted error keys from the response body in
both the log line and the raised RuntimeError.

Only `code`, `message`, and `requestid` are reported: the request body
carries the message content and credentials, so an error response that
echoes it back must not reach the log. Detail is truncated to 200 chars.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(test): make the suite green on CI

- Point the cookbook claude_code model assertion at qwen3.7-plus, the
  default commit 9404e600 set, so the pre-existing red stops blocking
- Satisfy pylint on the auto-fin tests: prefer implicit booleaness for
  the recorded Agent calls and drop an unused tmp_path fixture

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(auto-fin): name notes after the topic, not the Agent title

The research Agent returned a whole paragraph as its title; that became a
filename and blew past the filesystem's 255-byte name limit, failing with
ENAMETOOLONG inside resolve_note_path. Topics are configured values, so
they are short and predictable - use them for file names and keep the
Agent title in frontmatter.

- Name topic notes after the topic and the digest after the run date
- Fold a byte budget into normalize_title as a safety net for long topics
- Take an AutoFinReportOutput in _write_report instead of loose fields
- Ask both prompts for a short title now that it is display-only

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: isolate per-topic research failures and scope frontmatter reads

A single failing topic used to fail the whole cron job and discard the news
already gathered for the topics that had not run yet -- the 09-19 09:24 run
lost its robot notes that way. Research now logs the failure, continues with
the remaining topics, and only fails the run when no topic produced a note.

frontmatter_read was the only frontmatter step without the _allowed_paths
check that read, write, edit, and frontmatter_update already honour, so an
Agent scoped to one file could still read another file's metadata.

- Isolate per-topic research failures and report them as failed_topics
- Fail loudly when every topic fails so an empty brief is never sent
- Apply _check_path_permission in FrontmatterReadStep

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(auto-fin): keep hand-edited notes and wikilink delimiters from breaking the run

Addresses three review findings on the topic-per-note rework.

- `find_note` and `read_note` now skip a note whose YAML frontmatter does
  not parse. A hand-edited note in the day directory raised
  `yaml.parser.ParserError`, which per-topic isolation surfaced as
  "Auto Fin research failed for every topic" and took the run down with it.
- `normalize_title` also strips `[`, `]` and `#`, which `WikilinkHandler`
  treats as target delimiters. `AI[算力]` used to emit a trailer link the
  parser could not read at all, and `C#` resolved to `.../C` plus an anchor.
- `_write_report` returns the body it actually wrote, and both callers
  propagate it, so the digest answer and the note handed to the digest Agent
  no longer carry links that validation had already downgraded on disk.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-20 14:52:56 +08:00
xyf2020
06fb46fa48
feat(tags): add configurable tag indexing and filtered hybrid search (#530)
* Add optional tag generation and normalization to auto memory

* Add tag index components and clean up temporary JSONL files

* Preserve tag index state when reconciliation fails

* Refactor and streamline application implementation

* Fix pylint C1803 warnings in tag normalization tests

* Document optional tag index configuration

* Make tag index failures non-blocking and disable auto-memory tags

* Add configurable tag indexing and tag listing

* Add tag-filtered hybrid search with exact candidate ranking

* Remove obsolete generated files

* Rename tag index key to tag_key and reject reserved fields

* Extract automatic tagging into a dedicated step

* Restrict frontmatter updates to authorized keys

* Refine tag filtering and automatic memory tagging

* Require underscore-separated tags in auto-tag prompts

- Forbid spaces in tags and require underscores (e.g. sam_altman) in both
  English and Chinese auto_tag prompts, with English examples switched to
  English entities (OpenAI, gold)
- Drop prompt-string assertions superseded by the new tagging rule
- Merge construction/runtime tag_key validation tests into one parametrized case

* Make max_tags_per_file configurable in auto-tag step

* Consolidate tag index tests

* Fix search test fixture lint warnings

* Align tag contracts and index health behavior

* Fall back when tag index is unavailable

---------

Co-authored-by: jinli.yl <jinli.yl@alibaba-inc.com>
2026-09-10 18:01:09 +08:00
imrewce
58276f740b
fix(file_io): auto appending suffix for all related steps (#430)
* fix(file_io): auto appending suffix for all related steps

* fix(file_io): covering boundary cases of potential directory input
2026-08-11 11:09:28 +08:00