diff --git a/cookbook/auto-fin/README.md b/cookbook/auto-fin/README.md index 1baa1595..cf92ccc9 100644 --- a/cookbook/auto-fin/README.md +++ b/cookbook/auto-fin/README.md @@ -2,323 +2,214 @@ [中文](README_ZH.md) -Auto Fin is a local-first, file-native ETF event-research workflow. It identifies market events in CLS news, selects -related liquid ETFs, studies similar historical events and subsequent returns, and produces a Chinese research report. +Auto Fin is a local-first, file-native ETF event-research workflow. It collects CLS news and market data through +Tushare, identifies current news related to a configured ETF list, retrieves comparable events from local ReMe memory, +calculates observed post-event returns, and writes a Chinese research report. -> Auto Fin provides event research and holding-period references only. It is not investment advice, does not -> connect to a broker, and does not place or simulate trades. - -## Capabilities - -- Download CLS news through Tushare and maintain up to 360 days of traceable local news records. -- Rank ETF candidates by previous-trading-day turnover, then select representative ETFs related to current events. -- Search ReMe memory and local news files for similar historical events, with strict source-path and news-ID checks. -- Calculate adjusted D1–D10 historical returns in deterministic code instead of asking an Agent to invent numbers. -- Let an Agent judge event similarity, then calculate weights, expected returns, and a reference holding period in code. -- Save readable Markdown and structured JSON/JSONL artifacts, refresh the daily index, and optionally deliver the report - to DingTalk. +> Auto Fin is for event research and holding-period reference only. It is not investment advice, does not connect to a +> broker, and does not place or simulate trades. The workflow is assembled by [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml). Its public schemas are in -[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py), and its steps are in +[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py), and its four steps are in [`reme/steps/cookbook/auto_fin/`](../../reme/steps/cookbook/auto_fin/). ## Quick start -Auto Fin requires Python 3.11 or newer, the `core` dependencies, a Tushare token, and credentials for the configured -Claude Code-compatible endpoint. - -From the repository root: +Auto Fin requires Python 3.11 or newer, the `core` dependencies, a Tushare token, and an available AgentScope LLM. ```bash python -m pip install -e ".[core]" export TUSHARE_TOKEN="your-tushare-token" -export CLAUDE_CODE_API_KEY="your-api-key" +export LLM_API_KEY="your-api-key" reme start config=daily_cookbook job=auto_fin ``` -The built-in configuration uses `qwen3.7-max` through DashScope's Anthropic-compatible endpoint. Override these -variables to use another compatible model or provider: +The built-in LLM component defaults to `qwen3.7-plus`. `LLM_BASE_URL` has no built-in value, so set it when your +provider requires a custom OpenAI-compatible endpoint. The model and endpoint can be overridden with +`LLM_MODEL_NAME` and `LLM_BASE_URL`. + +The default workspace is `reme_workspace/` beneath the process working directory. Override it with +`DAILY_PAPER_WORKSPACE_DIR`; Auto Fin and Daily Paper share this setting. + +Dates and times use `Asia/Shanghai`. An explicit `date` must be today's date: ```bash -export CLAUDE_CODE_MODEL_NAME="your-model" -export CLAUDE_CODE_BASE_URL="https://your-anthropic-compatible-endpoint" +reme start config=daily_cookbook job=auto_fin date=2026-08-07 ``` -The default workspace is `reme_workspace/`. This standalone cookbook shares its workspace setting with the daily-paper -workflow: +Auto Fin checks the SSE trading calendar first and skips the whole workflow on a closed market day. -```bash -export DAILY_PAPER_WORKSPACE_DIR="/absolute/path/to/reme-workspace" -``` - -To deliver the final Markdown report to DingTalk, set: - -```bash -export DINGTALK_APP_KEY="your-app-key" -export DINGTALK_APP_SECRET="your-app-secret" -export DINGTALK_ROBOT_CODE="your-robot-code" -export DINGTALK_CONVERSATION_IDS="conversation-id-1,conversation-id-2" -``` - -DingTalk delivery is skipped when the required values are empty. - -Dates and times use `Asia/Shanghai`. The optional `date` must be the current date: - -```bash -reme start config=daily_cookbook job=auto_fin date=2026-07-25 -``` - -To refresh every configured news day instead of reusing valid historical files: - -```bash -reme start config=daily_cookbook job=auto_fin force=true -``` - -This may issue many Tushare requests. A normal run reuses valid historical news files and always refreshes today's file. - -### Optional SSH proxy - -The outbound proxy is disabled by default. To enable it, uncomment `components.outbound_proxy.default` in -`daily_cookbook.yaml`, configure non-interactive SSH authentication, and set: - -```bash -export REME_PROXY_IP="your-ssh-proxy-host" -export REME_PROXY_ACCOUNT="your-ssh-account" -``` - -## How it works - -```mermaid -flowchart LR - A[Resolve run date and cutoff] --> B[Maintain CLS news files] - B --> C[Resolve previous A-share trading day] - C --> D[Build current event window] - D --> E[Filter liquid ETF candidates] - E --> F[Agent selects related ETFs] - F --> G{For each ETF} - G --> H[Agent searches historical events] - H --> I[Code resolves original news] - I --> J[Code calculates adjusted D1-D10 returns] - J --> K[Agent judges similarity] - K --> L[Code calculates weighted forecast] - L --> G - G --> M[Agent writes the combined report] - M --> N[Write artifacts and refresh daily index] - N --> O[Optional DingTalk delivery] -``` - -The top-level job contains four Auto Fin steps: - -| Step | Responsibility | Agent | -|-------------------------|-------------------------------------------------------------|-------| -| `auto_fin_data_step` | Maintain news files and resolve the previous trading day | No | -| `auto_fin_topic_step` | Build inputs and select related ETFs and current events | Yes | -| `auto_fin_history_step` | Orchestrate historical research and market analysis per ETF | Yes | -| `auto_fin_merge_step` | Validate results and produce the final Markdown report | Yes | - -For each selected ETF, `auto_fin_history_step` dispatches: - -- `auto_fin_history_search_step`, which asks the Agent for historical news references and then resolves the original - records and calculates their returns in code. -- `auto_fin_market_step`, which asks the Agent only for similarity judgments and then calculates weights and forecasts - in code. - -Agents handle semantic judgments; deterministic code handles source validation and financial calculations. - -## Data and time boundaries - -### News history - -`auto_fin_data_step` reads CLS news from Tushare's `major_news` endpoint: - -- The default lookback is 360 calendar days, including the run date. -- A valid historical file is reused unless `force=true`. -- Today's file is refreshed through the current `decision_at` on every run. -- Large responses are fetched through recursively split time windows. -- Records are ordered and deduplicated before being written with a stable `news_id`. - -The current event window is: +## Pipeline ```text -(previous A-share trading day at 15:00, decision_at] +Tushare trade calendar + │ + ├─ closed day ──► skip + ▼ +Collect CLS news + configured ETF history + ▼ +Update the ReMe index + ▼ +Select ETF/news relationships with an agent + ▼ +Search local memory for comparable historical news + ▼ +Select same/opposite events with an agent + calculate D1/D2/D3/D5 returns in code + ▼ +Generate report with an agent ──► refresh day index ──► DingTalk (optional) ``` -Each run rebuilds this complete window; midday and evening runs do not use only the increment since the previous run. +| Step | Responsibility | Agent | +|---|---|---| +| `auto_fin_data_step` | Check the trading day, maintain news, and cache configured ETF market history | No | +| `auto_fin_topic_step` | Select direct relationships between today's news and configured ETFs | Yes | +| `auto_fin_history_step` | Retrieve comparable news, validate selections, and calculate observed returns | Yes | +| `auto_fin_merge_step` | Combine prepared evidence and the previous report into the final Markdown | Yes | -### ETF candidates +All three model-facing steps use structured Pydantic output. Agents make semantic judgments; code owns identifier +validation, source resolution, market calculations, and file writes. -The candidate universe combines: +## Data and selection boundaries -- `etf_basic` for currently listed ETFs and their tracked indexes. -- `fund_daily` for turnover on the previous A-share trading day. +### News -Code sorts candidates by turnover, removes duplicates by ETF name and index identity, and provides at most 150 -candidates to the Topic Agent. The Agent may return at most 20 ETFs and must copy every ETF code, name, and news ID from -the generated candidate files. +`auto_fin_data_step` calls Tushare `major_news` with `src="财联社"`. The default lookback is 60 calendar days including +today. Existing files for earlier days are reused, while today's file is always overwritten with news from 00:00 +through the current decision time. Large responses are recursively split when a request returns at least 400 rows. -Turnover is used only to narrow the research universe; it is not a trading signal. +Each item is stored in `daily/YYYY-MM-DD/auto_fin_news.md` with a stable ID made from its publication timestamp and a +short content hash. The current-event set used by the Topic step is today's complete file, not an increment since an +earlier run. -## Historical research and forecasting +### Configured ETFs -### Source resolution +The built-in configuration currently enables: -The History Agent searches by event type, entities, transmission mechanism, and expected direction. It first uses -`memory_search` and may then scan: +- `518880.SH` +- `159530.SZ` +- `512760.SH` -```text -daily/YYYY-MM-DD/auto_fin_news_data.jsonl -``` +Other examples remain commented out in `daily_cookbook.yaml`. For each enabled code, the Data step resolves its name +through `etf_basic`, then pages backward through `fund_daily` and `fund_adj` and rewrites its complete local JSONL +history. A missing ETF name fails the run. -Its output contains only a reason, `news_id`, and workspace-relative `source_path`. Code rejects: +The Topic agent receives only the configured ETF code/name pairs and today's locally stored news. It may retain up to +`current_news_limit_per_etf` valid, unique news references per ETF (10 by default). Unknown ETF codes, unknown news IDs, +empty reasons, duplicates, and ETFs with no accepted event are removed by code. -- Current-window news presented as historical evidence. -- Absolute paths, `..` traversal, or paths outside the workspace. -- Sources not named `auto_fin_news_data.jsonl`. -- Missing files or IDs that do not resolve exactly once. -- Records without a usable publication time, title, or body. +## Historical comparison and returns -Historical Markdown may guide retrieval, but the original news JSONL is the source of truth. +For every accepted current ETF/news pair, `auto_fin_history_step` calls the configured `memory_search` job over the +60-day news window, ending yesterday. `historical_search_limit` controls the maximum search results requested per +current event. Only search hits whose path is named `auto_fin_news.md` contribute candidate IDs; the step rereads the +source Markdown and resolves those IDs before calling the History agent. -### Adjusted returns +The History agent may select at most five candidates by default and labels each relationship `same` or `opposite`. +Code discards unknown or duplicate IDs and empty reasons, then calculates adjusted cumulative returns for D1, D2, D3, +and D5: -For every resolved historical event, code reads `fund_daily` and `fund_adj` and calculates up to ten future closes: +- For an event before 15:00 on a trading day, the adjusted same-day close is the entry; D1 is the next trading close. +- For an event at or after 15:00, the adjusted next-trading-day open is the entry; D1 is that day's close. +- If an entry or horizon cannot be calculated from valid positive prices and adjustment factors, that value is `null`. -- Before 09:30 on a trading day: enter at that day's open. -- From 09:30 until before 15:00: enter at that day's close. -- At or after 15:00, or on a non-trading day: enter at the next trading day's open. -- A daily close later than the current `decision_at` is excluded. +The final agent receives the fixed ETF list, all current and historical evidence, `same`/`opposite` directions, computed +returns, and the most recent earlier `auto_fin.md`. It decides whether the evidence supports a recommendation or an +explicit wait-and-see conclusion; the code does not calculate a score, expected return, or mandatory holding period. -```text -adjusted_entry = raw_entry × entry_adjustment_factor -adjusted_close = raw_close × close_adjustment_factor -cumulative_return = adjusted_close / adjusted_entry - 1 -``` - -Missing prices, factors, trading days, or horizons become explicit limitations. They are never filled with Agent-made -values. - -### Similarity and forecast - -The Market Agent returns semantic similarity in `[-1, 1]`: - -- Positive values mean a similar mechanism and direction. -- Negative values mean a comparable mechanism but opposite direction. -- Zero means no useful relationship. - -Code clamps out-of-range values, ignores zero-similarity events, normalizes weights from absolute similarity, and -reverses the historical return direction for negative matches. Each D1–D10 horizon is calculated from the samples -available at that horizon. The suggested holding period is the positive-return horizon with the highest expected return, -or empty when none is positive. - -The result also records limited samples, missing horizons, conflicting return directions, and other data limitations. It -is a comparison with a small historical sample, not evidence of statistical significance. - -## Output layout +## Outputs ```text reme_workspace/ ├── daily/ │ ├── YYYY-MM-DD.md │ └── YYYY-MM-DD/ -│ ├── auto_fin_news_data.jsonl -│ ├── auto_fin_analysis.jsonl +│ ├── auto_fin_news.md │ └── auto_fin.md └── resource/ + ├── fin/ + │ ├── etfs.json + │ ├── 518880.SH.jsonl + │ └── .jsonl └── YYYY-MM-DD/ - ├── filtered_news.jsonl - ├── filtered_etf.jsonl - ├── auto_fin_topic_output.jsonl - ├── auto_fin_history___output.json - ├── auto_fin_market___output.json - ├── auto_fin_history_output.jsonl + ├── auto_fin_topic_output.json + ├── auto_fin_history_001_output.json + ├── ... + ├── auto_fin_analysis.jsonl └── auto_fin_merge_output.json ``` -Important artifacts: +The daily news and report are user-owned Markdown. `resource/fin/` contains the market cache used for deterministic +return calculations. Date-scoped JSON/JSONL files preserve structured agent replies and prepared analyses. Writes use +same-directory temporary files and atomic replacement; the day index is refreshed after the report is written. -- `auto_fin_news_data.jsonl` is the user-owned source used to resolve historical news. -- `filtered_news.jsonl` and `filtered_etf.jsonl` are bounded inputs for the Topic Agent. -- Per-ETF history files contain resolved source news and code-calculated return paths. -- Per-ETF market files contain code-calculated matches, weights, and D1–D10 forecasts. -- `auto_fin_analysis.jsonl` contains the final structured analysis for every selected ETF. -- `auto_fin.md` is the readable report and DingTalk payload. -- `daily/YYYY-MM-DD.md` is refreshed after report generation so the report is discoverable from the daily index. +## Parameters and defaults -News and reports remain ordinary user-owned files. Resource artifacts and search indexes can be rebuilt. +Public job parameters: -## Configuration +| Parameter | Default | Purpose | +|---|---:|---| +| `date` | `""` | Empty uses today in `Asia/Shanghai`; a value must be strict `YYYY-MM-DD` and equal today | +| `historical_search_limit` | `10` | Maximum `memory_search` results requested for each current event; minimum 1 | -### Job parameters +Relevant job settings in `daily_cookbook.yaml`: -| Parameter | Default | Meaning | -|-----------|-------------:|---------------------------------------------------------| -| `date` | Current date | Strict `YYYY-MM-DD`; only the current date is supported | -| `force` | `false` | Refresh all configured news days | +| Setting | Default | Purpose | +|---|---:|---| +| `etf_codes` | three enabled codes above | Fixed ETF research universe | +| `news_lookback_days` | `60` | Local news and historical-search window | +| `current_news_limit_per_etf` | `10` | Maximum accepted current events per ETF | +| `historical_news_limit` | `5` | Maximum comparable events retained per current event | -### Environment variables +There is no public `force` parameter. Earlier news files are reused, today's news and all configured ETF market files +are refreshed, and same-day report/resource paths are overwritten on each successful run. -| Variable | Required | Meaning | -|-----------------------------|----------|---------------------------------------------------------| -| `TUSHARE_TOKEN` | Yes | News, calendar, ETF daily data, and adjustment factors | -| `CLAUDE_CODE_API_KEY` | Yes | Auto Fin Agent credentials | -| `CLAUDE_CODE_MODEL_NAME` | No | Defaults to `qwen3.7-max` | -| `CLAUDE_CODE_BASE_URL` | No | Anthropic-compatible endpoint | -| `AUTO_FIN_AGENT_BACKEND` | No | Defaults to `claude_code` | -| `AUTO_FIN_PROJECT_PATH` | No | Agent project path; defaults to `..` | -| `REME_PROXY_IP` | No | SSH proxy host; used only when `ssh_http` is enabled | -| `REME_PROXY_ACCOUNT` | No | SSH proxy account; used only when `ssh_http` is enabled | -| `DAILY_PAPER_WORKSPACE_DIR` | No | Standalone cookbook workspace | -| `DINGTALK_*` | No | DingTalk application, robot, and conversation settings | +## Environment and scheduling -Unit tests can inject `tushare_provider` through the runtime context and do not require real credentials. +| Variable | Required | Purpose | +|---|---|---| +| `TUSHARE_TOKEN` | Yes | Trading calendar, CLS news, ETF metadata, prices, and adjustment factors | +| `LLM_API_KEY` | Provider-dependent | Shared AgentScope LLM credentials; config defaults to an empty value | +| `LLM_MODEL_NAME` | No | Defaults to `qwen3.7-plus` | +| `LLM_BASE_URL` | Provider-dependent | OpenAI-compatible endpoint; no built-in default | +| `TUSHARE_MIRROR_URL` | No | Replaces the Tushare SDK HTTP URL after trimming a trailing slash | +| `DAILY_PAPER_WORKSPACE_DIR` | No | Shared standalone cookbook workspace | +| `DINGTALK_*` | No | Optional DingTalk application, robot, and group settings | -### Scheduled jobs +The optional mirror can be configured, for example, as: -`daily_cookbook.yaml` defines: +```bash +export TUSHARE_MIRROR_URL="http://112.124.63.173:4000/tushare" +``` -| Job | Cron | Asia/Shanghai | -|----------------------|---------------|----------------| -| `auto_fin_0930_cron` | `30 9 * * *` | Daily at 09:30 | -| `auto_fin_1145_cron` | `45 11 * * *` | Daily at 11:45 | -| `auto_fin_1800_cron` | `0 18 * * *` | Daily at 18:00 | +`auto_fin_0930_cron`, `auto_fin_1130_cron`, and `auto_fin_1800_cron` run every day at 09:30, 11:30, and 18:00 in +`Asia/Shanghai`. The crons fire on weekends and holidays, but the Data step then skips the remaining workflow when +Tushare reports that the date is not an SSE trading day. Same-day reruns refine the existing report. -These cron expressions do not exclude weekends or market holidays. The workflow resolves the previous A-share trading -day but does not currently skip a run merely because the run date is not a trading day. +To send a completed report, configure `DINGTALK_APP_KEY`, `DINGTALK_APP_SECRET`, `DINGTALK_ROBOT_CODE`, and the +comma-separated `DINGTALK_CONVERSATION_IDS`. With no conversation IDs, delivery is a no-op. -## Agent and security boundaries +## Agent and failure boundaries -The Auto Fin wrapper loads the `tushare-data` skill, exposes the `memory_search` job tool, and defaults to -`bypassPermissions`. Prompts constrain each Agent's role, while code revalidates schemas, ETF identities, source paths, -news references, and calculated values. +Auto Fin and Daily Paper share the tool-free `default` AgentScope wrapper. Built-in and configured job tools are not +exposed to their model calls. Auto Fin itself invokes `memory_search` in deterministic step code; this is not an agent +tool call. The separate interactive `dingtalk_wait` step has its own `bash` and ReMe job-tool allowlist. -The standalone cookbook does not configure an embedding store by default, so `memory_search` normally uses BM25 recall. -Vector and BM25 fusion becomes available only when an embedding store is configured. +The standalone config has no embedding store enabled by default, so `memory_search` uses the available BM25 path; +vector/BM25 fusion requires enabling the commented embedding components. -`bypassPermissions` is not an operating-system sandbox. Review the configured project path, workspace, credentials, and -network boundary before deployment. +Invalid dates, missing credentials or services, invalid structured model output, unknown configured ETFs, missing +market files, and failed memory search stop the job. A market holiday is a successful skip. The workflow has no global +same-date execution lock or cross-file transaction, and repeated successful runs can resend DingTalk notifications. -## Reruns and limitations +## Tests -- Valid historical news files are reused; today's news is always refreshed. -- Outputs use stable per-day paths, so a later same-day run replaces the previous report and resource outputs. -- Auto Fin intentionally has no “report exists, skip” shortcut because its scheduled runs analyze updated news. -- Every successful run attempts DingTalk delivery when configured; notification deduplication is not implemented. -- Missing historical market horizons degrade one sample and are recorded as limitations. -- Invalid dates, missing required services, invalid Agent schemas, unknown ETFs or news IDs, unsafe paths, and - inconsistent cross-step ETF identities fail the job. - -The current implementation does not include stocks, US-market correlation, portfolio accounting, BUY/SELL/HOLD actions, -T+1 execution rules, fees, slippage, broker integration, or real/simulated order execution. - -## Development - -Install development dependencies and run the focused suite: +Focused unit tests mock model and market-data boundaries: ```bash python -m pip install -e ".[dev,core]" -PYTHONPATH=. pytest tests/unit/test_auto_fin.py -v +pytest tests/unit/test_auto_fin.py -v ``` -The unit suite mocks model and market-data boundaries. Tests requiring real Tushare, model, or DingTalk credentials -should be run separately and only with explicit authorization. +Tests requiring real Tushare, LLM, or DingTalk credentials should be run separately and only with explicit +authorization. diff --git a/cookbook/auto-fin/README_ZH.md b/cookbook/auto-fin/README_ZH.md index 35c897c3..38cf5f9e 100644 --- a/cookbook/auto-fin/README_ZH.md +++ b/cookbook/auto-fin/README_ZH.md @@ -2,303 +2,200 @@ [English](README.md) -Auto Fin 是一个本地优先、文件原生的 ETF 事件研究工作流。它从财联社新闻中识别市场事件,选择相关且流动性较好的 -ETF,研究相似历史事件及其后续收益,并生成中文研究报告。 +Auto Fin 是一个 local-first、file-native 的 ETF 事件研究工作流。它通过 Tushare 获取财联社新闻和行情数据,从固定 +ETF 列表中识别与当日新闻相关的标的,利用 ReMe 本地记忆检索可比历史事件,计算事件后的实际收益,并生成中文研究报告。 -> Auto Fin 只提供事件研究和持有时间参考,不构成投资建议。当前实现不连接券商、不提交委托,也不执行模拟交易。 +> Auto Fin 只提供事件研究和持有时间参考,不构成投资建议,不连接券商,也不会执行或模拟交易。 -## 能力 - -- 通过 Tushare 获取财联社新闻,并维护最多 360 天可追溯的本地新闻记录。 -- 按上一交易日成交额筛选 ETF 候选,再选择与当前事件直接相关的代表性 ETF。 -- 通过 ReMe 记忆检索和本地新闻文件查找相似历史事件,并严格校验来源路径和新闻 ID。 -- 由确定性代码计算复权后的 D1–D10 历史收益,不让 Agent 编造行情数值。 -- 由 Agent 判断事件相似度,再由代码计算权重、预期收益和参考持有时间。 -- 保存可读 Markdown 和结构化 JSON/JSONL,刷新每日索引,并支持可选钉钉投递。 - -工作流由 [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml) 装配,公共 schema 位于 -[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py),各步骤位于 +工作流由 [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml) 装配;公开 schema 位于 +[`reme/schema/auto_fin.py`](../../reme/schema/auto_fin.py),四个 Step 位于 [`reme/steps/cookbook/auto_fin/`](../../reme/steps/cookbook/auto_fin/)。 ## 快速开始 -Auto Fin 要求 Python 3.11 或更高版本、`core` 依赖、Tushare token,以及所配置 Claude Code 兼容 endpoint 的凭据。 - -在仓库根目录运行: +要求 Python 3.11 或更高版本、`core` 依赖、Tushare token 和可用的 AgentScope LLM。 ```bash python -m pip install -e ".[core]" export TUSHARE_TOKEN="your-tushare-token" -export CLAUDE_CODE_API_KEY="your-api-key" +export LLM_API_KEY="your-api-key" reme start config=daily_cookbook job=auto_fin ``` -内置配置默认通过 DashScope 的 Anthropic 兼容 endpoint 使用 `qwen3.7-max`。如需更换兼容模型或服务商: +内置 LLM 组件默认使用 `qwen3.7-plus`。`LLM_BASE_URL` 没有内置默认值;如果服务商要求自定义 OpenAI 兼容 +endpoint,需要显式设置。可通过 `LLM_MODEL_NAME` 和 `LLM_BASE_URL` 覆盖模型与 endpoint。 + +默认 workspace 是进程启动目录下的 `reme_workspace/`。可通过 `DAILY_PAPER_WORKSPACE_DIR` 覆盖;Auto Fin 与 +Daily Paper 共用该设置。 + +日期和时间使用 `Asia/Shanghai`。显式传入的 `date` 必须是当天: ```bash -export CLAUDE_CODE_MODEL_NAME="your-model" -export CLAUDE_CODE_BASE_URL="https://your-anthropic-compatible-endpoint" +reme start config=daily_cookbook job=auto_fin date=2026-08-07 ``` -默认 workspace 是 `reme_workspace/`。该 standalone cookbook 与每日论文工作流共用 workspace 配置: +Auto Fin 首先检查上交所交易日历;休市日会跳过整个工作流。 -```bash -export DAILY_PAPER_WORKSPACE_DIR="/absolute/path/to/reme-workspace" -``` - -如需把最终 Markdown 报告发送到钉钉: - -```bash -export DINGTALK_APP_KEY="your-app-key" -export DINGTALK_APP_SECRET="your-app-secret" -export DINGTALK_ROBOT_CODE="your-robot-code" -export DINGTALK_CONVERSATION_IDS="conversation-id-1,conversation-id-2" -``` - -相关配置为空时会跳过钉钉投递。 - -日期和时间均使用 `Asia/Shanghai`。可选的 `date` 必须是当天: - -```bash -reme start config=daily_cookbook job=auto_fin date=2026-07-25 -``` - -如需刷新全部新闻日期,而不是复用有效历史文件: - -```bash -reme start config=daily_cookbook job=auto_fin force=true -``` - -这可能产生大量 Tushare 请求。普通运行会复用有效历史新闻,并始终刷新当天文件。 - -### 可选 SSH 代理 - -出站代理默认关闭。如需启用,请取消 `daily_cookbook.yaml` 中 -`components.outbound_proxy.default` 的注释,配置免交互 SSH 认证,并设置: - -```bash -export REME_PROXY_IP="your-ssh-proxy-host" -export REME_PROXY_ACCOUNT="your-ssh-account" -``` - -## 工作原理 - -```mermaid -flowchart LR - A[确定运行日和截止时间] --> B[维护财联社新闻文件] - B --> C[确定上一 A 股交易日] - C --> D[构建当前事件窗口] - D --> E[筛选高流动性 ETF 候选] - E --> F[Agent 选择相关 ETF] - F --> G{逐只 ETF} - G --> H[Agent 搜索历史事件] - H --> I[代码回查原始新闻] - I --> J[代码计算复权 D1-D10 收益] - J --> K[Agent 判断相似度] - K --> L[代码计算加权预测] - L --> G - G --> M[Agent 编写合并报告] - M --> N[写入产物并刷新每日索引] - N --> O[可选钉钉投递] -``` - -顶层 Job 包含四个 Auto Fin Step: - -| Step | 职责 | Agent | -|-------------------------|-----------------------------------|-------| -| `auto_fin_data_step` | 维护新闻文件并确定上一交易日 | 否 | -| `auto_fin_topic_step` | 构建输入并选择相关 ETF 和当前事件 | 是 | -| `auto_fin_history_step` | 逐只 ETF 编排历史研究和行情分析 | 是 | -| `auto_fin_merge_step` | 校验结果并生成最终 Markdown 报告 | 是 | - -对于每只已选 ETF,`auto_fin_history_step` 会派发: - -- `auto_fin_history_search_step`:Agent 返回历史新闻引用,然后由代码解析原始记录并计算收益。 -- `auto_fin_market_step`:Agent 只判断相似度,然后由代码计算权重和预测。 - -Agent 负责语义判断,确定性代码负责来源校验和金融数值计算。 - -## 数据和时间边界 - -### 新闻历史 - -`auto_fin_data_step` 使用 Tushare `major_news` 接口读取财联社新闻: - -- 默认回看包含运行日在内的 360 个自然日。 -- 有效历史文件会复用,除非设置 `force=true`。 -- 每次运行都会把当天文件刷新到当前 `decision_at`。 -- 数据量接近接口上限时会递归拆分时间窗口。 -- 写入前会排序和去重,并生成稳定 `news_id`。 - -当前事件窗口为: +## 工作流 ```text -(上一 A 股交易日 15:00, decision_at] +Tushare 交易日历 + │ + ├─ 休市 ──► 跳过 + ▼ +采集财联社新闻 + 固定 ETF 行情历史 + ▼ +更新 ReMe 索引 + ▼ +Agent 筛选 ETF/当日新闻关系 + ▼ +从本地记忆检索可比历史新闻 + ▼ +Agent 选择 same/opposite 事件 + 代码计算 D1/D2/D3/D5 收益 + ▼ +Agent 生成报告 ──► 刷新当日索引 ──► 钉钉(可选) ``` -每次运行都会重建完整窗口;午间和晚间运行不是只读取上次运行后的增量。 +| Step | 职责 | Agent | +|---|---|---| +| `auto_fin_data_step` | 检查交易日、维护新闻并缓存固定 ETF 的完整行情历史 | 否 | +| `auto_fin_topic_step` | 筛选当日新闻与固定 ETF 的直接关系 | 是 | +| `auto_fin_history_step` | 检索可比新闻、校验选择并计算实际收益 | 是 | +| `auto_fin_merge_step` | 汇总证据和上一份报告,生成最终 Markdown | 是 | -### ETF 候选 +三个模型 Step 都使用 Pydantic 结构化输出。Agent 负责语义判断;标识校验、来源解析、行情计算和文件写入由代码负责。 -候选池组合使用: +## 数据与筛选边界 -- `etf_basic`:当前上市 ETF 及其跟踪指数。 -- `fund_daily`:上一 A 股交易日成交额。 +### 新闻 -代码按成交额排序,并按 ETF 名称和指数标识去重,最多向 Topic Agent 提供 150 个候选。Agent 最多返回 20 只 ETF,且 ETF 代码、名称和新闻 -ID 都必须逐字来自候选文件。 +`auto_fin_data_step` 调用 Tushare `major_news`,并固定传入 `src="财联社"`。默认回看 60 个自然日(包含当天)。 +更早日期已有的文件会复用;当天文件始终覆盖为 00:00 至当前决策时刻的新闻。单次请求返回至少 400 条时,时间区间会递归拆分。 -成交额仅用于缩小研究范围,不是交易信号。 +每条新闻写入 `daily/YYYY-MM-DD/auto_fin_news.md`,其稳定 ID 由发布时间和短内容哈希组成。Topic Step 使用当天 +完整文件,不是从上一次运行到本次运行之间的增量。 -## 历史研究和预测 +### 固定 ETF -### 来源回查 +内置配置当前启用: -History Agent 按事件类型、关键实体、传导机制和预期方向搜索。它优先使用 `memory_search`,必要时扫描: +- `518880.SH` +- `159530.SZ` +- `512760.SH` -```text -daily/YYYY-MM-DD/auto_fin_news_data.jsonl -``` +`daily_cookbook.yaml` 中还保留了其他被注释的示例。Data Step 通过 `etf_basic` 解析每个启用代码的名称,然后对 +`fund_daily` 和 `fund_adj` 向前分页,并覆盖写入完整本地 JSONL 行情历史。任一 ETF 无法解析名称都会终止运行。 -Agent 只返回理由、`news_id` 和 workspace 相对 `source_path`。代码会拒绝: +Topic Agent 只接收固定 ETF 的 code/name 和当天本地新闻。每只 ETF 默认最多保留 +`current_news_limit_per_etf=10` 条有效且唯一的新闻引用。未知 ETF、未知 news ID、空理由和重复项会被代码移除; +没有有效事件的 ETF 不进入后续步骤。 -- 把当前事件窗口内的新闻当作历史证据; -- 绝对路径、`..` 路径穿越或 workspace 外路径; -- 文件名不是 `auto_fin_news_data.jsonl` 的来源; -- 不存在的文件或不能唯一解析的 ID; -- 缺少有效发布时间、标题或正文的记录。 +## 历史比较与收益 -历史 Markdown 只能作为检索线索,原始新闻 JSONL 才是事实来源。 +对每个有效的 ETF/当日新闻组合,`auto_fin_history_step` 会在 60 日新闻窗口内调用配置中的 `memory_search`, +结束日期为昨天。`historical_search_limit` 控制每个当前事件最多请求多少条检索结果。只有路径名为 +`auto_fin_news.md` 的命中才会贡献候选 ID;Step 会重新读取源 Markdown 并解析 ID,再调用 History Agent。 -### 复权收益 +History Agent 默认最多选择五条候选,并将关系标记为 `same` 或 `opposite`。代码会移除未知或重复 ID 以及空理由, +随后计算 D1、D2、D3、D5 的复权累计收益: -对每条已回查的历史事件,代码读取 `fund_daily` 和 `fund_adj`,计算最多十个未来收盘点: +- 交易日 15:00 前发生的事件,以当日复权收盘价为入场价,D1 是下一交易日收盘价; +- 15:00 或之后发生的事件,以下一交易日复权开盘价为入场价,D1 是该日收盘价; +- 如果无法从有效正价格和复权因子计算入场价或某个期限,该值为 `null`。 -- 交易日 09:30 前发生:以当日开盘价为 entry。 -- 09:30 至 15:00 前发生:以当日收盘价为 entry。 -- 15:00 或之后、以及非交易日发生:以下一交易日开盘价为 entry。 -- 晚于当前 `decision_at` 的日线收盘数据不会参与计算。 +最终 Agent 接收固定 ETF 列表、所有当前/历史证据、`same`/`opposite` 方向、代码计算的收益,以及此前最近一份 +`auto_fin.md`。它自行判断证据是否支持推荐或应明确观望;代码不会计算评分、期望收益,也不强制给出持有期限。 -```text -adjusted_entry = raw_entry × entry_adjustment_factor -adjusted_close = raw_close × close_adjustment_factor -cumulative_return = adjusted_close / adjusted_entry - 1 -``` - -缺少价格、复权因子、交易日或 horizon 时会记录明确限制,不会用 Agent 生成的数值补齐。 - -### 相似度和预测 - -Market Agent 返回 `[-1, 1]` 范围内的语义相似度: - -- 正值表示机制和方向相似。 -- 负值表示机制可比但方向相反。 -- `0` 表示没有有效关系。 - -代码会截断越界值、忽略零相似度事件,并按相似度绝对值归一化权重。负相似度样本会反转历史收益方向。 每个 D1–D10 horizon 只使用该 -horizon 有数据的样本。参考持有时间取正预期收益中最高的 horizon; 没有正值时留空。 - -结果还会记录样本不足、horizon 缺失、收益方向冲突等限制。它只是有限历史样本比较,不代表统计显著性。 - -## 输出布局 +## 产物 ```text reme_workspace/ ├── daily/ │ ├── YYYY-MM-DD.md │ └── YYYY-MM-DD/ -│ ├── auto_fin_news_data.jsonl -│ ├── auto_fin_analysis.jsonl +│ ├── auto_fin_news.md │ └── auto_fin.md └── resource/ + ├── fin/ + │ ├── etfs.json + │ ├── 518880.SH.jsonl + │ └── <其他固定 ETF>.jsonl └── YYYY-MM-DD/ - ├── filtered_news.jsonl - ├── filtered_etf.jsonl - ├── auto_fin_topic_output.jsonl - ├── auto_fin_history_<序号>__output.json - ├── auto_fin_market_<序号>__output.json - ├── auto_fin_history_output.jsonl + ├── auto_fin_topic_output.json + ├── auto_fin_history_001_output.json + ├── ... + ├── auto_fin_analysis.jsonl └── auto_fin_merge_output.json ``` -主要产物: +每日新闻和报告是用户拥有的 Markdown。`resource/fin/` 是确定性收益计算所用的行情缓存;日期目录下的 JSON/JSONL +保留结构化 Agent 回复和整理后的分析。写入通过同目录临时文件原子替换;报告写完后会刷新当日索引。 -- `auto_fin_news_data.jsonl`:用户拥有的历史新闻回查来源。 -- `filtered_news.jsonl` 和 `filtered_etf.jsonl`:边界明确的 Topic Agent 输入。 -- 各 ETF history 文件:已回查的原始新闻和代码计算的收益路径。 -- 各 ETF market 文件:代码计算的匹配、权重和 D1–D10 预测。 -- `auto_fin_analysis.jsonl`:全部已选 ETF 的最终结构化分析。 -- `auto_fin.md`:可读报告和钉钉投递内容。 -- `daily/YYYY-MM-DD.md`:报告生成后会刷新,确保每日索引能够发现 Auto Fin 报告。 +## 参数与默认值 -新闻和报告都是用户拥有的普通文件;resource 中间产物和搜索索引均可重建。 +公开 Job 参数: -## 配置 +| 参数 | 默认值 | 作用 | +|---|---:|---| +| `date` | `""` | 空值使用 `Asia/Shanghai` 当天;非空值必须是严格 `YYYY-MM-DD` 且等于当天 | +| `historical_search_limit` | `10` | 每个当前事件请求的 `memory_search` 结果上限;最小值为 1 | -### Job 参数 +`daily_cookbook.yaml` 中相关的 Job 级配置: -| 参数 | 默认值 | 说明 | -|---------|--------:|-----------------------------------| -| `date` | 当天 | 严格 `YYYY-MM-DD`;当前只支持当天 | -| `force` | `false` | 是否刷新全部已配置新闻日期 | +| 配置 | 默认值 | 作用 | +|---|---:|---| +| `etf_codes` | 上述三个启用代码 | 固定 ETF 研究范围 | +| `news_lookback_days` | `60` | 本地新闻及历史检索窗口 | +| `current_news_limit_per_etf` | `10` | 每只 ETF 最多保留的当前事件数 | +| `historical_news_limit` | `5` | 每个当前事件最多保留的可比历史事件数 | -### 环境变量 +当前没有公开 `force` 参数。更早的新闻文件会复用;当天新闻和所有固定 ETF 行情文件会刷新;同一天再次成功运行会覆盖 +当天报告和 resource 产物。 -| 变量 | 必需 | 说明 | -|-----------------------------|------|-------------------------------------| -| `TUSHARE_TOKEN` | 是 | 新闻、交易日历、ETF 日线和复权因子 | -| `CLAUDE_CODE_API_KEY` | 是 | Auto Fin Agent 凭据 | -| `CLAUDE_CODE_MODEL_NAME` | 否 | 默认 `qwen3.7-max` | -| `CLAUDE_CODE_BASE_URL` | 否 | Anthropic 兼容 endpoint | -| `AUTO_FIN_AGENT_BACKEND` | 否 | 默认 `claude_code` | -| `AUTO_FIN_PROJECT_PATH` | 否 | Agent project path,默认 `..` | -| `REME_PROXY_IP` | 否 | 仅启用 `ssh_http` 时使用的 SSH 主机 | -| `REME_PROXY_ACCOUNT` | 否 | 仅启用 `ssh_http` 时使用的 SSH 账户 | -| `DAILY_PAPER_WORKSPACE_DIR` | 否 | standalone cookbook workspace | -| `DINGTALK_*` | 否 | 钉钉应用、机器人和会话设置 | +## 环境变量与定时任务 -单元测试可通过 RuntimeContext 注入 `tushare_provider`,不需要真实凭据。 +| 变量 | 必需 | 作用 | +|---|---|---| +| `TUSHARE_TOKEN` | 是 | 交易日历、财联社新闻、ETF 元数据、价格与复权因子 | +| `LLM_API_KEY` | 取决于服务商 | 共享 AgentScope LLM 凭据;配置默认值为空 | +| `LLM_MODEL_NAME` | 否 | 默认 `qwen3.7-plus` | +| `LLM_BASE_URL` | 取决于服务商 | OpenAI 兼容 endpoint;无内置默认值 | +| `TUSHARE_MIRROR_URL` | 否 | 去掉末尾 `/` 后替换 Tushare SDK HTTP URL | +| `DAILY_PAPER_WORKSPACE_DIR` | 否 | standalone cookbook 的共享 workspace | +| `DINGTALK_*` | 否 | 可选的钉钉应用、机器人和群设置 | -### 定时任务 +镜像可按需配置,例如: -`daily_cookbook.yaml` 定义: +```bash +export TUSHARE_MIRROR_URL="http://112.124.63.173:4000/tushare" +``` -| Job | Cron | Asia/Shanghai | -|----------------------|---------------|---------------| -| `auto_fin_0930_cron` | `30 9 * * *` | 每天 09:30 | -| `auto_fin_1145_cron` | `45 11 * * *` | 每天 11:45 | -| `auto_fin_1800_cron` | `0 18 * * *` | 每天 18:00 | +`auto_fin_0930_cron`、`auto_fin_1130_cron` 和 `auto_fin_1800_cron` 按 `Asia/Shanghai` 时区每天 09:30、11:30 和 +18:00 触发。Cron 在周末和节假日仍会启动,但如果 Tushare 返回当天不是上交所交易日,Data Step 会跳过后续工作流; +同一天的后续运行会在已有报告基础上继续完善。 -这些 cron 表达式不会排除周末或休市日。工作流会确定上一 A 股交易日,但当前不会仅因为运行日不是交易日而跳过。 +要发送完成的报告,需要配置 `DINGTALK_APP_KEY`、`DINGTALK_APP_SECRET`、`DINGTALK_ROBOT_CODE` 和逗号分隔的 +`DINGTALK_CONVERSATION_IDS`。没有会话 ID 时发送步骤无副作用。 -## Agent 和安全边界 +## Agent 与失败边界 -Auto Fin wrapper 会加载 `tushare-data` skill、暴露 `memory_search` Job 工具,并默认使用 -`bypassPermissions`。Prompt 会约束各 Agent 的职责,代码则再次校验 schema、ETF 身份、来源路径、 新闻引用和计算值。 +Auto Fin 和 Daily Paper 共用无工具的 `default` AgentScope wrapper,其模型调用不会暴露内置工具或配置型 Job +工具。Auto Fin 由确定性的 Step 代码主动调用 `memory_search`,这不是 Agent 工具调用。独立的交互式 +`dingtalk_wait` Step 才有自己的 `bash` 和 ReMe Job tool allowlist。 -standalone cookbook 默认没有配置 embedding store,因此 `memory_search` 通常使用 BM25 召回; 配置 embedding store 后才能使用向量与 -BM25 融合。 +standalone 配置默认未启用 embedding store,因此 `memory_search` 使用可用的 BM25 路径;只有启用被注释的 +embedding 组件后才有向量/BM25 融合。 -`bypassPermissions` 不是操作系统沙箱。部署前应检查 project path、workspace、凭据和网络边界。 +非法日期、缺少凭据或服务、模型结构化输出无效、固定 ETF 未知、行情文件缺失、记忆检索失败都会终止 Job;休市日是 +成功跳过。工作流没有同日期全局执行锁或跨文件事务;重复成功运行也可能重复发送钉钉通知。 -## 重跑和限制 +## 测试 -- 有效历史新闻会复用,当天新闻始终刷新。 -- 输出使用稳定的每日路径,因此同一天后一次运行会替换前一次报告和 resource 输出。 -- Auto Fin 不使用“报告已存在则跳过”,因为定时运行需要分析更新后的新闻。 -- 配置钉钉后,每次成功运行都会尝试投递;当前没有通知去重。 -- 历史样本缺少部分 horizon 时会降级该样本并记录限制。 -- 非法日期、缺少必要服务、Agent schema 错误、未知 ETF/新闻、危险路径或跨步骤 ETF 身份不一致会使 Job 失败。 - -当前没有实现个股、美股关联、组合账本、BUY/SELL/HOLD、T+1 执行、手续费、滑点、券商连接, 也不会执行真实或模拟委托。 - -## 开发 - -安装开发依赖并运行聚焦测试: +聚焦单元测试会 mock 模型和行情数据边界: ```bash python -m pip install -e ".[dev,core]" -PYTHONPATH=. pytest tests/unit/test_auto_fin.py -v +pytest tests/unit/test_auto_fin.py -v ``` -单元测试会 mock 模型和行情数据边界。需要真实 Tushare、模型或钉钉凭据的测试应单独运行,且需要显式授权。 +需要真实 Tushare、LLM 或钉钉凭据的测试应单独运行,且需要显式授权。 diff --git a/cookbook/daily_paper/README.md b/cookbook/daily_paper/README.md index b3130c0e..dd355009 100644 --- a/cookbook/daily_paper/README.md +++ b/cookbook/daily_paper/README.md @@ -2,334 +2,118 @@ [中文](README_ZH.md) -Daily Paper is a local-first, file-native workflow for turning research rankings into a daily reading package. - -## Capabilities - -- Collect papers from the Hugging Face weekly and monthly rankings while excluding yesterday's papers and recent - recommendations. -- Rank and select candidates, then use Claude Code to produce detailed Chinese notes and a five-minute Chinese brief. -- Keep PDFs, notes, and memories as ordinary user-owned files; indexes and caches remain rebuildable. -- Support daily scheduling, optional DingTalk delivery, conversation memory, auto-dream consolidation, and BM25 recall - for the background DingTalk agent. - -The workflow is assembled by [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml). Its schemas live in -[`reme/schema/daily_paper.py`](../../reme/schema/daily_paper.py), and its steps live in -[`reme/steps/cookbook/daily_paper/`](../../reme/steps/cookbook/daily_paper/). +Daily Paper selects three papers from the Hugging Face Papers weekly and monthly rankings, downloads their arXiv PDFs, +and produces detailed Chinese reading notes plus a roughly five-minute Chinese brief. The implementation lives in +[`reme/steps/cookbook/daily_paper/`](../../reme/steps/cookbook/daily_paper/) and is assembled by +[`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml). ## Quick start -Daily Paper requires Python 3.11 or later, the `core` dependencies, network access to Hugging Face and arXiv, and -credentials for the configured Claude Code endpoint. Auto-memory and auto-dream additionally require the AgentScope LLM -credentials. - -From the repository root: +The workflow requires Python 3.11 or later, the `core` dependencies, an available AgentScope LLM, and network access to +Hugging Face Papers and arXiv. ```bash python -m pip install -e ".[core]" -export CLAUDE_CODE_API_KEY="your-api-key" +export LLM_API_KEY="your-api-key" reme start config=daily_cookbook job=daily_paper ``` -The built-in configuration uses `qwen3.7-max` through DashScope's Anthropic-compatible endpoint. Override -`CLAUDE_CODE_MODEL_NAME` and `CLAUDE_CODE_BASE_URL` when using another compatible model or provider. +The built-in LLM component defaults to: -This is enough to generate paper notes and the daily brief. To use auto-memory and auto-dream, also configure: +- model: `qwen3.7-plus` +- endpoint: no built-in `LLM_BASE_URL`; set the OpenAI-compatible endpoint required by your provider +- environment variables: `LLM_API_KEY`, `LLM_MODEL_NAME`, and `LLM_BASE_URL` -```bash -export LLM_API_KEY="your-api-key" -``` +Auto Fin and Daily Paper share this single `default` LLM and the single tool-free `default` AgentScope wrapper. +Only the interactive `dingtalk_wait` step overrides the wrapper per call with AgentScope `bash` and an explicit ReMe +job allowlist; Daily Paper calls remain tool-free. -By default, outputs are written under `reme_workspace/` in the directory where ReMe starts. +The default workspace is `reme_workspace/` beneath the process working directory. Override it with +`DAILY_PAPER_WORKSPACE_DIR`. -### Optional SSH proxy - -The outbound proxy is disabled by default. To enable it, uncomment `components.outbound_proxy.default` in -`daily_cookbook.yaml`, configure non-interactive SSH authentication, and set: - -```bash -export REME_PROXY_IP="your-ssh-proxy-host" -export REME_PROXY_ACCOUNT="your-ssh-account" -``` - -## What it creates - -A successful run writes ordinary PDFs and Markdown files beneath `workspace_dir`: +## Pipeline ```text -reme_workspace/ -├── daily/ -│ ├── YYYY-MM-DD.md -│ └── YYYY-MM-DD/ -│ ├── daily-paper-brief.md -│ ├── paper-.md -│ └── ... -├── resource/ -│ └── papers/ -│ ├── .pdf -│ └── ... -├── digest/ -│ ├── personal/ -│ ├── project/ -│ ├── resource/ -│ └── wiki/ -├── metadata/ -│ └── ... derived catalogs, indexes, and caches -└── mem_session/ - ├── agentscope/ - └── claude_config/ +Hugging Face weekly/monthly rankings + │ + ▼ +Collect ──► Rank ──► Select 3 ──► Analyze PDFs ──► Digest ──► DingTalk (optional) + │ │ + ├─ PDFs ├─ daily brief + └─ paper notes └─ day index ``` -- `paper-.md` is a detailed Chinese reading note with YAML frontmatter linking back to the source PDF and - paper pages. -- `daily-paper-brief.md` is a roughly five-minute Chinese digest with wikilinks to every selected paper note. -- `daily/YYYY-MM-DD.md` is a derived day index rebuilt from the Markdown files for that date. -- `resource/papers/` holds reusable source PDFs. -- `digest/` contains durable auto-dream output; files there remain ordinary user-owned Markdown. -- `metadata/` and search caches are derived state. `reindex` rebuilds the file store, BM25 index, and graph from source - files. +### 1. Collect -The paper notes are the source of truth for recommendation history: their frontmatter contains the `arxiv_id` values -used for future deduplication. The day index is derived and can be rebuilt. The workflow does not currently write a -separate run manifest. +`daily_paper_collect_step` concurrently fetches: -## How the workflow works +- the Hugging Face weekly ranking for the run date's ISO week; +- the monthly ranking for the run date's calendar month; and +- Hugging Face Daily Papers for exactly the previous calendar day. -```mermaid -flowchart LR - HF[Hugging Face
weekly + monthly] --> C[1. Collect] - Y[Yesterday's papers] --> C - H[Recent local notes] --> C - C --> R[2. Rank] - R --> S[3. Select] - S --> A[4. Analyze PDFs] - A --> D[5. Build brief] - D --> N[6. Notify DingTalk] - A --> P[PDFs + paper notes] - D --> B[Brief + day index] -``` +The weekly and monthly results are merged by arXiv ID while preserving both ranks. The step then excludes papers found +in yesterday's list or in the `arxiv_id` frontmatter of `daily//*.md` within the previous `history_days`. -### 1. Collect and deduplicate +If a Markdown file with `kind: daily-paper-brief` already exists and `force=false`, generation is skipped; the saved +brief can still proceed to DingTalk delivery. The job fails when no eligible papers remain. -The Collect step fetches the weekly ranking for the run date's ISO week, the monthly ranking for its calendar month, and -the Hugging Face Daily Papers IDs for exactly the previous calendar day. It merges weekly and monthly metadata by arXiv -ID and preserves each list's display rank. +### 2. Rank -It then scans `daily//paper-*.md` over the configured history window and excludes IDs found in note -frontmatter. The job fails clearly if no eligible papers remain. - -### 2. Rank candidates - -The Rank step uses reciprocal-rank fusion: +`daily_paper_rank_step` uses reciprocal-rank fusion: ```text score = 1 / (rrf_k + monthly_rank) + weekly_weight / (rrf_k + weekly_rank) ``` -A missing rank contributes zero. Candidates are ordered by fused score, upvotes, and arXiv ID. The bounded candidate -pool also reserves several positions for papers whose titles or summaries match memory-related terms such as agent -memory, memory retrieval, continual learning, context compression, knowledge graphs, and RAG. This reserve is a simple -keyword heuristic, not a semantic classifier. +A missing rank contributes zero. Papers are ordered by fused score, upvotes, and arXiv ID. The pool is capped at +`candidate_limit`, and Rank applies no topic preference. -### 3. Select papers +### 3. Select -Claude Code receives the bounded candidate pool and returns a structured `PaperSelection`. The implementation requires -exactly `top_k` unique in-pool IDs with consecutive ranks. Invalid output is returned to the agent once as validation -feedback; a second invalid response fails the job. +`daily_paper_select_step` sends candidate metadata to a tool-free AgentScope agent and requires exactly three items: -### 4. Download and analyze PDFs - -Selected papers are processed sequentially. For each paper, the workflow: - -1. validates the modern arXiv ID format; -2. downloads and validates the PDF, or reuses an existing file with a valid `%PDF-` header; -3. extracts text with `pypdf`, adding page markers and applying page and character limits; -4. asks Claude Code for a structured detailed reading; and -5. writes normalized frontmatter plus the generated Markdown body. - -The current extractor requires a usable PDF text layer. Scanned or image-only PDFs fail because there is no OCR -fallback. If extraction exceeds a configured limit, the note records that the input was truncated. - -### 5. Build the brief and index - -Claude Code reads every detailed note and produces the daily brief. The code verifies that each source-note wikilink is -present and appends any missing links before writing the file. It then rebuilds `daily/YYYY-MM-DD.md` from that day's -Markdown frontmatter. - -### 6. Optionally notify DingTalk - -The final step sends the brief body, without YAML frontmatter, to each configured DingTalk group in order. With no -conversation IDs it is a no-op. If one group fails, the step still attempts the remaining groups and reports the -combined failure afterward. - -## Memory and search - -The standalone configuration separates agent wrappers by responsibility: - -- `daily_paper` selects papers, analyzes them, and builds the brief. It keeps Claude Code's normal local tools and - disables `WebSearch`, but currently has no memory-retrieval job configured. -- `dingtalk_wait` runs the background DingTalk agent and exposes `memory_search` as a callable tool. -- `memory` runs auto-memory and the LLM-backed auto-dream steps through AgentScope. Its built-in shell and file tools - are disabled; memory changes go through the narrower ReMe jobs such as `daily_write`, `read`, `edit`, and `write`. - -The built-in `memory_search` job uses BM25 over Markdown under `daily/` and `digest/`. ReMe's search step can fuse -vector results, but this cookbook does not configure an embedding store by default, so vector retrieval is not run. -`node_search` is a narrower digest recall tool used internally by auto-dream. - -`index_update_loop` indexes existing memory files when the service starts and watches those directories for later -changes. Run `reindex` when recovering the derived file store or forcing a complete index rebuild. Source Markdown and -PDFs are not deleted by `reindex`. - -`auto_memory` writes or updates one daily note from caller-supplied conversation messages and a stable `session_id`. -`auto_dream` scans recent daily notes, integrates durable units under `digest/`, and writes interest topics. Both are -on-demand jobs in this cookbook; no auto-dream cron is configured. The DingTalk agent can recall through -`memory_search`, but it does not automatically call `auto_memory` after a conversation. - -## Dates, reruns, and idempotency - -- `date` must be an exact `YYYY-MM-DD` value. When omitted, the job uses today in the application timezone, which is - `Asia/Shanghai` in the built-in configuration. -- “Yesterday” means `date - 1 day`, not the previous 24 hours. -- `history_days` considers prior dated note directories only; the current run date is never part of its history scan. -- If `daily//daily-paper-brief.md` already exists and `force=false`, collection, ranking, model calls, PDF work, - and digest generation are skipped. The existing brief remains available to the DingTalk notification step. -- `force=true` regenerates the notes and brief. Existing valid PDFs are still reused. - -Each PDF, detailed note, and final brief uses a temporary file followed by replacement so callers do not see a partially -written file. The complete multi-file workflow is not transactional, and there is no global lock for two concurrent runs -of the same date. - -## Running the cookbook - -The main jobs in the standalone configuration are: - -| Job | Behavior | -|---------------------|-----------------------------------------------------------------| -| `daily_paper` | On-demand generation through the CLI or HTTP service | -| `daily_paper_cron` | The same pipeline every day at 08:00 in `Asia/Shanghai` | -| `dingtalk_wait` | A supervised background DingTalk agent with `memory_search` | -| `auto_memory` | Write or update a daily note from conversation messages | -| `auto_dream` | Consolidate recent daily notes into digest memory and interests | -| `memory_search` | BM25 recall over daily and digest Markdown | -| `reindex` | Rebuild derived search state from existing memory files | -| `index_update_loop` | Initialize and continuously update search state in service mode | - -Supporting jobs such as `node_search`, `daily_list`, `daily_write`, `read`, `write`, `edit`, and frontmatter updates -provide the constrained tools used by the memory agent. - -### One-time runs - -The quick-start command generates today's brief. To generate a specific date with selected overrides: - -```bash -reme start \ - config=daily_cookbook \ - job=daily_paper \ - date=2026-07-21 \ - top_k=3 \ - history_days=30 +```json +{"papers": [{"arxiv_id": "2601.01234", "reasoning": "A specific, verifiable reason"}]} ``` -Regenerate a date whose brief already exists: +All IDs must be unique and belong to the candidate pool, and every reason must be non-empty. A validation failure is +returned to the agent for one retry. Only a non-empty `topics` value injects a personalized subject preference into +the selection prompt; it does not change the fixed count of three papers. -```bash -reme start config=daily_cookbook job=daily_paper date=2026-07-21 force=true -``` +### 4. Analyze -Add `service.show_metadata=true` to a one-time command when the response metadata is useful for diagnostics. +`daily_paper_analyze_step` processes the selected papers in order: -### Long-running service and cron +1. validates a modern `YYYY.NNNN` or `YYYY.NNNNN` arXiv ID; +2. downloads the PDF to `resource/papers/.pdf`; +3. reuses an existing target whose header is `%PDF-`; +4. extracts paginated text with `pypdf`, bounded by `max_pdf_pages` and `max_pdf_chars`; +5. sends metadata, selection reasoning, and PDF text to a tool-free agent; and +6. writes a Chinese note to `daily//.md`. -Start the standalone HTTP service and its scheduled/background jobs: +Downloads use a temporary file and atomically replace the target only after validating the PDF header. They are also +bounded by `max_pdf_bytes`. There is no OCR fallback, so scanned or textless PDFs fail. When extraction is truncated, +the note records `pdf_text_truncated: true` in its frontmatter. -```bash -reme start config=daily_cookbook -``` +### 5. Digest -It listens on `127.0.0.1:8001` by default, so it can run beside the default ReMe service. Call the on-demand job from -another terminal with either the ReMe client or HTTP: +`daily_paper_digest_step` builds the Chinese brief directly from the three in-memory analyses; it does not reread or +search other material. The agent returns `title`, `desc`, and `body`. The code then: -```bash -reme daily_paper host=127.0.0.1 port=8001 -``` +- strips model-generated YAML frontmatter if present; +- normalizes the Chinese title for use as a filename; +- deterministically appends wikilinks to all three source notes; +- writes `daily//.md`; and +- rebuilds the `daily/.md` day index. -```bash -curl -s http://127.0.0.1:8001/daily_paper \ - -H 'Content-Type: application/json' \ - -d '{"date":"2026-07-21","top_k":3,"force":false}' -``` +Final response metadata includes the date, week/month scopes, selected arXiv IDs, selection reasons, note/PDF/brief +paths, source counts, and exclusion counts. -Recall memory, record a conversation, consolidate it, or explicitly rebuild the search index: +### 6. DingTalk -```bash -reme memory_search host=127.0.0.1 port=8001 query="agent memory" limit=5 - -reme auto_memory host=127.0.0.1 port=8001 \ - session_id=example-session \ - messages='[{"name":"user","role":"user","content":"I prefer concise paper summaries."}]' - -reme auto_dream host=127.0.0.1 port=8001 date=2026-07-21 -reme reindex host=127.0.0.1 port=8001 -``` - -Service and schedule settings can be overridden at startup: - -```bash -reme start \ - config=daily_cookbook \ - service.host=0.0.0.0 \ - service.port=8101 \ - jobs.daily_paper_cron.cron="30 7 * * *" -``` - -## Configuration - -The most useful job settings are: - -| Setting | Default | Purpose | -|-------------------|-------------:|--------------------------------------------------------------| -| `candidate_limit` | `20` | Maximum number of papers sent to selection | -| `memory_reserve` | `5` | Candidate positions reserved by the memory-keyword heuristic | -| `top_k` | `3` | Number of papers selected and analyzed | -| `rrf_k` | `60` | Reciprocal-rank fusion constant | -| `weekly_weight` | `0.7` | Weight of the weekly ranking in fusion | -| `history_days` | `30` | Prior recommendation window excluded by arXiv ID | -| `hf_timeout` | `30` seconds | Hugging Face request timeout | -| `hf_max_retries` | `3` | Maximum Hugging Face request attempts | -| `pdf_timeout` | `90` seconds | arXiv download timeout | -| `max_pdf_bytes` | `52428800` | Maximum PDF size (50 MiB) | -| `max_pdf_pages` | `80` | Maximum pages extracted for analysis | -| `max_pdf_chars` | `240000` | Maximum extracted characters sent for one paper | - -The public job parameters are `date`, `force`, `top_k`, `weekly_weight`, and `history_days`. Explicit invocation values -take precedence over the job defaults. - -The standalone application also accepts these environment variables: - -| Variable | Purpose | -|-----------------------------------------|------------------------------------------------| -| `DAILY_PAPER_WORKSPACE_DIR` | Overrides the default `reme_workspace` | -| `DAILY_PAPER_PROJECT_PATH` | Repository/project path visible to Claude Code | -| `REME_PROXY_IP` | Optional SSH proxy host | -| `REME_PROXY_ACCOUNT` | Optional SSH proxy account | -| `DAILY_PAPER_HOST` / `DAILY_PAPER_PORT` | HTTP bind address | -| `CLAUDE_CODE_API_KEY` | API key for the Claude Code endpoint | -| `CLAUDE_CODE_MODEL_NAME` | Claude Code model; default `qwen3.7-max` | -| `CLAUDE_CODE_BASE_URL` | Claude Code Anthropic-compatible endpoint | -| `LLM_API_KEY` | API key for the AgentScope memory model | -| `LLM_MODEL_NAME` | Memory model; default `qwen3.7-max` | -| `LLM_BASE_URL` | Memory model's Anthropic-compatible endpoint | - -`DAILY_PAPER_PROJECT_PATH` defaults to `..` relative to the workspace. With the default `reme_workspace`, starting from -the repository root resolves it back to the repository. If the workspace lives elsewhere, set both paths explicitly. - -ReMe loads an uncommitted `.env` file found from the current directory upward, so the same values may be placed there -instead of exported in the shell. - -## DingTalk configuration - -DingTalk is optional. Configure it only when brief delivery or the background DingTalk agent is needed: +The final `dingtalk_markdown_send_step` is optional. With no conversation IDs it is a no-op. When configured, it strips +frontmatter and sends the brief body to each group in order: ```dotenv DINGTALK_APP_KEY=your-app-key @@ -338,39 +122,119 @@ DINGTALK_ROBOT_CODE=your-robot-code DINGTALK_CONVERSATION_IDS=cid-group-one,cid-group-two ``` -`DINGTALK_CONVERSATION_IDS` is required only for proactive brief delivery. The background `dingtalk_wait` job uses the -first three credentials but not the conversation list. +A failed recipient does not prevent later attempts; the step reports a combined failure after trying every group. -## Failure recovery and boundaries +## Outputs -| Situation | Behavior | -|--------------------------------------|----------------------------------------------------------------| -| Temporary Hugging Face failure | Retries with exponential delay up to `hf_max_retries` attempts | -| No eligible papers | Fails before ranking | -| Invalid `top_k` or selection output | Fails after validation; selection output gets one retry | -| Oversized, invalid, or textless PDF | Stops during analysis | -| PDF exceeds page or character limits | Continues with truncated text and records the truncation | -| One paper analysis fails | Stops the job; earlier PDFs and notes remain on disk | -| Brief misses a source-note link | Appends the missing wikilink before writing | -| Auto-dream partial integration | Successful units remain; failed paths are not checkpointed | +```text +reme_workspace/ +├── daily/ +│ ├── YYYY-MM-DD.md +│ └── YYYY-MM-DD/ +│ ├── .md # three, kind: daily-paper-analysis +│ └── .md # one, kind: daily-paper-brief +└── resource/ + └── papers/ + └── .pdf +``` -To recover, inspect the date's notes and PDFs, fix the network, credential, model, or PDF issue, then rerun the same -date with `force=true`. Valid cached PDFs will be reused. +Filenames come from the agent's Chinese titles. The implementation removes unsafe path characters and resolves title +collisions. Markdown and PDF outputs are written through same-directory temporary files and atomic replacement. -The built-in Claude Code components run with `permission_mode: bypassPermissions` and disable `WebSearch`. -`dingtalk_wait` can call the local `memory_search` job; `daily_paper` currently has no job tools. The analysis and brief -prompts constrain what the agent should read, but these steps do not set a strict per-call tool allowlist or an -operating-system sandbox. The AgentScope memory wrapper disables its built-in shell and filesystem tools, but runs its -ReMe job tools in bypass permission mode. Run the cookbook only with a trusted project and workspace, and tighten the -agent configuration before shared or production use. +## Parameters and defaults + +Public job parameters: + +| Parameter | Default | Purpose | +|---|---:|---| +| `date` | `""` | Run date; empty uses today in the app timezone, otherwise requires `YYYY-MM-DD` | +| `force` | `false` | Regenerate even when the day's brief exists | +| `topics` | `""` | Topics to prioritize during selection | +| `weekly_weight` | `0.7` | Weekly contribution to RRF | +| `history_days` | `30` | Prior recommendation exclusion window | + +Step-level settings on the `daily_paper` job: + +| Setting | Default | Purpose | +|---|---:|---| +| `candidate_limit` | `20` | Maximum candidates sent to Select | +| `rrf_k` | `60` | RRF constant | +| `hf_timeout` | `600` seconds | Timeout for one Hugging Face request | +| `hf_max_retries` | `3` | Maximum Hugging Face attempts | +| `pdf_timeout` | `600` seconds | arXiv PDF download timeout | +| `max_pdf_bytes` | `52428800` | PDF limit, 50 MiB | +| `max_pdf_pages` | `20` | Maximum extracted pages | +| `max_pdf_chars` | `300000` | Maximum extracted PDF characters sent to the agent | + +## Mirrors + +The data clients use httpx's default environment handling, so `HTTP_PROXY`, `HTTPS_PROXY`, and `NO_PROXY` take effect +when present. Mirror environment variables independently replace each service's base URL: + +```dotenv +# Defaults to https://huggingface.co when unset +HF_MIRROR_URL=https://hf-mirror.com + +# Defaults to https://arxiv.org when unset +ARXIV_MIRROR_URL=https://export.arxiv.org + +# Path-prefixed relay URLs are also supported +# HF_MIRROR_URL=http://relay-host:18080/hf +# ARXIV_MIRROR_URL=http://relay-host:18080/arxiv +``` + +`HF_MIRROR_URL` must implement the `/papers/...`, `/api/daily_papers`, and `/api/papers/...` routes used by the current +client. `ARXIV_MIRROR_URL` must implement `/pdf/`. A path prefix in either base URL is preserved, and a +trailing slash is optional. When a variable is unset, the official service is used directly; there is no fallback chain. + +## Running the workflow + +Generate a brief for a specific date: + +```bash +reme start \ + config=daily_cookbook \ + job=daily_paper \ + date=2026-08-06 \ + topics="Agent memory" \ + history_days=30 +``` + +Force a rerun; valid local PDFs are still reused: + +```bash +reme start config=daily_cookbook job=daily_paper date=2026-08-06 force=true +``` + +Start the HTTP service and scheduled jobs: + +```bash +reme start config=daily_cookbook +``` + +The built-in service listens on `127.0.0.1:8001`. `daily_paper_cron` runs every day at 08:00 in the +`Asia/Shanghai` timezone. Override the bind address with `DAILY_PAPER_HOST`, `DAILY_PAPER_PORT`, or startup arguments. + +```bash +curl -s http://127.0.0.1:8001/daily_paper \ + -H 'Content-Type: application/json' \ + -d '{"date":"2026-08-06","force":false,"topics":"Agent memory"}' +``` + +## Failures and reruns + +- Hugging Face failures use exponential backoff up to `hf_max_retries` attempts. +- Fewer than three candidates, invalid agent selection, invalid/oversized/textless PDFs, or empty agent output stop the + job. +- Papers are analyzed sequentially; PDFs and notes completed before a failure remain on disk. +- `force=true` regenerates notes and the brief while reusing valid PDFs. +- The multi-file workflow is not transactional and has no global per-date execution lock. ## Tests -The focused unit suite mocks Hugging Face, arXiv, Claude Code, and DingTalk boundaries: +The focused unit tests mock Hugging Face, arXiv, AgentScope, and DingTalk boundaries and do not call real services: ```bash python -m pip install -e ".[dev,core]" pytest tests/unit/test_daily_paper.py -v ``` - -Real runs access external services and may incur model costs; they should not be used as ordinary unit tests. diff --git a/cookbook/daily_paper/README_ZH.md b/cookbook/daily_paper/README_ZH.md index 2d605eb7..cfa3a5ef 100644 --- a/cookbook/daily_paper/README_ZH.md +++ b/cookbook/daily_paper/README_ZH.md @@ -2,312 +2,108 @@ [English](README.md) -每日论文是一个本地优先、文件原生的研究资讯工作流,用于把研究榜单转化为每日阅读材料。 - -## 能力 - -- 从 Hugging Face 周榜和月榜采集论文,并排除昨日论文和近期已经推荐过的论文。 -- 对候选论文进行排序和精选,再使用 Claude Code 生成中文详细论文笔记和约五分钟可读完的中文简报。 -- 将 PDF、笔记和记忆保存为由用户拥有的普通文件;索引和缓存均可重建。 -- 支持每日定时运行、可选钉钉投递、对话记忆、auto-dream 整理,以及供后台钉钉 Agent 使用的 BM25 检索。 - -工作流由 [`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml) 装配,公共 schema 位于 -[`reme/schema/daily_paper.py`](../../reme/schema/daily_paper.py),各步骤位于 -[`reme/steps/cookbook/daily_paper/`](../../reme/steps/cookbook/daily_paper/)。 +每日论文工作流从 Hugging Face Papers 的周榜和月榜中筛选三篇论文,下载 arXiv PDF,生成中文论文解读和一篇约五分钟可读完的中文简报。当前实现位于 +[`reme/steps/cookbook/daily_paper/`](../../reme/steps/cookbook/daily_paper/),由 +[`daily_cookbook.yaml`](../../reme/config/daily_cookbook.yaml) 装配。 ## 快速开始 -每日论文要求 Python 3.11 或更高版本、`core` 依赖、可访问 Hugging Face 和 arXiv 的网络,以及所配置 Claude Code endpoint -的凭据。auto-memory 和 auto-dream 还需要 AgentScope LLM 凭据。 - -在仓库根目录运行: +要求 Python 3.11 或更高版本、`core` 依赖、可用的 AgentScope LLM,以及能访问 Hugging Face Papers 和 arXiv 的网络。 ```bash python -m pip install -e ".[core]" -export CLAUDE_CODE_API_KEY="your-api-key" +export LLM_API_KEY="your-api-key" reme start config=daily_cookbook job=daily_paper ``` -内置配置默认通过 DashScope 的 Anthropic 兼容 endpoint 使用 `qwen3.7-max`。如需使用其他兼容模型或服务商, 请覆盖 -`CLAUDE_CODE_MODEL_NAME` 和 `CLAUDE_CODE_BASE_URL`。 +内置 LLM 组件默认配置为: -以上配置足以生成论文笔记和每日简报。要使用 auto-memory 和 auto-dream,还需配置: +- 模型:`qwen3.7-plus` +- endpoint:无内置 `LLM_BASE_URL`;请设置服务商要求的 OpenAI 兼容 endpoint +- 环境变量:`LLM_API_KEY`、`LLM_MODEL_NAME`、`LLM_BASE_URL` -```bash -export LLM_API_KEY="your-api-key" -``` +Auto Fin 和 Daily Paper 共用这一个 `default` LLM,以及唯一的无工具 `default` AgentScope wrapper。 +只有交互式 `dingtalk_wait` Step 会在调用时覆盖默认值,启用 AgentScope `bash` 和明确的 ReMe Job allowlist;Daily +Paper 的模型调用仍然无工具。 -默认情况下,产物写入 ReMe 启动目录下的 `reme_workspace/`。 +默认 workspace 是启动目录下的 `reme_workspace/`,可通过 `DAILY_PAPER_WORKSPACE_DIR` 覆盖。 -### 可选 SSH 代理 - -出站代理默认关闭。如需启用,请取消 `daily_cookbook.yaml` 中 -`components.outbound_proxy.default` 的注释,配置免交互 SSH 认证,并设置: - -```bash -export REME_PROXY_IP="your-ssh-proxy-host" -export REME_PROXY_ACCOUNT="your-ssh-account" -``` - -## 文件产物 - -一次成功运行会在 `workspace_dir` 下写入普通 PDF 和 Markdown 文件: +## 工作流 ```text -reme_workspace/ -├── daily/ -│ ├── YYYY-MM-DD.md -│ └── YYYY-MM-DD/ -│ ├── daily-paper-brief.md -│ ├── paper-.md -│ └── ... -├── resource/ -│ └── papers/ -│ ├── .pdf -│ └── ... -├── digest/ -│ ├── personal/ -│ ├── project/ -│ ├── resource/ -│ └── wiki/ -├── metadata/ -│ └── ... 派生 catalog、索引和缓存 -└── mem_session/ - ├── agentscope/ - └── claude_config/ +Hugging Face 周榜/月榜 + │ + ▼ +Collect ──► Rank ──► Select 3 篇 ──► Analyze PDF ──► Digest ──► DingTalk(可选) + │ │ + ├─ PDF ├─ 每日简报 + └─ 论文解读 └─ 当日索引 ``` -- `paper-.md` 是中文详细论文解读,其 YAML frontmatter 会链接原始 PDF 和论文页面。 -- `daily-paper-brief.md` 是约五分钟可读完的中文简报,并包含每篇入选论文笔记的 wikilink。 -- `daily/YYYY-MM-DD.md` 是从当日 Markdown 文件重建的派生日索引。 -- `resource/papers/` 保存可复用的原始 PDF。 -- `digest/` 保存持久的 auto-dream 产物,其中仍然是由用户拥有的普通 Markdown 文件。 -- `metadata/` 和搜索缓存属于派生状态;`reindex` 会根据源文件重建 file store、BM25 索引和图。 +### 1. Collect -论文笔记是推荐历史的事实来源:后续排重会读取其 frontmatter 中的 `arxiv_id`。日索引属于可重建的派生文件。 当前工作流不会另外写入运行 -manifest。 +`daily_paper_collect_step` 根据运行日期并发读取: -## 工作流程 +- 该日期所在 ISO week 的 Hugging Face 周榜; +- 该日期所在自然月的 Hugging Face 月榜; +- 严格前一个自然日的 Hugging Face Daily Papers。 -```mermaid -flowchart LR - HF[Hugging Face
周榜 + 月榜] --> C[1. Collect] - Y[昨日论文] --> C - H[近期本地笔记] --> C - C --> R[2. Rank] - R --> S[3. Select] - S --> A[4. Analyze PDFs] - A --> D[5. Build brief] - D --> N[6. Notify DingTalk] - A --> P[PDF + 论文笔记] - D --> B[简报 + 日索引] -``` +周榜和月榜按 arXiv ID 合并,并保留各自排名。随后排除: -### 1. 采集与排重 +- 昨日 Daily Papers 中的论文; +- `history_days` 窗口内,已出现在 `daily//*.md` frontmatter `arxiv_id` 中的论文。 -Collect 会获取运行日所在 ISO week 的周榜、所在自然月的月榜,以及严格前一个自然日的 Hugging Face Daily Papers ID。周榜和月榜元数据按 -arXiv ID 合并,同时保留两个榜单各自的展示排名。 +如果当天已经存在 `kind: daily-paper-brief` 的 Markdown 且 `force=false`,整个生成流程会跳过;已有简报仍可进入钉钉发送步骤。没有剩余候选论文时,Job 直接失败。 -随后,它会在配置的历史窗口内扫描 `daily//paper-*.md`,排除笔记 frontmatter 中已有的 ID。如果排重后 没有任何可选论文,Job -会明确失败。 +### 2. Rank -### 2. 候选排序 - -Rank 使用 reciprocal-rank fusion(RRF): +`daily_paper_rank_step` 使用 reciprocal-rank fusion: ```text score = 1 / (rrf_k + monthly_rank) + weekly_weight / (rrf_k + weekly_rank) ``` -论文缺少某个榜单排名时,该项贡献为零。候选按融合分、upvotes 和 arXiv ID 排序。有界候选池还会为标题或摘要命中 Agent -memory、memory retrieval、continual learning、context compression、knowledge graph、RAG 等记忆相关关键词的 -论文保留若干位置。这个保留策略只是关键词启发式,不是语义分类器。 +缺失的榜单排名贡献为零。论文按融合分、upvotes、arXiv ID 排序,候选池最多保留 `candidate_limit` 篇。Rank 阶段不应用任何主题倾向。 -### 3. 精选论文 +### 3. Select -Claude Code 接收有界候选池,并返回结构化的 `PaperSelection`。实现要求恰好选择 `top_k` 个候选池内的唯一 ID, 且 rank -必须连续。输出不合法时,校验错误会反馈给 Agent 并重试一次;第二次仍不合法则 Job 失败。 +`daily_paper_select_step` 将候选元数据交给无工具的 AgentScope Agent,并要求返回恰好三项: -### 4. 下载并解读 PDF - -入选论文按顺序逐篇处理。每篇论文都会经过: - -1. 校验当前支持的新版 arXiv ID 格式; -2. 下载并校验 PDF,或复用文件头为 `%PDF-` 的已有文件; -3. 使用 `pypdf` 提取文本、插入页码标记,并应用页数和字符数限制; -4. 请求 Claude Code 返回结构化的详细解读; -5. 写入规范化 frontmatter 和生成的 Markdown 正文。 - -当前提取器依赖可用的 PDF 文本层。扫描版或纯图片 PDF 会失败,因为没有 OCR fallback。提取内容超过配置限制时, 笔记会记录输入已被截断。 - -### 5. 生成简报与索引 - -Claude Code 会读取全部详细笔记并生成当日简报。代码会检查每篇源笔记的 wikilink;如有遗漏,会在写入前自动补齐。 随后,工作流根据当日 -Markdown frontmatter 重建 `daily/YYYY-MM-DD.md`。 - -### 6. 可选的钉钉通知 - -最后一步会去掉 YAML frontmatter,把简报正文按顺序发送到每个已配置的钉钉群。未配置群会话 ID 时,该步骤无副作用 -跳过。某个群发送失败不会阻止继续尝试其他群,所有发送完成后再汇总报告失败。 - -## 记忆与检索 - -独立配置按职责分离 agent wrapper: - -- `daily_paper` 执行论文精选、解读和简报生成。它保留 Claude Code 的常规本地工具并禁用 `WebSearch`,但当前没有 配置记忆检索 - Job。 -- `dingtalk_wait` 运行后台钉钉 Agent,并把 `memory_search` 作为可调用工具。 -- `memory` 通过 AgentScope 执行 auto-memory 和 auto-dream 中依赖 LLM 的步骤。它禁用内置 shell 和文件工具; 记忆变更只能经过 - `daily_write`、`read`、`edit`、`write` 等更窄的 ReMe job。 - -内置配置的 `memory_search` 使用 BM25 检索 `daily/` 和 `digest/` 下的 Markdown。ReMe 的搜索步骤支持融合向量结果, 但本 -cookbook 默认没有配置 embedding store,因此不会执行向量检索。`node_search` 是 auto-dream 内部使用的 digest 节点检索工具。 - -`index_update_loop` 会在服务启动时索引已有记忆文件,并持续监听这些目录的后续变化。修复派生 file store 或需要强制 -完整重建索引时,可运行 `reindex`。`reindex` 不会删除源 Markdown 或 PDF。 - -`auto_memory` 根据调用方传入的对话消息和稳定 `session_id` 写入或更新一篇 daily note。`auto_dream` 扫描近期 daily -notes,把持久记忆单元整合到 `digest/`,并生成兴趣主题。这两个 job 在本 cookbook 中都是按需执行;当前没有配置 auto-dream -cron。钉钉 Agent 可以通过 `memory_search` 召回记忆,但对话结束后不会自动调用 `auto_memory`。 - -## 日期、重跑与幂等 - -- `date` 必须严格符合 `YYYY-MM-DD`。省略时使用应用配置时区中的今天;内置配置为 `Asia/Shanghai`。 -- “昨日”表示 `date - 1 day`,不是模糊的最近 24 小时。 -- `history_days` 只扫描此前的日期目录,不会把本次运行日纳入历史窗口。 -- 如果 `daily//daily-paper-brief.md` 已存在且 `force=false`,采集、排序、模型调用、PDF 处理和简报生成都会 - 跳过;已有简报仍会交给钉钉通知步骤。 -- `force=true` 会重新生成笔记和简报,但仍会复用已有且有效的 PDF。 - -每个 PDF、详细笔记和最终简报都会先写临时文件再替换,避免读取方看到半写状态。整个多文件工作流不是事务, 同一日期的并发运行也没有全局锁。 - -## 运行方式 - -独立配置中的主要 Job 如下: - -| Job | 行为 | -|---------------------|---------------------------------------------------------------| -| `daily_paper` | 通过 CLI 或 HTTP 服务按需生成 | -| `daily_paper_cron` | 每天 08:00(`Asia/Shanghai`)执行相同 pipeline | -| `dingtalk_wait` | 由 supervisor 管理、具有 `memory_search` 能力的后台钉钉 Agent | -| `auto_memory` | 根据对话消息写入或更新 daily note | -| `auto_dream` | 把近期 daily notes 整理为 digest 记忆和兴趣主题 | -| `memory_search` | 对 daily 和 digest Markdown 执行 BM25 检索 | -| `reindex` | 根据已有记忆文件重建派生检索状态 | -| `index_update_loop` | 在服务模式下初始化并持续更新检索状态 | - -`node_search`、`daily_list`、`daily_write`、`read`、`write`、`edit` 和 frontmatter 更新等辅助 job 构成 memory agent 使用的受约束工具。 - -### 一次性运行 - -快速开始中的命令会生成今天的简报。要生成指定日期并覆盖部分参数: - -```bash -reme start \ - config=daily_cookbook \ - job=daily_paper \ - date=2026-07-21 \ - top_k=3 \ - history_days=30 +```json +{"papers": [{"arxiv_id": "2601.01234", "reasoning": "具体且可核验的选择理由"}]} ``` -重新生成已有简报的日期: +三个 ID 必须唯一且都属于候选池,理由不能为空。校验失败后,错误信息会反馈给 Agent 并重试一次。只有非空 `topics` 会向精选提示注入个性化主题,且不会改变固定的三篇数量。 -```bash -reme start config=daily_cookbook job=daily_paper date=2026-07-21 force=true -``` +### 4. Analyze -需要查看响应 metadata 进行诊断时,可在一次性命令中加入 `service.show_metadata=true`。 +`daily_paper_analyze_step` 按精选顺序逐篇处理: -### 常驻服务与 cron +1. 校验新版 arXiv ID 格式 `YYYY.NNNN` 或 `YYYY.NNNNN`; +2. 下载 PDF 到 `resource/papers/.pdf`; +3. 如果目标文件已存在且以 `%PDF-` 开头,直接复用; +4. 用 `pypdf` 提取分页文本,受 `max_pdf_pages` 和 `max_pdf_chars` 限制; +5. 将论文元数据、选择理由和 PDF 文本交给无工具 Agent; +6. 将中文解读写入 `daily//<中文标题>.md`。 -启动独立 HTTP 服务以及定时、后台 Job: +下载采用临时文件并在校验 PDF 文件头后原子替换,同时限制 `max_pdf_bytes`。当前没有 OCR;扫描版或无文本层 PDF 会失败。提取被截断时,笔记 frontmatter 中的 `pdf_text_truncated` 会记录为 `true`。 -```bash -reme start config=daily_cookbook -``` +### 5. Digest -服务默认监听 `127.0.0.1:8001`,因此可以和默认 ReMe 服务并行运行。在另一个终端中通过 ReMe client 或 HTTP 调用按需任务: +`daily_paper_digest_step` 直接使用内存中的三篇解读生成中文简报,不会重新读取或搜索其他资料。输出必须包含 `title`、`desc` 和 `body`。代码会: -```bash -reme daily_paper host=127.0.0.1 port=8001 -``` +- 去掉模型可能生成的 YAML frontmatter; +- 规范化中文标题并用作文件名; +- 确定性追加三篇源笔记的 wikilink; +- 写入 `daily//<中文简报标题>.md`; +- 重建 `daily/.md` 当日索引。 -```bash -curl -s http://127.0.0.1:8001/daily_paper \ - -H 'Content-Type: application/json' \ - -d '{"date":"2026-07-21","top_k":3,"force":false}' -``` +最终响应 metadata 包含日期、周/月范围、入选 arXiv ID、选择理由、笔记/PDF/简报路径、源榜单数量和排重数量。 -检索记忆、记录对话、整理记忆,或显式重建检索索引: +### 6. DingTalk -```bash -reme memory_search host=127.0.0.1 port=8001 query="agent memory" limit=5 - -reme auto_memory host=127.0.0.1 port=8001 \ - session_id=example-session \ - messages='[{"name":"user","role":"user","content":"I prefer concise paper summaries."}]' - -reme auto_dream host=127.0.0.1 port=8001 date=2026-07-21 -reme reindex host=127.0.0.1 port=8001 -``` - -监听地址和调度时间可以在启动时覆盖: - -```bash -reme start \ - config=daily_cookbook \ - service.host=0.0.0.0 \ - service.port=8101 \ - jobs.daily_paper_cron.cron="30 7 * * *" -``` - -## 配置 - -最常用的 Job 配置如下: - -| 配置项 | 默认值 | 用途 | -|-------------------|-----------:|------------------------------------| -| `candidate_limit` | `20` | 送入精选阶段的最大论文数 | -| `memory_reserve` | `5` | 记忆关键词启发式保留的候选位置数 | -| `top_k` | `3` | 最终精选和解读的论文数 | -| `rrf_k` | `60` | RRF 常数 | -| `weekly_weight` | `0.7` | 周榜在融合排序中的权重 | -| `history_days` | `30` | 按 arXiv ID 排除近期推荐的时间窗口 | -| `hf_timeout` | `30` 秒 | Hugging Face 请求 timeout | -| `hf_max_retries` | `3` | Hugging Face 请求最多尝试次数 | -| `pdf_timeout` | `90` 秒 | arXiv 下载 timeout | -| `max_pdf_bytes` | `52428800` | PDF 大小上限(50 MiB) | -| `max_pdf_pages` | `80` | 最多提取的 PDF 页数 | -| `max_pdf_chars` | `240000` | 单篇论文送入模型的最大提取字符数 | - -公开 Job 参数为 `date`、`force`、`top_k`、`weekly_weight` 和 `history_days`。调用时显式传入的值优先于 Job 默认值。 - -独立应用还支持以下环境变量: - -| 环境变量 | 用途 | -|-----------------------------------------|----------------------------------------| -| `DAILY_PAPER_WORKSPACE_DIR` | 覆盖默认 `reme_workspace` | -| `DAILY_PAPER_PROJECT_PATH` | Claude Code 可见的仓库或项目路径 | -| `REME_PROXY_IP` | 可选 SSH 代理主机 | -| `REME_PROXY_ACCOUNT` | 可选 SSH 代理账户 | -| `DAILY_PAPER_HOST` / `DAILY_PAPER_PORT` | HTTP 监听地址 | -| `CLAUDE_CODE_API_KEY` | Claude Code endpoint 的 API key | -| `CLAUDE_CODE_MODEL_NAME` | Claude Code 模型;默认 `qwen3.7-max` | -| `CLAUDE_CODE_BASE_URL` | Claude Code 的 Anthropic 兼容 endpoint | -| `LLM_API_KEY` | AgentScope memory 模型的 API key | -| `LLM_MODEL_NAME` | memory 模型;默认 `qwen3.7-max` | -| `LLM_BASE_URL` | memory 模型的 Anthropic 兼容 endpoint | - -`DAILY_PAPER_PROJECT_PATH` 默认是相对于 workspace 的 `..`。使用默认 `reme_workspace` 并从仓库根目录启动时, 它会解析回仓库根目录。如果 -workspace 位于其他位置,请显式设置这两个路径。 - -ReMe 会从当前目录向上查找未提交的 `.env`,因此也可以把相同变量放在其中,而不是在 shell 中导出。 - -## 钉钉配置 - -钉钉是可选能力。仅在需要投递简报或运行后台钉钉 Agent 时配置: +最后的 `dingtalk_markdown_send_step` 是可选步骤。未设置群会话 ID 时无副作用跳过;配置后会去掉 frontmatter,并把简报正文依次发送给所有群: ```dotenv DINGTALK_APP_KEY=your-app-key @@ -316,36 +112,113 @@ DINGTALK_ROBOT_CODE=your-robot-code DINGTALK_CONVERSATION_IDS=cid-group-one,cid-group-two ``` -只有主动投递简报需要 `DINGTALK_CONVERSATION_IDS`。后台 `dingtalk_wait` Job 使用前三项凭据,不使用群会话列表。 +任一群发送失败不会阻止继续尝试后续群,全部尝试结束后统一报告失败。 -## 故障恢复与边界 +## 产物 -| 场景 | 当前行为 | -|----------------------------|----------------------------------------------| -| Hugging Face 暂时失败 | 按指数间隔重试,最多尝试 `hf_max_retries` 次 | -| 没有 eligible 论文 | 在排序前失败 | -| `top_k` 或精选结果不合法 | 校验后失败;精选结果可重试一次 | -| PDF 太大、无效或没有文本层 | 在解读阶段停止 | -| PDF 超过页数或字符数限制 | 使用截断文本继续,并记录截断状态 | -| 某篇论文解读失败 | Job 停止;此前写入的 PDF 和笔记保留 | -| 简报遗漏源笔记链接 | 写入前自动补齐 wikilink | -| auto-dream 部分整合失败 | 成功单元保留,失败路径不会被 checkpoint | +```text +reme_workspace/ +├── daily/ +│ ├── YYYY-MM-DD.md +│ └── YYYY-MM-DD/ +│ ├── <中文论文标题>.md # 三篇,kind: daily-paper-analysis +│ └── <中文简报标题>.md # 一篇,kind: daily-paper-brief +└── resource/ + └── papers/ + └── .pdf +``` -恢复时,先检查该日期已有的笔记和 PDF,修复网络、凭据、模型或 PDF 问题,再使用相同日期和 `force=true` 重跑。 有效的缓存 PDF -会被复用。 +文件名来自 Agent 返回的中文标题。代码会清理路径不安全字符,并处理同名文件。Markdown 和 PDF 都通过同目录临时文件写入后原子替换。 -内置 Claude Code 组件使用 `permission_mode: bypassPermissions`,并禁用 `WebSearch`。`dingtalk_wait` 可以调用本地 -`memory_search`;`daily_paper` 当前没有配置 Job 工具。Analyze 和 Brief prompt 会限制 Agent 应读取的内容,但这些 -步骤没有设置严格的逐次调用工具 allowlist,也不是操作系统级沙箱。AgentScope memory wrapper 禁用了内置 shell 和 文件系统工具,但其 -ReMe job tools 运行于 bypass permission mode。请只在可信的项目和 workspace 中运行;用于共享 或生产环境前,应进一步收紧配置。 +## 参数与默认值 + +可在调用时传入的 Job 参数: + +| 参数 | 默认值 | 作用 | +|---|---:|---| +| `date` | `""` | 运行日期;空值使用应用时区当天,非空值必须为 `YYYY-MM-DD` | +| `force` | `false` | 已有当日简报时仍重新生成 | +| `topics` | `""` | 精选论文时优先考虑的主题 | +| `weekly_weight` | `0.7` | RRF 中周榜权重 | +| `history_days` | `30` | 历史推荐排重窗口 | + +`daily_paper` Job 的步骤级配置: + +| 配置 | 默认值 | 作用 | +|---|---:|---| +| `candidate_limit` | `20` | 送入 Select 的最大候选数 | +| `rrf_k` | `60` | RRF 常数 | +| `hf_timeout` | `600` 秒 | Hugging Face 单次请求超时 | +| `hf_max_retries` | `3` | Hugging Face 最大尝试次数 | +| `pdf_timeout` | `600` 秒 | arXiv PDF 下载超时 | +| `max_pdf_bytes` | `52428800` | PDF 上限,50 MiB | +| `max_pdf_pages` | `20` | 最多提取页数 | +| `max_pdf_chars` | `300000` | 最多送入 Agent 的 PDF 字符数 | + +## 镜像站 + +数据客户端使用 httpx 默认的环境处理,因此存在 `HTTP_PROXY`、`HTTPS_PROXY` 或 `NO_PROXY` 时会自动生效。镜像环境变量独立替换对应数据源的 base URL: + +```dotenv +# 未设置时使用 https://huggingface.co +HF_MIRROR_URL=https://hf-mirror.com + +# 未设置时使用 https://arxiv.org +ARXIV_MIRROR_URL=https://export.arxiv.org + +# 也支持带路径前缀的中转地址 +# HF_MIRROR_URL=http://relay-host:18080/hf +# ARXIV_MIRROR_URL=http://relay-host:18080/arxiv +``` + +`HF_MIRROR_URL` 必须提供当前代码使用的 `/papers/...`、`/api/daily_papers` 和 `/api/papers/...` 路径。`ARXIV_MIRROR_URL` 必须支持 `/pdf/`。两种 base URL 都会保留路径前缀,末尾 `/` 可有可无;不配置就直接访问官方站点,不会执行备用地址回退。 + +## 运行方式 + +生成指定日期的简报: + +```bash +reme start \ + config=daily_cookbook \ + job=daily_paper \ + date=2026-08-06 \ + topics="Agent memory" \ + history_days=30 +``` + +强制重跑;有效的本地 PDF 仍会复用: + +```bash +reme start config=daily_cookbook job=daily_paper date=2026-08-06 force=true +``` + +启动 HTTP 服务和定时任务: + +```bash +reme start config=daily_cookbook +``` + +内置服务监听 `127.0.0.1:8001`,`daily_paper_cron` 按 `Asia/Shanghai` 时区每天 08:00 运行。可通过 `DAILY_PAPER_HOST`、`DAILY_PAPER_PORT` 或启动参数覆盖监听地址和端口。 + +```bash +curl -s http://127.0.0.1:8001/daily_paper \ + -H 'Content-Type: application/json' \ + -d '{"date":"2026-08-06","force":false,"topics":"Agent memory"}' +``` + +## 失败与重跑 + +- Hugging Face 请求失败会指数退避重试,最多尝试 `hf_max_retries` 次。 +- 候选少于三篇、Agent 精选不合法、PDF 无效/过大/无文本或 Agent 输出为空都会终止 Job。 +- 三篇论文按顺序处理;中途失败时,之前已完成的 PDF 和笔记会保留。 +- `force=true` 会重新生成笔记和简报,但会复用有效 PDF。 +- 多文件流程不是事务,也没有同一日期的全局运行锁。 ## 测试 -聚焦的单元测试会 mock Hugging Face、arXiv、Claude Code 和钉钉边界: +单元测试会 mock Hugging Face、arXiv、AgentScope 和 DingTalk 边界,不访问真实服务: ```bash python -m pip install -e ".[dev,core]" pytest tests/unit/test_daily_paper.py -v ``` - -真实运行会访问外部服务并可能产生模型费用,不应把它当作普通单元测试执行。 diff --git a/example.env b/example.env index 5910169e..72bfd5e4 100644 --- a/example.env +++ b/example.env @@ -3,6 +3,7 @@ EMBEDDING_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 LLM_API_KEY=sk-xxx LLM_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1 -# Optional SSH jump host for Hugging Face daily-paper downloads. -# REME_PROXY_IP=proxy.example.com -# REME_PROXY_ACCOUNT=your-ssh-account +# Optional data-source mirrors. Unset variables use the official services. +# HF_MIRROR_URL=https://hf-mirror.com +# ARXIV_MIRROR_URL=https://export.arxiv.org +# TUSHARE_MIRROR_URL=http://112.124.63.173:4000/tushare diff --git a/reme/config/daily_cookbook.yaml b/reme/config/daily_cookbook.yaml index 031999b8..55831dd9 100644 --- a/reme/config/daily_cookbook.yaml +++ b/reme/config/daily_cookbook.yaml @@ -57,18 +57,12 @@ jobs: default: 7 steps: - backend: dream_extract_step - as_llm: memory - agent_wrapper: memory file_catalog: dream topic_session_id: interests scan_days: 2 max_units: 5 - backend: dream_integrate_step - as_llm: memory - agent_wrapper: memory - backend: dream_topics_step - as_llm: memory - agent_wrapper: memory topic_count: 3 topic_diversity_days: 7 - backend: dream_finish_step @@ -99,7 +93,6 @@ jobs: required: [messages] steps: - backend: auto_memory_step - agent_wrapper: memory reindex: backend: base @@ -325,7 +318,25 @@ jobs: auto_fin: backend: base - description: "Maintain configured daily news history, analyze current topics and historical ETF reactions, then deliver the report." + description: "Analyze configured ETFs from current news and comparable historical events." + etf_codes: &auto_fin_etf_codes + - 518880.SH + # - 159516.SZ + # - 512800.SH + # - 512890.SH + # - 159992.SZ + - 159530.SZ + # - 159869.SZ + # - 512690.SH + # - 159755.SZ + # - 159611.SZ + # - 159652.SZ + - 512760.SH + # - 159732.SZ + news_lookback_days: 60 + current_news_limit_per_etf: 10 + historical_news_limit: 5 + historical_search_limit: 10 parameters: type: object properties: @@ -333,22 +344,26 @@ jobs: type: string description: "Current date in YYYY-MM-DD; empty means today in Asia/Shanghai." default: "" - force: - type: boolean - description: "Refresh all configured news days; today's news is always refreshed." - default: false + now: + type: string + description: "Optional simulated current time in ISO 8601 format; empty means the real current time." + default: "" + historical_search_limit: + type: integer + description: "Maximum historical-news candidates retrieved for each current event." + default: 10 + minimum: 1 steps: &auto_fin_steps - backend: auto_fin_data_step - outbound_proxy: default - lookback_days: 360 - progress_interval: 30 + etf_codes: *auto_fin_etf_codes + news_lookback_days: 60 + - backend: update_index_step - backend: auto_fin_topic_step - agent_wrapper: auto_fin + current_news_limit_per_etf: 10 - backend: auto_fin_history_step - agent_wrapper: auto_fin - dispatch_steps: [auto_fin_history_search_step, auto_fin_market_step] + historical_news_limit: 5 + historical_search_limit: 10 - backend: auto_fin_merge_step - agent_wrapper: auto_fin - backend: dingtalk_markdown_send_step input_mapping: auto_fin_digest_path: markdown_path @@ -359,14 +374,17 @@ jobs: title: ReMe Auto Fin timeout: 15 + # Three intraday runs (Asia/Shanghai). Non-trading days are skipped inside + # auto_fin_data_step, so a daily trigger is fine. Each rerun refines the same + # day's report on top of the earlier run rather than replacing it wholesale. auto_fin_0930_cron: backend: cron cron: "30 9 * * *" steps: *auto_fin_steps - auto_fin_1145_cron: + auto_fin_1130_cron: backend: cron - cron: "45 11 * * *" + cron: "30 11 * * *" steps: *auto_fin_steps auto_fin_1800_cron: @@ -378,17 +396,15 @@ jobs: backend: base description: "Build detailed readings and a five-minute brief from Hugging Face weekly/monthly papers." candidate_limit: &candidate_limit 20 - memory_reserve: &memory_reserve 5 - top_k: &top_k 3 rrf_k: &rrf_k 60 weekly_weight: &weekly_weight 0.7 history_days: &history_days 30 - hf_timeout: &hf_timeout 30 + hf_timeout: &hf_timeout 600 hf_max_retries: &hf_max_retries 3 - pdf_timeout: &pdf_timeout 90 + pdf_timeout: &pdf_timeout 600 max_pdf_bytes: &max_pdf_bytes 52428800 - max_pdf_pages: &max_pdf_pages 80 - max_pdf_chars: &max_pdf_chars 240000 + max_pdf_pages: &max_pdf_pages 20 + max_pdf_chars: &max_pdf_chars 300000 parameters: type: object properties: @@ -400,10 +416,10 @@ jobs: type: boolean description: "Regenerate even when that day's final brief already exists." default: false - top_k: - type: integer - description: "Number of papers selected by Claude Code." - default: 3 + topics: + type: string + description: "Optional topics to prioritize when selecting papers." + default: "" weekly_weight: type: number description: "Weekly contribution in reciprocal-rank fusion." @@ -414,15 +430,10 @@ jobs: default: 30 steps: &daily_paper_steps - backend: daily_paper_collect_step - outbound_proxy: default - backend: daily_paper_rank_step - backend: daily_paper_select_step - agent_wrapper: daily_paper - backend: daily_paper_analyze_step - agent_wrapper: daily_paper - outbound_proxy: default - backend: daily_paper_digest_step - agent_wrapper: daily_paper - backend: dingtalk_markdown_send_step input_mapping: daily_paper_digest_path: markdown_path @@ -437,8 +448,6 @@ jobs: backend: cron cron: "0 8 * * *" candidate_limit: *candidate_limit - memory_reserve: *memory_reserve - top_k: *top_k rrf_k: *rrf_k weekly_weight: *weekly_weight history_days: *history_days @@ -456,87 +465,45 @@ jobs: close_timeout: 10 steps: - backend: dingtalk_wait_step - agent_wrapper: dingtalk_wait app_key: ${DINGTALK_APP_KEY:-} app_secret: ${DINGTALK_APP_SECRET:-} robot_code: ${DINGTALK_ROBOT_CODE:-} worker_count: 4 + builtin_tools: [bash] + job_tools: + - memory_search + - read + - write + - edit + - daily_list + - daily_write + - frontmatter_read + - frontmatter_update components: - outbound_proxy: {} -# default: -# backend: ssh_http -# host: ${REME_PROXY_IP:-} -# account: ${REME_PROXY_ACCOUNT:-} -# connect_timeout: 10 -# monitor_interval: 1 -# restart_initial_delay: 1 -# restart_max_delay: 30 - tokenizer: default: backend: regex as_llm: - memory: - backend: anthropic - model: ${LLM_MODEL_NAME:-qwen3.7-max} + default: + backend: openai + model: ${LLM_MODEL_NAME:-qwen3.7-plus} stream: true context_size: 200000 max_retries: 3 credential: api_key: ${LLM_API_KEY:-} - base_url: ${LLM_BASE_URL:-https://dashscope.aliyuncs.com/apps/anthropic} + base_url: ${LLM_BASE_URL:-} parameters: max_tokens: 65536 thinking_enable: false agent_wrapper: - auto_fin: - backend: ${AUTO_FIN_AGENT_BACKEND:-claude_code} - project_path: ${AUTO_FIN_PROJECT_PATH:-..} - model: ${CLAUDE_CODE_MODEL_NAME:-qwen3.7-max} - api_key: ${CLAUDE_CODE_API_KEY:-} - base_url: ${CLAUDE_CODE_BASE_URL:-https://dashscope.aliyuncs.com/apps/anthropic} - skills: [tushare-data] - job_tools: [memory_search] - permission_mode: bypassPermissions - daily_paper: - backend: claude_code - project_path: ${DAILY_PAPER_PROJECT_PATH:-..} - model: ${CLAUDE_CODE_MODEL_NAME:-qwen3.7-max} - api_key: ${CLAUDE_CODE_API_KEY:-} - base_url: ${CLAUDE_CODE_BASE_URL:-https://dashscope.aliyuncs.com/apps/anthropic} - permission_mode: bypassPermissions - dingtalk_wait: - backend: claude_code - project_path: ${DAILY_PAPER_PROJECT_PATH:-..} - model: ${CLAUDE_CODE_MODEL_NAME:-qwen3.7-max} - api_key: ${CLAUDE_CODE_API_KEY:-} - base_url: ${CLAUDE_CODE_BASE_URL:-https://dashscope.aliyuncs.com/apps/anthropic} - skills: [tushare-data] - job_tools: [memory_search] - system_prompt: - type: preset - preset: claude_code - append: >- - Daily-paper Markdown is stored under the ReMe workspace. Detailed notes, including historical notes, are at - daily/YYYY-MM-DD/paper-.md; daily briefs are at daily/YYYY-MM-DD/daily-paper-brief.md. Use - memory_search to retrieve relevant long-term notes across dates. - permission_mode: bypassPermissions - memory: + default: backend: agentscope - as_llm: memory + as_llm: default builtin_tools: false - permission_mode: bypass - react_config: - max_iters: 30 - context_config: - trigger_ratio: 0.8 - reserve_ratio: 0.1 - tool_result_limit: 50000 - model_config: - max_retries: 1 # as_embedding: # default: diff --git a/reme/schema/__init__.py b/reme/schema/__init__.py index 19c8efa8..e69e831b 100644 --- a/reme/schema/__init__.py +++ b/reme/schema/__init__.py @@ -2,27 +2,24 @@ from .application_config import ApplicationConfig, ComponentConfig, JobConfig from .auto_fin import ( - AutoFinEtfEventReference, - AutoFinEtfHistoryDetail, - AutoFinEtfHistoricalEvents, - AutoFinEtfHistoricalResearch, + AutoFinCurrentEvent, + AutoFinEtfAnalysis, AutoFinEtfSelection, AutoFinEtfsOutput, - AutoFinDailyEntry, - AutoFinForecastReturnPoint, - AutoFinFutureReturnPoint, + AutoFinEventReference, AutoFinHistoricalEvent, - AutoFinHistoricalEventReference, - AutoFinHistoricalDirectionReference, - AutoFinHistoricalMatch, - AutoFinMarketSelection, - AutoFinMarketSample, + AutoFinHistoricalOutput, + AutoFinHistoricalReference, AutoFinReportOutput, - AutoFinSelectedEvent, - AutoFinSelectedEtfAnalysis, - AutoFinWeightedForecast, + AutoFinReturns, +) +from .daily_paper import ( + AnalyzedPaper, + DailyPaperMarkdownOutput, + PaperInfo, + PaperPick, + PaperPickList, ) -from .daily_paper import DailyBriefOutput, PaperInfo, PaperNoteOutput, PaperSelection, SelectedPaper from .dream import ( DreamExtractOutput, DreamState, @@ -46,27 +43,19 @@ from .traverse_graph import TraverseGraph, TraverseGraphEdge, TraverseGraphNode __all__ = [ "ApplicationConfig", - "AutoFinEtfEventReference", - "AutoFinEtfHistoryDetail", - "AutoFinEtfHistoricalEvents", - "AutoFinEtfHistoricalResearch", + "AutoFinCurrentEvent", + "AutoFinEtfAnalysis", "AutoFinEtfSelection", "AutoFinEtfsOutput", - "AutoFinDailyEntry", - "AutoFinForecastReturnPoint", - "AutoFinFutureReturnPoint", + "AutoFinEventReference", "AutoFinHistoricalEvent", - "AutoFinHistoricalEventReference", - "AutoFinHistoricalDirectionReference", - "AutoFinHistoricalMatch", - "AutoFinMarketSelection", - "AutoFinMarketSample", + "AutoFinHistoricalOutput", + "AutoFinHistoricalReference", "AutoFinReportOutput", - "AutoFinSelectedEvent", - "AutoFinSelectedEtfAnalysis", - "AutoFinWeightedForecast", + "AutoFinReturns", "ComponentConfig", - "DailyBriefOutput", + "AnalyzedPaper", + "DailyPaperMarkdownOutput", "DreamExtractOutput", "DreamState", "DreamTopic", @@ -82,12 +71,11 @@ __all__ = [ "IntegrateOutcome", "JobConfig", "PaperInfo", - "PaperNoteOutput", - "PaperSelection", + "PaperPick", + "PaperPickList", "ProactiveResult", "Request", "Response", - "SelectedPaper", "StreamChunk", "TokenUsage", "TopicSelectionOutput", diff --git a/reme/schema/auto_fin.py b/reme/schema/auto_fin.py index 55680a24..caa8b31f 100644 --- a/reme/schema/auto_fin.py +++ b/reme/schema/auto_fin.py @@ -2,298 +2,102 @@ from __future__ import annotations -from datetime import date, datetime -from math import isclose -from typing import Annotated, Literal -from zoneinfo import ZoneInfo +from datetime import datetime +from typing import Literal -from pydantic import BaseModel, BeforeValidator, ConfigDict, Field, model_validator - -_SHANGHAI = ZoneInfo("Asia/Shanghai") - - -def _shanghai_local_time(value): - """Normalize aware input to naive Shanghai wall-clock time.""" - if not isinstance(value, (str, datetime)): - return value - parsed = value if isinstance(value, datetime) else datetime.fromisoformat(value) - if parsed.tzinfo is not None and parsed.utcoffset() is not None: - parsed = parsed.astimezone(_SHANGHAI).replace(tzinfo=None) - return parsed - - -ShanghaiDateTime = Annotated[datetime, BeforeValidator(_shanghai_local_time)] +from pydantic import BaseModel, ConfigDict, Field class AutoFinModel(BaseModel): - """Strict base for program-owned Auto Fin data.""" + """Strict program-owned Auto Fin data.""" model_config = ConfigDict(extra="forbid") class AutoFinAgentModel(AutoFinModel): - """Tolerant base for raw Agent output.""" + """Agent output tolerant of harmless extra fields.""" model_config = ConfigDict(extra="ignore") -class AutoFinEtfEventReference(AutoFinAgentModel): - """One selected news item and why it is relevant to an ETF.""" +class AutoFinEventReference(AutoFinAgentModel): + """One current news item related to an ETF.""" - reason: str news_id: str - - -class AutoFinSelectedEvent(AutoFinModel): - """A selected current event with its source news reference.""" - - event_time: ShanghaiDateTime - event_content: str reason: str - news_id: str - event_title: str = "" class AutoFinEtfSelection(AutoFinAgentModel): - """One ETF selection returned by the Topic Agent.""" + """One ETF selected from the configured codes.""" etf_code: str etf_name: str = "" - events: list[AutoFinEtfEventReference] = Field(default_factory=list) + events: list[AutoFinEventReference] = Field(default_factory=list) class AutoFinEtfsOutput(AutoFinAgentModel): - """ETF selections returned by the Topic Agent before normalization.""" + """Selections returned by the first Agent.""" etfs: list[AutoFinEtfSelection] = Field(default_factory=list) -class AutoFinHistoricalEventReference(AutoFinAgentModel): - """One historical news item selected by the search Agent.""" +class AutoFinHistoricalReference(AutoFinAgentModel): + """One historical event selected by the second Agent.""" - reason: str news_id: str - source_path: str = "" + reason: str + direction: Literal["same", "opposite"] -class AutoFinEtfHistoricalEvents(AutoFinAgentModel): - """Historical news references returned by the search Agent.""" +class AutoFinHistoricalOutput(AutoFinAgentModel): + """Historical matches for one current news item.""" - etf_code: str = "" - etf_name: str = "" - historical_events: list[AutoFinHistoricalEventReference] = Field(default_factory=list) + historical_events: list[AutoFinHistoricalReference] = Field(default_factory=list) -class AutoFinDailyEntry(AutoFinModel): - """First daily open or close that can be traded after an event.""" +class AutoFinReturns(AutoFinModel): + """Adjusted cumulative ETF returns after one historical event.""" - entry_time: ShanghaiDateTime - trade_date: date - price_type: Literal["open", "close"] - raw_price: float = Field(gt=0) - adj_factor: float = Field(gt=0) - - @model_validator(mode="after") - def valid_entry_timestamp(self) -> "AutoFinDailyEntry": - """Require a Shanghai-local timestamp matching the daily price.""" - if self.entry_time.date() != self.trade_date: - raise ValueError("entry time and trade date must match") - expected_clock = (9, 30) if self.price_type == "open" else (15, 0) - if ( - (self.entry_time.hour, self.entry_time.minute) != expected_clock - or self.entry_time.second - or self.entry_time.microsecond - ): - raise ValueError(f"{self.price_type} entry time must use the official daily price timestamp") - return self - - -class AutoFinFutureReturnPoint(AutoFinModel): - """Cumulative adjusted return at one future valid close.""" - - horizon: int = Field(ge=1, le=10) - trade_date: date - raw_close: float = Field(gt=0) - adj_factor: float = Field(gt=0) - cumulative_return: float - - -class AutoFinMarketSample(AutoFinModel): - """Daily adjusted ETF returns following one historical event.""" - - event_time: ShanghaiDateTime - entry: AutoFinDailyEntry | None = None - future_returns: list[AutoFinFutureReturnPoint] = Field(default_factory=list, max_length=10) - reaction_summary: str - - @model_validator(mode="after") - def valid_daily_return_path(self) -> "AutoFinMarketSample": - """Reject look-ahead entries and inconsistent adjusted returns.""" - if self.entry is None: - if self.future_returns: - raise ValueError("future returns require an entry") - return self - if self.entry.entry_time <= self.event_time: - raise ValueError("entry must be strictly after the event") - - expected_horizons = list(range(1, len(self.future_returns) + 1)) - if [point.horizon for point in self.future_returns] != expected_horizons: - raise ValueError("future return horizons must be contiguous and start at 1") - trade_dates = [point.trade_date for point in self.future_returns] - if trade_dates != sorted(set(trade_dates)): - raise ValueError("future return trade dates must be unique and ascending") - if trade_dates: - first_trade_date = trade_dates[0] - if self.entry.price_type == "open" and first_trade_date < self.entry.trade_date: - raise ValueError("an open entry cannot use an earlier close") - if self.entry.price_type == "close" and first_trade_date <= self.entry.trade_date: - raise ValueError("a close entry requires a later close") - - adjusted_entry = self.entry.raw_price * self.entry.adj_factor - for point in self.future_returns: - expected_return = point.raw_close * point.adj_factor / adjusted_entry - 1 - if not isclose(point.cumulative_return, expected_return, rel_tol=1e-6, abs_tol=1e-6): - raise ValueError(f"incorrect adjusted return at horizon {point.horizon}") - return self + d1: float | None = None + d2: float | None = None + d3: float | None = None + d5: float | None = None class AutoFinHistoricalEvent(AutoFinModel): - """One resolved historical news item and its calculated ETF return path.""" + """A resolved historical event and its observed ETF performance.""" - reason: str news_id: str - source_path: str - event_time: ShanghaiDateTime - event_title: str - event_content: str - market_entry: AutoFinDailyEntry | None = None - future_returns: list[AutoFinFutureReturnPoint] = Field(default_factory=list, max_length=10) - - @model_validator(mode="after") - def valid_historical_event(self) -> "AutoFinHistoricalEvent": - """Require source identity and validate the embedded market reaction.""" - for field in ("reason", "news_id", "source_path", "event_title", "event_content"): - value = getattr(self, field).strip() - if not value: - raise ValueError(f"historical event {field} must not be empty") - setattr(self, field, value) - AutoFinMarketSample( - event_time=self.event_time, - entry=self.market_entry, - future_returns=self.future_returns, - reaction_summary="", - ) - return self - - -class AutoFinEtfHistoricalResearch(AutoFinModel): - """Resolved historical events with embedded calculated ETF return paths.""" - - etf_code: str - etf_name: str - historical_events: list[AutoFinHistoricalEvent] = Field(default_factory=list) - limitations: list[str] = Field(default_factory=list) - - @model_validator(mode="after") - def unique_historical_news(self) -> "AutoFinEtfHistoricalResearch": - """Reject duplicate resolved source records.""" - news_ids = [event.news_id for event in self.historical_events] - if len(news_ids) != len(set(news_ids)): - raise ValueError("historical event news IDs must be unique") - return self - - -class AutoFinHistoricalDirectionReference(AutoFinAgentModel): - """One direction-classified historical event returned by the Market Agent.""" - + event_time: datetime + title: str + content: str reason: str - news_id: str - - -class AutoFinMarketSelection(AutoFinAgentModel): - """Same- and opposite-direction historical events returned by the Market Agent.""" - - same_direction_events: list[AutoFinHistoricalDirectionReference] = Field(default_factory=list) - opposite_direction_events: list[AutoFinHistoricalDirectionReference] = Field(default_factory=list) - - -class AutoFinHistoricalMatch(AutoFinModel): - """One direction-classified historical event used by the equal-weight forecast.""" - - reason: str - news_id: str - event_time: ShanghaiDateTime direction: Literal["same", "opposite"] - weight: float = Field(ge=0.0, le=1.0) + returns: AutoFinReturns -class AutoFinForecastReturnPoint(AutoFinModel): - """Weighted expected cumulative return for one holding horizon.""" +class AutoFinCurrentEvent(AutoFinModel): + """One current event with comparable historical evidence.""" - horizon: int = Field(ge=1, le=10) - expected_return: float | None = None + news_id: str + event_time: datetime + title: str + content: str + reason: str + historical_events: list[AutoFinHistoricalEvent] = Field(default_factory=list) -class AutoFinWeightedForecast(AutoFinModel): - """Program-calculated forecast derived from similar historical events.""" - - returns: list[AutoFinForecastReturnPoint] = Field(min_length=10, max_length=10) - suggested_holding_days: int | None = Field(default=None, ge=1, le=10) - - @model_validator(mode="after") - def complete_horizons(self) -> "AutoFinWeightedForecast": - """Require one ordered forecast point for every D1-D10 horizon.""" - if [point.horizon for point in self.returns] != list(range(1, 11)): - raise ValueError("forecast horizons must be ordered D1-D10") - return self - - -class AutoFinSelectedEtfAnalysis(AutoFinModel): - """Program-calculated weighted forecast for one selected ETF.""" +class AutoFinEtfAnalysis(AutoFinModel): + """All evidence prepared for the final Agent for one ETF.""" etf_code: str etf_name: str - matched_historical_events: list[AutoFinHistoricalMatch] = Field(default_factory=list) - forecast: AutoFinWeightedForecast - limitations: list[str] = Field(default_factory=list) - - @model_validator(mode="after") - def valid_historical_weights(self) -> "AutoFinSelectedEtfAnalysis": - """Reject duplicate matches and invalid normalized weights.""" - news_ids = [event.news_id for event in self.matched_historical_events] - if len(news_ids) != len(set(news_ids)): - raise ValueError("matched historical events must be unique") - if news_ids and not isclose( - sum(event.weight for event in self.matched_historical_events), - 1.0, - rel_tol=1e-6, - abs_tol=1e-6, - ): - raise ValueError("matched historical event weights must sum to 1") - return self - - -class AutoFinEtfHistoryDetail(AutoFinModel): - """Complete historical research and market result for one selected ETF.""" - - etf: AutoFinEtfSelection - current_events: list[AutoFinSelectedEvent] = Field(min_length=1) - historical_research: AutoFinEtfHistoricalResearch - market_analysis: AutoFinSelectedEtfAnalysis - - @model_validator(mode="after") - def consistent_etf_and_events(self) -> "AutoFinEtfHistoryDetail": - """Reject stale or cross-ETF outputs from dispatched steps.""" - identity = (self.etf.etf_code, self.etf.etf_name) - if (self.historical_research.etf_code, self.historical_research.etf_name) != identity: - raise ValueError("historical research ETF must match the selected ETF") - if (self.market_analysis.etf_code, self.market_analysis.etf_name) != identity: - raise ValueError("market analysis ETF must match the selected ETF") - return self + events: list[AutoFinCurrentEvent] = Field(default_factory=list) class AutoFinReportOutput(AutoFinAgentModel): - """Final Markdown title and body for all selected ETFs.""" + """Final Markdown returned by the third Agent.""" title: str = "" + description: str = "" body: str = "" diff --git a/reme/schema/daily_paper.py b/reme/schema/daily_paper.py index 025ceb40..67713b59 100644 --- a/reme/schema/daily_paper.py +++ b/reme/schema/daily_paper.py @@ -1,7 +1,5 @@ """Typed contracts for the daily-paper cookbook workflow.""" -from typing import Literal - from pydantic import BaseModel, Field @@ -40,32 +38,31 @@ class PaperInfo(BaseModel): return f"https://arxiv.org/pdf/{self.arxiv_id}" -class SelectedPaper(BaseModel): - """One agent-selected paper.""" +class PaperPick(BaseModel): + """One minimal paper selection returned by the agent.""" arxiv_id: str - rank: int - reason: str - memory_relevance: Literal["high", "medium", "low"] + reasoning: str -class PaperSelection(BaseModel): - """Structured paper selection result.""" +class PaperPickList(BaseModel): + """The ordered papers selected for detailed analysis.""" - selection_reasoning: str - selected: list[SelectedPaper] - alternates: list[str] = Field(default_factory=list) + papers: list[PaperPick] -class PaperNoteOutput(BaseModel): - """Structured Claude Code output for one detailed paper note.""" +class DailyPaperMarkdownOutput(BaseModel): + """One Chinese Markdown document returned by a tool-free agent.""" - description: str + title: str + desc: str body: str -class DailyBriefOutput(BaseModel): - """Structured Claude Code output for the final five-minute brief.""" +class AnalyzedPaper(DailyPaperMarkdownOutput): + """A persisted paper analysis passed directly to the digest step.""" - description: str - body: str + arxiv_id: str + reasoning: str + note_path: str + pdf_path: str diff --git a/reme/steps/cookbook/auto_fin/__init__.py b/reme/steps/cookbook/auto_fin/__init__.py index f78ef316..a3a5076d 100644 --- a/reme/steps/cookbook/auto_fin/__init__.py +++ b/reme/steps/cookbook/auto_fin/__init__.py @@ -1,17 +1,13 @@ """Auto Fin news research workflow.""" from .data import AutoFinDataStep -from .history_search import AutoFinHistorySearchStep from .history import AutoFinHistoryStep -from .market import AutoFinMarketStep from .merge import AutoFinMergeStep from .topic import AutoFinTopicStep __all__ = [ "AutoFinDataStep", - "AutoFinHistorySearchStep", "AutoFinHistoryStep", - "AutoFinMarketStep", "AutoFinMergeStep", "AutoFinTopicStep", ] diff --git a/reme/steps/cookbook/auto_fin/_base.py b/reme/steps/cookbook/auto_fin/_base.py index 5f98df17..34646822 100644 --- a/reme/steps/cookbook/auto_fin/_base.py +++ b/reme/steps/cookbook/auto_fin/_base.py @@ -4,6 +4,7 @@ from __future__ import annotations import asyncio import hashlib +from html.parser import HTMLParser import json import math import os @@ -16,16 +17,44 @@ from zoneinfo import ZoneInfo from pydantic import BaseModel -from ....components.outbound_proxy import BaseOutboundProxy -from ....enumeration import ComponentEnum from ....utils.tushare import create_tushare_api -from ...base_step import BaseStep, Ref +from ...base_step import BaseStep AGENT_INPUT_LOG_LIMIT = 2000 AGENT_OUTPUT_LOG_LIMIT = 4000 SHANGHAI_TIMEZONE = ZoneInfo("Asia/Shanghai") +class _TextExtractor(HTMLParser): + def __init__(self) -> None: + super().__init__(convert_charrefs=True) + self.parts: list[str] = [] + self.hidden = 0 + + def handle_starttag(self, tag: str, _attrs) -> None: + if tag in {"script", "style"}: + self.hidden += 1 + elif tag in {"br", "div", "li", "p"}: + self.parts.append(" ") + + def handle_endtag(self, tag: str) -> None: + if tag in {"script", "style"} and self.hidden: + self.hidden -= 1 + elif tag in {"div", "li", "p"}: + self.parts.append(" ") + + def handle_data(self, data: str) -> None: + if not self.hidden: + self.parts.append(data) + + +def _plain_text(value: str) -> str: + parser = _TextExtractor() + parser.feed(value) + parser.close() + return " ".join("".join(parser.parts).split()) + + def _news_hash(row: dict[str, Any]) -> str: src = str(row.get("src") or "") content = str(row.get("content") or "") @@ -81,8 +110,6 @@ def _records(value: Any) -> list[dict[str, Any]]: class AutoFinStep(BaseStep): """Shared Auto Fin helpers.""" - outbound_proxy: BaseOutboundProxy | None = Ref(BaseOutboundProxy, ComponentEnum.OUTBOUND_PROXY, optional=True) - def _value(self, key: str, default: Any = None) -> Any: assert self.context is not None return self.context.get(key, self.kwargs.get(key, default)) @@ -98,10 +125,6 @@ class AutoFinStep(BaseStep): prompt_name: str, resource_name: str, model: type[BaseModel], - *, - output_suffix: str = ".json", - jsonl_field: str | None = None, - tool_context_id: str | None = None, **values: str, ) -> tuple[BaseModel, Path]: """Send a complete prompt directly and persist its structured reply.""" @@ -109,15 +132,9 @@ class AutoFinStep(BaseStep): raise RuntimeError("Auto Fin analysis requires an agent_wrapper") if Path(resource_name).name != resource_name: raise ValueError(f"Invalid Auto Fin resource name: {resource_name}") - if output_suffix not in {".json", ".jsonl"}: - raise ValueError(f"Invalid Auto Fin output suffix: {output_suffix}") - prompt = self.prompt_format(prompt_name, **values) output_path = ( - self.workspace_path - / "resource" - / str(self._required("auto_fin_date")) - / f"{resource_name}_output{output_suffix}" + self.workspace_path / "resource" / str(self._required("auto_fin_date")) / f"{resource_name}_output.json" ) started_at = perf_counter() serialized_input = json.dumps(prompt, ensure_ascii=False) @@ -126,25 +143,14 @@ class AutoFinStep(BaseStep): f"[{self.name}] agent input prompt={prompt_name} schema={model.__name__} " f"query_chars={len(prompt)} truncated={str(input_truncated).lower()} query={input_preview}", ) - agent_kwargs: dict[str, Any] = {"output_schema": model} - if tool_context_id: - agent_kwargs["tool_context_id"] = tool_context_id - result = await self.agent_wrapper.reply(prompt, **agent_kwargs) + result = await self.agent_wrapper.reply(prompt, output_schema=model) if not isinstance(result, dict): raise TypeError("Auto Fin Agent reply must be a dictionary") value = result.get("structured_output") if value is None: raise ValueError(f"Auto Fin Agent returned no structured output: {self._preview(result)}") output = value if isinstance(value, model) else model.model_validate(value) - payload = output.model_dump(mode="json") - serialized_output = json.dumps(payload, ensure_ascii=False, separators=(",", ":")) - if jsonl_field is None: - _write(output_path, f"{serialized_output}\n") - else: - records = payload.get(jsonl_field) - if not isinstance(records, list) or not all(isinstance(record, dict) for record in records): - raise ValueError(f"Auto Fin JSONL field must contain objects: {jsonl_field}") - _write_jsonl(output_path, records) + serialized_output = self._write_output(output_path, output) output_preview, output_truncated = self._text_preview(serialized_output, AGENT_OUTPUT_LOG_LIMIT) self.logger.info( f"[{self.name}] agent output prompt={prompt_name} schema={model.__name__} " @@ -153,6 +159,13 @@ class AutoFinStep(BaseStep): ) return output, output_path + @staticmethod + def _write_output(path: Path, model: BaseModel) -> str: + """Persist a model as compact JSON and return the serialized text.""" + serialized = json.dumps(model.model_dump(mode="json"), ensure_ascii=False, separators=(",", ":")) + _write(path, f"{serialized}\n") + return serialized + @staticmethod def _text_preview(text: str, limit: int) -> tuple[str, bool]: limit = max(0, limit) @@ -162,11 +175,7 @@ class AutoFinStep(BaseStep): @staticmethod def _preview(value: Any, limit: int = 1000) -> str: text = json.dumps(value, ensure_ascii=False, default=str) - return text if len(text) <= limit else f"{text[:limit]}..." - - @property - def _proxy_url(self) -> str | None: - return self.outbound_proxy.http_url if self.outbound_proxy is not None else None + return AutoFinStep._text_preview(text, limit)[0] async def _fetch(self, endpoint: str, **kwargs) -> list[dict[str, Any]]: provider = self._value("tushare_provider") @@ -176,8 +185,7 @@ class AutoFinStep(BaseStep): provider_name = "injected" if provider is not None else "sdk" started_at = perf_counter() self.logger.debug( - f"[{self.name}] tushare fetch start endpoint={endpoint} provider={provider_name} " - f"proxy={bool(self._proxy_url)} {details}", + f"[{self.name}] tushare fetch start endpoint={endpoint} provider={provider_name} {details}", ) try: if provider is not None: @@ -187,7 +195,7 @@ class AutoFinStep(BaseStep): token = os.getenv("TUSHARE_TOKEN", "").strip() if not token: raise RuntimeError("TUSHARE_TOKEN is required for Auto Fin") - api = create_tushare_api(token, proxy_url=self._proxy_url) + api = create_tushare_api(token) rows = _records(await asyncio.to_thread(getattr(api, endpoint), **kwargs)) except Exception: self.logger.exception( @@ -201,10 +209,6 @@ class AutoFinStep(BaseStep): ) return rows - def _news_path(self, day: date) -> Path: - daily_dir = str(self.config_value("daily_dir")) - return self.workspace_path / daily_dir / day.isoformat() / "auto_fin_news_data.jsonl" - @staticmethod def _days(start: date, end: date) -> list[date]: return [start + timedelta(days=offset) for offset in range((end - start).days + 1)] diff --git a/reme/steps/cookbook/auto_fin/data.py b/reme/steps/cookbook/auto_fin/data.py index ed1c9bec..e1ea8510 100644 --- a/reme/steps/cookbook/auto_fin/data.py +++ b/reme/steps/cookbook/auto_fin/data.py @@ -1,24 +1,29 @@ -"""Download the news required by Auto Fin.""" +"""Prepare local news and ETF market data for Auto Fin.""" from __future__ import annotations import asyncio +import json from datetime import date, datetime, time, timedelta from pathlib import Path from typing import Any from ....components import R -from ._base import SHANGHAI_TIMEZONE, AutoFinStep, _news_id, _write_jsonl +from ._base import SHANGHAI_TIMEZONE, AutoFinStep, _news_id, _plain_text, _write, _write_jsonl + +NEWS_FILENAME = "auto_fin_news.md" +MAJOR_NEWS_PAGE_LIMIT = 400 # major_news caps a single response; split the window when hit. +FUND_PAGE_LIMIT = 2000 # fund_daily / fund_adj cap a single response; page backwards past it. @R.register("auto_fin_data_step") class AutoFinDataStep(AutoFinStep): - """Fill missing daily news files and always refresh today's news.""" + """Skip closed markets, then overwrite today's news and all configured ETF data.""" def _schedule(self) -> tuple[date, datetime]: - now_value = self._value("now") - now = datetime.fromisoformat(str(now_value)) if now_value is not None else datetime.now(SHANGHAI_TIMEZONE) - if now.tzinfo is not None and now.utcoffset() is not None: + value = str(self._value("now", "")).strip() + now = datetime.fromisoformat(value) if value else datetime.now(SHANGHAI_TIMEZONE) + if now.tzinfo is not None: now = now.astimezone(SHANGHAI_TIMEZONE).replace(tzinfo=None) requested = str(self._value("date", "")).strip() run_date = date.fromisoformat(requested) if requested else now.date() @@ -26,43 +31,20 @@ class AutoFinDataStep(AutoFinStep): raise ValueError("Auto Fin only supports the current date") return run_date, now - async def _previous_trade_date(self, run_date: date) -> date: - supplied = self._value("trade_dates") - if supplied is not None: - dates = [date.fromisoformat(str(value)) for value in supplied] - else: - start = run_date - timedelta(days=30) - rows = await self._fetch( - "trade_cal", - exchange="SSE", - start_date=start.strftime("%Y%m%d"), - end_date=run_date.strftime("%Y%m%d"), - fields="cal_date,is_open", - ) - dates = [ - datetime.strptime(str(row["cal_date"]), "%Y%m%d").date() - for row in rows - if int(row.get("is_open", 0)) == 1 - ] - previous = [day for day in dates if day < run_date] - if not previous: - raise ValueError("Auto Fin requires a previous A-share trade date") - return max(previous) + async def _is_trade_day(self, day: date) -> bool: + rows = await self._fetch( + "trade_cal", + exchange="SSE", + start_date=day.strftime("%Y%m%d"), + end_date=day.strftime("%Y%m%d"), + fields="cal_date,is_open", + ) + return any( + str(row.get("cal_date")) == day.strftime("%Y%m%d") and int(row.get("is_open", 0)) == 1 for row in rows + ) - async def _valid_news(self, path: Path) -> bool: - if not path.is_file(): - return False - try: - rows = await self._read_jsonl(path) - except (OSError, ValueError) as exc: - self.logger.warning( - f"[{self.name}] invalid news cache path={path} error={type(exc).__name__}: {exc}", - ) - return False - valid = all(str(row.get("src") or "") == "财联社" for row in rows) - if not valid: - self.logger.warning(f"[{self.name}] invalid news cache source path={path}") - return valid + def _news_path(self, day: date) -> Path: + return self.workspace_path / str(self.config_value("daily_dir")) / day.isoformat() / NEWS_FILENAME async def _fetch_news(self, start: datetime, end: datetime) -> list[dict[str, Any]]: rows = await self._fetch( @@ -72,83 +54,181 @@ class AutoFinDataStep(AutoFinStep): end_date=end.strftime("%Y-%m-%d %H:%M:%S"), fields="title,pub_time,src,content", ) - if len(rows) < 400 or end - start <= timedelta(minutes=1): + if len(rows) < MAJOR_NEWS_PAGE_LIMIT or end - start <= timedelta(minutes=1): return rows - midpoint = start + (end - start) / 2 - self.logger.debug( - f"[{self.name}] news fetch split start={start.isoformat()} end={end.isoformat()} " - f"midpoint={midpoint.isoformat()} records={len(rows)}", - ) - left, right = await asyncio.gather(self._fetch_news(start, midpoint), self._fetch_news(midpoint, end)) - self.logger.debug( - f"[{self.name}] news fetch split done start={start.isoformat()} end={end.isoformat()} " - f"records={len(left) + len(right)}", - ) + middle = start + (end - start) / 2 + left, right = await asyncio.gather(self._fetch_news(start, middle), self._fetch_news(middle, end)) return left + right - async def _cache_news(self, day: date, decision_at: datetime, refresh: bool) -> bool: - path = self._news_path(day) - if not refresh and await self._valid_news(path): - self.logger.debug(f"[{self.name}] news cache hit date={day.isoformat()} path={path}") - return False + async def _write_news(self, day: date, decision_at: datetime) -> str: start = datetime.combine(day, time.min) end = decision_at if day == decision_at.date() else start + timedelta(days=1) - candidates = [] + records: dict[str, dict[str, str]] = {} for row in await self._fetch_news(start, end): published_at = self._published_at(row) - in_range = ( - published_at is not None - and start <= published_at - and (published_at <= end if day == decision_at.date() else published_at < end) + if published_at is None or str(row.get("src") or "") != "财联社": + continue + if not start <= published_at <= end or (day != decision_at.date() and published_at == end): + continue + news_id = _news_id(row, published_at) + records.setdefault( + news_id, + { + "news_id": news_id, + "event_time": published_at.isoformat(), + "title": _plain_text(str(row.get("title") or "")), + "content": _plain_text(str(row.get("content") or "")), + }, ) - if in_range and str(row.get("src") or "") == "财联社": - candidates.append((published_at, _news_id(row, published_at), row)) - news = {} - for _published_at, news_id, row in sorted(candidates, key=lambda item: item[:2]): - news.setdefault(news_id, {**row, "news_id": news_id}) - _write_jsonl(path, list(news.values())) - self.logger.debug(f"[{self.name}] news written date={day.isoformat()} records={len(news)} path={path}") - return True + ordered = sorted(records.values(), key=lambda row: (row["event_time"], row["news_id"])) + path = self._news_path(day) + change = "modified" if path.exists() else "added" + _write(path, self._render_news(day, ordered)) + return change + + @staticmethod + def _render_news(day: date, rows: list[dict[str, str]]) -> str: + blocks = [f"# 财联社新闻 {day.isoformat()}\n"] + for row in rows: + blocks.append( + "\n".join( + [ + f"## {row['title'] or '无标题'}", + "", + f"- news_id: `{row['news_id']}`", + f"- 时间: {row['event_time']}", + "- 来源: 财联社", + "", + row["content"] or row["title"], + "", + ], + ), + ) + return "\n".join(blocks).rstrip() + "\n" + + @staticmethod + def read_news(path: Path) -> list[dict[str, str]]: + """Parse an Auto Fin news Markdown file written by `_render_news`.""" + text = path.read_text(encoding="utf-8") + rows = [] + for block in text.split("\n## ")[1:]: + lines = block.splitlines() + if len(lines) < 5: + continue + news_line = next((line for line in lines if line.startswith("- news_id: `")), "") + time_line = next((line for line in lines if line.startswith("- 时间: ")), "") + news_id = news_line.removeprefix("- news_id: `").removesuffix("`").strip() + event_time = time_line.removeprefix("- 时间: ").strip() + content_start = next( + (index + 1 for index, line in enumerate(lines) if line == "- 来源: 财联社"), + len(lines), + ) + content = "\n".join(lines[content_start:]).strip() + if news_id and event_time: + rows.append( + { + "news_id": news_id, + "event_time": event_time, + "title": lines[0].strip(), + "content": content, + }, + ) + return rows + + async def _fetch_all(self, endpoint: str, code: str, end: date) -> list[dict[str, Any]]: + """Page backwards because TuShare fund endpoints cap one response.""" + rows_by_date: dict[str, dict[str, Any]] = {} + end_date = end + while True: + page = await self._fetch( + endpoint, + ts_code=code, + start_date="19900101", + end_date=end_date.strftime("%Y%m%d"), + ) + for row in page: + if trade_date := str(row.get("trade_date") or ""): + rows_by_date[trade_date] = row + if len(page) < FUND_PAGE_LIMIT: + break + dates = [ + datetime.strptime(str(row["trade_date"]), "%Y%m%d").date() for row in page if row.get("trade_date") + ] + if not dates or min(dates) <= date(1990, 1, 1): + break + next_end = min(dates) - timedelta(days=1) + if next_end >= end_date: + raise RuntimeError(f"TuShare pagination did not advance for {endpoint} {code}") + end_date = next_end + return [rows_by_date[key] for key in sorted(rows_by_date)] + + @staticmethod + def _etf_name(row: dict[str, Any]) -> str: + return str(row.get("csname") or row.get("extname") or row.get("cname") or "").strip() + + async def _cache_etfs(self, codes: list[str], run_date: date) -> dict[str, str]: + basics = await self._fetch( + "etf_basic", + list_status="L", + fields="ts_code,csname,extname,cname,list_status", + ) + names = {str(row.get("ts_code") or "").strip().upper(): self._etf_name(row) for row in basics} + missing = [code for code in codes if not names.get(code)] + if missing: + raise ValueError(f"TuShare returned no ETF name for: {', '.join(missing)}") + + fin_dir = self.workspace_path / str(self.config_value("resource_dir")) / "fin" + mapping = [{"etf_code": code, "etf_name": names[code]} for code in codes] + _write( + fin_dir / "etfs.json", + json.dumps(mapping, ensure_ascii=False, indent=2) + "\n", + ) + for code in codes: + daily, factors = await asyncio.gather( + self._fetch_all("fund_daily", code, run_date), + self._fetch_all("fund_adj", code, run_date), + ) + factor_by_date = {str(row.get("trade_date")): row.get("adj_factor") for row in factors} + merged = [{**row, "adj_factor": factor_by_date.get(str(row.get("trade_date")))} for row in daily] + _write_jsonl(fin_dir / f"{code}.jsonl", merged) + return {code: names[code] for code in codes} async def execute(self): assert self.context is not None run_date, decision_at = self._schedule() - news_days = int(self._value("lookback_days")) - progress_interval = int(self._value("progress_interval")) - if news_days < 1: - raise ValueError("lookback_days must be at least 1") - if progress_interval < 1: - raise ValueError("progress_interval must be at least 1") - start = run_date - timedelta(days=news_days - 1) - force = bool(self._value("force", False)) - self.logger.info( - f"[{self.name}] start date={run_date.isoformat()} range={start.isoformat()}..{run_date.isoformat()} " - f"days={news_days} force={force} decision_at={decision_at.isoformat()}", - ) - previous_trade_date = await self._previous_trade_date(run_date) - self.logger.info( - f"[{self.name}] trade date resolved date={run_date.isoformat()} " - f"previous_trade_date={previous_trade_date.isoformat()}", - ) - downloaded = 0 - for processed, day in enumerate(self._days(start, run_date), start=1): - downloaded += int(await self._cache_news(day, decision_at, force or day == run_date)) - if processed % progress_interval == 0 and processed < news_days: - self.logger.info( - f"[{self.name}] progress processed={processed}/{news_days} downloaded={downloaded} " - f"cached={processed - downloaded} last_date={day.isoformat()}", - ) + if not await self._is_trade_day(run_date): + reason = f"{run_date.isoformat()} 不是交易日,Auto Fin 已跳过。" + self.context["auto_fin_skipped"] = True + self.context["auto_fin_skip_reason"] = reason + self.context.response.answer = reason + self.context.response.metadata.update({"date": run_date.isoformat(), "skipped": True}) + return self.context.response + + lookback = int(self._value("news_lookback_days", 60)) + if lookback < 1: + raise ValueError("news_lookback_days must be positive") + news_start = run_date - timedelta(days=lookback - 1) + changes = [] + for day in self._days(news_start, run_date): + path = self._news_path(day) + if day != run_date and path.is_file(): + continue + change = await self._write_news(day, decision_at) + changes.append({"change": change, "path": str(path)}) + + codes = self._value("etf_codes") + if not codes: + raise ValueError("auto_fin_data_step requires a non-empty etf_codes") + codes = [str(code).strip().upper() for code in codes] + names = await self._cache_etfs(codes, run_date) self.context.update( { + "changes": changes, "auto_fin_date": run_date.isoformat(), "auto_fin_decision_at": decision_at.isoformat(), - "auto_fin_news_start": start.isoformat(), - "auto_fin_previous_trade_date": previous_trade_date.isoformat(), + "auto_fin_news_start": news_start.isoformat(), + "auto_fin_etf_names": names, }, ) - self.context.response.metadata.update({"date": run_date.isoformat(), "news_downloaded": downloaded}) - self.logger.info( - f"[{self.name}] done downloaded={downloaded} cached={news_days - downloaded} " - f"previous_trade_date={previous_trade_date.isoformat()}", - ) + self.context.response.metadata.update({"date": run_date.isoformat(), "news_downloaded": len(changes)}) return self.context.response diff --git a/reme/steps/cookbook/auto_fin/history.py b/reme/steps/cookbook/auto_fin/history.py index e9da473e..68653faa 100644 --- a/reme/steps/cookbook/auto_fin/history.py +++ b/reme/steps/cookbook/auto_fin/history.py @@ -1,93 +1,200 @@ -"""Orchestrate historical research and market analysis for selected ETFs.""" +"""Prepare historical news and observed ETF returns.""" from __future__ import annotations -from ....components import R -from ....schema import ( - AutoFinEtfHistoricalResearch, - AutoFinEtfHistoryDetail, - AutoFinEtfSelection, - AutoFinSelectedEtfAnalysis, - AutoFinSelectedEvent, -) -from ._base import AutoFinStep, _write_jsonl +import json +import re +from datetime import date, datetime, time, timedelta +from pathlib import Path +from typing import Any -DEFAULT_DISPATCH_STEPS = ["auto_fin_history_search_step", "auto_fin_market_step"] +from ....components import R +from ....schema import AutoFinEtfAnalysis, AutoFinHistoricalOutput +from ._base import AutoFinStep, _write_jsonl +from .data import AutoFinDataStep, NEWS_FILENAME + +NEWS_ID_RE = re.compile(r"\b\d{14}_[0-9a-fA-F]{4}\b") +HORIZONS = (1, 2, 3, 5) @R.register("auto_fin_history_step") class AutoFinHistoryStep(AutoFinStep): - """Dispatch historical and market steps for each selected ETF.""" + """Search with ReMe, then let a tool-free Agent select comparable events.""" - def __init__(self, *args, dispatch_steps=None, **kwargs): - super().__init__( - *args, - dispatch_steps=DEFAULT_DISPATCH_STEPS if dispatch_steps is None else dispatch_steps, - **kwargs, + def _market_rows(self, code: str) -> list[dict[str, Any]]: + path = self.workspace_path / str(self.config_value("resource_dir")) / "fin" / f"{code}.jsonl" + rows = self._read_jsonl_sync(path) + return sorted(rows, key=lambda row: str(row.get("trade_date") or "")) + + @staticmethod + def _read_news( + path: Path, + cache: dict[str, list[dict[str, str]]], + ) -> list[dict[str, str]]: + """Parse a news file once per run; the same file recurs across events/ETFs.""" + key = str(path) + if key not in cache: + cache[key] = AutoFinDataStep.read_news(path) + return cache[key] + + @staticmethod + def _number(value: Any) -> float | None: + try: + number = float(value) + except (TypeError, ValueError): + return None + return number if number > 0 else None + + @classmethod + def _returns(cls, event_time: datetime, rows: list[dict[str, Any]]) -> dict[str, float | None]: + usable = [] + for row in rows: + try: + trade_day = datetime.strptime(str(row.get("trade_date")), "%Y%m%d").date() + except ValueError: + continue + if (close := cls._number(row.get("close"))) is None or ( + factor := cls._number(row.get("adj_factor")) + ) is None: + continue + usable.append((trade_day, row, close, factor)) + + event_day = event_time.date() + before_close = event_time.time() < time(15) + entry_index = entry_price = None + entered_at_close = False + for index, (trade_day, row, close, factor) in enumerate(usable): + # React before the close only when the event day is itself a trading day. + if before_close and trade_day == event_day: + entry_index, entry_price, entered_at_close = index, close * factor, True + break + # Otherwise (post-close, or any time on a non-trading day) enter at the + # open of the first trading day after the event. + if trade_day > event_day: + if (opened := cls._number(row.get("open"))) is not None: + entry_index, entry_price = index, opened * factor + break + result = {f"d{horizon}": None for horizon in HORIZONS} + if entry_index is None or entry_price is None: + return result + first_close = entry_index + 1 if entered_at_close else entry_index + for horizon in HORIZONS: + target = first_close + horizon - 1 + if target < len(usable): + _, _, close, factor = usable[target] + result[f"d{horizon}"] = close * factor / entry_price - 1 + return result + + async def _candidates( + self, + event: dict[str, str], + start: date, + end: date, + news_cache: dict[str, list[dict[str, str]]], + ) -> list[dict[str, str]]: + query = " ".join(filter(None, [event.get("title"), event.get("content"), event.get("reason")]))[:2000] + response = await self.run_job( + "memory_search", + query=query, + limit=int(self._value("historical_search_limit", 10)), + start_date=start.isoformat(), + end_date=end.isoformat(), + strict_date_filter=True, ) + if not response.success: + raise RuntimeError(f"Auto Fin memory search failed: {response.answer}") + ids_by_path: dict[str, set[str]] = {} + for result in response.metadata.get("results", []): + path = str(result.get("path") or "") + if Path(path).name == NEWS_FILENAME: + ids_by_path.setdefault(path, set()).update(NEWS_ID_RE.findall(str(result.get("text") or ""))) + candidates = [] + for relative, news_ids in ids_by_path.items(): + path = self.workspace_path / relative + for row in self._read_news(path, news_cache): + if row["news_id"] in news_ids and row["news_id"] != event["news_id"]: + candidates.append(row) + unique = {row["news_id"]: row for row in candidates} + return sorted( + unique.values(), + key=lambda row: (row["event_time"], row["news_id"]), + reverse=True, + ) + + @staticmethod + def _normalize(output: AutoFinHistoricalOutput, candidates: list[dict[str, str]], limit: int): + by_id = {row["news_id"]: row for row in candidates} + selected = [] + seen = set() + for item in output.historical_events: + news_id, reason = item.news_id.strip(), item.reason.strip() + if news_id in by_id and news_id not in seen and reason: + selected.append(item.model_copy(update={"news_id": news_id, "reason": reason})) + seen.add(news_id) + selected.sort(key=lambda item: by_id[item.news_id]["event_time"], reverse=True) + return selected[:limit] async def execute(self): assert self.context is not None - etfs = [AutoFinEtfSelection.model_validate(item) for item in self._required("auto_fin_etfs")] - history_details = [] - news_rows = await self._read_jsonl(self.workspace_path / str(self._required("auto_fin_filtered_news"))) - news_by_id = {str(row["news_id"]): row for row in news_rows} - current_keys = ( - "auto_fin_current_history", - "auto_fin_current_history_resource", - "auto_fin_current_analysis", - ) - self.logger.info(f"[{self.name}] start etfs={len(etfs)}") - for index, item in enumerate(etfs, 1): - label = f"{item.etf_code}({item.etf_name})" - self.logger.info( - f"[{self.name}] etf start index={index}/{len(etfs)} etf={label!r} events={len(item.events)}", - ) - events = [ - AutoFinSelectedEvent( - reason=event.reason, - news_id=event.news_id, - event_time=news_by_id[event.news_id]["event_time"], - event_title=str(news_by_id[event.news_id].get("title") or ""), - event_content=str( - news_by_id[event.news_id].get("content") or news_by_id[event.news_id].get("title") or "", - ), + if self.context.get("auto_fin_skipped"): + return self.context.response + news_by_id = {row["news_id"]: row for row in self._required("auto_fin_news")} + run_date = date.fromisoformat(str(self._required("auto_fin_date"))) + search_start = date.fromisoformat(str(self._required("auto_fin_news_start"))) + limit = int(self._value("historical_news_limit", 5)) + analyses = [] + call_index = 0 + news_cache: dict[str, list[dict[str, str]]] = {} + for etf in self._required("auto_fin_etfs"): + market_rows = self._market_rows(etf["etf_code"]) + events = [] + for reference in etf["events"]: + current = { + **news_by_id[reference["news_id"]], + "reason": reference["reason"], + } + candidates = await self._candidates( + current, + search_start, + run_date - timedelta(days=1), + news_cache, ) - for event in item.events - ] - for key in current_keys: - if key in self.context: - del self.context[key] - await self.dispatch_steps( - self.dispatch_step_specs, - agent_wrapper=self.agent_wrapper, - auto_fin_current_index=index, - auto_fin_current_etf=item.model_dump(mode="json"), - auto_fin_current_events=[event.model_dump(mode="json") for event in events], + call_index += 1 + output, _ = await self._reply( + "history_user", + f"auto_fin_history_{call_index:03d}", + AutoFinHistoricalOutput, + etf=json.dumps( + {"etf_code": etf["etf_code"], "etf_name": etf["etf_name"]}, + ensure_ascii=False, + ), + current_event=json.dumps(current, ensure_ascii=False), + candidates=json.dumps(candidates, ensure_ascii=False), + ) + historical = [] + candidates_by_id = {row["news_id"]: row for row in candidates} + for match in self._normalize(output, candidates, limit): + row = candidates_by_id[match.news_id] + historical.append( + { + **row, + "reason": match.reason, + "direction": match.direction, + "returns": self._returns(datetime.fromisoformat(row["event_time"]), market_rows), + }, + ) + events.append({**current, "historical_events": historical}) + analyses.append( + AutoFinEtfAnalysis.model_validate( + { + "etf_code": etf["etf_code"], + "etf_name": etf["etf_name"], + "events": events, + }, + ).model_dump(mode="json"), ) - history = AutoFinEtfHistoricalResearch.model_validate(self._required("auto_fin_current_history")) - analysis = AutoFinSelectedEtfAnalysis.model_validate(self._required("auto_fin_current_analysis")) - detail = AutoFinEtfHistoryDetail( - etf=item, - current_events=events, - historical_research=history, - market_analysis=analysis, - ) - history_details.append(detail.model_dump(mode="json")) - self.logger.info( - f"[{self.name}] etf done index={index}/{len(etfs)} etf={label!r}", - ) - for key in current_keys: - if key in self.context: - del self.context[key] - history_path = ( - self.workspace_path / "resource" / str(self._required("auto_fin_date")) / "auto_fin_history_output.jsonl" - ) - _write_jsonl(history_path, history_details) - self.context["auto_fin_history_details"] = history_details - self.context["auto_fin_history_resource"] = str(history_path) - self.context.response.metadata["analysis_count"] = len(history_details) - self.logger.info( - f"[{self.name}] done analyses={len(history_details)} history_resource={history_path}", - ) + path = self.workspace_path / str(self.config_value("resource_dir")) / str(run_date) / "auto_fin_analysis.jsonl" + _write_jsonl(path, analyses) + self.context["auto_fin_analyses"] = analyses + self.context.response.metadata["analysis_count"] = len(analyses) return self.context.response diff --git a/reme/steps/cookbook/auto_fin/history.yaml b/reme/steps/cookbook/auto_fin/history.yaml new file mode 100644 index 00000000..b8cb622b --- /dev/null +++ b/reme/steps/cookbook/auto_fin/history.yaml @@ -0,0 +1,14 @@ +history_user: | + 你是历史事件比较 Agent。上下文已完整提供,不得搜索、调用工具或补充外部信息。 + + ETF:{etf} + 当前新闻:{current_event} + ReMe 已检索出的历史候选新闻:{candidates} + + 从候选新闻中选择与当前新闻在事件类型、关键实体或价格传导机制上真正可比的历史事件。 + direction 表示历史事件对该 ETF 的影响方向与当前事件相比: + - same:两次事件对 ETF 的影响方向相同; + - opposite:影响方向相反。例如当前降息利好黄金、历史加息利空黄金,应为 opposite。 + + 只复制候选中完整的 news_id;没有看到完整 news_id 就不要输出。不要根据事后 ETF 涨跌判断方向。 + 按结构化输出契约返回 historical_events;每项包含完整 news_id、reason 和 direction。 diff --git a/reme/steps/cookbook/auto_fin/history_search.py b/reme/steps/cookbook/auto_fin/history_search.py deleted file mode 100644 index 706dba6c..00000000 --- a/reme/steps/cookbook/auto_fin/history_search.py +++ /dev/null @@ -1,339 +0,0 @@ -"""Find historical events relevant to one selected ETF.""" - -from __future__ import annotations - -import asyncio -import json -from datetime import date, datetime, time -from pathlib import Path -from typing import Any -from uuid import uuid4 - -from ....components import R -from ....schema import ( - AutoFinEtfHistoricalEvents, - AutoFinEtfHistoricalResearch, - AutoFinEtfSelection, - AutoFinHistoricalEvent, - AutoFinHistoricalEventReference, - AutoFinMarketSample, - AutoFinSelectedEvent, -) -from ...index._dedup import _ToolContextDedupMixin -from ._base import AutoFinStep, _write -from .topic import _plain_text - - -@R.register("auto_fin_history_search_step") -class AutoFinHistorySearchStep(AutoFinStep): - """Find historical events and calculate their adjusted ETF returns.""" - - @staticmethod - def _trade_date(value: Any) -> date | None: - text = str(value or "").replace("-", "") - try: - return datetime.strptime(text, "%Y%m%d").date() - except ValueError: - return None - - @staticmethod - def _positive_float(value: Any) -> float | None: - try: - number = float(value) - except (TypeError, ValueError): - return None - return number if number > 0 else None - - def _historical_source_candidates(self, source_path_value: str, news_id: str) -> list[Path]: - """Return safe declared and date-derived source candidates within the workspace.""" - workspace = self.workspace_path.resolve() - relative_path = Path(source_path_value) - candidates = [] - if not relative_path.is_absolute() and ".." not in relative_path.parts: - source_path = (workspace / relative_path).resolve() - try: - source_path.relative_to(workspace) - except ValueError: - pass - else: - if source_path.name == "auto_fin_news_data.jsonl": - candidates.append(source_path) - - news_date = news_id.partition("_")[0][:8] - try: - parsed_date = datetime.strptime(news_date, "%Y%m%d").date() - except ValueError: - parsed_date = None - if parsed_date is not None: - inferred_path = workspace / "daily" / parsed_date.isoformat() / "auto_fin_news_data.jsonl" - if inferred_path not in candidates: - candidates.append(inferred_path) - if not candidates: - raise ValueError(f"Historical source cannot be inferred safely: {source_path_value}") - return candidates - - async def _resolve_historical_event( - self, - reference: AutoFinHistoricalEventReference, - current_news_ids: set[str], - window_start: datetime, - rows_by_path: dict[Path, list[dict[str, Any]]], - ) -> AutoFinHistoricalEvent: - """Resolve one Agent-selected identity from user-owned source files.""" - workspace = self.workspace_path.resolve() - if reference.news_id in current_news_ids: - raise ValueError(f"History Agent returned a current news item: {reference.news_id}") - - candidates = self._historical_source_candidates(reference.source_path, reference.news_id) - matches: list[tuple[Path, dict[str, Any]]] = [] - for source_path in candidates: - if not source_path.is_file(): - continue - rows = rows_by_path.get(source_path) - if rows is None: - rows = await self._read_jsonl(source_path) - rows_by_path[source_path] = rows - matches.extend((source_path, row) for row in rows if str(row.get("news_id") or "") == reference.news_id) - if len(matches) != 1: - raise ValueError( - f"Historical news_id must resolve exactly once from {reference.source_path} " - f"or its ID-derived daily file: {reference.news_id}", - ) - - source_path, row = matches[0] - event_time = self._published_at(row) - if event_time is None: - raise ValueError(f"Historical news has no valid publication time: {reference.news_id}") - if event_time >= window_start: - raise ValueError(f"History Agent returned an event inside the current news window: {reference.news_id}") - event_title = str(row.get("title") or "").strip() - event_content = _plain_text(str(row.get("content") or event_title)) - if not event_title or not event_content: - raise ValueError(f"Historical news has no usable title or content: {reference.news_id}") - - return AutoFinHistoricalEvent( - reason=reference.reason, - news_id=reference.news_id, - source_path=source_path.relative_to(workspace).as_posix(), - event_time=event_time, - event_title=event_title, - event_content=event_content, - ) - - async def _resolve_historical_events( - self, - references: list[AutoFinHistoricalEventReference], - current_news_ids: set[str], - window_start: datetime, - ) -> tuple[list[AutoFinHistoricalEvent], list[str]]: - """Resolve valid references and report invalid ones without stopping the ETF.""" - rows_by_path: dict[Path, list[dict[str, Any]]] = {} - events_by_news_id: dict[str, AutoFinHistoricalEvent] = {} - limitations = [] - for reference in references: - reference = reference.model_copy( - update={ - "reason": reference.reason.strip(), - "news_id": reference.news_id.strip(), - "source_path": reference.source_path.strip(), - }, - ) - if not reference.reason or not reference.news_id: - limitation = "跳过缺少 reason 或 news_id 的历史新闻" - self.logger.warning(f"[{self.name}] {limitation}") - limitations.append(limitation) - continue - try: - event = await self._resolve_historical_event( - reference, - current_news_ids, - window_start, - rows_by_path, - ) - except (OSError, ValueError) as exc: - limitation = f"跳过无法解析的历史新闻 {reference.news_id}: {exc}" - self.logger.warning(f"[{self.name}] {limitation}") - limitations.append(limitation) - continue - events_by_news_id.setdefault(reference.news_id, event) - resolved = sorted(events_by_news_id.values(), key=lambda event: (event.event_time, event.news_id)) - return resolved, limitations - - async def _calculate_samples( - self, - etf_code: str, - events: list[AutoFinHistoricalEvent], - decision_at: datetime, - ) -> tuple[list[AutoFinMarketSample], list[str]]: - if not events: - return [], [] - start_date = min(event.event_time.date() for event in events).strftime("%Y%m%d") - end_date = decision_at.date().strftime("%Y%m%d") - daily, factors = await asyncio.gather( - self._fetch("fund_daily", ts_code=etf_code, start_date=start_date, end_date=end_date), - self._fetch("fund_adj", ts_code=etf_code, start_date=start_date, end_date=end_date), - ) - factors_by_date = { - trade_date: self._positive_float(row.get("adj_factor")) - for row in factors - if (trade_date := self._trade_date(row.get("trade_date"))) is not None - } - daily_by_date = { - trade_date: { - "trade_date": trade_date, - "open": self._positive_float(row.get("open")), - "close": self._positive_float(row.get("close")), - "adj_factor": factors_by_date.get(trade_date), - } - for row in daily - if (trade_date := self._trade_date(row.get("trade_date"))) is not None - and datetime.combine(trade_date, time(15, 0)) <= decision_at - } - rows = [daily_by_date[trade_date] for trade_date in sorted(daily_by_date)] - row_indexes = {row["trade_date"]: index for index, row in enumerate(rows)} - samples = [] - limitations = [] - for event in events: - event_time = event.event_time - event_date = event_time.date() - row_index = row_indexes.get(event_date) - price_type = None - if row_index is not None and event_time.time() < time(9, 30): - price_type = "open" - elif row_index is not None and event_time.time() < time(15, 0): - price_type = "close" - else: - row_index = next( - (index for index, row in enumerate(rows) if row["trade_date"] > event_date), - None, - ) - price_type = "open" if row_index is not None else None - - if row_index is None or price_type is None: - limitations.append(f"{event_time.isoformat()} 之后没有已完成的 ETF 日线") - samples.append( - AutoFinMarketSample( - event_time=event.event_time, - reaction_summary="事件之后没有可用的已完成日线。", - ), - ) - continue - - entry_row = rows[row_index] - raw_price = entry_row[price_type] - adj_factor = entry_row["adj_factor"] - entry_clock = time(9, 30) if price_type == "open" else time(15, 0) - entry_time = datetime.combine(entry_row["trade_date"], entry_clock) - if raw_price is None or adj_factor is None: - limitations.append(f"{entry_row['trade_date']} 缺少有效的 {price_type} 或 ETF 复权因子") - samples.append( - AutoFinMarketSample( - event_time=event.event_time, - reaction_summary="买入点缺少有效价格或复权因子。", - ), - ) - continue - - future_returns = [] - first_close_index = row_index if price_type == "open" else row_index + 1 - adjusted_entry = raw_price * adj_factor - for future_row in rows[first_close_index : first_close_index + 10]: - raw_close = future_row["close"] - close_factor = future_row["adj_factor"] - if raw_close is None or close_factor is None: - limitations.append(f"{future_row['trade_date']} 缺少有效 close 或 ETF 复权因子") - break - future_returns.append( - { - "horizon": len(future_returns) + 1, - "trade_date": future_row["trade_date"], - "raw_close": raw_close, - "adj_factor": close_factor, - "cumulative_return": raw_close * close_factor / adjusted_entry - 1, - }, - ) - if len(future_returns) < 10: - limitations.append(f"{event_time.isoformat()} 只有 {len(future_returns)} 个已完成的未来收盘点") - samples.append( - AutoFinMarketSample.model_validate( - { - "event_time": event.event_time, - "entry": { - "entry_time": entry_time, - "trade_date": entry_row["trade_date"], - "price_type": price_type, - "raw_price": raw_price, - "adj_factor": adj_factor, - }, - "future_returns": future_returns, - "reaction_summary": f"按复权日线计算了 {len(future_returns)} 个未来收盘点。", - }, - ), - ) - return samples, list(dict.fromkeys(limitations)) - - async def execute(self): - assert self.context is not None - item = AutoFinEtfSelection.model_validate(self._required("auto_fin_current_etf")) - events = [AutoFinSelectedEvent.model_validate(event) for event in self._required("auto_fin_current_events")] - index = int(self._required("auto_fin_current_index")) - window_start = datetime.fromisoformat(str(self._required("auto_fin_window_start"))) - decision_at = datetime.fromisoformat(str(self._required("auto_fin_decision_at"))) - search_events = [event.model_dump(mode="json", exclude={"news_id"}) for event in events] - label = f"{item.etf_code}({item.etf_name})" - tool_context_id = f"auto_fin_history_{index:02d}_{item.etf_code}_{uuid4().hex}" - try: - history, history_path = await self._reply( - "history_search_user", - f"auto_fin_history_{index:02d}_{item.etf_code}", - AutoFinEtfHistoricalEvents, - tool_context_id=tool_context_id, - etf_code=item.etf_code, - etf_name=item.etf_name, - events=str(search_events), - window_start=window_start.isoformat(), - workspace_root=str(self.workspace_path), - ) - finally: - if self.app_context is not None: - contexts = self.app_context.metadata.get(_ToolContextDedupMixin.TOOL_CONTEXTS_KEY) - if isinstance(contexts, dict): - contexts.pop(tool_context_id, None) - if not contexts: - self.app_context.metadata.pop(_ToolContextDedupMixin.TOOL_CONTEXTS_KEY, None) - resolved_events, resolution_limitations = await self._resolve_historical_events( - history.historical_events, - {event.news_id for event in events}, - window_start, - ) - samples, market_limitations = await self._calculate_samples( - item.etf_code, - resolved_events, - decision_at, - ) - enriched_events = [ - event.model_copy( - update={ - "market_entry": sample.entry, - "future_returns": sample.future_returns, - }, - ) - for event, sample in zip(resolved_events, samples, strict=True) - ] - enriched_history = AutoFinEtfHistoricalResearch( - etf_code=item.etf_code, - etf_name=item.etf_name, - historical_events=enriched_events, - limitations=list(dict.fromkeys([*resolution_limitations, *market_limitations])), - ) - _write( - history_path, - json.dumps(enriched_history.model_dump(mode="json"), ensure_ascii=False, indent=2) + "\n", - ) - self.context["auto_fin_current_history"] = enriched_history.model_dump(mode="json") - self.context["auto_fin_current_history_resource"] = str(history_path) - self.logger.info( - f"[{self.name}] ready etf={label!r} events={len(enriched_events)} " - f"limitations={len(enriched_history.limitations)}", - ) - return self.context.response diff --git a/reme/steps/cookbook/auto_fin/history_search.yaml b/reme/steps/cookbook/auto_fin/history_search.yaml deleted file mode 100644 index b94a13a9..00000000 --- a/reme/steps/cookbook/auto_fin/history_search.yaml +++ /dev/null @@ -1,25 +0,0 @@ -history_search_user: | - 为 {etf_code}({etf_name})寻找与当前事件相似的历史新闻。不查询行情、不计算收益或预测。 - - 当前事件:{events} - 历史截止时间:{window_start} - ReMe workspace:{workspace_root} - - 工作要求: - 1. 使用 memory_search 从事件类型、关键实体、传导机制和影响方向等角度搜索截止时间之前的 - 相似新闻;结果不足时调整角度继续搜索。必要时扫描 workspace 过去 360 天的 - `daily/YYYY-MM-DD/auto_fin_news_data.jsonl`。 - 2. 选择机制可比的历史事件,说明相似原因,并返回原始新闻的 news_id 和 source_path。 - 当前事件不属于历史事件。程序会回查、过滤、去重、排序并补充行情。 - 3. 返回以下 JSON;没有合适新闻时 historical_events 为空: - ```json - {{ - "historical_events": [ - {{ - "reason": "事件类型、关键实体、传导机制和影响方向相似", - "news_id": "20260601100000_a3f8", - "source_path": "daily/2026-06-01/auto_fin_news_data.jsonl" - }} - ] - }} - ``` diff --git a/reme/steps/cookbook/auto_fin/market.py b/reme/steps/cookbook/auto_fin/market.py deleted file mode 100644 index 8b77ff99..00000000 --- a/reme/steps/cookbook/auto_fin/market.py +++ /dev/null @@ -1,164 +0,0 @@ -"""Forecast one selected ETF from calculated historical samples.""" - -from __future__ import annotations - -import json - -from ....components import R -from ....schema import ( - AutoFinEtfHistoricalResearch, - AutoFinEtfSelection, - AutoFinMarketSelection, - AutoFinSelectedEtfAnalysis, - AutoFinSelectedEvent, -) -from ._base import AutoFinStep, _write - - -@R.register("auto_fin_market_step") -class AutoFinMarketStep(AutoFinStep): - """Classify historical event directions and calculate one ETF forecast.""" - - @staticmethod - def _normalize_selection( - selection: AutoFinMarketSelection, - history: AutoFinEtfHistoricalResearch, - ) -> tuple[AutoFinMarketSelection, bool]: - """Filter unknown, blank, and duplicate direction references.""" - known_news_ids = {event.news_id for event in history.historical_events} - seen_news_ids: set[str] = set() - - def valid_events(events): - normalized = [] - for event in events: - reason = event.reason.strip() - news_id = event.news_id.strip() - if not reason or news_id not in known_news_ids or news_id in seen_news_ids: - continue - normalized.append({"reason": reason, "news_id": news_id}) - seen_news_ids.add(news_id) - return normalized - - normalized_selection = AutoFinMarketSelection.model_validate( - { - "same_direction_events": valid_events(selection.same_direction_events), - "opposite_direction_events": valid_events(selection.opposite_direction_events), - }, - ) - changed = normalized_selection.model_dump(mode="json") != selection.model_dump(mode="json") - return normalized_selection, changed - - @staticmethod - def _calculate_analysis( - item: AutoFinEtfSelection, - history: AutoFinEtfHistoricalResearch, - selection: AutoFinMarketSelection, - ) -> AutoFinSelectedEtfAnalysis: - """Build all deterministic market fields from Agent-selected news IDs.""" - history_by_news_id = {event.news_id: event for event in history.historical_events} - selected = [ - *((match, "same", 1.0) for match in selection.same_direction_events), - *((match, "opposite", -1.0) for match in selection.opposite_direction_events), - ] - - weight = 1.0 / len(selected) if selected else 0.0 - matches = [ - { - "reason": match.reason, - "news_id": match.news_id, - "event_time": history_by_news_id[match.news_id].event_time, - "direction": direction, - "weight": weight, - } - for match, direction, _ in selected - ] - - returns = [] - has_missing_horizon = False - has_direction_conflict = False - for horizon in range(1, 11): - available = [] - for match, _, direction_coefficient in selected: - event = history_by_news_id[match.news_id] - point = next((point for point in event.future_returns if point.horizon == horizon), None) - if point is not None: - available.append(direction_coefficient * point.cumulative_return) - if not available: - has_missing_horizon = True - expected_return = None - else: - expected_return = sum(available) / len(available) - has_direction_conflict |= any(value > 0 for value in available) and any( - value < 0 for value in available - ) - returns.append({"horizon": horizon, "expected_return": expected_return}) - - positive_returns = [point for point in returns if (point["expected_return"] or 0) > 0] - suggested_holding_days = ( - max(positive_returns, key=lambda point: (point["expected_return"], -point["horizon"]))["horizon"] - if positive_returns - else None - ) - limitations = list(history.limitations) - if not selected: - limitations.append("没有匹配的历史事件") - elif len(selected) < 2: - limitations.append("相似历史样本少于 2 个") - if has_missing_horizon: - limitations.append("部分持有期缺少可用历史收益") - if has_direction_conflict: - limitations.append("相似历史样本的收益方向存在分歧") - if selected and suggested_holding_days is None: - limitations.append("加权预期收益没有正值") - - return AutoFinSelectedEtfAnalysis.model_validate( - { - "etf_code": item.etf_code, - "etf_name": item.etf_name, - "matched_historical_events": matches, - "forecast": { - "returns": returns, - "suggested_holding_days": suggested_holding_days, - }, - "limitations": list(dict.fromkeys(limitations)), - }, - ) - - async def execute(self): - assert self.context is not None - item = AutoFinEtfSelection.model_validate(self._required("auto_fin_current_etf")) - events = [AutoFinSelectedEvent.model_validate(event) for event in self._required("auto_fin_current_events")] - history = AutoFinEtfHistoricalResearch.model_validate(self._required("auto_fin_current_history")) - index = int(self._required("auto_fin_current_index")) - event_lines = "\n".join( - f"- [{event.event_time.isoformat()}] {event.event_title or event.reason}: {event.event_content}" - for event in events - ) - resource_name = f"auto_fin_market_{index:02d}_{item.etf_code}" - if history.historical_events: - selection, selection_path = await self._reply( - "market_user", - resource_name, - AutoFinMarketSelection, - etf_code=item.etf_code, - etf_name=item.etf_name, - events=event_lines, - history_path=str(self._required("auto_fin_current_history_resource")), - decision_at=str(self._required("auto_fin_decision_at")), - ) - else: - selection = AutoFinMarketSelection() - selection_path = ( - self.workspace_path / "resource" / str(self._required("auto_fin_date")) / f"{resource_name}_output.json" - ) - self.logger.warning(f"[{self.name}] skip direction Agent for {item.etf_code}: no valid history") - selection, normalized = self._normalize_selection(selection, history) - if normalized: - self.logger.info(f"[{self.name}] normalized historical direction selections") - analysis = self._calculate_analysis(item, history, selection) - _write( - selection_path, - json.dumps(analysis.model_dump(mode="json"), ensure_ascii=False, indent=2) + "\n", - ) - self.context["auto_fin_current_analysis"] = analysis.model_dump(mode="json") - return self.context.response diff --git a/reme/steps/cookbook/auto_fin/market.yaml b/reme/steps/cookbook/auto_fin/market.yaml deleted file mode 100644 index 56273c89..00000000 --- a/reme/steps/cookbook/auto_fin/market.yaml +++ /dev/null @@ -1,34 +0,0 @@ -market_user: | - 筛选与当前事件机制可比的历史事件,并判断其影响方向与当前事件相同还是相反。 - 不计算收益或生成预测。 - - ETF:{etf_code}({etf_name}) - 分析截止时间:{decision_at} - 当前事件: - {events} - 历史事件文件:{history_path} - - 工作要求: - 1. 根据历史事件的时间、标题、正文和 reason,综合事件类型、关键实体、传导机制和影响方向: - - 机制可比且影响方向相同,放入 same_direction_events; - - 机制可比但影响方向相反,放入 opposite_direction_events; - - 机制不可比则不选择。 - 不要依据 market_entry 或 future_returns 选择事件,避免使用事后行情。 - 2. 返回 news_id 和判断理由。程序会过滤、去重并完成计算。 - 3. 返回以下 JSON: - ```json - {{ - "same_direction_events": [ - {{ - "reason": "事件类型、关键实体、传导机制和影响方向相似", - "news_id": "20260601100000_a3f8" - }} - ], - "opposite_direction_events": [ - {{ - "reason": "传导机制相似,但影响方向相反", - "news_id": "20260501100000_b4c9" - }} - ] - }} - ``` diff --git a/reme/steps/cookbook/auto_fin/merge.py b/reme/steps/cookbook/auto_fin/merge.py index 21c7c3ac..33826947 100644 --- a/reme/steps/cookbook/auto_fin/merge.py +++ b/reme/steps/cookbook/auto_fin/merge.py @@ -1,97 +1,93 @@ -"""Merge all selected ETF analyses into the final report.""" +"""Generate and save the final Auto Fin recommendation.""" from __future__ import annotations import json import re +from datetime import date +from pathlib import Path from types import SimpleNamespace from ....components import R -from ....schema import AutoFinEtfHistoryDetail, AutoFinReportOutput +from ....schema import AutoFinReportOutput from ...file_io import refresh_day_index -from ._base import AutoFinStep, _write, _write_jsonl +from ._base import AutoFinStep, _write @R.register("auto_fin_merge_step") class AutoFinMergeStep(AutoFinStep): - """Ask a fresh Agent for the final Markdown and persist it directly.""" + """Call the final tool-free Agent with all evidence already prepared.""" + + def _report_path(self, run_date: date) -> Path: + return self.workspace_path / str(self.config_value("daily_dir")) / str(run_date) / "auto_fin.md" + + def _previous_report(self, run_date: date) -> str: + """Return the most recent report from a *prior* day (yesterday's, typically).""" + daily = self.workspace_path / str(self.config_value("daily_dir")) + candidates = [] + for path in daily.glob("*/auto_fin.md"): + try: + day = date.fromisoformat(path.parent.name) + except ValueError: + continue + if day < run_date: + candidates.append((day, path)) + return max(candidates)[1].read_text(encoding="utf-8") if candidates else "无历史推荐。" + + def _current_report(self, run_date: date) -> str: + """Return today's existing report so intra-day reruns refine it, not replace it.""" + path = self._report_path(run_date) + if path.is_file(): + return path.read_text(encoding="utf-8") + return "今日暂无更早时段的推荐,本次为当日首次生成。" @staticmethod - def _normalize_report(output: AutoFinReportOutput) -> AutoFinReportOutput: - """Normalize cosmetic report fields and provide safe empty fallbacks.""" + def _normalize(output: AutoFinReportOutput) -> AutoFinReportOutput: title = re.sub(r"^#+\s*", "", output.title.strip()) or "Auto Fin ETF 结论" + description = output.description.strip() or "基于当前事件与相似历史表现的 ETF 观察。" body = output.body.strip() or "## 结论\n\n暂无可用结论。" - first_line, separator, remainder = body.partition("\n") - if first_line.lstrip().startswith("# "): - body = remainder.lstrip() if separator else "## 结论\n\n暂无可用结论。" - return AutoFinReportOutput(title=title, body=body) - - @staticmethod - def _calculation_results(history_details: list[AutoFinEtfHistoryDetail]) -> list[dict]: - """Return the program-calculated forecast for every analyzed ETF.""" - results = [] - for item in history_details: - holding_days = item.market_analysis.forecast.suggested_holding_days - results.append( - { - "etf_code": item.etf.etf_code, - "etf_name": item.etf.etf_name, - "suggested_holding_days": holding_days, - "returns": [point.model_dump(mode="json") for point in item.market_analysis.forecast.returns], - }, - ) - return results + if body.startswith("# "): + body = body.partition("\n")[2].lstrip() or "## 结论\n\n暂无可用结论。" + return AutoFinReportOutput(title=title, description=description, body=body) async def execute(self): assert self.context is not None - etfs = list(self._required("auto_fin_etfs")) - history_details = [ - AutoFinEtfHistoryDetail.model_validate(item) for item in self._required("auto_fin_history_details") - ] - selected = [item.etf.model_dump(mode="json") for item in history_details] - if selected != etfs: - raise ValueError("Auto Fin merge history details must match the selected ETFs") - analyses = [item.market_analysis.model_dump(mode="json") for item in history_details] - calculation_results = self._calculation_results(history_details) - self.logger.info( - f"[{self.name}] start etfs={len(etfs)}", - ) + if self.context.get("auto_fin_skipped"): + self.context.response.answer = str(self.context.get("auto_fin_skip_reason") or "Auto Fin 已跳过。") + return self.context.response + run_date = date.fromisoformat(str(self._required("auto_fin_date"))) output, output_path = await self._reply( "merge_user", "auto_fin_merge", AutoFinReportOutput, decision_at=str(self._required("auto_fin_decision_at")), - window_start=str(self._required("auto_fin_window_start")), - etfs_path=str(self._required("auto_fin_etfs_resource")), - history_path=str(self._required("auto_fin_history_resource")), - calculation_results=json.dumps(calculation_results, ensure_ascii=False), + etfs=json.dumps( + [ + {"etf_code": code, "etf_name": name} + for code, name in dict(self._required("auto_fin_etf_names")).items() + ], + ensure_ascii=False, + ), + analyses=json.dumps(self._required("auto_fin_analyses"), ensure_ascii=False), + previous_report=self._previous_report(run_date), + current_report=self._current_report(run_date), ) - normalized_output = self._normalize_report(output) - if normalized_output != output: - _write( - output_path, - json.dumps(normalized_output.model_dump(mode="json"), ensure_ascii=False, separators=(",", ":")) + "\n", - ) - output = normalized_output - markdown = f"# {output.title}\n\n{output.body}\n\n" - markdown += "> 仅为事件研究和持有时间参考,不构成投资建议,不会执行交易。\n" - day_dir = self.workspace_path / str(self.config_value("daily_dir")) / str(self._required("auto_fin_date")) - report_path = day_dir / "auto_fin.md" - _write_jsonl(day_dir / "auto_fin_analysis.jsonl", analyses) - _write(report_path, markdown) - relative = report_path.relative_to(self.workspace_path).as_posix() + output = self._normalize(output) + self._write_output(output_path, output) + markdown = ( + f"# {output.title}\n\n> {output.description}\n\n{output.body}\n\n" + "> 仅为事件研究和持有时间参考,不构成投资建议,不会执行交易。\n" + ) + report = self._report_path(run_date) + _write(report, markdown) await refresh_day_index( SimpleNamespace(workspace_path=self.workspace_path), - str(self._required("auto_fin_date")), + str(run_date), str(self.config_value("daily_dir")), ) + relative = report.relative_to(self.workspace_path).as_posix() self.context["markdown_path"] = relative self.context["auto_fin_digest_path"] = relative self.context.response.answer = output.body - self.context.response.metadata.update( - {"markdown_path": relative, "digest_path": relative, "etf_count": len(history_details)}, - ) - self.logger.info( - f"[{self.name}] done path={relative} etfs={self.context.response.metadata['etf_count']}", - ) + self.context.response.metadata.update({"markdown_path": relative, "digest_path": relative}) return self.context.response diff --git a/reme/steps/cookbook/auto_fin/merge.yaml b/reme/steps/cookbook/auto_fin/merge.yaml index b4c156f6..ca05a826 100644 --- a/reme/steps/cookbook/auto_fin/merge.yaml +++ b/reme/steps/cookbook/auto_fin/merge.yaml @@ -1,23 +1,16 @@ merge_user: | - 把已经完成的结构化分析写成中文 Markdown 结论,不重新搜索新闻、下载行情或修改数值。 + 你是最终 ETF 建议 Agent。上下文已完整提供,不得搜索、调用工具、修改历史收益或虚构数据。 - 分析截止时间:{decision_at} - 新闻窗口:({window_start}, {decision_at}] - ETF 与当前事件:{etfs_path} - 历史事件、行情样本及分析:{history_path} - 程序计算结果:{calculation_results} + 分析时间:{decision_at} + 固定 ETF code/name:{etfs} + 当前新闻、相似历史新闻、same/opposite 方向及复权 D1/D2/D3/D5 表现:{analyses} + 最近一份历史推荐(往日): + {previous_report} + 今天早些时段的推荐(如有,请在其基础上结合截至当前的最新证据继续修订,而非推倒重来; + 保留仍然成立的判断,只更新已变化或新出现的部分): + {current_report} - 工作要求: - 1. 根据当前事件内容和传导关系判断影响方向。只推荐同时满足以下条件的 ETF: - - 经你判断,当前事件对该 ETF 所代表的资产、行业或主题影响明确为正向; - - 程序给出了最佳持有天数,且该天数对应的 expected_return 大于 0。 - 不使用计算结果反推事件方向;不满足条件时给出观望结论。 - 2. 推荐结论包含 ETF code、name、最佳持有天数和对应预估收益;负向或无正收益的情况可以合并 - 简述。 - 3. 返回以下 JSON: - ```json - {{ - "title": "Auto Fin ETF 结论", - "body": "## 结论\n\n推荐 518880.SH(黄金ETF),参考持有 3 个交易日,当前加权预估收益 +1.2%;......。\n\n负向提示:相关能源 ETF 事件影响偏负,不推荐。" - }} - ``` + 综合事件传导方向、历史样本数量、不同期限表现、样本一致性、往日推荐以及今天早些的推荐,自行判断 + 是否推荐 ETF 及参考持有期限。direction=opposite 的历史表现应反向理解。没有历史样本、样本冲突或 + 证据不足时可以明确观望,不要求机械打分,也不要为了凑数推荐。 + 按结构化输出契约返回 title、description 和完整中文 Markdown body。 diff --git a/reme/steps/cookbook/auto_fin/topic.py b/reme/steps/cookbook/auto_fin/topic.py index bd59c591..5909ce82 100644 --- a/reme/steps/cookbook/auto_fin/topic.py +++ b/reme/steps/cookbook/auto_fin/topic.py @@ -1,270 +1,67 @@ -"""Build the current Auto Fin topic timelines.""" +"""Select configured ETFs that are directly related to today's news.""" from __future__ import annotations -import math -import re -from datetime import date, datetime, time -from html.parser import HTMLParser +import json from typing import Any from ....components import R from ....schema import AutoFinEtfsOutput -from ._base import AutoFinStep, _news_id, _write_jsonl - -NEWS_TITLE_MAX_CHARS = 200 -NEWS_CONTENT_MAX_CHARS = 1200 -NEWS_TOTAL_CONTENT_MAX_CHARS = 60_000 -ETF_CANDIDATE_LIMIT = 150 -ETF_OUTPUT_LIMIT = 20 - - -class _NewsTextExtractor(HTMLParser): - """Extract visible text without retaining markup, links, or image URLs.""" - - def __init__(self) -> None: - super().__init__(convert_charrefs=True) - self.parts: list[str] = [] - self.ignored_depth = 0 - - def handle_starttag(self, tag: str, _attrs) -> None: - if tag in {"script", "style"}: - self.ignored_depth += 1 - elif tag in {"br", "div", "h1", "h2", "h3", "h4", "li", "p"}: - self.parts.append(" ") - - def handle_endtag(self, tag: str) -> None: - if tag in {"script", "style"} and self.ignored_depth: - self.ignored_depth -= 1 - elif tag in {"div", "h1", "h2", "h3", "h4", "li", "p"}: - self.parts.append(" ") - - def handle_data(self, data: str) -> None: - if not self.ignored_depth: - self.parts.append(data) - - -def _plain_text(value: str) -> str: - parser = _NewsTextExtractor() - parser.feed(value) - parser.close() - return " ".join("".join(parser.parts).split()) - - -def _normalized_key(value: Any) -> str: - """Normalize display labels used only for deterministic deduplication.""" - return re.sub(r"\s+", "", str(value or "")).casefold() +from ._base import AutoFinStep @R.register("auto_fin_topic_step") class AutoFinTopicStep(AutoFinStep): - """Match current news to a small set of liquid, representative ETFs.""" - - async def _current_news(self, start: datetime, end: datetime) -> list[dict]: - news = {} - title_limit = max(0, int(self._value("news_title_max_chars", NEWS_TITLE_MAX_CHARS))) - content_limit = max(0, int(self._value("news_content_max_chars", NEWS_CONTENT_MAX_CHARS))) - total_content_limit = max( - 0, - int(self._value("news_total_content_max_chars", NEWS_TOTAL_CONTENT_MAX_CHARS)), - ) - first_day = date.fromisoformat(str(self._required("auto_fin_news_start"))) - last_day = date.fromisoformat(str(self._required("auto_fin_date"))) - for day in self._days(first_day, last_day): - for row in await self._read_jsonl(self._news_path(day)): - published_at = self._published_at(row) - if published_at is None or not start < published_at <= end: - continue - news_id = str(row.get("news_id") or _news_id(row, published_at)) - news.setdefault( - news_id, - { - "news_id": news_id, - "event_time": published_at.isoformat(), - "title": str(row.get("title") or "").strip()[:title_limit], - "content": _plain_text(str(row.get("content") or "")), - }, - ) - rows = sorted(news.values(), key=lambda row: (row["event_time"], row["news_id"])) - per_news_limit = min(content_limit, total_content_limit // len(rows)) if rows else 0 - for row in rows: - row["content"] = row["content"][:per_news_limit] - return rows - - async def _filtered_etfs(self, trade_date: date) -> list[dict[str, str]]: - basics = await self._fetch( - "etf_basic", - list_status="L", - fields="ts_code,csname,extname,cname,index_code,index_name,list_status", - ) - daily = await self._fetch( - "fund_daily", - trade_date=trade_date.strftime("%Y%m%d"), - fields="ts_code,trade_date,amount", - ) - - basic_by_code: dict[str, dict[str, Any]] = {} - for row in basics: - code = str(row.get("ts_code") or row.get("code") or "").strip().upper() - name = str( - row.get("csname") or row.get("name") or row.get("extname") or row.get("cname") or "", - ).strip() - if code and name and str(row.get("list_status") or "L").upper() == "L": - basic_by_code.setdefault(code, {**row, "code": code, "name": name}) - - amount_by_code: dict[str, float] = {} - for row in daily: - code = str(row.get("ts_code") or row.get("code") or "").strip().upper() - try: - amount = float(row.get("amount")) - except (TypeError, ValueError): - continue - if code in basic_by_code and math.isfinite(amount) and amount >= 0: - amount_by_code[code] = max(amount, amount_by_code.get(code, -math.inf)) - - ranked = sorted( - ({**basic_by_code[code], "amount": amount} for code, amount in amount_by_code.items()), - key=lambda row: (-row["amount"], row["code"]), - ) - selected: list[dict[str, str]] = [] - seen_names: set[str] = set() - seen_indexes: set[str] = set() - limit = max(0, int(self._value("etf_candidate_limit", ETF_CANDIDATE_LIMIT))) - if not limit: - return selected - for row in ranked: - name_key = _normalized_key(row["name"]) - index_keys = { - key - for key in ( - _normalized_key(row.get("index_code")), - _normalized_key(row.get("index_name")), - ) - if key - } - if name_key in seen_names or index_keys & seen_indexes: - continue - selected.append({"code": row["code"], "name": row["name"]}) - seen_names.add(name_key) - seen_indexes.update(index_keys) - if len(selected) >= limit: - break - return selected + """Call the first tool-free Agent with complete current-news context.""" @staticmethod - def _repair_news_ids( + def _normalize( output: AutoFinEtfsOutput, - news: list[dict[str, Any]], - ) -> tuple[AutoFinEtfsOutput, dict[str, str]]: - """Repair a mistyped timestamp only when the content hash is unambiguous.""" - news_ids = {str(row["news_id"]) for row in news} - ids_by_suffix: dict[str, list[str]] = {} - for news_id in news_ids: - _, separator, suffix = news_id.rpartition("_") - if separator and suffix: - ids_by_suffix.setdefault(suffix, []).append(news_id) - - data = output.model_dump(mode="json") - repairs: dict[str, str] = {} - for item in data["etfs"]: - for event in item["events"]: - event["reason"] = event["reason"].strip() - news_id = event["news_id"].strip() - event["news_id"] = news_id - if news_id in news_ids: - continue - _, separator, suffix = news_id.rpartition("_") - candidates = ids_by_suffix.get(suffix, []) if separator else [] - if len(candidates) == 1: - event["news_id"] = candidates[0] - repairs[news_id] = candidates[0] - return AutoFinEtfsOutput.model_validate(data), repairs - - @staticmethod - def _normalize_selection( - output: AutoFinEtfsOutput, - news: list[dict[str, Any]], - etfs: list[dict[str, str]], - ) -> tuple[AutoFinEtfsOutput, bool]: - """Canonicalize, filter, deduplicate, sort, and limit Agent selections.""" - news_order = {str(row["news_id"]): index for index, row in enumerate(news)} - news_ids = set(news_order) - candidates = {str(row["code"]).strip().upper(): str(row["name"]).strip() for row in etfs} + news: list[dict[str, str]], + names: dict[str, str], + limit: int, + ): + news_ids = {row["news_id"] for row in news} selected: dict[str, dict[str, Any]] = {} for item in output.etfs: code = item.etf_code.strip().upper() - name = candidates.get(code) - if name is None: + if code not in names: continue - normalized = selected.setdefault( - code, - {"etf_code": code, "etf_name": name, "events": [], "seen_news_ids": set()}, - ) + target = selected.setdefault(code, {"etf_code": code, "etf_name": names[code], "events": []}) + seen = {event["news_id"] for event in target["events"]} for event in item.events: - if not event.reason or event.news_id not in news_ids or event.news_id in normalized["seen_news_ids"]: - continue - normalized["events"].append(event.model_dump(mode="json")) - normalized["seen_news_ids"].add(event.news_id) - - normalized_items = [] - for item in selected.values(): - events = sorted(item["events"], key=lambda event: news_order[event["news_id"]]) - if events: - normalized_items.append( - { - "etf_code": item["etf_code"], - "etf_name": item["etf_name"], - "events": events, - }, - ) - if len(normalized_items) >= ETF_OUTPUT_LIMIT: - break - normalized_output = AutoFinEtfsOutput.model_validate({"etfs": normalized_items}) - changed = normalized_output.model_dump(mode="json") != output.model_dump(mode="json") - return normalized_output, changed + news_id, reason = event.news_id.strip(), event.reason.strip() + if news_id in news_ids and news_id not in seen and reason and len(target["events"]) < limit: + target["events"].append({"news_id": news_id, "reason": reason}) + seen.add(news_id) + return AutoFinEtfsOutput.model_validate({"etfs": [item for item in selected.values() if item["events"]]}) async def execute(self): assert self.context is not None - decision_at = datetime.fromisoformat(str(self._required("auto_fin_decision_at"))) - previous = date.fromisoformat(str(self._required("auto_fin_previous_trade_date"))) - window_start = datetime.combine(previous, time(15)) - news = await self._current_news(window_start, decision_at) - self.logger.info( - f"[{self.name}] start window=({window_start.isoformat()},{decision_at.isoformat()}] news={len(news)} " - f"content_chars={sum(len(row['content']) for row in news)}", - ) - resource_dir = self.workspace_path / "resource" / decision_at.date().isoformat() - news_path = resource_dir / "filtered_news.jsonl" - etf_path = resource_dir / "filtered_etf.jsonl" - etfs = await self._filtered_etfs(previous) - _write_jsonl(news_path, news) - _write_jsonl(etf_path, etfs) + if self.context.get("auto_fin_skipped"): + return self.context.response + from .data import AutoFinDataStep # Avoid a module import cycle. + + day = str(self._required("auto_fin_date")) + news_path = self.workspace_path / str(self.config_value("daily_dir")) / day / "auto_fin_news.md" + news = AutoFinDataStep.read_news(news_path) + names = dict(self._required("auto_fin_etf_names")) output, output_path = await self._reply( "topic_user", "auto_fin_topic", AutoFinEtfsOutput, - output_suffix=".jsonl", - jsonl_field="etfs", - window_start=window_start.isoformat(), - decision_at=decision_at.isoformat(), - filtered_news_path=str(news_path), - filtered_etf_path=str(etf_path), + news=json.dumps(news, ensure_ascii=False), + etfs=json.dumps( + [{"etf_code": code, "etf_name": name} for code, name in names.items()], + ensure_ascii=False, + ), ) - output, repairs = self._repair_news_ids(output, news) - output, normalized = self._normalize_selection(output, news, etfs) - if repairs: - self.logger.warning(f"[{self.name}] repaired mistyped news IDs: {repairs}") - if normalized: - self.logger.info(f"[{self.name}] normalized ETF selections") - if repairs or normalized: - _write_jsonl(output_path, output.model_dump(mode="json")["etfs"]) - self.context["auto_fin_window_start"] = window_start.isoformat() - self.context["auto_fin_etfs"] = output.model_dump(mode="json")["etfs"] - self.context["auto_fin_etfs_resource"] = str(output_path) - self.context["auto_fin_filtered_news"] = str(news_path) - self.context.response.metadata.update({"news_count": len(news), "etf_count": len(output.etfs)}) - self.logger.info( - f"[{self.name}] done etfs={len(output.etfs)} " - f"events={sum(len(item.events) for item in output.etfs)} news_path={news_path} etf_path={etf_path}", + normalized = self._normalize(output, news, names, int(self._value("current_news_limit_per_etf", 10))) + if normalized != output: + self._write_output(output_path, normalized) + self.context["auto_fin_news"] = news + self.context["auto_fin_etfs"] = normalized.model_dump(mode="json")["etfs"] + self.context.response.metadata.update( + {"news_count": len(news), "etf_count": len(normalized.etfs)}, ) return self.context.response diff --git a/reme/steps/cookbook/auto_fin/topic.yaml b/reme/steps/cookbook/auto_fin/topic.yaml index 7f1bbda0..ea1d540e 100644 --- a/reme/steps/cookbook/auto_fin/topic.yaml +++ b/reme/steps/cookbook/auto_fin/topic.yaml @@ -1,32 +1,10 @@ topic_user: | - 你只负责筛选与当前新闻直接相关的代表性 ETF,并说明每条相关新闻与 ETF 的关系。 - 不搜索或补充新闻、不下载数据、不计算收益、不提供预测或投资建议。 + 你是 ETF 新闻关联分析 Agent。上下文已完整提供,不得搜索、调用工具或补充外部信息。 - 输入文件: - - 新闻(时间窗口为 ({window_start}, {decision_at}]):{filtered_news_path} - - 候选 ETF:{filtered_etf_path} + 固定 ETF:{etfs} + 今日 00:00 至当前时间的财联社新闻:{news} - 输出要求: - 1. 从候选 ETF 中选择与新闻有明确传导关系的代表性 ETF,不要为了凑数纳入弱相关 ETF。 - 2. 为每只 ETF 返回相关新闻的 news_id,并用 reason 简洁说明直接传导关系。 - 3. 返回以下 JSON: - ```json - {{ - "etfs": [ - {{ - "etf_code": "518880.SH", - "etf_name": "华安易富黄金ETF", - "events": [ - {{ - "reason": "避险需求直接影响黄金价格", - "news_id": "20260724070000_4d19" - }}, - {{ - "reason": "美元变化影响黄金计价", - "news_id": "20260724072145_6fa5" - }} - ] - }} - ] - }} - ``` + 只保留与 ETF 所代表资产、行业或主题真实且直接相关的新闻。弱相关、仅提及公司名称、无法说明 + 明确传导关系的新闻不要输出;没有相关新闻的 ETF 不要输出。etf_code、etf_name 和 news_id 必须逐字 + 复制输入,不得猜测或缩写。 + 按结构化输出契约返回 ETF 列表,以及每只 ETF 对应的 news_id 和 reason。 diff --git a/reme/steps/cookbook/daily_paper/_common.py b/reme/steps/cookbook/daily_paper/_common.py index f475cc2c..7d8dcd33 100644 --- a/reme/steps/cookbook/daily_paper/_common.py +++ b/reme/steps/cookbook/daily_paper/_common.py @@ -1,7 +1,9 @@ """Shared state and file helpers for daily-paper steps.""" +import datetime as dt import os import re +from collections.abc import Iterator from pathlib import Path from typing import Any, TypeVar from uuid import uuid4 @@ -10,13 +12,16 @@ import aiofiles import frontmatter from pydantic import BaseModel -from ....components.outbound_proxy import BaseOutboundProxy -from ....enumeration import ComponentEnum -from ...base_step import BaseStep, Ref -from ...file_io import get_path_lock +from ...base_step import BaseStep +from ...file_io import get_path_lock, validate_filename_component +# Number of papers selected, analyzed, and digested each run. Shared across steps. +PAPER_COUNT = 3 _STATE_PREFIX = "daily_paper_" _FRONTMATTER_PATTERN = re.compile(r"^---\s*\n.*?\n---\s*\n", re.DOTALL) +_MARKDOWN_HEADING_PATTERN = re.compile(r"^#+\s*") +_UNSAFE_FILENAME_CHARS = re.compile(r'[<>:"/\\|?*\x00-\x1f]') +_CHINESE_PATTERN = re.compile(r"[\u3400-\u4dbf\u4e00-\u9fff]") _OutputT = TypeVar("_OutputT", bound=BaseModel) @@ -31,6 +36,56 @@ def strip_frontmatter(body: str) -> str: return _FRONTMATTER_PATTERN.sub("", body.strip(), count=1).strip() +def normalize_chinese_title(raw: str, fallback: str) -> str: + """Return one safe Chinese title that can also be used as the filename stem.""" + title = _MARKDOWN_HEADING_PATTERN.sub("", str(raw or "").strip()) + if title.lower().endswith(".md"): + title = title[:-3] + title = _UNSAFE_FILENAME_CHARS.sub("-", title) + title = re.sub(r"\s+", " ", title).strip(" .-") + if not title or not _CHINESE_PATTERN.search(title): + title = fallback + title = _UNSAFE_FILENAME_CHARS.sub("-", title).strip(" .-") + if error := validate_filename_component(title, kind="title"): + raise ValueError(f"Unable to produce a safe daily-paper title from {raw!r}: {error}") + return title + + +def utc_now_iso() -> str: + """Return the current UTC time as an ISO-8601 string for note metadata.""" + return dt.datetime.now(dt.timezone.utc).isoformat() + + +def iter_note_metadata(day_dir: Path) -> Iterator[tuple[Path, dict[str, Any]]]: + """Yield ``(path, frontmatter metadata)`` for each readable Markdown note in a day.""" + if not day_dir.is_dir(): + return + for path in sorted(day_dir.glob("*.md")): + try: + yield path, frontmatter.load(path).metadata + except (OSError, UnicodeError, ValueError): + continue + + +def resolve_unique_note_path( + day_dir: Path, + title: str, + *, + taken: set[str], + taken_suffix: str, + disk_suffix: str, + existing: Path | None, +) -> tuple[str, Path]: + """Disambiguate a note title against already-used titles and on-disk collisions.""" + if title in taken: + title = f"{title}{taken_suffix}" + path = day_dir / f"{title}.md" + if path.exists() and path != existing: + title = f"{title}{disk_suffix}" + path = day_dir / f"{title}.md" + return title, path + + async def write_atomic(path: Path, content: str | bytes) -> None: """Write through a sibling temporary file under the repository path lock.""" path.parent.mkdir(parents=True, exist_ok=True) @@ -56,12 +111,6 @@ async def write_markdown(path: Path, body: str, metadata: dict[str, Any]) -> Non class DailyPaperStep(BaseStep): """Shared helpers for steps in one daily-paper RuntimeContext.""" - outbound_proxy: BaseOutboundProxy | None = Ref( - BaseOutboundProxy, - ComponentEnum.OUTBOUND_PROXY, - optional=True, - ) - def _skip(self) -> bool: assert self.context is not None return bool(self.context.get(f"{_STATE_PREFIX}skip", False)) diff --git a/reme/steps/cookbook/daily_paper/analyze.py b/reme/steps/cookbook/daily_paper/analyze.py index bb4195ee..f15b763b 100644 --- a/reme/steps/cookbook/daily_paper/analyze.py +++ b/reme/steps/cookbook/daily_paper/analyze.py @@ -1,26 +1,41 @@ """Download and analyze selected daily-paper PDFs.""" import asyncio -import datetime as dt import json from pathlib import Path from ....components import R -from ....schema import PaperInfo, PaperNoteOutput, PaperSelection, SelectedPaper +from ....schema import AnalyzedPaper, DailyPaperMarkdownOutput, PaperInfo, PaperPick from ....utils.arxiv import ArxivPdfClient -from ._common import DailyPaperStep, strip_frontmatter, structured_output, write_markdown +from ._common import ( + PAPER_COUNT, + DailyPaperStep, + iter_note_metadata, + normalize_chinese_title, + resolve_unique_note_path, + strip_frontmatter, + structured_output, + utc_now_iso, + write_markdown, +) @R.register("daily_paper_analyze_step") class DailyPaperAnalyzeStep(DailyPaperStep): - """Download each selected PDF and use Claude Code for detailed reading.""" + """Download and analyze the three papers selected for the daily brief.""" @staticmethod - def _extract_pdf_text_sync(path: Path, max_pages: int, max_chars: int) -> tuple[str, int, bool]: + def _extract_pdf_text_sync( + path: Path, + max_pages: int, + max_chars: int, + ) -> tuple[str, int, bool]: try: from pypdf import PdfReader except ImportError as exc: # pragma: no cover - dependency error has an explicit message - raise RuntimeError("pypdf is required for the daily-paper workflow") from exc + raise RuntimeError( + "pypdf is required for the daily-paper workflow", + ) from exc reader = PdfReader(str(path)) chunks: list[str] = [] @@ -41,33 +56,43 @@ class DailyPaperAnalyzeStep(DailyPaperStep): raise ValueError(f"No extractable text found in PDF: {path.name}") return content, len(reader.pages), truncated + @staticmethod + def _find_existing_note(day_dir: Path, arxiv_id: str) -> Path | None: + """Find a prior generated note independently of its title filename.""" + for path, metadata in iter_note_metadata(day_dir): + if metadata.get("arxiv_id") == arxiv_id and ( + metadata.get("kind") == "daily-paper-analysis" or path.name == f"paper-{arxiv_id}.md" + ): + return path + return None + async def _analyze_one( self, downloader: ArxivPdfClient, paper: PaperInfo, - selected: SelectedPaper, - ) -> tuple[str, str]: + selected: PaperPick, + used_titles: set[str], + ) -> AnalyzedPaper: if self.agent_wrapper is None: - raise RuntimeError("Claude Code agent_wrapper is required for paper analysis") + raise RuntimeError("An agent_wrapper is required for paper analysis") day = self._run_day() daily_dir, resource_dir = ( str(self.config_value("daily_dir")).strip("/"), str(self.config_value("resource_dir")).strip("/"), ) - pdf_rel, note_rel = ( - f"{resource_dir}/papers/{paper.arxiv_id}.pdf", - f"{daily_dir}/{day}/paper-{paper.arxiv_id}.md", - ) - pdf_path, note_path = self.workspace_path / pdf_rel, self.workspace_path / note_rel + pdf_rel = f"{resource_dir}/papers/{paper.arxiv_id}.pdf" + pdf_path = self.workspace_path / pdf_rel self.logger.info(f"[{self.name}] paper start arxiv_id={paper.arxiv_id}") await downloader.download(paper.arxiv_id, pdf_path) - self.logger.info(f"[{self.name}] pdf ready arxiv_id={paper.arxiv_id} path={pdf_rel}") + self.logger.info( + f"[{self.name}] pdf ready arxiv_id={paper.arxiv_id} path={pdf_rel}", + ) pdf_text, page_count, truncated = await asyncio.to_thread( self._extract_pdf_text_sync, pdf_path, - int(self._value("max_pdf_pages", 80)), - int(self._value("max_pdf_chars", 240_000)), + int(self._value("max_pdf_pages", 20)), + int(self._value("max_pdf_chars", 300_000)), ) self.logger.info( f"[{self.name}] pdf extracted arxiv_id={paper.arxiv_id} pages={page_count} " @@ -78,27 +103,42 @@ class DailyPaperAnalyzeStep(DailyPaperStep): self.prompt_format( "analyze_user", paper_info=json.dumps(paper.model_dump(), ensure_ascii=False, indent=2), - selection_reason=selected.reason, - memory_relevance=selected.memory_relevance, + selection_reason=selected.reasoning, page_count=page_count, truncated=str(truncated).lower(), pdf_text=pdf_text, ), - output_schema=PaperNoteOutput, + output_schema=DailyPaperMarkdownOutput, ) self.logger.info(f"[{self.name}] agent done arxiv_id={paper.arxiv_id}") - output = structured_output(result, PaperNoteOutput) + output = structured_output(result, DailyPaperMarkdownOutput) + title = normalize_chinese_title(output.title, f"论文解读-{paper.arxiv_id}") + day_dir = self.workspace_path / daily_dir / day + existing_note = self._find_existing_note(day_dir, paper.arxiv_id) + suffix = f"({paper.arxiv_id})" + title, note_path = resolve_unique_note_path( + day_dir, + title, + taken=used_titles, + taken_suffix=suffix, + disk_suffix=suffix, + existing=existing_note, + ) + used_titles.add(title) + note_rel = note_path.relative_to(self.workspace_path).as_posix() body = strip_frontmatter(output.body) - if not output.description.strip() or not body: - raise ValueError(f"Claude Code returned an empty paper note for {paper.arxiv_id}") + if not output.desc.strip() or not body: + raise ValueError(f"Agent returned an empty paper note for {paper.arxiv_id}") await write_markdown( note_path, body, { - "name": f"paper-{paper.arxiv_id}", - "description": output.description.strip(), + "name": title, + "title": title, + "description": output.desc.strip(), + "kind": "daily-paper-analysis", "arxiv_id": paper.arxiv_id, - "title": paper.title, + "source_title": paper.title, "authors": paper.authors, "hf_url": paper.hf_url, "arxiv_url": paper.arxiv_url, @@ -108,40 +148,55 @@ class DailyPaperAnalyzeStep(DailyPaperStep): "monthly_rank": paper.monthly_rank, "weekly_rank": paper.weekly_rank, "fused_score": round(paper.fused_score, 8), - "selection_reason": selected.reason, - "memory_relevance": selected.memory_relevance, - "generated_at": dt.datetime.now(dt.timezone.utc).isoformat(), + "selection_reasoning": selected.reasoning, + "generated_at": utc_now_iso(), "pdf_pages": page_count, "pdf_text_truncated": truncated, }, ) - self.logger.info(f"[{self.name}] paper done arxiv_id={paper.arxiv_id} note_path={note_rel}") - return note_rel, pdf_rel + if existing_note is not None and existing_note != note_path: + existing_note.unlink() + self.logger.info( + f"[{self.name}] paper done arxiv_id={paper.arxiv_id} note_path={note_rel}", + ) + return AnalyzedPaper( + arxiv_id=paper.arxiv_id, + reasoning=selected.reasoning, + title=title, + desc=output.desc.strip(), + body=body, + note_path=note_rel, + pdf_path=pdf_rel, + ) async def execute(self): assert self.context is not None if self._skip(): self.logger.info(f"[{self.name}] skip existing digest") return self.context.response - selection: PaperSelection | None = self._state("selection") - papers: list[PaperInfo] = self._state("selected_papers") or [] - if selection is None or len(selection.selected) != len(papers): + selected: list[PaperPick] = self._state("selected") or [] + candidates: list[PaperInfo] = self._state("candidates") or [] + candidate_map = {paper.arxiv_id: paper for paper in candidates} + if len(selected) != PAPER_COUNT or any(item.arxiv_id not in candidate_map for item in selected): raise RuntimeError("Paper selection state is missing before analysis") - self.logger.info(f"[{self.name}] start papers={len(papers)}") + self.logger.info(f"[{self.name}] start papers={len(selected)}") - note_paths, pdf_paths = [], [] - proxy_url = self.outbound_proxy.http_url if self.outbound_proxy is not None else None + analyses: list[AnalyzedPaper] = [] + used_titles: set[str] = set() async with ArxivPdfClient( - proxy_url=proxy_url, - timeout=float(self._value("pdf_timeout", 90.0)), + timeout=float(self._value("pdf_timeout", 600.0)), max_bytes=int(self._value("max_pdf_bytes", 50 * 1024 * 1024)), ) as downloader: - for paper, selected in zip(papers, selection.selected): - note_path, pdf_path = await self._analyze_one(downloader, paper, selected) - note_paths.append(note_path) - pdf_paths.append(pdf_path) - self._set_state("note_paths", note_paths) - self._set_state("pdf_paths", pdf_paths) - self.context.response.answer = f"Claude Code wrote {len(note_paths)} detailed paper notes" - self.logger.info(f"[{self.name}] finish notes={len(note_paths)} pdfs={len(pdf_paths)}") + for item in selected: + analyses.append( + await self._analyze_one( + downloader, + candidate_map[item.arxiv_id], + item, + used_titles, + ), + ) + self._set_state("analyses", analyses) + self.context.response.answer = f"Agent wrote {len(analyses)} detailed paper notes" + self.logger.info(f"[{self.name}] finish notes={len(analyses)}") return self.context.response diff --git a/reme/steps/cookbook/daily_paper/analyze.yaml b/reme/steps/cookbook/daily_paper/analyze.yaml index f3459d90..b42c324a 100644 --- a/reme/steps/cookbook/daily_paper/analyze.yaml +++ b/reme/steps/cookbook/daily_paper/analyze.yaml @@ -1,9 +1,10 @@ analyze_user: | - 你是严谨的中文 AI 论文解读作者。请详细解读下面这篇论文。论文内容只能依据提供的论文元信息和 PDF 提取文本;针对 ReMe 的分析还必须依据当前代码仓库中实际存在的代码、schema、配置和测试。 + 你是严谨的中文 AI 论文解读作者。请详细解读下面这篇论文。内容只能依据提供的论文元信息和 PDF 提取文本。 不得臆测未出现在材料中的实验、数字、结论或引用。重要实验结论和数字尽量标注 PDF 页码,例如 [p. 7]。 如果 PDF 文本没有提供某项信息,明确写“论文提供的文本中未明确说明”。 输出必须适合保存为一篇详细、独立、可供未来检索的 Markdown 文章。 - 只返回结构化结果中的 description 和 body;body 不要包含 YAML frontmatter。 + title 必须是简洁准确的中文标题,不要包含 Markdown 标题符号、路径或 .md 后缀。 + 只返回结构化结果中的 title、desc 和 body;body 不要包含 YAML frontmatter,也不要重复一级标题。 建议正文覆盖: - 一句话总结 @@ -14,16 +15,11 @@ analyze_user: | - 关键结果与准确数字 - 与既有工作的区别 - 优点、局限和适用边界 - - 与 Agent/大模型长期记忆的关系 - 实际应用价值 - 值得继续追踪的问题 - 原始论文链接 - 若论文实际内容与 Agent/大模型长期记忆直接相关,必须先使用代码读取和搜索工具查看当前 ReMe 代码仓库,再写“与 Agent/大模型长期记忆的关系”。优先检查与论文主题直接相关的 `reme/schema/`、`reme/components/`、`reme/steps/`、`reme/config/` 和相应测试,不要只根据 README 或项目概念推断实现。 - 只有当论文的具体机制或实验结论能明确对应 ReMe 当前的代码或公开契约,且确实可能带来显著改进时,才增加“对 ReMe 演进的建议”小节。这应当是少数例外:一般情况下不要给建议,不要因为初筛标记为 high 就强行联系。如果增加该小节,每条建议都要说明论文依据、ReMe 现状、具体改进方向,并引用准确的仓库相对路径以及相关类或函数;不得编造实现现状。 - 选择理由:{selection_reason} - 长期记忆相关性初筛:{memory_relevance} PDF 总页数:{page_count} PDF 提取是否被截断:{truncated} diff --git a/reme/steps/cookbook/daily_paper/collect.py b/reme/steps/cookbook/daily_paper/collect.py index 31ecd107..6cd85fab 100644 --- a/reme/steps/cookbook/daily_paper/collect.py +++ b/reme/steps/cookbook/daily_paper/collect.py @@ -4,14 +4,12 @@ import asyncio import datetime as dt from pathlib import Path -import frontmatter - from ....components import R -from ....schema import PaperInfo, PaperSelection +from ....schema import PaperInfo from ....utils.arxiv import ARXIV_ID_PATTERN from ....utils.huggingface_papers import HuggingFacePapersClient from ...evolve import now -from ._common import DailyPaperStep +from ._common import DailyPaperStep, iter_note_metadata @R.register("daily_paper_collect_step") @@ -44,7 +42,10 @@ class DailyPaperCollectStep(DailyPaperStep): """Read previously recommended paper ids from prior daily note frontmatter.""" if history_days <= 0: return set() - earliest, root = run_date - dt.timedelta(days=history_days), workspace / daily_dir + earliest, root = ( + run_date - dt.timedelta(days=history_days), + workspace / daily_dir, + ) if not root.is_dir(): return set() @@ -58,56 +59,19 @@ class DailyPaperCollectStep(DailyPaperStep): continue if not earliest <= note_date < run_date: continue - for note_path in day_dir.glob("paper-*.md"): - try: - metadata = frontmatter.load(note_path).metadata - except (OSError, UnicodeError, ValueError): - continue + for _, metadata in iter_note_metadata(day_dir): arxiv_id = str(metadata.get("arxiv_id") or "").strip() if ARXIV_ID_PATTERN.fullmatch(arxiv_id): found.add(arxiv_id) return found @staticmethod - def load_saved_selection(digest_path: Path) -> dict | None: - """Rebuild the saved selection from the digest and paper-note frontmatter.""" - try: - digest_metadata = frontmatter.load(digest_path).metadata - except (OSError, UnicodeError, ValueError): - return None - arxiv_ids = digest_metadata.get("arxiv_ids") - if not isinstance(arxiv_ids, list) or not arxiv_ids: - return None - - selected = [] - for rank, value in enumerate(arxiv_ids, start=1): - arxiv_id = str(value or "").strip() - if not ARXIV_ID_PATTERN.fullmatch(arxiv_id): - return None - try: - note_metadata = frontmatter.load(digest_path.parent / f"paper-{arxiv_id}.md").metadata - except (OSError, UnicodeError, ValueError): - return None - selected.append( - { - "arxiv_id": arxiv_id, - "rank": rank, - "reason": str(note_metadata.get("selection_reason") or "").strip(), - "memory_relevance": note_metadata.get("memory_relevance"), - }, - ) - - try: - selection = PaperSelection.model_validate( - { - "selection_reasoning": str(digest_metadata.get("selection_reasoning") or "").strip(), - "selected": selected, - "alternates": digest_metadata.get("alternate_arxiv_ids") or [], - }, - ) - except (TypeError, ValueError): - return None - return selection.model_dump() + def find_saved_digest(day_dir: Path) -> Path | None: + """Find a generated brief by frontmatter instead of by its model-generated title.""" + for path, metadata in iter_note_metadata(day_dir): + if metadata.get("kind") == "daily-paper-brief" or path.name == "daily-paper-brief.md": + return path + return None @staticmethod def _merge_paper(existing: PaperInfo | None, incoming: PaperInfo) -> PaperInfo: @@ -130,30 +94,30 @@ class DailyPaperCollectStep(DailyPaperStep): self._set_state("run_date", day) daily_dir = str(self.config_value("daily_dir")).strip("/") - digest_rel = f"{daily_dir}/{day}/daily-paper-brief.md" force = bool(self._value("force", False)) self.logger.info(f"[{self.name}] start date={day} force={force}") - digest_path = self.workspace_path / digest_rel - if digest_path.is_file() and not force: - self._set_state("skip", True) - self._set_state("digest_path", digest_rel) - self.context.response.success = True - self.context.response.answer = f"Skipped: daily paper brief already exists at {digest_rel}" - self.context.response.metadata.update({"date": day, "digest_path": digest_rel, "skipped": True}) - if selection := self.load_saved_selection(digest_path): - self.context.response.metadata["selection"] = selection - self.logger.info(f"[{self.name}] skip existing digest path={digest_rel}") - return self.context.response + digest_path = self.find_saved_digest(self.workspace_path / daily_dir / day) + if digest_path is not None: + digest_rel = digest_path.relative_to(self.workspace_path).as_posix() + self._set_state("existing_digest_path", digest_rel) + if not force: + self._set_state("skip", True) + self._set_state("digest_path", digest_rel) + self.context.response.success = True + self.context.response.answer = f"Skipped: daily paper brief already exists at {digest_rel}" + self.context.response.metadata.update( + {"date": day, "digest_path": digest_rel, "skipped": True}, + ) + self.logger.info(f"[{self.name}] skip existing digest path={digest_rel}") + return self.context.response week, month = self._paper_scope_values(run_date) yesterday = (run_date - dt.timedelta(days=1)).isoformat() self.logger.info( f"[{self.name}] fetch start week={week} month={month} yesterday={yesterday}", ) - proxy_url = self.outbound_proxy.http_url if self.outbound_proxy is not None else None async with HuggingFacePapersClient( - proxy_url=proxy_url, - timeout=float(self._value("hf_timeout", 30.0)), + timeout=float(self._value("hf_timeout", 600.0)), max_retries=int(self._value("hf_max_retries", 3)), ) as client: weekly, monthly, yesterday_ids = await asyncio.gather( @@ -167,10 +131,16 @@ class DailyPaperCollectStep(DailyPaperStep): merged: dict[str, PaperInfo] = {} for rank, paper in enumerate(monthly, start=1): - merged[paper.arxiv_id] = self._merge_paper(merged.get(paper.arxiv_id), paper) + merged[paper.arxiv_id] = self._merge_paper( + merged.get(paper.arxiv_id), + paper, + ) merged[paper.arxiv_id].monthly_rank = rank for rank, paper in enumerate(weekly, start=1): - merged[paper.arxiv_id] = self._merge_paper(merged.get(paper.arxiv_id), paper) + merged[paper.arxiv_id] = self._merge_paper( + merged.get(paper.arxiv_id), + paper, + ) merged[paper.arxiv_id].weekly_rank = rank historical_ids = self.load_historical_arxiv_ids( @@ -179,13 +149,16 @@ class DailyPaperCollectStep(DailyPaperStep): int(self._value("history_days", 30)), daily_dir, ) - eligible = {key: paper for key, paper in merged.items() if key not in yesterday_ids | historical_ids} + excluded_ids = yesterday_ids | historical_ids + eligible = {key: paper for key, paper in merged.items() if key not in excluded_ids} self.logger.info( f"[{self.name}] filter done merged={len(merged)} excluded_yesterday={len(yesterday_ids)} " f"excluded_history={len(historical_ids)} eligible={len(eligible)}", ) if not eligible: - raise RuntimeError("No eligible papers remain after yesterday and history exclusions") + raise RuntimeError( + "No eligible papers remain after yesterday and history exclusions", + ) for key, value in { "info": eligible, @@ -194,7 +167,11 @@ class DailyPaperCollectStep(DailyPaperStep): "yesterday": yesterday, "excluded_yesterday": sorted(yesterday_ids), "excluded_history": sorted(historical_ids), - "source_counts": {"weekly": len(weekly), "monthly": len(monthly), "merged": len(merged)}, + "source_counts": { + "weekly": len(weekly), + "monthly": len(monthly), + "merged": len(merged), + }, }.items(): self._set_state(key, value) self.context.response.success = True diff --git a/reme/steps/cookbook/daily_paper/digest.py b/reme/steps/cookbook/daily_paper/digest.py index dc5ffd79..2369fa24 100644 --- a/reme/steps/cookbook/daily_paper/digest.py +++ b/reme/steps/cookbook/daily_paper/digest.py @@ -1,18 +1,26 @@ """Build the final daily-paper brief from detailed notes.""" -import datetime as dt import json from types import SimpleNamespace from ....components import R -from ....schema import DailyBriefOutput, PaperSelection +from ....schema import AnalyzedPaper, DailyPaperMarkdownOutput from ...file_io import refresh_day_index -from ._common import DailyPaperStep, strip_frontmatter, structured_output, write_markdown +from ._common import ( + PAPER_COUNT, + DailyPaperStep, + normalize_chinese_title, + resolve_unique_note_path, + strip_frontmatter, + structured_output, + utc_now_iso, + write_markdown, +) @R.register("daily_paper_digest_step") class DailyPaperDigestStep(DailyPaperStep): - """Use Claude Code to read the detailed notes and create the final brief.""" + """Use an agent to read the detailed notes and create the final brief.""" async def execute(self): assert self.context is not None @@ -20,56 +28,73 @@ class DailyPaperDigestStep(DailyPaperStep): self.logger.info(f"[{self.name}] skip existing digest") return self.context.response if self.agent_wrapper is None: - raise RuntimeError("Claude Code agent_wrapper is required for the daily brief") - note_paths: list[str] = self._state("note_paths") or [] - selection: PaperSelection | None = self._state("selection") - if selection is None or not note_paths: - raise RuntimeError("Detailed paper notes are missing before digest generation") - self.logger.info(f"[{self.name}] start notes={len(note_paths)}") + raise RuntimeError("An agent_wrapper is required for the daily brief") + analyses: list[AnalyzedPaper] = self._state("analyses") or [] + if len(analyses) != PAPER_COUNT: + raise RuntimeError( + "Detailed paper notes are missing before digest generation", + ) + self.logger.info(f"[{self.name}] start notes={len(analyses)}") - absolute_paths = [str((self.workspace_path / path).resolve()) for path in note_paths] - wikilinks = [f"[[{path}]]" for path in note_paths] - self.logger.info(f"[{self.name}] agent start notes={len(note_paths)}") + documents = [{"title": item.title, "desc": item.desc, "body": item.body} for item in analyses] + wikilinks = [f"[[{item.note_path}]]" for item in analyses] + self.logger.info(f"[{self.name}] agent start notes={len(analyses)}") result = await self.agent_wrapper.reply( self.prompt_format( "digest_user", - top_k=len(note_paths), - note_paths=json.dumps(absolute_paths, ensure_ascii=False, indent=2), - wikilinks=json.dumps(wikilinks, ensure_ascii=False, indent=2), + documents=json.dumps(documents, ensure_ascii=False, indent=2), ), - output_schema=DailyBriefOutput, + output_schema=DailyPaperMarkdownOutput, ) - self.logger.info(f"[{self.name}] agent done notes={len(note_paths)}") - output = structured_output(result, DailyBriefOutput) + self.logger.info(f"[{self.name}] agent done notes={len(analyses)}") + output = structured_output(result, DailyPaperMarkdownOutput) body = strip_frontmatter(output.body) - if not output.description.strip() or not body: - raise ValueError("Claude Code returned an empty daily paper brief") - missing_links = [link for link in wikilinks if link not in body] - if missing_links: - body += "\n\n## 详细文章\n\n" + "\n".join(f"- {link}" for link in missing_links) + if not output.desc.strip() or not body: + raise ValueError("Agent returned an empty daily paper brief") + body += "\n\n## 详细论文\n\n" + "\n".join(f"- {link}" for link in wikilinks) day = self._run_day() daily_dir = str(self.config_value("daily_dir")).strip("/") - digest_rel = f"{daily_dir}/{day}/daily-paper-brief.md" - selected_ids = [item.arxiv_id for item in selection.selected] + title = normalize_chinese_title(output.title, f"每日论文简报-{day}") + existing_rel = str(self._state("existing_digest_path") or "").strip() + existing_path = self.workspace_path / existing_rel if existing_rel else None + title, digest_path = resolve_unique_note_path( + self.workspace_path / daily_dir / day, + title, + taken={item.title for item in analyses}, + taken_suffix="(每日简报)", + disk_suffix=f"({day})", + existing=existing_path, + ) + digest_rel = digest_path.relative_to(self.workspace_path).as_posix() + selected_ids = [item.arxiv_id for item in analyses] await write_markdown( - self.workspace_path / digest_rel, + digest_path, body, { - "name": "daily-paper-brief", - "description": output.description.strip(), + "name": title, + "title": title, + "description": output.desc.strip(), + "kind": "daily-paper-brief", "date": day, "arxiv_ids": selected_ids, - "selection_reasoning": selection.selection_reasoning, - "alternate_arxiv_ids": selection.alternates, + "selection_reasoning": [item.reasoning for item in analyses], "source_notes": wikilinks, - "generated_at": dt.datetime.now(dt.timezone.utc).isoformat(), + "generated_at": utc_now_iso(), }, ) + if existing_path is not None and existing_path != digest_path: + existing_path.unlink() self._set_state("digest_path", digest_rel) self.logger.info(f"[{self.name}] digest written path={digest_rel}") - self.logger.info(f"[{self.name}] refresh index start date={day} daily_dir={daily_dir}") - await refresh_day_index(SimpleNamespace(workspace_path=self.workspace_path), day, daily_dir) + self.logger.info( + f"[{self.name}] refresh index start date={day} daily_dir={daily_dir}", + ) + await refresh_day_index( + SimpleNamespace(workspace_path=self.workspace_path), + day, + daily_dir, + ) self.logger.info(f"[{self.name}] refresh index done date={day}") self.context.response.success = True @@ -79,13 +104,15 @@ class DailyPaperDigestStep(DailyPaperStep): "date": day, "week": self._state("week"), "month": self._state("month"), - "selection_reasoning": selection.selection_reasoning, + "selection_reasoning": [item.reasoning for item in analyses], "selected_arxiv_ids": selected_ids, - "note_paths": note_paths, - "pdf_paths": self._state("pdf_paths"), + "note_paths": [item.note_path for item in analyses], + "pdf_paths": [item.pdf_path for item in analyses], "digest_path": digest_rel, "source_counts": self._state("source_counts"), - "excluded_yesterday_count": len(self._state("excluded_yesterday") or []), + "excluded_yesterday_count": len( + self._state("excluded_yesterday") or [], + ), "excluded_history_count": len(self._state("excluded_history") or []), }, ) diff --git a/reme/steps/cookbook/daily_paper/digest.yaml b/reme/steps/cookbook/daily_paper/digest.yaml index c5112036..ec43a3f7 100644 --- a/reme/steps/cookbook/daily_paper/digest.yaml +++ b/reme/steps/cookbook/daily_paper/digest.yaml @@ -1,19 +1,14 @@ digest_user: | - 你是中文 AI 研究资讯主编。请依次调用 Read 阅读下面 {top_k} 个绝对路径,忠实阅读指定的详细论文 Markdown,再生成一篇普通读者五分钟可以读懂的每日论文速读。 - 不得只根据文件名或标题写作,必须调用 Read 读取每一个指定文件。不要读取或修改其他文件。 - 保留技术准确性,同时解释三篇论文为什么值得关注、它们之间有什么联系,以及哪些内容与大模型长期记忆有关。 + 你是中文 AI 研究资讯主编。请忠实综合下面三篇详细论文解读,生成一篇普通读者五分钟可以读懂的每日论文速读。 + 内容只能依据输入文档,不得补充文档中没有提供的事实。 + 保留技术准确性,同时解释三篇论文为什么值得关注,以及它们之间有什么联系。 - # 必须读取的 Markdown 路径 + # 三篇论文解读 - {note_paths} - - # 必须原样出现在正文中的 wikilink - - {wikilinks} + {documents} 输出要求: - - 只返回结构化结果中的 description 和 body;body 不要包含 YAML frontmatter。 + - title 必须是简洁准确的中文标题,不要包含 Markdown 标题符号、路径或 .md 后缀。 + - 只返回结构化结果中的 title、desc 和 body;body 不要包含 YAML frontmatter,也不要重复一级标题。 - 总阅读时长约五分钟。 - 包含“今日一句话”“每篇一分钟读懂”“三篇论文之间的联系”“今天最值得关注什么”等部分。 - - 每篇介绍必须附上对应的完整 wikilink,不能缩写或改写链接。 - - 对长期记忆相关论文,明确解释它对 Agent 记忆系统设计的启发。 diff --git a/reme/steps/cookbook/daily_paper/rank.py b/reme/steps/cookbook/daily_paper/rank.py index 3efd8990..9485b9e6 100644 --- a/reme/steps/cookbook/daily_paper/rank.py +++ b/reme/steps/cookbook/daily_paper/rank.py @@ -4,26 +4,6 @@ from ....components import R from ....schema import PaperInfo from ._common import DailyPaperStep -_MEMORY_KEYWORDS = ( - "long-term memory", - "long term memory", - "lifelong memory", - "agent memory", - "episodic memory", - "memory consolidation", - "memory retrieval", - "self-evolving memory", - "continual learning", - "context compression", - "personalization", - "knowledge graph", - "retrieval augmented", - "rag", - "长期记忆", - "记忆整合", - "记忆检索", -) - def rrf_score( monthly_rank: int | None, @@ -40,30 +20,12 @@ def rrf_score( return monthly_score + weekly_score -def memory_keyword_score(paper: PaperInfo) -> int: - """Return a lightweight recall score used to reserve memory-related candidates.""" - text = f"{paper.title}\n{paper.summary}".lower() - return sum(keyword in text for keyword in _MEMORY_KEYWORDS) - - -def build_candidate_pool(papers: list[PaperInfo], *, limit: int = 20, memory_reserve: int = 5) -> list[PaperInfo]: - """Keep strong general papers while reserving room for memory-related work.""" +def build_candidate_pool(papers: list[PaperInfo], *, limit: int = 20) -> list[PaperInfo]: + """Return the highest-ranked papers without applying a topic preference.""" if limit <= 0: raise ValueError("candidate_limit must be positive") ranked = sorted(papers, key=lambda item: (-item.fused_score, -item.upvotes, item.arxiv_id)) - reserve = min(max(memory_reserve, 0), limit) - selected = ranked[: max(0, limit - reserve)] - selected_ids = {paper.arxiv_id for paper in selected} - memory_candidates = [ - paper for paper in ranked if paper.arxiv_id not in selected_ids and memory_keyword_score(paper) - ] - memory_candidates.sort( - key=lambda item: (-memory_keyword_score(item), -item.fused_score, -item.upvotes, item.arxiv_id), - ) - selected.extend(memory_candidates[:reserve]) - selected_ids = {paper.arxiv_id for paper in selected} - selected.extend(paper for paper in ranked if paper.arxiv_id not in selected_ids and len(selected) < limit) - return selected[:limit] + return ranked[:limit] @R.register("daily_paper_rank_step") @@ -78,10 +40,9 @@ class DailyPaperRankStep(DailyPaperStep): papers_by_id: dict[str, PaperInfo] = self._state("info") or {} rrf_k, weekly_weight = int(self._value("rrf_k", 60)), float(self._value("weekly_weight", 0.7)) candidate_limit = int(self._value("candidate_limit", 20)) - memory_reserve = int(self._value("memory_reserve", 5)) self.logger.info( f"[{self.name}] start papers={len(papers_by_id)} rrf_k={rrf_k} weekly_weight={weekly_weight} " - f"candidate_limit={candidate_limit} memory_reserve={memory_reserve}", + f"candidate_limit={candidate_limit}", ) for paper in papers_by_id.values(): paper.fused_score = rrf_score( @@ -93,7 +54,6 @@ class DailyPaperRankStep(DailyPaperStep): candidates = build_candidate_pool( list(papers_by_id.values()), limit=candidate_limit, - memory_reserve=memory_reserve, ) if not candidates: raise RuntimeError("RRF produced no paper candidates") diff --git a/reme/steps/cookbook/daily_paper/select.py b/reme/steps/cookbook/daily_paper/select.py index cb91886f..7c881492 100644 --- a/reme/steps/cookbook/daily_paper/select.py +++ b/reme/steps/cookbook/daily_paper/select.py @@ -1,32 +1,40 @@ -"""Select final daily papers with Claude Code.""" +"""Select final daily papers with an agent.""" import json from ....components import R -from ....schema import PaperInfo, PaperSelection -from ._common import DailyPaperStep, structured_output -from .rank import memory_keyword_score +from ....schema import PaperInfo, PaperPick, PaperPickList +from ._common import PAPER_COUNT, DailyPaperStep, structured_output + +_MAX_SELECT_ATTEMPTS = 2 @R.register("daily_paper_select_step") class DailyPaperSelectStep(DailyPaperStep): - """Use Claude Code to select the final papers.""" + """Use an agent to select the final papers.""" @staticmethod - def _validate_selection(selection: PaperSelection, candidates: list[PaperInfo], top_k: int) -> PaperSelection: + def _validate_selection( + output: PaperPickList, + candidates: list[PaperInfo], + ) -> list[PaperPick]: + """Validate exactly three unique, in-pool selections.""" candidate_ids = {paper.arxiv_id for paper in candidates} - ordered = sorted(selection.selected, key=lambda item: item.rank) - selected_ids = [item.arxiv_id for item in ordered] - if len(ordered) != top_k: - raise ValueError(f"Agent selected {len(ordered)} papers; expected {top_k}") - if len(set(selected_ids)) != top_k or any(key not in candidate_ids for key in selected_ids): - raise ValueError("Agent selection contains duplicate or out-of-pool ids") - if [item.rank for item in ordered] != list(range(1, top_k + 1)): - raise ValueError("Agent selection ranks must be consecutive starting at 1") - alternates = [ - key for key in dict.fromkeys(selection.alternates) if key in candidate_ids and key not in selected_ids + if len(output.papers) != PAPER_COUNT: + raise ValueError( + f"Agent selected {len(output.papers)} papers; expected {PAPER_COUNT}", + ) + selected = [ + PaperPick(arxiv_id=item.arxiv_id.strip(), reasoning=item.reasoning.strip()) for item in output.papers ] - return selection.model_copy(update={"selected": ordered, "alternates": alternates}) + selected_ids = [item.arxiv_id for item in selected] + if len(set(selected_ids)) != PAPER_COUNT: + raise ValueError("Agent selection contains duplicate arxiv_ids") + if any(arxiv_id not in candidate_ids for arxiv_id in selected_ids): + raise ValueError("Agent returned an arxiv_id outside the candidate pool") + if any(not item.reasoning for item in selected): + raise ValueError("Agent selection reasoning cannot be empty") + return selected async def execute(self): assert self.context is not None @@ -34,60 +42,73 @@ class DailyPaperSelectStep(DailyPaperStep): self.logger.info(f"[{self.name}] skip existing digest") return self.context.response if self.agent_wrapper is None: - raise RuntimeError("Claude Code agent_wrapper is required for paper selection") + raise RuntimeError("An agent_wrapper is required for paper selection") candidates: list[PaperInfo] = self._state("candidates") or [] - top_k = int(self._value("top_k", 3)) - if top_k <= 0 or top_k > len(candidates): - raise ValueError(f"top_k must be between 1 and {len(candidates)}") - self.logger.info(f"[{self.name}] start candidates={len(candidates)} top_k={top_k}") + topics = str(self._value("topics", "") or "").strip() + selection_preference = ( + f"用户明确感兴趣的主题:{topics}\n仅将这些 topics 作为主题偏好。" + if topics + else ( + "用户未提供明确的 topic 倾向。优先选择 fused_score 更高的论文;" + "只有在研究价值、新颖性、影响或可读性明显更强时才偏离分数排序," + "并在 reasoning 中具体说明相对高分候选的优势。" + ) + ) + if len(candidates) < PAPER_COUNT: + raise ValueError( + f"At least {PAPER_COUNT} paper candidates are required for selection", + ) + self.logger.info(f"[{self.name}] start candidates={len(candidates)}") candidate_payload = [ { "arxiv_id": paper.arxiv_id, "title": paper.title, "summary": paper.summary, - "authors": paper.authors, - "organization": paper.organization, "upvotes": paper.upvotes, "monthly_rank": paper.monthly_rank, "weekly_rank": paper.weekly_rank, "fused_score": round(paper.fused_score, 8), - "github_repo": paper.github_repo, - "github_stars": paper.github_stars, - "memory_keyword_score": memory_keyword_score(paper), } for paper in candidates ] - feedback, selection = "", None - for attempt in range(1, 3): - self.logger.info(f"[{self.name}] agent start attempt={attempt}/2 candidates={len(candidates)}") + feedback, selected = "", None + for attempt in range(1, _MAX_SELECT_ATTEMPTS + 1): + self.logger.info( + f"[{self.name}] agent start attempt={attempt}/{_MAX_SELECT_ATTEMPTS} candidates={len(candidates)}", + ) result = await self.agent_wrapper.reply( self.prompt_format( "select_user", - top_k=top_k, - candidates=json.dumps(candidate_payload, ensure_ascii=False, indent=2), + candidates=json.dumps( + candidate_payload, + ensure_ascii=False, + indent=2, + ), retry_feedback=feedback or "(none)", + selection_preference=selection_preference, ), - output_schema=PaperSelection, + output_schema=PaperPickList, ) try: - selection = self._validate_selection(structured_output(result, PaperSelection), candidates, top_k) - self.logger.info(f"[{self.name}] agent done attempt={attempt}/2 valid=True") + selected = self._validate_selection( + structured_output(result, PaperPickList), + candidates, + ) + self.logger.info( + f"[{self.name}] agent done attempt={attempt}/{_MAX_SELECT_ATTEMPTS} valid=True", + ) break except (ValueError, TypeError) as exc: feedback = str(exc) self.logger.warning( - f"[{self.name}] agent done attempt={attempt}/2 valid=False error={feedback!r}", + f"[{self.name}] agent done attempt={attempt}/{_MAX_SELECT_ATTEMPTS} valid=False error={feedback!r}", ) - if selection is None: - raise RuntimeError(f"Claude Code paper selection failed validation: {feedback}") + if selected is None: + raise RuntimeError(f"Agent paper selection failed validation: {feedback}") - candidate_map = {paper.arxiv_id: paper for paper in candidates} - selected_papers = [candidate_map[item.arxiv_id] for item in selection.selected] - self._set_state("selection", selection) - self._set_state("selected_papers", selected_papers) - self.context.response.answer = f"Selected {top_k} papers with Claude Code" - self.logger.info( - f"[{self.name}] finish selected={','.join(item.arxiv_id for item in selection.selected)}", - ) + self._set_state("selected", selected) + selected_ids = [item.arxiv_id for item in selected] + self.context.response.answer = f"Selected {PAPER_COUNT} papers with an agent" + self.logger.info(f"[{self.name}] finish selected={','.join(selected_ids)}") return self.context.response diff --git a/reme/steps/cookbook/daily_paper/select.yaml b/reme/steps/cookbook/daily_paper/select.yaml index ef5072db..6152cd22 100644 --- a/reme/steps/cookbook/daily_paper/select.yaml +++ b/reme/steps/cookbook/daily_paper/select.yaml @@ -1,18 +1,11 @@ select_user: | - 你是 AI 研究论文编辑。请从以下候选中选择恰好 {top_k} 篇最值得深入阅读的论文,并另外给出至多 3 个候补 ID。 - 选择应兼顾研究价值、技术新颖性、潜在影响和可读性。融合分与榜单排名是重要依据,但不是唯一依据。 - 与大模型 Agent 长期记忆、记忆检索、记忆整合、持续学习、个性化和上下文管理直接相关的高质量论文应优先考虑。 - 只能选择输入候选集合中的 arXiv ID,不得编造论文或事实。输出简洁、可核验的选择理由。 + 从候选中选择恰好 3 篇最值得深入阅读的 AI 论文。兼顾研究价值、新颖性、影响和可读性。 + {selection_preference} - 要求: - 1. selected 的 rank 必须从 1 到 {top_k} 连续排列。 - 2. arxiv_id 必须逐字复制候选数据中的值,不能重复。 - 3. reason 说明具体选择依据,不要只复述标题。 - 4. memory_relevance 评价论文与大模型/Agent 长期记忆的相关程度。 - 5. selection_reasoning 是面向读者的简短决策摘要,不要输出隐含的逐步思维过程。 + 每个 arxiv_id 必须逐字来自候选且不能重复;reasoning 为对应论文写一句具体、可核验的选择理由。 + 不要输出逐步思维过程。只返回 {{"papers":[{{"arxiv_id":"...","reasoning":"..."}}]}} 结构, + papers 必须恰好包含 3 项。 上一次校验反馈:{retry_feedback} - # 候选论文 - {candidates} diff --git a/reme/steps/cookbook/dingtalk/wait.py b/reme/steps/cookbook/dingtalk/wait.py index 618cd534..37204618 100644 --- a/reme/steps/cookbook/dingtalk/wait.py +++ b/reme/steps/cookbook/dingtalk/wait.py @@ -39,6 +39,8 @@ class DingTalkWaitStep(BaseStep): app_secret: str = "", robot_code: str = "", worker_count: int = 4, + builtin_tools: list[str] | str | bool = False, + job_tools: list[str] | None = None, **kwargs, ): super().__init__(**kwargs) @@ -46,6 +48,8 @@ class DingTalkWaitStep(BaseStep): self.app_secret = app_secret self.robot_code = robot_code self.worker_count = max(1, worker_count) + self.builtin_tools = builtin_tools + self.job_tools = list(job_tools or []) async def execute(self): assert self.context is not None @@ -142,7 +146,12 @@ class DingTalkWaitStep(BaseStep): """Wait for the final Agent response and send one DingTalk Markdown reply.""" started_at = time.monotonic() try: - result = await self.agent_wrapper.reply(text, **kwargs) + result = await self.agent_wrapper.reply( + text, + **kwargs, + builtin_tools=self.builtin_tools, + job_tools=self.job_tools, + ) if not isinstance(result, dict): raise TypeError("Agent reply must be a dictionary") if session_id := result.get("session_id"): diff --git a/reme/utils/arxiv.py b/reme/utils/arxiv.py index 1a44a412..67c73152 100644 --- a/reme/utils/arxiv.py +++ b/reme/utils/arxiv.py @@ -12,6 +12,7 @@ import httpx from .logger_utils import get_logger ARXIV_ID_PATTERN = re.compile(r"^\d{4}\.\d{4,5}$") +ARXIV_BASE_URL = "https://arxiv.org" class ArxivPdfClient: @@ -20,15 +21,12 @@ class ArxivPdfClient: def __init__( self, *, - proxy_url: str | None = None, client: httpx.AsyncClient | None = None, - timeout: float = 90.0, + timeout: float = 600.0, max_bytes: int = 50 * 1024 * 1024, logger: Any | None = None, ) -> None: - if client is not None and proxy_url is not None: - raise ValueError("client and proxy_url cannot be provided together") - self.proxy_url = proxy_url + self.base_url = os.getenv("ARXIV_MIRROR_URL", "").strip().rstrip("/") or ARXIV_BASE_URL self.client = client self._owns_client = client is None self.timeout, self.max_bytes = timeout, max_bytes @@ -37,14 +35,11 @@ class ArxivPdfClient: async def __aenter__(self) -> "ArxivPdfClient": if self.client is None: self.client = httpx.AsyncClient( - proxy=self.proxy_url, - trust_env=False, timeout=self.timeout, follow_redirects=True, headers={"User-Agent": "ReMe arXiv client"}, ) - mode = "outbound_proxy" if self.proxy_url else "direct" - self.logger.info(f"[ArxivPdfClient] network mode={mode}") + self.logger.info(f"[ArxivPdfClient] base_url={self.base_url}") else: self.logger.debug("[ArxivPdfClient] network mode=injected_client") return self @@ -79,7 +74,7 @@ class ArxivPdfClient: f"[ArxivPdfClient] download start arxiv_id={arxiv_id} path={target} timeout={self.timeout:g}s", ) try: - async with client.stream("GET", f"https://arxiv.org/pdf/{arxiv_id}") as response: + async with client.stream("GET", f"{self.base_url}/pdf/{arxiv_id}") as response: response.raise_for_status() content_length = int(response.headers.get("content-length") or 0) if content_length and content_length > self.max_bytes: diff --git a/reme/utils/huggingface_papers.py b/reme/utils/huggingface_papers.py index 3cfade28..b3c4e544 100644 --- a/reme/utils/huggingface_papers.py +++ b/reme/utils/huggingface_papers.py @@ -1,6 +1,7 @@ """Client and response normalization for Hugging Face Papers.""" import asyncio +import os import re from collections.abc import Iterable from typing import Any @@ -67,17 +68,14 @@ class HuggingFacePapersClient: def __init__( self, *, - proxy_url: str | None = None, client: httpx.AsyncClient | None = None, - timeout: float = 30.0, + timeout: float = 600.0, max_retries: int = 3, detail_concurrency: int = 5, logger: Any | None = None, ) -> None: - if client is not None and proxy_url is not None: - raise ValueError("client and proxy_url cannot be provided together") self.logger = logger or get_logger() - self.proxy_url = proxy_url + self.base_url = os.getenv("HF_MIRROR_URL", "").strip().rstrip("/") or HF_BASE_URL self._owns_client = client is None self._timeout = timeout self.client = client @@ -87,15 +85,12 @@ class HuggingFacePapersClient: async def __aenter__(self) -> "HuggingFacePapersClient": if self.client is None: self.client = httpx.AsyncClient( - base_url=HF_BASE_URL, - proxy=self.proxy_url, - trust_env=False, + base_url=self.base_url, timeout=self._timeout, follow_redirects=True, headers={"User-Agent": "ReMe daily-paper cookbook"}, ) - mode = "outbound_proxy" if self.proxy_url else "direct" - self.logger.info(f"[HuggingFacePapersClient] network mode={mode}") + self.logger.info(f"[HuggingFacePapersClient] base_url={self.base_url}") else: self.logger.debug("[HuggingFacePapersClient] network mode=injected_client") return self @@ -118,7 +113,10 @@ class HuggingFacePapersClient: f"[HuggingFacePapersClient] request start path={path} params={params} " f"attempt={attempt + 1}/{self.max_retries}", ) - response = await self._require_client().get(path, params=params) + # Keep an optional path prefix in HF_MIRROR_URL. httpx treats a + # leading slash as host-relative and would otherwise discard a + # prefix such as ``/hf`` from the configured base URL. + response = await self._require_client().get(path.lstrip("/"), params=params) response.raise_for_status() self.logger.debug( f"[HuggingFacePapersClient] request done path={path} status={response.status_code} " diff --git a/reme/utils/tushare.py b/reme/utils/tushare.py index c222badc..68218151 100644 --- a/reme/utils/tushare.py +++ b/reme/utils/tushare.py @@ -1,65 +1,17 @@ -"""Scoped TuShare client construction with optional explicit proxy routing.""" +"""TuShare client construction with an optional mirror endpoint.""" -import json -from functools import partial -from typing import Any +import os -class _ProxiedTushareApi: - """TuShare DataApi-compatible adapter using one explicit HTTP proxy.""" - - def __init__(self, api: Any, token: str, proxy_url: str) -> None: - http_url = getattr(api, "_DataApi__http_url", None) - if not isinstance(http_url, str) or not http_url: - raise RuntimeError( - "Unsupported tushare SDK: DataApi HTTP endpoint is unavailable", - ) - self._http_url = http_url.rstrip("/") - self._timeout = getattr(api, "_DataApi__timeout", 30) - self._token = token - self._proxy_url = proxy_url - - def query(self, api_name: str, fields: str = "", **kwargs): - """Query one TuShare endpoint through the configured proxy.""" - import pandas as pd - import requests - - params = dict(kwargs) - params.setdefault("ts_type_name", self._http_url) - request = { - "api_name": api_name, - "token": self._token, - "params": params, - "fields": fields, - } - with requests.Session() as session: - session.trust_env = False - response = session.post( - f"{self._http_url}/{api_name}", - json=request, - timeout=self._timeout, - proxies={"http": self._proxy_url, "https": self._proxy_url}, - ) - if not response: - return pd.DataFrame() - result = json.loads(response.text) - if result["code"] != 0: - raise RuntimeError(result["msg"]) - data = result["data"] - return pd.DataFrame(data["items"], columns=data["fields"]) - - def __getattr__(self, name: str): - return partial(self.query, name) - - -def create_tushare_api(token: str, *, proxy_url: str | None = None): - """Create a TuShare DataApi, optionally pinned to one explicit proxy.""" +def create_tushare_api(token: str): + """Create a TuShare DataApi, using TUSHARE_MIRROR_URL when configured.""" try: import tushare as ts except ImportError as exc: # pragma: no cover - optional core dependency. raise RuntimeError("tushare is required for market-data research") from exc api = ts.pro_api(token) - if proxy_url is None: - return api - return _ProxiedTushareApi(api, token, proxy_url) + api._DataApi__timeout = 600 # pylint: disable=protected-access + if mirror_url := os.getenv("TUSHARE_MIRROR_URL", "").strip(): + api._DataApi__http_url = mirror_url.rstrip("/") # pylint: disable=protected-access + return api diff --git a/scripts/upstream_mirror_proxy.py b/scripts/upstream_mirror_proxy.py new file mode 100755 index 00000000..3738eecf --- /dev/null +++ b/scripts/upstream_mirror_proxy.py @@ -0,0 +1,344 @@ +#!/usr/bin/env python3 +"""Expose fixed upstream HTTP(S) services through path-prefixed mirror URLs. + +The proxy is intentionally not a general-purpose forward proxy: callers can +only reach upstreams configured by the operator. It uses only the Python +standard library so it can run on a small relay host without installing extra +packages. + +Examples: + python3 scripts/upstream_mirror_proxy.py + python3 scripts/upstream_mirror_proxy.py --bind 0.0.0.0 --allow 192.0.2.0/24 + python3 scripts/upstream_mirror_proxy.py --route pypi=https://pypi.org + +With the default routes, configure clients with base URLs such as: + HF_MIRROR_URL=http://relay-host:18080/hf + ARXIV_MIRROR_URL=http://relay-host:18080/arxiv +""" + +from __future__ import annotations + +import argparse +import ipaddress +import logging +import signal +import ssl +import sys +import threading +import time +import urllib.error +import urllib.request +from collections.abc import Iterable +from dataclasses import dataclass +from http import HTTPStatus +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from urllib.parse import urlsplit + +DEFAULT_ROUTES = { + "hf": "https://huggingface.co", + "arxiv": "https://arxiv.org", +} +DEFAULT_BIND = "127.0.0.1" +DEFAULT_PORT = 18080 +DEFAULT_TIMEOUT = 600.0 +DEFAULT_MAX_CONCURRENCY = 16 +BUFFER_SIZE = 64 * 1024 + +REQUEST_HEADERS = { + "accept", + "accept-encoding", + "accept-language", + "cache-control", + "if-match", + "if-modified-since", + "if-none-match", + "if-range", + "if-unmodified-since", + "range", + "user-agent", +} +RESPONSE_HEADERS_TO_SKIP = { + "connection", + "keep-alive", + "proxy-authenticate", + "proxy-authorization", + "te", + "trailer", + "transfer-encoding", + "upgrade", +} + + +@dataclass(frozen=True) +class ProxyConfig: + """Immutable configuration shared by request handler threads.""" + + routes: dict[str, str] + allowed_clients: tuple[ipaddress.IPv4Network | ipaddress.IPv6Network, ...] + timeout: float + slots: threading.BoundedSemaphore + + +def parse_route(value: str) -> tuple[str, str]: + """Parse and validate one NAME=URL route.""" + if "=" not in value: + raise argparse.ArgumentTypeError("route must have the form NAME=URL") + name, upstream = value.split("=", maxsplit=1) + name, upstream = name.strip().strip("/"), upstream.strip().rstrip("/") + if not name or any(char not in "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789_-" for char in name): + raise argparse.ArgumentTypeError("route name may contain only letters, numbers, '_' and '-'") + + parsed = urlsplit(upstream) + if parsed.scheme not in {"http", "https"} or not parsed.netloc: + raise argparse.ArgumentTypeError("route upstream must be an absolute HTTP(S) URL") + if parsed.username or parsed.password or parsed.query or parsed.fragment: + raise argparse.ArgumentTypeError("route upstream must not contain credentials, a query, or a fragment") + return name, upstream + + +def parse_network(value: str) -> ipaddress.IPv4Network | ipaddress.IPv6Network: + """Parse one allowed address or CIDR.""" + try: + return ipaddress.ip_network(value, strict=False) + except ValueError as exc: + raise argparse.ArgumentTypeError(str(exc)) from exc + + +def client_is_allowed( + address: str, + networks: Iterable[ipaddress.IPv4Network | ipaddress.IPv6Network], +) -> bool: + """Return whether an address belongs to any configured network.""" + try: + client = ipaddress.ip_address(address) + except ValueError: + return False + return any(client in network for network in networks) + + +class MirrorProxyHandler(BaseHTTPRequestHandler): + """Proxy GET and HEAD requests to a fixed, path-selected upstream.""" + + protocol_version = "HTTP/1.0" + server_version = "ReMeMirrorProxy/1.0" + + @property + def config(self) -> ProxyConfig: + """Return the server-wide immutable proxy configuration.""" + return self.server.proxy_config # type: ignore[attr-defined,no-any-return] + + def do_GET(self) -> None: # noqa: N802 - BaseHTTPRequestHandler API + """Proxy a GET request.""" + self._handle_request(send_body=True) + + def do_HEAD(self) -> None: # noqa: N802 - BaseHTTPRequestHandler API + """Proxy a HEAD request.""" + self._handle_request(send_body=False) + + def do_POST(self) -> None: # noqa: N802 - BaseHTTPRequestHandler API + """Reject unsupported methods.""" + self.send_error(HTTPStatus.METHOD_NOT_ALLOWED, "Only GET and HEAD are supported") + + do_PUT = do_POST + do_PATCH = do_POST + do_DELETE = do_POST + do_CONNECT = do_POST + + def _handle_request(self, *, send_body: bool) -> None: + started = time.monotonic() + if not client_is_allowed(self.client_address[0], self.config.allowed_clients): + logging.warning("rejected client address=%s path=%s", self.client_address[0], self.path) + self.send_error(HTTPStatus.FORBIDDEN, "Client address is not allowed") + return + + parsed = urlsplit(self.path) + if parsed.path == "/healthz": + self._send_text(HTTPStatus.OK, "ok\n", send_body=send_body) + return + if parsed.path == "/routes": + body = "".join(f"/{name} -> {upstream}\n" for name, upstream in sorted(self.config.routes.items())) + self._send_text(HTTPStatus.OK, body, send_body=send_body) + return + + route_name, separator, suffix = parsed.path.lstrip("/").partition("/") + upstream = self.config.routes.get(route_name) + if not separator or upstream is None: + self.send_error(HTTPStatus.NOT_FOUND, "Unknown mirror route") + return + + target = f"{upstream}/{suffix.lstrip('/')}" + if parsed.query: + target = f"{target}?{parsed.query}" + + if not self.config.slots.acquire(blocking=False): + self.send_error(HTTPStatus.SERVICE_UNAVAILABLE, "Proxy concurrency limit reached") + return + try: + self._proxy(target, send_body=send_body, started=started) + finally: + self.config.slots.release() + + def _proxy(self, target: str, *, send_body: bool, started: float) -> None: + headers = {name: value for name, value in self.headers.items() if name.lower() in REQUEST_HEADERS} + headers.setdefault("User-Agent", "ReMe mirror proxy") + request = urllib.request.Request(target, headers=headers, method=self.command) + + response = None + try: + # HTTPError is also a readable response and is handled by the same + # cleanup block below, so a single with statement is not suitable. + # pylint: disable=consider-using-with + response = urllib.request.urlopen( # noqa: S310 - targets are operator-configured + request, + timeout=self.config.timeout, + context=ssl.create_default_context(), + ) + except urllib.error.HTTPError as exc: + response = exc + except (urllib.error.URLError, TimeoutError, OSError) as exc: + logging.error("upstream request failed target=%s error=%s", target, exc) + self.send_error( + HTTPStatus.BAD_GATEWAY, + f"Upstream request failed: {exc.reason if isinstance(exc, urllib.error.URLError) else exc}", + ) + return + + try: + self.send_response(response.status) + for name, value in response.headers.items(): + if name.lower() not in RESPONSE_HEADERS_TO_SKIP: + self.send_header(name, value) + self.send_header("Connection", "close") + self.end_headers() + + transferred = 0 + if send_body: + while chunk := response.read(BUFFER_SIZE): + self.wfile.write(chunk) + transferred += len(chunk) + elapsed = time.monotonic() - started + logging.info( + "%s %s -> %s status=%s bytes=%s elapsed=%.3fs", + self.client_address[0], + self.path, + target, + response.status, + transferred, + elapsed, + ) + except (BrokenPipeError, ConnectionResetError): + logging.warning("client disconnected path=%s target=%s", self.path, target) + finally: + response.close() + + def _send_text(self, status: HTTPStatus, body: str, *, send_body: bool) -> None: + encoded = body.encode() + self.send_response(status) + self.send_header("Content-Type", "text/plain; charset=utf-8") + self.send_header("Content-Length", str(len(encoded))) + self.send_header("Connection", "close") + self.end_headers() + if send_body: + self.wfile.write(encoded) + + def log_message(self, format_string: str, *args: object) -> None: + """Send the built-in access message to debug logging.""" + logging.debug("%s - %s", self.client_address[0], format_string % args) + + +class MirrorProxyServer(ThreadingHTTPServer): + """Threading server carrying immutable proxy configuration.""" + + daemon_threads = True + allow_reuse_address = True + + def __init__(self, address: tuple[str, int], config: ProxyConfig): + self.proxy_config = config + super().__init__(address, MirrorProxyHandler) + + +def build_parser() -> argparse.ArgumentParser: + """Build the command-line parser.""" + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--bind", default=DEFAULT_BIND, help=f"listen address (default: {DEFAULT_BIND})") + parser.add_argument( + "--port", + type=int, + default=DEFAULT_PORT, + help=f"listen port (default: {DEFAULT_PORT})", + ) + parser.add_argument( + "--route", + action="append", + default=[], + type=parse_route, + metavar="NAME=URL", + help="add or replace a route; may be repeated", + ) + parser.add_argument( + "--allow", + action="append", + default=[], + type=parse_network, + metavar="IP_OR_CIDR", + help="allow a client address/network in addition to loopback; may be repeated", + ) + parser.add_argument( + "--timeout", + type=float, + default=DEFAULT_TIMEOUT, + help="upstream timeout in seconds", + ) + parser.add_argument( + "--max-concurrency", + type=int, + default=DEFAULT_MAX_CONCURRENCY, + help=f"maximum simultaneous upstream requests (default: {DEFAULT_MAX_CONCURRENCY})", + ) + parser.add_argument("--verbose", action="store_true", help="enable debug logging") + return parser + + +def main() -> int: + """Run the mirror proxy until SIGINT or SIGTERM.""" + args = build_parser().parse_args() + if not 1 <= args.port <= 65535: + raise SystemExit("--port must be between 1 and 65535") + if args.timeout <= 0: + raise SystemExit("--timeout must be positive") + if args.max_concurrency < 1: + raise SystemExit("--max-concurrency must be at least 1") + + routes = dict(DEFAULT_ROUTES) + routes.update(args.route) + loopback_networks = [parse_network("127.0.0.0/8"), parse_network("::1/128")] + config = ProxyConfig( + routes=routes, + allowed_clients=tuple(loopback_networks + args.allow), + timeout=args.timeout, + slots=threading.BoundedSemaphore(args.max_concurrency), + ) + logging.basicConfig( + level=logging.DEBUG if args.verbose else logging.INFO, + format="%(asctime)s | %(levelname)s | %(message)s", + ) + + server = MirrorProxyServer((args.bind, args.port), config) + for signum in (signal.SIGINT, signal.SIGTERM): + signal.signal( + signum, + lambda _signum, _frame: threading.Thread(target=server.shutdown).start(), + ) + + logging.info("listening on http://%s:%s routes=%s", args.bind, args.port, sorted(routes)) + logging.info("allowed clients=%s", [str(network) for network in config.allowed_clients]) + try: + server.serve_forever() + finally: + server.server_close() + logging.info("stopped") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/unit/test_auto_fin.py b/tests/unit/test_auto_fin.py index 9a19505a..53bb2f4e 100644 --- a/tests/unit/test_auto_fin.py +++ b/tests/unit/test_auto_fin.py @@ -1,43 +1,30 @@ -"""Focused tests for the four-step Auto Fin workflow.""" +"""Focused tests for the tool-free Auto Fin workflow.""" # pylint: disable=missing-function-docstring,protected-access import hashlib -import json from datetime import date, datetime from pathlib import Path -from types import SimpleNamespace import pytest from reme.components import ApplicationContext from reme.components.agent_wrapper.base_agent_wrapper import BaseAgentWrapper from reme.components.runtime_context import RuntimeContext -from reme.config.config_parser import _load_config from reme.schema import ( - AutoFinEtfHistoricalEvents, - AutoFinEtfHistoricalResearch, - AutoFinEtfSelection, AutoFinEtfsOutput, - AutoFinHistoricalEvent, - AutoFinHistoricalEventReference, - AutoFinMarketSelection, - AutoFinMarketSample, + AutoFinHistoricalOutput, AutoFinReportOutput, + Response, ) -from reme.steps.cookbook.auto_fin._base import _write +from reme.steps.cookbook.auto_fin._base import _plain_text, _write from reme.steps.cookbook.auto_fin.data import AutoFinDataStep from reme.steps.cookbook.auto_fin.history import AutoFinHistoryStep -from reme.steps.cookbook.auto_fin.history_search import AutoFinHistorySearchStep from reme.steps.cookbook.auto_fin.merge import AutoFinMergeStep -from reme.steps.cookbook.auto_fin.market import AutoFinMarketStep -from reme.steps.cookbook.auto_fin.topic import AutoFinTopicStep, _plain_text +from reme.steps.cookbook.auto_fin.topic import AutoFinTopicStep -def test_atomic_write_preserves_existing_file_and_cleans_temporary_file_on_failure( - tmp_path: Path, - monkeypatch, -): +def test_atomic_write_preserves_existing_file_on_failure(tmp_path: Path, monkeypatch): path = tmp_path / "result.json" path.write_text("existing", encoding="utf-8") @@ -45,440 +32,146 @@ def test_atomic_write_preserves_existing_file_and_cleans_temporary_file_on_failu raise OSError("replace failed") monkeypatch.setattr("reme.steps.cookbook.auto_fin._base.os.replace", fail_replace) - with pytest.raises(OSError, match="replace failed"): _write(path, "replacement") - assert path.read_text(encoding="utf-8") == "existing" assert not list(tmp_path.glob(".*.tmp")) -@pytest.mark.asyncio -async def test_read_jsonl_preserves_unicode_line_separator(tmp_path: Path): - path = tmp_path / "news.jsonl" - rows = [{"title": "包含\u2028行分隔符"}, {"title": "下一条"}] - path.write_text( - "".join(json.dumps(row, ensure_ascii=False) + "\n" for row in rows), - encoding="utf-8", - ) - - assert await AutoFinDataStep._read_jsonl(path) == rows - - -def test_plain_news_text_removes_markup_images_and_hidden_content(): - content = '

甲&乙

' - - assert _plain_text(content) == "甲&乙 丙" - - -def test_topic_repairs_unknown_news_id_by_unique_content_hash(): - output = AutoFinEtfsOutput.model_validate( - { - "etfs": [ - { - "etf_code": "159819.SZ", - "etf_name": "人工智能ETF", - "events": [{"reason": "穆迪警告AI投资风险", "news_id": "20260725061826_1c76"}], - }, - ], - }, - ) - news = [ - {"news_id": "20260725050638_1c76"}, - {"news_id": "20260725061826_765a"}, - ] - - repaired, repairs = AutoFinTopicStep._repair_news_ids(output, news) - - assert repaired.etfs[0].events[0].news_id == "20260725050638_1c76" - assert repairs == {"20260725061826_1c76": "20260725050638_1c76"} - - -def test_topic_does_not_repair_ambiguous_content_hash(): - output = AutoFinEtfsOutput.model_validate( - { - "etfs": [ - { - "etf_code": "159819.SZ", - "etf_name": "人工智能ETF", - "events": [{"reason": "相关事件", "news_id": "20260725061826_abcd"}], - }, - ], - }, - ) - news = [{"news_id": "20260725050638_abcd"}, {"news_id": "20260725070000_abcd"}] - - repaired, repairs = AutoFinTopicStep._repair_news_ids(output, news) - - assert repaired.etfs[0].events[0].news_id == "20260725061826_abcd" - assert not repairs - - -def test_topic_normalizes_duplicate_and_invalid_selections(): - output = AutoFinEtfsOutput.model_validate( - { - "etfs": [ - { - "etf_code": "159516.SZ", - "etf_name": "半导体设备ETF", - "events": [ - {"reason": "较晚新闻", "news_id": "20260724164534_9368"}, - {"reason": "较早新闻", "news_id": "20260724164043_7332"}, - {"reason": "重复新闻", "news_id": "20260724164043_7332"}, - {"reason": "未知新闻", "news_id": "20260724170000_ffff"}, - {"reason": " ", "news_id": "20260724170000_abcd"}, - ], - }, - { - "etf_code": "159516.SZ", - "etf_name": "模型返回的错误名称", - "events": [{"reason": "另一条新闻", "news_id": "20260724170000_abcd"}], - }, - { - "etf_code": "000000.SZ", - "etf_name": "候选范围外ETF", - "events": [{"reason": "范围外", "news_id": "20260724164043_7332"}], - }, - ], - }, - ) - news = [ - {"news_id": "20260724164043_7332"}, - {"news_id": "20260724164534_9368"}, - {"news_id": "20260724170000_abcd"}, - ] - etfs = [{"code": "159516.SZ", "name": "国泰中证半导体材料设备主题ETF"}] - - normalized, changed = AutoFinTopicStep._normalize_selection(output, news, etfs) - - assert changed - assert len(normalized.etfs) == 1 - assert normalized.etfs[0].etf_name == "国泰中证半导体材料设备主题ETF" - assert [event.news_id for event in normalized.etfs[0].events] == [ - "20260724164043_7332", - "20260724164534_9368", - "20260724170000_abcd", - ] - - -def test_topic_limits_normalized_output_in_code(): - output = AutoFinEtfsOutput.model_validate( - { - "etfs": [ - { - "etf_code": f"{index:06d}.SZ", - "etf_name": f"ETF {index}", - "events": [{"reason": "相关事件", "news_id": "20260724164043_7332"}], - } - for index in range(21) - ], - }, - ) - news = [{"news_id": "20260724164043_7332"}] - etfs = [{"code": f"{index:06d}.SZ", "name": f"ETF {index}"} for index in range(21)] - - normalized, changed = AutoFinTopicStep._normalize_selection(output, news, etfs) - - assert changed - assert len(normalized.etfs) == 20 - assert [item.etf_code for item in normalized.etfs] == [f"{index:06d}.SZ" for index in range(20)] - - -def test_published_time_is_normalized_once_to_shanghai_local_time(): - parsed = AutoFinDataStep._published_at({"published_at": "2026-07-24T01:00:00+00:00"}) - - assert parsed == datetime(2026, 7, 24, 9) - assert parsed.tzinfo is None - - -@pytest.mark.asyncio -async def test_current_news_keeps_all_items_with_per_item_and_total_content_limits(tmp_path: Path): - day_dir = tmp_path / "daily" / "2026-07-24" - day_dir.mkdir(parents=True) +def test_news_markdown_round_trip_and_plain_text(tmp_path: Path): rows = [ { - "title": f"新闻标题{index}", - "pub_time": f"2026-07-24 09:0{index}:00", - "src": "财联社", - "content": f"

正文内容{index}ABCDEFGHIJ

", - } - for index in range(3) + "news_id": "20260724070000_abcd", + "event_time": "2026-07-24T07:00:00", + "title": "黄金上涨", + "content": "避险需求增强", + }, ] - (day_dir / "auto_fin_news_data.jsonl").write_text( - "".join(json.dumps(row, ensure_ascii=False) + "\n" for row in rows), - encoding="utf-8", - ) - step = AutoFinTopicStep(app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai")) - step.context = RuntimeContext( - auto_fin_news_start="2026-07-24", - auto_fin_date="2026-07-24", - news_title_max_chars=5, - news_content_max_chars=10, - news_total_content_max_chars=12, - ) + path = tmp_path / "news.md" + path.write_text(AutoFinDataStep._render_news(date(2026, 7, 24), rows), encoding="utf-8") - news = await step._current_news( - datetime.fromisoformat("2026-07-24T08:59:00"), - datetime.fromisoformat("2026-07-24T10:00:00"), - ) - - assert len(news) == 3 - assert all(len(row["title"]) <= 5 for row in news) - assert all(len(row["content"]) <= 4 for row in news) - assert sum(len(row["content"]) for row in news) <= 12 - assert all("<" not in row["content"] for row in news) - assert {row["news_id"] for row in news} == { - f"20260724090{index}00_" f"{hashlib.sha256(f'财联社

正文内容{index}ABCDEFGHIJ

'.encode()).hexdigest()[:4]}" - for index in range(3) - } + assert AutoFinDataStep.read_news(path) == rows + assert _plain_text("

甲&乙

") == "甲&乙 丙" @pytest.mark.asyncio -async def test_data_fills_missing_news_refreshes_today_and_force_refreshes_history( - tmp_path: Path, -): +async def test_non_trade_day_stops_after_trade_calendar(tmp_path: Path): calls = [] - def provider(endpoint: str, **kwargs): + def provider(endpoint, **kwargs): calls.append((endpoint, kwargs)) - if endpoint == "major_news": - day = kwargs["start_date"][:10] - return [ - { - "title": day, - "pub_time": f"{day} 08:00:00", - "src": "财联社", - "content": day, - }, - ] - raise AssertionError(endpoint) + return [{"cal_date": "20260725", "is_open": 0}] - app_context = ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai") - context = RuntimeContext( - date="2026-07-24", - now="2026-07-24T09:30:00+08:00", - lookback_days=3, - progress_interval=30, - trade_dates=["2026-07-23"], - tushare_provider=provider, - ) - - await AutoFinDataStep(app_context=app_context)(context) - assert [endpoint for endpoint, _ in calls] == ["major_news"] * 3 - for day in ("2026-07-22", "2026-07-23", "2026-07-24"): - [row] = AutoFinDataStep._read_jsonl_sync( - tmp_path / "daily" / day / "auto_fin_news_data.jsonl", - ) - news_hash = hashlib.sha256(f"财联社{day}".encode()).hexdigest()[:4] - assert row["news_id"] == f"{day.replace('-', '')}080000_{news_hash}" - - calls.clear() - await AutoFinDataStep(app_context=app_context)(context) - assert [endpoint for endpoint, _ in calls] == ["major_news"] - assert calls[0][1]["start_date"] == "2026-07-24 00:00:00" - assert calls[0][1]["end_date"] == "2026-07-24 09:30:00" - - calls.clear() - context["force"] = True - await AutoFinDataStep(app_context=app_context)(context) - assert [endpoint for endpoint, _ in calls] == ["major_news"] * 3 - assert context["auto_fin_previous_trade_date"] == "2026-07-23" - - -@pytest.mark.asyncio -async def test_data_deduplicates_by_publish_time_and_short_hash(tmp_path: Path): - def provider(endpoint: str, **_kwargs): - assert endpoint == "major_news" - return [ - { - "title": "重复新闻的较晚记录", - "pub_time": "2026-07-24 09:00:00", - "src": "财联社", - "content": "相同正文", - }, - { - "title": "另一条新闻", - "pub_time": "2026-07-24 08:00:00", - "src": "财联社", - "content": "另一篇正文", - }, - { - "title": "重复新闻的最早记录", - "pub_time": "2026-07-24 07:00:00", - "src": "财联社", - "content": "相同正文", - }, - { - "title": "同一时间和正文的重复记录", - "pub_time": "2026-07-24 07:00:00", - "src": "财联社", - "content": "相同正文", - }, - ] - - context = RuntimeContext( - date="2026-07-24", - now="2026-07-24T09:30:00+08:00", - lookback_days=1, - progress_interval=30, - trade_dates=["2026-07-23"], - tushare_provider=provider, - ) - await AutoFinDataStep( + context = RuntimeContext(date="2026-07-25", now="2026-07-25T09:00:00+08:00", tushare_provider=provider) + response = await AutoFinDataStep( app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai"), )(context) - rows = AutoFinDataStep._read_jsonl_sync( - tmp_path / "daily" / "2026-07-24" / "auto_fin_news_data.jsonl", - ) - assert [row["title"] for row in rows] == ["重复新闻的最早记录", "另一条新闻", "重复新闻的较晚记录"] - assert len({row["news_id"] for row in rows}) == len(rows) + assert [call[0] for call in calls] == ["trade_cal"] + assert context["auto_fin_skipped"] is True + assert response.metadata["skipped"] is True -@pytest.mark.asyncio -async def test_topic_etfs_join_rank_deduplicate_name_and_index(tmp_path: Path): - def provider(endpoint: str, **kwargs): - if endpoint == "etf_basic": - assert kwargs["list_status"] == "L" - return [ - {"ts_code": "510001.SH", "csname": "同名 ETF", "index_code": "I1", "index_name": "指数一"}, - {"ts_code": "510002.SH", "csname": "同名ETF", "index_code": "I2", "index_name": "指数二"}, - {"ts_code": "510003.SH", "csname": "另一名称", "index_code": "I1", "index_name": "指数一"}, - {"ts_code": "510004.SH", "csname": "独立ETF", "index_code": "I4", "index_name": "指数四"}, - ] - if endpoint == "fund_daily": - assert kwargs["trade_date"] == "20260723" - return [ - {"ts_code": "510002.SH", "amount": 90}, - {"ts_code": "510004.SH", "amount": 70}, - {"ts_code": "510001.SH", "amount": 100}, - {"ts_code": "510003.SH", "amount": 80}, - {"ts_code": "510001.SH", "amount": 95}, - ] - raise AssertionError(endpoint) - - step = AutoFinTopicStep(app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai")) - step.context = RuntimeContext(tushare_provider=provider, etf_candidate_limit=150) - - assert await step._filtered_etfs(date(2026, 7, 23)) == [ - {"code": "510001.SH", "name": "同名 ETF"}, - {"code": "510004.SH", "name": "独立ETF"}, +def test_adjusted_returns_use_close_before_1500_and_next_open_after_1500(): + rows = [ + {"trade_date": "20260601", "open": 9, "close": 10, "adj_factor": 1}, + {"trade_date": "20260602", "open": 10, "close": 11, "adj_factor": 1}, + {"trade_date": "20260603", "open": 11, "close": 12, "adj_factor": 1}, + {"trade_date": "20260604", "open": 12, "close": 13, "adj_factor": 1}, + {"trade_date": "20260605", "open": 13, "close": 14, "adj_factor": 1}, + {"trade_date": "20260608", "open": 14, "close": 15, "adj_factor": 1}, ] + before = AutoFinHistoryStep._returns(datetime(2026, 6, 1, 14), rows) + after = AutoFinHistoryStep._returns(datetime(2026, 6, 1, 16), rows) + + assert before == pytest.approx({"d1": 0.1, "d2": 0.2, "d3": 0.3, "d5": 0.5}) + assert after == pytest.approx({"d1": 0.1, "d2": 0.2, "d3": 0.3, "d5": 0.5}) + + +def test_returns_enter_next_session_for_before_close_non_trade_day(): + # A Saturday 14:00 event has no same-day close, so it must fall back to the + # first following trading day's open rather than yielding no returns. + rows = [ + {"trade_date": "20260605", "open": 8, "close": 9, "adj_factor": 1}, + {"trade_date": "20260608", "open": 10, "close": 11, "adj_factor": 1}, + {"trade_date": "20260609", "open": 11, "close": 12, "adj_factor": 1}, + {"trade_date": "20260610", "open": 12, "close": 13, "adj_factor": 1}, + {"trade_date": "20260611", "open": 13, "close": 14, "adj_factor": 1}, + {"trade_date": "20260612", "open": 14, "close": 15, "adj_factor": 1}, + ] + + returns = AutoFinHistoryStep._returns(datetime(2026, 6, 6, 14), rows) + + assert returns == pytest.approx({"d1": 0.1, "d2": 0.2, "d3": 0.3, "d5": 0.5}) + + +class _SearchJob: + def __init__(self, path: str, news_id: str): + self.path = path + self.news_id = news_id + self.calls = [] + + async def __call__(self, **kwargs): + self.calls.append(kwargs) + return Response( + metadata={ + "results": [ + { + "path": self.path, + "text": f"- news_id: `{self.news_id}`\n历史降息新闻", + }, + ], + }, + ) + class _Agent(BaseAgentWrapper): - """Return deterministic structured replies for every analysis stage.""" - - def __init__(self, **kwargs): + def __init__(self, current_id: str, historical_id: str, **kwargs): super().__init__(**kwargs) - self.calls = [] + self.current_id = current_id + self.historical_id = historical_id + self.schemas = [] async def reply(self, inputs, **kwargs): + assert kwargs.keys() == {"output_schema"} schema = kwargs["output_schema"] + self.schemas.append(schema) prompt = str(inputs) - self.calls.append((schema, prompt, kwargs)) - assert "resume" not in kwargs - assert "session_id" not in kwargs - task = prompt - assert "schema" not in task.casefold() - assert "```json" in task + assert "不得搜索、调用工具" in prompt + assert "```json" not in prompt if schema is AutoFinEtfsOutput: - assert "filtered_news.jsonl" in task - assert "filtered_etf.jsonl" in task - assert "最多返回 20" not in task value = { "etfs": [ { - "etf_code": "159018.SZ", - "etf_name": "油气ETF", - "events": [ - { - "reason": "供应中断直接影响油气产业链盈利预期", - "news_id": ( - f"20260723160000_" - f"{hashlib.sha256('财联社主要产油区供应中断'.encode()).hexdigest()[:4]}" - ), - }, - { - "reason": "供应恢复时间影响油气价格预期", - "news_id": ( - f"20260724090000_" - f"{hashlib.sha256('财联社供应恢复时间仍不确定'.encode()).hexdigest()[:4]}" - ), - }, - ], + "etf_code": "518880.SH", + "etf_name": "错误名称", + "events": [{"news_id": self.current_id, "reason": "降息预期利好黄金"}], }, ], } - elif schema is AutoFinEtfHistoricalEvents: - assert "memory_search" in task - assert "不查询行情" in task - assert "159018.SZ(油气ETF)" in task - current_news_id = ( - f"20260724090000_" f"{hashlib.sha256('财联社供应恢复时间仍不确定'.encode()).hexdigest()[:4]}" - ) - assert current_news_id not in task - assert "程序会回查、过滤、去重、排序并补充行情" in task - tool_context_id = kwargs.get("tool_context_id", "") - assert tool_context_id.startswith("auto_fin_history_01_159018.SZ_") - assert tool_context_id not in task - self.app_context.metadata.setdefault("tool_contexts", {})[tool_context_id] = { - "search_seen_chunk_ids": {}, - } + elif schema is AutoFinHistoricalOutput: value = { - "etf_code": "changed", - "etf_name": "changed", "historical_events": [ { - "reason": "供应中断的事件类型和传导机制相同", - "news_id": ( - f"20260601100000_" f"{hashlib.sha256('财联社历史供应中断'.encode()).hexdigest()[:4]}" - ), - "source_path": "daily/2026-06-01/auto_fin_news_data.jsonl", + "news_id": self.historical_id, + "reason": "同属利率政策变化且对黄金影响相同", + "direction": "same", }, - ], - } - elif schema is AutoFinMarketSelection: - assert "ETF:159018.SZ(油气ETF)" in task - assert "[2026-07-23T16:00:00] 原油供应中断" in task - assert "影响方向与当前事件相同还是相反" in task - assert "不要依据" in task - assert "程序会过滤、去重并完成计算" in task - assert "$tushare-data" not in task - history_path = Path( - next( - line.rsplit(":", 1)[-1].strip() - for line in task.splitlines() - if line.strip().endswith("_output.json") - ), - ) - history = json.loads(history_path.read_text(encoding="utf-8")) - assert "historical_samples" not in history - assert len(history["historical_events"]) == 1 - assert history["historical_events"][0]["event_title"] == "历史供应中断" - assert len(history["historical_events"][0]["future_returns"]) == 10 - value = { - "same_direction_events": [ { - "reason": "供应中断的事件类型和传导机制相同", - "news_id": history["historical_events"][0]["news_id"], + "news_id": "20260101000000_ffff", + "reason": "模型虚构", + "direction": "same", }, ], - "opposite_direction_events": [], } elif schema is AutoFinReportOutput: - assert "不重新搜索新闻" in task - assert "auto_fin_history_output.jsonl" in task - assert '"etf_code": "159018.SZ"' in task - assert '"suggested_holding_days": 10' in task - assert '"horizon": 1' in task - assert '"horizon": 10' in task - assert "不使用计算结果反推事件方向" in task - assert "负向或无正收益的情况可以合并" in task + assert '"d5": 0.5' in prompt + assert "上一份建议" in prompt value = { - "title": "Auto Fin ETF 结论", - "body": "## 结论\n\n推荐 159018.SZ(油气ETF),参考持有 10 个交易日," - "当前加权预估收益 +10%;核心风险:供应恢复。", + "title": "# 黄金观察", + "description": "降息事件偏利好黄金。", + "body": "## 建议\n\n关注黄金ETF。", } else: # pragma: no cover raise AssertionError(schema) @@ -486,613 +179,155 @@ class _Agent(BaseAgentWrapper): @pytest.mark.asyncio -async def test_four_step_pipeline_writes_plain_markdown_and_cleans_temporary_data( +async def test_new_pipeline_prepares_context_and_uses_three_tool_free_agents( tmp_path: Path, ): - def provider(endpoint: str, **_kwargs): + current_content = "美联储释放降息信号" + historical_content = "美联储宣布降息" + current_id = f"20260724090000_{hashlib.sha256(f'财联社{current_content}'.encode()).hexdigest()[:4]}" + historical_id = f"20260601100000_{hashlib.sha256(f'财联社{historical_content}'.encode()).hexdigest()[:4]}" + market_dates = [ + "20260601", + "20260602", + "20260603", + "20260604", + "20260605", + "20260608", + ] + + def provider(endpoint, **kwargs): + if endpoint == "trade_cal": + return [{"cal_date": "20260724", "is_open": 1}] if endpoint == "major_news": - return [ - { - "title": "原油供应中断", - "pub_time": "2026-07-23 16:00:00", - "src": "财联社", - "content": "主要产油区供应中断", - }, - { - "title": "供应恢复时间不确定", - "pub_time": "2026-07-24 09:00:00", - "src": "财联社", - "content": "供应恢复时间仍不确定", - }, - ] + day = kwargs["start_date"][:10] + if day == "2026-07-24": + return [ + { + "title": "降息信号", + "pub_time": "2026-07-24 09:00:00", + "src": "财联社", + "content": current_content, + }, + ] + return [] if endpoint == "etf_basic": - return [ - { - "ts_code": "159018.SZ", - "csname": "油气ETF", - "index_code": "930987.CSI", - "index_name": "中证油气产业指数", - "list_status": "L", - }, - ] + return [{"ts_code": "518880.SH", "csname": "黄金ETF", "list_status": "L"}] if endpoint == "fund_daily": - if "trade_date" in _kwargs: - return [{"ts_code": "159018.SZ", "trade_date": "20260723", "amount": 1000}] - trade_dates = [ - "20260601", - "20260602", - "20260603", - "20260604", - "20260605", - "20260608", - "20260609", - "20260610", - "20260611", - "20260612", - "20260615", - ] return [ { - "ts_code": "159018.SZ", + "ts_code": "518880.SH", "trade_date": trade_date, - "open": 0.995 + index / 100, - "close": 1.0 + index / 100, + "open": 9 + index, + "close": 10 + index, + "pct_chg": 1, } - for index, trade_date in enumerate(trade_dates) + for index, trade_date in enumerate(market_dates) ] if endpoint == "fund_adj": - return [ - {"ts_code": "159018.SZ", "trade_date": trade_date, "adj_factor": 1.0} - for trade_date in ( - "20260601", - "20260602", - "20260603", - "20260604", - "20260605", - "20260608", - "20260609", - "20260610", - "20260611", - "20260612", - "20260615", - ) - ] + return [{"trade_date": trade_date, "adj_factor": 1} for trade_date in market_dates] raise AssertionError(endpoint) + historical_path = tmp_path / "daily" / "2026-06-01" / "auto_fin_news.md" + historical_path.parent.mkdir(parents=True) + historical_path.write_text( + AutoFinDataStep._render_news( + date(2026, 6, 1), + [ + { + "news_id": historical_id, + "event_time": "2026-06-01T10:00:00", + "title": "历史降息", + "content": historical_content, + }, + ], + ), + encoding="utf-8", + ) + previous = tmp_path / "daily" / "2026-07-23" / "auto_fin.md" + previous.parent.mkdir(parents=True) + previous.write_text("# 上一份建议\n", encoding="utf-8") + app_context = ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai") - agent = _Agent(app_context=app_context) + search = _SearchJob("daily/2026-06-01/auto_fin_news.md", historical_id) + app_context.jobs["memory_search"] = search + agent = _Agent(current_id, historical_id, app_context=app_context) context = RuntimeContext( date="2026-07-24", now="2026-07-24T09:30:00+08:00", - lookback_days=2, - progress_interval=30, - trade_dates=["2026-07-23"], tushare_provider=provider, + etf_codes=["518880.SH"], + news_lookback_days=1, + historical_search_limit=7, ) await AutoFinDataStep(app_context=app_context)(context) - historical_path = tmp_path / "daily" / "2026-06-01" / "auto_fin_news_data.jsonl" - historical_path.parent.mkdir(parents=True) - historical_content = "历史供应中断" - historical_news_id = f"20260601100000_" f"{hashlib.sha256(f'财联社{historical_content}'.encode()).hexdigest()[:4]}" - historical_path.write_text( - json.dumps( - { - "title": "历史供应中断", - "pub_time": "2026-06-01 10:00:00", - "src": "财联社", - "content": historical_content, - "news_id": historical_news_id, - }, - ensure_ascii=False, - ) - + "\n", - encoding="utf-8", - ) - logs = [] - topic_step = AutoFinTopicStep(app_context=app_context, agent_wrapper=agent) - history_step = AutoFinHistoryStep(app_context=app_context, agent_wrapper=agent) - merge_step = AutoFinMergeStep(app_context=app_context, agent_wrapper=agent) - for step in (topic_step, history_step, merge_step): - step.logger = SimpleNamespace(info=logs.append, debug=lambda _message: None) - await topic_step(context) - await history_step(context) - response = await merge_step(context) + await AutoFinTopicStep(app_context=app_context, agent_wrapper=agent)(context) + await AutoFinHistoryStep(app_context=app_context, agent_wrapper=agent)(context) + response = await AutoFinMergeStep(app_context=app_context, agent_wrapper=agent)(context) - assert [schema for schema, _, _ in agent.calls] == [ + assert agent.schemas == [ AutoFinEtfsOutput, - AutoFinEtfHistoricalEvents, - AutoFinMarketSelection, + AutoFinHistoricalOutput, AutoFinReportOutput, ] - assert "tool_contexts" not in app_context.metadata - report = (tmp_path / "daily" / "2026-07-24" / "auto_fin.md").read_text(encoding="utf-8") - assert report.startswith("# Auto Fin ETF 结论\n\n") - assert not report.startswith("---") - detail = context["auto_fin_history_details"][0] - assert detail["etf"]["etf_code"] == "159018.SZ" - first_event = detail["etf"]["events"][0] - assert first_event["reason"] == "供应中断直接影响油气产业链盈利预期" - assert first_event["news_id"].startswith("20260723160000_") - analysis = detail["market_analysis"] - assert analysis["matched_historical_events"][0]["weight"] == 1.0 - assert analysis["matched_historical_events"][0]["news_id"] == historical_news_id - assert analysis["matched_historical_events"][0]["direction"] == "same" - assert analysis["forecast"]["suggested_holding_days"] == 10 - assert analysis["forecast"]["returns"][-1]["expected_return"] == pytest.approx(0.1) - assert "calculation_code" not in analysis - assert detail["historical_research"]["historical_events"][0]["event_content"] == "历史供应中断" - assert detail["historical_research"]["historical_events"][0]["reason"] == "供应中断的事件类型和传导机制相同" - assert "当前加权预估收益 +10%" in report - assert "历史事件" not in report - assert not (tmp_path / "daily" / "2026-07-24" / "auto_fin_brief.md").exists() - assert response.answer.startswith("## 结论") - assert response.metadata["etf_count"] == 1 - assert context["markdown_path"] == "daily/2026-07-24/auto_fin.md" - assert context["auto_fin_digest_path"] == "daily/2026-07-24/auto_fin.md" + assert search.calls[0]["end_date"] == "2026-07-23" + assert search.calls[0]["limit"] == 7 + assert context["auto_fin_etfs"][0]["etf_name"] == "黄金ETF" + historical = context["auto_fin_analyses"][0]["events"][0]["historical_events"] + assert [event["news_id"] for event in historical] == [historical_id] + assert historical[0]["returns"]["d5"] == pytest.approx(0.5) assert response.metadata["digest_path"] == "daily/2026-07-24/auto_fin.md" - daily_index = (tmp_path / "daily" / "2026-07-24.md").read_text(encoding="utf-8") - assert "[[daily/2026-07-24/auto_fin.md]]" in daily_index - assert sum("agent input prompt=" in line for line in logs) == 2 - assert sum("agent output prompt=" in line for line in logs) == 2 - assert all("agent start prompt=" not in line and "agent done prompt=" not in line for line in logs) - assert any('query="你只负责筛选与当前新闻直接相关' in line for line in logs) - assert any('output={"etfs":[{"etf_code":"159018.SZ"' in line for line in logs) - resource_dir = tmp_path / "resource" / "2026-07-24" - filtered_news = AutoFinDataStep._read_jsonl_sync(resource_dir / "filtered_news.jsonl") - filtered_etfs = AutoFinDataStep._read_jsonl_sync(resource_dir / "filtered_etf.jsonl") - topic_etfs = AutoFinDataStep._read_jsonl_sync(resource_dir / "auto_fin_topic_output.jsonl") - history_output = json.loads( - (resource_dir / "auto_fin_history_01_159018.SZ_output.json").read_text(encoding="utf-8"), - ) - history_details = AutoFinDataStep._read_jsonl_sync(resource_dir / "auto_fin_history_output.jsonl") - assert not list(resource_dir.glob("*_input.md")) - assert [row["news_id"] for row in filtered_news] == [event["news_id"] for event in detail["etf"]["events"]] - assert filtered_etfs == [{"code": "159018.SZ", "name": "油气ETF"}] - assert topic_etfs == [detail["etf"]] - historical_event = history_output["historical_events"][0] - assert historical_event["news_id"] == historical_news_id - assert historical_event["event_time"] == "2026-06-01T10:00:00" - assert historical_event["event_title"] == "历史供应中断" - assert historical_event["market_entry"]["price_type"] == "close" - assert [point["horizon"] for point in historical_event["future_returns"]] == list( - range(1, 11), - ) - assert historical_event["future_returns"][-1]["cumulative_return"] == pytest.approx(0.1) - assert '\n "etf_code"' in (resource_dir / "auto_fin_history_01_159018.SZ_output.json").read_text(encoding="utf-8") - assert history_details == context["auto_fin_history_details"] + report = (tmp_path / "daily" / "2026-07-24" / "auto_fin.md").read_text(encoding="utf-8") + assert report.startswith("# 黄金观察\n\n> 降息事件偏利好黄金。") -def test_historical_market_sample_rejects_look_ahead_and_incorrect_adjusted_returns(): - sample = { - "event_time": "2026-06-01T10:00:00", - "entry": { - "entry_time": "2026-06-01T15:00:00", - "trade_date": "2026-06-01", - "price_type": "close", - "raw_price": 1.0, - "adj_factor": 1.2, - }, - "future_returns": [ - { - "horizon": 1, - "trade_date": "2026-06-02", - "raw_close": 1.1, - "adj_factor": 1.2, - "cumulative_return": 0.1, - }, - ], - "reaction_summary": "第一个有效收盘点上涨。", - } - - assert AutoFinMarketSample.model_validate(sample).future_returns[0].cumulative_return == pytest.approx(0.1) - - sample["event_time"] = "2026-06-01T16:00:00" - with pytest.raises(ValueError, match="entry must be strictly after the event"): - AutoFinMarketSample.model_validate(sample) - - sample["event_time"] = "2026-06-01T10:00:00" - sample["future_returns"][0]["cumulative_return"] = 0.2 - with pytest.raises(ValueError, match="incorrect adjusted return"): - AutoFinMarketSample.model_validate(sample) - - -def test_market_calculation_equal_weights_and_reverses_opposite_direction_event(): - item = AutoFinEtfSelection.model_validate( - { - "etf_code": "518880.SH", - "etf_name": "黄金ETF", - "events": [{"reason": "黄金涨价", "news_id": "20260724090000_abcd"}], - }, - ) - history = AutoFinEtfHistoricalResearch.model_validate( - { - "etf_code": "518880.SH", - "etf_name": "黄金ETF", - "historical_events": [ - { - "reason": "黄金价格方向相反", - "news_id": "20260601100000_abcd", - "source_path": "daily/2026-06-01/auto_fin_news_data.jsonl", - "event_time": "2026-06-01T10:00:00", - "event_title": "黄金价格下跌", - "event_content": "黄金价格出现明显下跌。", - "market_entry": { - "entry_time": "2026-06-01T15:00:00", - "trade_date": "2026-06-01", - "price_type": "close", - "raw_price": 1.0, - "adj_factor": 1.0, - }, - "future_returns": [ - { - "horizon": 1, - "trade_date": "2026-06-02", - "raw_close": 1.1, - "adj_factor": 1.0, - "cumulative_return": 0.1, - }, - ], - }, - { - "reason": "黄金价格方向相同", - "news_id": "20260602100000_efgh", - "source_path": "daily/2026-06-02/auto_fin_news_data.jsonl", - "event_time": "2026-06-02T10:00:00", - "event_title": "黄金价格上涨", - "event_content": "黄金价格出现明显上涨。", - "market_entry": { - "entry_time": "2026-06-02T15:00:00", - "trade_date": "2026-06-02", - "price_type": "close", - "raw_price": 1.0, - "adj_factor": 1.0, - }, - "future_returns": [ - { - "horizon": 1, - "trade_date": "2026-06-03", - "raw_close": 1.3, - "adj_factor": 1.0, - "cumulative_return": 0.3, - }, - ], - }, - ], - }, - ) - selection = AutoFinMarketSelection.model_validate( - { - "same_direction_events": [ - { - "reason": "机制和价格方向均相同", - "news_id": "20260602100000_efgh", - }, - ], - "opposite_direction_events": [ - { - "reason": "机制可比但价格方向相反", - "news_id": "20260601100000_abcd", - }, - ], - }, - ) - - analysis = AutoFinMarketStep._calculate_analysis(item, history, selection) - - assert [match.direction for match in analysis.matched_historical_events] == ["same", "opposite"] - assert [match.weight for match in analysis.matched_historical_events] == [0.5, 0.5] - assert analysis.forecast.returns[0].expected_return == pytest.approx(0.1) - assert analysis.forecast.suggested_holding_days == 1 - assert "相似历史样本的收益方向存在分歧" in analysis.limitations - - -def test_market_selection_filters_duplicate_unknown_and_blank_references(): - duplicate = {"reason": "方向判断", "news_id": "20260601100000_abcd"} - selection = AutoFinMarketSelection.model_validate( - { - "same_direction_events": [ - duplicate, - {"reason": " ", "news_id": "20260602100000_efgh"}, - {"reason": "不存在", "news_id": "20260603100000_dead"}, - ], - "opposite_direction_events": [duplicate], - "ignored_extra_field": True, - }, - ) - history = AutoFinEtfHistoricalResearch.model_validate( - { - "etf_code": "518880.SH", - "etf_name": "黄金ETF", - "historical_events": [ - { - "reason": "历史事件", - "news_id": "20260601100000_abcd", - "source_path": "daily/2026-06-01/auto_fin_news_data.jsonl", - "event_time": "2026-06-01T10:00:00", - "event_title": "黄金上涨", - "event_content": "黄金价格上涨。", - }, - { - "reason": "历史事件", - "news_id": "20260602100000_efgh", - "source_path": "daily/2026-06-02/auto_fin_news_data.jsonl", - "event_time": "2026-06-02T10:00:00", - "event_title": "美元变化", - "event_content": "美元发生变化。", - }, - ], - }, - ) - - normalized, changed = AutoFinMarketStep._normalize_selection(selection, history) - - assert changed - assert [event.news_id for event in normalized.same_direction_events] == ["20260601100000_abcd"] - assert not normalized.opposite_direction_events - - -def test_merge_normalizes_cosmetic_or_empty_report_fields(): - normalized = AutoFinMergeStep._normalize_report( - AutoFinReportOutput(title="## Auto Fin ETF 结论 ", body="# 重复标题\n\n## 结论\n\n观望。"), - ) - fallback = AutoFinMergeStep._normalize_report(AutoFinReportOutput.model_validate({})) - - assert normalized.title == "Auto Fin ETF 结论" - assert normalized.body == "## 结论\n\n观望。" - assert fallback.title == "Auto Fin ETF 结论" - assert fallback.body == "## 结论\n\n暂无可用结论。" - - -@pytest.mark.asyncio -async def test_market_skips_agent_when_all_historical_references_were_invalid(tmp_path: Path): - class _UnexpectedAgent(BaseAgentWrapper): - async def reply(self, inputs, **kwargs): - raise AssertionError((inputs, kwargs)) - - app_context = ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai") - step = AutoFinMarketStep(app_context=app_context, agent_wrapper=_UnexpectedAgent(app_context=app_context)) - step.logger = SimpleNamespace(warning=lambda _message: None) - step.context = RuntimeContext( - auto_fin_current_etf={ - "etf_code": "159819.SZ", - "etf_name": "人工智能ETF", - "events": [{"reason": "AI事件", "news_id": "20260725050638_1c76"}], - }, - auto_fin_current_events=[ - { - "reason": "AI事件", - "news_id": "20260725050638_1c76", - "event_time": "2026-07-25T05:06:38", - "event_title": "当前新闻", - "event_content": "当前内容", - }, - ], - auto_fin_current_history={ - "etf_code": "159819.SZ", - "etf_name": "人工智能ETF", - "historical_events": [], - "limitations": ["跳过无法解析的历史新闻 20260427100000_dead"], - }, - auto_fin_current_history_resource=str(tmp_path / "history.json"), - auto_fin_current_index=1, - auto_fin_date="2026-07-26", - auto_fin_decision_at="2026-07-26T18:00:00", - ) - - await step.execute() - - analysis = step.context["auto_fin_current_analysis"] - assert not analysis["matched_historical_events"] - assert "没有匹配的历史事件" in analysis["limitations"] - assert (tmp_path / "resource" / "2026-07-26" / "auto_fin_market_01_159819.SZ_output.json").is_file() - - -@pytest.mark.asyncio -async def test_history_search_calculates_adjusted_returns_for_event_time_boundaries(tmp_path: Path): - calls = [] - - def provider(endpoint: str, **kwargs): - calls.append((endpoint, kwargs)) - if endpoint == "fund_daily": - return [ - {"trade_date": "20260609", "open": 7.5, "close": 8.0}, - {"trade_date": "20260608", "open": 6.0, "close": 7.0}, - {"trade_date": "20260605", "open": 10.0, "close": 11.0}, - ] - if endpoint == "fund_adj": - return [ - {"trade_date": "20260609", "adj_factor": 2.0}, - {"trade_date": "20260608", "adj_factor": 2.0}, - {"trade_date": "20260605", "adj_factor": 1.0}, - ] - raise AssertionError(endpoint) - - step = AutoFinHistorySearchStep( +def test_previous_and_current_reports_feed_merge_context(tmp_path: Path): + step = AutoFinMergeStep( app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai"), ) - step.context = RuntimeContext(tushare_provider=provider) - events = [ - AutoFinHistoricalEvent( - reason="历史行情边界测试", - news_id=f"{event_time.replace('-', '').replace(':', '').replace('T', '')}_abcd", - source_path=f"daily/{event_time[:10]}/auto_fin_news_data.jsonl", - event_time=event_time, - event_title=label, - event_content=label, - ) - for event_time, label in ( - ("2026-06-05T08:00:00", "盘前事件"), - ("2026-06-05T10:00:00", "盘中事件"), - ("2026-06-05T16:00:00", "盘后事件"), - ("2026-06-06T10:00:00", "休市日事件"), - ) - ] + run_date = date(2026, 7, 24) - samples, limitations = await step._calculate_samples( - "159018.SZ", - events, - datetime.fromisoformat("2026-06-10T09:00:00"), - ) + # Nothing on disk yet: yesterday falls back, today marks a first run. + assert step._previous_report(run_date) == "无历史推荐。" + assert "首次生成" in step._current_report(run_date) - assert [sample.entry.price_type for sample in samples if sample.entry] == ["open", "close", "open", "open"] - assert [sample.entry.trade_date.isoformat() for sample in samples if sample.entry] == [ - "2026-06-05", - "2026-06-05", - "2026-06-08", - "2026-06-08", - ] - assert samples[0].future_returns[0].cumulative_return == pytest.approx(0.1) - assert samples[1].future_returns[0].cumulative_return == pytest.approx(14 / 11 - 1) - assert samples[2].future_returns[0].cumulative_return == pytest.approx(14 / 12 - 1) - assert samples[3].future_returns[0].cumulative_return == pytest.approx(14 / 12 - 1) - assert limitations - assert [endpoint for endpoint, _ in calls] == ["fund_daily", "fund_adj"] + (tmp_path / "daily" / "2026-07-23").mkdir(parents=True) + (tmp_path / "daily" / "2026-07-23" / "auto_fin.md").write_text("# 昨日建议\n", encoding="utf-8") + step._report_path(run_date).parent.mkdir(parents=True) + step._report_path(run_date).write_text("# 今晨建议\n", encoding="utf-8") + + # A later same-day rerun sees yesterday's report and its own earlier report. + assert step._previous_report(run_date) == "# 昨日建议\n" + assert step._current_report(run_date) == "# 今晨建议\n" -@pytest.mark.asyncio -async def test_history_search_recovers_source_path_from_news_id_date(tmp_path: Path): - actual_path = tmp_path / "daily" / "2026-04-26" / "auto_fin_news_data.jsonl" - actual_path.parent.mkdir(parents=True) - actual_path.write_text( - json.dumps( - { - "news_id": "20260426221449_de86", - "pub_time": "2026-04-26 22:14:49", - "title": "PCB厂商一季度业绩增长", - "content": "AI硬件需求带动PCB订单增长。", - }, - ensure_ascii=False, - ) - + "\n", - encoding="utf-8", - ) - step = AutoFinHistorySearchStep( - app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai"), - ) - references = [ - AutoFinHistoricalEventReference( - reason="PCB订单增长的传导机制相同", - news_id="20260426221449_de86", - source_path="daily/2026-04-27/auto_fin_news_data.jsonl", - ), - ] +def test_config_has_fixed_codes_and_no_agent_tools(): + from reme.config.config_parser import _load_config - events, limitations = await step._resolve_historical_events( - references, - set(), - datetime.fromisoformat("2026-07-24T15:00:00"), - ) - - assert len(events) == 1 - assert events[0].source_path == "daily/2026-04-26/auto_fin_news_data.jsonl" - assert not limitations - - -@pytest.mark.asyncio -async def test_history_search_ignores_unsafe_source_path_and_uses_news_id_date(tmp_path: Path): - actual_path = tmp_path / "daily" / "2026-04-26" / "auto_fin_news_data.jsonl" - actual_path.parent.mkdir(parents=True) - actual_path.write_text( - json.dumps( - { - "news_id": "20260426221449_de86", - "pub_time": "2026-04-26 22:14:49", - "title": "有效历史新闻", - "content": "有效内容", - }, - ensure_ascii=False, - ) - + "\n", - encoding="utf-8", - ) - step = AutoFinHistorySearchStep( - app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai"), - ) - references = [ - AutoFinHistoricalEventReference( - reason="相似事件", - news_id="20260426221449_de86", - source_path="/tmp/not-allowed.jsonl", - ), - ] - - events, limitations = await step._resolve_historical_events( - references, - set(), - datetime.fromisoformat("2026-07-24T15:00:00"), - ) - - assert [event.news_id for event in events] == ["20260426221449_de86"] - assert events[0].source_path == "daily/2026-04-26/auto_fin_news_data.jsonl" - assert not limitations - - -@pytest.mark.asyncio -async def test_history_search_skips_one_invalid_reference_and_continues(tmp_path: Path): - actual_path = tmp_path / "daily" / "2026-04-26" / "auto_fin_news_data.jsonl" - actual_path.parent.mkdir(parents=True) - actual_path.write_text( - json.dumps( - { - "news_id": "20260426221449_de86", - "pub_time": "2026-04-26 22:14:49", - "title": "有效历史新闻", - "content": "有效内容", - }, - ensure_ascii=False, - ) - + "\n", - encoding="utf-8", - ) - step = AutoFinHistorySearchStep( - app_context=ApplicationContext(workspace_dir=str(tmp_path), timezone="Asia/Shanghai"), - ) - step.logger = SimpleNamespace(warning=lambda _message: None) - references = [ - AutoFinHistoricalEventReference( - reason="有效引用", - news_id="20260426221449_de86", - source_path="daily/2026-04-26/auto_fin_news_data.jsonl", - ), - AutoFinHistoricalEventReference( - reason="无效引用", - news_id="20260427100000_dead", - source_path="daily/2026-04-27/auto_fin_news_data.jsonl", - ), - ] - - events, limitations = await step._resolve_historical_events( - references, - set(), - datetime.fromisoformat("2026-07-24T15:00:00"), - ) - - assert [event.news_id for event in events] == ["20260426221449_de86"] - assert len(limitations) == 1 - assert "20260427100000_dead" in limitations[0] - - -def test_daily_cookbook_wires_enabled_auto_fin_steps_and_tushare_skill(): config = _load_config("daily_cookbook") - steps = config["jobs"]["auto_fin"]["steps"] - - assert [step["backend"] for step in steps] == [ - "auto_fin_data_step", - "auto_fin_topic_step", - "auto_fin_history_step", - "auto_fin_merge_step", - "dingtalk_markdown_send_step", + job = config["jobs"]["auto_fin"] + assert job["etf_codes"] == [ + "518880.SH", + "159530.SZ", + "512760.SH", ] - expected_cron_jobs = { + search_limit = job["parameters"]["properties"]["historical_search_limit"] + assert job["parameters"]["properties"]["now"]["default"] == "" + assert search_limit["default"] == 10 + assert search_limit["minimum"] == 1 + assert job["historical_search_limit"] == 10 + assert job["steps"][3]["historical_search_limit"] == 10 + crons = { "auto_fin_0930_cron": "30 9 * * *", - "auto_fin_1145_cron": "45 11 * * *", + "auto_fin_1130_cron": "30 11 * * *", "auto_fin_1800_cron": "0 18 * * *", } - for job_name, cron in expected_cron_jobs.items(): - assert config["jobs"][job_name]["cron"] == cron - assert config["jobs"][job_name]["steps"] == steps - assert steps[0]["lookback_days"] == 360 - assert steps[0]["progress_interval"] == 30 - assert steps[-1]["input_mapping"] == {"auto_fin_digest_path": "markdown_path"} - assert config["components"]["agent_wrapper"]["auto_fin"]["skills"] == ["tushare-data"] + for name, schedule in crons.items(): + assert config["jobs"][name]["cron"] == schedule + assert config["jobs"][name]["steps"] == job["steps"] + assert "auto_fin_1200_cron" not in config["jobs"] + wrapper = config["components"]["agent_wrapper"]["default"] + assert wrapper["backend"] == "agentscope" + assert wrapper["builtin_tools"] is False + assert "skills" not in wrapper and "job_tools" not in wrapper + assert "agent_wrapper" not in job["steps"][2] diff --git a/tests/unit/test_daily_paper.py b/tests/unit/test_daily_paper.py index db437144..1bc4fd3d 100644 --- a/tests/unit/test_daily_paper.py +++ b/tests/unit/test_daily_paper.py @@ -14,10 +14,9 @@ import pytest from reme.components import ApplicationContext from reme.components.agent_wrapper.base_agent_wrapper import BaseAgentWrapper -from reme.components.outbound_proxy import FixedHttpOutboundProxy from reme.components.runtime_context import RuntimeContext from reme.config.config_parser import _load_config -from reme.schema import DailyBriefOutput, PaperInfo, PaperNoteOutput, PaperSelection +from reme.schema import AnalyzedPaper, DailyPaperMarkdownOutput, PaperInfo, PaperPick, PaperPickList from reme.steps.cookbook.daily_paper import ( DailyPaperAnalyzeStep, DailyPaperCollectStep, @@ -32,7 +31,6 @@ from reme.steps.cookbook.dingtalk import send as dingtalk_send from reme.utils import arxiv as arxiv_utils from reme.utils import huggingface_papers as hf_utils from reme.utils.huggingface_papers import paper_ids_from_html, paper_info_from_payload -from reme.enumeration import ComponentEnum class _QueuedAgentWrapper(BaseAgentWrapper): @@ -86,8 +84,8 @@ def test_hf_payload_and_html_normalization(): @pytest.mark.asyncio -async def test_hf_client_uses_explicit_outbound_proxy(monkeypatch): - """The owned HTTP client uses only the explicitly supplied HTTP proxy.""" +async def test_hf_client_uses_configured_mirror(monkeypatch): + """The owned HTTP client uses HF_MIRROR_URL as its source.""" events: list[str] = [] client_kwargs: dict = {} logger = MagicMock() @@ -103,20 +101,47 @@ async def test_hf_client_uses_explicit_outbound_proxy(monkeypatch): events.append("client-close") monkeypatch.setattr(hf_utils.httpx, "AsyncClient", FakeAsyncClient) + monkeypatch.setenv("HF_MIRROR_URL", "https://hf-mirror.com/") client = hf_utils.HuggingFacePapersClient( - proxy_url="http://127.0.0.1:43124", timeout=12.0, logger=logger, ) assert client.client is None async with client: - assert client_kwargs["proxy"] == "http://127.0.0.1:43124" - assert client_kwargs["trust_env"] is False + assert client_kwargs["base_url"] == "https://hf-mirror.com" + assert "trust_env" not in client_kwargs + assert "proxy" not in client_kwargs assert events == ["client-close"] info_messages = [call.args[0] for call in logger.info.call_args_list] - assert info_messages == ["[HuggingFacePapersClient] network mode=outbound_proxy"] + assert info_messages == ["[HuggingFacePapersClient] base_url=https://hf-mirror.com"] + + +@pytest.mark.asyncio +async def test_hf_client_preserves_mirror_path_prefix(monkeypatch): + """Relative requests retain a path-prefixed HF_MIRROR_URL.""" + requests: list[httpx.Request] = [] + + async def handler(request: httpx.Request) -> httpx.Response: + requests.append(request) + return httpx.Response(200, json=[]) + + transport = httpx.MockTransport(handler) + async_client = httpx.AsyncClient + monkeypatch.setattr( + hf_utils.httpx, + "AsyncClient", + lambda **kwargs: async_client(transport=transport, **kwargs), + ) + monkeypatch.setenv("HF_MIRROR_URL", "http://relay.example:18080/hf/") + + async with hf_utils.HuggingFacePapersClient() as client: + assert await client.fetch_daily_ids("2026-07-22") == set() + + assert [str(request.url) for request in requests] == [ + "http://relay.example:18080/hf/api/daily_papers?date=2026-07-22&limit=100", + ] @pytest.mark.asyncio @@ -154,17 +179,17 @@ async def test_hf_client_logs_retry_after_http_error(monkeypatch): assert "error=ConnectTimeout detail=timed out" in warning -def test_rrf_and_memory_candidate_reserve(): - """RRF is exact and the candidate pool preserves a memory-related slot.""" - general = _paper("2607.10001", title="General model", upvotes=100) - memory = _paper("2607.10002", title="Long-term memory for agents", upvotes=1) - general.fused_score = rrf_score(1, None, rrf_k=60, weekly_weight=0.7) - memory.fused_score = rrf_score(100, None, rrf_k=60, weekly_weight=0.7) +def test_rrf_candidate_pool_has_no_topic_preference(): + """The candidate pool follows RRF without reserving slots for a topic.""" + higher = _paper("2607.10001", title="General model", upvotes=100) + lower = _paper("2607.10002", title="Long-term memory for agents", upvotes=1) + higher.fused_score = rrf_score(1, None, rrf_k=60, weekly_weight=0.7) + lower.fused_score = rrf_score(100, None, rrf_k=60, weekly_weight=0.7) - candidates = build_candidate_pool([general, memory], limit=2, memory_reserve=1) + candidates = build_candidate_pool([lower, higher], limit=1) - assert candidates == [general, memory] - assert general.fused_score == pytest.approx(1 / 61) + assert candidates == [higher] + assert higher.fused_score == pytest.approx(1 / 61) def test_history_exclusion_reads_prior_frontmatter_only(tmp_path: Path): @@ -216,7 +241,9 @@ async def test_arxiv_pdf_downloads_missing_cache_once(tmp_path: Path, monkeypatc assert await client.download("2607.10001", target) == target assert target.read_bytes() == b"%PDF-downloaded" - assert [str(request.url) for request in requests] == ["https://arxiv.org/pdf/2607.10001"] + assert [str(request.url) for request in requests] == [ + "https://arxiv.org/pdf/2607.10001", + ] info_messages = [call.args[0] for call in logger.info.call_args_list] assert any("download start arxiv_id=2607.10001" in message for message in info_messages) assert any("download done arxiv_id=2607.10001" in message for message in info_messages) @@ -224,8 +251,8 @@ async def test_arxiv_pdf_downloads_missing_cache_once(tmp_path: Path, monkeypatc @pytest.mark.asyncio -async def test_arxiv_pdf_uses_explicit_outbound_proxy(tmp_path: Path, monkeypatch): - """The owned arXiv client uses and closes its explicit HTTP proxy client.""" +async def test_arxiv_pdf_uses_configured_mirror(tmp_path: Path, monkeypatch): + """The owned arXiv client downloads from ARXIV_MIRROR_URL.""" events: list[str] = [] client_kwargs: dict = {} logger = MagicMock() @@ -264,56 +291,51 @@ async def test_arxiv_pdf_uses_explicit_outbound_proxy(tmp_path: Path, monkeypatc def stream(self, method, url): """Return the fake streaming response context.""" - assert (method, url) == ("GET", "https://arxiv.org/pdf/2607.10001") + assert (method, url) == ("GET", "https://export.arxiv.org/pdf/2607.10001") return FakeStream() monkeypatch.setattr(arxiv_utils.httpx, "AsyncClient", FakeAsyncClient) + monkeypatch.setenv("ARXIV_MIRROR_URL", "https://export.arxiv.org/") target = tmp_path / "2607.10001.pdf" async with arxiv_utils.ArxivPdfClient( - proxy_url="http://127.0.0.1:43124", timeout=12.0, logger=logger, ) as client: assert await client.download("2607.10001", target) == target assert target.read_bytes() == b"%PDF-proxied" - assert client_kwargs["proxy"] == "http://127.0.0.1:43124" - assert client_kwargs["trust_env"] is False + assert "trust_env" not in client_kwargs + assert "proxy" not in client_kwargs assert events == ["client-close"] info_messages = [call.args[0] for call in logger.info.call_args_list] - assert info_messages[0] == "[ArxivPdfClient] network mode=outbound_proxy" + assert info_messages[0] == "[ArxivPdfClient] base_url=https://export.arxiv.org" @pytest.mark.asyncio async def test_paper_clients_enforce_context_and_unambiguous_ownership(tmp_path: Path): - """Requests require context entry and injected clients cannot also receive proxy_url.""" + """Requests require context entry and injected clients remain caller-owned.""" with pytest.raises(RuntimeError, match="async context manager"): await hf_utils.HuggingFacePapersClient().fetch_daily_ids("2026-07-22") with pytest.raises(RuntimeError, match="async context manager"): - await arxiv_utils.ArxivPdfClient().download("2607.10001", tmp_path / "paper.pdf") + await arxiv_utils.ArxivPdfClient().download( + "2607.10001", + tmp_path / "paper.pdf", + ) async with httpx.AsyncClient() as raw_client: async with arxiv_utils.ArxivPdfClient(client=raw_client): pass assert raw_client.is_closed is False - with pytest.raises(ValueError, match="client and proxy_url"): - hf_utils.HuggingFacePapersClient( - client=raw_client, - proxy_url="http://127.0.0.1:43124", - ) - with pytest.raises(ValueError, match="client and proxy_url"): - arxiv_utils.ArxivPdfClient( - client=raw_client, - proxy_url="http://127.0.0.1:43124", - ) - @pytest.mark.asyncio -async def test_daily_paper_steps_forward_one_managed_proxy_endpoint(tmp_path: Path, monkeypatch): - """Collection and one shared PDF client receive the same application proxy URL.""" - paper = _paper("2607.10001", title="Managed proxy paper") +async def test_daily_paper_steps_construct_source_clients_without_proxy( + tmp_path: Path, + monkeypatch, +): + """Collection and analysis construct direct source clients.""" + papers = [_paper(f"2607.1000{index}", title=f"Managed proxy paper {index}") for index in range(1, 4)] hf_kwargs: list[dict] = [] arxiv_kwargs: list[dict] = [] @@ -331,7 +353,7 @@ async def test_daily_paper_steps_forward_one_managed_proxy_endpoint(tmp_path: Pa async def fetch_scope(self, _scope: str, _value: str): """Return one ranked paper.""" - return [paper] + return papers async def fetch_daily_ids(self, _day: str): """Return no yesterday exclusions.""" @@ -363,44 +385,27 @@ async def test_daily_paper_steps_forward_one_managed_proxy_endpoint(tmp_path: Pa lambda *_args: ("--- PAGE 1 ---\nPaper content", 1, False), ) - proxy = FixedHttpOutboundProxy(url="http://127.0.0.1:43124") - await proxy.start() app_context = ApplicationContext(workspace_dir=str(tmp_path)) - app_context.components = {ComponentEnum.OUTBOUND_PROXY: {"default": proxy}} context = RuntimeContext(date="2026-07-21") agent = _QueuedAgentWrapper( [ { - "description": "Detailed note", + "title": f"代理论文解读{index}", + "desc": "Detailed note", "body": "# Detailed reading\n\nEvidence [p. 1].", - }, + } + for index in range(1, 4) ], ) - try: - await DailyPaperCollectStep(app_context=app_context)(context) - context["daily_paper_selection"] = PaperSelection.model_validate( - { - "selection_reasoning": "Only candidate.", - "selected": [ - { - "arxiv_id": paper.arxiv_id, - "rank": 1, - "reason": "Relevant", - "memory_relevance": "low", - }, - ], - "alternates": [], - }, - ) - context["daily_paper_selected_papers"] = [paper] - await DailyPaperAnalyzeStep(app_context=app_context, agent_wrapper=agent)(context) - finally: - await proxy.close() + await DailyPaperCollectStep(app_context=app_context)(context) + context["daily_paper_selected"] = [PaperPick(arxiv_id=paper.arxiv_id, reasoning="Relevant") for paper in papers] + context["daily_paper_candidates"] = papers + await DailyPaperAnalyzeStep(app_context=app_context, agent_wrapper=agent)(context) - assert hf_kwargs[0]["proxy_url"] == "http://127.0.0.1:43124" + assert "proxy_url" not in hf_kwargs[0] assert len(arxiv_kwargs) == 1 - assert arxiv_kwargs[0]["proxy_url"] == "http://127.0.0.1:43124" + assert "proxy_url" not in arxiv_kwargs[0] def test_daily_paper_config_passes_dingtalk_environment(monkeypatch): @@ -424,6 +429,97 @@ def test_daily_paper_config_passes_dingtalk_environment(monkeypatch): } +def test_daily_paper_uses_agentscope_without_tools(): + """Daily Paper uses the shared tool-free agent.""" + config = _load_config("daily_cookbook") + wrapper = config["components"]["agent_wrapper"]["default"] + + assert wrapper == { + "backend": "agentscope", + "as_llm": "default", + "builtin_tools": False, + } + assert "agent_wrapper" not in config["jobs"]["daily_paper"]["steps"][2] + + +def test_daily_paper_topics_parameter_defaults_to_empty(): + """Topics are an optional selection preference in the public job schema.""" + topics = _load_config("daily_cookbook")["jobs"]["daily_paper"]["parameters"]["properties"]["topics"] + + assert topics == { + "type": "string", + "description": "Optional topics to prioritize when selecting papers.", + "default": "", + } + + +def test_paper_pick_list_uses_an_object_root_for_tool_output(): + """AgentScope function arguments require an object-root JSON schema.""" + schema = PaperPickList.model_json_schema() + + assert schema["type"] == "object" + assert schema["required"] == ["papers"] + assert schema["properties"]["papers"]["type"] == "array" + + +def test_daily_paper_selects_three_papers_and_bounds_pdf_context(): + """The public job has no paper-count option and bounds extracted PDF text.""" + job = _load_config("daily_cookbook")["jobs"]["daily_paper"] + + assert "top_k" not in job + assert "top_k" not in job["parameters"]["properties"] + assert job["max_pdf_pages"] == 20 + assert job["max_pdf_chars"] == 300_000 + + +@pytest.mark.asyncio +async def test_selection_retries_invalid_id_and_keeps_three_candidates(tmp_path: Path): + """Selection retries an out-of-pool id and stores three validated candidate ids.""" + candidates = [ + _paper("2607.10001", title="First paper"), + _paper("2607.10002", title="Second paper"), + _paper("2607.10003", title="Third paper"), + ] + agent = _QueuedAgentWrapper( + [ + { + "papers": [ + {"arxiv_id": "not-a-candidate", "reasoning": "Invalid"}, + {"arxiv_id": "2607.10002", "reasoning": "Second"}, + {"arxiv_id": "2607.10003", "reasoning": "Third"}, + ], + }, + { + "papers": [ + {"arxiv_id": " 2607.10001 ", "reasoning": "First"}, + {"arxiv_id": "2607.10002", "reasoning": "Second"}, + {"arxiv_id": "2607.10003", "reasoning": "Third"}, + ], + }, + ], + ) + context = RuntimeContext(topics="context engineering") + context["daily_paper_candidates"] = candidates + + response = await DailyPaperSelectStep( + app_context=ApplicationContext(workspace_dir=str(tmp_path)), + agent_wrapper=agent, + )(context) + + selected = context["daily_paper_selected"] + assert [item.arxiv_id for item in selected] == [ + "2607.10001", + "2607.10002", + "2607.10003", + ] + assert [item.reasoning for item in selected] == ["First", "Second", "Third"] + assert response.answer == "Selected 3 papers with an agent" + assert len(agent.calls) == 2 + assert "outside the candidate pool" in agent.calls[1]["inputs"] + assert "用户明确感兴趣的主题:context engineering" in agent.calls[1]["inputs"] + assert "仅将这些 topics 作为主题偏好" in agent.calls[1]["inputs"] + + def test_reme_import_does_not_require_optional_dingtalk_stream(): """Importing ReMe must not eagerly load the core-only DingTalk dependency.""" script = """ @@ -461,6 +557,8 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs( "2607.10001": _paper("2607.10001", title="Best monthly paper", upvotes=100), "2607.10002": _paper("2607.10002", title="Yesterday paper", upvotes=90), "2607.10003": _paper("2607.10003", title="Previously recommended", upvotes=80), + "2607.10004": _paper("2607.10004", title="Second eligible paper", upvotes=70), + "2607.10005": _paper("2607.10005", title="Third eligible paper", upvotes=60), } prior_dir = tmp_path / "daily" / "2026-07-19" @@ -488,7 +586,12 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs( assert value == "2026-07" return list(papers.values()) assert value == "2026-W30" - return [papers["2607.10001"], papers["2607.10002"]] + return [ + papers["2607.10001"], + papers["2607.10002"], + papers["2607.10004"], + papers["2607.10005"], + ] async def fetch_daily_ids(self, day: str): """Record and return the exact requested day.""" @@ -501,35 +604,49 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs( target.write_bytes(b"%PDF-fake") return target - def fake_extract(_self, _path: Path, _max_pages: int, _max_chars: int): + extraction_limits: list[tuple[int, int]] = [] + + def fake_extract(_self, _path: Path, max_pages: int, max_chars: int): """Return deterministic extracted text.""" + extraction_limits.append((max_pages, max_chars)) return "--- PAGE 1 ---\nPaper content", 1, False monkeypatch.setattr(collect, "HuggingFacePapersClient", _FakeHfClient) monkeypatch.setattr(analyze.ArxivPdfClient, "download", fake_download) - monkeypatch.setattr(analyze.DailyPaperAnalyzeStep, "_extract_pdf_text_sync", fake_extract) + monkeypatch.setattr( + analyze.DailyPaperAnalyzeStep, + "_extract_pdf_text_sync", + fake_extract, + ) cc_wrapper = _QueuedAgentWrapper( [ { - "selection_reasoning": "Best remaining ranked paper.", - "selected": [ - { - "arxiv_id": "2607.10001", - "rank": 1, - "reason": "Strong result", - "memory_relevance": "low", - }, + "papers": [ + {"arxiv_id": "2607.10001", "reasoning": "Strong result"}, + {"arxiv_id": "2607.10004", "reasoning": "Useful method"}, + {"arxiv_id": "2607.10005", "reasoning": "Clear evidence"}, ], - "alternates": [], }, { - "description": "Detailed note", - "body": "# Detailed reading\n\nEvidence [p. 1].", + "title": "记忆代理研究", + "desc": "Detailed note one", + "body": "Evidence one [p. 1].", }, { - "description": "Five-minute brief", - "body": "# 今日论文速读\n\n[[daily/2026-07-21/paper-2607.10001.md]]", + "title": "上下文压缩研究", + "desc": "Detailed note two", + "body": "Evidence two [p. 1].", + }, + { + "title": "持续学习研究", + "desc": "Detailed note three", + "body": "Evidence three [p. 1].", + }, + { + "title": "今日智能体论文简报", + "desc": "Five-minute brief", + "body": "# 今日一句话\n\nSummary.", }, ], ) @@ -541,71 +658,132 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs( ) context = RuntimeContext( date="2026-07-21", - top_k=1, - candidate_limit=2, - memory_reserve=0, + candidate_limit=3, ) await DailyPaperCollectStep(app_context=app_context)(context) await DailyPaperRankStep(app_context=app_context)(context) - await DailyPaperSelectStep(app_context=app_context, agent_wrapper=cc_wrapper)(context) - await DailyPaperAnalyzeStep(app_context=app_context, agent_wrapper=cc_wrapper)(context) - await DailyPaperDigestStep(app_context=app_context, agent_wrapper=cc_wrapper)(context) + await DailyPaperSelectStep(app_context=app_context, agent_wrapper=cc_wrapper)( + context, + ) + await DailyPaperAnalyzeStep(app_context=app_context, agent_wrapper=cc_wrapper)( + context, + ) + await DailyPaperDigestStep(app_context=app_context, agent_wrapper=cc_wrapper)( + context, + ) assert _FakeHfClient.requested_daily == ["2026-07-20"] - assert context.response.metadata["selected_arxiv_ids"] == ["2607.10001"] + assert context.response.metadata["selected_arxiv_ids"] == [ + "2607.10001", + "2607.10004", + "2607.10005", + ] assert context.response.metadata["excluded_yesterday_count"] == 1 assert context.response.metadata["excluded_history_count"] == 1 - note_path = tmp_path / "daily" / "2026-07-21" / "paper-2607.10001.md" - digest_path = tmp_path / "daily" / "2026-07-21" / "daily-paper-brief.md" + note_path = tmp_path / "daily" / "2026-07-21" / "记忆代理研究.md" + digest_path = tmp_path / "daily" / "2026-07-21" / "今日智能体论文简报.md" note = frontmatter.load(note_path) assert note.metadata["arxiv_id"] == "2607.10001" assert note.metadata["source_pdf"] == "[[external-assets/papers/2607.10001.pdf]]" assert (tmp_path / "external-assets" / "papers" / "2607.10001.pdf").is_file() - assert "[[daily/2026-07-21/paper-2607.10001.md]]" in digest_path.read_text( + assert note.metadata["title"] == "记忆代理研究" + assert extraction_limits == [(20, 300_000)] * 3 + assert "[[daily/2026-07-21/记忆代理研究.md]]" in digest_path.read_text( encoding="utf-8", ) assert not (tmp_path / "metadata" / "daily_paper" / "2026-07-21.json").exists() digest = frontmatter.load(digest_path) - assert digest.metadata["selection_reasoning"] == "Best remaining ranked paper." - assert digest.metadata["arxiv_ids"] == ["2607.10001"] - assert all(set(call["kwargs"]) == {"output_schema"} for call in cc_wrapper.calls) + assert digest.metadata["title"] == "今日智能体论文简报" + assert digest.metadata["selection_reasoning"] == [ + "Strong result", + "Useful method", + "Clear evidence", + ] + assert digest.metadata["arxiv_ids"] == ["2607.10001", "2607.10004", "2607.10005"] + assert cc_wrapper.calls[0]["kwargs"] == {"output_schema": PaperPickList} + assert "用户感兴趣的主题" not in cc_wrapper.calls[0]["inputs"] + assert "用户未提供明确的 topic 倾向" in cc_wrapper.calls[0]["inputs"] + assert "优先选择 fused_score 更高的论文" in cc_wrapper.calls[0]["inputs"] + assert "memory_keyword_score" not in cc_wrapper.calls[0]["inputs"] + assert "Agent 长期记忆" not in cc_wrapper.calls[0]["inputs"] + assert all(call["kwargs"] == {"output_schema": DailyPaperMarkdownOutput} for call in cc_wrapper.calls[1:]) assert [call["kwargs"]["output_schema"] for call in cc_wrapper.calls] == [ - PaperSelection, - PaperNoteOutput, - DailyBriefOutput, + PaperPickList, + DailyPaperMarkdownOutput, + DailyPaperMarkdownOutput, + DailyPaperMarkdownOutput, + DailyPaperMarkdownOutput, ] analysis_prompt = cc_wrapper.calls[1]["inputs"] - assert "长期记忆相关性初筛:low" in analysis_prompt - assert "必须先使用代码读取和搜索工具查看当前 ReMe 代码仓库" in analysis_prompt - assert "这应当是少数例外:一般情况下不要给建议" in analysis_prompt + assert "长期记忆相关性初筛" not in analysis_prompt + assert "ReMe" not in analysis_prompt + assert "# PDF 分页文本" in analysis_prompt + digest_prompt = cc_wrapper.calls[-1]["inputs"] + assert "Evidence one [p. 1]." in digest_prompt + assert "调用 Read" not in digest_prompt + assert "daily/2026-07-21" not in digest_prompt + assert "长期记忆" not in digest_prompt rerun = RuntimeContext(date="2026-07-21") await DailyPaperCollectStep(app_context=app_context)(rerun) assert rerun.response.metadata["skipped"] is True - assert rerun.response.metadata["selection"] == { - "selection_reasoning": "Best remaining ranked paper.", - "selected": [ - { - "arxiv_id": "2607.10001", - "rank": 1, - "reason": "Strong result", - "memory_relevance": "low", - }, - ], - "alternates": [], - } - assert rerun.get("daily_paper_digest_path") == "daily/2026-07-21/daily-paper-brief.md" + assert rerun.get("daily_paper_digest_path") == "daily/2026-07-21/今日智能体论文简报.md" assert _FakeHfClient.requested_daily == ["2026-07-20"] @pytest.mark.asyncio -async def test_dingtalk_markdown_sends_groups_serially_in_configured_order(tmp_path: Path, monkeypatch): +async def test_digest_force_migrates_old_fixed_filename_to_chinese_title(tmp_path: Path): + """A forced regeneration replaces the old generated brief without leaving a stale copy.""" + old_path = tmp_path / "daily" / "2026-07-21" / "daily-paper-brief.md" + old_path.parent.mkdir(parents=True) + old_path.write_text( + frontmatter.dumps(frontmatter.Post("old", kind="daily-paper-brief")), + encoding="utf-8", + ) + analyses = [ + AnalyzedPaper( + arxiv_id=f"2607.1000{index}", + reasoning=f"Reason {index}", + title=f"论文解读{index}", + desc=f"Description {index}", + body=f"Body {index}", + note_path=f"daily/2026-07-21/论文解读{index}.md", + pdf_path=f"resource/papers/2607.1000{index}.pdf", + ) + for index in range(1, 4) + ] + context = RuntimeContext() + context["daily_paper_run_date"] = "2026-07-21" + context["daily_paper_existing_digest_path"] = "daily/2026-07-21/daily-paper-brief.md" + context["daily_paper_analyses"] = analyses + agent = _QueuedAgentWrapper( + [{"title": "全新论文简报", "desc": "Digest", "body": "Digest body"}], + ) + + await DailyPaperDigestStep( + app_context=ApplicationContext(workspace_dir=str(tmp_path)), + agent_wrapper=agent, + )(context) + + new_path = tmp_path / "daily" / "2026-07-21" / "全新论文简报.md" + assert new_path.is_file() + assert not old_path.exists() + assert context["daily_paper_digest_path"] == "daily/2026-07-21/全新论文简报.md" + + +@pytest.mark.asyncio +async def test_dingtalk_markdown_sends_groups_serially_in_configured_order( + tmp_path: Path, + monkeypatch, +): """The notifier gets one app token and posts once per group in list order.""" digest_path = tmp_path / "daily" / "2026-07-21" / "daily-paper-brief.md" digest_path.parent.mkdir(parents=True) digest_path.write_text( - frontmatter.dumps(frontmatter.Post("# 今日论文\n\n测试内容", name="daily-paper-brief")), + frontmatter.dumps( + frontmatter.Post("# 今日论文\n\n测试内容", name="daily-paper-brief"), + ), encoding="utf-8", ) token_calls = 0 @@ -622,7 +800,10 @@ async def test_dingtalk_markdown_sends_groups_serially_in_configured_order(tmp_p assert request.url.path == "/v1.0/robot/groupMessages/send" assert request.headers["x-acs-dingtalk-access-token"] == "app-access-token" seen_payloads.append(json.loads(request.content)) - return httpx.Response(200, json={"processQueryKey": f"query-{len(seen_payloads)}"}) + return httpx.Response( + 200, + json={"processQueryKey": f"query-{len(seen_payloads)}"}, + ) transport = httpx.MockTransport(handler) transport_kwargs: dict = {} @@ -632,7 +813,11 @@ async def test_dingtalk_markdown_sends_groups_serially_in_configured_order(tmp_p return transport dingtalk_stream = importlib.import_module("dingtalk_stream") - monkeypatch.setattr(dingtalk_stream.DingTalkStreamClient, "get_access_token", get_access_token) + monkeypatch.setattr( + dingtalk_stream.DingTalkStreamClient, + "get_access_token", + get_access_token, + ) monkeypatch.setattr(dingtalk_send.httpx, "AsyncHTTPTransport", ipv4_transport) app_context = ApplicationContext(workspace_dir=str(tmp_path)) context = RuntimeContext(markdown_path="daily/2026-07-21/daily-paper-brief.md") @@ -650,7 +835,10 @@ async def test_dingtalk_markdown_sends_groups_serially_in_configured_order(tmp_p assert token_calls == 1 assert transport_kwargs == {"local_address": "0.0.0.0"} - assert [payload["openConversationId"] for payload in seen_payloads] == ["group-one", "group-two"] + assert [payload["openConversationId"] for payload in seen_payloads] == [ + "group-one", + "group-two", + ] assert all(payload["robotCode"] == "robot-code" for payload in seen_payloads) assert all(payload["msgKey"] == "sampleMarkdown" for payload in seen_payloads) assert [json.loads(payload["msgParam"]) for payload in seen_payloads] == [ @@ -669,7 +857,9 @@ async def test_dingtalk_markdown_without_conversations_is_a_noop(tmp_path: Path) """An empty conversation list keeps daily-paper generation usable without DingTalk.""" context = RuntimeContext(markdown_path="missing.md") - response = await DingTalkMarkdownSendStep(app_context=ApplicationContext(workspace_dir=str(tmp_path)))(context) + response = await DingTalkMarkdownSendStep( + app_context=ApplicationContext(workspace_dir=str(tmp_path)), + )(context) assert response.success is True assert response.metadata["dingtalk_configured_count"] == 0 @@ -677,12 +867,17 @@ async def test_dingtalk_markdown_without_conversations_is_a_noop(tmp_path: Path) @pytest.mark.asyncio -async def test_existing_daily_paper_is_reused_and_sent_to_dingtalk(tmp_path: Path, monkeypatch): +async def test_existing_daily_paper_is_reused_and_sent_to_dingtalk( + tmp_path: Path, + monkeypatch, +): """An idempotent daily-paper run skips generation but still notifies DingTalk.""" digest_path = tmp_path / "daily" / "2026-07-22" / "daily-paper-brief.md" digest_path.parent.mkdir(parents=True) digest_path.write_text( - frontmatter.dumps(frontmatter.Post("# 已有日报\n\n复用正文", name="daily-paper-brief")), + frontmatter.dumps( + frontmatter.Post("# 已有日报\n\n复用正文", name="daily-paper-brief"), + ), encoding="utf-8", ) seen_payloads: list[dict] = [] @@ -700,7 +895,11 @@ async def test_existing_daily_paper_is_reused_and_sent_to_dingtalk(tmp_path: Pat transport = httpx.MockTransport(handler) - monkeypatch.setattr(dingtalk_send.httpx, "AsyncHTTPTransport", lambda **_kwargs: transport) + monkeypatch.setattr( + dingtalk_send.httpx, + "AsyncHTTPTransport", + lambda **_kwargs: transport, + ) app_context = ApplicationContext(workspace_dir=str(tmp_path)) context = RuntimeContext(date="2026-07-22") diff --git a/tests/unit/test_dingtalk_wait.py b/tests/unit/test_dingtalk_wait.py index 09768971..5bbefade 100644 --- a/tests/unit/test_dingtalk_wait.py +++ b/tests/unit/test_dingtalk_wait.py @@ -130,7 +130,11 @@ async def test_final_reply_resumes_session_and_clear_only_removes_combined_key( await step._handle_message(message, key, sessions, handler) assert sessions == {key: "session-1"} - assert wrapper.reply_calls == [("hello", {}), ("hello", {"resume": "session-1"})] + tool_kwargs = {"builtin_tools": False, "job_tools": []} + assert wrapper.reply_calls == [ + ("hello", tool_kwargs), + ("hello", {"resume": "session-1", **tool_kwargs}), + ] assert handler.markdown_replies == [("ReMe Agent", "回答"), ("ReMe Agent", "回答")] await step._handle_message(_message(text="/compact"), key, sessions, handler) @@ -174,6 +178,31 @@ async def test_final_reply_rejects_empty_agent_reply_and_dingtalk_send_failure( await step._handle_message(message, key, {}, handler) +@pytest.mark.asyncio +async def test_final_reply_injects_only_configured_tools(tmp_path): + app_context = ApplicationContext(workspace_dir=str(tmp_path)) + wrapper = _AgentWrapper(app_context=app_context) + step = DingTalkWaitStep( + app_context=app_context, + agent_wrapper=wrapper, + builtin_tools=["bash"], + job_tools=["read", "write", "edit"], + ) + + message = _message() + await step._handle_message(message, _session_key(message), {}, _Handler()) + + assert wrapper.reply_calls == [ + ( + "hello", + { + "builtin_tools": ["bash"], + "job_tools": ["read", "write", "edit"], + }, + ), + ] + + def test_daily_cookbook_registers_one_step_background_wait_job(monkeypatch): for name in ("DINGTALK_APP_KEY", "DINGTALK_APP_SECRET", "DINGTALK_ROBOT_CODE"): monkeypatch.delenv(name, raising=False) @@ -183,25 +212,29 @@ def test_daily_cookbook_registers_one_step_background_wait_job(monkeypatch): assert job["steps"] == [ { "backend": "dingtalk_wait_step", - "agent_wrapper": "dingtalk_wait", "app_key": "", "app_secret": "", "robot_code": "", "worker_count": 4, + "builtin_tools": ["bash"], + "job_tools": [ + "memory_search", + "read", + "write", + "edit", + "daily_list", + "daily_write", + "frontmatter_read", + "frontmatter_update", + ], }, ] - dingtalk_wait = config["components"]["agent_wrapper"]["dingtalk_wait"] - assert dingtalk_wait["skills"] == ["tushare-data"] - assert dingtalk_wait["job_tools"] == ["memory_search"] - assert dingtalk_wait["system_prompt"] == { - "type": "preset", - "preset": "claude_code", - "append": ( - "Daily-paper Markdown is stored under the ReMe workspace. Detailed notes, including historical notes, " - "are at daily/YYYY-MM-DD/paper-.md; daily briefs are at " - "daily/YYYY-MM-DD/daily-paper-brief.md. Use memory_search to retrieve relevant long-term notes " - "across dates." - ), + assert config["components"]["agent_wrapper"] == { + "default": { + "backend": "agentscope", + "as_llm": "default", + "builtin_tools": False, + }, } assert R.get(ComponentEnum.STEP, "dingtalk_wait_step") is DingTalkWaitStep diff --git a/tests/unit/test_tushare.py b/tests/unit/test_tushare.py new file mode 100644 index 00000000..2bbdbe4e --- /dev/null +++ b/tests/unit/test_tushare.py @@ -0,0 +1,21 @@ +"""Focused tests for TuShare source selection.""" + +# pylint: disable=protected-access + +import sys +from types import SimpleNamespace + +from reme.utils.tushare import create_tushare_api + + +def test_tushare_uses_configured_mirror_or_sdk_default(monkeypatch): + """A mirror overrides the SDK endpoint while an empty setting leaves it unchanged.""" + api = SimpleNamespace(_DataApi__http_url="https://api.tushare.pro") + monkeypatch.setitem(sys.modules, "tushare", SimpleNamespace(pro_api=lambda _token: api)) + + monkeypatch.delenv("TUSHARE_MIRROR_URL", raising=False) + assert create_tushare_api("token")._DataApi__http_url == "https://api.tushare.pro" + assert api._DataApi__timeout == 600 + + monkeypatch.setenv("TUSHARE_MIRROR_URL", "http://112.124.63.173:4000/tushare/") + assert create_tushare_api("token")._DataApi__http_url == "http://112.124.63.173:4000/tushare"