fix(release): harden embedding store and plugins for ReMe 0.4.1.9 (#503)
Some checks failed
CI / Python packages / Build and verify distributions (push) Has been cancelled
CI / Python tests / Unit Tests - py3.12 (push) Has been cancelled
CI / Python tests / Unit Tests - py3.13 (push) Has been cancelled
CI / ReMe Studio / Studio checks (push) Has been cancelled
CI / TypeScript integrations / Type-check, test, and pack (push) Has been cancelled
CI / Windows / CLI smoke - py3.11 (push) Has been cancelled
Deploy / Documentation / Build documentation (push) Has been cancelled
CI / Python quality / Pre-commit (push) Has been cancelled
Security / CodeQL / Analyze javascript-typescript (push) Has been cancelled
Security / CodeQL / Analyze python (push) Has been cancelled
CI / Documentation / Test and build documentation (push) Has been cancelled
CI / Python tests / Unit Tests - py3.11 (push) Has been cancelled
Deploy / Documentation / deploy (push) Has been cancelled

* chore(release): prepare ReMe 0.4.1.9

* refactor(config): remove daily_cookbook and streamline plugin configs

- Delete the entire daily_cookbook.yaml standalone application config
- Remove qwenpaw dependencies verification and related CI workflow steps
- Simplify release workflows by removing qwenpaw verification and enforcing reme-ai >=0.4.1.9
- Update plugin start commands and examples to use 'default' or 'demo' configs instead of daily_cookbook
- Adjust imports and tests related to daily_cookbook removal and injected_job_kwargs enhancements
- Refactor agent wrapper to support injected_job_kwargs for job parameter injection in auto-fin and daily-paper
- Improve daily_paper digest prompt to include configured daily directory and correct historical search constraints
- Update dependency versions in pyproject.toml files to require reme-ai >=0.4.1.9 and remove qwenpaw optional dependencies
- Clean up unused environment variables and obsolete test cases related to daily_cookbook and verification steps

* fix(local_embedding_store): retry batch computation on vector space changes

- Add up to 3 attempts to recompute embedding batch if vector space changes during processing
- Log warnings when maximum retries reached and discard stale results
- Prevent caching results from outdated vector spaces to maintain consistency
- Add tests to verify retry behavior and abort after continuous vector space churn

fix(daily_paper): update digest search logic and tests

- Change search to query existing memory, not only previous articles in daily_dir
- Allow multiple searches outside daily_dir but limit links to dated markdown in daily_dir before today
- Update test assertions to reflect revised search and linking rules

* fix(embedding): retry vector space changes per request
This commit is contained in:
jinliyl 2026-08-28 11:35:04 +08:00 • committed by GitHub
parent 2dd2255760
commit 99afc2604f
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
31 changed files with 293 additions and 642 deletions

View file

@ -13,11 +13,6 @@ on:
required: false
default: false
type: boolean
verify_qwenpaw_dependencies:
description: Verify that independently published qwenpaw plugins are installable
required: false
default: true
type: boolean
permissions:
contents: read
@ -84,22 +79,6 @@ jobs:
assert (static_dir() / "index.html").is_file()
PY
# qwenpaw composes independently released plugins. Bootstrap releases may
# skip this check to publish the reme-ai version required by those plugins.
- name: Verify released qwenpaw dependencies
if: inputs.expected_version != '' && inputs.verify_qwenpaw_dependencies
run: |
REME_WHEEL="$(pwd)/$(ls dist/reme/reme_ai-[0-9]*.whl)"
python -m venv "${RUNNER_TEMP}/reme-qwenpaw-package-smoke"
"${RUNNER_TEMP}/reme-qwenpaw-package-smoke/bin/python" -m pip install "${REME_WHEEL}[qwenpaw]"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-qwenpaw-package-smoke/bin/python" - <<'PY'
from importlib.metadata import distribution
assert distribution("reme-auto-fin")
assert distribution("reme-daily-paper")
PY
- name: Upload ReMe distributions
if: inputs.upload_artifacts
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6

View file

@ -73,8 +73,10 @@ jobs:
if len(requirements) != 1:
raise SystemExit(f"Expected one reme-ai dependency, found {requirements!r}")
reme_requirement = Requirement(requirements[0])
if reme_requirement.name != "reme-ai" or set(reme_requirement.extras) != {"core"}:
raise SystemExit(f"Expected a reme-ai[core] dependency, found {requirements[0]!r}")
if reme_requirement.name != "reme-ai" or reme_requirement.extras:
raise SystemExit(f"Expected a base reme-ai dependency, found {requirements[0]!r}")
if Version("0.4.1.8") in reme_requirement.specifier or Version("0.4.1.9") not in reme_requirement.specifier:
raise SystemExit(f"Expected reme-ai>=0.4.1.9, found {requirements[0]!r}")
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
print(f"reme_requirement={reme_requirement}", file=output)
print(f"Publishing {project['name']} {actual}")
@ -88,7 +90,7 @@ jobs:
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
run: |
python -m pip download --no-deps \
--dest "${RUNNER_TEMP}/reme-auto-fin-core" \
--dest "${RUNNER_TEMP}/reme-auto-fin-base" \
"${REME_REQUIREMENT}"
- name: Build and check distributions
@ -104,7 +106,8 @@ jobs:
python -m zipfile -l "${AUTO_FIN_WHEEL}" | grep 'dist-info/licenses/LICENSE'
python -m tarfile -l "${AUTO_FIN_SDIST}" | grep '/LICENSE'
python -m venv "${RUNNER_TEMP}/reme-auto-fin-smoke"
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" -m pip install "${AUTO_FIN_WHEEL}"
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" -m pip install \
"agentscope[model-ollama]==2.0.7" "${AUTO_FIN_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-auto-fin-smoke/bin/python" - <<'PY'
from importlib.metadata import distribution
@ -122,9 +125,7 @@ jobs:
}
assert set(manifest.application_defaults["jobs"]) == {
"auto_fin",
"auto_fin_0930_cron",
"auto_fin_1130_cron",
"auto_fin_1800_cron",
"auto_fin_cron",
}
PY

View file

@ -70,8 +70,10 @@ jobs:
raise SystemExit(f"Package version is {actual}, but workflow input is {expected}")
requirements = [Requirement(value) for value in project["dependencies"]]
reme_requirements = [requirement for requirement in requirements if requirement.name == "reme-ai"]
if len(reme_requirements) != 1 or set(reme_requirements[0].extras) != {"core"}:
raise SystemExit(f"Expected one reme-ai[core] dependency, found {reme_requirements!r}")
if len(reme_requirements) != 1 or reme_requirements[0].extras:
raise SystemExit(f"Expected one base reme-ai dependency, found {reme_requirements!r}")
if Version("0.4.1.8") in reme_requirements[0].specifier or Version("0.4.1.9") not in reme_requirements[0].specifier:
raise SystemExit(f"Expected reme-ai>=0.4.1.9, found {reme_requirements!r}")
if sum(requirement.name == "pypdf" for requirement in requirements) != 1:
raise SystemExit("Expected exactly one pypdf dependency")
with Path(os.environ["GITHUB_OUTPUT"]).open("a", encoding="utf-8") as output:
@ -87,7 +89,7 @@ jobs:
REME_REQUIREMENT: ${{ steps.package.outputs.reme_requirement }}
run: |
python -m pip download --no-deps \
--dest "${RUNNER_TEMP}/reme-daily-paper-core" \
--dest "${RUNNER_TEMP}/reme-daily-paper-base" \
"${REME_REQUIREMENT}"
- name: Build and check distributions
@ -105,7 +107,8 @@ jobs:
python -m zipfile -l "${DAILY_PAPER_WHEEL}" | grep 'dist-info/licenses/LICENSE'
python -m tarfile -l "${DAILY_PAPER_SDIST}" | grep '/LICENSE'
python -m venv "${RUNNER_TEMP}/reme-daily-paper-smoke"
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" -m pip install "${DAILY_PAPER_WHEEL}"
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" -m pip install \
"agentscope[model-ollama]==2.0.7" "${DAILY_PAPER_WHEEL}"
cd "${RUNNER_TEMP}"
"${RUNNER_TEMP}/reme-daily-paper-smoke/bin/python" - <<'PY'
from importlib.metadata import distribution

View file

@ -1,8 +1,5 @@
name: Release / Python packages
# reme-ai[qwenpaw] is normally verified before publication. For a bootstrap
# release where the plugins require this new reme-ai version, disable that
# verification, publish reme-ai first, and then publish the plugins.
# Configure a PyPI Trusted Publisher for this repository, workflow, and its
# pypi environment before running the manual release.
@ -13,11 +10,6 @@ on:
description: Release version
required: true
type: string
verify_qwenpaw_dependencies:
description: Verify already-published Auto Fin and Daily Paper packages
required: true
default: true
type: boolean
permissions:
contents: read
@ -33,7 +25,6 @@ jobs:
with:
expected_version: ${{ inputs.version }}
upload_artifacts: true
verify_qwenpaw_dependencies: ${{ inputs.verify_qwenpaw_dependencies }}
publish-reme:
needs: build

View file

@ -64,7 +64,7 @@ reme plugins list --json
To compare installed plugins with one application config:
```bash
reme plugins list --config daily_cookbook
reme plugins list --config default
```
The optional `ENABLED` column reflects only the `plugins` list resolved from that config. A command-line override used
@ -173,11 +173,8 @@ curl -s http://127.0.0.1:2333/auto_fin \
When the application uses an MCP service, service-enabled plugin Jobs appear as MCP tools instead.
To add the plugin to another application config, select it explicitly:
```bash
reme start config=daily_cookbook plugins='["auto-fin"]'
```
Custom application configs must provide the plugin's runtime dependencies, including an `agent_wrapper.default` and
the `search` and `read` Jobs used by Auto Fin.
## Uninstall a plugin

View file

@ -61,7 +61,7 @@ reme plugins list --json
对照某个应用配置查看启用状态:
```bash
reme plugins list --config daily_cookbook
reme plugins list --config default
```
可选的 `ENABLED` 列只反映该配置解析出的 `plugins` 列表。其他运行中进程使用的 CLI override 不是全局启用状态。
@ -167,11 +167,7 @@ curl -s http://127.0.0.1:2333/auto_fin \
当应用使用 MCP service 时,允许对外服务的插件 Job 会显示为 MCP tool。
如果需要将插件叠加到其他应用配置,则显式选择该配置:
```bash
reme start config=daily_cookbook plugins='["auto-fin"]'
```
自定义应用配置需要提供插件的运行依赖,包括 `agent_wrapper.default`,以及 Auto Fin 使用的 `search` 和 `read` Jobs。
## 卸载插件

View file

@ -17,7 +17,7 @@ through `plugins=["auto-fin"]`.
### 1. Install ReMe and Auto Fin
```bash
python -m pip install "reme-ai[core]>=0.4.1.8"
python -m pip install "reme-ai[core]>=0.4.1.9"
reme plugins install reme-auto-fin
```
@ -59,11 +59,7 @@ reme start plugins='["auto-fin"]' \
service.backend=http
```
To add Auto Fin to another application instead, select that config explicitly, for example:
```bash
reme start config=daily_cookbook plugins='["auto-fin"]'
```
Custom application configs must provide `agent_wrapper.default` and the `search` and `read` Jobs used by Auto Fin.
## Pipeline
@ -74,7 +70,7 @@ normalize and deduplicate in RuntimeContext
↓
topic Agent selects real news IDs in bounded batches
↓
research Agent uses memory_search + read on historical memory
research Agent uses search + read on historical memory
↓
validate historical wikilinks in code
↓
@ -89,8 +85,8 @@ records outside the window are discarded.
IDs and deduplicates repeated IDs, then preserves the source-news order. If nothing is relevant, the job succeeds as a
skip without writing or sending a report.
`auto_fin_merge_step` receives only selected current news. It exposes `memory_search` and `read`, instructs the Agent to
search no later than yesterday, and keeps current CLS IDs, times, and titles as plain evidence. The prompt limits
`auto_fin_merge_step` receives only selected current news. It exposes `search` and `read`, and keeps current CLS IDs,
times, and titles as plain evidence. The prompt limits
wikilinks to historical Markdown actually used by the Agent; the code-level boundary independently keeps only existing,
workspace-relative Markdown targets. Missing, absolute, escaping, backslash, and self-referential targets are degraded
to their readable aliases.
@ -109,7 +105,7 @@ refreshes the daily index. No JSONL, intermediate Markdown, or structured Agent
| `request_interval` | `10` | Minimum delay in seconds after every CLS request attempt; may be zero |
| `max_retries` | `3` | Maximum attempts for each CLS page request; must be at least one |
The three plugin cron Jobs start with the application and run daily at 09:30, 11:30, and 18:00 in `Asia/Shanghai`.
The plugin cron Job starts with the application and runs daily at 18:00 in the application timezone.
## Output

View file

@ -14,7 +14,7 @@ distribution:单个 `reme.plugins` entry point 暴露 `plugin.yaml`,其中
### 1. 安装 ReMe 和 Auto Fin
```bash
python -m pip install "reme-ai[core]>=0.4.1.8"
python -m pip install "reme-ai[core]>=0.4.1.9"
reme plugins install reme-auto-fin
```
@ -54,11 +54,7 @@ reme start plugins='["auto-fin"]' \
service.backend=http
```
如果需要将 Auto Fin 叠加到其他应用,则显式选择相应配置,例如:
```bash
reme start config=daily_cookbook plugins='["auto-fin"]'
```
自定义应用配置需要提供 `agent_wrapper.default`,以及 Auto Fin 使用的 `search` 和 `read` Jobs。
## 流程
@ -69,7 +65,7 @@ reme start config=daily_cookbook plugins='["auto-fin"]'
↓
Topic Agent 分批选择真实 news_id
↓
Research Agent 使用 memory_search + read 检索历史记忆
Research Agent 使用 search + read 检索历史记忆
↓
代码校验历史 wikilink
↓
@ -82,8 +78,8 @@ daily/YYYY-MM-DD/auto_fin.md
`auto_fin_topic_step` 分批接收当前新闻,只返回相关的 `news_id`。代码会忽略未知 ID、去除重复 ID,并保持源新闻顺序。如果没有相关新闻,Job
会成功跳过,不写报告也不发送通知。
`auto_fin_merge_step` 只接收筛选后的当前新闻,并向 Agent 开放 `memory_search` 和 `read`。历史检索截止到昨天;当前新闻以
CLS ID、时间和标题作为普通证据。Prompt 要求 Agent 只链接实际使用过的历史 Markdown;代码边界则独立保证只保留真实存在、相对
`auto_fin_merge_step` 只接收筛选后的当前新闻,并向 Agent 开放 `search` 和 `read`。当前新闻以 CLS ID、时间和标题作为普通证据。
Prompt 要求 Agent 只链接实际使用过的历史 Markdown;代码边界则独立保证只保留真实存在、相对
workspace 的 Markdown 目标。不存在、绝对路径、越界、带反斜杠和自引用的目标都会降级为可读 alias。
同日重跑会参考当天已有报告并覆盖为修订结果。最终写入使用原子替换并刷新当天索引;流程不会写入 JSONL、中间 Markdown 或 Agent
@ -100,7 +96,7 @@ workspace 的 Markdown 目标。不存在、绝对路径、越界、带反斜杠
| `request_interval` | `10` | 每次财联社请求尝试后的最小等待秒数,可设为 0 |
| `max_retries` | `3` | 每页财联社请求的最大尝试次数,至少为 1 |
插件的三个 cron Job 随应用启动,并按 `Asia/Shanghai` 时区在每天 09:30、11:30 和 18:00 运行。
插件的 cron Job 随应用启动,并按应用配置的时区在每天 18:00 运行。
## 产物

View file

@ -1,13 +1,13 @@
[project]
name = "reme-auto-fin"
version = "0.1.1"
version = "0.1.2"
description = "Auto Fin example plugin for ReMe."
readme = "README.md"
license = "Apache-2.0"
license-files = ["LICENSE"]
requires-python = ">=3.11"
dependencies = [
"reme-ai[core]>=0.4.1.8",
"reme-ai>=0.4.1.9",
]
[project.entry-points."reme.plugins"]

View file

@ -76,6 +76,7 @@ class AutoFinStep(BaseStep):
prompt_name: str,
model: type[BaseModel],
job_tools: list[str] | None = None,
injected_job_kwargs: dict[str, Any] | None = None,
**values: str,
) -> BaseModel:
if self.agent_wrapper is None:
@ -89,6 +90,8 @@ class AutoFinStep(BaseStep):
kwargs: dict[str, Any] = {"output_schema": model}
if job_tools:
kwargs["job_tools"] = job_tools
if injected_job_kwargs:
kwargs["injected_job_kwargs"] = injected_job_kwargs
result = await self.agent_wrapper.reply(prompt, **kwargs)
if not isinstance(result, dict) or result.get("structured_output") is None:
raise ValueError(f"Auto Fin Agent returned no structured output: {self._preview(result)}")

View file

@ -102,17 +102,24 @@ class AutoFinMergeStep(AutoFinStep):
)
async def execute(self):
"""Research the selected news and persist the validated report."""
assert self.context is not None
if self.context.get("auto_fin_skipped"):
return self.context.response
run_date = date.fromisoformat(str(self._required("auto_fin_date")))
historical_search = {
"limit": 5,
"min_score": 0.0,
"start_date": None,
"end_date": (run_date - timedelta(days=1)).isoformat(),
}
output = await self._reply(
"merge_user",
AutoFinReportOutput,
job_tools=list(self.kwargs.get("job_tools") or []),
injected_job_kwargs=historical_search,
decision_at=str(self._required("auto_fin_decision_at")),
window_start=str(self._required("auto_fin_window_start")),
historical_end=(run_date - timedelta(days=1)).isoformat(),
topics=json.dumps(self._required("auto_fin_topics"), ensure_ascii=False),
news=json.dumps(self._required("auto_fin_selected_news"), ensure_ascii=False),
current_report=self._current_report(run_date),

View file

@ -1,5 +1,5 @@
merge_user: |
你是主题新闻研究 Agent。当前新闻已经按 topics 做过语义筛选。你可以使用 `memory_search` 搜索历史记忆,
你是主题新闻研究 Agent。当前新闻已经按 topics 做过语义筛选。你可以使用 `search` 搜索历史记忆,
并使用 `read` 阅读可能相关的完整 Markdown。不得使用外部搜索,不得虚构行情、收益、价格或未提供的数据。
研究窗口:{window_start} 至 {decision_at}
@ -8,7 +8,7 @@ merge_user: |
今天早些时段的报告(如有,请保留仍成立的判断,只修订变化部分):
{current_report}
先围绕 topics 和当前重要事件多次调用 `memory_search`,并将 end_date 设为 {historical_end},避免召回今天的旧报告。
先围绕 topics 和当前重要事件多次调用 `search` 检索历史记忆。
只对明显相关的结果调用 `read`。说明历史事件与当前事件的相同点、关键差异,以及旧判断是否仍适用。
给出值得回顾的新闻、应继续观察的信息,以及哪些条件会强化或推翻判断,但不要给出投资建议。

View file

@ -42,19 +42,9 @@ application_defaults:
- backend: auto_fin_data_step
- backend: auto_fin_topic_step
- backend: auto_fin_merge_step
job_tools: [memory_search, read]
job_tools: [search, read]
auto_fin_0930_cron:
backend: cron
cron: "30 9 * * *"
steps: *auto_fin_steps
auto_fin_1130_cron:
backend: cron
cron: "30 11 * * *"
steps: *auto_fin_steps
auto_fin_1800_cron:
auto_fin_cron:
backend: cron
cron: "0 18 * * *"
steps: *auto_fin_steps

View file

@ -195,14 +195,23 @@ async def test_merge_writes_only_final_report_and_validates_historical_links(tmp
response = await AutoFinMergeStep(
app_context=app_context,
agent_wrapper=agent,
job_tools=["memory_search", "read"],
job_tools=["search", "read"],
)(context)
prompt, kwargs = agent.calls[0]
assert "end_date 设为 2026-08-09" in prompt
assert "调用 `memory_search`" in prompt
assert "end_date" not in prompt
assert "调用 `search`" in prompt
assert "调用 `read`" in prompt
assert kwargs == {"output_schema": AutoFinReportOutput, "job_tools": ["memory_search", "read"]}
assert kwargs == {
"output_schema": AutoFinReportOutput,
"job_tools": ["search", "read"],
"injected_job_kwargs": {
"limit": 5,
"min_score": 0.0,
"start_date": None,
"end_date": "2026-08-09",
},
}
report = (tmp_path / "daily" / "2026-08-10" / "auto_fin.md").read_text(encoding="utf-8")
assert "[[daily/2026-08-01/auto_fin.md|历史黄金观察]]" in report
assert "](daily/2026-08-01/auto_fin.md)" not in report
@ -254,14 +263,17 @@ def test_plugin_config_has_default_topics_and_no_intermediate_index_step():
"auto_fin_topic_step",
"auto_fin_merge_step",
]
assert job["steps"][2]["job_tools"] == ["memory_search", "read"]
for name, schedule in {
"auto_fin_0930_cron": "30 9 * * *",
"auto_fin_1130_cron": "30 11 * * *",
"auto_fin_1800_cron": "0 18 * * *",
}.items():
assert jobs[name]["cron"] == schedule
assert jobs[name]["steps"] == job["steps"]
assert job["steps"][2]["job_tools"] == ["search", "read"]
assert jobs["auto_fin_cron"]["cron"] == "0 18 * * *"
assert jobs["auto_fin_cron"]["steps"] == job["steps"]
assert (
not {
"auto_fin_0930_cron",
"auto_fin_1130_cron",
"auto_fin_1800_cron",
}
& jobs.keys()
)
def test_agent_schemas_are_small_and_required():

View file

@ -13,7 +13,7 @@ their Job configuration under `application_defaults`. Enable the installed plugi
### 1. Install ReMe and Daily Paper
```bash
python -m pip install "reme-ai[core]>=0.4.1.8"
python -m pip install "reme-ai[core]>=0.4.1.9"
reme plugins install reme-daily-paper
```
@ -62,7 +62,7 @@ rank with RRF and let an Agent select three papers
↓
download and parse arXiv PDFs, then write three Chinese analyses
↓
use memory_search + read to connect prior memory and generate a brief
use search + read to connect prior memory and generate a brief
↓
refresh the daily index and optionally send the brief to DingTalk
```
@ -80,7 +80,7 @@ the configured page, character, and file-size limits. It writes the three Chines
PDFs and files without a text layer fail explicitly.
`daily_paper_digest_step` treats those three analyses as the factual source and receives only the read-only
`memory_search` and `read` tools for linking earlier memory. Code validates historical wikilinks, appends links to all
`search` and `read` tools for linking earlier memory. Code validates historical wikilinks, appends links to all
three source notes, and rebuilds the daily index. The optional `dingtalk_markdown_send_step` sends the final brief when
conversation IDs are configured and otherwise skips without side effects.

View file

@ -11,7 +11,7 @@ Step backend,并在 `application_defaults` 下提供 Job 配置;通过 `plug
### 1. 安装 ReMe 和每日论文插件
```bash
python -m pip install "reme-ai[core]>=0.4.1.8"
python -m pip install "reme-ai[core]>=0.4.1.9"
reme plugins install reme-daily-paper
```
@ -58,7 +58,7 @@ RRF 排序后由 Agent 精选三篇
↓
下载并解析 arXiv PDF,生成三篇中文解读
↓
使用 memory_search + read 关联历史记忆并生成简报
使用 search + read 关联历史记忆并生成简报
↓
写入当日索引,并按需发送到钉钉
```
@ -72,7 +72,7 @@ RRF 排序后由 Agent 精选三篇
`daily_paper_analyze_step` 下载 PDF 到 `resource/papers/`,复用已有的有效文件,并在页数、字符数和文件大小限制内提取
文本。三篇中文解读按精选顺序写入当天目录;扫描版或没有文本层的 PDF 会明确失败。
`daily_paper_digest_step` 以本次生成的三篇解读为事实来源,只开放只读的 `memory_search` 和 `read` 来关联较早记忆。
`daily_paper_digest_step` 以本次生成的三篇解读为事实来源,只开放只读的 `search` 和 `read` 来关联较早记忆。
代码会校验历史 wikilink、追加三篇源笔记链接,并重建当日索引。可选的 `dingtalk_markdown_send_step` 在配置群会话后
发送最终简报;未配置时无副作用跳过。

View file

@ -1,6 +1,6 @@
[project]
name = "reme-daily-paper"
version = "0.1.1"
version = "0.1.2"
description = "Daily Paper research and reading-note plugin for ReMe."
readme = "README.md"
license = "Apache-2.0"
@ -8,7 +8,7 @@ license-files = ["LICENSE"]
requires-python = ">=3.11"
dependencies = [
"pypdf>=5.0.0",
"reme-ai[core]>=0.4.1.8",
"reme-ai>=0.4.1.9",
]
[project.entry-points."reme.plugins"]

View file

@ -64,6 +64,7 @@ class DailyPaperDigestStep(DailyPaperStep):
return _WIKILINK_RE.sub(replace, body)
async def execute(self):
"""Generate and persist the final brief from analyzed papers."""
assert self.context is not None
if self._skip():
self.logger.info(f"[{self.name}] skip existing digest")
@ -79,16 +80,23 @@ class DailyPaperDigestStep(DailyPaperStep):
documents = [{"title": item.title, "desc": item.desc, "body": item.body} for item in analyses]
wikilinks = [f"[[{item.note_path}]]" for item in analyses]
previous_day = (dt.date.fromisoformat(self._run_day()) - dt.timedelta(days=1)).isoformat()
run_day = dt.date.fromisoformat(self._run_day())
daily_dir = str(self.config_value("daily_dir")).strip("/")
self.logger.info(f"[{self.name}] agent start notes={len(analyses)}")
result = await self.agent_wrapper.reply(
self.prompt_format(
"digest_user",
documents=json.dumps(documents, ensure_ascii=False, indent=2),
previous_day=previous_day,
daily_dir=daily_dir,
),
output_schema=DailyPaperMarkdownOutput,
job_tools=list(self.kwargs.get("job_tools") or []),
injected_job_kwargs={
"limit": 20,
"min_score": 0.0,
"start_date": None,
"end_date": (run_day - dt.timedelta(days=1)).isoformat(),
},
)
self.logger.info(f"[{self.name}] agent done notes={len(analyses)}")
output = structured_output(result, DailyPaperMarkdownOutput)
@ -96,8 +104,7 @@ class DailyPaperDigestStep(DailyPaperStep):
if not output.desc.strip() or not body:
raise ValueError("Agent returned an empty daily paper brief")
day = self._run_day()
daily_dir = str(self.config_value("daily_dir")).strip("/")
day = run_day.isoformat()
title = normalize_chinese_title(output.title, f"每日论文简报-{day}")
existing_rel = str(self._state("existing_digest_path") or "").strip()
existing_path = self.workspace_path / existing_rel if existing_rel else None
@ -110,7 +117,7 @@ class DailyPaperDigestStep(DailyPaperStep):
existing=existing_path,
)
digest_rel = digest_path.relative_to(self.workspace_path).as_posix()
body = self._validate_historical_wikilinks(body, dt.date.fromisoformat(day), digest_path)
body = self._validate_historical_wikilinks(body, run_day, digest_path)
body += "\n\n## 详细论文\n\n" + "\n".join(f"- {link}" for link in wikilinks)
selected_ids = [item.arxiv_id for item in analyses]
await write_markdown(

View file

@ -4,12 +4,13 @@ digest_user: |
内容只能依据输入文档,不得补充文档中没有提供的事实。
保留技术准确性,同时解释三篇论文为什么值得关注,以及它们之间有什么联系。
在写作前,先调用 `memory_search` 检索以前的文章:围绕三篇论文的核心问题、方法、关键词和同义表达组织查询,
使用 end_date={previous_day}、limit=20。主题跨度较大时可以多次检索。只把 `daily/` 下日期早于今天、
且与本期内容确实相似或互补的 Markdown 文章作为候选;必要时调用 `read` 核验全文,不要仅凭标题判断。
在写作前,先调用 `search` 检索已有记忆:围绕三篇论文的核心问题、方法、关键词和同义表达组织查询。
主题跨度较大时可以多次检索,搜索结果不必局限于 `{daily_dir}/`。只有 `{daily_dir}/` 下日期早于今天、
且与本期内容确实相似或互补的 Markdown 文章才可作为正文中的历史链接候选;必要时调用 `read` 核验全文,
不要仅凭标题判断。
将确认相关的旧文章以 Wikilink 自然织入正文,并用句子说明关联(延续、对比、补充或方法相似);
链接必须采用带 `.md` 的完整 workspace-relative 路径,例如
`[[daily/2026-07-01/旧文章.md|此前的相关解读]]`。不要输出裸链接、独立关系字段,也不要虚构搜索未命中的路径。
`[[{daily_dir}/2026-07-01/旧文章.md|此前的相关解读]]`。不要输出裸链接、独立关系字段,也不要虚构搜索未命中的路径。
旧文章只用于判断关联和建立链接,不得用来补充本期事实。如果没有真正相关的旧文章,不要强行添加;
当日三篇详细解读的链接会由系统统一附在文末。

View file

@ -53,7 +53,7 @@ application_defaults:
- backend: daily_paper_select_step
- backend: daily_paper_analyze_step
- backend: daily_paper_digest_step
job_tools: [memory_search, read]
job_tools: [search, read]
- backend: dingtalk_markdown_send_step
input_mapping:
daily_paper_digest_path: markdown_path

View file

@ -627,6 +627,27 @@ def test_daily_paper_cron_hf_mirror_defaults_enabled_with_environment_override(m
assert _plugin_config()["jobs"]["daily_paper_cron"]["use_hf_mirror"] is False
def test_digest_prompt_uses_configured_daily_directory(tmp_path: Path):
"""Use the host application's daily directory in historical-link guidance."""
step = DailyPaperDigestStep(
app_context=ApplicationContext(
workspace_dir=str(tmp_path),
daily_dir="memory",
),
)
prompt = step.prompt_format(
"digest_user",
documents="[]",
daily_dir=str(step.config_value("daily_dir")).strip("/"),
)
assert "`memory/`" in prompt
assert "搜索结果不必局限于 `memory/`" in prompt
assert "[[memory/2026-07-01/旧文章.md" in prompt
assert "[[daily/2026-07-01/" not in prompt
def test_paper_pick_list_uses_an_object_root_for_tool_output():
"""AgentScope function arguments require an object-root JSON schema."""
schema = PaperPickList.model_json_schema()
@ -846,7 +867,7 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs(
await DailyPaperDigestStep(
app_context=app_context,
agent_wrapper=cc_wrapper,
job_tools=["memory_search", "read"],
job_tools=["search", "read"],
)(context)
assert _FakeHfClient.requested_daily == ["2026-07-20"]
@ -886,7 +907,13 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs(
assert all(call["kwargs"] == {"output_schema": DailyPaperMarkdownOutput} for call in cc_wrapper.calls[1:-1])
assert cc_wrapper.calls[-1]["kwargs"] == {
"output_schema": DailyPaperMarkdownOutput,
"job_tools": ["memory_search", "read"],
"job_tools": ["search", "read"],
"injected_job_kwargs": {
"limit": 20,
"min_score": 0.0,
"start_date": None,
"end_date": "2026-07-20",
},
}
assert [call["kwargs"]["output_schema"] for call in cc_wrapper.calls] == [
PaperPickList,
@ -906,8 +933,10 @@ async def test_pipeline_filters_strict_yesterday_and_writes_outputs(
assert "调用 Read" not in digest_prompt
assert "daily/2026-07-21" not in digest_prompt
assert "长期记忆" not in digest_prompt
assert "先调用 `memory_search` 检索以前的文章" in digest_prompt
assert "end_date=2026-07-20" in digest_prompt
assert "先调用 `search` 检索已有记忆" in digest_prompt
assert "搜索结果不必局限于 `daily/`" in digest_prompt
assert "end_date" not in digest_prompt
assert "limit=" not in digest_prompt
assert "Wikilink" in digest_prompt
rerun = RuntimeContext(date="2026-07-21")

View file

@ -42,7 +42,7 @@ dependencies = [
[project.optional-dependencies]
as = [
"agentscope[model-ollama]==2.0.6",
"agentscope[model-ollama]==2.0.7",
]
web = [
"reme_studio",
@ -62,11 +62,6 @@ core = [
"polars>=1.43.0",
"reme_studio",
]
qwenpaw = [
"reme-ai[core]",
"reme-auto-fin>=0.1.1",
"reme-daily-paper>=0.1.1",
]
dev = [
"packaging>=24.2",
"pre-commit>=4.6.1",

View file

@ -1,6 +1,6 @@
"""ReMe CLI package."""
__version__ = "0.4.1.8"
__version__ = "0.4.1.9"
from . import config
from . import constants

View file

@ -12,6 +12,7 @@ from ..component_registry import R
from ..as_embedding import BaseAsEmbedding
Miss = tuple[int, str, str] # (result_index, text, cache_key)
_MAX_VECTOR_SPACE_ATTEMPTS = 3
@R.register("local")
@ -103,12 +104,25 @@ class LocalEmbeddingStore(BaseEmbeddingStore):
# -- Public API --
async def get_embeddings(self, input_text: list[str], **kwargs) -> list[np.ndarray | None]:
await self._sync_cache_space()
texts = [self._truncate(t) for t in input_text]
results, misses = self._partition_by_cache(texts)
if misses:
await self._fill_misses(misses, results, **kwargs)
return results
for attempt in range(1, _MAX_VECTOR_SPACE_ATTEMPTS + 1):
await self._sync_cache_space()
vector_space_id = self._cache_space
results, misses = self._partition_by_cache(texts)
stable = not misses or await self._fill_misses(misses, results, vector_space_id, **kwargs)
if stable and vector_space_id == self.vector_space_id == self._cache_space:
return results
if attempt == _MAX_VECTOR_SPACE_ATTEMPTS:
self.logger.warning(
f"Embedding vector space kept changing while computing a request; "
f"discarding all result(s) after {attempt} attempts",
)
else:
self.logger.info(
f"Embedding vector space changed while computing a request; "
f"discarding all result(s) and retrying ({attempt}/{_MAX_VECTOR_SPACE_ATTEMPTS})",
)
return [None] * len(texts)
# -- Batching --
@ -124,15 +138,26 @@ class LocalEmbeddingStore(BaseEmbeddingStore):
misses.append((idx, text, key))
return results, misses
async def _fill_misses(self, misses: list[Miss], results: list[np.ndarray | None], **kwargs) -> None:
vector_space_id = self._cache_space
async def _fill_misses(
self,
misses: list[Miss],
results: list[np.ndarray | None],
vector_space_id: str,
**kwargs,
) -> bool:
"""Fill every miss only while the request remains in one vector space."""
size = self.max_batch_size
for start in range(0, len(misses), size):
if vector_space_id != self.vector_space_id or vector_space_id != self._cache_space:
return False
batch = misses[start : start + size]
for idx, key, emb in await self._compute_batch(batch, **kwargs):
computed = await self._compute_batch(batch, **kwargs)
if vector_space_id != self.vector_space_id or vector_space_id != self._cache_space:
return False
for idx, key, emb in computed:
results[idx] = emb
if vector_space_id == self.vector_space_id == self._cache_space:
self._cache_put(key, emb)
self._cache_put(key, emb)
return True
async def _compute_batch(self, batch: list[Miss], **kwargs) -> list[tuple[int, str, np.ndarray]]:
texts = [text for _, text, _ in batch]

View file

@ -1,416 +0,0 @@
app_name: ReMe Daily Cookbook
workspace_dir: ${DAILY_PAPER_WORKSPACE_DIR:-reme_workspace}
timezone: Asia/Shanghai
language: zh
# This is a standalone application config. It intentionally does not inherit
# default.yaml and listens on a separate port so it can run beside ReMe.
service:
backend: http
host: ${DAILY_PAPER_HOST:-127.0.0.1}
port: ${DAILY_PAPER_PORT:-8001}
jobs:
index_update_loop:
backend: background
watch_dirs: [daily_dir, digest_dir]
watch_suffixes: [md, jsonl]
steps:
- backend: init_changes_step
monitor_type: file_store
monitor_name: default
dispatch_steps: [update_index_step]
- backend: watch_changes_step
dispatch_steps:
- backend: update_index_step
persist: false
auto_dream:
backend: base
description: "Auto-dream: consolidate recent daily notes into digest memory and interest topics."
parameters:
type: object
properties:
date:
type: string
description: "YYYY-MM-DD to scan; defaults to today in the configured timezone."
default: ""
hint:
type: string
description: "Optional guidance for extraction and integration."
default: ""
scan_days:
type: integer
description: "Number of recent daily directories to scan."
default: 2
max_units:
type: integer
description: "Maximum number of extracted memory units."
default: 5
topic_count:
type: integer
description: "Maximum number of interest topics to write."
default: 3
topic_diversity_days:
type: integer
description: "Previous interest-topic days used for de-duplication."
default: 7
steps:
- backend: dream_extract_step
file_catalog: dream
topic_session_id: interests
scan_days: 2
max_units: 5
- backend: dream_integrate_step
- backend: dream_topics_step
topic_count: 3
topic_diversity_days: 7
- backend: dream_finish_step
file_catalog: dream
auto_memory:
backend: base
description: "Auto-memory: record conversation facts into a daily note."
parameters:
type: object
properties:
messages:
type: array
description: "Conversation messages."
items:
type: object
session_id:
type: string
description: "Source conversation session identifier."
default: ""
memory_hint:
type: string
description: "Optional memory-writing guidance."
date:
type: string
description: "YYYY-MM-DD daily-note date; empty infers it from messages or current time."
default: ""
required: [messages]
steps:
- backend: auto_memory_step
reindex:
backend: base
description: "Wipe the derived search store and rebuild it from memory files."
watch_dirs: [daily_dir, digest_dir]
watch_suffixes: [md, jsonl]
parameters:
type: object
properties: {}
steps:
- backend: clear_store_step
- backend: init_changes_step
monitor_type: file_store
monitor_name: default
dispatch_steps: [update_index_step]
memory_search:
backend: base
description: "Long-term memory retrieval via hybrid workspace search (vector + BM25, RRF-fused)."
parameters:
type: object
properties:
query:
type: string
description: "Search query."
limit:
type: integer
description: "Maximum number of results."
default: 5
min_score:
type: number
description: "Minimum fused score."
default: 0.0
start_date:
type: string
description: "Optional inclusive start date (YYYY-MM-DD)."
end_date:
type: string
description: "Optional inclusive end date (YYYY-MM-DD)."
required: [query]
steps:
- backend: search_step
vector_weight: 0.7
candidate_multiplier: 5.0
expand_links: true
max_links_per_direction: 10
node_search:
backend: base
description: "Recall digest nodes for auto-dream de-duplication and linking."
parameters:
type: object
properties:
query:
type: string
description: "Candidate memory-node name and description."
limit:
type: integer
description: "Maximum number of digest nodes."
default: 20
required: [query]
steps:
- backend: node_search_step
vector_weight: 0.7
candidate_multiplier: 5.0
daily_list:
backend: base
description: "List notes under one day."
parameters:
type: object
properties:
date:
type: string
description: "YYYY-MM-DD; empty means today."
default: ""
steps:
- backend: daily_list_step
frontmatter_read:
backend: base
description: "Read a file's frontmatter."
parameters:
type: object
properties:
path:
type: string
description: "Workspace-relative path."
required: [path]
steps:
- backend: frontmatter_read_step
frontmatter_update:
backend: base
description: "Merge key-values into a file's frontmatter."
parameters:
type: object
properties:
path:
type: string
description: "Workspace-relative path."
metadata:
type: object
description: "Key-values to merge."
required: [path, metadata]
steps:
- backend: frontmatter_update_step
move:
backend: base
description: "Move or rename a workspace file and retarget inbound wikilinks."
parameters:
type: object
properties:
src_path:
type: string
description: "Workspace-relative source path."
dst_path:
type: string
description: "Workspace-relative destination path."
overwrite:
type: boolean
default: false
retarget:
type: boolean
default: true
required: [src_path, dst_path]
steps:
- backend: move_step
read:
backend: base
description: "Read a markdown file under the workspace."
parameters:
type: object
properties:
path:
type: string
description: "Workspace-relative markdown path."
start_line:
type: integer
end_line:
type: integer
required: [path]
steps:
- backend: read_step
with_neighbors: false
max_neighbors_per_direction: 10
write:
backend: base
description: "Create or overwrite a markdown file with frontmatter."
parameters:
type: object
properties:
path:
type: string
description: "Workspace-relative markdown path."
name:
type: string
description: "Frontmatter name."
description:
type: string
description: "Frontmatter description."
content:
type: string
description: "Markdown body."
metadata:
type: object
description: "Optional extra frontmatter fields."
required: [path, name, description, content]
steps:
- backend: write_step
daily_write:
backend: base
description: "Write a daily markdown note linked to its source conversation."
parameters:
type: object
properties:
name:
type: string
description: "Filename stem and frontmatter name."
description:
type: string
description: "Frontmatter description."
session_id:
type: string
description: "Source conversation session identifier."
content:
type: string
description: "Markdown body."
date:
type: string
description: "YYYY-MM-DD; empty means today."
default: ""
metadata:
type: object
description: "Optional extra frontmatter fields."
required: [name, description, session_id, content]
steps:
- backend: daily_write_step
edit:
backend: base
description: "Find and replace text in a markdown file."
parameters:
type: object
properties:
path:
type: string
description: "Workspace-relative path."
old:
type: string
description: "Text to replace."
new:
type: string
description: "Replacement text."
default: ""
required: [path, old, new]
steps:
- backend: edit_step
dingtalk_wait:
backend: background
supervisor: true
close_timeout: 10
steps:
- backend: dingtalk_wait_step
app_key: ${DINGTALK_APP_KEY:-}
app_secret: ${DINGTALK_APP_SECRET:-}
robot_code: ${DINGTALK_ROBOT_CODE:-}
worker_count: 4
builtin_tools: [bash]
job_tools:
- memory_search
- read
- write
- edit
- daily_list
- daily_write
- frontmatter_read
- frontmatter_update
components:
tokenizer:
default:
backend: regex
as_llm:
default:
backend: openai
model: ${LLM_MODEL_NAME:-qwen3.7-plus}
stream: true
context_size: 200000
max_retries: 3
credential:
api_key: ${LLM_API_KEY:-}
base_url: ${LLM_BASE_URL:-}
parameters:
max_tokens: 65536
thinking_enable: false
agent_wrapper:
default:
backend: agentscope
as_llm: default
builtin_tools: false
# as_embedding:
# default:
# backend: openai
# model: ${EMBEDDING_MODEL_NAME:-text-embedding-v4}
# dimensions: 1024
# max_retries: 0
# credential:
# api_key: ${EMBEDDING_API_KEY:-}
# base_url: ${EMBEDDING_BASE_URL:-https://dashscope.aliyuncs.com/compatible-mode/v1}
# parameters: {}
#
# embedding_store:
# default:
# backend: local
# as_embedding: default
# max_retries: 3
# quota_retry_delay: 60.0
file_graph:
default:
backend: local
file_catalog:
dream:
backend: local
file_chunker:
markdown:
backend: markdown
supported_extensions: [md]
embed_toc: true
max_ast_sections: 100
include_frontmatter_in_metadata: false
include_frontmatter_keys_in_metadata: []
jsonl:
backend: jsonl
supported_extensions: [jsonl]
max_lines_per_chunk: 1
keyword_index:
default:
backend: bm25
tokenizer: default
file_store:
default:
backend: local
store_name: local
# embedding_store: default
embedding_store: ""
keyword_index: default
file_graph: default

View file

@ -123,14 +123,6 @@ def test_default_config_keeps_frontmatter_chunk_metadata_opt_in():
) in (None, [])
def test_daily_cookbook_chunks_jsonl_one_line_at_a_time():
"""Daily cookbook keeps JSONL records as individually addressable chunks."""
cfg = _load_config("daily_cookbook.yaml")
jsonl = cfg["components"]["file_chunker"]["jsonl"]
assert jsonl["max_lines_per_chunk"] == 1
def test_parse_args_rejects_non_key_value_extra_argument():
"""Extra CLI arguments must use key=value syntax."""
with pytest.raises(ValueError, match="expected key=value"):

View file

@ -9,10 +9,8 @@ from unittest.mock import MagicMock
import pytest
from reme.components import ApplicationContext, R
from reme.components import ApplicationContext
from reme.components.agent_wrapper.base_agent_wrapper import BaseAgentWrapper
from reme.config.config_parser import _load_config
from reme.enumeration import ComponentEnum
from reme.steps.cookbook.dingtalk.wait import DingTalkWaitStep, _session_key
@ -203,54 +201,6 @@ async def test_final_reply_injects_only_configured_tools(tmp_path):
]
def test_daily_cookbook_registers_one_step_background_wait_job(monkeypatch):
for name in ("DINGTALK_APP_KEY", "DINGTALK_APP_SECRET", "DINGTALK_ROBOT_CODE"):
monkeypatch.delenv(name, raising=False)
config = _load_config("daily_cookbook")
job = config["jobs"]["dingtalk_wait"]
assert job["backend"] == "background"
assert job["steps"] == [
{
"backend": "dingtalk_wait_step",
"app_key": "",
"app_secret": "",
"robot_code": "",
"worker_count": 4,
"builtin_tools": ["bash"],
"job_tools": [
"memory_search",
"read",
"write",
"edit",
"daily_list",
"daily_write",
"frontmatter_read",
"frontmatter_update",
],
},
]
assert config["components"]["agent_wrapper"] == {
"default": {
"backend": "agentscope",
"as_llm": "default",
"builtin_tools": False,
},
}
assert R.get(ComponentEnum.STEP, "dingtalk_wait_step") is DingTalkWaitStep
def test_daily_cookbook_passes_dingtalk_environment_to_step(monkeypatch):
monkeypatch.setenv("DINGTALK_APP_KEY", "app-key")
monkeypatch.setenv("DINGTALK_APP_SECRET", "app-secret")
monkeypatch.setenv("DINGTALK_ROBOT_CODE", "robot-code")
step = _load_config("daily_cookbook")["jobs"]["dingtalk_wait"]["steps"][0]
assert (step["app_key"], step["app_secret"], step["robot_code"]) == (
"app-key",
"app-secret",
"robot-code",
)
@pytest.mark.asyncio
async def test_stream_client_closes_when_background_stop_is_set(monkeypatch):
websocket = _WebSocket()

View file

@ -65,6 +65,31 @@ def test_strip_injected_parameters_hides_keys_from_schema():
assert "date" in job.parameters["properties"]
def test_search_injection_exposes_only_query():
parameters = {
"type": "object",
"properties": {
"query": {"type": "string"},
"limit": {"type": "integer"},
"min_score": {"type": "number"},
"start_date": {"type": "string"},
"end_date": {"type": "string"},
},
"required": ["query"],
}
injected = {
"limit": 20,
"min_score": 0.0,
"start_date": None,
"end_date": "2026-07-20",
}
stripped = BaseAgentWrapper._strip_injected_parameters(parameters, injected)
assert stripped["properties"] == {"query": {"type": "string"}}
assert stripped["required"] == ["query"]
# -- AgentScope wrapper -----------------------------------------------------------
@ -139,7 +164,7 @@ class _RecordingWrapper(BaseAgentWrapper):
super().__init__(**kwargs)
self.calls: list[dict] = []
async def reply(self, inputs, **kwargs) -> dict:
async def reply(self, _inputs, **kwargs) -> dict:
self.calls.append(kwargs)
return {"session_id": "s-1", "last_message": {}, "result": "ok"}

View file

@ -509,24 +509,95 @@ def test_cache_space_is_rechecked_after_async_load(monkeypatch, tmp_path):
run(go())
def test_completed_request_only_writes_to_its_active_cache_space():
"""A v3 request must not populate v4 after the provider switches back to v3."""
def test_whole_request_retries_after_vector_space_changes_between_batches():
"""Completed batches must be discarded when a later batch changes vector space."""
async def go():
embedding = OpenAIAsEmbedding(name="t_space_write_race", backend="openai", model="v3", dimensions=2)
store = LocalEmbeddingStore(name="t_local_write_race")
embedding = FakeAsEmbedding()
embedding.vector_space_id = "v3"
store = LocalEmbeddingStore(name="t_local_write_race", max_batch_size=1, enable_cache=False)
store.as_embedding = embedding
store._cache_space = embedding.vector_space_id
calls = 0
async def compute_after_round_trip(_batch, **_kwargs):
embedding.model = FakeProviderModel("v4")
store._cache_space = embedding.vector_space_id
embedding.model = FakeProviderModel("v3")
return [(0, "key", np.array([3.0, 0.0], dtype=np.float16))]
async def switch_during_second_batch(batch, **_kwargs):
nonlocal calls
calls += 1
if calls == 2:
embedding.vector_space_id = "v4"
idx, _text, key = batch[0]
version = 3.0 if calls < 3 else 4.0
return [(idx, key, np.array([version, 0.0], dtype=np.float16))]
store._compute_batch = compute_after_round_trip
await store._fill_misses([(0, "text", "key")], [None])
store._compute_batch = switch_during_second_batch
results = await store.get_embeddings(["first", "second"])
assert calls == 4
assert store._cache_space == embedding.vector_space_id
for result in results:
np.testing.assert_array_equal(result, np.array([4.0, 0.0], dtype=np.float16))
run(go())
def test_whole_request_rereads_cache_after_vector_space_changes(monkeypatch, tmp_path):
"""A cache hit from the old space must not survive a later provider switch."""
async def go():
monkeypatch.setattr(
LocalEmbeddingStore,
"component_metadata_path",
property(lambda _self: tmp_path),
)
embedding = FakeAsEmbedding()
embedding.vector_space_id = "v3"
store = LocalEmbeddingStore(name="t_local_cache_race", enable_cache=True)
store.as_embedding = embedding
store._cache_space = embedding.vector_space_id
first_key = store._cache_key("first")
store._cache[first_key] = np.array([3.0, 0.0], dtype=np.float16)
calls = 0
async def switch_on_miss(batch, **_kwargs):
nonlocal calls
calls += 1
if calls == 1:
embedding.vector_space_id = "v4"
version = 3.0 if calls == 1 else 4.0
return [(idx, key, np.array([version, 0.0], dtype=np.float16)) for idx, _text, key in batch]
store._compute_batch = switch_on_miss
results = await store.get_embeddings(["first", "second"])
assert calls == 2
for result in results:
np.testing.assert_array_equal(result, np.array([4.0, 0.0], dtype=np.float16))
run(go())
def test_whole_request_stops_retrying_when_vector_space_keeps_changing():
"""Continuous configuration churn must discard the whole request instead of blocking forever."""
async def go():
embedding = FakeAsEmbedding()
embedding.vector_space_id = "v3"
store = LocalEmbeddingStore(name="t_local_write_churn")
store.as_embedding = embedding
store._cache_space = embedding.vector_space_id
calls = 0
async def change_space_every_time(_batch, **_kwargs):
nonlocal calls
calls += 1
embedding.vector_space_id = f"v{calls + 3}"
return [(0, "key", np.array([float(calls), 0.0], dtype=np.float16))]
store._compute_batch = change_space_every_time
results = await store.get_embeddings(["text"])
assert calls == 3
assert results == [None]
assert "key" not in store._cache
run(go())

View file

@ -7,6 +7,7 @@ import tomllib
from types import ModuleType
from packaging.requirements import Requirement
from packaging.version import Version
import pytest
REPOSITORY = Path(__file__).resolve().parents[2]
@ -41,17 +42,13 @@ def test_studio_packages_have_independent_identity() -> None:
assert studio_config["project"]["name"] == "reme_studio"
assert npm_config["name"] == "@agentscope-ai/reme_studio"
assert studio_config["project"]["version"] == npm_config["version"]
assert main_config["project"]["optional-dependencies"]["as"] == ["agentscope[model-ollama]==2.0.6"]
assert main_config["project"]["optional-dependencies"]["as"] == ["agentscope[model-ollama]==2.0.7"]
assert main_config["project"]["optional-dependencies"]["web"] == ["reme_studio"]
assert main_config["project"]["optional-dependencies"]["core"].count("reme-ai[as]") == 1
assert main_config["project"]["optional-dependencies"]["core"].count("reme_studio") == 1
assert main_config["project"]["optional-dependencies"]["qwenpaw"] == [
"reme-ai[core]",
"reme-auto-fin>=0.1.1",
"reme-daily-paper>=0.1.1",
]
assert auto_fin_config["project"]["version"] == "0.1.1"
assert daily_paper_config["project"]["version"] == "0.1.1"
assert "qwenpaw" not in main_config["project"]["optional-dependencies"]
assert auto_fin_config["project"]["version"] == "0.1.2"
assert daily_paper_config["project"]["version"] == "0.1.2"
assert main_config["tool"]["setuptools"]["packages"]["find"]["include"] == ["reme", "reme.*"]
assert "reme_studio*" in main_config["tool"]["setuptools"]["packages"]["find"]["exclude"]
@ -153,14 +150,16 @@ def test_auto_fin_license_matches_repository() -> None:
).read_text(encoding="utf-8")
def test_auto_fin_requires_reme_core() -> None:
"""Install the optional runtime packages needed while loading Auto Fin's entry points."""
def test_auto_fin_requires_reme_base() -> None:
"""Keep the plugin dependency limited to ReMe's public base package."""
config = tomllib.loads((REPOSITORY / "plugins" / "auto-fin" / "pyproject.toml").read_text(encoding="utf-8"))
requirements = [Requirement(value) for value in config["project"]["dependencies"]]
reme_requirements = [requirement for requirement in requirements if requirement.name == "reme-ai"]
assert len(reme_requirements) == 1
assert set(reme_requirements[0].extras) == {"core"}
assert not reme_requirements[0].extras
assert Version("0.4.1.8") not in reme_requirements[0].specifier
assert Version("0.4.1.9") in reme_requirements[0].specifier
def test_daily_paper_license_matches_repository() -> None:
@ -171,12 +170,14 @@ def test_daily_paper_license_matches_repository() -> None:
def test_daily_paper_declares_runtime_dependencies() -> None:
"""Keep Daily Paper's ReMe feature set and PDF parser explicit in its own distribution."""
"""Keep Daily Paper's minimal ReMe and PDF dependencies explicit."""
config = tomllib.loads((REPOSITORY / "plugins" / "daily_paper" / "pyproject.toml").read_text(encoding="utf-8"))
requirements = [Requirement(value) for value in config["project"]["dependencies"]]
by_name = {requirement.name: requirement for requirement in requirements}
assert set(by_name["reme-ai"].extras) == {"core"}
assert not by_name["reme-ai"].extras
assert Version("0.4.1.8") not in by_name["reme-ai"].specifier
assert Version("0.4.1.9") in by_name["reme-ai"].specifier
assert "pypdf" in by_name

View file

@ -76,7 +76,7 @@ def test_list_plugins_marks_configured_plugins(monkeypatch, tmp_path, capsys):
)
monkeypatch.setattr(plugin_cli_module, "_enabled_plugins", lambda _config: {"auto-fin"})
assert plugin_cli_module.plugin_cli(["list", "--config", "daily_cookbook"]) == 0
assert plugin_cli_module.plugin_cli(["list", "--config", "default"]) == 0
output = capsys.readouterr().out
assert "ENABLED" in output