From 8c487981648025576fa08f30c595ee7720deded5 Mon Sep 17 00:00:00 2001 From: Sen Huang <48879559+ployts@users.noreply.github.com> Date: Thu, 28 May 2026 15:17:23 +0800 Subject: [PATCH] feat(jobs): add digester step for knowledge distillation from daily notes (#261) * feat(jobs): add digester step for knowledge distillation from daily notes * feat: add initial implementation * refactor(jobs): remove unnecessary blank lines in digester and synchronizer * feat: add initial implementation --- docs4/auto_dream_design.md | 578 +++++++++++++++++++++ docs4/auto_link_design.md | 209 ++++++++ docs4/auto_maintain_design.md | 190 +++++++ docs4/auto_memory_design.md | 197 ++++++++ docs4/structure.md | 782 +++++++++++++++++++++++++++++ reme4/steps/jobs/__init__.py | 10 + reme4/steps/jobs/digester.py | 300 +++++++++++ reme4/steps/jobs/digester.yaml | 114 +++++ reme4/steps/jobs/protocol.md | 112 +++++ reme4/steps/jobs/synchronizer.py | 276 ++++++++++ reme4/steps/jobs/synchronizer.yaml | 144 ++++++ 11 files changed, 2912 insertions(+) create mode 100644 docs4/auto_dream_design.md create mode 100644 docs4/auto_link_design.md create mode 100644 docs4/auto_maintain_design.md create mode 100644 docs4/auto_memory_design.md create mode 100644 docs4/structure.md create mode 100644 reme4/steps/jobs/__init__.py create mode 100644 reme4/steps/jobs/digester.py create mode 100644 reme4/steps/jobs/digester.yaml create mode 100644 reme4/steps/jobs/protocol.md create mode 100644 reme4/steps/jobs/synchronizer.py create mode 100644 reme4/steps/jobs/synchronizer.yaml diff --git a/docs4/auto_dream_design.md b/docs4/auto_dream_design.md new file mode 100644 index 00000000..4a548160 --- /dev/null +++ b/docs4/auto_dream_design.md @@ -0,0 +1,578 @@ +# auto-dream 设计(digest 沉淀:物理 + 图) + +> 本文档记录 reme4 中 **auto-dream**(digest 沉淀知识层)的设计讨论 —— 含物理布局、图模型(节点 + 边)、入流端 digester(G\*)。 +> +> 配套阅读: +> - `structure.md` §1.2(数据视角)/ §2(三层存储)/ §3.5(digest 动作)/ §7(L4 模块) +> - `auto_memory_design.md`:auto-memory(daily 实时事件)是 dream 的入流之一(G1 scope 拉取);auto-memory 写完即对 dream 可见 +> - `auto_maintain_design.md`:digest 的组织 / 重组 / 写入运行时(M split / D 检测 / CAS 写入协议) —— 本文档定义节点 + 边模型与 G\*,maintain 定义运行时 +> - `auto_link_design.md`:auto-link 是 dream 写完之后的后置增强(实体识别 + wikilink 写回);复用 maintain 的 CAS 协议 +> +> **四层对应**:reme 服务整体四份设计 —— auto-memory(daily 入流)/ **auto-dream(本文档:digest 沉淀 + content link + G\*)** / auto-maintain(digest 组织端:split + 检测 + CAS)/ auto-link(背景图关系增强)。`structure.md` §3.5-3.6 的 L4 action 视角:dream 对应 `digest` action(`resource + daily → digest`),maintain 对应 `maintain` action(`digest → digest`)。auto-dream 承担"空闲整理"(报告 §5.2):把 daily / resource 的散点抽出共性、合并重复、生成总结,落到 `digest//.md`。 +> +> **核心立场**:dream 定义**模型层**(节点 + 边 / 守恒规则)+ **生成侧**(G\* create_or_update)+ **召回**(SearchStep);组织端(M split / D 检测 / CAS / 时序)归 `auto_maintain_design.md`;后置增强(实体识别 / wikilink 写回)归 `auto_link_design.md`。三方共享 §1.5 节点 + 边模型 + §1.5.5 边守恒 + §1.5.4 F-invariants。 +> +> **关键收敛**:digest 不分"逻辑层"。整个 digest = **物理布局(浅桶 + flat .md)** + **一张图(节点 + 边)**。没有 hub / topic / leaf 之分 —— 所有 .md 文件都是同一种节点,内容决定它扮演什么角色(主题概览 / 概念定义 / 方法描述 / 实体记录 ...)。"主题"是从图中涌现的,不是结构性宣告的。 + +--- + +## 0. 问题陈述 + +digest 是 agent 长期记忆的"组织化沉淀"层,与三层架构的另两层职责互补: + +| 层 | 组织主轴 | 形态 | +|---|---|---| +| resource/ | 时间(`/`) | 外部原始资料,不可变 | +| daily/ | 时间 + 任务(`//`) | agent 任务过程,半可变 | +| **digest/** | **语义** | **跨任务知识,可重组** | + +digest 的核心问题: +- **物理布局**:文件系统怎么组织?(目录 / 文件 / 路径) +- **节点与边**:概念粒度 / 链接形态 / 主题如何涌现? +- **演化机制**:从碎片到网络的过程谁负责?(digester 创建或更新节点 / maintainer 拆分过载节点 / D 信号写后 inline 检测) + +--- + +## 1. 已对齐决策 + +### 1.1 节点粒度:Atomic 节点优先 + +| 项 | 决策 | +|---|---| +| **粒度** | 一个文件 = 一个原子单元(概念 / 方法 / 实体 / 案例 / 原则 / 主题概览) | +| **节点角色** | 由内容决定,不由 frontmatter 类型标记;一个节点扮演"主题概览"还是"具体方法",看它的 body 写了什么 | +| **风格参考** | Zettelkasten:atomic note + 高密度 wikilink 网络 | + +**理由**: +1. **节点粒度 = retrieve 精度上限** —— semantic 检索召回 "一个原子单元" 远比召回 "一个 5000 字的主题文档" 信噪比高;agent 上下文窗口经不起粗粒度文档塞满 +2. **wikilink 在 atomic 粒度才真有意义** —— `[[jwt-rotation]]` 指向"一个具体方法"比指向"auth 主题文档"精确一个数量级,这也是 I-4(wikilink 唯一跨层载体)能撑起来的前提 +3. **物理目录纯做 navigate,图承担关系** —— 两套机制各司其职,不互相绑架 + +**Tradeoff**: +- 文件数量爆炸(一个领域几百节点)→ 浅桶物理归档 + 路径作 ID,wikilink 走完整路径写法 +- 节点会频繁演化(新材料 update 已有节点 / 节点过载触发 split)→ G\* 去重质量、SearchStep 召回、写后 inline 检测都得到位 + +### 1.2 物理几何:浅桶(shallow bucket,**固定集合**) + +| 项 | 决策 | +|---|---| +| **物理布局** | `digest//.md`,bucket 一层(顶多两层),桶内 flat | +| **bucket 角色** | **仅承担物理归档与 OS-level 浏览锚点**;不承担语义本体角色 —— 主题由图中的节点表达 | +| **bucket 内** | 不再分子目录,所有节点平铺 | +| **bucket 集合** | **固定预定义,不由 digester / maintainer 动态生成** | +| **bucket 主页节点** | **不强制存在**;若 split 在该 bucket 内累积出层级(parent 节点天然中心性高),parent 节点天然成为浏览主页(纯约定,非架构必需) | +| **新节点归属** | digester (G4) 只能从已有桶集合里**挑选**;LLM 不能造新桶 | +| **集合来源** | vault 配置(opinionated default + 消费层可改),与 schema/prompt 同属服务消费层 | + +**理由**: +- 物理浏览有"主题轮廓"(打开 `digest/auth/` 能看到这一族节点),不像纯 flat 那样毫无锚点 +- 节点不被深路径绑死("属 auth/jwt 还是 auth/session"这种归属焦虑被消解 —— 一个节点可以同时被多个主题通过 wikilink 引用) +- maintain 操作面坍缩到只剩 split:节点过载就拆,不做跨节点重组(详 §1.5 / §2) +- **固定集合的关键意义**: + - LLM 在 G4 桶决定时只做"分类",不做"造类" —— 决策面坍缩,错率大幅下降 + - bucket 集合作为**预先约定的物理归档规则**,跨任务跨时间稳定;不会出现 "auth-stuff" / "auth" / "authentication" 三个语义重叠的桶共存 + - 与 reme 核心立场一致:bucket 集合是消费层契约,reme 不自主造桶 +- **bucket 主页节点不强制的关键意义**: + - 旧设计中 root hub 是架构必需(检测稳定锚点 / vault 总览结构性入口);新设计中这些角色都由图中心性自然承担,不需要为每个 bucket 强制创建一个空架子节点 + - 主页节点的"主页"地位是涌现的:某节点入度高 / 中心性高 → 它就是浏览入口 + - 若一个 bucket 完全没节点,它就只是个空目录;不需要先造一个 placeholder + +**未归类节点**:digester 抽到一个原子单元但找不到合适的专属桶时,**不允许造新桶**;统一落入兜底桶 `digest/general/`。general 桶在固定集合内是一等公民,详见 §3.7。 + +### 1.3 节点身份:路径即 ID + +| 项 | 决策 | +|---|---| +| **ID 载体** | **vault-relative 路径**(含 `.md`)即节点身份 —— `digest/auth/jwt-rotation.md` | +| **wikilink 写法** | `[[digest/auth/jwt-rotation.md]]`(literal,与 `wikilink_handler.py` 默认形态对齐;不隐含 `.md`,无 short-form 补全) | +| **`name` frontmatter** | 文件名 basename(不含扩展名),与文件名同步 —— 检索 hint / 人读标签,**不当 ID 用** | +| **同名冲突** | 同 bucket 内文件名冲突 → 文件系统层断言;**不需要独立 D6 信号** | +| **rename 成本** | 一次 `wikilink_handler.retarget_links(old_path, new_path)`,机制现成 | +| **跨桶移动** | F-1 已禁止;若必须做(人工介入修错桶),走一次 retarget | + +**理由**: +- F-1(0 文件移动)+ 平铺 + 下层 immutable 后,slug abstraction 的核心价值(移动鲁棒性)蒸发;只剩下"wikilink 短形式"这一项收益,但代价是 file_graph slug 索引 + D6 冲突检测 + alias 表 + retrieve 透明展开,**净亏** +- `wikilink_handler.py` docstring 自己写的就是 *Recommended form: full path relative to the vault with extension* —— literal 匹配,无 short-link 补全;路径作 ID 与核心库默认完全对齐 +- 完整路径前缀 `digest/auth/` 给 LLM 读写时提供语义 context(知道节点在哪个桶),不全是负担 +- provenance wikilink 反指 daily/resource 本来就用路径,统一后整个 vault 一种 wikilink 形态,不必区分"slug 形态 vs 路径形态" + +### 1.4 链接语义:基础 wikilink + 可选 Dataview 谓词 + +参考实现:`reme4/utils/wikilink_handler.py` + `reme4/schema/file_link.py`。 + +| 项 | 决策 | +|---|---| +| **link 基础形态** | `[[.md]]` —— target 字面取(literal,不隐含 `.md`,不自动短链补全);路径即 ID(详 §1.3) | +| **alias / image** | `[[path.md\|alias]]`(显示文本)/ `![[image.png]]`(图片资源)—— rewrite 时 alias 保持 | +| **anchor 不引入** | digest 设计层**不使用** `[[path.md#section]]` —— atomic 节点 + child 边界已充当精度替代品(详 §1.3 / §3.13);`wikilink_handler` 仍可解析 anchor 字面(供其它消费层),但 digest 不生成、不依赖、不在 split 时迁移 anchor | +| **可选谓词(Dataview 风格)** | 行级: `predicate:: [[path.md]]` / 内联: `[predicate:: [[path.md]]]`;**谓词写在 `[[]]` 外**,不是 `[[predicate::path]]` | +| **谓词标识符** | `[A-Za-z][A-Za-z0-9_]*`(如 `is_a` / `extends` / `causes` / `references`);词表**开放**,任意标识符 | +| **未类型化合法** | 绝大多数 wikilink **不加** predicate;`predicate=None` 是默认 / 常态 | +| **edge 唯一性键** | `(target_path, predicate)`(二元组);同源同标但不同 predicate = 不同边。`FileLink.target_anchor` 字段在 schema 中保留(供其它消费层),digest 层永远写 `None` | +| **类型信息载体** | 节点 frontmatter `kind` + 边 `predicate`(双轨可选);二者都是**消费层 schema 提示**,reme 核心解析 / 存储 / 索引,但**不读它们做结构决策** | +| **plugin 层扩展** | transclusion / 引用图谱视图等留给消费层加,reme 核心不固化语义 | + +**理由**: +- 与 I-4 "wikilink 是唯一跨层载体" 对齐 —— reme 核心永远只看机械拓扑 +- 与 [[reme4_schema_layering]] 一致 —— 类型语义(无论是节点 `kind` 还是边 `predicate`)都是消费层契约,reme 核心不固化 +- predicate 走 Dataview 而非内嵌:`[[]]` 内容保持"纯目标"(rewrite / retarget 不必感知 predicate);predicate 是文本上的**装饰位**,与 wikilink 解耦 +- `kind` 与 predicate 严格只是内容标签:reme 核心**只有节点这一种结构类型 + 边这一种结构关系**,kind / predicate 永远不参与"hub / topic / leaf"这类结构角色判断 + +**reme 核心对 predicate 的"透明"边界**(关键): +- G7(横向 link)、retrieve 中心性 —— 都**聚合所有 predicate** 算,不分桶 +- 只有 edge 唯一性 / 反向索引会用到 predicate(否则 `[[A]]` 和 `is_a:: [[A]]` 会被当作同一条边互相覆盖) +- 消费层若要按 predicate 做更精细的推理(如"taxonomic 路径只走 `is_a` 边"),自己读 `FileLink.predicate` 即可 + +### 1.5 图模型与节点演化 + +**核心模型**:digest = **物理布局(浅桶 + flat .md)** + **一张图(节点 + 边)**。 + +| 维度 | 形态 | +|---|---| +| **节点** | 每个 .md 文件 = 一个节点;无结构性 kind,角色由 body 内容决定(主题概览 / 概念定义 / 方法描述 / 实体记录 / 案例 ...) | +| **边** | 基础 `[[.md]]` wikilink(路径即 target,详 §1.3 / §1.4);可选 Dataview 谓词 `predicate:: [[path.md]]` / `[predicate:: [[path.md]]]` 写在 `[[]]` 外;边唯一性键 = `(target, predicate)`;**digest 层不引入 anchor**(详 §1.4 / §3.13);**reme 核心结构决策不读 predicate**(详 §1.4) | +| **多归属** | 一个节点可被多个其它节点引用,也可指向多个其它节点;**不存在"单父"约束** | + +**演化只做两件事**: +1. **G\* create_or_update**:新材料进入,LLM 提取原子单元 → 命中已有节点就 update 该节点 body(语义守恒地融合新旧),否则新建节点 +2. **M split**:节点累积过载(token / 主题离散度超阈值)→ LLM 把它拆成 parent overview + N 个 children,parent 文件原地保留作 overview,children 是新文件 + +"主题概览节点 / 摘要节点"不是一种 kind,也不是 maintainer 主动涌现的产物 —— 它是 split 的副产品(parent 节点天然成为该 cluster 的 overview)。 + +#### 1.5.1 单一节点 / 单一边 + +**节点 frontmatter** —— 只有保留字段: + +| 字段 | 内容 | 用途 | +|---|---|---| +| `name` | 文件名 basename(不含扩展名),与文件名同步 | 检索 hint / 人读标签(I-4 不再用它做身份;路径才是 ID,详 §1.3) | +| `description` | 一句话 | 标题 / 检索 hint | +| (可选)`kind` | concept / method / case / entity / topic / ... | **消费层 schema 提示**,reme 核心透明,不读它做结构决策 | +| body | 任意内容 | 一句定义 / 一段方法 / 一篇主题概览 / 一份案例,皆可 | + +`hub__` / `topic__` 前缀**不存在**;文件名自然命名(`auth-fundamentals.md` / `jwt-rotation.md` / `jwt-overview.md`)。"overview 节点"靠内容形态识别,不靠前缀。 + +**边的形态** —— 与 §1.4 一致,这里给最小汇总: + +| 维度 | 形态 | +|---|---| +| 基础 | `[[.md]]`(无谓词;常态;`predicate=None`) | +| 可选谓词 | `predicate:: [[path.md]]`(行级)/ `[predicate:: [[path.md]]]`(内联);谓词在 `[[]]` 外 | +| alias | `[[path.md|display-text]]`(rewrite 时 alias 保持) | +| image | `![[image.png]]`(资源引用,不是知识边) | +| **不引入 anchor** | digest 层不使用 `#section`;详 §1.4 / §3.13 | +| **边唯一性键** | `(target_path, predicate)` —— 同源同标不同 predicate = 不同边 | +| **角色识别** | 默认无谓词时由端点内容形态推断;有谓词时谓词即角色标签(消费层语义,核心不读) | + +> **关键收敛**:reme 核心**只有节点 + 边两种结构类型**;kind / predicate 都是内容标签,绝不参与 hub / topic / leaf 这类结构角色。 + +#### 1.5.2 节点演化:create_or_update + split + +``` +[入流] material 进入(daily / resource) + │ + ▼ + G1 scope:选哪些 daily/resource 进入本轮 + │ + ▼ + G2 提取原子单元(LLM,可能产 N 个候选) + │ + ▼ + 对每个候选: + │ + ▼ + G* create_or_update(LLM 决策点) + ├─ 语义相似查 → 拉相似候选节点(top-k) + ├─ LLM 判:候选中有"同概念节点"吗? + │ ├─ 有 → update 路径 + │ │ (a) 把新内容融入已有 body(语义守恒重写) + │ │ (b) 加 provenance 反指 + │ │ (c) 必要时加 / 改 wikilink + │ └─ 无 → create 路径 + │ G3 路径 / G4 bucket / G6 provenance / G7 横向 link / 写 body + │ + ▼ + 写入(机械) + +[D3 检测] 写后立即:G\* / split 写完 body 顺手 inline 检测(token 阈值 → 超阈值则 LLM 判离散度;详 `auto_maintain_design.md` §4) + │ + ▼ + D3 派发候选节点(F-4 一次一个) + │ + ▼ + M split(LLM + 机械) + ├─ LLM 把节点 body 拆成 N 个 cluster(每个是个原子单元) + ├─ parent 文件原地保留 → body 重写为 overview + 列出 children wikilinks + ├─ 每个 child 创建新文件(文件名 / bucket / body 由 LLM 给) + ├─ children 各自加 [[.md]] 反向链接 + ├─ 边守恒机械校验:`(parent_new ∪ ∪children) ⊇ parent_old` 出边集合(F-11 / E-2) + └─ inbound 链 `[[.md]]` 不动(F-10 / E-3) —— digest 层无 anchor 链,无需 retarget +``` + +**关键**: +- **G\* update 改 subject body(语义守恒重写)** —— 新材料融入已有节点正文,要求 LLM 守住"只增不删 / 不改原意",老内容不能丢;**写入前机械校验出边强守恒**(`new outbound ⊇ old outbound`,详 §1.5.5 E-1) +- **G\* 不改其它节点正文** —— 只动 subject;不像旧 M-E 会到邻居 body append wikilink +- **M split 不改其它节点正文** —— 只动 parent(重写为 overview)+ 新建 children +- **inbound 在 split 时一律不动** —— digest 设计不引入 anchor,inbound 全是裸链 `[[.md]]`,parent 路径未变即天然有效;后续 G\* 进入若 LLM 觉得 child 粒度更合适,直接加新边到 child(F-10) + +#### 1.5.3 走一个具体例子 + +**场景**:`digest/auth/` 桶,初始只有几个零散 auth 节点,没有 jwt-rotation。 + +**第 1 轮 G\***:某 daily 提到 "JWT rotation:每 24 小时换密钥,旧密钥保留 1 小时窗口给未过期 token"。 +- 语义查 → 没找到 jwt-rotation 节点 +- 走 create 路径 → 新建 `jwt-rotation.md`,body = 一段 200 字的 rotation 描述 + +**第 2 轮 G\***:另一个 daily 提到 "JWT rotation 的 grace period 通常是 1-2 小时"。 +- 语义查 → 命中 `jwt-rotation`(高相似) +- LLM 判:这是同概念,走 update 路径 +- 把 grace period 信息**融入** `jwt-rotation.md` body(不只是 append): + ``` + Before: "...旧密钥保留 1 小时窗口..." + After: "...旧密钥保留 1-2 小时 grace period(典型值,具体看 token 寿命)..." + ``` +- body 略增长,加一条 provenance 反指 + +**第 N 轮 G\***:经过几个月,各种 daily 持续 update `jwt-rotation` —— 加了密钥派生算法、加了 RS256/HS256 区别、加了 key rotation 失败处理、加了 with-leeway 实践、加了 monitoring 建议 ... + +`jwt-rotation.md` body 现在 ~3500 token,涵盖:轮换策略 / 密钥派生 / 算法选择 / 失败处理 / 监控。 + +**触发**:D3 检测 token > 2000 阈值 → 派发候选。 + +**M split**:LLM 拉 `jwt-rotation.md` body + frontmatter,判断主题离散度(5 个相对独立的子主题),决定拆: +- parent: `jwt-rotation`(留下,body 重写为 overview) +- children: `jwt-key-derivation` / `jwt-algorithm-selection` / `jwt-rotation-failure-handling` / `jwt-rotation-monitoring`(4 个新文件) +- "with-leeway 实践"内容并入 parent overview(粒度太细不单独拆) + +**执行后**: + +``` +digest/auth/ +├── jwt-rotation.md ← 文件原地;body 重写为 overview +│ description: JWT 轮换策略总览 +│ body: JWT 轮换的核心是 X,主要环节包括: +│ 密钥派生 [[digest/auth/jwt-key-derivation.md]] +│ 算法选择 [[digest/auth/jwt-algorithm-selection.md]] +│ 失败处理 [[digest/auth/jwt-rotation-failure-handling.md]] +│ 监控告警 [[digest/auth/jwt-rotation-monitoring.md]] +│ (参考:with-leeway 实践 ...) +├── jwt-key-derivation.md ← 新建,body 来自原 jwt-rotation 拆出片段 +│ → [[digest/auth/jwt-rotation.md]] ← child 反指 parent +├── jwt-algorithm-selection.md ← 同上 +│ → [[digest/auth/jwt-rotation.md]] +├── jwt-rotation-failure-handling.md ← 同上 +│ → [[digest/auth/jwt-rotation.md]] +├── jwt-rotation-monitoring.md ← 同上 +│ → [[digest/auth/jwt-rotation.md]] +└── ...(其它 auth 节点不变) +``` + +**inbound 不动**:之前指 `jwt-rotation.md` 的所有外部 wikilink(无论裸链还是 typed)都仍然指 `[[digest/auth/jwt-rotation.md]]`。如果后续某外部节点写新材料时 LLM 觉得 child 粒度更合适,直接 G\* 时新加 `[[digest/auth/jwt-key-derivation.md]]` 这种边即可 —— 不强求 split 时即时重定向。 + +**继续演化**: +- 若某个 child(如 `jwt-key-derivation`)被持续 update,某天也长到过载 → D3 又触发 → 它再次 split,自然涌现第三层 +- 若某 child 长期空 / 0 入度 / 0 update —— 不主动删(没有 dissolve 操作);除非人工介入 + +#### 1.5.4 操作的核心约束 + +| # | 约束 | 含义 | +|---|---|---| +| **F-1** | **0 文件移动** | G\* / split 都不移动现有文件;split 创建的是**新文件**,parent 文件原地 | +| **F-2** | **改正文限定 subject** | G\* update 改 subject node body(语义守恒重写,不改其它节点);M split 改 parent body(重写为 overview)+ 创建 children body;**没有任何操作改"其它节点正文"** | +| **F-3** | **maintainer 只做 split** | 没有 summarize / merge / re-edge / link / unify / dissolve;过载 → 拆 | +| **F-4** | **一次一个候选** | M split 一次拆一个节点;G\* 一次处理一个原子单元(N 个候选 = N 次 G\*) | +| **F-5** | **不确定时不动** | G\* 拿不准是 create 还是 update → 倾向 create(不污染已有节点);split 拿不准 cluster 边界 → 不拆 | +| **F-7** | **多归属合法** | 一个节点可被多个其它节点引用,也可指向多个其它节点;**没有"单父"约束** | +| **F-10** | **inbound 目标节点不动** | 所有 inbound 都是裸链 `[[.md]]`(digest 不引入 anchor,详 §1.4 / §3.13);split 时全部保持不动,parent 路径未变即天然有效;后续 G\* 进入时 LLM 可自由选择更精细 target(直接加新边到 child) | +| **F-11** | **wikilink 是 body 的一部分** | 不存在"独立的边" —— 边的所有迁移都是 body 文本变化的副作用;reme 核心机械算子只感知字符层,语义责任在 LLM(G\* / split prompt)+ 守恒校验(outbound diff 机械验证;详 §1.5.5) | + +#### 1.5.5 边的迁移规则 + +**前提**:wikilink 是 body 的一部分(F-11)。"边"不是独立抽象 —— body 一变,边就跟着变。reme 核心**没有"修边"算子**,边的所有变化都是 body 文本编辑的副作用。 + +但语义守恒不能放任 LLM:守恒责任在 prompt + 机械校验,不在算子。 + +**3 类边按"在哪类操作中变化"区分**: + +| # | 类别 | 规则 | 谁负责 | +|---|---|---|---| +| **E-1** | **G\* update 节点出边**(subject 自身) | **强守恒**:新 body 出边集合 ⊇ 原 body 出边集合(`(target, predicate)` 二元组比对,**predicate 一并守住**);不满足 → LLM 重试或拒写 | LLM(prompt 强约束)+ 机械校验(outbound diff) | +| **E-2** | **split parent 出边**(parent body 拆解) | parent overview + N 个 children 各持一段,原 parent 出边按内容自然分配到 parent overview + children;**机械校验合计守恒**:`(parent_new ∪ ∪children_outbound) ⊇ parent_old` | LLM(split prompt)+ 机械校验 | +| **E-3** | **inbound wikilink** `[[.md]]` | split 时**不动** —— 仍指 parent;后续 G\* 进入若 LLM 觉得 child 粒度更合适,直接加新边到 child(F-10) | 不动 | + +**provenance 不单列一类**:节点反指上游 daily/resource 的 wikilink 是 body 正文的一部分(§3.9),由 LLM 在 G\* / split prompt 中自然写出 —— 跟其它 body wikilink 走同一套规则:G\* update 走 E-1 强守恒(老 provenance 链不能丢,新材料追加新 provenance),split 走 E-2 合计守恒(parent 全量 provenance ⊆ parent_overview ∪ ∪children_outbound)。reme 核心**没有** provenance 专用算子。 + +**inbound anchor 这一类不存在**:digest 设计层不引入 anchor(详 §1.4 / §3.13),所有 inbound 都是裸链,走 E-3 即可,无需机械 retarget 子流程。 + +> **关键拆分**: +> - **守恒**(E-1 / E-2):LLM 写正文时不能丢边;靠 prompt + 写后 outbound diff 校验 +> - **保守**(E-3):没有信号说一定要变;不变的代价 = 后续 G\* 自然纠正,变的代价 = 错信号大量假阳;选不变 + +**机械 outbound diff 校验**(E-1 / E-2)伪码: + +``` +write_subject_body(subject, new_body): + old_outbound = extract_links(old_body) # set of (target, predicate) + new_outbound = extract_links(new_body) + missing = old_outbound - new_outbound + if missing: + # LLM 漏了原边 —— 重试一次 + new_body = llm_retry_with_missing(missing) + new_outbound = extract_links(new_body) + if old_outbound - new_outbound: + raise ConservationViolation(...) # 拒写,记 audit,等人介入 + write(subject, new_body) +``` + +机械层只做集合比对,**不判断"为什么丢了"** —— 那是 LLM 的事。 + +**强守恒(集合包含)而非等价**:`new ⊇ old` 是"新材料融入,老知识保留"的最小契约 —— 允许加新边(新关联),不允许减边(老内容不能丢);等价(`new == old`)会拒绝任何新出边,update 失去意义。 + +**predicate 守住** —— `[[A]]` ↔ `is_a:: [[A]]` 视为不同 key,升降级走显式 audit 路径,不走默认。重排 / 改 alias / 加新边都不被拦下(集合相同或只增)。 + +#### 1.5.6 图模型 vs 建子目录 + +| 维度 | 建子目录(深树) | 图模型 + split 演化 | +|---|---|---| +| 物理变化 | 移动文件,改路径 | 0 文件移动(F-1);split 只创建新文件 | +| wikilink 影响 | 路径 ID 模型下要全图 retarget(代价大) | 0 影响(parent 路径未动);split 不触发任何 retarget | +| 主题归属 | 一个节点只能属一棵子树 | 一个节点可同时属多个主题(被多源 wikilink) | +| 撤销成本 | 移回文件 + 重建上下文 | 删 children + parent body 还原(手工) | +| 演化路径 | 子树重组痛苦 | parent 只增不减,children 是 parent 拆出的快照 | +| navigate | 浏览目录树 | 任意节点入手沿出边漫游;parent 节点是天然中心 | +| retrieve 精度 | 路径反映主题但与 link 无关 | 节点中心性 + 内容形态共同决定权重 | + +#### 1.5.7 多级结构自然涌现 + +split 是节点的**局部操作**(只看一个过载节点),多级深度自然涌现: + +``` +digest/auth/ +├── auth-fundamentals.md ← 早期写下,~1500 token,稳定 +├── jwt-rotation.md ← 第一次 split:body 从 3500 token 重写为 overview +├── jwt-key-derivation.md ← 第一次 split 的 child +├── jwt-key-derivation-hkdf.md ← 二次 split:jwt-key-derivation 累积 update 后过载,再拆 +├── jwt-key-derivation-pbkdf2.md ← 二次 split 的 child +├── ... +``` + +**物理仍是浅桶(1 层),"层级"由 split 链 + 节点中心性自然承载**。每一级的过载条件、决策机制、执行步骤完全相同 —— 没有"二级 split"特殊逻辑,只有"过载节点的 body 可以被 split 进一步拆"。 + +#### 1.5.8 retrieve 时节点怎么参与 + +| query | 期望返回 | +|---|---| +| "JWT 怎么轮换" | 优先 `jwt-rotation`(具体 overview)+ children(如 `jwt-key-derivation`) | +| "auth 体系" | 优先中心性高的节点(`auth-fundamentals` / `jwt-rotation` 等被多次 update / 是 split parent 的节点) | +| "auth 有什么子主题" | 沿高中心性节点的入/出邻居遍历;parent 节点优先返回 | +| "vault 里都有什么" | 各 bucket 中心性最高的节点(自然形成 vault 总览) | + +**加权策略**(opinionated default,消费层可改): +- 节点权重 = base(=1.0) × intent 调节 × 中心性增益 +- query 含"概览 / 主题 / 入门 / 全景"等**元意图**时,中心性高的节点加权(intent 调节 > 1) +- 中心性低 / body 短的具体节点权重稳定(默认 1.0,不被压低) + +**topological traverse**: +- 沿 wikilink 自由走(不区分边类型 / predicate) +- 经过中心性高的节点默认**不强行展开**(否则一次 traverse 把整族 children 拉进来);agent 可显式深入 + +**中心性的天然来源 = split parent**:被拆过的节点是 parent,自然有 children 反向链接它,中心性自然高 —— 不需要单独维护 `kind: hub` 标记。 + +--- + +## 2. 完整能力集 + +### 2.0 设计目标 + +**让图的形状持续匹配实际知识的语义结构,在最小变更面 + 渐进演化的前提下,使任意尺度的知识访问都能命中合适粒度的节点。** + +这个目标直接来自结构本身的设计意图 —— 浅桶 + 单一节点类型 + 单一边类型 + create_or_update + split 的组合,每一项都是为它服务。能力集的入选标准:**对至少一个验证维度有贡献**。 + +| 维度 | 含义 | 失败示例 | +|---|---|---| +| **形状匹配** | 节点中心性 / 边连接 / 节点邻域反映知识间的实际语义关系 | 一个节点 token 5000+ 长期不拆;同主题节点彼此 0 链接;同概念被建成多个独立节点 | +| **最小变更面** | 不重写其它节点正文,不大规模移文件,不破坏现有 wikilink | 任何"全图重组"或"批量改其它节点正文"的方案 | +| **任意尺度访问** | 具体方法节点 / overview 节点 / 节点邻居遍历都能命中 | 全 flat,主题级 query 命中不到东西 | + +**显式排除**(不在目标内,避免能力集内卷): +- ❌ "完美归簇" —— F-5 留白,不确定就不动 +- ❌ "实时一致" —— 异步 / eventual,节点写完不必立刻 split +- ❌ "零冲突 / 零违反" —— invariants 检测 + 事后修复,不追求永不发生 +- ❌ 替消费层做检索 / 决策 —— digest 自治边界止于"维持图的形状" +- ❌ 跨节点重组(merge / re-edge / unify / dissolve)—— 简化模型不做这些;同概念二次进入由 G\* update 路径处理 + +**演化只有两件事**:G\* create_or_update(入流型,新材料融入)+ M split(后台,过载就拆)。detection 派生信号驱动这套循环。 + +### 2.1 生成侧:digester(入流型) + +| # | 能力 | 服务 | 性质 | 何时发生 | +|---|---|---|---|---| +| **G1** | **scope 决定**:选哪组 daily/resource 进入本轮蒸馏 | 形状匹配(决定形状从哪生长) | LLM | digester 启动 | +| **G2** | **原子单元抽取**:从 scope 中识别值得沉淀的原子单元(N 个候选) | 形状匹配 + 任意尺度 | LLM | 核心环节 | +| **G\*** | **create_or_update**:对每个候选,**多路召回(SearchStep:vector + keyword + 邻接展开,RRF 融合,scope 限 `digest/`)** → LLM 看完整候选池 → 终判 create / update / drop;create 路径走 G3/G4/G6/G7 + 写 body;update 路径融入已有 body(语义守恒重写)+ 自然追加 provenance(详 §3.10) | 形状匹配(去重内置)+ 最小变更面 | LLM(决策)+ 机械(召回 + 守恒写入) | 每个候选 | +| **G3** | **路径命名**(create 路径):在 G4 选定 bucket 内,文件名同 bucket 唯一(fs 层断言);风格与同主题节点一致 | 任意尺度(可寻址) | LLM(命名)+ 机械(同 bucket 文件名冲突 → 拒写) | create 时 | +| **G4** | **bucket 落地**(create 路径):从固定集合中挑选;找不到合适专属桶 → 落 `general/`(§3.7) | 形状匹配(物理归档) | LLM(读 bucket 列表) | create 时 | +| **G6** | **provenance 写入**(create 与 update):新节点 body 内联反指上游 daily/resource 的 wikilink;update 时 LLM 在融入新材料时自然追加新 provenance 链,旧 provenance 链由 E-1 守恒校验保住(§3.9) | 任意尺度(跨层访问) | LLM(prompt 引导写出 `[[daily/...]]` / `[[resource/...]]`)+ 机械(outbound diff 校验) | 写入时 | +| **G7** | **横向 link**(create 时):新节点链到相关的已有节点(出边);update 时也可加新 link | 形状匹配 + 任意尺度 | LLM | 写入时 | + +**关键边界**: +- **G\* 是入流唯一改 body 的操作**,且**只改 subject node** —— update 改的是同概念那个节点自己,不改其它节点 +- **G\* update 必须语义守恒**:LLM 重写 body 时只能"融入"新内容,不能删除已有信息(只增不删 / 不改原意) +- **0 出边节点合法**(G7 没识别到合适邻居),后续 G\* 进入时其它节点可以反向链回来 —— 不强求 LLM 一次性给全 +- **G\* 漏判去重**(把同概念建成新节点)→ 不主动兜底,接受重复;若 vault 累积明显的同概念重复,可由 auto-link 离线 audit 工具产报告(详 `auto_link_design.md` §1.3 L4) +- **digester 不做 split** —— split 是后台 M 操作 + +### 2.2 组织侧 / 检测 / 写入并发 → `auto_maintain_design.md` + +M split / D 检测信号(D1 / D3 / D10)/ 阈值校准 / D3 写后触发模型 / G\* / split / auto-link L1 三方共用的 CAS 写入协议 / split provenance / 时序 / 后门 —— 全部归 `auto_maintain_design.md`。 + +dream 保留**模型层**(§1 节点 + 边 + 守恒规则)+ **生成侧**(§2.1 G\*)+ **召回**(§3.10 SearchStep);maintain 负责**组织 / 运行时**(split + D + CAS + 时序)。两份文档共享 §1.5 节点 + 边模型、§1.5.5 边守恒、§1.5.4 F-invariants。 + +| 在 maintain 文档中 | 内容 | +|---|---| +| §1 | M split 能力卡 + 关键边界 | +| §2 | 检测信号 D1 / D3 / D10 | +| §3 | 阈值校准 | +| §4 | D3 写后触发模型 | +| §5 | CAS 写入协议(三方共用) | +| §6 | split 时 provenance | +| §7 | G\* / split / auto-link L1 时序 | +| §8 | 后门(暂缓) | + +### 2.4 边界协议(谁不能做什么) + +| 边界 | 内容 | 来源 | +|---|---|---| +| digester ∩ maintainer | digester 不做 split;maintainer 不做原子单元抽取 / 新具体节点 create | 入流 vs 自维护职责分离 | +| digester → 其它节点 | G\* update 改 subject node body,**不改其它任何节点正文** | F-2 | +| digester → "摘要 / overview" | digester 不为做 overview 而创建节点;它产的节点都是具体原子单元;overview 是后续 split 的副产品 | F-3 | +| maintainer → 其它节点 | M split 改 parent body(重写为 overview)+ 创建 N 个 children body;**不改任何其它节点** | F-2 | +| maintainer → inbound 链 | split 时**全部不动** —— digest 不引入 anchor,inbound 一律是裸链 `[[.md]]`,parent 路径未变 | F-10 / E-3 | +| digester → 边守恒 | G\* update 写新 body 前,机械对比 old/new outbound:`new ⊇ old`((target, predicate) 二元组);失败 → LLM 重试一次,再失败拒写 | F-11 / E-1 | +| maintainer → 边守恒 | split 写新 parent body + N children body 前,机械对比:`(parent_new ∪ ∪children_outbound) ⊇ parent_old`;失败 → LLM 重试或拒写 | F-11 / E-2 | +| 全员 → typed link predicate | wikilink 的 predicate 是 edge identity 的一部分;G\* update / split 不能丢 predicate(`is_a:: [[A]]` 必须保持;否则被守恒校验当作 drop edge + add edge 拦下);predicate 升 / 降级走显式 audit 路径 | F-11 / §1.4 | +| 全员 → resource/daily | 都不能改 | I-2 / I-3 | +| 全员 → 节点 rename | rename = 一次 `wikilink_handler.retarget_links(old_path, new_path)`;无 alias 表,无透明展开 | §1.3 | +| 全员 → provenance link | 永远必须可达(I 不变量 + D10 检测) | I-1 / I-4 | +| 全员 → kind 字段 | reme 核心**透明**:不读取 frontmatter `kind` 做结构决策;`kind` 是消费层 schema 提示 | [[reme4_schema_layering]] | +| 全员 → predicate 谓词 | reme 核心**结构决策不读**:G7 / 中心性都聚合所有 predicate 算;edge 唯一性 / 反向索引会用到 predicate(防同源同标不同 predicate 互相覆盖);未类型化 link 是默认形态 | [[reme4_schema_layering]] / §1.4 | + +--- + +## 3. 待对齐边界点(后续讨论清单) + +### 3.1 G\* update 的语义守恒边界(已收敛) + +**决策**:**LLM 重写整段**(prompt 强约束"语义守恒,只增不删 / 不改原意;冲突标注 `> 注:不同来源记载...`,不擅自仲裁")+ **机械守恒校验**(详 §1.5.5 E-1)。校验失败 LLM 重试一次,再失败拒写 + audit。 + +首版可先用 append 起步(出边集合天然 ⊇,守恒校验自动通过),prompt 工程量小;成熟后切到重写。 + +### 3.2 maintainer 的人 / agent 后门 → `auto_maintain_design.md` §8 + +### 3.3 G\* 与 split 的时序 → `auto_maintain_design.md` §7 + +### 3.4 节点 kind / 边 predicate(已收敛) + +reme 核心**只有节点 + 边两种结构类型**: +- frontmatter `kind` 字段(若存在)= 消费层的**节点内容标签**(concept / method / case / entity / topic / ...),reme 不读它做结构决策 +- 边 `predicate`(Dataview 风格,若存在)= 消费层的**边关系标签**(`is_a` / `extends` / `causes` / ...),reme 解析 / 存储 / 参与 edge 唯一性,但**结构决策不读**(G7 不区分 predicate;中心性不区分) +- "overview 节点"角色靠图位置(高中心性 / 是 split parent)+ body 内容形态识别,不靠 frontmatter 或 predicate 标记 +- 未类型化 wikilink 是默认 / 常态形态 + +详见 §1.4 / §1.5.1 / §2.4。 + +### 3.5 retrieve 时的权重策略(部分收敛 → §1.5.8) + +- 加权策略:节点权重 = base(=1.0) × intent 调节 × 中心性增益;query 含元意图("概览 / 主题 / 入门 / 全景"等)时,中心性高的节点加权 +- traverse 默认不强行展开高中心性节点(防止整族 children 拉进来);agent 可显式深入 +- 不按 frontmatter `kind` 加权;中心性天然来源 = split parent(详 §1.5.8) + +剩余待定:**中心性算法选型**(eigenvector / PageRank / 简单入度,初期可用入度,后续校准)。 + +### 3.6 M split 时的 provenance 处理 → `auto_maintain_design.md` §6 + +### 3.7 bucket 集合管理 + +§1.2 已定:bucket 集合**固定预定义**,不由 digester / maintainer 动态生成。补足细节: + +- **定义位置**:`vault.yaml` 顶层 `digest.buckets:` 是源 + 自动生成 `digest/_buckets.md` 作为人/LLM 可读视图;digester G4 时读后者作为 prompt context +- **初始化**:opinionated default(通用桶 `concept` / `method` / `pattern` / `tool` / `domain` 等 + 必带 `general`);消费层可改桶名,但 **`general` 不可删**(否则 G4 失去兜底) +- **扩展路径**:reme 不主动提议扩 bucket(对比旧设计的 maintainer 周期建议已 DROPPED);用户编辑 `vault.yaml` 后下次 G4 即生效 + +**未归类节点处理**(G4 找不到合适专属桶时):**统一落入 `digest/general/`**。 + +| 维度 | 内容 | +|---|---| +| **bucket 名** | `general`(固定集合一等公民,默认包含) | +| **语义** | "通用主题 / 暂无专属归属" —— 合法常态,非故障状态 | +| **路径** | `digest/general/.md`,与其它 bucket 完全等同 | +| **节点演化** | 与其它 bucket 一致 | +| **错桶后续** | 不主动跨桶 move(无 D9 / M-D);若严重,人工 mv + `retarget_links(old, new)` | + +**为什么是 `general` 而不是 `_unclassified`**:`_unclassified` 暗示待处理状态,LLM/人都想清理掉;`general` 是合法常态,G4 选桶时是显式合法选项而非 fallback 故障路径。 + +**已排除**:拒绝写入(候选丢失)/ 强行选最近似专属桶(本体污染,general 反而更安全)。 + +### 3.8 检测阈值校准 → `auto_maintain_design.md` §3 + +### 3.9 provenance 载体形态(已收敛) + +**决策**:**provenance wikilink 嵌在节点 body 正文中**(inline body prose),由 LLM 在 G\* / split prompt 里自然写出,跟其它 body wikilink 完全同形,**靠语义维护**。reme 核心没有 provenance 专用算子。 + +**写出形态**: +- 行文中自然带出处:"... 该模式最早出现在 [[daily/2026/05/15.md]] 的实践中" +- 或专门一段总结式段落,内含若干 wikilink 指向上游 +- 可选 predicate(`derived_from:: [[daily/2026/05/15.md]]`),不强制 + +**机械保护**: +- E-1 守恒(G\* update):旧 provenance 不在新 outbound 集合 → 重试或拒写,机械兜底 +- E-2 守恒(split):provenance 跟着对应内容段自然分配到 parent overview / children,合计守恒 +- D10 检测:provenance 断裂 = D1 断链子集(target 命中 `daily/` / `resource/` 前缀);D10 严重程度高于普通 D1(I 不变量) + +**E-5 / G6 等"provenance 专用机制"全部坍缩** —— 不再单列。Prompt 必须要求"出处用 `[[...]]` 形式表达"(纯散文会被守恒校验视为丢边)。 + +### 3.10 G\* 语义查后端(已收敛) + +**决策**:**直接复用 `SearchStep`(`reme4/steps/index/search.py`)** —— 多路召回并发(vector + keyword)+ RRF 融合 + file_graph 邻接展开,把**完整候选池交给 LLM 终判**;G\* 入口不做 bucket 粗筛(LLM 拥有完整跨桶视野,可识别"概念错分到 general"或"跨桶同概念";三路信号 RRF 融合后噪声可控)。 + +**召回链路**: +1. 候选原子单元(摘要 / 关键词)→ `SearchStep`(`search_filter={"path_prefix": "digest/"}`,I-2/I-3 daily/resource 不入池) +2. `SearchStep` 内部:`vector_search` + `keyword_search` 并发 → RRF 融合 → `expand_links` 邻接展开 → 返回 top-`limit` FileChunks(含 path / 行号 / 邻接节点) +3. 候选池整体喂 LLM,按 path 自然聚合(同节点多 chunk 命中 = 强信号);终判输出节点路径 +4. LLM 终判 create / update / drop;update 选定 subject node → 走 E-1 守恒重写 + +**provenance 不依赖召回** —— G\* / split prompt 让 LLM 直接写 `[[daily/...]]` / `[[resource/...]]`(§3.9)。 + +**索引维护**:沿用 `update_index` step,G\* / split 写 body 后调一次刷该节点索引;启动一次全建(`clear_and_scan` 已就绪),损坏走全建兜底。 + +**默认参数**(可按 dogfooding 调):`limit` 5~10 / `vector_weight` 0.7 / `expand_links` on / `min_score` 0(初版不过滤,LLM 兜底)。 + +### 3.11 G\* / split 写入并发 / 原子性 → `auto_maintain_design.md` §5 + +### 3.12 D3 触发模型 → `auto_maintain_design.md` §4 + +### 3.13 anchor 不引入 wikilink 设计(已收敛) + +**决策**:**digest 设计层不使用 `[[path.md#section]]` 形态** —— wikilink 只有 `[[path.md]]`(可选 alias / 谓词),anchor 不进入 digest。当 LLM 想"指向某个具体子主题"时,正确做法是让那个子主题升级为独立节点(必要时通过 split),而不是在过载 parent 内部用 anchor 凑合。 + +**连锁简化**: +- E-4(inbound anchor 机械 retarget)整类**消失**;split 流程末尾不再扫 inbound anchor 子流程;`{anchor → child}` 映射输出从 split prompt 中移除 +- 边唯一性键从三元组 `(target, predicate, anchor)` 简化为二元组 `(target, predicate)` +- §1.5.5 边迁移类别从 4 类(E-1..E-4)简化为 3 类(E-1..E-3) +- `FileLink.target_anchor` 字段在 schema 中保留(供其它消费层),digest 层永远写 `None` + +**Prompt 约束**:G\* / split 的 prompt 必须明确告知 LLM 写 wikilink 时不带 `#section`。若 LLM 仍写出 `[[path.md#section]]`,wikilink_handler 仍能解析,守恒校验只看 `(target, predicate)`,不会形成"丢边"风险 —— 但 anchor 在 digest 层无语义。若引用方依赖某 anchor 锚定具体段落,表明该内容应升级为 child 节点。 + +--- + +## 4. 下一步 + +本文档覆盖 dream 模型 + 生成侧 + 召回(G\* / 节点+边模型 / SearchStep)。组织端实现清单(M split / D 检测 / CAS)见 `auto_maintain_design.md` §10。 + +1. **digester 流程图**(G1 / G2 / G\* 的实际编排;G\* 内 create / update 路径分流;**召回直接复用 `SearchStep`**(vector + keyword + 邻接展开,RRF 融合,scope `digest/`)—— 详 §3.10;**G\* update 写入前 outbound diff 守恒校验** — E-1) +2. **rename 路径设计**(`wikilink_handler.retarget_links(old_path, new_path)` 已就绪;封装为单步 step 入口,无 alias 表 / 无透明展开) +3. **bucket 集合配置**(`vault.yaml` schema / 默认桶模板 / `general` 兜底机制 / `_buckets.md` 视图生成) +4. **边守恒校验工具**(`extract_links` 已就绪;新增 outbound diff 比较器 + LLM 重试编排 + ConservationViolation audit 事件) +5. **provenance prompt 规范**(G\* / split 引导 LLM 写 `[[daily/...]]` / `[[resource/...]]` —— §3.9) + +实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_memory_design.md` / `auto_maintain_design.md` / `auto_link_design.md` 共同作为契约依据。 diff --git a/docs4/auto_link_design.md b/docs4/auto_link_design.md new file mode 100644 index 00000000..a4ec3dbc --- /dev/null +++ b/docs4/auto_link_design.md @@ -0,0 +1,209 @@ +# auto-link 设计(背景实体识别 + wikilink 写回) + +> 本文档记录 reme4 中 **auto-link** 的设计讨论 —— 在已写入节点之间发现隐含关系,把这些关系作为 `[[...]]` wikilink **写回 body**,形成可见、可编辑的图结构增强。 +> +> 配套阅读: +> - `structure.md` §1.2(三层数据视角)/ §4(retrieve 三种问法) +> - `auto_memory_design.md`:auto-link 可反向扫 daily event,补实体 wikilink(daily → digest) +> - `auto_dream_design.md`:wikilink 模型(§1.4 边语法 / §1.5 演化 / §1.5.5 边守恒 E-1 / E-2 / E-3);auto-link 借这套基础设施 +> - `auto_maintain_design.md`:CAS 写入协议(§5);auto-link L1 写回与 dream G\* / maintain split 三方共用同一套 CAS +> +> **三层对应**:reme 服务整体三层 —— auto-memory / auto-dream / **auto-link(本文档)**。auto-link 是图关系的**后置增强** —— 在已落地的 vault 上做实体识别 + wikilink 写回,补足 content link(写记忆时由 LLM 直接产生的 `[[...]]`)在长 tail 隐含关系上的盲区。 +> +> **核心立场**:auto-link **写回 body**,不只是产报告。生成的 wikilink 是**可见、可编辑**的(写在 Markdown 文件里),agent / 人可后续 curate。auto-link 不引入新材料,纯 additive 插入 wikilink,天然满足 E-1 守恒;复用 dream 的 CAS 写入协议,不引入新基础设施。 + +--- + +## 0. 问题陈述 + +content link(`auto_dream_design.md` G\* / split 写入时由 LLM inline 产生的 `[[...]]`)解决了"写记忆时显式的关系"。但有一类关系不会在 inline 写入时自然涌现,需要后台扫描已写入的 vault 才能识别: + +1. **历史 body 的实体未链接** —— G\* update 时 LLM 关注新材料融入,可能忽略已有 body 中某个未链接的实体(例如 body 提到 "JWT" 但没写 `[[digest/auth/jwt-overview.md]]`) +2. **跨节点 / 跨桶的隐含关联** —— 节点 A 提到 "rate limit",但 `digest/api/rate-limit.md` 是后来才被 split 创建 → A 写入时没机会建立这条边 +3. **同主题未连 / 同概念重复** —— G\* 漏判去重把同概念建成两个节点;或两个主题相关但 0 链接的节点彼此不知晓 + +auto-link 承担这部分:**后台扫描已写入节点 → 实体识别 / 候选挖掘 → wikilink 写回 body**。 + +--- + +## 1. 已对齐决策 + +### 1.1 与 content link 的边界 + +| 维度 | content link(在 dream) | auto-link(本文档) | +|---|---|---| +| 何时产生 | 写记忆 inline:G\* update / M split prompt | 后台扫描:离线 / 周期 / 触发后异步 | +| 由谁产生 | LLM 在 dream 写入流中顺手写出 | LLM 在 auto-link 扫描流中识别后写出 | +| 输入 | 新材料 + 召回候选节点 | 已写入 body + 全 vault 索引 | +| 改 body | 是(重写整段 body) | 是(纯 additive 插入 wikilink,不改文字) | +| 守恒 | E-1 强守恒(out ⊇ old) | E-1 天然满足(纯增) | +| 用途 | 写入即关系明示 | 弥补 inline 漏判,挖掘长 tail 关系 | + +### 1.2 写回模型:纯 additive,复用 dream CAS + +auto-link 写回是**纯 additive** 操作 —— 在已有 body 文字中找到实体 mention,替换为 wikilink 形态: + +``` +Before: "JWT 轮换的核心是密钥派生 ..." +After: "[[digest/auth/jwt-rotation.md|JWT 轮换]]的核心是[[digest/auth/jwt-key-derivation.md|密钥派生]] ..." +``` + +| 维度 | 决策 | +|---|---| +| **alias 必须保留原文** | `[[path.md\|<原文>]]` 形态;原文一字不改 —— 守住"不改写其它节点正文" (`auto_dream_design.md` §1.5.4 F-2) 的精神 | +| **predicate 默认为空** | auto-link 默认产生无谓词 wikilink;升 typed link 走 L3(详 §1.3) | +| **不引入 anchor** | 与 dream 一致(`auto_dream_design.md` §1.4 / §3.13);target 永远是节点路径 | +| **CAS 写入** | 完全复用 `auto_maintain_design.md` §5 的 read-stamp + CAS-write 协议(冲突重做 ≤ 3 次) | +| **E-1 守恒** | 纯 additive:`new outbound = old outbound ∪ {new wikilinks}`;`new ⊇ old` 天然满足,守恒校验默认通过 | +| **rollback** | 若 auto-link 误插入(例如 entity mention 是同名歧义),走标准 edit 或 retarget 撤销;auto-link 不维护"我插过哪些"audit log(留给 SDK 决定) | + +**为什么是 additive 而不是重写**: +- additive = 0 文字风险(原文不变,只在原 mention 周围加 `[[ | ]]` 包装) +- 重写 = 触发完整 E-1 守恒校验 + LLM 重写整段语义守恒 prompt + 多次 LLM 调用 = 跟 G\* update 重复 +- additive 失败可见:产生坏 wikilink 时,人/agent 直接编辑 body 修就行 + +### 1.3 候选挖掘类型(L1-L4) + +| # | 类型 | 描述 | 写回形态 | +|---|---|---|---| +| **L1** | **实体识别**(主路径) | 扫 body,识别已是 digest 节点的实体名(模糊匹配 + 语义召回);未被 wikilink 化的 mention → 加 `[[path.md\|]]` | additive wikilink 插入 | +| **L2** | **同主题未连**(旧 D7) | 两个 digest 节点谈相关主题但 0 wikilink → 候选 add link;LLM 判后在 body 末尾追加一句引用 | additive(在合适位置 / 节末追加 `参见 [[other.md\|other]]`)| +| **L3** | **隐含 predicate 推导** | 已有 `[[A]]` 但 LLM 可推断关系类型(`is_a` / `causes` / `extends` / ...)→ 升级为 typed link | 改 `[[A]]` → `is_a:: [[A]]`(predicate 升降级走显式 audit,详 §2.1)| +| **L4** | **重复语义检测**(旧 D8) | 两个节点描述同一概念但被独立 create(G\* 漏判去重)→ 候选 merge | **不写回**;产报告 + 提示人/agent 触发 G\* update 路径手工合并 | + +**L1 是主路径** —— 它是 auto-link 最核心、最频繁、最高 ROI 的操作:每个 digest 节点写完后,后台扫一遍 body,找未链接的已知实体,additive 加 wikilink。 + +**L2-L3 是辅助** —— 周期扫,产候选,LLM 终判,写回部分(L2 节末追加 / L3 升 predicate)。 + +**L4 不写回** —— 节点合并是结构改动,影响 E-1 守恒边界 + inbound 链路 + provenance 链路,不适合自动写;auto-link 只产报告,人/agent 决定走 G\* update 路径解决。 + +### 1.4 触发节奏 + +| 模式 | 何时 | 适用 | +|---|---|---| +| **inline post-write**(默认) | 每次 G\* update / M split 写完 body → enqueue auto-link L1 job(异步,FIFO,CAS 保护)| L1 实体识别;反应即时,与 D3 写后检测同节奏 | +| **周期 batch**(可选)| cron(daily / weekly)扫全 vault | L2 / L3 候选挖掘;成本可控 | +| **手动触发** | SDK / 人显式调用 | 全量重扫 / 修复 | + +**L1 inline 的必要性**:新 split 出的 child 节点立即被既有 body 引用(用 wikilink 而非纯 mention)的关键 = 写入即扫描;不 inline 会让"刚创建的 child 节点"在很长时间内只有 split parent 一个 inbound,中心性失真。 + +**已排除**: +- inline 时同步 auto-link(阻塞 G\* return)—— 时延不可接受;auto-link 始终异步 +- 所有 L\* 都 inline —— L2-L3 候选挖掘 RTL 跨节点,成本高,只适合 batch +- 全 cron 唯一触发 —— L1 滞后过久,新节点孤岛 + +### 1.5 中心性算法(retrieve 加权依赖) + +retrieve 时节点权重 = base × intent 调节 × **中心性增益**(详 `auto_dream_design.md` §1.5.8)。中心性需要 auto-link 这一层提供 —— content link 给底子,auto-link 补 long tail,二者合起来才是完整的图。 + +| 选项 | 优点 | 缺点 | +|---|---|---| +| **简单入度** | 实现最简;split parent 入度天然高;auto-link L1 加边后入度即时反映 | 不区分"权威节点"vs"被随手提的节点";高入度 ≠ 高权威 | +| **PageRank** | 经典;权威性传递 | 实现复杂 + 增量更新成本(每次写边重算成本高,需 incremental algorithm)| +| **eigenvector centrality** | 与 PageRank 相近 | 同上 | + +**首版决策**:**简单入度**(file_graph 已有 inbound 链表,O(1) 查);auto-link L1 加边后入度立刻更新,split parent 自然涌现高入度。dogfooding 后视 retrieve 质量演进。 + +中心性是 retrieve 时**查询时计算**,不预存: +- file_graph 已建反向索引(inbound),计算 `len(inbound(node))` 是 O(1) +- 不预存避免"加边后中心性陈旧"问题 +- PageRank 演进时可加增量计算 + 周期 refresh + +--- + +## 2. 待对齐边界点 + +### 2.1 L3 predicate 升降级的 audit + +L3 把 `[[A]]` 升级为 `is_a:: [[A]]` 时,**改了 edge identity** —— `(target, None)` 变成 `(target, "is_a")`,在 E-1 守恒视角下 = 删一条边 + 加一条边: + +``` +old outbound: {(A, None)} +new outbound: {(A, "is_a")} +diff: missing = {(A, None)}; added = {(A, "is_a")} +``` + +不打 audit 走默认会被守恒校验拦下(`missing != ∅` → 重试 / 拒写)。 + +**决策方向**: +- L3 写入必须打 audit flag(消费层意图:升级 predicate,允许 drop + add 同时发生) +- audit flag 由 reme4 step 暴露(`maintainer_step(action="predicate_upgrade", from=..., to=...)`),不放在普通 write 路径 +- 普通 G\* / auto-link L1 写入永远不带 audit flag,守恒校验照常严格 + +详细 audit flag 接口形态留到 SDK 阶段。 + +### 2.2 多歧义实体识别 + +L1 扫 body 找 "JWT" 这个 mention,vault 中有 `digest/auth/jwt-overview.md` 和 `digest/payment/jwt-payment-flow.md` 两个 candidate: + +候选方案: +- LLM 上下文判 —— 把 body 周围段落给 LLM,选最相关 target +- 跳过模糊 case —— L1 只处理 unambiguous mention,歧义 case 留人/agent +- 全部链 —— `[[overview]][[payment-flow]]`,后续人 curate + +**首版**:LLM 上下文判(每个候选 candidate 提供 description / 周围若干节点 summary,LLM 选择 top-1 或 drop);成本可接受(扫描已是离线 batch)。 + +### 2.3 auto-link 写回与 G\* / split 的并发 + +auto-link 写 body 走 §1.2 CAS,但有特殊情况: +- 同节点同时被 G\* update 与 auto-link L1 写入 → CAS 协议自动序列化 (`auto_maintain_design.md` §5):后到者重做 +- auto-link L1 写完后立刻被 G\* update 覆盖(G\* 重写 body) → 看 G\* prompt 是否守住 auto-link 加的 wikilink(E-1 强守恒 → 守住) +- auto-link L1 与 D3 派发的 split job 同节点并发 → split 先到 / 后到都不影响最终拓扑(split 把 body 拆成 parent + children,auto-link 加的 wikilink 跟着对应内容段自然分配到 parent / child) + +**结论**:CAS + E-1 + E-2 守恒已覆盖所有并发场景,auto-link 不需要新协调机制。 + +### 2.4 跨 vault / 跨进程 + +M0 单 reme 实例 + 单 vault,auto-link 走内进程 enqueue;多实例 / 跨进程留 M1+(同 `auto_maintain_design.md` §5)。 + +### 2.5 实体识别 vs 现成 NER 库 + +L1 实体识别可选: +- LLM 直接扫(贵但灵活,与 digest 节点同构) +- 现成 NER 库(spaCy 等)预筛 + LLM 终判(快但 entity 类型与 digest 节点形态可能不匹配) +- 纯字符串匹配(已知节点名字 + 简单变体)+ LLM 终判 ambiguity + +**倾向**:从纯字符串匹配 + LLM 终判 ambiguity 起步(实现最简,效果可能已经够好);视 dogfooding 决定是否引入 NER 库。 + +--- + +## 3. 与其它层的协作 + +| 上下游 | 关系 | +|---|---| +| ← **auto-dream** | dream 写完一个节点 → 通过 inline post-write enqueue auto-link L1(§1.4);auto-link 用 dream 的 CAS 协议 | +| ← **auto-memory** | auto-link 可反向扫 daily event,把实体识别成 `[[digest/...]]`(daily → digest);auto-memory 写入端不主动调 auto-link,触发同 dream 路径 | +| → **digest body** | 主要写入对象 —— L1 additive 加 wikilink / L2 节末追加引用 / L3 升 predicate(走 audit) | +| → **daily body** | auto-link 扫 daily event 时同样可加 `[[digest/...]]`(I-2 daily 单作者需协调:auto-link 应在 event 关闭后才动该 event,不与 active event 并发改;实现细节留 step 层处理) | +| → **resource body** | I-3 immutable;auto-link **不写 resource**(reading-only) | +| → **L4 候选 report** | L4 重复语义检测产报告,落 `audit//auto_link_l4.md`(具体路径 / 形态留 step 层) | + +--- + +## 4. 与 auto-dream 模型的引用关系 + +本文档复用 dream 定义的底层模型,所有具体规则在 `auto_dream_design.md` 中: + +| 引用 | 来源 | +|---|---| +| wikilink 基础语法(`[[path.md\|alias]]` / predicate) | `auto_dream_design.md` §1.4 | +| 节点 / 边模型 | `auto_dream_design.md` §1.5 / §1.5.1 | +| F-invariants(F-1..F-11)| `auto_dream_design.md` §1.5.4 | +| 边守恒 E-1 / E-2 / E-3 | `auto_dream_design.md` §1.5.5 | +| 路径即 ID / rename | `auto_dream_design.md` §1.3 | +| CAS 写入协议 | `auto_maintain_design.md` §5 | +| anchor 不引入 | `auto_dream_design.md` §1.4 / §3.13 | +| SearchStep 召回 | `auto_dream_design.md` §3.10 | + +--- + +## 5. 下一步 + +1. **L1 实体识别 step 实现** —— 字符串匹配 + 语义召回 + LLM ambiguity 终判 + additive wikilink 写回(§1.2 / §1.3) +2. **inline post-write trigger 接入** —— G\* update / M split CAS 写入成功后 enqueue auto-link L1 job(§1.4) +3. **L2 / L3 周期 batch 框架** —— cron(daily / weekly)+ 候选挖掘 prompt + 写回路径(§1.3) +4. **L3 audit flag 接口** —— `maintainer_step` 提供 `predicate_upgrade` 操作,带 audit context 走特殊守恒规则(§2.1) +5. **中心性 retrieve 增益** —— file_graph inbound count → retrieve 加权乘子(§1.5) +6. **L4 报告框架** —— 重复语义检测产报告,提供 SDK / 人介入入口(§1.3 / §3) + +实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` 共同作为契约依据。 diff --git a/docs4/auto_maintain_design.md b/docs4/auto_maintain_design.md new file mode 100644 index 00000000..19b4fb52 --- /dev/null +++ b/docs4/auto_maintain_design.md @@ -0,0 +1,190 @@ +# auto-maintain 设计(digest 组织端:M split / 检测 / 写入并发) + +> 本文档记录 reme4 中 **auto-maintain** 的设计讨论 —— digest 层的组织 / 重组 / 写入运行时,含 M split、D 检测信号、写后触发模型、CAS 写入协议。 +> +> 配套阅读: +> - `structure.md` §3.6(maintain 动作语义)/ §7.3(maintainer 模块) +> - `auto_dream_design.md`:节点 + 边模型(§1.1-1.5)/ F-invariants(§1.5.4)/ 边守恒 E-1/E-2/E-3(§1.5.5)/ G\* 操作(§2.1)—— maintain 复用这套底层模型 +> - `auto_link_design.md`:auto-link 写回也走本文档的 CAS 协议(§5) +> - `auto_memory_design.md`:auto-memory 不直接复用 maintain,但事件级"拆"与节点级 split 在概念上同构(都把过载粒度切小) +> +> **三层框架的位置**:报告 §5 三层为 auto-memory / auto-dream / auto-link。maintain 严格按 `structure.md` §3.5-3.6 的 L4 action 分类是独立 action(`maintain: digest → digest`),不属 `digest` action(`digest: resource + daily → digest`)。本文档作为四方分工的**第四份**,专门覆盖 dream 写完之后 digest 的组织 / 重组 / 写入运行时。 +> +> **核心立场**: +> - **maintain 与 dream 同 pace**(idle background)、同模型(节点 + 边 / 守恒规则),但**语义边界不同**:dream 是 compose(资料 → digest),maintain 是 reorganize(digest → digest) +> - **maintain 只做 split**,不做 merge / dissolve / re-edge / unify;过载就拆,其它跨节点重组留给消费层 / 人工 +> - **CAS 写入协议是基础设施**,被 dream G\* / maintain split / auto-link L1 共用,统一编排在本文档(§5) + +--- + +## 0. 问题陈述 + +dream 模型(`auto_dream_design.md` §1.5)规定 digest 的演化只做两件事:G\* create_or_update(入流型,新材料融入)+ M split(后台,过载就拆)。dream 文档负责 G\* 与节点 / 边模型;**本文档负责 M split 与运行时机制**(D 检测 / 触发模型 / 写入并发协议)。 + +| 输入 | 输出 | +|---|---| +| dream 写入后的 digest 状态 + 触发信号(D3 过载,inline) | parent overview 重写 + N 个新 children 文件;边守恒 E-2 通过 | + +**设计目标**: +1. **形状匹配** —— 让节点粒度持续与实际语义结构对齐(过载节点拆;不过载不动) +2. **最小变更面** —— split 改 parent + 创建 N children,不动其它节点(F-2) +3. **不引入新基础设施** —— 复用 dream 的节点 + 边模型 / 守恒规则;CAS 写协议自洽 + +**显式排除**: +- ❌ merge / dissolve / re-edge / unify —— 跨节点重组不做(简化模型;同概念二次进入靠 G\* update) +- ❌ 改其它节点正文 —— split 只改 parent body(重写为 overview)+ 创建 children body +- ❌ 重建 inbound —— split 时 inbound 一律不动(F-10) + +--- + +## 1. M split(maintainer 唯一 op) + +| # | 能力 | 服务 | 触发 | graph | file | body | +|---|---|---|---|---|---|---| +| **M split** | 节点过载 → LLM 拆成 parent overview + N 个 children;parent 文件原地保留,children 是新文件;children 加 `[[parent]]` 反向链接;inbound 边不动 | 形状匹配(粒度对齐)+ 任意尺度(涌现层级) | D3 过载 | parent 0 拓扑改;新 children 节点 + 各自加 `[[parent]]` 出边 | 创建 N 个 children 文件;parent 文件原地 | parent body 重写为 overview;children 各自有新 body | + +**关键边界**: +- **M split 改两类 body**:parent body(重写为 overview)+ N 个新 children body;不改任何**其它**节点(`auto_dream_design.md` §1.5.4 F-2) +- **inbound 不重定向** —— 外部对 parent 的 wikilink 全部保留指 parent;后续 G\* 进入时若 LLM 觉得 child 粒度更合适,直接加新边到 child 即可(F-10) +- **没有 dissolve 操作** —— children 长期空也不主动删;消费层 / 人工显式介入 +- **没有 merge / re-edge / unify** —— 跨节点重组不做;同概念二次进入靠 G\* update;错桶节点不主动 move(若严重,人工介入) +- **边守恒** —— split 写新 parent body + N children body 前,机械对比 outbound:`(parent_new ∪ ∪children_outbound) ⊇ parent_old`;失败 → LLM 重试或拒写(F-11 / E-2,详 `auto_dream_design.md` §1.5.5) + +--- + +## 2. 检测信号 D1 / D3 / D10 + +| # | 信号 | 服务 | 服务能力 | +|---|---|---|---| +| **D1** | 断链(wikilink → 不存在的 path) | 任意尺度(可达性) | 告警 / 简单修复(就地删 wikilink 或保留 alias 文本) | +| **D3** | 过载节点(token 阈值 → LLM 判离散度) | 形状匹配(粒度) | maintainer(M split) | +| **D10** | provenance 断裂(digest 节点反指的 daily/resource 不可达) | 任意尺度(跨层不变量) | 严重告警(I 不变量违反) | + +**触发模型**:**写后立即** —— G\* / split 写完 body inline 检测;无后台 watcher / 无周期 tick / 无 dirty 队列(详 §4)。D1 / D10 是 wikilink 断链的子集,跟 file_graph 链路一起在写时检测。 + +> **简化模型砍掉的信号**: +> - **D2 隔离 / D4 过疏 / D5 高入度 / D5b 低入度摘要 / D6 slug 冲突 / D7 相似未链 / D8 重复语义 / D9 邻居异质** —— 全部 DROPPED +> - 旧 D5 高入度涌现 → 由 split 副产品(parent + children)等价覆盖;触发源换成节点过载(D3) +> - 旧 D6 slug 冲突 → 路径即 ID 后,同 bucket 内文件名冲突由文件系统层断言(写入即拒),不需要独立信号(详 `auto_dream_design.md` §1.3) +> - 旧 D7 / D8 → 简化模型不做 link / merge 提议;若 vault 累积明显的同概念重复,由 `auto_link_design.md` §1.3 L4 离线 audit 工具产报告 +> - 旧 D9 邻居异质 → 简化模型不做跨桶 move;桶选择只在 G4 一次性决定,后续不重排 +> +> **D3 过载的判据**:token 阈值机械检查 + LLM 判离散度;**写后立即 inline**。阈值见 §3,触发模型见 §4。 + +--- + +## 3. 检测阈值校准 + +简化模型只剩 D3(过载)是核心阈值,其它都是 invariant 触发(无可调阈值)或 informational(无 maintenance 联动)。 + +| 信号 | 阈值类型 | 默认 | 备注 | +|---|---|---|---| +| **D3 过载** | token + 主题离散度 | token 2000 / 离散度由 LLM 写后 inline 判 | **唯一驱动 split 的阈值**(详 §4) | +| **D1 断链** | 0 容忍 | 任意 1 条断链 → 告警 | 修复策略简单(就地删 wikilink) | +| **D10 provenance 断裂** | 0 容忍 | 任意 1 条断裂 → 严重告警 | I 不变量 | + +D3 阈值作为 `vault.yaml` 配置项(opinionated default,reme 核心提供机制不写死阈值),消费层可改;dogfooding 后调优。token 阈值起点 2000(对应"约 5 个独立子主题"的常见过载点),首版可调。 + +--- + +## 4. D3 触发模型(已收敛) + +**决策**:**写后立即检测,无 watcher 抽象,无 batch 窗口** —— 每次 G\* update / split 写 body 成功后,**inline** 在同一 job 内跑 D3:token 阈值 + LLM 离散度判定 → 必要时 enqueue split job(异步,走 §5 CAS 队列)。 + +``` +G* / split 写 body 成功(CAS 通过) + └─ if len(body) > T: + └─ LLM 判离散度 + └─ if is_overloaded: + └─ enqueue split job (FIFO, CAS-protected) + └─ return +``` + +**协议**: +- token 阈值默认 `2000`(§3 已定,`vault.yaml` 可配) +- 离散度 prompt 输出 `{is_overloaded: bool, suggested_clusters: [...]}`(若 overloaded 直接供 split job 吃,不重判) +- 启动无全扫(避免长启动);新写入立即检测覆盖增长路径;历史遗留过载随下次 update 自然检出 +- 无 dirty 标 / 无 dirty 集合 / 无后台 worker —— D3 是写路径的合成函数 + +**为什么 inline**:反应即时(不等下一次 ingest);实现最简(无批处理窗口 / dirty 状态 / 独立 worker);LLM 判定成本可接受(大多写入 < T 不触发,触发后 split 切小后续不再越界);不引入 watcher = 少一层部署/监控。 + +**已排除**:定时 cron tick(静止 vault 浪费扫描)/ ingest-after batch(引入 dirty 集合)/ 独立 L2 watcher worker(多余部署层)。 + +**演进路径(M1+)**:若 inline LLM 阻塞 G\* 时延成问题 → D3 改为 fire-and-forget enqueue;若同节点重复触发 LLM 成本高 → 加节点级 body hash 缓存。 + +--- + +## 5. CAS 写入协议(共享基础设施) + +**位置说明**:CAS 是 G\* update(`auto_dream_design.md` §2.1)、M split(本文档 §1)、auto-link L1 写回(`auto_link_design.md` §1.2)**三方共用**的写入协议。归在本文档是因为 maintain 是 digest 的"组织 / 运行时"端,运行时机制(检测 / 触发 / 写入)集中在一处方便对照。 + +**决策**:**并行决策 + 乐观冲突重做(CAS)** —— 所有 G\* / split / auto-link L1 决策并发跑,写入前用 body 版本戳(hash / mtime)做 CAS 比对;变了就丢弃 planned body 重做。无锁,无 ingest 级互斥。冲突率低 + E-1 / E-2 守恒校验顺手承担 race 兜底,无需新基础设施。 + +**协议(单个写入调用)**: +1. **读 + 记戳**:读 subject body → `version_stamp = sha256(body) | mtime` +2. **决策**:LLM 看候选池 → 决定 create / update / drop / split / additive-link;产 planned new_body +3. **CAS 写入**:重读 body 比 version_stamp + - **未变**:跑 E-1 / E-2 守恒校验 → 通过则 atomic write(write-temp + rename)→ done + - **已变**:丢弃 planned new_body,带最新 body 重走 step 1 +4. **守恒校验失败**:走 `auto_dream_design.md` §1.5.5 既有重试路径(LLM 重试一次,二次失败拒写 + audit) +5. **重做次数上限**:CAS-冲突重做最多 3 次;超出 → 跳过候选 + audit log(避免活锁) + +**create 路径 race**:两个 G\* 都决定 `create digest/auth/jwt-rotation.md` → atomic create(`O_CREAT | O_EXCL`)只让一个赢;输者拿 EEXIST → 重走 step 1(此时大概率改判 update)。 + +**适用范围**(全部走同一套 CAS):同 ingest 内 N 个候选并发 / 跨 ingest job 并发 / 后台 split 与前台 G\* 命中同节点(split 同样走 CAS)/ auto-link 写回(`auto_link_design.md` §1.2)。 + +**不解决的**:高冲突 workload(同概念被反复 ingest)→ 重做上限触发后 audit;跨进程并发(多 reme 实例同 vault)→ 不在 M0,需 fs lock(M1+)。 + +--- + +## 6. split 时的 provenance 处理(已收敛) + +**坍缩到 E-2 合计守恒** —— provenance 是 body 内联 wikilink(`auto_dream_design.md` §3.9),split 时跟其它 body 边完全同形:LLM 把 parent body 拆成 parent overview + N children,provenance wikilink 跟着对应内容段自然分配;机械层 outbound 合计守恒校验保证 `(parent_new ∪ ∪children_outbound) ⊇ parent_old`,旧 provenance 不可能丢。无需专门的 provenance 分配逻辑或"全部复制到 child / parent 保留全量"等特殊策略 —— LLM 按"哪个 child 谈到了哪段上游就带走哪条 provenance"自然处理。 + +--- + +## 7. G\* / split / auto-link L1 时序(已收敛) + +时序由 §4 / §5 与 `auto_dream_design.md` §3.10 共同规定,这里给最小汇总: + +- **G\* 调用本身同步** —— material 进来就走 G\* 决策(召回 + LLM 终判)+ CAS 写入(§5) +- **D3 检测 inline** —— G\* / split 写完 body 顺手跑 token 阈值 + LLM 判离散度(§4),无 tick / batch / watcher +- **split 异步** —— D3 触发后 enqueue split job 进 §5 CAS 队列,跟其它 ingest / split job FIFO 共享,异步消费;**不阻塞 G\* return** +- **auto-link L1 异步** —— 写入成功后 enqueue auto-link L1 job(`auto_link_design.md` §1.4),与 split job 同 CAS 队列、FIFO 共享;不阻塞 G\* return + +检测延迟 ≈ 0(inline);split 执行延迟 ≈ 队列等待时间(typically 数秒~数十秒);新建 / update 节点不必等 split 完成,体验连续。 + +--- + +## 8. maintainer 的人 / agent 后门(暂缓 — 非底层) + +消费层 / SDK 接口问题,不影响底层机制。底层只需保证 split / rename / delete 等 op 走同一套 §5 CAS + 守恒校验链路:F-3 仍成立(maintainer 自动路径只做 split);merge / dissolve / re-edge 在底层**不存在**(无对应机械算子)。后门接口形态推迟到 SDK 阶段再定。 + +--- + +## 9. 与 dream 模型的引用关系 + +本文档复用 dream 定义的底层模型,所有具体规则在 `auto_dream_design.md` 中: + +| 引用 | 来源 | +|---|---| +| wikilink 基础语法(`[[path.md\|alias]]` / predicate) | `auto_dream_design.md` §1.4 | +| 节点 / 边模型 | `auto_dream_design.md` §1.5 / §1.5.1 | +| F-invariants(F-1..F-11) | `auto_dream_design.md` §1.5.4 | +| 边守恒 E-1 / E-2 / E-3 | `auto_dream_design.md` §1.5.5 | +| 路径即 ID / rename | `auto_dream_design.md` §1.3 | +| anchor 不引入 | `auto_dream_design.md` §1.4 / §3.13 | +| provenance 载体形态 | `auto_dream_design.md` §3.9 | +| G\* 行为 | `auto_dream_design.md` §2.1 | + +--- + +## 10. 下一步 + +1. **M split step 实现** —— D3 触发 → 候选 → split prompt → 写入 + E-2 守恒(§1 / §4 / §5) +2. **D 检测信号实现清单**(D1 断链 / D3 写后 inline / D10 provenance 哪些已就绪 / 缺哪些)—— §2 +3. **D3 阈值配置**(`vault.yaml` 中 D3 token / 离散度阈值)—— §3 +4. **CAS 写入框架** —— per-path body version_stamp + CAS 写入 + EEXIST create race + 重做上限 + audit;对外暴露给 dream G\* / auto-link L1 复用 —— §5 +5. **后门 SDK 接口形态**(暂缓 M1+)—— §8 + +实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` / `auto_link_design.md` 共同作为契约依据。 diff --git a/docs4/auto_memory_design.md b/docs4/auto_memory_design.md new file mode 100644 index 00000000..ac594306 --- /dev/null +++ b/docs4/auto_memory_design.md @@ -0,0 +1,197 @@ +# auto-memory 设计(实时事件拆分 / 写入 daily) + +> 本文档记录 reme4 中 **auto-memory** 的设计讨论 —— 把 agent 连续的对话 / 任务流切成离散的 daily 事件原子,inline 落到 `daily/` 层。 +> +> 配套阅读: +> - `structure.md` §2.1-2.2(daily 层定位)/ §3.4(sync 动作语义)/ §7.1(synchronizer 模块) +> - `auto_dream_design.md`:auto-memory 产物如何被 dream 消化(G\* 读 daily 作为入流之一) +> - `auto_maintain_design.md`:digest 的组织端 / CAS 写入协议;auto-memory 不直接复用,但事件级"拆"与节点级 split 在概念上同构(都把过载粒度切小) +> - `auto_link_design.md`:auto-link 可反向扫 daily 事件,补充实体 wikilink(daily → digest) +> +> **四份分工**:reme 服务整体四份设计 —— **auto-memory(本文档)** / auto-dream / auto-maintain / auto-link。auto-memory 是入流端,把 agent 实时事件流切成 daily 事件原子;它的产物是 dream 消化的两路输入之一(另一路是 resource)。 +> +> **核心立场**:auto-memory 是 `structure.md` §3.4 `sync` 动作的实现侧 —— 强调 **inline 实时**与**事件边界检测**。是不是改名 sync → auto-memory 留给上层文档对齐,本文档聚焦机制。 + +--- + +## 0. 问题陈述 + +agent 的对话与任务过程是连续事件流(用户回合、工具调用、上下文切换、中断恢复),但记忆系统需要离散的、可独立检索的事件单元。auto-memory 解决这个切分问题。 + +| 输入 | 输出 | +|---|---| +| agent 当前事件流(对话回合 / 工具调用 / 任务切换信号);可选 `notify` 候选作为 cue | `daily///.md` 事件原子;`daily/.md` 主索引 | + +**设计目标**: +1. **事件边界尽量与 agent 语义意图一致** —— 同一个意图(同一个任务 / 同一段思路)→ 同一个事件;意图切换 → 新事件 +2. **inline 实时写入** —— 不滞后,不批处理;agent 一边工作,记忆一边落地 +3. **保持 daily 写权契约** —— I-2 单作者(同 folder 不并发改);folder 名 = summary note 名(I-3 可移动单元) + +**显式排除**(不属于 auto-memory 职责): +- ❌ 蒸馏 / 沉淀:那是 auto-dream(`auto_dream_design.md`)的事 +- ❌ 实体识别 / wikilink 自动补全:那是 auto-link(`auto_link_design.md`)的事 +- ❌ 改写 resource / digest:auto-memory 只写 daily(I-1 / I-3) + +--- + +## 1. 已对齐决策 + +### 1.1 物理布局:与 `structure.md` §2.2 对齐 + +| 项 | 决策 | +|---|---| +| **主轴** | 时间 + 任务:`daily///` | +| **event-slug** | LLM 抽取的事件短名(snake_case / dash-case;不强制 schema),同 `` 下唯一 | +| **folder 内** | 一个事件可有 N 个 note(`progress.md` / `decision.md` / `references.md` 等),由消费层 schema 决定;最少含一个 summary note,与 folder 同名 | +| **主索引** | `daily/.md`:当天事件列表(机械写入,wikilink 指向各 event folder)| +| **跨日索引** | 不强制;dream 消费时按 `` 范围拉取即可 | + +**为什么不是单文件 event**(report §5.1 一种简化方向): +- 单文件 event = `daily//.md` 比 folder 模型简单,但失去"一个事件可包含多个视角 note"的灵活度 +- 现行 `structure.md` 已定 folder 单位模型;auto-memory 沿用,不破坏既有 I-2 / I-3 +- 若后续 dogfooding 验证单事件普遍只有一份 note,可演进为 folder 内只放一份 summary,机械上等价于单文件方案 —— 演进路径平滑,不需要现在选 + +### 1.2 事件边界:语义意图切换驱动 + +事件边界由 LLM 在 inline 写入时判:**当前回合的意图是否仍属上一个 event**。 + +| 维度 | 决策 | +|---|---| +| **决策时机** | 每个 agent 回合写入前 inline 判 | +| **决策依据** | 上一个 active event 的 summary + 当前回合内容;LLM 输出 `{continue: bool, new_event_slug?: str, summary_patch?: str}` | +| **continue=true** | append 当前回合到 active event(append-only 或 LLM 重写 summary,详 §1.3) | +| **continue=false** | 关闭 active event(写最终 summary)+ 开新 event folder(slug 由 LLM 给)| +| **同时 active 多事件** | 不允许(I-2 单作者)—— 一时刻只一个 active event;真要并行任务,agent 自己 sync 切换 | + +**已排除**: +- 时间窗口切分(N 分钟无活动则切)—— 对话节奏因任务而异,时间窗口噪声大 +- 关键词切分(出现"切换 / 现在做 X"等触发词)—— 假阳性高,且不所有切换都明显说出 +- 后置 batch 切分 —— inline 写入要求 event 必须当下可决定归属,不能等 + +### 1.3 事件内写入模型 + +active event 内,每个回合的内容写到 event folder 下,有两种模式可选(消费层 schema 决定): + +| 模式 | 形态 | 适用 | +|---|---|---| +| **append-only** | 一份 `.md`,新回合 append 到末尾(章节 / 时间戳 / 等)| 实现最简;事件短(< 几十回合)时可读性 OK | +| **多 note 重写** | summary note(folder 同名)+ 各视角 note(`progress.md` / `decision.md`);LLM 把新内容融到对应 note,summary note 重写为当下概览 | 事件长 / 多视角时可读性高;LLM 成本高 | + +**默认 opinionated default**:append-only(最简启动)。消费层可改 prompt + schema 走多 note。 + +**与 E-1 守恒的关系**:daily 不强制 E-1 守恒(它是工作记录,允许 LLM 删旧加新);只在 multi-note 重写模式下,可选启用类似守恒(保留所有 wikilink),具体由消费层决定。 + +### 1.4 主索引 `daily/.md` + +当天事件 list 视图,机械维护(无需 LLM): + +| 触发 | 操作 | +|---|---| +| 新建 event folder | 主索引 append 一行 `[[daily///.md|]]` | +| 事件关闭(被切下一个 event) | 主索引该行 append 最终 summary 摘要(可选,LLM 写最终 summary 时附带写入) | +| 索引文件不存在 | 写入第一个 event 时创建 | + +主索引**仅承担当天浏览锚点**:文件系统 `ls daily//` 也能看见,但有主索引人/agent 可直接 `read daily/2026/05/28.md` 拿到 list 视图 + summary 一览。 + +不维护跨日索引(`daily/2026/05.md` 或 `daily.md`):dream 消费时按时间范围拉取即可;`list daily//` 已经覆盖浏览需求。 + +### 1.5 与 notify 的协作 + +`notify` 是 reme → agent 的虚边推送(`structure.md` §3.3),把"有新 resource 值得看"传递给 agent。auto-memory 在以下两点与 notify 协作: + +| 维度 | 协作方式 | +|---|---| +| **新事件 cue** | agent 收到 notify 后,如果决定响应(开始处理这个候选),通常会触发**新 event** —— auto-memory 把 notify payload 作为 hint(候选 resource 路径)写入新 event 的 summary,顺手用 wikilink 引上 | +| **acknowledge 派生** | event note 里出现指向 `[[resource/...]]` 的 wikilink → L1 watcher 将该 resource 推送状态置 `acknowledged`(`structure.md` §3.3 / §6.2);auto-memory 自身不调任何 ack API | + +**关键约束**:auto-memory **不强制** agent 用 wikilink 引 notify 候选 —— agent 可能略过、也可能不通过 wikilink 而是直接读 resource。ack 是 daily → resource wikilink 的副产品,不是 auto-memory 显式负责的事。 + +--- + +## 2. 待对齐边界点 + +### 2.1 LLM 决策频率与成本 + +inline 边界检测的最朴素形态是每回合调一次 LLM。在长对话 + 高频回合下成本可观。可选优化: +- **continue 假设默认**:大多数回合是 continue(同一意图内),LLM 可能只在"看似切换"启发(token 跨度大 / 工具种类突变 / 用户显式说"接下来")时跑;否则默认 continue 不调 LLM +- **批回合**:每 N 回合批一次,延迟切分(代价:active event 边界滞后,首版可接受) + +首版默认每回合调一次(最简,正确率高),M1+ 视成本优化。 + +### 2.2 中断恢复 / 跨进程 active event + +agent 进程重启 / Service 重启后,如何识别"还有 active event"? + +候选方案: +- **L2 自治状态**:L1 watcher 派生 `daily///` 中最新 mtime 的 event 为 active(默认 N 分钟内有写入) +- **状态文件**:`.daily-active` 维护 active event slug,Service 启动时读 +- **每次重建**:agent 进程重启视为新 event,旧的关闭(切到 §1.2 continue=false 路径)—— 最简但会增加事件数 + +倾向 §1.2 自然路径(进程重启 = LLM 下次判 continue=false 概率高)+ 不维护状态文件,详细 worker recovery 留给 Service 实现。 + +### 2.3 多 agent 同 vault 的 active event 隔离 + +I-2 daily 单作者契约在多 agent 场景下需细化。候选: +- per-agent date subfolder:`daily////` +- 单 agent 模式 + agent ID 进 event-slug:`/_/` + +第二种破坏 slug 短名习惯;第一种引入额外层级。倾向后者作为消费层契约,reme 核心不固化。 + +### 2.4 事件粒度的 prompt 引导 + +边界检测的 prompt 决定切分粒度。粗 = event 大 / dream 看每个 event 时容易 overflow;细 = event 数爆炸 / 主索引拥挤。 + +**opinionated default prompt 倾向**: +- 一个意图 = 一个事件(用户提了 X 问题 / agent 开了 Y 任务 → 直到这个意图收尾) +- 跨意图的"附带工作"(查资料 / 算个数)归入当前意图,不开新 event +- 真新意图("好,现在我们做下一件事")才切 + +详细 prompt 落 `reme4/steps/jobs/protocol.md` 或 synchronizer 的 prompt 模板。 + +### 2.5 与 resource ingest 的时序 + +如果 ingest 与 auto-memory 同时活跃(External push 推 resource 进来 + agent 在 sync),且 agent 想响应这个新 resource: +- ingest 写完 resource → L1 watcher 派生 L2 → notifier 决策推送(`structure.md` §5.3)→ Service MCP 推给 agent +- agent 在当前回合或下一回合响应 → auto-memory 判 continue=false 开新 event,wikilink 引上 resource + +整条链 sub-second 到 seconds(notify 节奏);auto-memory 不直接知道 ingest,只在 agent 决定响应时被动接收 notify payload。 + +### 2.6 跨日任务延续 + +event 物理路径含日期(`daily///`),同一意图跨日的任务无法用同一 event folder 承载。候选模型: + +| 模式 | 形态 | 适用 | +|---|---|---| +| **每日新 event,wikilink 反指前日** | new day 起新 folder;summary note frontmatter 加 `inherits: [[daily///.md]]`;新 event body 不复制旧内容,仅引用 | event-slug 短,日切口干净;查 backlinks 拼出整条任务链 | +| **同 event 重复写不同日** | 不允许(I-2 single author + event folder date 在路径上,跨日写违反路径不可变) | × | +| **任务 ID 跨 daily 抽象** | 引入 `task-id` 维度,daily event 只是某 task 的某一日切片;额外维护 task index | 复杂度高,M0 不引入 | + +**倾向**:第一种(`inherits` frontmatter wikilink)—— 与 §1.4 主索引一致(机械维护),实现侧 LLM 在 §1.2 boundary 判定时若发现意图与最近 N 天某个 active 任务一致,直接写入 inherits 即可。详细 boundary prompt 落 §2.4。 + +INHERIT 行为细节(扫描窗口、predecessor 是否关闭、Plan/Objective 是否拷贝)归消费层 schema 决定;reme 核心只承认 `inherits:` frontmatter wikilink 作为跨日链路载体。 + +--- + +## 3. 与其它层的协作 + +| 上下游 | 关系 | +|---|---| +| ← **notify** | 接收 notify payload 作为新 event cue;不强制响应,不强制 wikilink 引 | +| ← **resource** | 只读(通过 wikilink 引);不写 | +| → **daily** | **唯一写者**(I-2);写 event folder + 主索引 | +| → **auto-dream** | dream 的 G\* 读 daily 作为入流(`auto_dream_design.md` §2.1 G1 scope);auto-memory 写完即对 dream 可见(走 L2 索引,有 eventual 窗口) | +| → **auto-link** | auto-link 可反向扫 daily event,做实体识别 + wikilink 写回(`auto_link_design.md` §1.3)—— 与 auto-memory 写入不冲突(双方写不同字段段落 / CAS 协议保护)| + +**关键边界**:auto-memory 是 daily 写入端的**唯一**入口;dream / link 不写 daily 主路径,只通过 auto-link 走 §1.3 写回(read-only audit-then-write,CAS 保护)。 + +--- + +## 4. 下一步 + +1. **synchronizer step 实现**:event 边界检测 prompt + active event 状态管理 + inline 写入(append-only 默认) +2. **主索引维护**:`daily/.md` 机械维护(新 event 时 append、关闭时附 summary)—— 走 crud/daily 基础工具 +3. **notify ack 派生验证**:L1 watcher 派生 acknowledged 状态(`structure.md` §6.2),与 auto-memory 的 wikilink 写入端到端跑通 +4. **多 agent 隔离 schema**(M1+):若实际有并发 agent,确定 daily 子目录 / slug 命名约定 +5. **粗 / 细粒度 prompt 调参**:dogfooding 后看实际 event 数 / dream 消化效率,调 boundary prompt + +实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` / `auto_link_design.md` 共同作为契约依据。 diff --git a/docs4/structure.md b/docs4/structure.md new file mode 100644 index 00000000..56bb89dc --- /dev/null +++ b/docs4/structure.md @@ -0,0 +1,782 @@ +# reme4 系统架构 — 设计文档 + +## 文档定位 + +本文档定义 reme4 的**架构设计**:概念边界、数据流契约、职责划分。 + +- 不涉及代码路径 / 实现进度 / API 具体形态 +- **执行栈与架构角色**(分层、模块切分、职责划分)在第 6-7 节;源码映射与落地状态见 `docs4/reme4_report.md` 与源码 +- 不规定怎么做,只规定**是什么、谁负责、输入输出** + +> 阅读顺序:**第 1 节**给出完整的架构总览(数据视角 + 运行时视角 + 核心机制 + 不变量 + 导航);**第 2-5 节**逐层展开三层存储 / 6 类 L4 动作语义 / 反向回流 / 触发节奏;**第 6 节**描述底层执行栈(L0→L5);**第 7 节**把 6 类 Action 落到 L4 实现模块;**第 8 节**讲跨切面 schema 契约;**第 9 节**列出明确不属于本架构的反例。 + +--- + +## 1. 架构总览 + +本章给出 reme4 完整的设计骨架,后续 §2-§9 逐项展开细节。 + +### 1.1 解决什么问题 + +reme4 是 agent 的**长期记忆系统**。它把 agent 的工作过程沉淀为可检索、可演化的知识结构。 + +设计要解耦两件事: + +| 关注点 | 由谁负责 | +|---|---| +| **agent 写什么 / 读什么** | agent 的工作流自决 | +| **vault 自身如何健康演化** | reme 自治,agent 不感知 | + +实现方式:三层存储拓扑为"两路并行写入(原始资料 / 任务过程) → 双源合流到沉淀知识层",reme 提供这三层的容器、动作语义、反向检索与自治维护。 + +### 1.2 数据视角:并行起点 + 双源合流 + 反向回流 + +数据有**两个并行起点** —— **External source**(webhook/upload/pull)与 **Agent**(外部主体,自身任务驱动)。两条独立通道各自落地到**并行的材料层**:External 经 `ingest` 沉到 `resource/`,Agent 经 `sync` 写到 `daily/`。两层材料**合流**到 `digest/`,由 reme 通过 `digest` 动作完成"消化"。`digest/` 自身由 `maintain` 做 in-place 重组。Agent 通过 `retrieve` 从 resource + daily + digest **三层并行**回流。`notify` 是 reme 跨过 vault 直接提醒 Agent 的**虚边**(控制信号,不写任何文件)。每段路径都对应一个 L4 动作语义(完整动作详见 §5.2 / §7): + +| 路径 | L4 动作 | 说明 | +|---|---|---| +| External source → resource | **ingest** | 外部信源落到 vault | +| External source ╌╌► Agent | **notify** | reme 推送通知(虚边,只走 L2 推送队列,不写文件) | +| Agent → daily | **sync** | Agent 写 daily(响应 notify 或自身任务驱动) | +| resource + daily → digest | **digest** | reme 内部 LLM 抽取沉淀(双源合流) | +| digest → digest | **maintain** | reme 内部 LLM 折叠重组(in-place, fold-only) | +| resource + daily + digest → Agent | **retrieve** | 三层并行回流(state / semantic / topological 三种正交问法) | + +``` + notify(虚边,控制信号) +External source ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌► AGENT +(webhook/upload/pull) (外部主体/自身任务) + │ │ ▲ + │ ingest │ │ retrieve + │ ┌────── sync ────────────────────-┘ │ (state / + ▼ ▼ │ semantic / + ┌──────────┐ ┌──────────┐ │ topological) + │resource/ │ │ daily/ │ │ + │ 原始资料 │ │任务工作区 │───────── retrieve ─────────────┤ + │ 不可变 │ │ 半可变 │ │ + └───┬──┬───┘ └────┬─────┘ │ + │ │ │ │ + │ │ digest digest │ │ + │ └────────┐ ┌────┘ │ + │ ▼ ▼ │ + │ ┌────────────────────┐ │ + │ │ digest/ │ ◄──╮ │ + │ │ 沉淀知识 │ │ maintain │ + │ │ 可重组(语义索引) │ ───╯ (in-place, fold-only) │ + │ └──────────┬─────────┘ │ + │ │ retrieve │ + │ └─────────────────────────────────────-─┤ + │ │ + │ retrieve │ + └──────────────────────────────────────────────────────────┘ +``` + +**模型要点**: + +- **两个起点平行,不存在主从** —— External 与 Agent 各自独立驱动;Agent 既可响应 `notify` 也可由自身任务直接 `sync`。 +- **两层材料平行,不存在传递** —— resource 与 daily 是**两条独立的写入通道**,不互相穿越:Agent 不写 resource,ingester 不写 daily。 +- **digest 是双源合流的产物** —— `digest` 动作的输入是 resource + daily 的组合(不是仅 daily);相应地,digest 节点的 provenance 可同时指向 resource 与 daily。 +- **digest 自循环** —— `maintain` 在 digest 内部做密度折叠,不与上游材料层交互。 +- **notify 是虚边** —— reme 用它提醒 Agent "有新 resource 值得看",但不落任何文件;Agent 的响应通过 `sync` 落 daily(并可选地用 wikilink 引 resource)。 + +retrieve 三种问法正交: + +| 问法 | 工具 | 主要看哪层 | +|---|---|---| +| **state**(谁在 / 是什么状态) | `list` / `frontmatter` | 各层平等 | +| **semantic**(我想到一个意思) | `search` | digest > daily > resource(默认权重) | +| **topological**(从一个点向外摸) | `traverse` | 沿 wikilink 跨层平等 | + +### 1.3 运行时视角:六层执行栈 + 双进程 + +reme4 的功能不是堆在一层,而是从文件系统底层往上栈式堆叠。顶层 Service 与 Runtime 是同一套 vault 上的两个进程角色,共享 L0-L4 全栈(详见 §5.3 / §6)。 + +``` + ┌─────────────────────┐ ┌─────────────────────┐ + │ L5 Service │ │ L5 Runtime │ + │ (HTTP / MCP) │ │ (scheduler 自治) │ + └──────────┬───────────┘ └──────────┬──────────┘ + │ │ + └─────────────┬────────────────┘ + ▼ + ┌──────────────────────────────────────────────────┐ + │ L4 6 类 Action(动作语义) │ + │ ingest notify sync retrieve digest maintain│ + └──────────────────────┬───────────────────────────┘ + ▼ + ┌──────────────────────────────────────────────────┐ + │ L3 原子工具 │ + │ ┌─────────────────────┐ ┌──────────────────┐ │ + │ │ 基础工具 │ │ 高级工具 │ │ + │ │ create/append/edit/ │ │ search │ │ + │ │ read/write/move/ │ │ traverse │ │ + │ │ delete/list/stat │ │ frontmatter │ │ + │ └──────────┬──────────┘ └────────┬─────────┘ │ + └─────────────│──────────────────────│─────────────┘ + │ 直读 / 直写 │ 走索引读 + │ (eventual,有滞后) │ + │ ▼ + │ ┌────────────────────────────┐ + │ │ L2 文件状态 │ + │ │ · file_store(chunk+vec) │ + │ │ · file_graph(node+link) │ + │ │ · 自治状态(scheduler 用): │ + │ │ - resource: 入流批次/ │ + │ │ 未消化(orphan) │ + │ │ - daily: 任务索引 │ + │ │ (进行中/stale/完成) │ + │ │ - digest: 密度水位/ │ + │ │ 断链(broken wikilink) │ + │ │ · 推送队列(notify): │ + │ │ pending/notified/ │ + │ │ acknowledged │ + │ └─────────────▲──────────────┘ + │ │ 派生 / 更新 + │ ┌─────────────┴──────────────┐ + │ │ L1 file_watcher │ + │ │ fs event → state delta │ + │ │ (唯一 fs→state 桥) │ + │ └─────────────▲──────────────┘ + │ │ 监听 + ▼ │ + ┌──────────────────────────────────────────────────┐ + │ L0 vault 文件系统 │ + │ resource/ daily/ digest/ │ + └──────────────────────────────────────────────────┘ +``` + +Service 与 Runtime 是同一份 vault 上的两个进程角色: + +| 进程 | 触发源 | 时延敏感 | 典型动作 | +|---|---|---|---| +| **Service** | 外部 push / 外部 pull / agent 同步请求 / **MCP 推送通道** | 是 | ingest / sync / retrieve / **notify-out(MCP)** | +| **Runtime** | scheduler 周期 + L2 自治状态阈值 | 否(eventual) | **notify 决策** / digest / maintain | + +### 1.4 核心机制总览 + +| 机制 | 一句话 | 详见 | +|---|---|---| +| 三层存储 | resource(冷) / daily(温) / digest(冷,组织化) | §2 | +| 6 类 L4 动作 | ingest / notify / sync / retrieve / digest / maintain | §1.2 / §3 | +| notify+sync 链 | reme 主动从 L2 资源自治状态选候选,经 MCP 推给 agent;agent sync 落 daily | §3.3-3.4 / §7.1 | +| 反向 retrieval | state / semantic / topological 三种正交问法 | §4 | +| 触发四源 | 外部 push / 外部 pull / agent on-demand / reme 后台 | §5.1-§5.2 | +| Service + Runtime | 双进程角色,共享 L0-L4,职责按时延分 | §5.3 | +| 执行栈(L0-L5) | filesystem → file_watcher → 文件状态 → 原子工具 → Action → Service/Runtime | §6 | +| 原子工具:基础 vs 高级 | 基础直 fs;高级走 L2 索引(eventual) | §6.4 / §5.5 | +| Action 模块映射 | 6 类 Action 由 5 个模块实现(retrieve 直走原子工具) | §7 | +| Schema 跨切面 | name+description 是核心强约束,其余 opinionated default 可重载 | §8 | + +### 1.5 核心不变量速览 + +写入拓扑(架构脊梁,来自 §2.3):两路并行写入(External→resource、Agent→daily)→ 双源合流到 digest;resource 与 daily 之间互不写入;任何一层都不能反向改写它的上游。 + +| 不变量 | 内容 | 来源 | +|---|---|---| +| **I-1** | agent 不直接写 digest(digest 写权只属 digester / maintainer) | §2.4 | +| **I-2** | daily folder 单作者(同 folder 不并发改) | §2.4 | +| **I-3** | resource 内容不可变,只允许 metadata appendable | §2.4 | +| **I-4** | 三层共用同一套 wikilink 索引,跨层引用全靠 wikilink | §2.4 | +| **R-1** | retrieve 三种问法分立,不合并为单一 read verb | §4.3 | +| **M-1** | Maintainer 只做一件事:密度折叠(把碎片叶子折叠到新的中间节点下) | §7.3 | +| **F-1** | L1 `file_watcher` 是 L0→L2 的唯一派生桥 | §6.5 | +| **F-2** | L3 基础工具直接读写 L0;高级工具只走 L2 | §6.5 | +| **F-5** | L5 Service / Runtime 共享 L0-L4,不直接通信 | §6.5 | +| **F-6** | L0↔L2 存在 eventual 窗口,agent 上下文承担近期信息 | §5.5 / §6.5 | + +### 1.6 文档导航 + +| 想了解… | 看 | +|---|---| +| 三层各自的定位、不变量 | §2 | +| 6 类 L4 动作语义(notify / sync / digest / maintain 的输入产出不变量) | §3 | +| Retrieval 的三种问法与跨层语义 | §4 | +| 谁来触发、什么节奏、为什么分两个进程 | §5 | +| 系统从文件系统到 Service 的分层(底层基础) | §6 | +| L4 五个模块的对称结构与 Maintainer 折叠设计 | §7 | +| Schema 协议与重载机制 | §8 | +| 哪些设计不属于本架构(反例与边界) | §9 | +| 术语回查 | 附录 | + +--- + +## 2. 三层存储 + +### 2.1 一句话定位 + +| 层 | 一句话 | +|---|---| +| **resource/** | 外部原始资料的**不可变快照**。reme 是容器,不是作者。 | +| **daily/** | agent 的**任务工作区**。folder 是单位,以"日 + 任务"为索引。 | +| **digest/** | 跨任务沉淀的**有组织知识**。以语义(概念/实体/方法)为索引,与时间无关。 | + +### 2.2 五维度对照 + +| 维度 | resource/ | daily/ | digest/ | +|---|---|---|---| +| **组织主轴** | 时间(`/`) | 时间 + 任务(`//`) | 语义(`//...`,任意嵌套) | +| **写权归属** | 入流通道唯一(webhook / upload / pull) | agent(写入任务过程) | digester / maintainer(无 agent 直写) | +| **可变性** | 不可变,只追加新文件 | folder 内可反复更新 | 单节点可演化,可被合并/拆分/移动 | +| **不变量** | 写入即冻结,原文永不变 | folder 名 = summary note 名(可移动单元);同 slug 同日只一份 | slug 全局唯一;每 folder 有 canonical entry;wikilink 全路径 | +| **谁在用** | agent(查原文)、digester(双源输入之一) | agent(自己的工作记录)、digester(双源输入之一) | agent(召回主目标)、maintainer(自维护对象) | + +### 2.3 写入纪律:并行写入 + 双源合流 + +``` + External source Agent + │ │ + │ ingest │ sync + ▼ ▼ + ┌──────────┐ ┌──────────┐ + │resource/ │ │ daily/ │ + │ 不可变 │ │ 半可变 │ + │(ingester)│ │ (agent) │ + └─────┬────┘ └─────┬────┘ + │ │ + │ digest digest │ + └──────────────┐ ┌─────────────┘ + ▼ ▼ + ┌──────────────┐ ◄──╮ + │ digest/ │ │ maintain + │ 可重组 │ ────╯ (in-place, + │ (digester + │ fold-only) + │ maintainer) │ + └──────────────┘ +``` + +写权按这个**两层并行 → 单层合流**的拓扑分配:resource 写权专属 ingester(外部入流通道),daily 写权专属 agent(sync 落入,可响应 notify 或自身任务驱动),digest 写权专属 digester + maintainer。**resource 与 daily 之间互不写入**(agent 不动 resource,ingester 不动 daily);任何一层都不能反向改写它的上游。这是整个架构的脊梁。 + +### 2.4 不变量(永远成立) + +| # | 不变量 | 否则后果 | +|---|---|---| +| **I-1** | agent 不直接写 digest | digest 的"有组织"性失守,沉淀质量退化 | +| **I-2** | daily folder 单作者(同 folder 不并发改) | 任务边界模糊,sync/digest 竞态 | +| **I-3** | resource content immutable,只允许 metadata appendable | 原文可能消失/被覆写,citation 不可信 | +| **I-4** | 三层共用同一套 wikilink 索引,跨层引用全靠 wikilink | 引入第二套引用机制 → 索引重建复杂 / 跨层关系不可达 | + +--- + +## 3. L4 动作语义详解 + +§1.2 给出了 6 个 L4 动作在数据视角下的整体形态。本节按动作逐个展开输入 / 产出 / 不变量 / 反例。`ingest`(外部→resource,机械)和 `retrieve`(三层并行回流,只读)分别在 §5/§7.4 与 §4 详述,本节聚焦四个**写动作**:`notify` / `sync` / `digest` / `maintain`。 + +### 3.1 统一原则 + +四个写动作都遵守: + +| 原则 | 内容 | +|---|---| +| **Monotonic content** | 上游内容不可变,下游只能新建节点或加链接,不能改写上游 | +| **Provenance 必须可达** | 任何下游节点必须能通过 wikilink 反查到上游来源 | +| **Wikilink 是新结构的唯一载体** | 跨层关系靠 wikilink,不靠内容拷贝 | + +它们都不是"数据搬家",而是"在下游新生成有引用关系的节点"。 + +### 3.2 四个写动作的本质对照 + +| 动作 | 上游 → 下游 | 性质 | 上游变化 | 下游变化 | +|---|---|---|---|---| +| **notify** | resource → Agent | **Attention**(推送注意力) | 不变 | 不写 vault;仅入 L2 推送队列 | +| **sync** | Agent → daily | **Reference**(引用落地) | 不变 | daily 中新增工作记录 + 对 resource 的 wikilink | +| **digest** | resource + daily → digest | **Crystallize**(双源合流结晶) | 不变 | digest 新增节点,wikilink 反指上游来源(daily 与/或 resource) | +| **maintain** | digest → digest | **Reorganize**(重组) | 结构变,内容守恒 | fold-only:引入子中间节点搬叶子,改变拓扑 | + +> `notify` 与 `sync` 共同实现"resource 中的候选被 Agent 看见并织入 daily"这条**Reference 链**;它们是两个独立的 L4 动作,主体不同(notify 由 Reme 自治触发,sync 由 Agent 触发)。 + +### 3.3 notify:reme → Agent 推送候选 + +| 维度 | 内容 | +|---|---| +| 主体 | Reme Runtime(`notifier` 模块) | +| 输入 | L2 资源自治状态:orphan(无 inbound wikilink)/ 入流批次 / 未消化老于 N | +| 产出 | L2 推送队列条目;通过 Service MCP **server-initiated notification** 推到 Agent;**不写任何 vault 文件** | +| 不变量 | resource 原文 0 修改;**完全单向**,不维护任何反向元数据;`notify` 决策在 Runtime,Service 只作 MCP transport | +| Acknowledge 机制 | L1 watcher 检测到 daily→resource 新 wikilink → L2 推送状态 `notified` → `acknowledged`,避免重复推送 | +| 反例 | (a) Service 自决推什么 notify(✗-17);(b) notify 写入 vault(✗-16);(c) Agent 主动调 notify(✗-18) | + +### 3.4 sync:Agent → daily 落地 + +| 维度 | 内容 | +|---|---| +| 主体 | Agent(`synchronizer` 模块在 Service 内编排) | +| 输入 | Agent 当前事件流(响应 `notify` 的候选,**或**自身任务直接驱动) | +| 产出 | daily folder 内的工作叙事;可选地用全路径 wikilink 引 resource(agent 自决,reme 不强制) | +| 不变量 | resource 原文 0 修改;daily 单作者(I-2);folder 名 = summary note 名(可移动单元) | +| Provenance | daily → resource 可达(通过 daily body 中的 wikilink) | +| 反例 | "agent 把 resource 内容拷进 daily" —— 不允许,daily 只持有引用 + 自己的工作记录 | + +### 3.5 digest:resource + daily → digest 双源合流结晶 + +| 维度 | 内容 | +|---|---| +| 主体 | Reme Runtime(`digester` 模块,LLM-driven) | +| 输入 | 一组待蒸馏的 daily folder + 相关 resource(双源合流;通常以 daily 任务为线索,顺着 wikilink / 同主题搜索拉入相关 resource 原文) | +| 产出 | digest 中 0~N 个新节点 或 已有节点的更新;新节点必须用 wikilink 反指至少一个上游来源 | +| 不变量 | resource / daily 正文 0 修改;digest 新节点必须 wikilink 反指上游(provenance);digest 节点遵守第 2.4 节列的不变量 | +| Provenance | digest → daily / resource 双源链条可达(资料源是 resource 时直接反指,任务过程是 daily 时反指 daily 进而可达 resource) | +| 反例 | digester 改写 resource;digester 改写 daily 正文 | + +### 3.6 maintain:digest → digest 折叠 + +| 维度 | 内容 | +|---|---| +| 主体 | Reme Runtime(`maintainer` 模块,LLM-driven,**fold-only**) | +| 输入 | digest/ 当前整体状态 | +| 产出 | 同一 digest/ 树的**密度折叠**(fold):某中间节点下叶子过多时,引入子中间节点把相关叶子归簇 + 写"高密度摘要" | +| 不变量 | 树**只向下生长**(从不反向);叶子内容 0 修改,只被搬位置;新中间节点 = 一个高密度摘要文件;任何节点移动**原子重写所有入边**(retarget);slug 全局唯一在折叠后仍成立 | +| Provenance | digest → daily 的反指链接在折叠后仍有效(retarget 保证) | +| 反例 | merge / move / promote / demote 等改写既有拓扑的操作;改写既有叶子内容 | + +> 折叠操作的承诺与决策点见 §7.3。 + +### 3.7 链接重定向例外 + +`maintain` 的 retarget 会改写其它节点中指向被移动节点的 wikilink。从字面看,这违反了"上游内容不可变"。 + +实际上这是 wikilink 系统的**机械性副作用**,不算下游写上游: + +| 字面 | 实质 | +|---|---| +| daily 里的 `[[digest/old.md]]` 被改成 `[[digest/new.md]]` | 作者意图("我引用了 X 这个 digest 节点")没变,只是 X 的物理位置变了 | + +只要 retarget 保持 wikilink 的**目标语义不变**,就允许它作为机械维护副作用穿越层界。这是这条规则的唯一例外。 + +--- + +## 4. Retrieval 反向回流 + +Retrieval 是把 §3 几条正向写动作反着读:agent 站在结果端,沿 wikilink 反查源头。 + +### 4.1 三种问法 + +按 agent 意图分,有三类完全不同的读需求,**正交**,各自独立: + +| 问法 | 例子 | 本质 | +|---|---|---| +| **状态问** (State) | "我有哪些 in-progress 的任务?""哪些 resource 还没被引用?" | 在某层做 list + frontmatter 过滤 | +| **语义问** (Semantic) | "关于 auth 重构我知道什么?" | 跨层全文/向量检索 | +| **拓扑问** (Topological) | "auth 概念周围都连了什么?" | 从某节点沿 wikilink 走 | + +### 4.2 三层 × 三问法 矩阵 + +| 问法 | resource/ | daily/ | digest/ | +|---|---|---|---| +| **状态问** | "未处理 resource 清单" | "active / pending-digest 清单" | "孤儿节点 / canonical 缺失 清单"(给 maintain 用) | +| **语义问** | 兜底(原文,信噪比低) | 次优(最新,但未沉淀) | **首选**(沉淀过,信噪比高) | +| **拓扑问** | 通常是叶子(被指向) | daily → resource / digest | digest 内部连接最密 | + +### 4.3 设计原则 + +| # | 原则 | 含义 | +|---|---|---| +| **R-1** | 三种问法分立,不合并为单一 "read" verb | 不同问法的索引、过滤、排序逻辑完全不同 | +| **R-2** | 语义问的默认权重 `digest > daily > resource`,**可被显式覆盖** | 默认体现"沉淀质量",但 agent 可指定单层或调权 | +| **R-3** | 拓扑问与层无关 | traverse 沿 wikilink 走,天然跨三层(I-4) | +| **R-4** | Provenance expansion 默认 lazy,eager 是上层便利封装 | 原子 retrieval 不自动展开;agent 需要时再 traverse | +| **R-5** | Cold start 不是新的 retrieval mode | 只是状态问 + 语义问的组合,reme 不为它单设 verb | + +### 4.4 在主干图里的位置 + +Retrieval 不引入新存储,不引入新层。它是 **agent 与三层存储之间的读视图**,通过三种正交问法暴露,共享同一套 wikilink 索引(I-4)。 + +--- + +## 5. 节奏与触发 + +锁定"谁推动每个动作发生"。这一步定 reme 是纯被动 service 还是带后台 runtime。 + +### 5.1 触发源四分类 + +| 触发源 | 性质 | 例子 | +|---|---|---| +| **External push** | 外部事件主动推 | webhook / 用户 upload | +| **External pull** | reme 主动去外部拉 | scheduled fetcher(RSS / 邮件 / API 轮询) | +| **Agent on-demand** | agent 在请求里显式调用 | "sync 我的对话" / "搜 X" | +| **Reme background** | reme 自己的 watcher / scheduler | file_watcher / cron-like | + +### 5.2 六个动作的触发归属 + +| 动作 | 主触发 | 备用触发 | 备注 | +|---|---|---|---| +| **ingest** | External push / pull | — | 外部→resource,机械 | +| **notify** | **Reme background**(cron + L2 资源自治状态阈值) | — | Runtime 决策,Service MCP 推送 | +| **sync** | Agent on-demand | — | agent→daily;notify 的响应也走这里 | +| **digest** | **Reme background** | Agent 显式(后门) | resource + daily 双源合流到 digest | +| **maintain** | **Reme background**(cron + threshold) | Agent / 人工 显式(后门) | digest 内部折叠 | +| **retrieve** | Agent on-demand | — | 三层只读 | + +**关键定性**:`notify` / `digest` / `maintain` 三个 reme 自治动作的主控制权**在 reme,不在 agent**。Agent 只负责"响应 notify + 写 daily(响应或自身任务驱动)+ 主动读";不需要记得"该看哪些 resource""该蒸了""该整理了"。 + +### 5.3 Service + Runtime 双进程结构 + +把 5.2 的归属直接推出 reme 的基本架构。两进程共享 L0-L4 全栈,只在 L5(进程入口)分叉: + +``` +┌──────────────────────────────────────────────────────────┐ +│ Reme System │ +│ │ +│ ┌────────────────────┐ ┌────────────────────────┐ │ +│ │ L5 Service │ │ L5 Runtime │ │ +│ │ (HTTP / MCP) │ │ (scheduler 自治) │ │ +│ │ 服务 agent 请求: │ │ 服务 vault 健康: │ │ +│ │ · ingest │ │ · notify(决策) │ │ +│ │ · sync │ │ · digest │ │ +│ │ · retrieve │ │ · maintain │ │ +│ │ · notify(MCP 推送)│ │ │ │ +│ └─────────┬──────────┘ └───────────┬────────────┘ │ +│ │ │ │ +│ └──────────────┬───────────────┘ │ +│ ▼ │ +│ ┌──────────────────────────────────────────────────┐ │ +│ │ L4 Action / L3 原子工具 │ │ +│ │ Action 编排 → 基础工具 + 高级工具 │ │ +│ └──────────────────────┬───────────────────────────┘ │ +│ ▼ │ +│ ┌──────────────────────────────────────────────────┐ │ +│ │ L2 文件状态(file_store + file_graph + 自治) │ │ +│ └──────────────────────▲───────────────────────────┘ │ +│ │ 派生 │ +│ ┌──────────────────────┴───────────────────────────┐ │ +│ │ L1 file_watcher(fs → state 的唯一桥) │ │ +│ └──────────────────────▲───────────────────────────┘ │ +│ │ 监听 │ +│ ┌──────────────────────┴───────────────────────────┐ │ +│ │ L0 vault filesystem │ │ +│ │ resource/ daily/ digest/ │ │ +│ └──────────────────────────────────────────────────┘ │ +└──────────────────────────────────────────────────────────┘ +``` + +两进程职责正交、共享 L0-L4 基础设施(详见 §6): + +| 维度 | L5 Service | L5 Runtime | +|---|---|---| +| 触发方式 | 请求-响应 + MCP server-initiated 推送 | 周期 + L2 自治状态阈值 | +| 服务对象 | agent | vault 自身 | +| 暴露给 agent | 是 | 否(agent 不感知) | +| 主要动作 | ingest / sync / retrieve / **notify 推送通道**(MCP) | **notify 决策** / digest / maintain | +| 与对端的耦合 | 通过 L2 推送队列读 notifier 产出 | 通过 L2 推送队列写,**不直接调 Service** | + +### 5.4 节奏(latency tolerance) + +| 动作 | 节奏 | latency 容忍 | +|---|---|---| +| retrieve | request-driven | sub-second | +| ingest | event-driven | seconds | +| sync | agent on-demand | seconds | +| **notify** | reactive(L2 资源状态变化后) | seconds ~ minutes | +| digest | reactive(状态变化后) | minutes ~ hours(eventual consistency) | +| maintain | periodic | days(无紧迫) | + +实时性需求差三个数量级。这是 digest / maintain 必须放后台异步的根本原因 —— 不能阻塞 agent 的 retrieve / sync 请求。 + +### 5.5 排序约束与一致性模型 + +部分动作对**不能并发**,background runtime 内部要保证排序: + +| 约束 | 原因 | +|---|---| +| **sync(同一 daily)→ digest(同一 daily)** | digest 不能看到 sync 半成品 | +| **digest(同一 scope)→ maintain(同一 scope)** | maintain 重组的拓扑不应被 digest 中途插入 | + +ingest / retrieve 跟所有动作都可并发(纯入流 + 纯读)。 + +**一致性模型 = Eventual consistency on digest/maintain**。Agent 不能依赖"我刚写完 daily 就能查到对应 digest"。digest / maintain 都是后台异步,有可见的延迟窗口。 + +**watcher 滞后契约(L0 ↔ L2)**:基础工具直接写 L0 文件系统,L2 文件状态由 L1 file_watcher 派生,二者之间存在 eventual 窗口 —— 写完一份 daily 后,search / traverse 这类走 L2 索引的高级工具不一定立刻能看到。这是设计意图,不是 bug:agent 本身有上下文窗口,近期信息靠 agent 自带的对话上下文承接,不依赖 reme 索引立即可见。需要"写后立刻可读"的场景请用基础工具(read 直接读 fs)。 + +--- + +## 6. 执行栈:六层结构 + +L3 原子工具、L4 Action、L5 进程都不直接操作文件系统。它们坐在 L0-L2 的**底层基础**上 —— 这套基础是 reme 的"动力源",决定了为什么上层能解耦成 Service + Runtime 两进程,且两者既正交又共享状态。 + +### 6.1 概念分层(L0 → L5) + +``` +┌──────────────────────────────────────────────────────────┐ +│ L5 Service ‖ Runtime │ +│ 进程入口:Service 服务 agent;Runtime 自治维护 │ +├──────────────────────────────────────────────────────────┤ +│ L4 Action(6 类动作语义) │ +│ ingest / notify / sync / retrieve / digest / maintain │ +├──────────────────────────────────────────────────────────┤ +│ L3 原子工具 │ +│ 基础工具(直接 fs) + 高级工具(走 L2 索引) │ +├──────────────────────────────────────────────────────────┤ +│ L2 文件状态 │ +│ file_store + file_graph + 自治状态(scheduler 用) │ +├──────────────────────────────────────────────────────────┤ +│ L1 file_watcher │ +│ fs event → state delta(唯一 fs→state 桥) │ +├──────────────────────────────────────────────────────────┤ +│ L0 vault filesystem │ +│ resource/ + daily/ + digest/ │ +└──────────────────────────────────────────────────────────┘ + +数据流(主要关系): + · L3 基础工具 ──写──► L0 + · L3 基础工具 ──读──► L0(无需经 L2) + · L0 变化 ──► L1 监听到 ──派生──► L2 state delta + · L3 高级工具 ──读──► L2(走索引) + · L4 Action ──编排──► L3 工具组合 + · L5 进程 ──触发──► L4 Action +``` + +| 层 | 角色 | 关键约束 | +|---|---|---| +| L0 | vault 文件系统 | 唯一真相源;任何 L2 状态都可由 reindex 从 L0 重建 | +| L1 | `file_watcher` | **唯一**与 fs 事件直接耦合的组件;fs→state 的唯一派生桥 | +| L2 | 文件状态(`file_store` + `file_graph` + 自治状态) | 高级工具的读视图;由 L1 单向更新,L3+ 只读不写 | +| L3 | 原子工具(基础 / 高级) | 基础直读写 L0;高级只走 L2 | +| L4 | Action(6 类语义动作) | N:M 编排 L3 工具;不直接碰 L0 / L2 | +| L5 | Service / Runtime 进程 | 共享 L0-L4 全栈,**不直接通信**,只通过 L0 / L2 状态间接耦合 | + +### 6.2 L1 file_watcher:fs → state 的唯一桥 + +`file_watcher` 是 reme 唯一与 OS filesystem 事件直接耦合的组件。它把 fs 变化翻译为 L2 文件状态的 delta,承担**双重职责**: + +``` +filesystem events (create / modify / move / delete) + │ + ▼ + file_watcher ─┬─► 索引同步:写完文件,L2 file_store/file_graph 自动更新 + │ (L3 基础工具不需要显式调用"入索引") + │ + └─► 自治状态派生:维护 scheduler 用的可推导状态 + · resource: 入流批次 / 未消化(orphan) / + 推送状态(pending/notified/acknowledged) + · daily: 任务索引(进行中 / stale / 完成) + · digest: 密度水位 / 断链(broken wikilink) +``` + +**关键设计**:可派生的状态由 watcher 在外部索引中维护,**不写回 frontmatter**。L3 工具只管写内容文件,状态由 watcher 独立派生。这是 ✗-14 反例(L3 写 `status` 字段)成立的基础。 + +**Acknowledge 派生例**:`notify` 的"推送状态"由 watcher 维护 —— 当 watcher 检测到一条新 wikilink 从 daily 指向某 resource,即把该 resource 的推送状态从 `notified` 改为 `acknowledged`。notifier / Service 都不需要显式 ack。 + +### 6.3 L2 文件状态 + +| 组件 | 职责 | 由谁更新 | +|---|---|---| +| `file_store` | chunk 分块 + 向量持久化,提供 search / read API | L1 watcher 派生 | +| `file_graph` | wikilink 有向图,提供 upsert / traverse(双向)API | L1 watcher 派生 | +| **自治状态** | scheduler 自治决策的输入(入流批次 / 任务索引 / 密度水位 / 断链) | L1 watcher 派生 | +| **推送队列** | `notify` 的待推送 / 已推送 / 已确认条目 | notifier 写 pending;Service 推送后置 notified;L1 watcher 检测到 ack 后置 acknowledged | + +L2 只关心"vault 当前是什么样",**无业务语义** —— 不知道 daily / digest / 动作语义的存在。L3+ 只读 L2,不写(**例外**:notifier 写推送队列,这是 Runtime 与 Service 之间唯一的间接耦合通道,见 F-5)。 + +> **现状提示**:当前实现中 file_store / file_graph 之外的"自治状态"和"推送队列"尚不完整,这是 L1 watcher 与 notifier 待补齐的能力。完整化后 scheduler 才能从"周期扫描"切换为"事件驱动",`notify` 才能从隐式变为显式。 + +### 6.4 L3 原子工具:基础 vs 高级 + +两组工具的切分依据只有一条:**是否必须经过 L2 索引**。 + +| 组别 | 工具 | 数据通路 | 一致性 | +|---|---|---|---| +| **基础工具** | create / append / edit / read / write / move / delete / list / stat | 直接对接 L0 | 写后立即可读(同一工具) | +| **高级工具** | search / traverse / frontmatter | 必须走 L2 索引 | 受 watcher 滞后影响(eventual) | + +**写路径全部走基础工具**(L4 Action 编排基础工具完成写入)。高级工具是**只读**的索引查询入口。 + +**eventual 窗口**:基础工具写 L0 后,L2 索引由 L1 watcher 异步追平。在窗口内,高级工具看到的是滞后的视图。详见 §5.5"watcher 滞后契约"。 + +### 6.5 设计含义 + +| # | 不变量 | 推论 | +|---|---|---| +| **F-1** | L1 `file_watcher` 是 L0 → L2 的**唯一**派生桥 | 状态一致性是 L1 的事;L3 工具不要"自己更新索引" | +| **F-2** | L3 基础工具直接读写 L0;L3 高级工具只走 L2 | 写路径无需"先 reindex";读路径接受 eventual | +| **F-3** | L0 是唯一真相源 | 任何 L2 状态都可由 reindex 从 L0 重建,L2 是缓存而非数据库 | +| **F-4** | L4 Action 与 L3 工具是 N:M 编排关系 | Action 不直接碰 L0 / L2 | +| **F-5** | L5 Service 与 Runtime 共享 L0-L4 全栈,**不直接通信** | 只通过 L0 / L2 状态间接耦合;一边崩了不影响另一边的读 | +| **F-6** | L0 与 L2 之间存在 eventual 窗口 | agent 上下文承担近期信息,不依赖 L2 立即可见(详见 §5.5) | + +--- + +## 7. L4 Action 模块映射 + +第 5 节列了 6 类 Action 及其触发源,本节把这些 Action 落到 **L4 实现模块**(架构角色,不指代源码路径);触发机制(on-demand 路径与 background 路径)见 §5.3。 + +### 7.1 五个 L4 模块 + +| 模块 | 实现动作 | 触发 | LLM-driven | 单一职责 | +|---|---|---|---|---| +| **ingester** | ingest | External push / pull | × | 原样落 resource + 抽 frontmatter + 入索引 | +| **notifier** | notify | Reme background(cron + L2 资源自治状态阈值) | × | 从 L2 资源自治状态选候选 → 写 L2 推送队列;Service MCP 拿走 | +| **synchronizer** | sync | Agent on-demand | ✓ | 把当下事件织入 daily 工作叙事 | +| **digester** | digest | Reme background | ✓ | resource + daily 双源合流成 digest 长期条目 | +| **maintainer** | maintain | Reme background | ✓ | digest topic tree 的**密度折叠**(fold-only) | + +`retrieve` 不构成独立 L4 模块,理由见 §7.4。 + +**三个 reme 自治模块**:notifier(机械)、digester(LLM)、maintainer(LLM)。三者都由 scheduler 触发,都消费 L2 自治状态,但只有 notifier 是机械的 —— 候选选择不需要 LLM,LLM 决策在 agent 侧的 sync。 + +### 7.2 对称结构 + +``` + 跨表征层翻译 结构性纪律 + (LLM-driven) (机械) + + Inbound: ingester + Attention: notifier + Working: synchronizer + Sink: digester + Organization: maintainer (fold-only) +``` + +五类不同方向的"翻译": + +| 模块 | 翻译方向 | +|---|---| +| ingester | 外部异构格式 → vault 统一文件 | +| notifier | L2 资源自治状态 → agent 注意力(`notify` 推送) | +| synchronizer | agent 事件流 → 工作过程叙事(写 hot) | +| digester | 工作过程 + 原始资料 → 长期知识(双源合流,写 cold) | +| maintainer | 散乱叶子 → 有层次的 topic tree(组织 cold) | + +ingester 和 notifier 是机械(确定性阈值/流水线);其它三个是 LLM 决策模块,各自跨越一层语义鸿沟。 + +### 7.3 Maintainer:Topic Tree 密度折叠 + +`digest/` 整体视为一颗 **topic tree**:文件夹 = 中间节点,文件 = 叶子。Maintainer 唯一职责:随写入持续,某中间节点下叶子过密时,**折叠**为新的子中间节点 + 高密度摘要。 + +``` +触发前:某中间节点叶子过多 / 太碎 + digest/infra/ + ├── logging.md + ├── tracing.md + ├── metrics.md + ├── alerting.md + ├── dashboards.md + └── slo.md + +折叠后:LLM 判断聚类,引入子中间节点 + 摘要 + digest/infra/ + ├── observability/ ← 新中间节点 + │ ├── _index.md ← 新生成的高密度摘要 + │ ├── logging.md ← 内容不变,只搬位置 + │ ├── tracing.md + │ ├── metrics.md + │ ├── alerting.md + │ ├── dashboards.md + │ └── slo.md + └── ...(未被折叠的叶子原位) +``` + +**设计承诺**(在 `maintain` 通用不变量之上进一步收紧): + +| # | 承诺 | 含义 | +|---|---|---| +| **M-1** | **Fold-only**,无 merge / move / promote / demote / introduce | 树只向下生长,从不反向 | +| **M-2** | 叶子内容 0 修改,只被搬位置 | 与 `maintain` 内容守恒一致 | +| **M-3** | 新中间节点带一个高密度摘要文件,读摘要就能决定要不要深入 | 折叠后可读性不降反升 | +| **M-4** | 每次只处理一个候选节点 | 最小化变更面 | +| **M-5** | 不能聚类时,**不动**(默认保守) | 宁可不折,不要错折 | + +LLM 唯一的决策点: + +1. 这些叶子能不能聚类(if not → 不动) +2. 新中间节点叫什么、摘要怎么写 + +其它都机械:阈值判断(L1 file_watcher 派生 L2 自治状态提供信号)、移动文件(crud)、wikilink 重定向(graph/retarget)。 + +### 7.4 为什么没有 retriever 模块 + +L4 模块的存在条件 = "有跨原子编排 / 需要 LLM 决策"。Retrieve 不满足: + +- 三种问法(state / semantic / topological)各自被 **L3 原子工具**直接覆盖(list+filter / search / traverse) +- 没有跨原子状态、没有 LLM 决策点 +- Agent 直接调用 L3 原子即可 + +--- + +## 8. 跨切面:Schema + +Schema(资料的 frontmatter / wikilink / 章节约定)是横跨三层、各写动作的共同契约。reme4 的核心立场: + +| 立场 | 说明 | +|---|---| +| **reme 核心只保留 `name` / `description` 两个字段** | 其它都是 opinionated convention,服务消费层可以替换 | +| **Schema 是"协议"不是"代码"** | 用 markdown 文字描述,LLM agent 自我约束;不内嵌 schema validator | +| **三层共用同一套 wikilink 协议** | 全路径引用,无 short-link / no-ext 解析 | + +### 8.1 协议文档(opinionated default) + +| 内容 | 谁规定 | +|---|---| +| 目录结构(三层 + folder 单位) | 第 2 节本文档 | +| 动作语义契约(notify / sync / digest / maintain 的输入产出不变量) | 第 3 节本文档 | +| Frontmatter 推荐字段(4 轴等) | `reme4/steps/jobs/protocol.md` (opinionated) | +| 章节约定(Objective/Plan/Progress/...)| sync / digest 各自的 prompt(opinionated) | + +### 8.2 重载入口 + +服务消费层(plugin / 自定义 caller)无需 fork reme,可通过以下方式替换 schema: + +| 入口 | 适用场景 | +|---|---| +| 替换 protocol 文档 | 改 frontmatter / wikilink / 章节约定 | +| 替换 prompt 模板 | 改 sync / digest 的决策流程 | +| 替换 toolkit | 改 ReAct agent 可见的工具集 | + +--- + +## 9. 反例:不属于本架构的设计 + +明确画出**不允许**的设计,免得后续讨论或扩展时滑回去: + +| # | 反例 | 违反的不变量 | +|---|---|---| +| ✗-1 | Agent 通过任意 verb 直接写 digest | I-1(digest 写权只属 digester / maintainer) | +| ✗-2 | 多 agent 并发改同一个 daily folder | I-2(daily 单作者) | +| ✗-3 | 任何动作改写 resource 的原文 | I-3(resource immutable) | +| ✗-4 | 跨层引用引入第二套机制(hash-id / external ref / SQL) | I-4(wikilink 是唯一跨层载体) | +| ✗-5 | digester 改写 daily 正文 | `digest` 不变量(§3.5) | +| ✗-6 | maintain 改写 daily / resource 的语义内容 | `maintain` 不变量(§3.6) | +| ✗-7 | `notify` 维护 resource 上的 `referenced_by` 反指 | `notify` 完全单向(§3.3) | +| ✗-8 | 把 state / semantic / topological 合并成单一 read verb | R-1 | +| ✗-9 | Retrieve 自动 eager-expand provenance | R-4 | +| ✗-10 | digest / maintain 同步阻塞 agent 请求 | 5.4 节奏分级 | +| ✗-11 | digest / maintain 强一致(agent 写完 daily 立即可查 digest) | 5.5 eventual consistency | +| ✗-12 | maintainer 做 merge / move / promote / demote 等"通用重组" | M-1(fold-only) | +| ✗-13 | maintainer 改写既有叶子的内容(不只是搬位置) | M-2(叶子内容 0 修改) | +| ✗-14 | L4 模块在 frontmatter 里写 `status` / `pending` 等可派生状态字段 | 状态由 L1 file_watcher 派生到 L2 自治状态,L4 不重复 | +| ✗-15 | 为 retrieve 单设 L4 模块或聚合 verb | §7.4(L3 原子已足够) | +| ✗-16 | notify 写入 vault(在 resource 上加 `notified` frontmatter 或新建 daily 占位) | notify 只写 L2 推送队列,**不落任何文件**;ack 由 L1 watcher 检测 wikilink 派生 | +| ✗-17 | Service 自决推什么 notify 候选 | notify 决策在 Runtime(notifier);Service 只是 MCP transport,从 L2 推送队列读取(F-5) | +| ✗-18 | agent 主动调用 `notify` 想"标记这个 resource 我要看" | notify 是 reme→agent 单向,反向是 agent 用 sync 写 wikilink(自然 ack) | + +--- + +## 附录:术语索引 + +| 术语 | 定义 | +|---|---| +| **resource/** | 不可变原始资料层 | +| **daily/** | agent 任务工作区层 | +| **digest/** | 沉淀知识层 | +| **State 问** | 在某层做 list + 过滤的状态查询 | +| **Semantic 问** | 跨层全文/向量检索 | +| **Topological 问** | 沿 wikilink 走的拓扑查询 | +| **Provenance** | 下游节点反查到上游来源的能力 | +| **Retarget** | 节点移动时对所有入向 wikilink 的原子重写 | +| **L5 Service** | 服务 agent 请求的进程(HTTP / MCP);执行栈最上层;也是 notify 的 MCP transport | +| **L5 Runtime** | 自治维护 vault 的进程;scheduler 在其中按 L2 自治状态阈值触发 background Action | +| **L4 Action** | 6 类动作语义:ingest / notify / sync / retrieve / digest / maintain | +| **L4 模块** | 实现 Action 的架构角色;五个:ingester / notifier / synchronizer / digester / maintainer(retrieve 不构成独立模块) | +| **ingester** | L4 模块,机械:外部源原样落 resource + 抽 frontmatter + 入索引 | +| **notifier** | L4 模块,机械:从 L2 资源自治状态选 notify 候选 → 写 L2 推送队列;Service MCP 拿走推给 agent | +| **synchronizer** | L4 模块,LLM-driven:agent 事件织入 daily 工作叙事;响应 notify 的也走这里 | +| **digester** | L4 模块,LLM-driven:resource + daily 双源合流为 digest 长期条目 | +| **maintainer** | L4 模块,LLM-driven,**fold-only**:digest topic tree 的密度折叠 | +| **scheduler** | L5 Runtime 内部触发器:按 cron + L2 自治状态阈值拉起 background L4 模块(notifier / digester / maintainer) | +| **Topic tree** | digest/ 的心智模型:文件夹 = 中间节点,文件 = 叶子 | +| **Fold(密度折叠)** | maintainer 唯一操作:把过密叶子归簇到新子中间节点 + 写高密度摘要 | +| **L3 原子工具** | 基础(create/append/edit/read/write/move/delete/list/stat,直 fs)+ 高级(search/traverse/frontmatter,走 L2)两组 | +| **L2 文件状态** | `file_store` + `file_graph` + 自治状态(resource 入流批次/orphan、daily 任务索引、digest 密度水位/断链)+ 推送队列;由 L1 派生(推送队列由 notifier 写) | +| **L2 推送队列** | `notify` 的 L2 状态条目;notifier 写 pending,Service MCP 推送后置 notified,L1 watcher 检测到 daily→resource wikilink 后置 acknowledged | +| **L1 file_watcher** | fs event → L2 state delta 的唯一派生桥;承担索引同步 + 自治状态派生 + `notify` ack 派生 | +| **L0 vault filesystem** | 物理目录:resource/ + daily/ + digest/;唯一真相源 | +| **Eventual consistency** | digest / maintain 异步处理,有可见延迟窗口;L0↔L2 之间 watcher 滞后窗口同理 | +| **Opinionated default** | reme 提供的参考实现,服务层可替换 | diff --git a/reme4/steps/jobs/__init__.py b/reme4/steps/jobs/__init__.py new file mode 100644 index 00000000..952c3cd4 --- /dev/null +++ b/reme4/steps/jobs/__init__.py @@ -0,0 +1,10 @@ +"""Jobs steps — composite ReAct-agent-driven workflows. + +Two steps: + + digester — cold-write: distill daily notes into digest/ (R-M-W via a ReAct agent). + synchronizer — hot-write: persist in-progress task as a daily note. +""" + +from . import digester # noqa: F401 -- @R.register("digester") +from . import synchronizer # noqa: F401 -- @R.register("synchronizer") diff --git a/reme4/steps/jobs/digester.py b/reme4/steps/jobs/digester.py new file mode 100644 index 00000000..641ae39e --- /dev/null +++ b/reme4/steps/jobs/digester.py @@ -0,0 +1,300 @@ +"""Smart Digester — knowledge distillation from daily notes to digest/. + +The Digester is the **cold-write** counterpart to Synchronizer (hot-write). +It reads completed work in ``daily//.md`` note files, +identifies entities / concepts / claims / methods worth preserving +long-term, and sinks them into ``digest/`` as canonical-entry nodes so +the main agent can retrieve them later via search and graph traversal. + +``digest/`` is the cold-tier root. Per ``protocol.md``, scope folders +under it may nest arbitrarily; each folder's canonical entry is +``/.md``, and slug (folder name) is globally unique across +the whole tree. Pending detection and graph machinery treat nodes at any +depth uniformly. + +Drives a ReAct agent with a read/lookup/graph/write toolkit; the +agent follows the protocol in ``protocol.md`` (the opinionated +default schema + R-M-W decision tree). The schema is convention-driven +— reme core only reserves ``name`` / ``description``, so the agent +owns its own discipline rather than relying on a post-write linter. + +Distillation state lives in the daily note's ``status`` frontmatter +— a **daily-tier convention owned by this digester** (reme core +reserves only ``name`` / ``description``; ``status`` is just an +extra). After processing each daily, the agent must call +``frontmatter_update`` with ``metadata={"status": "completed"}`` +(or ``metadata={"status": "skipped"}`` when intentionally bypassed). Convention: absent +≡ ``pending``, so the next pass finds residual work via +``file_list path=daily recursive=true`` + per-item ``frontmatter_read`` to filter for absent ``status``. +Only the digester writes ``status``; Synchronizer / hand-edits must +leave it alone. + +No degraded path — distillation strictly requires an LLM. When ``as_llm`` +is unavailable, the step short-circuits with ``skipped=True`` and an +error message. + +Override interface — schema is a service-consumption concern, not a +core invariant, so this step ships an **opinionated default** that any +caller can fully replace without touching reme4: + +* ``protocol`` / ``protocol_path`` constructor args replace the + ``protocol.md`` injected as ``{protocol}`` in the system prompt + (use when keeping the default prompt template but swapping schema). +* ``prompt_dict`` (inherited from ``BaseStep``) replaces the + ``system_prompt`` / ``user_message`` templates wholesale (use when + the prompt structure itself needs to change). +* ``toolkit`` replaces the tool surface ``_DIGESTER_TOOLS`` builds. + +Service layers (e.g. plugin-side configs) wire these in via component +config; ``digester.py`` / ``protocol.md`` shipped here are just a +reference implementation of one viable convention. + +Toolkit. Each entry in ``_DIGESTER_TOOLS`` is a job name registered +in the active config; ``add_as_tool`` wraps ``job(**kwargs)`` into a +``ToolResponse``. The job indirection means the agent sees the same +tool surface (rich descriptions + JSON schema) as the L2 MCP layer. +""" + +import datetime +import zoneinfo +from pathlib import Path + +from agentscope.agent import ReActAgent +from agentscope.message import Msg +from agentscope.tool import Toolkit +from pydantic import BaseModel, Field + +from ..base_step import BaseStep + +from ...components import R + + +_DIGESTER_TOOLS: tuple[str, ...] = ( + "file_list", + "file_read", + "file_stat", + "frontmatter_read", + "traverse", + "file_write", + "file_append", + "file_move", + "frontmatter_update", + "frontmatter_delete", +) + + +def _pack_daily(file_store, daily_path: str) -> str: + """Render one daily note file's body into a prompt-friendly block. + + A daily note is a single self-contained markdown file at + ``daily//.md`` — everything the originating task wanted + the digester to see is inline (no sibling materials). External + assets land in ``resource//`` and are linked from the note's + ``## References`` section; the LLM opens those on demand via + ``file_read``. + """ + try: + absolute = (Path(file_store.vault_path or ".") / daily_path).resolve() + except Exception as e: + return f"### {daily_path}\n(error resolving path: {type(e).__name__}: {e})\n" + + if not absolute.is_file(): + return f"### {daily_path}\n(note file not found)\n" + + parts: list[str] = [f"### {daily_path}"] + try: + parts.append(absolute.read_text(encoding="utf-8")) + except Exception as e: + parts.append(f"(error reading note: {type(e).__name__}: {e})") + return "\n".join(parts) + "\n" + + +class DistillResult(BaseModel): + """Outcome of a single distillation call. + + Without per-tool audit (the agent's toolkit is the job surface, + which doesn't expose per-call records back to the orchestrator), + the structured outcome is just the inputs the call was asked to + process plus the LLM's free-form summary. Per-file write outcomes + can be verified by re-reading vault_dir afterwards if needed. + + Field semantics: + * ``daily_read`` — daily paths actually processed (input + order, deduped) + * ``summary`` — LLM's free-form one-paragraph summary + * ``skipped`` — True when no LLM available, no daily + paths provided, or the LLM reported ``SKIP`` + * ``error`` — short error string when skipped due + to misconfiguration (e.g. no LLM) + """ + + used_llm: bool = False + skipped: bool = False + daily_read: list[str] = Field(default_factory=list) + summary: str = "" + error: str = "" + + +@R.register("digester") +class Digester(BaseStep): + """Knowledge digester: daily/ → digest/ via a ReAct agent. + + Inputs (from RuntimeContext): + daily_paths (list[str], required): vault-relative paths to + daily note files (``daily//.md``) to distill. + Pass ``[]`` to no-op. + hint (str, optional): caller guidance to the LLM + (e.g. "focus on the auth-related decisions"). + + Output (written to context.response.answer): + DistillResult JSON — see model docstring. + """ + + def __init__( + self, + toolkit: Toolkit | None = None, + console_enabled: bool = False, + timezone: str | None = None, + protocol: str | None = None, + protocol_path: str | None = None, + **kwargs, + ): + """Constructor overrides (service layer customization points): + + * ``toolkit`` — replace the agent's tool surface; default builds + one from ``_DIGESTER_TOOLS``. + * ``protocol`` — inline protocol document (highest precedence); + overrides whatever the agent sees under ``{protocol}`` in the + system prompt. + * ``protocol_path`` — path (relative to the vault or absolute) to a protocol + markdown file; used when ``protocol`` is not given. + * ``prompt_dict`` (inherited via ``BaseStep``) — override the + ``system_prompt`` / ``user_message`` templates wholesale, e.g. + to swap in a service-layer prompt that hardcodes a different + schema entirely. + + With none of the above, falls back to the opinionated default + (the ``protocol.md`` and ``digester.yaml`` shipped alongside + this module). + """ + super().__init__(**kwargs) + self.toolkit = toolkit + self.console_enabled = console_enabled + self.timezone = timezone + self._protocol = self._load_protocol(protocol, protocol_path) + + @staticmethod + def _load_protocol(protocol: str | None, protocol_path: str | None) -> str: + """Resolve the protocol document; explicit string > path > default.""" + if protocol is not None: + return protocol + if protocol_path: + path = Path(protocol_path) + if path.exists(): + return path.read_text(encoding="utf-8") + default_path = Path(__file__).parent / "protocol.md" + return default_path.read_text(encoding="utf-8") if default_path.exists() else "" + + def _now(self) -> datetime.datetime: + if self.timezone: + try: + return datetime.datetime.now(zoneinfo.ZoneInfo(self.timezone)) + except Exception as e: + self.logger.error(f"Invalid timezone: {self.timezone}, error={e}") + return datetime.datetime.now() + + def _vault_dir(self) -> Path: + vr = getattr(self.file_store, "vault_path", None) + return Path(vr).resolve() if vr else Path.cwd().resolve() + + def _llm_available(self) -> bool: + """Pre-flight check: can ``self.as_llm`` resolve without raising? + + ``BaseStep.as_llm`` asserts when no model is registered, so we + wrap the access here to avoid hard-failing at the call site.""" + try: + return self.as_llm is not None + except Exception: + return False + + def _build_toolkit(self) -> Toolkit: + """Bind every digester-relevant job as a tool function.""" + toolkit = self.toolkit or Toolkit() + for job_name in _DIGESTER_TOOLS: + self.add_as_tool(toolkit, job_name) + return toolkit + + async def execute(self): + assert self.context is not None + daily_paths: list[str] = list(self.context.get("daily_paths") or []) + hint: str = (self.context.get("hint", "") or "").strip() + + # No work to do: no daily paths supplied. + if not daily_paths: + result = DistillResult(used_llm=False, skipped=True) + self.context.response.success = True + self.context.response.answer = "Skipped: no daily paths supplied" + self.context.response.metadata.update(result.model_dump()) + return + + # No LLM available: distillation strictly requires one. + if not self._llm_available(): + result = DistillResult( + used_llm=False, + skipped=True, + error="no as_llm configured; distillation requires an LLM", + ) + self.context.response.success = False + self.context.response.answer = f"Error: {result.error}" + self.context.response.metadata.update(result.model_dump()) + return + + # Dedupe daily_paths while preserving order. + seen: set[str] = set() + deduped: list[str] = [] + for p in daily_paths: + if p and p not in seen: + seen.add(p) + deduped.append(p) + daily_paths = deduped + + # Build the per-daily blob the agent will see. + daily_blob = "\n\n".join(_pack_daily(self.file_store, p) for p in daily_paths) + + vault_dir = self._vault_dir() + toolkit = self._build_toolkit() + + agent = ReActAgent( + name="reme_digester", + model=self.as_llm, + sys_prompt=self.prompt_format( + "system_prompt", + vault_dir=str(vault_dir), + protocol=self._protocol, + ), + formatter=self.as_llm_formatter, + toolkit=toolkit, + ) + agent.set_console_output_enabled(self.console_enabled) + + user_message: str = self.prompt_format( + "user_message", + today=self._now().strftime("%Y-%m-%d"), + hint=hint or "(none)", + daily_blob=daily_blob or "(none)", + ) + + final_msg: Msg = await agent.reply( + Msg(name="reme", role="user", content=user_message), + ) + summary = (final_msg.get_text_content() or "").strip() + + result = DistillResult( + used_llm=True, + daily_read=list(daily_paths), + summary=summary, + skipped=summary.upper().startswith("SKIP"), + ) + self.context.response.success = True + self.context.response.answer = summary or "Distillation completed" + self.context.response.metadata.update(result.model_dump()) diff --git a/reme4/steps/jobs/digester.yaml b/reme4/steps/jobs/digester.yaml new file mode 100644 index 00000000..cf311ec6 --- /dev/null +++ b/reme4/steps/jobs/digester.yaml @@ -0,0 +1,114 @@ +system_prompt: | + You are the digester. You read completed work in `daily//.md` + note files and lift entities, concepts, claims, and methods into + canonical entries under `digest/<…>//.md`. The main agent + retrieves them later via search and graph traversal. + + vault_dir: {vault_dir} + + ## Five steps per call + + ### Step 1 — Read each daily + + For every daily note passed in, the full body is already packed + in the user message below. Each note is a single self-contained + markdown file — anything the originating task wanted you to see is + inline. If a note's `## References` section points at + `[[resource//]]` items and you need them, open them via + `file_read` on demand. + + ### Step 2 — Lookup candidates + + Identify each entity / concept / claim / method named in the daily. + Before deciding to CREATE anything, look it up in `digest/`: + + - `file_list path=digest/ recursive=true` to scan the tree, or + - `graph_traverse path=.md depth=1` for neighborhood. + + **Slugs are globally unique under `digest/`** — a prior occurrence + at any nesting depth means the node already exists. Find and reuse + it; do not CREATE a duplicate under a new path. This is the worst + failure mode of the digester. + + ### Step 3 — R-M-W decision + + Per candidate, pick exactly one branch. Every branch writes only + the node being authored in this step — never sideways into other + nodes' bodies. A relation worth recording is captured by a typed + wikilink in the source body; the inbound view is queried later via + `graph_traverse direction=in`. + + - **CREATE** — no hit anywhere under `digest/`: + `file_write digest/<…>//.md`. Place it under the + closest existing semantic parent scope; top-level if none applies. + + - **UPDATE** — exact match exists: + Merge new facts into the right section. Prefer + `frontmatter_update` for metadata; `file_append` for purely + additive trailing sections; `file_read` + `file_write` for + mid-body edits. Override stale claims rather than stacking + contradictions. Don't restructure unrelated parts. + + - **MOVE / promote** — rename or relocate an existing node: + `file_move` (the retarget pass rewrites inbound wikilinks + atomically; never `file_write` to new + `file_delete` old). + + When CREATE / UPDATE writes a relation into the body, prefer a + typed wikilink (`predicate:: [[X]]` or `[predicate:: [[X]]]`) when + the relation has clear semantic weight; default to bare `[[X]]` for + plain mention. See protocol.md §4.1 for the recommended predicate + vocabulary. + + If a candidate is only a passing mention with no new fact to write, + do nothing — the mention stays in the daily, search will still + find it, and a future digester pass can lift it when it accumulates + substance worth CREATE/UPDATE. + + ### Step 4 — Flip status per daily (mandatory) + + After processing each daily, call: + + frontmatter_update path=daily//.md metadata={status: completed} + + Use `status=skipped` if the daily had nothing worth lifting (chitchat, + dead end). This flip is the daily-tier convention this digester owns + — absent ≡ pending, so the next pass uses + `file_list path=daily recursive=true` + per-item `frontmatter_read` + to find leftover work. Forgetting the flip leaves the daily in the + queue forever. + + ### Step 5 — Reply + + One short paragraph: which dailies you read, what you CREATEd / + UPDATEd / MOVEd (with paths), and which dailies you flipped to + `completed` vs `skipped`. Don't replay every tool call — those are + on the audit trail. + + ## Wikilinks + + Always full path relative to the vault with `.md`: + `[[digest/<…>//.md]]` for digest entries, + `[[daily//.md]]` for daily notes, + `[[resource//]]` for ingested resources. Short forms or + extension-less forms don't resolve. + + ## Boundaries + + - **Never write under `daily/`** except for the Step 4 `status` flip. + - **Only this digester writes `status`** — sync / hand-edits leave + it alone (absent ≡ pending is exactly how this digester finds its + workload). + + ## Memory protocol + + {protocol} + +user_message: | + today: {today} + hint: {hint} + + # Daily notes to distill + + {daily_blob} + + Run the five steps from the system prompt. Reply = one paragraph audit. diff --git a/reme4/steps/jobs/protocol.md b/reme4/steps/jobs/protocol.md new file mode 100644 index 00000000..58e42105 --- /dev/null +++ b/reme4/steps/jobs/protocol.md @@ -0,0 +1,112 @@ +# Memory Protocol + +Opinionated default contract for writing memory into a reme vault. +reme core reserves only `name` / `description` (both optional); +everything below is convention that consumers may replace. + +## 1. Directory architecture + +``` +/ +├── daily/ +│ └── / +│ └── / +│ ├── .md # hot summary note (markdown) +│ └── .* # sibling materials (any file type) +└── digest/ + └── / + ├── .md # cold canonical entry + ├── .md # supporting docs + └── / # nested narrower scope + └── .md +``` + +- **Hot tier** (`daily/…`) — streaming. One upstream writer per + folder; every other consumer treats it as read-only. The + summary note `.md` is markdown; siblings may be any + file type the writer chooses. +- **Cold tier** (`digest/…`) — curated. Each folder is a + scope and must contain `/.md` as its canonical + entry. A scope's other children are sibling material files or + narrower scope subfolders; nesting depth is unconstrained. + Slugs are globally unique — a folder name appears at most once + anywhere under `digest/`. + +Facts flow one-way, `daily/` → `digest/`. References use the +full path relative to the vault: `[[digest///.md]]`. + +## 2. Frontmatter + +Reserved (typed; all optional): + +| key | type | +|---|---| +| `name` | string | +| `description` | string | + +Opinionated default axes (closed enums): + +| key | values | +|---|---| +| `lifecycle` | `streaming` / `evolving` / `frozen` | +| `scope` | `instance` / `class` | +| `source` | `auto` / `curated` / `derived` | +| `role` | `profile` / `concept` / `claim` / `method` / `reference` / `observation` / `question` / `fundamentals` | + +Any other keys consumers want (e.g. a workflow `status` flag) live +as extras — write them, read them with the `where` filter on `list` +tools (`null` matches absent-or-null); the protocol does not name +or enumerate them. + +## 3. Body + +Section structure is **advisory** — `## Summary`, `## Key Facts`, +`## Decisions`, `## Related` are convenient defaults but the +protocol mandates no specific section. + +## 4. Wikilinks + +Three recognized forms: + +| Form | Example | Meaning | +|---|---|---| +| Bare | `See [[张三.md]]` | weakest layer — "mention" | +| Line-level Dataview | `colleague:: [[李四.md]]` | typed relation, queryable by predicate | +| Inline-bracketed Dataview | `主导 [负责:: [[项目X.md]]] 的重构` | typed relation, embedded inline | + +Targets are stored **verbatim** as full paths relative to the vault. +`[[digest/zhang-san/zhang-san.md]]` resolves; short or +extension-less forms do not — no implicit `.md` completion, no +basename search, no folder-note expansion. + +Renaming a node requires atomically rewriting every inbound +wikilink. + +### 4.1 Typed predicates (half-open) + +Recommended core vocabulary — writers may extend beyond this set, +but new predicates should be reused consistently: + +| predicate | meaning | +|---|---| +| `is_a` | hierarchical (X is a kind of Y) | +| `part_of` | containment (X is part of Y) | +| `depends_on` | dependency (X requires Y) | +| `manages` | authority / responsibility | +| `alias_of` | equivalence (X and Y are the same thing) | +| `references` | citation / external pointer | + +Use typed `predicate:: [[X]]` only when the relation has clear +semantic weight; default to bare `[[X]]` for plain mentions. Typed +edges become queryable via `graph_traverse predicate=`. + +### 4.2 One-way write rule + +A wikilink lives in the **source** node's body only — the node whose +prose introduces the relation. The **target** is never modified to +record the inbound relation. Backlinks are discovered at query time +via `graph_traverse direction=in`, never written into target bodies. + +This keeps every write authoritative: a node's body reflects only +what its own author/writer chose to say, never sideways annotations +from other nodes' writers. diff --git a/reme4/steps/jobs/synchronizer.py b/reme4/steps/jobs/synchronizer.py new file mode 100644 index 00000000..0eabaede --- /dev/null +++ b/reme4/steps/jobs/synchronizer.py @@ -0,0 +1,276 @@ +"""Synchronizer — daily-note sync ReAct agent. + +Watches the agent's recent conversation and persists in-progress +tasks as a daily note inside reme's vault_dir, so future agent +invocations can pick the work back up. Pure sync mechanism — does +**not** compress the agent's context (compression is the agent's +own concern). + +Note layout: a single markdown file ``daily//.md``. +Everything worth preserving (verbatim user prompt, key tool output, +intermediate data) goes inline inside this file — there are no +sibling materials. References to it use the full path relative to +the vault (``[[daily//.md]]``); short or no-extension +forms do not resolve. Frontmatter carries ``name`` / ``description`` +plus an optional ``inherits`` wikilink for cross-day continuation. +Body splits into ``Objective`` / ``Plan`` / ``Progress`` / +``Findings`` / ``Decisions`` / ``Next`` / ``References`` sections +(the last is a list of ``[[resource//]]`` wikilinks for +inbound assets landed via ``ingest`` — non-markdown +artefacts the task itself produced are summarized inline). + +Inputs (from RuntimeContext): + messages (list[Msg], required): conversation slice to inspect. + note (str, optional): caller-supplied note hint (task name + or ``daily//.md`` path) to bias slug selection + and disambiguate same-day tasks. + +Output (written to context.response.answer): + { + "skipped": True if the agent reported [SKIP], + "actions": one-line action statement from the agent, + "note": note file path relative to the vault, or None, + "summary": full markdown content of the synced note, + for the calling agent to reload into a + compacted context. None when SKIP / failed. + } + +The agent's toolkit is assembled by ``add_as_tool`` — each entry in +``_NOTE_TOOLS`` is a job name (registered in the active config); +the wrapper turns ``job(**kwargs)`` into a ``ToolResponse``. The job +indirection means the agent sees exactly the same tool surface +(name / description / parameter schema) as the L2 MCP layer. + +Override interface — note shape (single-file layout, frontmatter +fields, section discipline) is opinionated convention, not a core +invariant, so callers can fully replace it without touching reme4: + +* ``prompt_dict`` (inherited from ``BaseStep``) overrides the + ``system_prompt`` / ``user_message`` templates wholesale — this is + how a service layer swaps in its own note schema (e.g. a different + section list, different frontmatter fields). +* ``toolkit`` replaces the tool surface ``_NOTE_TOOLS`` builds. + +What ships here (``synchronizer.yaml``) is just one viable convention; +the service layer (plugin configs, custom callers) is the right place +to pin down the *deployment-specific* shape. +""" + +import datetime +import re +import zoneinfo +from pathlib import Path + +from agentscope.agent import ReActAgent +from agentscope.message import Msg +from agentscope.tool import Toolkit +from pydantic import BaseModel, Field + +from ..base_step import BaseStep + +from ...components import R + + +_NOTE_PATH_RE = re.compile(r"daily/\d{4}-\d{2}-\d{2}/[^/\s]+\.md") + + +_NOTE_TOOLS: tuple[str, ...] = ( + "file_list", + "file_read", + "file_write", + "file_append", + "file_edit", + "file_stat", + "frontmatter_read", + "frontmatter_update", + "frontmatter_delete", + "daily_read", + "daily_write", + "daily_reindex", +) + + +def _coerce_messages(raw) -> list[Msg]: + """Normalize incoming messages to ``Msg`` instances. + + The Python caller hands in ``list[Msg]`` directly; the MCP layer + delivers ``list[dict]`` (each dict shaped roughly ``{name?, role?, + content?}``). Both shapes land here. + """ + if not raw: + return [] + out: list[Msg] = [] + for item in raw: + if isinstance(item, Msg): + out.append(item) + continue + if isinstance(item, dict): + out.append( + Msg( + name=item.get("name") or item.get("role") or "user", + role=item.get("role") or "user", + content=item.get("content", ""), + ), + ) + return out + + +def _format_history(messages: list[Msg]) -> str: + """Render the conversation as a speaker-tagged transcript. + + Skips messages whose text content is empty (tool-only frames + don't help the LLM judge task state). + """ + if not messages: + return "(empty)" + lines: list[str] = [] + for msg in messages: + speaker = msg.name or msg.role or "?" + text = (msg.get_text_content() or "").strip() + if not text: + continue + lines.append(f"[{speaker}]\n{text}") + return "\n\n".join(lines) or "(no text)" + + +class SynchronizerResult(BaseModel): + """Outcome of a single note-sync call. + + Without per-tool audit (the agent's toolkit is the job surface, + which doesn't expose per-call records back to the orchestrator), + the structured outcome is just what the agent reports plus what + we re-read from disk after it returns. + """ + + used_llm: bool = Field(default=False) + skipped: bool = Field(default=False) + actions: str = Field( + default="", + description="One-line action statement from the agent (e.g. " + "'updated daily/2026-05-15/auth-refactor.md' or '[SKIP]').", + ) + note: str | None = Field( + default=None, + description="Note file path relative to the vault, " + "e.g. 'daily/2026-05-15/auth-refactor.md'. None when SKIP / failed.", + ) + summary: str | None = Field( + default=None, + description="Full markdown content of the note file. Lets the calling " + "agent reload the warm summary into a freshly compacted context without an extra read.", + ) + + +@R.register("synchronizer") +class Synchronizer(BaseStep): + """Drive daily-note sync via a ReAct agent.""" + + def __init__( + self, + toolkit: Toolkit | None = None, + console_enabled: bool = False, + timezone: str | None = None, + inherit_window_days: int = 7, + **kwargs, + ): + super().__init__(**kwargs) + self.toolkit = toolkit + self.console_enabled = console_enabled + self.timezone = timezone + self.inherit_window_days = inherit_window_days + + def _now(self) -> datetime.datetime: + if self.timezone: + try: + return datetime.datetime.now(zoneinfo.ZoneInfo(self.timezone)) + except Exception as e: + self.logger.error( + f"Invalid timezone {self.timezone!r}, falling back to local time: {e}", + ) + return datetime.datetime.now() + + def _vault_dir(self) -> Path: + wd = getattr(self.file_store, "vault_path", None) + return Path(wd).resolve() if wd else Path.cwd().resolve() + + def _build_toolkit(self) -> Toolkit: + """Bind every note-relevant job as a tool function. + + Each entry in ``_NOTE_TOOLS`` is a job name registered in the + active config; ``add_as_tool`` wraps ``job(**kwargs)`` into a + ``ToolResponse``. The job indirection means the agent sees exactly + the same tool surface (name / description / parameter schema) as + the L2 MCP layer. + """ + toolkit = self.toolkit or Toolkit() + for job_name in _NOTE_TOOLS: + self.add_as_tool(toolkit, job_name) + return toolkit + + async def execute(self): + assert self.context is not None + messages: list[Msg] = _coerce_messages(self.context.get("messages")) + note_hint: str = self.context.get("note", "") or "" + + if not messages: + result = SynchronizerResult(used_llm=False, skipped=True) + self.context.response.success = True + self.context.response.answer = "Skipped: no messages supplied" + self.context.response.metadata.update(result.model_dump()) + return + + toolkit = self._build_toolkit() + + agent = ReActAgent( + name="reme_synchronizer", + model=self.as_llm, + sys_prompt=self.prompt_format("system_prompt"), + formatter=self.as_llm_formatter, + toolkit=toolkit, + ) + agent.set_console_output_enabled(self.console_enabled) + + user_message: str = self.prompt_format( + "user_message", + today=self._now().strftime("%Y-%m-%d"), + vault_dir=str(self._vault_dir()), + inherit_window_days=self.inherit_window_days, + note=note_hint or "(none)", + history=_format_history(messages), + ) + + final_msg: Msg = await agent.reply( + Msg(name="reme", role="user", content=user_message), + ) + actions = (final_msg.get_text_content() or "").strip() + + result = SynchronizerResult(used_llm=True, actions=actions) + if "[SKIP]" in actions.upper(): + result.skipped = True + + # Reload the freshly written note so the calling agent can drop + # it back into a compacted context without an extra read trip. + if not result.skipped: + self._reload_note(result, actions) + + self.context.response.success = True + self.context.response.answer = actions or "Synchronization completed" + self.context.response.metadata.update(result.model_dump()) + + def _reload_note(self, result: SynchronizerResult, actions: str) -> None: + """Parse the agent's action line for the note path and read the + full file back into ``result.summary``. Best-effort: a parse miss + leaves the context-management fields as None but does not fail + the step (persistence already succeeded).""" + match = _NOTE_PATH_RE.search(actions) + if not match: + return + note_path = match.group(0) + try: + absolute = (Path(self.file_store.vault_path or ".") / note_path).resolve() + text = absolute.read_text(encoding="utf-8") + except Exception as e: + self.logger.warning(f"synchronizer: could not reload note {note_path!r}: {e}") + return + result.note = note_path + result.summary = text diff --git a/reme4/steps/jobs/synchronizer.yaml b/reme4/steps/jobs/synchronizer.yaml new file mode 100644 index 00000000..137e130f --- /dev/null +++ b/reme4/steps/jobs/synchronizer.yaml @@ -0,0 +1,144 @@ +system_prompt: | + You persist an in-progress task as a daily note so the next session + can pick it up. One call writes at most one note. + + ## Note shape + + /daily//.md # the whole note, one file + + One note per task per day; cross-day continuation uses an + `inherits:` frontmatter wikilink. Anything worth preserving + (verbatim user prompt, key tool output, intermediate data) goes + inline inside this file — there are no sibling materials. + + ## Five steps per call + + ### Step 1 — Skip check + + Is the conversation a real multi-step task in progress? Casual Q&A, a + single-shot answer, or idle chat → reply `[SKIP]` (alone, literal) and + stop. When genuinely ambiguous, default to writing — losing work is + worse than an extra note. + + ### Step 2 — Pick the slug + + - Note hint already a kebab-case slug → use verbatim. + - Otherwise mint a stable kebab-case from the task topic (≤60 chars). + - **Reuse the same slug across calls for the same logical thread** — + same slug = same file = upsert. Fragmenting one thread across + multiple slugs is the worst failure mode. + + ### Step 3 — Discover the branch + + Call `file_list path=daily/{today}` (and `daily_read` / + `frontmatter_read` on candidates as needed) to determine one of three + branches: + + - **UPDATE** — `daily/{today}/.md` already exists: + `daily_read slug=` to fetch body + frontmatter, merge per + Step 4, then `daily_write slug= overwrite=true` with the + full new body + frontmatter. + + - **INHERIT** — no file today, but within the last + {inherit_window_days} days an active note with the same name + exists: confirm via `daily_read` on the earlier note, then + `daily_write slug=` (default `overwrite=false`) with a + fresh body that sets `inherits: [[daily//.md]]` + in frontmatter and copies the predecessor's `Objective` + `Plan`. + Progress / Findings / Decisions / Next start empty. **Do not + modify the predecessor.** + + - **CREATE** — neither: `daily_write slug=` (default + `overwrite=false`) with the full body + frontmatter. Idempotent — + no-ops if a same-slug note already exists today (caller falls + back to the UPDATE branch). + + ### Step 4 — Write + + Pick the smallest write for each change: + + | Change | Tool | + |---|---| + | New / replacement full note | `daily_write` (one shot: body + frontmatter + index refresh) | + | Append to a trailing append-only section | `file_append` — cheaper than R-M-W | + | One frontmatter key | `frontmatter_update` (call `daily_reindex` after if `name`/`description` changed) | + | Mid-body restructure | `daily_read` + `daily_write overwrite=true` | + + Suggested body shape (sections are convention — adapt as fits): + + ```markdown + --- + name: + description: <2-3 sentences: what + why> + inherits: [[daily//.md]] # INHERIT only + --- + ## Objective + + + ## Plan + + + ## Progress + - + + ## Findings + - + + ## Decisions + - + + ## Next + - [ ] + + ## References + - [[resource//]] — # external assets the task consumed + ``` + + Section discipline: Progress / Findings / Decisions are append-only + (never delete history). Plan / Next are wholesale-rewritten each call. + Objective is set once. `## References` lists `[[resource//]]` + wikilinks for inbound assets that arrived through an external channel + and were landed by `ingest`; non-markdown content the task + itself produced (raw outputs, screenshots) should be summarized inline + or skipped — there is no per-note sibling folder anymore. + + Wikilink form is fixed: note refs use the full vault-relative + path with `.md` (`[[daily//.md]]`); resource refs use the + canonical resource path (`[[resource//]]`). Short forms + don't resolve. + + ### Step 5 — Emit one line + + Your final reply must be exactly one line in this form: + + daily//.md + + - `` ∈ `created` / `inherited` / `updated` + + Or `[SKIP]` (literal, alone) if Step 1 said skip. This line is parsed + mechanically — match the format exactly. + + ## Boundaries + + - **Never write under `digest/`** — that tier is downstream. + - **Never write the `status` frontmatter** — `status` is reserved for + the downstream distillation pass, which uses absence to find pending + work. Touching it from here makes the note invisible to the + next distill run. + - **`daily_write` defaults to `overwrite=false`** (idempotent skip-if-exists, + mirroring the old `daily_resolve` probe). Pass `overwrite=true` only when + you've already read the file via `daily_read` and intend to replace it + (the UPDATE branch). On a surprise collision (`created: false`) — fall + back to UPDATE rather than blindly overwriting. + +user_message: | + Today: {today} + Vault dir: {vault_dir} + Inherit window: last {inherit_window_days} days + Note hint: {note} + + # Recent conversation + + {history} + + Run the five steps from the system prompt. Final reply = one line.