feat(jobs): add digester step for knowledge distillation from daily notes (#261)
Some checks are pending
Pre-commit / run (ubuntu-latest) (push) Waiting to run
Tests ReMe / Unit Tests - py3.10 (push) Waiting to run
Tests ReMe / Unit Tests - py3.13 (push) Waiting to run

* feat(jobs): add digester step for knowledge distillation from daily notes

* feat: add initial implementation

* refactor(jobs): remove unnecessary blank lines in digester and synchronizer

* feat: add initial implementation
This commit is contained in:
Sen Huang 2026-05-28 15:17:23 +08:00 • committed by GitHub
parent a4efc0f776
commit 8c48798164
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
11 changed files with 2912 additions and 0 deletions

578
docs4/auto_dream_design.md Normal file
View file

@ -0,0 +1,578 @@
# auto-dream 设计(digest 沉淀:物理 + 图)
> 本文档记录 reme4 中 **auto-dream**(digest 沉淀知识层)的设计讨论 —— 含物理布局、图模型(节点 + 边)、入流端 digester(G\*)。
>
> 配套阅读:
> - `structure.md` §1.2(数据视角)/ §2(三层存储)/ §3.5(digest 动作)/ §7(L4 模块)
> - `auto_memory_design.md`:auto-memory(daily 实时事件)是 dream 的入流之一(G1 scope 拉取);auto-memory 写完即对 dream 可见
> - `auto_maintain_design.md`:digest 的组织 / 重组 / 写入运行时(M split / D 检测 / CAS 写入协议) —— 本文档定义节点 + 边模型与 G\*,maintain 定义运行时
> - `auto_link_design.md`:auto-link 是 dream 写完之后的后置增强(实体识别 + wikilink 写回);复用 maintain 的 CAS 协议
>
> **四层对应**:reme 服务整体四份设计 —— auto-memory(daily 入流)/ **auto-dream(本文档:digest 沉淀 + content link + G\*)** / auto-maintain(digest 组织端:split + 检测 + CAS)/ auto-link(背景图关系增强)。`structure.md` §3.5-3.6 的 L4 action 视角:dream 对应 `digest` action(`resource + daily → digest`),maintain 对应 `maintain` action(`digest → digest`)。auto-dream 承担"空闲整理"(报告 §5.2):把 daily / resource 的散点抽出共性、合并重复、生成总结,落到 `digest/<bucket>/<slug>.md`。
>
> **核心立场**:dream 定义**模型层**(节点 + 边 / 守恒规则)+ **生成侧**(G\* create_or_update)+ **召回**(SearchStep);组织端(M split / D 检测 / CAS / 时序)归 `auto_maintain_design.md`;后置增强(实体识别 / wikilink 写回)归 `auto_link_design.md`。三方共享 §1.5 节点 + 边模型 + §1.5.5 边守恒 + §1.5.4 F-invariants。
>
> **关键收敛**:digest 不分"逻辑层"。整个 digest = **物理布局(浅桶 + flat .md)** + **一张图(节点 + 边)**。没有 hub / topic / leaf 之分 —— 所有 .md 文件都是同一种节点,内容决定它扮演什么角色(主题概览 / 概念定义 / 方法描述 / 实体记录 ...)。"主题"是从图中涌现的,不是结构性宣告的。
---
## 0. 问题陈述
digest 是 agent 长期记忆的"组织化沉淀"层,与三层架构的另两层职责互补:
| 层 | 组织主轴 | 形态 |
|---|---|---|
| resource/ | 时间(`<date>/<name>`) | 外部原始资料,不可变 |
| daily/ | 时间 + 任务(`<date>/<slug>/`) | agent 任务过程,半可变 |
| **digest/** | **语义** | **跨任务知识,可重组** |
digest 的核心问题:
- **物理布局**:文件系统怎么组织?(目录 / 文件 / 路径)
- **节点与边**:概念粒度 / 链接形态 / 主题如何涌现?
- **演化机制**:从碎片到网络的过程谁负责?(digester 创建或更新节点 / maintainer 拆分过载节点 / D 信号写后 inline 检测)
---
## 1. 已对齐决策
### 1.1 节点粒度:Atomic 节点优先
| 项 | 决策 |
|---|---|
| **粒度** | 一个文件 = 一个原子单元(概念 / 方法 / 实体 / 案例 / 原则 / 主题概览) |
| **节点角色** | 由内容决定,不由 frontmatter 类型标记;一个节点扮演"主题概览"还是"具体方法",看它的 body 写了什么 |
| **风格参考** | Zettelkasten:atomic note + 高密度 wikilink 网络 |
**理由**:
1. **节点粒度 = retrieve 精度上限** —— semantic 检索召回 "一个原子单元" 远比召回 "一个 5000 字的主题文档" 信噪比高;agent 上下文窗口经不起粗粒度文档塞满
2. **wikilink 在 atomic 粒度才真有意义** —— `[[jwt-rotation]]` 指向"一个具体方法"比指向"auth 主题文档"精确一个数量级,这也是 I-4(wikilink 唯一跨层载体)能撑起来的前提
3. **物理目录纯做 navigate,图承担关系** —— 两套机制各司其职,不互相绑架
**Tradeoff**:
- 文件数量爆炸(一个领域几百节点)→ 浅桶物理归档 + 路径作 ID,wikilink 走完整路径写法
- 节点会频繁演化(新材料 update 已有节点 / 节点过载触发 split)→ G\* 去重质量、SearchStep 召回、写后 inline 检测都得到位
### 1.2 物理几何:浅桶(shallow bucket,**固定集合**)
| 项 | 决策 |
|---|---|
| **物理布局** | `digest/<bucket>/<slug>.md`,bucket 一层(顶多两层),桶内 flat |
| **bucket 角色** | **仅承担物理归档与 OS-level 浏览锚点**;不承担语义本体角色 —— 主题由图中的节点表达 |
| **bucket 内** | 不再分子目录,所有节点平铺 |
| **bucket 集合** | **固定预定义,不由 digester / maintainer 动态生成** |
| **bucket 主页节点** | **不强制存在**;若 split 在该 bucket 内累积出层级(parent 节点天然中心性高),parent 节点天然成为浏览主页(纯约定,非架构必需) |
| **新节点归属** | digester (G4) 只能从已有桶集合里**挑选**;LLM 不能造新桶 |
| **集合来源** | vault 配置(opinionated default + 消费层可改),与 schema/prompt 同属服务消费层 |
**理由**:
- 物理浏览有"主题轮廓"(打开 `digest/auth/` 能看到这一族节点),不像纯 flat 那样毫无锚点
- 节点不被深路径绑死("属 auth/jwt 还是 auth/session"这种归属焦虑被消解 —— 一个节点可以同时被多个主题通过 wikilink 引用)
- maintain 操作面坍缩到只剩 split:节点过载就拆,不做跨节点重组(详 §1.5 / §2)
- **固定集合的关键意义**:
- LLM 在 G4 桶决定时只做"分类",不做"造类" —— 决策面坍缩,错率大幅下降
- bucket 集合作为**预先约定的物理归档规则**,跨任务跨时间稳定;不会出现 "auth-stuff" / "auth" / "authentication" 三个语义重叠的桶共存
- 与 reme 核心立场一致:bucket 集合是消费层契约,reme 不自主造桶
- **bucket 主页节点不强制的关键意义**:
- 旧设计中 root hub 是架构必需(检测稳定锚点 / vault 总览结构性入口);新设计中这些角色都由图中心性自然承担,不需要为每个 bucket 强制创建一个空架子节点
- 主页节点的"主页"地位是涌现的:某节点入度高 / 中心性高 → 它就是浏览入口
- 若一个 bucket 完全没节点,它就只是个空目录;不需要先造一个 placeholder
**未归类节点**:digester 抽到一个原子单元但找不到合适的专属桶时,**不允许造新桶**;统一落入兜底桶 `digest/general/`。general 桶在固定集合内是一等公民,详见 §3.7。
### 1.3 节点身份:路径即 ID
| 项 | 决策 |
|---|---|
| **ID 载体** | **vault-relative 路径**(含 `.md`)即节点身份 —— `digest/auth/jwt-rotation.md` |
| **wikilink 写法** | `[[digest/auth/jwt-rotation.md]]`(literal,与 `wikilink_handler.py` 默认形态对齐;不隐含 `.md`,无 short-form 补全) |
| **`name` frontmatter** | 文件名 basename(不含扩展名),与文件名同步 —— 检索 hint / 人读标签,**不当 ID 用** |
| **同名冲突** | 同 bucket 内文件名冲突 → 文件系统层断言;**不需要独立 D6 信号** |
| **rename 成本** | 一次 `wikilink_handler.retarget_links(old_path, new_path)`,机制现成 |
| **跨桶移动** | F-1 已禁止;若必须做(人工介入修错桶),走一次 retarget |
**理由**:
- F-1(0 文件移动)+ 平铺 + 下层 immutable 后,slug abstraction 的核心价值(移动鲁棒性)蒸发;只剩下"wikilink 短形式"这一项收益,但代价是 file_graph slug 索引 + D6 冲突检测 + alias 表 + retrieve 透明展开,**净亏**
- `wikilink_handler.py` docstring 自己写的就是 *Recommended form: full path relative to the vault with extension* —— literal 匹配,无 short-link 补全;路径作 ID 与核心库默认完全对齐
- 完整路径前缀 `digest/auth/` 给 LLM 读写时提供语义 context(知道节点在哪个桶),不全是负担
- provenance wikilink 反指 daily/resource 本来就用路径,统一后整个 vault 一种 wikilink 形态,不必区分"slug 形态 vs 路径形态"
### 1.4 链接语义:基础 wikilink + 可选 Dataview 谓词
参考实现:`reme4/utils/wikilink_handler.py` + `reme4/schema/file_link.py`。
| 项 | 决策 |
|---|---|
| **link 基础形态** | `[[<vault-relative-path>.md]]` —— target 字面取(literal,不隐含 `.md`,不自动短链补全);路径即 ID(详 §1.3) |
| **alias / image** | `[[path.md\|alias]]`(显示文本)/ `![[image.png]]`(图片资源)—— rewrite 时 alias 保持 |
| **anchor 不引入** | digest 设计层**不使用** `[[path.md#section]]` —— atomic 节点 + child 边界已充当精度替代品(详 §1.3 / §3.13);`wikilink_handler` 仍可解析 anchor 字面(供其它消费层),但 digest 不生成、不依赖、不在 split 时迁移 anchor |
| **可选谓词(Dataview 风格)** | 行级: `predicate:: [[path.md]]` / 内联: `[predicate:: [[path.md]]]`;**谓词写在 `[[]]` 外**,不是 `[[predicate::path]]` |
| **谓词标识符** | `[A-Za-z][A-Za-z0-9_]*`(如 `is_a` / `extends` / `causes` / `references`);词表**开放**,任意标识符 |
| **未类型化合法** | 绝大多数 wikilink **不加** predicate;`predicate=None` 是默认 / 常态 |
| **edge 唯一性键** | `(target_path, predicate)`(二元组);同源同标但不同 predicate = 不同边。`FileLink.target_anchor` 字段在 schema 中保留(供其它消费层),digest 层永远写 `None` |
| **类型信息载体** | 节点 frontmatter `kind` + 边 `predicate`(双轨可选);二者都是**消费层 schema 提示**,reme 核心解析 / 存储 / 索引,但**不读它们做结构决策** |
| **plugin 层扩展** | transclusion / 引用图谱视图等留给消费层加,reme 核心不固化语义 |
**理由**:
- 与 I-4 "wikilink 是唯一跨层载体" 对齐 —— reme 核心永远只看机械拓扑
- 与 [[reme4_schema_layering]] 一致 —— 类型语义(无论是节点 `kind` 还是边 `predicate`)都是消费层契约,reme 核心不固化
- predicate 走 Dataview 而非内嵌:`[[]]` 内容保持"纯目标"(rewrite / retarget 不必感知 predicate);predicate 是文本上的**装饰位**,与 wikilink 解耦
- `kind` 与 predicate 严格只是内容标签:reme 核心**只有节点这一种结构类型 + 边这一种结构关系**,kind / predicate 永远不参与"hub / topic / leaf"这类结构角色判断
**reme 核心对 predicate 的"透明"边界**(关键):
- G7(横向 link)、retrieve 中心性 —— 都**聚合所有 predicate** 算,不分桶
- 只有 edge 唯一性 / 反向索引会用到 predicate(否则 `[[A]]` 和 `is_a:: [[A]]` 会被当作同一条边互相覆盖)
- 消费层若要按 predicate 做更精细的推理(如"taxonomic 路径只走 `is_a` 边"),自己读 `FileLink.predicate` 即可
### 1.5 图模型与节点演化
**核心模型**:digest = **物理布局(浅桶 + flat .md)** + **一张图(节点 + 边)**。
| 维度 | 形态 |
|---|---|
| **节点** | 每个 .md 文件 = 一个节点;无结构性 kind,角色由 body 内容决定(主题概览 / 概念定义 / 方法描述 / 实体记录 / 案例 ...) |
| **边** | 基础 `[[<vault-path>.md]]` wikilink(路径即 target,详 §1.3 / §1.4);可选 Dataview 谓词 `predicate:: [[path.md]]` / `[predicate:: [[path.md]]]` 写在 `[[]]` 外;边唯一性键 = `(target, predicate)`;**digest 层不引入 anchor**(详 §1.4 / §3.13);**reme 核心结构决策不读 predicate**(详 §1.4) |
| **多归属** | 一个节点可被多个其它节点引用,也可指向多个其它节点;**不存在"单父"约束** |
**演化只做两件事**:
1. **G\* create_or_update**:新材料进入,LLM 提取原子单元 → 命中已有节点就 update 该节点 body(语义守恒地融合新旧),否则新建节点
2. **M split**:节点累积过载(token / 主题离散度超阈值)→ LLM 把它拆成 parent overview + N 个 children,parent 文件原地保留作 overview,children 是新文件
"主题概览节点 / 摘要节点"不是一种 kind,也不是 maintainer 主动涌现的产物 —— 它是 split 的副产品(parent 节点天然成为该 cluster 的 overview)。
#### 1.5.1 单一节点 / 单一边
**节点 frontmatter** —— 只有保留字段:
| 字段 | 内容 | 用途 |
|---|---|---|
| `name` | 文件名 basename(不含扩展名),与文件名同步 | 检索 hint / 人读标签(I-4 不再用它做身份;路径才是 ID,详 §1.3) |
| `description` | 一句话 | 标题 / 检索 hint |
| (可选)`kind` | concept / method / case / entity / topic / ... | **消费层 schema 提示**,reme 核心透明,不读它做结构决策 |
| body | 任意内容 | 一句定义 / 一段方法 / 一篇主题概览 / 一份案例,皆可 |
`hub__` / `topic__` 前缀**不存在**;文件名自然命名(`auth-fundamentals.md` / `jwt-rotation.md` / `jwt-overview.md`)。"overview 节点"靠内容形态识别,不靠前缀。
**边的形态** —— 与 §1.4 一致,这里给最小汇总:
| 维度 | 形态 |
|---|---|
| 基础 | `[[<vault-path>.md]]`(无谓词;常态;`predicate=None`) |
| 可选谓词 | `predicate:: [[path.md]]`(行级)/ `[predicate:: [[path.md]]]`(内联);谓词在 `[[]]` 外 |
| alias | `[[path.md|display-text]]`(rewrite 时 alias 保持) |
| image | `![[image.png]]`(资源引用,不是知识边) |
| **不引入 anchor** | digest 层不使用 `#section`;详 §1.4 / §3.13 |
| **边唯一性键** | `(target_path, predicate)` —— 同源同标不同 predicate = 不同边 |
| **角色识别** | 默认无谓词时由端点内容形态推断;有谓词时谓词即角色标签(消费层语义,核心不读) |
> **关键收敛**:reme 核心**只有节点 + 边两种结构类型**;kind / predicate 都是内容标签,绝不参与 hub / topic / leaf 这类结构角色。
#### 1.5.2 节点演化:create_or_update + split
```
[入流] material 进入(daily / resource)
│
▼
G1 scope:选哪些 daily/resource 进入本轮
│
▼
G2 提取原子单元(LLM,可能产 N 个候选)
│
▼
对每个候选:
│
▼
G* create_or_update(LLM 决策点)
├─ 语义相似查 → 拉相似候选节点(top-k)
├─ LLM 判:候选中有"同概念节点"吗?
│ ├─ 有 → update 路径
│ │ (a) 把新内容融入已有 body(语义守恒重写)
│ │ (b) 加 provenance 反指
│ │ (c) 必要时加 / 改 wikilink
│ └─ 无 → create 路径
│ G3 路径 / G4 bucket / G6 provenance / G7 横向 link / 写 body
│
▼
写入(机械)
[D3 检测] 写后立即:G\* / split 写完 body 顺手 inline 检测(token 阈值 → 超阈值则 LLM 判离散度;详 `auto_maintain_design.md` §4)
│
▼
D3 派发候选节点(F-4 一次一个)
│
▼
M split(LLM + 机械)
├─ LLM 把节点 body 拆成 N 个 cluster(每个是个原子单元)
├─ parent 文件原地保留 → body 重写为 overview + 列出 children wikilinks
├─ 每个 child 创建新文件(文件名 / bucket / body 由 LLM 给)
├─ children 各自加 [[<parent-path>.md]] 反向链接
├─ 边守恒机械校验:`(parent_new ∪ ∪children) ⊇ parent_old` 出边集合(F-11 / E-2)
└─ inbound 链 `[[<parent-path>.md]]` 不动(F-10 / E-3) —— digest 层无 anchor 链,无需 retarget
```
**关键**:
- **G\* update 改 subject body(语义守恒重写)** —— 新材料融入已有节点正文,要求 LLM 守住"只增不删 / 不改原意",老内容不能丢;**写入前机械校验出边强守恒**(`new outbound ⊇ old outbound`,详 §1.5.5 E-1)
- **G\* 不改其它节点正文** —— 只动 subject;不像旧 M-E 会到邻居 body append wikilink
- **M split 不改其它节点正文** —— 只动 parent(重写为 overview)+ 新建 children
- **inbound 在 split 时一律不动** —— digest 设计不引入 anchor,inbound 全是裸链 `[[<parent-path>.md]]`,parent 路径未变即天然有效;后续 G\* 进入若 LLM 觉得 child 粒度更合适,直接加新边到 child(F-10)
#### 1.5.3 走一个具体例子
**场景**:`digest/auth/` 桶,初始只有几个零散 auth 节点,没有 jwt-rotation。
**第 1 轮 G\***:某 daily 提到 "JWT rotation:每 24 小时换密钥,旧密钥保留 1 小时窗口给未过期 token"。
- 语义查 → 没找到 jwt-rotation 节点
- 走 create 路径 → 新建 `jwt-rotation.md`,body = 一段 200 字的 rotation 描述
**第 2 轮 G\***:另一个 daily 提到 "JWT rotation 的 grace period 通常是 1-2 小时"。
- 语义查 → 命中 `jwt-rotation`(高相似)
- LLM 判:这是同概念,走 update 路径
- 把 grace period 信息**融入** `jwt-rotation.md` body(不只是 append):
```
Before: "...旧密钥保留 1 小时窗口..."
After: "...旧密钥保留 1-2 小时 grace period(典型值,具体看 token 寿命)..."
```
- body 略增长,加一条 provenance 反指
**第 N 轮 G\***:经过几个月,各种 daily 持续 update `jwt-rotation` —— 加了密钥派生算法、加了 RS256/HS256 区别、加了 key rotation 失败处理、加了 with-leeway 实践、加了 monitoring 建议 ...
`jwt-rotation.md` body 现在 ~3500 token,涵盖:轮换策略 / 密钥派生 / 算法选择 / 失败处理 / 监控。
**触发**:D3 检测 token > 2000 阈值 → 派发候选。
**M split**:LLM 拉 `jwt-rotation.md` body + frontmatter,判断主题离散度(5 个相对独立的子主题),决定拆:
- parent: `jwt-rotation`(留下,body 重写为 overview)
- children: `jwt-key-derivation` / `jwt-algorithm-selection` / `jwt-rotation-failure-handling` / `jwt-rotation-monitoring`(4 个新文件)
- "with-leeway 实践"内容并入 parent overview(粒度太细不单独拆)
**执行后**:
```
digest/auth/
├── jwt-rotation.md ← 文件原地;body 重写为 overview
│ description: JWT 轮换策略总览
│ body: JWT 轮换的核心是 X,主要环节包括:
│ 密钥派生 [[digest/auth/jwt-key-derivation.md]]
│ 算法选择 [[digest/auth/jwt-algorithm-selection.md]]
│ 失败处理 [[digest/auth/jwt-rotation-failure-handling.md]]
│ 监控告警 [[digest/auth/jwt-rotation-monitoring.md]]
│ (参考:with-leeway 实践 ...)
├── jwt-key-derivation.md ← 新建,body 来自原 jwt-rotation 拆出片段
│ → [[digest/auth/jwt-rotation.md]] ← child 反指 parent
├── jwt-algorithm-selection.md ← 同上
│ → [[digest/auth/jwt-rotation.md]]
├── jwt-rotation-failure-handling.md ← 同上
│ → [[digest/auth/jwt-rotation.md]]
├── jwt-rotation-monitoring.md ← 同上
│ → [[digest/auth/jwt-rotation.md]]
└── ...(其它 auth 节点不变)
```
**inbound 不动**:之前指 `jwt-rotation.md` 的所有外部 wikilink(无论裸链还是 typed)都仍然指 `[[digest/auth/jwt-rotation.md]]`。如果后续某外部节点写新材料时 LLM 觉得 child 粒度更合适,直接 G\* 时新加 `[[digest/auth/jwt-key-derivation.md]]` 这种边即可 —— 不强求 split 时即时重定向。
**继续演化**:
- 若某个 child(如 `jwt-key-derivation`)被持续 update,某天也长到过载 → D3 又触发 → 它再次 split,自然涌现第三层
- 若某 child 长期空 / 0 入度 / 0 update —— 不主动删(没有 dissolve 操作);除非人工介入
#### 1.5.4 操作的核心约束
| # | 约束 | 含义 |
|---|---|---|
| **F-1** | **0 文件移动** | G\* / split 都不移动现有文件;split 创建的是**新文件**,parent 文件原地 |
| **F-2** | **改正文限定 subject** | G\* update 改 subject node body(语义守恒重写,不改其它节点);M split 改 parent body(重写为 overview)+ 创建 children body;**没有任何操作改"其它节点正文"** |
| **F-3** | **maintainer 只做 split** | 没有 summarize / merge / re-edge / link / unify / dissolve;过载 → 拆 |
| **F-4** | **一次一个候选** | M split 一次拆一个节点;G\* 一次处理一个原子单元(N 个候选 = N 次 G\*) |
| **F-5** | **不确定时不动** | G\* 拿不准是 create 还是 update → 倾向 create(不污染已有节点);split 拿不准 cluster 边界 → 不拆 |
| **F-7** | **多归属合法** | 一个节点可被多个其它节点引用,也可指向多个其它节点;**没有"单父"约束** |
| **F-10** | **inbound 目标节点不动** | 所有 inbound 都是裸链 `[[<parent-path>.md]]`(digest 不引入 anchor,详 §1.4 / §3.13);split 时全部保持不动,parent 路径未变即天然有效;后续 G\* 进入时 LLM 可自由选择更精细 target(直接加新边到 child) |
| **F-11** | **wikilink 是 body 的一部分** | 不存在"独立的边" —— 边的所有迁移都是 body 文本变化的副作用;reme 核心机械算子只感知字符层,语义责任在 LLM(G\* / split prompt)+ 守恒校验(outbound diff 机械验证;详 §1.5.5) |
#### 1.5.5 边的迁移规则
**前提**:wikilink 是 body 的一部分(F-11)。"边"不是独立抽象 —— body 一变,边就跟着变。reme 核心**没有"修边"算子**,边的所有变化都是 body 文本编辑的副作用。
但语义守恒不能放任 LLM:守恒责任在 prompt + 机械校验,不在算子。
**3 类边按"在哪类操作中变化"区分**:
| # | 类别 | 规则 | 谁负责 |
|---|---|---|---|
| **E-1** | **G\* update 节点出边**(subject 自身) | **强守恒**:新 body 出边集合 ⊇ 原 body 出边集合(`(target, predicate)` 二元组比对,**predicate 一并守住**);不满足 → LLM 重试或拒写 | LLM(prompt 强约束)+ 机械校验(outbound diff) |
| **E-2** | **split parent 出边**(parent body 拆解) | parent overview + N 个 children 各持一段,原 parent 出边按内容自然分配到 parent overview + children;**机械校验合计守恒**:`(parent_new ∪ ∪children_outbound) ⊇ parent_old` | LLM(split prompt)+ 机械校验 |
| **E-3** | **inbound wikilink** `[[<parent-path>.md]]` | split 时**不动** —— 仍指 parent;后续 G\* 进入若 LLM 觉得 child 粒度更合适,直接加新边到 child(F-10) | 不动 |
**provenance 不单列一类**:节点反指上游 daily/resource 的 wikilink 是 body 正文的一部分(§3.9),由 LLM 在 G\* / split prompt 中自然写出 —— 跟其它 body wikilink 走同一套规则:G\* update 走 E-1 强守恒(老 provenance 链不能丢,新材料追加新 provenance),split 走 E-2 合计守恒(parent 全量 provenance ⊆ parent_overview ∪ ∪children_outbound)。reme 核心**没有** provenance 专用算子。
**inbound anchor 这一类不存在**:digest 设计层不引入 anchor(详 §1.4 / §3.13),所有 inbound 都是裸链,走 E-3 即可,无需机械 retarget 子流程。
> **关键拆分**:
> - **守恒**(E-1 / E-2):LLM 写正文时不能丢边;靠 prompt + 写后 outbound diff 校验
> - **保守**(E-3):没有信号说一定要变;不变的代价 = 后续 G\* 自然纠正,变的代价 = 错信号大量假阳;选不变
**机械 outbound diff 校验**(E-1 / E-2)伪码:
```
write_subject_body(subject, new_body):
old_outbound = extract_links(old_body) # set of (target, predicate)
new_outbound = extract_links(new_body)
missing = old_outbound - new_outbound
if missing:
# LLM 漏了原边 —— 重试一次
new_body = llm_retry_with_missing(missing)
new_outbound = extract_links(new_body)
if old_outbound - new_outbound:
raise ConservationViolation(...) # 拒写,记 audit,等人介入
write(subject, new_body)
```
机械层只做集合比对,**不判断"为什么丢了"** —— 那是 LLM 的事。
**强守恒(集合包含)而非等价**:`new ⊇ old` 是"新材料融入,老知识保留"的最小契约 —— 允许加新边(新关联),不允许减边(老内容不能丢);等价(`new == old`)会拒绝任何新出边,update 失去意义。
**predicate 守住** —— `[[A]]` ↔ `is_a:: [[A]]` 视为不同 key,升降级走显式 audit 路径,不走默认。重排 / 改 alias / 加新边都不被拦下(集合相同或只增)。
#### 1.5.6 图模型 vs 建子目录
| 维度 | 建子目录(深树) | 图模型 + split 演化 |
|---|---|---|
| 物理变化 | 移动文件,改路径 | 0 文件移动(F-1);split 只创建新文件 |
| wikilink 影响 | 路径 ID 模型下要全图 retarget(代价大) | 0 影响(parent 路径未动);split 不触发任何 retarget |
| 主题归属 | 一个节点只能属一棵子树 | 一个节点可同时属多个主题(被多源 wikilink) |
| 撤销成本 | 移回文件 + 重建上下文 | 删 children + parent body 还原(手工) |
| 演化路径 | 子树重组痛苦 | parent 只增不减,children 是 parent 拆出的快照 |
| navigate | 浏览目录树 | 任意节点入手沿出边漫游;parent 节点是天然中心 |
| retrieve 精度 | 路径反映主题但与 link 无关 | 节点中心性 + 内容形态共同决定权重 |
#### 1.5.7 多级结构自然涌现
split 是节点的**局部操作**(只看一个过载节点),多级深度自然涌现:
```
digest/auth/
├── auth-fundamentals.md ← 早期写下,~1500 token,稳定
├── jwt-rotation.md ← 第一次 split:body 从 3500 token 重写为 overview
├── jwt-key-derivation.md ← 第一次 split 的 child
├── jwt-key-derivation-hkdf.md ← 二次 split:jwt-key-derivation 累积 update 后过载,再拆
├── jwt-key-derivation-pbkdf2.md ← 二次 split 的 child
├── ...
```
**物理仍是浅桶(1 层),"层级"由 split 链 + 节点中心性自然承载**。每一级的过载条件、决策机制、执行步骤完全相同 —— 没有"二级 split"特殊逻辑,只有"过载节点的 body 可以被 split 进一步拆"。
#### 1.5.8 retrieve 时节点怎么参与
| query | 期望返回 |
|---|---|
| "JWT 怎么轮换" | 优先 `jwt-rotation`(具体 overview)+ children(如 `jwt-key-derivation`) |
| "auth 体系" | 优先中心性高的节点(`auth-fundamentals` / `jwt-rotation` 等被多次 update / 是 split parent 的节点) |
| "auth 有什么子主题" | 沿高中心性节点的入/出邻居遍历;parent 节点优先返回 |
| "vault 里都有什么" | 各 bucket 中心性最高的节点(自然形成 vault 总览) |
**加权策略**(opinionated default,消费层可改):
- 节点权重 = base(=1.0) × intent 调节 × 中心性增益
- query 含"概览 / 主题 / 入门 / 全景"等**元意图**时,中心性高的节点加权(intent 调节 > 1)
- 中心性低 / body 短的具体节点权重稳定(默认 1.0,不被压低)
**topological traverse**:
- 沿 wikilink 自由走(不区分边类型 / predicate)
- 经过中心性高的节点默认**不强行展开**(否则一次 traverse 把整族 children 拉进来);agent 可显式深入
**中心性的天然来源 = split parent**:被拆过的节点是 parent,自然有 children 反向链接它,中心性自然高 —— 不需要单独维护 `kind: hub` 标记。
---
## 2. 完整能力集
### 2.0 设计目标
**让图的形状持续匹配实际知识的语义结构,在最小变更面 + 渐进演化的前提下,使任意尺度的知识访问都能命中合适粒度的节点。**
这个目标直接来自结构本身的设计意图 —— 浅桶 + 单一节点类型 + 单一边类型 + create_or_update + split 的组合,每一项都是为它服务。能力集的入选标准:**对至少一个验证维度有贡献**。
| 维度 | 含义 | 失败示例 |
|---|---|---|
| **形状匹配** | 节点中心性 / 边连接 / 节点邻域反映知识间的实际语义关系 | 一个节点 token 5000+ 长期不拆;同主题节点彼此 0 链接;同概念被建成多个独立节点 |
| **最小变更面** | 不重写其它节点正文,不大规模移文件,不破坏现有 wikilink | 任何"全图重组"或"批量改其它节点正文"的方案 |
| **任意尺度访问** | 具体方法节点 / overview 节点 / 节点邻居遍历都能命中 | 全 flat,主题级 query 命中不到东西 |
**显式排除**(不在目标内,避免能力集内卷):
- ❌ "完美归簇" —— F-5 留白,不确定就不动
- ❌ "实时一致" —— 异步 / eventual,节点写完不必立刻 split
- ❌ "零冲突 / 零违反" —— invariants 检测 + 事后修复,不追求永不发生
- ❌ 替消费层做检索 / 决策 —— digest 自治边界止于"维持图的形状"
- ❌ 跨节点重组(merge / re-edge / unify / dissolve)—— 简化模型不做这些;同概念二次进入由 G\* update 路径处理
**演化只有两件事**:G\* create_or_update(入流型,新材料融入)+ M split(后台,过载就拆)。detection 派生信号驱动这套循环。
### 2.1 生成侧:digester(入流型)
| # | 能力 | 服务 | 性质 | 何时发生 |
|---|---|---|---|---|
| **G1** | **scope 决定**:选哪组 daily/resource 进入本轮蒸馏 | 形状匹配(决定形状从哪生长) | LLM | digester 启动 |
| **G2** | **原子单元抽取**:从 scope 中识别值得沉淀的原子单元(N 个候选) | 形状匹配 + 任意尺度 | LLM | 核心环节 |
| **G\*** | **create_or_update**:对每个候选,**多路召回(SearchStep:vector + keyword + 邻接展开,RRF 融合,scope 限 `digest/`)** → LLM 看完整候选池 → 终判 create / update / drop;create 路径走 G3/G4/G6/G7 + 写 body;update 路径融入已有 body(语义守恒重写)+ 自然追加 provenance(详 §3.10) | 形状匹配(去重内置)+ 最小变更面 | LLM(决策)+ 机械(召回 + 守恒写入) | 每个候选 |
| **G3** | **路径命名**(create 路径):在 G4 选定 bucket 内,文件名同 bucket 唯一(fs 层断言);风格与同主题节点一致 | 任意尺度(可寻址) | LLM(命名)+ 机械(同 bucket 文件名冲突 → 拒写) | create 时 |
| **G4** | **bucket 落地**(create 路径):从固定集合中挑选;找不到合适专属桶 → 落 `general/`(§3.7) | 形状匹配(物理归档) | LLM(读 bucket 列表) | create 时 |
| **G6** | **provenance 写入**(create 与 update):新节点 body 内联反指上游 daily/resource 的 wikilink;update 时 LLM 在融入新材料时自然追加新 provenance 链,旧 provenance 链由 E-1 守恒校验保住(§3.9) | 任意尺度(跨层访问) | LLM(prompt 引导写出 `[[daily/...]]` / `[[resource/...]]`)+ 机械(outbound diff 校验) | 写入时 |
| **G7** | **横向 link**(create 时):新节点链到相关的已有节点(出边);update 时也可加新 link | 形状匹配 + 任意尺度 | LLM | 写入时 |
**关键边界**:
- **G\* 是入流唯一改 body 的操作**,且**只改 subject node** —— update 改的是同概念那个节点自己,不改其它节点
- **G\* update 必须语义守恒**:LLM 重写 body 时只能"融入"新内容,不能删除已有信息(只增不删 / 不改原意)
- **0 出边节点合法**(G7 没识别到合适邻居),后续 G\* 进入时其它节点可以反向链回来 —— 不强求 LLM 一次性给全
- **G\* 漏判去重**(把同概念建成新节点)→ 不主动兜底,接受重复;若 vault 累积明显的同概念重复,可由 auto-link 离线 audit 工具产报告(详 `auto_link_design.md` §1.3 L4)
- **digester 不做 split** —— split 是后台 M 操作
### 2.2 组织侧 / 检测 / 写入并发 → `auto_maintain_design.md`
M split / D 检测信号(D1 / D3 / D10)/ 阈值校准 / D3 写后触发模型 / G\* / split / auto-link L1 三方共用的 CAS 写入协议 / split provenance / 时序 / 后门 —— 全部归 `auto_maintain_design.md`。
dream 保留**模型层**(§1 节点 + 边 + 守恒规则)+ **生成侧**(§2.1 G\*)+ **召回**(§3.10 SearchStep);maintain 负责**组织 / 运行时**(split + D + CAS + 时序)。两份文档共享 §1.5 节点 + 边模型、§1.5.5 边守恒、§1.5.4 F-invariants。
| 在 maintain 文档中 | 内容 |
|---|---|
| §1 | M split 能力卡 + 关键边界 |
| §2 | 检测信号 D1 / D3 / D10 |
| §3 | 阈值校准 |
| §4 | D3 写后触发模型 |
| §5 | CAS 写入协议(三方共用) |
| §6 | split 时 provenance |
| §7 | G\* / split / auto-link L1 时序 |
| §8 | 后门(暂缓) |
### 2.4 边界协议(谁不能做什么)
| 边界 | 内容 | 来源 |
|---|---|---|
| digester ∩ maintainer | digester 不做 split;maintainer 不做原子单元抽取 / 新具体节点 create | 入流 vs 自维护职责分离 |
| digester → 其它节点 | G\* update 改 subject node body,**不改其它任何节点正文** | F-2 |
| digester → "摘要 / overview" | digester 不为做 overview 而创建节点;它产的节点都是具体原子单元;overview 是后续 split 的副产品 | F-3 |
| maintainer → 其它节点 | M split 改 parent body(重写为 overview)+ 创建 N 个 children body;**不改任何其它节点** | F-2 |
| maintainer → inbound 链 | split 时**全部不动** —— digest 不引入 anchor,inbound 一律是裸链 `[[<parent-path>.md]]`,parent 路径未变 | F-10 / E-3 |
| digester → 边守恒 | G\* update 写新 body 前,机械对比 old/new outbound:`new ⊇ old`((target, predicate) 二元组);失败 → LLM 重试一次,再失败拒写 | F-11 / E-1 |
| maintainer → 边守恒 | split 写新 parent body + N children body 前,机械对比:`(parent_new ∪ ∪children_outbound) ⊇ parent_old`;失败 → LLM 重试或拒写 | F-11 / E-2 |
| 全员 → typed link predicate | wikilink 的 predicate 是 edge identity 的一部分;G\* update / split 不能丢 predicate(`is_a:: [[A]]` 必须保持;否则被守恒校验当作 drop edge + add edge 拦下);predicate 升 / 降级走显式 audit 路径 | F-11 / §1.4 |
| 全员 → resource/daily | 都不能改 | I-2 / I-3 |
| 全员 → 节点 rename | rename = 一次 `wikilink_handler.retarget_links(old_path, new_path)`;无 alias 表,无透明展开 | §1.3 |
| 全员 → provenance link | 永远必须可达(I 不变量 + D10 检测) | I-1 / I-4 |
| 全员 → kind 字段 | reme 核心**透明**:不读取 frontmatter `kind` 做结构决策;`kind` 是消费层 schema 提示 | [[reme4_schema_layering]] |
| 全员 → predicate 谓词 | reme 核心**结构决策不读**:G7 / 中心性都聚合所有 predicate 算;edge 唯一性 / 反向索引会用到 predicate(防同源同标不同 predicate 互相覆盖);未类型化 link 是默认形态 | [[reme4_schema_layering]] / §1.4 |
---
## 3. 待对齐边界点(后续讨论清单)
### 3.1 G\* update 的语义守恒边界(已收敛)
**决策**:**LLM 重写整段**(prompt 强约束"语义守恒,只增不删 / 不改原意;冲突标注 `> 注:不同来源记载...`,不擅自仲裁")+ **机械守恒校验**(详 §1.5.5 E-1)。校验失败 LLM 重试一次,再失败拒写 + audit。
首版可先用 append 起步(出边集合天然 ⊇,守恒校验自动通过),prompt 工程量小;成熟后切到重写。
### 3.2 maintainer 的人 / agent 后门 → `auto_maintain_design.md` §8
### 3.3 G\* 与 split 的时序 → `auto_maintain_design.md` §7
### 3.4 节点 kind / 边 predicate(已收敛)
reme 核心**只有节点 + 边两种结构类型**:
- frontmatter `kind` 字段(若存在)= 消费层的**节点内容标签**(concept / method / case / entity / topic / ...),reme 不读它做结构决策
- 边 `predicate`(Dataview 风格,若存在)= 消费层的**边关系标签**(`is_a` / `extends` / `causes` / ...),reme 解析 / 存储 / 参与 edge 唯一性,但**结构决策不读**(G7 不区分 predicate;中心性不区分)
- "overview 节点"角色靠图位置(高中心性 / 是 split parent)+ body 内容形态识别,不靠 frontmatter 或 predicate 标记
- 未类型化 wikilink 是默认 / 常态形态
详见 §1.4 / §1.5.1 / §2.4。
### 3.5 retrieve 时的权重策略(部分收敛 → §1.5.8)
- 加权策略:节点权重 = base(=1.0) × intent 调节 × 中心性增益;query 含元意图("概览 / 主题 / 入门 / 全景"等)时,中心性高的节点加权
- traverse 默认不强行展开高中心性节点(防止整族 children 拉进来);agent 可显式深入
- 不按 frontmatter `kind` 加权;中心性天然来源 = split parent(详 §1.5.8)
剩余待定:**中心性算法选型**(eigenvector / PageRank / 简单入度,初期可用入度,后续校准)。
### 3.6 M split 时的 provenance 处理 → `auto_maintain_design.md` §6
### 3.7 bucket 集合管理
§1.2 已定:bucket 集合**固定预定义**,不由 digester / maintainer 动态生成。补足细节:
- **定义位置**:`vault.yaml` 顶层 `digest.buckets:` 是源 + 自动生成 `digest/_buckets.md` 作为人/LLM 可读视图;digester G4 时读后者作为 prompt context
- **初始化**:opinionated default(通用桶 `concept` / `method` / `pattern` / `tool` / `domain` 等 + 必带 `general`);消费层可改桶名,但 **`general` 不可删**(否则 G4 失去兜底)
- **扩展路径**:reme 不主动提议扩 bucket(对比旧设计的 maintainer 周期建议已 DROPPED);用户编辑 `vault.yaml` 后下次 G4 即生效
**未归类节点处理**(G4 找不到合适专属桶时):**统一落入 `digest/general/`**。
| 维度 | 内容 |
|---|---|
| **bucket 名** | `general`(固定集合一等公民,默认包含) |
| **语义** | "通用主题 / 暂无专属归属" —— 合法常态,非故障状态 |
| **路径** | `digest/general/<slug>.md`,与其它 bucket 完全等同 |
| **节点演化** | 与其它 bucket 一致 |
| **错桶后续** | 不主动跨桶 move(无 D9 / M-D);若严重,人工 mv + `retarget_links(old, new)` |
**为什么是 `general` 而不是 `_unclassified`**:`_unclassified` 暗示待处理状态,LLM/人都想清理掉;`general` 是合法常态,G4 选桶时是显式合法选项而非 fallback 故障路径。
**已排除**:拒绝写入(候选丢失)/ 强行选最近似专属桶(本体污染,general 反而更安全)。
### 3.8 检测阈值校准 → `auto_maintain_design.md` §3
### 3.9 provenance 载体形态(已收敛)
**决策**:**provenance wikilink 嵌在节点 body 正文中**(inline body prose),由 LLM 在 G\* / split prompt 里自然写出,跟其它 body wikilink 完全同形,**靠语义维护**。reme 核心没有 provenance 专用算子。
**写出形态**:
- 行文中自然带出处:"... 该模式最早出现在 [[daily/2026/05/15.md]] 的实践中"
- 或专门一段总结式段落,内含若干 wikilink 指向上游
- 可选 predicate(`derived_from:: [[daily/2026/05/15.md]]`),不强制
**机械保护**:
- E-1 守恒(G\* update):旧 provenance 不在新 outbound 集合 → 重试或拒写,机械兜底
- E-2 守恒(split):provenance 跟着对应内容段自然分配到 parent overview / children,合计守恒
- D10 检测:provenance 断裂 = D1 断链子集(target 命中 `daily/` / `resource/` 前缀);D10 严重程度高于普通 D1(I 不变量)
**E-5 / G6 等"provenance 专用机制"全部坍缩** —— 不再单列。Prompt 必须要求"出处用 `[[...]]` 形式表达"(纯散文会被守恒校验视为丢边)。
### 3.10 G\* 语义查后端(已收敛)
**决策**:**直接复用 `SearchStep`(`reme4/steps/index/search.py`)** —— 多路召回并发(vector + keyword)+ RRF 融合 + file_graph 邻接展开,把**完整候选池交给 LLM 终判**;G\* 入口不做 bucket 粗筛(LLM 拥有完整跨桶视野,可识别"概念错分到 general"或"跨桶同概念";三路信号 RRF 融合后噪声可控)。
**召回链路**:
1. 候选原子单元(摘要 / 关键词)→ `SearchStep`(`search_filter={"path_prefix": "digest/"}`,I-2/I-3 daily/resource 不入池)
2. `SearchStep` 内部:`vector_search` + `keyword_search` 并发 → RRF 融合 → `expand_links` 邻接展开 → 返回 top-`limit` FileChunks(含 path / 行号 / 邻接节点)
3. 候选池整体喂 LLM,按 path 自然聚合(同节点多 chunk 命中 = 强信号);终判输出节点路径
4. LLM 终判 create / update / drop;update 选定 subject node → 走 E-1 守恒重写
**provenance 不依赖召回** —— G\* / split prompt 让 LLM 直接写 `[[daily/...]]` / `[[resource/...]]`(§3.9)。
**索引维护**:沿用 `update_index` step,G\* / split 写 body 后调一次刷该节点索引;启动一次全建(`clear_and_scan` 已就绪),损坏走全建兜底。
**默认参数**(可按 dogfooding 调):`limit` 5~10 / `vector_weight` 0.7 / `expand_links` on / `min_score` 0(初版不过滤,LLM 兜底)。
### 3.11 G\* / split 写入并发 / 原子性 → `auto_maintain_design.md` §5
### 3.12 D3 触发模型 → `auto_maintain_design.md` §4
### 3.13 anchor 不引入 wikilink 设计(已收敛)
**决策**:**digest 设计层不使用 `[[path.md#section]]` 形态** —— wikilink 只有 `[[path.md]]`(可选 alias / 谓词),anchor 不进入 digest。当 LLM 想"指向某个具体子主题"时,正确做法是让那个子主题升级为独立节点(必要时通过 split),而不是在过载 parent 内部用 anchor 凑合。
**连锁简化**:
- E-4(inbound anchor 机械 retarget)整类**消失**;split 流程末尾不再扫 inbound anchor 子流程;`{anchor → child}` 映射输出从 split prompt 中移除
- 边唯一性键从三元组 `(target, predicate, anchor)` 简化为二元组 `(target, predicate)`
- §1.5.5 边迁移类别从 4 类(E-1..E-4)简化为 3 类(E-1..E-3)
- `FileLink.target_anchor` 字段在 schema 中保留(供其它消费层),digest 层永远写 `None`
**Prompt 约束**:G\* / split 的 prompt 必须明确告知 LLM 写 wikilink 时不带 `#section`。若 LLM 仍写出 `[[path.md#section]]`,wikilink_handler 仍能解析,守恒校验只看 `(target, predicate)`,不会形成"丢边"风险 —— 但 anchor 在 digest 层无语义。若引用方依赖某 anchor 锚定具体段落,表明该内容应升级为 child 节点。
---
## 4. 下一步
本文档覆盖 dream 模型 + 生成侧 + 召回(G\* / 节点+边模型 / SearchStep)。组织端实现清单(M split / D 检测 / CAS)见 `auto_maintain_design.md` §10。
1. **digester 流程图**(G1 / G2 / G\* 的实际编排;G\* 内 create / update 路径分流;**召回直接复用 `SearchStep`**(vector + keyword + 邻接展开,RRF 融合,scope `digest/`)—— 详 §3.10;**G\* update 写入前 outbound diff 守恒校验** — E-1)
2. **rename 路径设计**(`wikilink_handler.retarget_links(old_path, new_path)` 已就绪;封装为单步 step 入口,无 alias 表 / 无透明展开)
3. **bucket 集合配置**(`vault.yaml` schema / 默认桶模板 / `general` 兜底机制 / `_buckets.md` 视图生成)
4. **边守恒校验工具**(`extract_links` 已就绪;新增 outbound diff 比较器 + LLM 重试编排 + ConservationViolation audit 事件)
5. **provenance prompt 规范**(G\* / split 引导 LLM 写 `[[daily/...]]` / `[[resource/...]]` —— §3.9)
实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_memory_design.md` / `auto_maintain_design.md` / `auto_link_design.md` 共同作为契约依据。

209
docs4/auto_link_design.md Normal file
View file

@ -0,0 +1,209 @@
# auto-link 设计(背景实体识别 + wikilink 写回)
> 本文档记录 reme4 中 **auto-link** 的设计讨论 —— 在已写入节点之间发现隐含关系,把这些关系作为 `[[...]]` wikilink **写回 body**,形成可见、可编辑的图结构增强。
>
> 配套阅读:
> - `structure.md` §1.2(三层数据视角)/ §4(retrieve 三种问法)
> - `auto_memory_design.md`:auto-link 可反向扫 daily event,补实体 wikilink(daily → digest)
> - `auto_dream_design.md`:wikilink 模型(§1.4 边语法 / §1.5 演化 / §1.5.5 边守恒 E-1 / E-2 / E-3);auto-link 借这套基础设施
> - `auto_maintain_design.md`:CAS 写入协议(§5);auto-link L1 写回与 dream G\* / maintain split 三方共用同一套 CAS
>
> **三层对应**:reme 服务整体三层 —— auto-memory / auto-dream / **auto-link(本文档)**。auto-link 是图关系的**后置增强** —— 在已落地的 vault 上做实体识别 + wikilink 写回,补足 content link(写记忆时由 LLM 直接产生的 `[[...]]`)在长 tail 隐含关系上的盲区。
>
> **核心立场**:auto-link **写回 body**,不只是产报告。生成的 wikilink 是**可见、可编辑**的(写在 Markdown 文件里),agent / 人可后续 curate。auto-link 不引入新材料,纯 additive 插入 wikilink,天然满足 E-1 守恒;复用 dream 的 CAS 写入协议,不引入新基础设施。
---
## 0. 问题陈述
content link(`auto_dream_design.md` G\* / split 写入时由 LLM inline 产生的 `[[...]]`)解决了"写记忆时显式的关系"。但有一类关系不会在 inline 写入时自然涌现,需要后台扫描已写入的 vault 才能识别:
1. **历史 body 的实体未链接** —— G\* update 时 LLM 关注新材料融入,可能忽略已有 body 中某个未链接的实体(例如 body 提到 "JWT" 但没写 `[[digest/auth/jwt-overview.md]]`)
2. **跨节点 / 跨桶的隐含关联** —— 节点 A 提到 "rate limit",但 `digest/api/rate-limit.md` 是后来才被 split 创建 → A 写入时没机会建立这条边
3. **同主题未连 / 同概念重复** —— G\* 漏判去重把同概念建成两个节点;或两个主题相关但 0 链接的节点彼此不知晓
auto-link 承担这部分:**后台扫描已写入节点 → 实体识别 / 候选挖掘 → wikilink 写回 body**。
---
## 1. 已对齐决策
### 1.1 与 content link 的边界
| 维度 | content link(在 dream) | auto-link(本文档) |
|---|---|---|
| 何时产生 | 写记忆 inline:G\* update / M split prompt | 后台扫描:离线 / 周期 / 触发后异步 |
| 由谁产生 | LLM 在 dream 写入流中顺手写出 | LLM 在 auto-link 扫描流中识别后写出 |
| 输入 | 新材料 + 召回候选节点 | 已写入 body + 全 vault 索引 |
| 改 body | 是(重写整段 body) | 是(纯 additive 插入 wikilink,不改文字) |
| 守恒 | E-1 强守恒(out ⊇ old) | E-1 天然满足(纯增) |
| 用途 | 写入即关系明示 | 弥补 inline 漏判,挖掘长 tail 关系 |
### 1.2 写回模型:纯 additive,复用 dream CAS
auto-link 写回是**纯 additive** 操作 —— 在已有 body 文字中找到实体 mention,替换为 wikilink 形态:
```
Before: "JWT 轮换的核心是密钥派生 ..."
After: "[[digest/auth/jwt-rotation.md|JWT 轮换]]的核心是[[digest/auth/jwt-key-derivation.md|密钥派生]] ..."
```
| 维度 | 决策 |
|---|---|
| **alias 必须保留原文** | `[[path.md\|<原文>]]` 形态;原文一字不改 —— 守住"不改写其它节点正文" (`auto_dream_design.md` §1.5.4 F-2) 的精神 |
| **predicate 默认为空** | auto-link 默认产生无谓词 wikilink;升 typed link 走 L3(详 §1.3) |
| **不引入 anchor** | 与 dream 一致(`auto_dream_design.md` §1.4 / §3.13);target 永远是节点路径 |
| **CAS 写入** | 完全复用 `auto_maintain_design.md` §5 的 read-stamp + CAS-write 协议(冲突重做 ≤ 3 次) |
| **E-1 守恒** | 纯 additive:`new outbound = old outbound ∪ {new wikilinks}`;`new ⊇ old` 天然满足,守恒校验默认通过 |
| **rollback** | 若 auto-link 误插入(例如 entity mention 是同名歧义),走标准 edit 或 retarget 撤销;auto-link 不维护"我插过哪些"audit log(留给 SDK 决定) |
**为什么是 additive 而不是重写**:
- additive = 0 文字风险(原文不变,只在原 mention 周围加 `[[ | ]]` 包装)
- 重写 = 触发完整 E-1 守恒校验 + LLM 重写整段语义守恒 prompt + 多次 LLM 调用 = 跟 G\* update 重复
- additive 失败可见:产生坏 wikilink 时,人/agent 直接编辑 body 修就行
### 1.3 候选挖掘类型(L1-L4)
| # | 类型 | 描述 | 写回形态 |
|---|---|---|---|
| **L1** | **实体识别**(主路径) | 扫 body,识别已是 digest 节点的实体名(模糊匹配 + 语义召回);未被 wikilink 化的 mention → 加 `[[path.md\|<mention>]]` | additive wikilink 插入 |
| **L2** | **同主题未连**(旧 D7) | 两个 digest 节点谈相关主题但 0 wikilink → 候选 add link;LLM 判后在 body 末尾追加一句引用 | additive(在合适位置 / 节末追加 `参见 [[other.md\|other]]`)|
| **L3** | **隐含 predicate 推导** | 已有 `[[A]]` 但 LLM 可推断关系类型(`is_a` / `causes` / `extends` / ...)→ 升级为 typed link | 改 `[[A]]` → `is_a:: [[A]]`(predicate 升降级走显式 audit,详 §2.1)|
| **L4** | **重复语义检测**(旧 D8) | 两个节点描述同一概念但被独立 create(G\* 漏判去重)→ 候选 merge | **不写回**;产报告 + 提示人/agent 触发 G\* update 路径手工合并 |
**L1 是主路径** —— 它是 auto-link 最核心、最频繁、最高 ROI 的操作:每个 digest 节点写完后,后台扫一遍 body,找未链接的已知实体,additive 加 wikilink。
**L2-L3 是辅助** —— 周期扫,产候选,LLM 终判,写回部分(L2 节末追加 / L3 升 predicate)。
**L4 不写回** —— 节点合并是结构改动,影响 E-1 守恒边界 + inbound 链路 + provenance 链路,不适合自动写;auto-link 只产报告,人/agent 决定走 G\* update 路径解决。
### 1.4 触发节奏
| 模式 | 何时 | 适用 |
|---|---|---|
| **inline post-write**(默认) | 每次 G\* update / M split 写完 body → enqueue auto-link L1 job(异步,FIFO,CAS 保护)| L1 实体识别;反应即时,与 D3 写后检测同节奏 |
| **周期 batch**(可选)| cron(daily / weekly)扫全 vault | L2 / L3 候选挖掘;成本可控 |
| **手动触发** | SDK / 人显式调用 | 全量重扫 / 修复 |
**L1 inline 的必要性**:新 split 出的 child 节点立即被既有 body 引用(用 wikilink 而非纯 mention)的关键 = 写入即扫描;不 inline 会让"刚创建的 child 节点"在很长时间内只有 split parent 一个 inbound,中心性失真。
**已排除**:
- inline 时同步 auto-link(阻塞 G\* return)—— 时延不可接受;auto-link 始终异步
- 所有 L\* 都 inline —— L2-L3 候选挖掘 RTL 跨节点,成本高,只适合 batch
- 全 cron 唯一触发 —— L1 滞后过久,新节点孤岛
### 1.5 中心性算法(retrieve 加权依赖)
retrieve 时节点权重 = base × intent 调节 × **中心性增益**(详 `auto_dream_design.md` §1.5.8)。中心性需要 auto-link 这一层提供 —— content link 给底子,auto-link 补 long tail,二者合起来才是完整的图。
| 选项 | 优点 | 缺点 |
|---|---|---|
| **简单入度** | 实现最简;split parent 入度天然高;auto-link L1 加边后入度即时反映 | 不区分"权威节点"vs"被随手提的节点";高入度 ≠ 高权威 |
| **PageRank** | 经典;权威性传递 | 实现复杂 + 增量更新成本(每次写边重算成本高,需 incremental algorithm)|
| **eigenvector centrality** | 与 PageRank 相近 | 同上 |
**首版决策**:**简单入度**(file_graph 已有 inbound 链表,O(1) 查);auto-link L1 加边后入度立刻更新,split parent 自然涌现高入度。dogfooding 后视 retrieve 质量演进。
中心性是 retrieve 时**查询时计算**,不预存:
- file_graph 已建反向索引(inbound),计算 `len(inbound(node))` 是 O(1)
- 不预存避免"加边后中心性陈旧"问题
- PageRank 演进时可加增量计算 + 周期 refresh
---
## 2. 待对齐边界点
### 2.1 L3 predicate 升降级的 audit
L3 把 `[[A]]` 升级为 `is_a:: [[A]]` 时,**改了 edge identity** —— `(target, None)` 变成 `(target, "is_a")`,在 E-1 守恒视角下 = 删一条边 + 加一条边:
```
old outbound: {(A, None)}
new outbound: {(A, "is_a")}
diff: missing = {(A, None)}; added = {(A, "is_a")}
```
不打 audit 走默认会被守恒校验拦下(`missing != ∅` → 重试 / 拒写)。
**决策方向**:
- L3 写入必须打 audit flag(消费层意图:升级 predicate,允许 drop + add 同时发生)
- audit flag 由 reme4 step 暴露(`maintainer_step(action="predicate_upgrade", from=..., to=...)`),不放在普通 write 路径
- 普通 G\* / auto-link L1 写入永远不带 audit flag,守恒校验照常严格
详细 audit flag 接口形态留到 SDK 阶段。
### 2.2 多歧义实体识别
L1 扫 body 找 "JWT" 这个 mention,vault 中有 `digest/auth/jwt-overview.md` 和 `digest/payment/jwt-payment-flow.md` 两个 candidate:
候选方案:
- LLM 上下文判 —— 把 body 周围段落给 LLM,选最相关 target
- 跳过模糊 case —— L1 只处理 unambiguous mention,歧义 case 留人/agent
- 全部链 —— `[[overview]][[payment-flow]]`,后续人 curate
**首版**:LLM 上下文判(每个候选 candidate 提供 description / 周围若干节点 summary,LLM 选择 top-1 或 drop);成本可接受(扫描已是离线 batch)。
### 2.3 auto-link 写回与 G\* / split 的并发
auto-link 写 body 走 §1.2 CAS,但有特殊情况:
- 同节点同时被 G\* update 与 auto-link L1 写入 → CAS 协议自动序列化 (`auto_maintain_design.md` §5):后到者重做
- auto-link L1 写完后立刻被 G\* update 覆盖(G\* 重写 body) → 看 G\* prompt 是否守住 auto-link 加的 wikilink(E-1 强守恒 → 守住)
- auto-link L1 与 D3 派发的 split job 同节点并发 → split 先到 / 后到都不影响最终拓扑(split 把 body 拆成 parent + children,auto-link 加的 wikilink 跟着对应内容段自然分配到 parent / child)
**结论**:CAS + E-1 + E-2 守恒已覆盖所有并发场景,auto-link 不需要新协调机制。
### 2.4 跨 vault / 跨进程
M0 单 reme 实例 + 单 vault,auto-link 走内进程 enqueue;多实例 / 跨进程留 M1+(同 `auto_maintain_design.md` §5)。
### 2.5 实体识别 vs 现成 NER 库
L1 实体识别可选:
- LLM 直接扫(贵但灵活,与 digest 节点同构)
- 现成 NER 库(spaCy 等)预筛 + LLM 终判(快但 entity 类型与 digest 节点形态可能不匹配)
- 纯字符串匹配(已知节点名字 + 简单变体)+ LLM 终判 ambiguity
**倾向**:从纯字符串匹配 + LLM 终判 ambiguity 起步(实现最简,效果可能已经够好);视 dogfooding 决定是否引入 NER 库。
---
## 3. 与其它层的协作
| 上下游 | 关系 |
|---|---|
| ← **auto-dream** | dream 写完一个节点 → 通过 inline post-write enqueue auto-link L1(§1.4);auto-link 用 dream 的 CAS 协议 |
| ← **auto-memory** | auto-link 可反向扫 daily event,把实体识别成 `[[digest/...]]`(daily → digest);auto-memory 写入端不主动调 auto-link,触发同 dream 路径 |
| → **digest body** | 主要写入对象 —— L1 additive 加 wikilink / L2 节末追加引用 / L3 升 predicate(走 audit) |
| → **daily body** | auto-link 扫 daily event 时同样可加 `[[digest/...]]`(I-2 daily 单作者需协调:auto-link 应在 event 关闭后才动该 event,不与 active event 并发改;实现细节留 step 层处理) |
| → **resource body** | I-3 immutable;auto-link **不写 resource**(reading-only) |
| → **L4 候选 report** | L4 重复语义检测产报告,落 `audit/<date>/auto_link_l4.md`(具体路径 / 形态留 step 层) |
---
## 4. 与 auto-dream 模型的引用关系
本文档复用 dream 定义的底层模型,所有具体规则在 `auto_dream_design.md` 中:
| 引用 | 来源 |
|---|---|
| wikilink 基础语法(`[[path.md\|alias]]` / predicate) | `auto_dream_design.md` §1.4 |
| 节点 / 边模型 | `auto_dream_design.md` §1.5 / §1.5.1 |
| F-invariants(F-1..F-11)| `auto_dream_design.md` §1.5.4 |
| 边守恒 E-1 / E-2 / E-3 | `auto_dream_design.md` §1.5.5 |
| 路径即 ID / rename | `auto_dream_design.md` §1.3 |
| CAS 写入协议 | `auto_maintain_design.md` §5 |
| anchor 不引入 | `auto_dream_design.md` §1.4 / §3.13 |
| SearchStep 召回 | `auto_dream_design.md` §3.10 |
---
## 5. 下一步
1. **L1 实体识别 step 实现** —— 字符串匹配 + 语义召回 + LLM ambiguity 终判 + additive wikilink 写回(§1.2 / §1.3)
2. **inline post-write trigger 接入** —— G\* update / M split CAS 写入成功后 enqueue auto-link L1 job(§1.4)
3. **L2 / L3 周期 batch 框架** —— cron(daily / weekly)+ 候选挖掘 prompt + 写回路径(§1.3)
4. **L3 audit flag 接口** —— `maintainer_step` 提供 `predicate_upgrade` 操作,带 audit context 走特殊守恒规则(§2.1)
5. **中心性 retrieve 增益** —— file_graph inbound count → retrieve 加权乘子(§1.5)
6. **L4 报告框架** —— 重复语义检测产报告,提供 SDK / 人介入入口(§1.3 / §3)
实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` 共同作为契约依据。

View file

@ -0,0 +1,190 @@
# auto-maintain 设计(digest 组织端:M split / 检测 / 写入并发)
> 本文档记录 reme4 中 **auto-maintain** 的设计讨论 —— digest 层的组织 / 重组 / 写入运行时,含 M split、D 检测信号、写后触发模型、CAS 写入协议。
>
> 配套阅读:
> - `structure.md` §3.6(maintain 动作语义)/ §7.3(maintainer 模块)
> - `auto_dream_design.md`:节点 + 边模型(§1.1-1.5)/ F-invariants(§1.5.4)/ 边守恒 E-1/E-2/E-3(§1.5.5)/ G\* 操作(§2.1)—— maintain 复用这套底层模型
> - `auto_link_design.md`:auto-link 写回也走本文档的 CAS 协议(§5)
> - `auto_memory_design.md`:auto-memory 不直接复用 maintain,但事件级"拆"与节点级 split 在概念上同构(都把过载粒度切小)
>
> **三层框架的位置**:报告 §5 三层为 auto-memory / auto-dream / auto-link。maintain 严格按 `structure.md` §3.5-3.6 的 L4 action 分类是独立 action(`maintain: digest → digest`),不属 `digest` action(`digest: resource + daily → digest`)。本文档作为四方分工的**第四份**,专门覆盖 dream 写完之后 digest 的组织 / 重组 / 写入运行时。
>
> **核心立场**:
> - **maintain 与 dream 同 pace**(idle background)、同模型(节点 + 边 / 守恒规则),但**语义边界不同**:dream 是 compose(资料 → digest),maintain 是 reorganize(digest → digest)
> - **maintain 只做 split**,不做 merge / dissolve / re-edge / unify;过载就拆,其它跨节点重组留给消费层 / 人工
> - **CAS 写入协议是基础设施**,被 dream G\* / maintain split / auto-link L1 共用,统一编排在本文档(§5)
---
## 0. 问题陈述
dream 模型(`auto_dream_design.md` §1.5)规定 digest 的演化只做两件事:G\* create_or_update(入流型,新材料融入)+ M split(后台,过载就拆)。dream 文档负责 G\* 与节点 / 边模型;**本文档负责 M split 与运行时机制**(D 检测 / 触发模型 / 写入并发协议)。
| 输入 | 输出 |
|---|---|
| dream 写入后的 digest 状态 + 触发信号(D3 过载,inline) | parent overview 重写 + N 个新 children 文件;边守恒 E-2 通过 |
**设计目标**:
1. **形状匹配** —— 让节点粒度持续与实际语义结构对齐(过载节点拆;不过载不动)
2. **最小变更面** —— split 改 parent + 创建 N children,不动其它节点(F-2)
3. **不引入新基础设施** —— 复用 dream 的节点 + 边模型 / 守恒规则;CAS 写协议自洽
**显式排除**:
- ❌ merge / dissolve / re-edge / unify —— 跨节点重组不做(简化模型;同概念二次进入靠 G\* update)
- ❌ 改其它节点正文 —— split 只改 parent body(重写为 overview)+ 创建 children body
- ❌ 重建 inbound —— split 时 inbound 一律不动(F-10)
---
## 1. M split(maintainer 唯一 op)
| # | 能力 | 服务 | 触发 | graph | file | body |
|---|---|---|---|---|---|---|
| **M split** | 节点过载 → LLM 拆成 parent overview + N 个 children;parent 文件原地保留,children 是新文件;children 加 `[[parent]]` 反向链接;inbound 边不动 | 形状匹配(粒度对齐)+ 任意尺度(涌现层级) | D3 过载 | parent 0 拓扑改;新 children 节点 + 各自加 `[[parent]]` 出边 | 创建 N 个 children 文件;parent 文件原地 | parent body 重写为 overview;children 各自有新 body |
**关键边界**:
- **M split 改两类 body**:parent body(重写为 overview)+ N 个新 children body;不改任何**其它**节点(`auto_dream_design.md` §1.5.4 F-2)
- **inbound 不重定向** —— 外部对 parent 的 wikilink 全部保留指 parent;后续 G\* 进入时若 LLM 觉得 child 粒度更合适,直接加新边到 child 即可(F-10)
- **没有 dissolve 操作** —— children 长期空也不主动删;消费层 / 人工显式介入
- **没有 merge / re-edge / unify** —— 跨节点重组不做;同概念二次进入靠 G\* update;错桶节点不主动 move(若严重,人工介入)
- **边守恒** —— split 写新 parent body + N children body 前,机械对比 outbound:`(parent_new ∪ ∪children_outbound) ⊇ parent_old`;失败 → LLM 重试或拒写(F-11 / E-2,详 `auto_dream_design.md` §1.5.5)
---
## 2. 检测信号 D1 / D3 / D10
| # | 信号 | 服务 | 服务能力 |
|---|---|---|---|
| **D1** | 断链(wikilink → 不存在的 path) | 任意尺度(可达性) | 告警 / 简单修复(就地删 wikilink 或保留 alias 文本) |
| **D3** | 过载节点(token 阈值 → LLM 判离散度) | 形状匹配(粒度) | maintainer(M split) |
| **D10** | provenance 断裂(digest 节点反指的 daily/resource 不可达) | 任意尺度(跨层不变量) | 严重告警(I 不变量违反) |
**触发模型**:**写后立即** —— G\* / split 写完 body inline 检测;无后台 watcher / 无周期 tick / 无 dirty 队列(详 §4)。D1 / D10 是 wikilink 断链的子集,跟 file_graph 链路一起在写时检测。
> **简化模型砍掉的信号**:
> - **D2 隔离 / D4 过疏 / D5 高入度 / D5b 低入度摘要 / D6 slug 冲突 / D7 相似未链 / D8 重复语义 / D9 邻居异质** —— 全部 DROPPED
> - 旧 D5 高入度涌现 → 由 split 副产品(parent + children)等价覆盖;触发源换成节点过载(D3)
> - 旧 D6 slug 冲突 → 路径即 ID 后,同 bucket 内文件名冲突由文件系统层断言(写入即拒),不需要独立信号(详 `auto_dream_design.md` §1.3)
> - 旧 D7 / D8 → 简化模型不做 link / merge 提议;若 vault 累积明显的同概念重复,由 `auto_link_design.md` §1.3 L4 离线 audit 工具产报告
> - 旧 D9 邻居异质 → 简化模型不做跨桶 move;桶选择只在 G4 一次性决定,后续不重排
>
> **D3 过载的判据**:token 阈值机械检查 + LLM 判离散度;**写后立即 inline**。阈值见 §3,触发模型见 §4。
---
## 3. 检测阈值校准
简化模型只剩 D3(过载)是核心阈值,其它都是 invariant 触发(无可调阈值)或 informational(无 maintenance 联动)。
| 信号 | 阈值类型 | 默认 | 备注 |
|---|---|---|---|
| **D3 过载** | token + 主题离散度 | token 2000 / 离散度由 LLM 写后 inline 判 | **唯一驱动 split 的阈值**(详 §4) |
| **D1 断链** | 0 容忍 | 任意 1 条断链 → 告警 | 修复策略简单(就地删 wikilink) |
| **D10 provenance 断裂** | 0 容忍 | 任意 1 条断裂 → 严重告警 | I 不变量 |
D3 阈值作为 `vault.yaml` 配置项(opinionated default,reme 核心提供机制不写死阈值),消费层可改;dogfooding 后调优。token 阈值起点 2000(对应"约 5 个独立子主题"的常见过载点),首版可调。
---
## 4. D3 触发模型(已收敛)
**决策**:**写后立即检测,无 watcher 抽象,无 batch 窗口** —— 每次 G\* update / split 写 body 成功后,**inline** 在同一 job 内跑 D3:token 阈值 + LLM 离散度判定 → 必要时 enqueue split job(异步,走 §5 CAS 队列)。
```
G* / split 写 body 成功(CAS 通过)
└─ if len(body) > T:
└─ LLM 判离散度
└─ if is_overloaded:
└─ enqueue split job (FIFO, CAS-protected)
└─ return
```
**协议**:
- token 阈值默认 `2000`(§3 已定,`vault.yaml` 可配)
- 离散度 prompt 输出 `{is_overloaded: bool, suggested_clusters: [...]}`(若 overloaded 直接供 split job 吃,不重判)
- 启动无全扫(避免长启动);新写入立即检测覆盖增长路径;历史遗留过载随下次 update 自然检出
- 无 dirty 标 / 无 dirty 集合 / 无后台 worker —— D3 是写路径的合成函数
**为什么 inline**:反应即时(不等下一次 ingest);实现最简(无批处理窗口 / dirty 状态 / 独立 worker);LLM 判定成本可接受(大多写入 < T 不触发,触发后 split 切小后续不再越界);不引入 watcher = 少一层部署/监控。
**已排除**:定时 cron tick(静止 vault 浪费扫描)/ ingest-after batch(引入 dirty 集合)/ 独立 L2 watcher worker(多余部署层)。
**演进路径(M1+)**:若 inline LLM 阻塞 G\* 时延成问题 → D3 改为 fire-and-forget enqueue;若同节点重复触发 LLM 成本高 → 加节点级 body hash 缓存。
---
## 5. CAS 写入协议(共享基础设施)
**位置说明**:CAS 是 G\* update(`auto_dream_design.md` §2.1)、M split(本文档 §1)、auto-link L1 写回(`auto_link_design.md` §1.2)**三方共用**的写入协议。归在本文档是因为 maintain 是 digest 的"组织 / 运行时"端,运行时机制(检测 / 触发 / 写入)集中在一处方便对照。
**决策**:**并行决策 + 乐观冲突重做(CAS)** —— 所有 G\* / split / auto-link L1 决策并发跑,写入前用 body 版本戳(hash / mtime)做 CAS 比对;变了就丢弃 planned body 重做。无锁,无 ingest 级互斥。冲突率低 + E-1 / E-2 守恒校验顺手承担 race 兜底,无需新基础设施。
**协议(单个写入调用)**:
1. **读 + 记戳**:读 subject body → `version_stamp = sha256(body) | mtime`
2. **决策**:LLM 看候选池 → 决定 create / update / drop / split / additive-link;产 planned new_body
3. **CAS 写入**:重读 body 比 version_stamp
- **未变**:跑 E-1 / E-2 守恒校验 → 通过则 atomic write(write-temp + rename)→ done
- **已变**:丢弃 planned new_body,带最新 body 重走 step 1
4. **守恒校验失败**:走 `auto_dream_design.md` §1.5.5 既有重试路径(LLM 重试一次,二次失败拒写 + audit)
5. **重做次数上限**:CAS-冲突重做最多 3 次;超出 → 跳过候选 + audit log(避免活锁)
**create 路径 race**:两个 G\* 都决定 `create digest/auth/jwt-rotation.md` → atomic create(`O_CREAT | O_EXCL`)只让一个赢;输者拿 EEXIST → 重走 step 1(此时大概率改判 update)。
**适用范围**(全部走同一套 CAS):同 ingest 内 N 个候选并发 / 跨 ingest job 并发 / 后台 split 与前台 G\* 命中同节点(split 同样走 CAS)/ auto-link 写回(`auto_link_design.md` §1.2)。
**不解决的**:高冲突 workload(同概念被反复 ingest)→ 重做上限触发后 audit;跨进程并发(多 reme 实例同 vault)→ 不在 M0,需 fs lock(M1+)。
---
## 6. split 时的 provenance 处理(已收敛)
**坍缩到 E-2 合计守恒** —— provenance 是 body 内联 wikilink(`auto_dream_design.md` §3.9),split 时跟其它 body 边完全同形:LLM 把 parent body 拆成 parent overview + N children,provenance wikilink 跟着对应内容段自然分配;机械层 outbound 合计守恒校验保证 `(parent_new ∪ ∪children_outbound) ⊇ parent_old`,旧 provenance 不可能丢。无需专门的 provenance 分配逻辑或"全部复制到 child / parent 保留全量"等特殊策略 —— LLM 按"哪个 child 谈到了哪段上游就带走哪条 provenance"自然处理。
---
## 7. G\* / split / auto-link L1 时序(已收敛)
时序由 §4 / §5 与 `auto_dream_design.md` §3.10 共同规定,这里给最小汇总:
- **G\* 调用本身同步** —— material 进来就走 G\* 决策(召回 + LLM 终判)+ CAS 写入(§5)
- **D3 检测 inline** —— G\* / split 写完 body 顺手跑 token 阈值 + LLM 判离散度(§4),无 tick / batch / watcher
- **split 异步** —— D3 触发后 enqueue split job 进 §5 CAS 队列,跟其它 ingest / split job FIFO 共享,异步消费;**不阻塞 G\* return**
- **auto-link L1 异步** —— 写入成功后 enqueue auto-link L1 job(`auto_link_design.md` §1.4),与 split job 同 CAS 队列、FIFO 共享;不阻塞 G\* return
检测延迟 ≈ 0(inline);split 执行延迟 ≈ 队列等待时间(typically 数秒~数十秒);新建 / update 节点不必等 split 完成,体验连续。
---
## 8. maintainer 的人 / agent 后门(暂缓 — 非底层)
消费层 / SDK 接口问题,不影响底层机制。底层只需保证 split / rename / delete 等 op 走同一套 §5 CAS + 守恒校验链路:F-3 仍成立(maintainer 自动路径只做 split);merge / dissolve / re-edge 在底层**不存在**(无对应机械算子)。后门接口形态推迟到 SDK 阶段再定。
---
## 9. 与 dream 模型的引用关系
本文档复用 dream 定义的底层模型,所有具体规则在 `auto_dream_design.md` 中:
| 引用 | 来源 |
|---|---|
| wikilink 基础语法(`[[path.md\|alias]]` / predicate) | `auto_dream_design.md` §1.4 |
| 节点 / 边模型 | `auto_dream_design.md` §1.5 / §1.5.1 |
| F-invariants(F-1..F-11) | `auto_dream_design.md` §1.5.4 |
| 边守恒 E-1 / E-2 / E-3 | `auto_dream_design.md` §1.5.5 |
| 路径即 ID / rename | `auto_dream_design.md` §1.3 |
| anchor 不引入 | `auto_dream_design.md` §1.4 / §3.13 |
| provenance 载体形态 | `auto_dream_design.md` §3.9 |
| G\* 行为 | `auto_dream_design.md` §2.1 |
---
## 10. 下一步
1. **M split step 实现** —— D3 触发 → 候选 → split prompt → 写入 + E-2 守恒(§1 / §4 / §5)
2. **D 检测信号实现清单**(D1 断链 / D3 写后 inline / D10 provenance 哪些已就绪 / 缺哪些)—— §2
3. **D3 阈值配置**(`vault.yaml` 中 D3 token / 离散度阈值)—— §3
4. **CAS 写入框架** —— per-path body version_stamp + CAS 写入 + EEXIST create race + 重做上限 + audit;对外暴露给 dream G\* / auto-link L1 复用 —— §5
5. **后门 SDK 接口形态**(暂缓 M1+)—— §8
实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` / `auto_link_design.md` 共同作为契约依据。

197
docs4/auto_memory_design.md Normal file
View file

@ -0,0 +1,197 @@
# auto-memory 设计(实时事件拆分 / 写入 daily)
> 本文档记录 reme4 中 **auto-memory** 的设计讨论 —— 把 agent 连续的对话 / 任务流切成离散的 daily 事件原子,inline 落到 `daily/` 层。
>
> 配套阅读:
> - `structure.md` §2.1-2.2(daily 层定位)/ §3.4(sync 动作语义)/ §7.1(synchronizer 模块)
> - `auto_dream_design.md`:auto-memory 产物如何被 dream 消化(G\* 读 daily 作为入流之一)
> - `auto_maintain_design.md`:digest 的组织端 / CAS 写入协议;auto-memory 不直接复用,但事件级"拆"与节点级 split 在概念上同构(都把过载粒度切小)
> - `auto_link_design.md`:auto-link 可反向扫 daily 事件,补充实体 wikilink(daily → digest)
>
> **四份分工**:reme 服务整体四份设计 —— **auto-memory(本文档)** / auto-dream / auto-maintain / auto-link。auto-memory 是入流端,把 agent 实时事件流切成 daily 事件原子;它的产物是 dream 消化的两路输入之一(另一路是 resource)。
>
> **核心立场**:auto-memory 是 `structure.md` §3.4 `sync` 动作的实现侧 —— 强调 **inline 实时**与**事件边界检测**。是不是改名 sync → auto-memory 留给上层文档对齐,本文档聚焦机制。
---
## 0. 问题陈述
agent 的对话与任务过程是连续事件流(用户回合、工具调用、上下文切换、中断恢复),但记忆系统需要离散的、可独立检索的事件单元。auto-memory 解决这个切分问题。
| 输入 | 输出 |
|---|---|
| agent 当前事件流(对话回合 / 工具调用 / 任务切换信号);可选 `notify` 候选作为 cue | `daily/<date>/<event-slug>/<note>.md` 事件原子;`daily/<date>.md` 主索引 |
**设计目标**:
1. **事件边界尽量与 agent 语义意图一致** —— 同一个意图(同一个任务 / 同一段思路)→ 同一个事件;意图切换 → 新事件
2. **inline 实时写入** —— 不滞后,不批处理;agent 一边工作,记忆一边落地
3. **保持 daily 写权契约** —— I-2 单作者(同 folder 不并发改);folder 名 = summary note 名(I-3 可移动单元)
**显式排除**(不属于 auto-memory 职责):
- ❌ 蒸馏 / 沉淀:那是 auto-dream(`auto_dream_design.md`)的事
- ❌ 实体识别 / wikilink 自动补全:那是 auto-link(`auto_link_design.md`)的事
- ❌ 改写 resource / digest:auto-memory 只写 daily(I-1 / I-3)
---
## 1. 已对齐决策
### 1.1 物理布局:与 `structure.md` §2.2 对齐
| 项 | 决策 |
|---|---|
| **主轴** | 时间 + 任务:`daily/<date>/<event-slug>/` |
| **event-slug** | LLM 抽取的事件短名(snake_case / dash-case;不强制 schema),同 `<date>` 下唯一 |
| **folder 内** | 一个事件可有 N 个 note(`progress.md` / `decision.md` / `references.md` 等),由消费层 schema 决定;最少含一个 summary note,与 folder 同名 |
| **主索引** | `daily/<date>.md`:当天事件列表(机械写入,wikilink 指向各 event folder)|
| **跨日索引** | 不强制;dream 消费时按 `<date>` 范围拉取即可 |
**为什么不是单文件 event**(report §5.1 一种简化方向):
- 单文件 event = `daily/<date>/<event-slug>.md` 比 folder 模型简单,但失去"一个事件可包含多个视角 note"的灵活度
- 现行 `structure.md` 已定 folder 单位模型;auto-memory 沿用,不破坏既有 I-2 / I-3
- 若后续 dogfooding 验证单事件普遍只有一份 note,可演进为 folder 内只放一份 summary,机械上等价于单文件方案 —— 演进路径平滑,不需要现在选
### 1.2 事件边界:语义意图切换驱动
事件边界由 LLM 在 inline 写入时判:**当前回合的意图是否仍属上一个 event**。
| 维度 | 决策 |
|---|---|
| **决策时机** | 每个 agent 回合写入前 inline 判 |
| **决策依据** | 上一个 active event 的 summary + 当前回合内容;LLM 输出 `{continue: bool, new_event_slug?: str, summary_patch?: str}` |
| **continue=true** | append 当前回合到 active event(append-only 或 LLM 重写 summary,详 §1.3) |
| **continue=false** | 关闭 active event(写最终 summary)+ 开新 event folder(slug 由 LLM 给)|
| **同时 active 多事件** | 不允许(I-2 单作者)—— 一时刻只一个 active event;真要并行任务,agent 自己 sync 切换 |
**已排除**:
- 时间窗口切分(N 分钟无活动则切)—— 对话节奏因任务而异,时间窗口噪声大
- 关键词切分(出现"切换 / 现在做 X"等触发词)—— 假阳性高,且不所有切换都明显说出
- 后置 batch 切分 —— inline 写入要求 event 必须当下可决定归属,不能等
### 1.3 事件内写入模型
active event 内,每个回合的内容写到 event folder 下,有两种模式可选(消费层 schema 决定):
| 模式 | 形态 | 适用 |
|---|---|---|
| **append-only** | 一份 `<event-slug>.md`,新回合 append 到末尾(章节 / 时间戳 / 等)| 实现最简;事件短(< 几十回合)时可读性 OK |
| **多 note 重写** | summary note(folder 同名)+ 各视角 note(`progress.md` / `decision.md`);LLM 把新内容融到对应 note,summary note 重写为当下概览 | 事件长 / 多视角时可读性高;LLM 成本高 |
**默认 opinionated default**:append-only(最简启动)。消费层可改 prompt + schema 走多 note。
**与 E-1 守恒的关系**:daily 不强制 E-1 守恒(它是工作记录,允许 LLM 删旧加新);只在 multi-note 重写模式下,可选启用类似守恒(保留所有 wikilink),具体由消费层决定。
### 1.4 主索引 `daily/<date>.md`
当天事件 list 视图,机械维护(无需 LLM):
| 触发 | 操作 |
|---|---|
| 新建 event folder | 主索引 append 一行 `[[daily/<date>/<event-slug>/<event-slug>.md|<event-slug>]]` |
| 事件关闭(被切下一个 event) | 主索引该行 append 最终 summary 摘要(可选,LLM 写最终 summary 时附带写入) |
| 索引文件不存在 | 写入第一个 event 时创建 |
主索引**仅承担当天浏览锚点**:文件系统 `ls daily/<date>/` 也能看见,但有主索引人/agent 可直接 `read daily/2026/05/28.md` 拿到 list 视图 + summary 一览。
不维护跨日索引(`daily/2026/05.md` 或 `daily.md`):dream 消费时按时间范围拉取即可;`list daily/<date>/` 已经覆盖浏览需求。
### 1.5 与 notify 的协作
`notify` 是 reme → agent 的虚边推送(`structure.md` §3.3),把"有新 resource 值得看"传递给 agent。auto-memory 在以下两点与 notify 协作:
| 维度 | 协作方式 |
|---|---|
| **新事件 cue** | agent 收到 notify 后,如果决定响应(开始处理这个候选),通常会触发**新 event** —— auto-memory 把 notify payload 作为 hint(候选 resource 路径)写入新 event 的 summary,顺手用 wikilink 引上 |
| **acknowledge 派生** | event note 里出现指向 `[[resource/...]]` 的 wikilink → L1 watcher 将该 resource 推送状态置 `acknowledged`(`structure.md` §3.3 / §6.2);auto-memory 自身不调任何 ack API |
**关键约束**:auto-memory **不强制** agent 用 wikilink 引 notify 候选 —— agent 可能略过、也可能不通过 wikilink 而是直接读 resource。ack 是 daily → resource wikilink 的副产品,不是 auto-memory 显式负责的事。
---
## 2. 待对齐边界点
### 2.1 LLM 决策频率与成本
inline 边界检测的最朴素形态是每回合调一次 LLM。在长对话 + 高频回合下成本可观。可选优化:
- **continue 假设默认**:大多数回合是 continue(同一意图内),LLM 可能只在"看似切换"启发(token 跨度大 / 工具种类突变 / 用户显式说"接下来")时跑;否则默认 continue 不调 LLM
- **批回合**:每 N 回合批一次,延迟切分(代价:active event 边界滞后,首版可接受)
首版默认每回合调一次(最简,正确率高),M1+ 视成本优化。
### 2.2 中断恢复 / 跨进程 active event
agent 进程重启 / Service 重启后,如何识别"还有 active event"?
候选方案:
- **L2 自治状态**:L1 watcher 派生 `daily/<date>/<event-slug>/` 中最新 mtime 的 event 为 active(默认 N 分钟内有写入)
- **状态文件**:`.daily-active` 维护 active event slug,Service 启动时读
- **每次重建**:agent 进程重启视为新 event,旧的关闭(切到 §1.2 continue=false 路径)—— 最简但会增加事件数
倾向 §1.2 自然路径(进程重启 = LLM 下次判 continue=false 概率高)+ 不维护状态文件,详细 worker recovery 留给 Service 实现。
### 2.3 多 agent 同 vault 的 active event 隔离
I-2 daily 单作者契约在多 agent 场景下需细化。候选:
- per-agent date subfolder:`daily/<date>/<agent-id>/<event-slug>/`
- 单 agent 模式 + agent ID 进 event-slug:`<date>/<agent-id>_<event-slug>/`
第二种破坏 slug 短名习惯;第一种引入额外层级。倾向后者作为消费层契约,reme 核心不固化。
### 2.4 事件粒度的 prompt 引导
边界检测的 prompt 决定切分粒度。粗 = event 大 / dream 看每个 event 时容易 overflow;细 = event 数爆炸 / 主索引拥挤。
**opinionated default prompt 倾向**:
- 一个意图 = 一个事件(用户提了 X 问题 / agent 开了 Y 任务 → 直到这个意图收尾)
- 跨意图的"附带工作"(查资料 / 算个数)归入当前意图,不开新 event
- 真新意图("好,现在我们做下一件事")才切
详细 prompt 落 `reme4/steps/jobs/protocol.md` 或 synchronizer 的 prompt 模板。
### 2.5 与 resource ingest 的时序
如果 ingest 与 auto-memory 同时活跃(External push 推 resource 进来 + agent 在 sync),且 agent 想响应这个新 resource:
- ingest 写完 resource → L1 watcher 派生 L2 → notifier 决策推送(`structure.md` §5.3)→ Service MCP 推给 agent
- agent 在当前回合或下一回合响应 → auto-memory 判 continue=false 开新 event,wikilink 引上 resource
整条链 sub-second 到 seconds(notify 节奏);auto-memory 不直接知道 ingest,只在 agent 决定响应时被动接收 notify payload。
### 2.6 跨日任务延续
event 物理路径含日期(`daily/<date>/<event-slug>/`),同一意图跨日的任务无法用同一 event folder 承载。候选模型:
| 模式 | 形态 | 适用 |
|---|---|---|
| **每日新 event,wikilink 反指前日** | new day 起新 folder;summary note frontmatter 加 `inherits: [[daily/<prev-date>/<prev-slug>/<prev-slug>.md]]`;新 event body 不复制旧内容,仅引用 | event-slug 短,日切口干净;查 backlinks 拼出整条任务链 |
| **同 event 重复写不同日** | 不允许(I-2 single author + event folder date 在路径上,跨日写违反路径不可变) | × |
| **任务 ID 跨 daily 抽象** | 引入 `task-id` 维度,daily event 只是某 task 的某一日切片;额外维护 task index | 复杂度高,M0 不引入 |
**倾向**:第一种(`inherits` frontmatter wikilink)—— 与 §1.4 主索引一致(机械维护),实现侧 LLM 在 §1.2 boundary 判定时若发现意图与最近 N 天某个 active 任务一致,直接写入 inherits 即可。详细 boundary prompt 落 §2.4。
INHERIT 行为细节(扫描窗口、predecessor 是否关闭、Plan/Objective 是否拷贝)归消费层 schema 决定;reme 核心只承认 `inherits:` frontmatter wikilink 作为跨日链路载体。
---
## 3. 与其它层的协作
| 上下游 | 关系 |
|---|---|
| ← **notify** | 接收 notify payload 作为新 event cue;不强制响应,不强制 wikilink 引 |
| ← **resource** | 只读(通过 wikilink 引);不写 |
| → **daily** | **唯一写者**(I-2);写 event folder + 主索引 |
| → **auto-dream** | dream 的 G\* 读 daily 作为入流(`auto_dream_design.md` §2.1 G1 scope);auto-memory 写完即对 dream 可见(走 L2 索引,有 eventual 窗口) |
| → **auto-link** | auto-link 可反向扫 daily event,做实体识别 + wikilink 写回(`auto_link_design.md` §1.3)—— 与 auto-memory 写入不冲突(双方写不同字段段落 / CAS 协议保护)|
**关键边界**:auto-memory 是 daily 写入端的**唯一**入口;dream / link 不写 daily 主路径,只通过 auto-link 走 §1.3 写回(read-only audit-then-write,CAS 保护)。
---
## 4. 下一步
1. **synchronizer step 实现**:event 边界检测 prompt + active event 状态管理 + inline 写入(append-only 默认)
2. **主索引维护**:`daily/<date>.md` 机械维护(新 event 时 append、关闭时附 summary)—— 走 crud/daily 基础工具
3. **notify ack 派生验证**:L1 watcher 派生 acknowledged 状态(`structure.md` §6.2),与 auto-memory 的 wikilink 写入端到端跑通
4. **多 agent 隔离 schema**(M1+):若实际有并发 agent,确定 daily 子目录 / slug 命名约定
5. **粗 / 细粒度 prompt 调参**:dogfooding 后看实际 event 数 / dream 消化效率,调 boundary prompt
实现进入 `reme4/steps/jobs/` 与 `reme4/file_graph/` 时,本文档与 `auto_dream_design.md` / `auto_link_design.md` 共同作为契约依据。

782
docs4/structure.md Normal file
View file

@ -0,0 +1,782 @@
# reme4 系统架构 — 设计文档
## 文档定位
本文档定义 reme4 的**架构设计**:概念边界、数据流契约、职责划分。
- 不涉及代码路径 / 实现进度 / API 具体形态
- **执行栈与架构角色**(分层、模块切分、职责划分)在第 6-7 节;源码映射与落地状态见 `docs4/reme4_report.md` 与源码
- 不规定怎么做,只规定**是什么、谁负责、输入输出**
> 阅读顺序:**第 1 节**给出完整的架构总览(数据视角 + 运行时视角 + 核心机制 + 不变量 + 导航);**第 2-5 节**逐层展开三层存储 / 6 类 L4 动作语义 / 反向回流 / 触发节奏;**第 6 节**描述底层执行栈(L0→L5);**第 7 节**把 6 类 Action 落到 L4 实现模块;**第 8 节**讲跨切面 schema 契约;**第 9 节**列出明确不属于本架构的反例。
---
## 1. 架构总览
本章给出 reme4 完整的设计骨架,后续 §2-§9 逐项展开细节。
### 1.1 解决什么问题
reme4 是 agent 的**长期记忆系统**。它把 agent 的工作过程沉淀为可检索、可演化的知识结构。
设计要解耦两件事:
| 关注点 | 由谁负责 |
|---|---|
| **agent 写什么 / 读什么** | agent 的工作流自决 |
| **vault 自身如何健康演化** | reme 自治,agent 不感知 |
实现方式:三层存储拓扑为"两路并行写入(原始资料 / 任务过程) → 双源合流到沉淀知识层",reme 提供这三层的容器、动作语义、反向检索与自治维护。
### 1.2 数据视角:并行起点 + 双源合流 + 反向回流
数据有**两个并行起点** —— **External source**(webhook/upload/pull)与 **Agent**(外部主体,自身任务驱动)。两条独立通道各自落地到**并行的材料层**:External 经 `ingest` 沉到 `resource/`,Agent 经 `sync` 写到 `daily/`。两层材料**合流**到 `digest/`,由 reme 通过 `digest` 动作完成"消化"。`digest/` 自身由 `maintain` 做 in-place 重组。Agent 通过 `retrieve` 从 resource + daily + digest **三层并行**回流。`notify` 是 reme 跨过 vault 直接提醒 Agent 的**虚边**(控制信号,不写任何文件)。每段路径都对应一个 L4 动作语义(完整动作详见 §5.2 / §7):
| 路径 | L4 动作 | 说明 |
|---|---|---|
| External source → resource | **ingest** | 外部信源落到 vault |
| External source ╌╌► Agent | **notify** | reme 推送通知(虚边,只走 L2 推送队列,不写文件) |
| Agent → daily | **sync** | Agent 写 daily(响应 notify 或自身任务驱动) |
| resource + daily → digest | **digest** | reme 内部 LLM 抽取沉淀(双源合流) |
| digest → digest | **maintain** | reme 内部 LLM 折叠重组(in-place, fold-only) |
| resource + daily + digest → Agent | **retrieve** | 三层并行回流(state / semantic / topological 三种正交问法) |
```
notify(虚边,控制信号)
External source ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌► AGENT
(webhook/upload/pull) (外部主体/自身任务)
│ │ ▲
│ ingest │ │ retrieve
│ ┌────── sync ────────────────────-┘ │ (state /
▼ ▼ │ semantic /
┌──────────┐ ┌──────────┐ │ topological)
│resource/ │ │ daily/ │ │
│ 原始资料 │ │任务工作区 │───────── retrieve ─────────────┤
│ 不可变 │ │ 半可变 │ │
└───┬──┬───┘ └────┬─────┘ │
│ │ │ │
│ │ digest digest │ │
│ └────────┐ ┌────┘ │
│ ▼ ▼ │
│ ┌────────────────────┐ │
│ │ digest/ │ ◄──╮ │
│ │ 沉淀知识 │ │ maintain │
│ │ 可重组(语义索引) │ ───╯ (in-place, fold-only) │
│ └──────────┬─────────┘ │
│ │ retrieve │
│ └─────────────────────────────────────-─┤
│ │
│ retrieve │
└──────────────────────────────────────────────────────────┘
```
**模型要点**:
- **两个起点平行,不存在主从** —— External 与 Agent 各自独立驱动;Agent 既可响应 `notify` 也可由自身任务直接 `sync`。
- **两层材料平行,不存在传递** —— resource 与 daily 是**两条独立的写入通道**,不互相穿越:Agent 不写 resource,ingester 不写 daily。
- **digest 是双源合流的产物** —— `digest` 动作的输入是 resource + daily 的组合(不是仅 daily);相应地,digest 节点的 provenance 可同时指向 resource 与 daily。
- **digest 自循环** —— `maintain` 在 digest 内部做密度折叠,不与上游材料层交互。
- **notify 是虚边** —— reme 用它提醒 Agent "有新 resource 值得看",但不落任何文件;Agent 的响应通过 `sync` 落 daily(并可选地用 wikilink 引 resource)。
retrieve 三种问法正交:
| 问法 | 工具 | 主要看哪层 |
|---|---|---|
| **state**(谁在 / 是什么状态) | `list` / `frontmatter` | 各层平等 |
| **semantic**(我想到一个意思) | `search` | digest > daily > resource(默认权重) |
| **topological**(从一个点向外摸) | `traverse` | 沿 wikilink 跨层平等 |
### 1.3 运行时视角:六层执行栈 + 双进程
reme4 的功能不是堆在一层,而是从文件系统底层往上栈式堆叠。顶层 Service 与 Runtime 是同一套 vault 上的两个进程角色,共享 L0-L4 全栈(详见 §5.3 / §6)。
```
┌─────────────────────┐ ┌─────────────────────┐
│ L5 Service │ │ L5 Runtime │
│ (HTTP / MCP) │ │ (scheduler 自治) │
└──────────┬───────────┘ └──────────┬──────────┘
│ │
└─────────────┬────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ L4 6 类 Action(动作语义) │
│ ingest notify sync retrieve digest maintain│
└──────────────────────┬───────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ L3 原子工具 │
│ ┌─────────────────────┐ ┌──────────────────┐ │
│ │ 基础工具 │ │ 高级工具 │ │
│ │ create/append/edit/ │ │ search │ │
│ │ read/write/move/ │ │ traverse │ │
│ │ delete/list/stat │ │ frontmatter │ │
│ └──────────┬──────────┘ └────────┬─────────┘ │
└─────────────│──────────────────────│─────────────┘
│ 直读 / 直写 │ 走索引读
│ (eventual,有滞后) │
│ ▼
│ ┌────────────────────────────┐
│ │ L2 文件状态 │
│ │ · file_store(chunk+vec) │
│ │ · file_graph(node+link) │
│ │ · 自治状态(scheduler 用): │
│ │ - resource: 入流批次/ │
│ │ 未消化(orphan) │
│ │ - daily: 任务索引 │
│ │ (进行中/stale/完成) │
│ │ - digest: 密度水位/ │
│ │ 断链(broken wikilink) │
│ │ · 推送队列(notify): │
│ │ pending/notified/ │
│ │ acknowledged │
│ └─────────────▲──────────────┘
│ │ 派生 / 更新
│ ┌─────────────┴──────────────┐
│ │ L1 file_watcher │
│ │ fs event → state delta │
│ │ (唯一 fs→state 桥) │
│ └─────────────▲──────────────┘
│ │ 监听
▼ │
┌──────────────────────────────────────────────────┐
│ L0 vault 文件系统 │
│ resource/ daily/ digest/ │
└──────────────────────────────────────────────────┘
```
Service 与 Runtime 是同一份 vault 上的两个进程角色:
| 进程 | 触发源 | 时延敏感 | 典型动作 |
|---|---|---|---|
| **Service** | 外部 push / 外部 pull / agent 同步请求 / **MCP 推送通道** | 是 | ingest / sync / retrieve / **notify-out(MCP)** |
| **Runtime** | scheduler 周期 + L2 自治状态阈值 | 否(eventual) | **notify 决策** / digest / maintain |
### 1.4 核心机制总览
| 机制 | 一句话 | 详见 |
|---|---|---|
| 三层存储 | resource(冷) / daily(温) / digest(冷,组织化) | §2 |
| 6 类 L4 动作 | ingest / notify / sync / retrieve / digest / maintain | §1.2 / §3 |
| notify+sync 链 | reme 主动从 L2 资源自治状态选候选,经 MCP 推给 agent;agent sync 落 daily | §3.3-3.4 / §7.1 |
| 反向 retrieval | state / semantic / topological 三种正交问法 | §4 |
| 触发四源 | 外部 push / 外部 pull / agent on-demand / reme 后台 | §5.1-§5.2 |
| Service + Runtime | 双进程角色,共享 L0-L4,职责按时延分 | §5.3 |
| 执行栈(L0-L5) | filesystem → file_watcher → 文件状态 → 原子工具 → Action → Service/Runtime | §6 |
| 原子工具:基础 vs 高级 | 基础直 fs;高级走 L2 索引(eventual) | §6.4 / §5.5 |
| Action 模块映射 | 6 类 Action 由 5 个模块实现(retrieve 直走原子工具) | §7 |
| Schema 跨切面 | name+description 是核心强约束,其余 opinionated default 可重载 | §8 |
### 1.5 核心不变量速览
写入拓扑(架构脊梁,来自 §2.3):两路并行写入(External→resource、Agent→daily)→ 双源合流到 digest;resource 与 daily 之间互不写入;任何一层都不能反向改写它的上游。
| 不变量 | 内容 | 来源 |
|---|---|---|
| **I-1** | agent 不直接写 digest(digest 写权只属 digester / maintainer) | §2.4 |
| **I-2** | daily folder 单作者(同 folder 不并发改) | §2.4 |
| **I-3** | resource 内容不可变,只允许 metadata appendable | §2.4 |
| **I-4** | 三层共用同一套 wikilink 索引,跨层引用全靠 wikilink | §2.4 |
| **R-1** | retrieve 三种问法分立,不合并为单一 read verb | §4.3 |
| **M-1** | Maintainer 只做一件事:密度折叠(把碎片叶子折叠到新的中间节点下) | §7.3 |
| **F-1** | L1 `file_watcher` 是 L0→L2 的唯一派生桥 | §6.5 |
| **F-2** | L3 基础工具直接读写 L0;高级工具只走 L2 | §6.5 |
| **F-5** | L5 Service / Runtime 共享 L0-L4,不直接通信 | §6.5 |
| **F-6** | L0↔L2 存在 eventual 窗口,agent 上下文承担近期信息 | §5.5 / §6.5 |
### 1.6 文档导航
| 想了解… | 看 |
|---|---|
| 三层各自的定位、不变量 | §2 |
| 6 类 L4 动作语义(notify / sync / digest / maintain 的输入产出不变量) | §3 |
| Retrieval 的三种问法与跨层语义 | §4 |
| 谁来触发、什么节奏、为什么分两个进程 | §5 |
| 系统从文件系统到 Service 的分层(底层基础) | §6 |
| L4 五个模块的对称结构与 Maintainer 折叠设计 | §7 |
| Schema 协议与重载机制 | §8 |
| 哪些设计不属于本架构(反例与边界) | §9 |
| 术语回查 | 附录 |
---
## 2. 三层存储
### 2.1 一句话定位
| 层 | 一句话 |
|---|---|
| **resource/** | 外部原始资料的**不可变快照**。reme 是容器,不是作者。 |
| **daily/** | agent 的**任务工作区**。folder 是单位,以"日 + 任务"为索引。 |
| **digest/** | 跨任务沉淀的**有组织知识**。以语义(概念/实体/方法)为索引,与时间无关。 |
### 2.2 五维度对照
| 维度 | resource/ | daily/ | digest/ |
|---|---|---|---|
| **组织主轴** | 时间(`<date>/<name>`) | 时间 + 任务(`<date>/<slug>/`) | 语义(`<slug>/<subslug>/...`,任意嵌套) |
| **写权归属** | 入流通道唯一(webhook / upload / pull) | agent(写入任务过程) | digester / maintainer(无 agent 直写) |
| **可变性** | 不可变,只追加新文件 | folder 内可反复更新 | 单节点可演化,可被合并/拆分/移动 |
| **不变量** | 写入即冻结,原文永不变 | folder 名 = summary note 名(可移动单元);同 slug 同日只一份 | slug 全局唯一;每 folder 有 canonical entry;wikilink 全路径 |
| **谁在用** | agent(查原文)、digester(双源输入之一) | agent(自己的工作记录)、digester(双源输入之一) | agent(召回主目标)、maintainer(自维护对象) |
### 2.3 写入纪律:并行写入 + 双源合流
```
External source Agent
│ │
│ ingest │ sync
▼ ▼
┌──────────┐ ┌──────────┐
│resource/ │ │ daily/ │
│ 不可变 │ │ 半可变 │
│(ingester)│ │ (agent) │
└─────┬────┘ └─────┬────┘
│ │
│ digest digest │
└──────────────┐ ┌─────────────┘
▼ ▼
┌──────────────┐ ◄──╮
│ digest/ │ │ maintain
│ 可重组 │ ────╯ (in-place,
│ (digester + │ fold-only)
│ maintainer) │
└──────────────┘
```
写权按这个**两层并行 → 单层合流**的拓扑分配:resource 写权专属 ingester(外部入流通道),daily 写权专属 agent(sync 落入,可响应 notify 或自身任务驱动),digest 写权专属 digester + maintainer。**resource 与 daily 之间互不写入**(agent 不动 resource,ingester 不动 daily);任何一层都不能反向改写它的上游。这是整个架构的脊梁。
### 2.4 不变量(永远成立)
| # | 不变量 | 否则后果 |
|---|---|---|
| **I-1** | agent 不直接写 digest | digest 的"有组织"性失守,沉淀质量退化 |
| **I-2** | daily folder 单作者(同 folder 不并发改) | 任务边界模糊,sync/digest 竞态 |
| **I-3** | resource content immutable,只允许 metadata appendable | 原文可能消失/被覆写,citation 不可信 |
| **I-4** | 三层共用同一套 wikilink 索引,跨层引用全靠 wikilink | 引入第二套引用机制 → 索引重建复杂 / 跨层关系不可达 |
---
## 3. L4 动作语义详解
§1.2 给出了 6 个 L4 动作在数据视角下的整体形态。本节按动作逐个展开输入 / 产出 / 不变量 / 反例。`ingest`(外部→resource,机械)和 `retrieve`(三层并行回流,只读)分别在 §5/§7.4 与 §4 详述,本节聚焦四个**写动作**:`notify` / `sync` / `digest` / `maintain`。
### 3.1 统一原则
四个写动作都遵守:
| 原则 | 内容 |
|---|---|
| **Monotonic content** | 上游内容不可变,下游只能新建节点或加链接,不能改写上游 |
| **Provenance 必须可达** | 任何下游节点必须能通过 wikilink 反查到上游来源 |
| **Wikilink 是新结构的唯一载体** | 跨层关系靠 wikilink,不靠内容拷贝 |
它们都不是"数据搬家",而是"在下游新生成有引用关系的节点"。
### 3.2 四个写动作的本质对照
| 动作 | 上游 → 下游 | 性质 | 上游变化 | 下游变化 |
|---|---|---|---|---|
| **notify** | resource → Agent | **Attention**(推送注意力) | 不变 | 不写 vault;仅入 L2 推送队列 |
| **sync** | Agent → daily | **Reference**(引用落地) | 不变 | daily 中新增工作记录 + 对 resource 的 wikilink |
| **digest** | resource + daily → digest | **Crystallize**(双源合流结晶) | 不变 | digest 新增节点,wikilink 反指上游来源(daily 与/或 resource) |
| **maintain** | digest → digest | **Reorganize**(重组) | 结构变,内容守恒 | fold-only:引入子中间节点搬叶子,改变拓扑 |
> `notify` 与 `sync` 共同实现"resource 中的候选被 Agent 看见并织入 daily"这条**Reference 链**;它们是两个独立的 L4 动作,主体不同(notify 由 Reme 自治触发,sync 由 Agent 触发)。
### 3.3 notify:reme → Agent 推送候选
| 维度 | 内容 |
|---|---|
| 主体 | Reme Runtime(`notifier` 模块) |
| 输入 | L2 资源自治状态:orphan(无 inbound wikilink)/ 入流批次 / 未消化老于 N |
| 产出 | L2 推送队列条目;通过 Service MCP **server-initiated notification** 推到 Agent;**不写任何 vault 文件** |
| 不变量 | resource 原文 0 修改;**完全单向**,不维护任何反向元数据;`notify` 决策在 Runtime,Service 只作 MCP transport |
| Acknowledge 机制 | L1 watcher 检测到 daily→resource 新 wikilink → L2 推送状态 `notified` → `acknowledged`,避免重复推送 |
| 反例 | (a) Service 自决推什么 notify(✗-17);(b) notify 写入 vault(✗-16);(c) Agent 主动调 notify(✗-18) |
### 3.4 sync:Agent → daily 落地
| 维度 | 内容 |
|---|---|
| 主体 | Agent(`synchronizer` 模块在 Service 内编排) |
| 输入 | Agent 当前事件流(响应 `notify` 的候选,**或**自身任务直接驱动) |
| 产出 | daily folder 内的工作叙事;可选地用全路径 wikilink 引 resource(agent 自决,reme 不强制) |
| 不变量 | resource 原文 0 修改;daily 单作者(I-2);folder 名 = summary note 名(可移动单元) |
| Provenance | daily → resource 可达(通过 daily body 中的 wikilink) |
| 反例 | "agent 把 resource 内容拷进 daily" —— 不允许,daily 只持有引用 + 自己的工作记录 |
### 3.5 digest:resource + daily → digest 双源合流结晶
| 维度 | 内容 |
|---|---|
| 主体 | Reme Runtime(`digester` 模块,LLM-driven) |
| 输入 | 一组待蒸馏的 daily folder + 相关 resource(双源合流;通常以 daily 任务为线索,顺着 wikilink / 同主题搜索拉入相关 resource 原文) |
| 产出 | digest 中 0~N 个新节点 或 已有节点的更新;新节点必须用 wikilink 反指至少一个上游来源 |
| 不变量 | resource / daily 正文 0 修改;digest 新节点必须 wikilink 反指上游(provenance);digest 节点遵守第 2.4 节列的不变量 |
| Provenance | digest → daily / resource 双源链条可达(资料源是 resource 时直接反指,任务过程是 daily 时反指 daily 进而可达 resource) |
| 反例 | digester 改写 resource;digester 改写 daily 正文 |
### 3.6 maintain:digest → digest 折叠
| 维度 | 内容 |
|---|---|
| 主体 | Reme Runtime(`maintainer` 模块,LLM-driven,**fold-only**) |
| 输入 | digest/ 当前整体状态 |
| 产出 | 同一 digest/ 树的**密度折叠**(fold):某中间节点下叶子过多时,引入子中间节点把相关叶子归簇 + 写"高密度摘要" |
| 不变量 | 树**只向下生长**(从不反向);叶子内容 0 修改,只被搬位置;新中间节点 = 一个高密度摘要文件;任何节点移动**原子重写所有入边**(retarget);slug 全局唯一在折叠后仍成立 |
| Provenance | digest → daily 的反指链接在折叠后仍有效(retarget 保证) |
| 反例 | merge / move / promote / demote 等改写既有拓扑的操作;改写既有叶子内容 |
> 折叠操作的承诺与决策点见 §7.3。
### 3.7 链接重定向例外
`maintain` 的 retarget 会改写其它节点中指向被移动节点的 wikilink。从字面看,这违反了"上游内容不可变"。
实际上这是 wikilink 系统的**机械性副作用**,不算下游写上游:
| 字面 | 实质 |
|---|---|
| daily 里的 `[[digest/old.md]]` 被改成 `[[digest/new.md]]` | 作者意图("我引用了 X 这个 digest 节点")没变,只是 X 的物理位置变了 |
只要 retarget 保持 wikilink 的**目标语义不变**,就允许它作为机械维护副作用穿越层界。这是这条规则的唯一例外。
---
## 4. Retrieval 反向回流
Retrieval 是把 §3 几条正向写动作反着读:agent 站在结果端,沿 wikilink 反查源头。
### 4.1 三种问法
按 agent 意图分,有三类完全不同的读需求,**正交**,各自独立:
| 问法 | 例子 | 本质 |
|---|---|---|
| **状态问** (State) | "我有哪些 in-progress 的任务?""哪些 resource 还没被引用?" | 在某层做 list + frontmatter 过滤 |
| **语义问** (Semantic) | "关于 auth 重构我知道什么?" | 跨层全文/向量检索 |
| **拓扑问** (Topological) | "auth 概念周围都连了什么?" | 从某节点沿 wikilink 走 |
### 4.2 三层 × 三问法 矩阵
| 问法 | resource/ | daily/ | digest/ |
|---|---|---|---|
| **状态问** | "未处理 resource 清单" | "active / pending-digest 清单" | "孤儿节点 / canonical 缺失 清单"(给 maintain 用) |
| **语义问** | 兜底(原文,信噪比低) | 次优(最新,但未沉淀) | **首选**(沉淀过,信噪比高) |
| **拓扑问** | 通常是叶子(被指向) | daily → resource / digest | digest 内部连接最密 |
### 4.3 设计原则
| # | 原则 | 含义 |
|---|---|---|
| **R-1** | 三种问法分立,不合并为单一 "read" verb | 不同问法的索引、过滤、排序逻辑完全不同 |
| **R-2** | 语义问的默认权重 `digest > daily > resource`,**可被显式覆盖** | 默认体现"沉淀质量",但 agent 可指定单层或调权 |
| **R-3** | 拓扑问与层无关 | traverse 沿 wikilink 走,天然跨三层(I-4) |
| **R-4** | Provenance expansion 默认 lazy,eager 是上层便利封装 | 原子 retrieval 不自动展开;agent 需要时再 traverse |
| **R-5** | Cold start 不是新的 retrieval mode | 只是状态问 + 语义问的组合,reme 不为它单设 verb |
### 4.4 在主干图里的位置
Retrieval 不引入新存储,不引入新层。它是 **agent 与三层存储之间的读视图**,通过三种正交问法暴露,共享同一套 wikilink 索引(I-4)。
---
## 5. 节奏与触发
锁定"谁推动每个动作发生"。这一步定 reme 是纯被动 service 还是带后台 runtime。
### 5.1 触发源四分类
| 触发源 | 性质 | 例子 |
|---|---|---|
| **External push** | 外部事件主动推 | webhook / 用户 upload |
| **External pull** | reme 主动去外部拉 | scheduled fetcher(RSS / 邮件 / API 轮询) |
| **Agent on-demand** | agent 在请求里显式调用 | "sync 我的对话" / "搜 X" |
| **Reme background** | reme 自己的 watcher / scheduler | file_watcher / cron-like |
### 5.2 六个动作的触发归属
| 动作 | 主触发 | 备用触发 | 备注 |
|---|---|---|---|
| **ingest** | External push / pull | — | 外部→resource,机械 |
| **notify** | **Reme background**(cron + L2 资源自治状态阈值) | — | Runtime 决策,Service MCP 推送 |
| **sync** | Agent on-demand | — | agent→daily;notify 的响应也走这里 |
| **digest** | **Reme background** | Agent 显式(后门) | resource + daily 双源合流到 digest |
| **maintain** | **Reme background**(cron + threshold) | Agent / 人工 显式(后门) | digest 内部折叠 |
| **retrieve** | Agent on-demand | — | 三层只读 |
**关键定性**:`notify` / `digest` / `maintain` 三个 reme 自治动作的主控制权**在 reme,不在 agent**。Agent 只负责"响应 notify + 写 daily(响应或自身任务驱动)+ 主动读";不需要记得"该看哪些 resource""该蒸了""该整理了"。
### 5.3 Service + Runtime 双进程结构
把 5.2 的归属直接推出 reme 的基本架构。两进程共享 L0-L4 全栈,只在 L5(进程入口)分叉:
```
┌──────────────────────────────────────────────────────────┐
│ Reme System │
│ │
│ ┌────────────────────┐ ┌────────────────────────┐ │
│ │ L5 Service │ │ L5 Runtime │ │
│ │ (HTTP / MCP) │ │ (scheduler 自治) │ │
│ │ 服务 agent 请求: │ │ 服务 vault 健康: │ │
│ │ · ingest │ │ · notify(决策) │ │
│ │ · sync │ │ · digest │ │
│ │ · retrieve │ │ · maintain │ │
│ │ · notify(MCP 推送)│ │ │ │
│ └─────────┬──────────┘ └───────────┬────────────┘ │
│ │ │ │
│ └──────────────┬───────────────┘ │
│ ▼ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ L4 Action / L3 原子工具 │ │
│ │ Action 编排 → 基础工具 + 高级工具 │ │
│ └──────────────────────┬───────────────────────────┘ │
│ ▼ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ L2 文件状态(file_store + file_graph + 自治) │ │
│ └──────────────────────▲───────────────────────────┘ │
│ │ 派生 │
│ ┌──────────────────────┴───────────────────────────┐ │
│ │ L1 file_watcher(fs → state 的唯一桥) │ │
│ └──────────────────────▲───────────────────────────┘ │
│ │ 监听 │
│ ┌──────────────────────┴───────────────────────────┐ │
│ │ L0 vault filesystem │ │
│ │ resource/ daily/ digest/ │ │
│ └──────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────┘
```
两进程职责正交、共享 L0-L4 基础设施(详见 §6):
| 维度 | L5 Service | L5 Runtime |
|---|---|---|
| 触发方式 | 请求-响应 + MCP server-initiated 推送 | 周期 + L2 自治状态阈值 |
| 服务对象 | agent | vault 自身 |
| 暴露给 agent | 是 | 否(agent 不感知) |
| 主要动作 | ingest / sync / retrieve / **notify 推送通道**(MCP) | **notify 决策** / digest / maintain |
| 与对端的耦合 | 通过 L2 推送队列读 notifier 产出 | 通过 L2 推送队列写,**不直接调 Service** |
### 5.4 节奏(latency tolerance)
| 动作 | 节奏 | latency 容忍 |
|---|---|---|
| retrieve | request-driven | sub-second |
| ingest | event-driven | seconds |
| sync | agent on-demand | seconds |
| **notify** | reactive(L2 资源状态变化后) | seconds ~ minutes |
| digest | reactive(状态变化后) | minutes ~ hours(eventual consistency) |
| maintain | periodic | days(无紧迫) |
实时性需求差三个数量级。这是 digest / maintain 必须放后台异步的根本原因 —— 不能阻塞 agent 的 retrieve / sync 请求。
### 5.5 排序约束与一致性模型
部分动作对**不能并发**,background runtime 内部要保证排序:
| 约束 | 原因 |
|---|---|
| **sync(同一 daily)→ digest(同一 daily)** | digest 不能看到 sync 半成品 |
| **digest(同一 scope)→ maintain(同一 scope)** | maintain 重组的拓扑不应被 digest 中途插入 |
ingest / retrieve 跟所有动作都可并发(纯入流 + 纯读)。
**一致性模型 = Eventual consistency on digest/maintain**。Agent 不能依赖"我刚写完 daily 就能查到对应 digest"。digest / maintain 都是后台异步,有可见的延迟窗口。
**watcher 滞后契约(L0 ↔ L2)**:基础工具直接写 L0 文件系统,L2 文件状态由 L1 file_watcher 派生,二者之间存在 eventual 窗口 —— 写完一份 daily 后,search / traverse 这类走 L2 索引的高级工具不一定立刻能看到。这是设计意图,不是 bug:agent 本身有上下文窗口,近期信息靠 agent 自带的对话上下文承接,不依赖 reme 索引立即可见。需要"写后立刻可读"的场景请用基础工具(read 直接读 fs)。
---
## 6. 执行栈:六层结构
L3 原子工具、L4 Action、L5 进程都不直接操作文件系统。它们坐在 L0-L2 的**底层基础**上 —— 这套基础是 reme 的"动力源",决定了为什么上层能解耦成 Service + Runtime 两进程,且两者既正交又共享状态。
### 6.1 概念分层(L0 → L5)
```
┌──────────────────────────────────────────────────────────┐
│ L5 Service ‖ Runtime │
│ 进程入口:Service 服务 agent;Runtime 自治维护 │
├──────────────────────────────────────────────────────────┤
│ L4 Action(6 类动作语义) │
│ ingest / notify / sync / retrieve / digest / maintain │
├──────────────────────────────────────────────────────────┤
│ L3 原子工具 │
│ 基础工具(直接 fs) + 高级工具(走 L2 索引) │
├──────────────────────────────────────────────────────────┤
│ L2 文件状态 │
│ file_store + file_graph + 自治状态(scheduler 用) │
├──────────────────────────────────────────────────────────┤
│ L1 file_watcher │
│ fs event → state delta(唯一 fs→state 桥) │
├──────────────────────────────────────────────────────────┤
│ L0 vault filesystem │
│ resource/ + daily/ + digest/ │
└──────────────────────────────────────────────────────────┘
数据流(主要关系):
· L3 基础工具 ──写──► L0
· L3 基础工具 ──读──► L0(无需经 L2)
· L0 变化 ──► L1 监听到 ──派生──► L2 state delta
· L3 高级工具 ──读──► L2(走索引)
· L4 Action ──编排──► L3 工具组合
· L5 进程 ──触发──► L4 Action
```
| 层 | 角色 | 关键约束 |
|---|---|---|
| L0 | vault 文件系统 | 唯一真相源;任何 L2 状态都可由 reindex 从 L0 重建 |
| L1 | `file_watcher` | **唯一**与 fs 事件直接耦合的组件;fs→state 的唯一派生桥 |
| L2 | 文件状态(`file_store` + `file_graph` + 自治状态) | 高级工具的读视图;由 L1 单向更新,L3+ 只读不写 |
| L3 | 原子工具(基础 / 高级) | 基础直读写 L0;高级只走 L2 |
| L4 | Action(6 类语义动作) | N:M 编排 L3 工具;不直接碰 L0 / L2 |
| L5 | Service / Runtime 进程 | 共享 L0-L4 全栈,**不直接通信**,只通过 L0 / L2 状态间接耦合 |
### 6.2 L1 file_watcher:fs → state 的唯一桥
`file_watcher` 是 reme 唯一与 OS filesystem 事件直接耦合的组件。它把 fs 变化翻译为 L2 文件状态的 delta,承担**双重职责**:
```
filesystem events (create / modify / move / delete)
│
▼
file_watcher ─┬─► 索引同步:写完文件,L2 file_store/file_graph 自动更新
│ (L3 基础工具不需要显式调用"入索引")
│
└─► 自治状态派生:维护 scheduler 用的可推导状态
· resource: 入流批次 / 未消化(orphan) /
推送状态(pending/notified/acknowledged)
· daily: 任务索引(进行中 / stale / 完成)
· digest: 密度水位 / 断链(broken wikilink)
```
**关键设计**:可派生的状态由 watcher 在外部索引中维护,**不写回 frontmatter**。L3 工具只管写内容文件,状态由 watcher 独立派生。这是 ✗-14 反例(L3 写 `status` 字段)成立的基础。
**Acknowledge 派生例**:`notify` 的"推送状态"由 watcher 维护 —— 当 watcher 检测到一条新 wikilink 从 daily 指向某 resource,即把该 resource 的推送状态从 `notified` 改为 `acknowledged`。notifier / Service 都不需要显式 ack。
### 6.3 L2 文件状态
| 组件 | 职责 | 由谁更新 |
|---|---|---|
| `file_store` | chunk 分块 + 向量持久化,提供 search / read API | L1 watcher 派生 |
| `file_graph` | wikilink 有向图,提供 upsert / traverse(双向)API | L1 watcher 派生 |
| **自治状态** | scheduler 自治决策的输入(入流批次 / 任务索引 / 密度水位 / 断链) | L1 watcher 派生 |
| **推送队列** | `notify` 的待推送 / 已推送 / 已确认条目 | notifier 写 pending;Service 推送后置 notified;L1 watcher 检测到 ack 后置 acknowledged |
L2 只关心"vault 当前是什么样",**无业务语义** —— 不知道 daily / digest / 动作语义的存在。L3+ 只读 L2,不写(**例外**:notifier 写推送队列,这是 Runtime 与 Service 之间唯一的间接耦合通道,见 F-5)。
> **现状提示**:当前实现中 file_store / file_graph 之外的"自治状态"和"推送队列"尚不完整,这是 L1 watcher 与 notifier 待补齐的能力。完整化后 scheduler 才能从"周期扫描"切换为"事件驱动",`notify` 才能从隐式变为显式。
### 6.4 L3 原子工具:基础 vs 高级
两组工具的切分依据只有一条:**是否必须经过 L2 索引**。
| 组别 | 工具 | 数据通路 | 一致性 |
|---|---|---|---|
| **基础工具** | create / append / edit / read / write / move / delete / list / stat | 直接对接 L0 | 写后立即可读(同一工具) |
| **高级工具** | search / traverse / frontmatter | 必须走 L2 索引 | 受 watcher 滞后影响(eventual) |
**写路径全部走基础工具**(L4 Action 编排基础工具完成写入)。高级工具是**只读**的索引查询入口。
**eventual 窗口**:基础工具写 L0 后,L2 索引由 L1 watcher 异步追平。在窗口内,高级工具看到的是滞后的视图。详见 §5.5"watcher 滞后契约"。
### 6.5 设计含义
| # | 不变量 | 推论 |
|---|---|---|
| **F-1** | L1 `file_watcher` 是 L0 → L2 的**唯一**派生桥 | 状态一致性是 L1 的事;L3 工具不要"自己更新索引" |
| **F-2** | L3 基础工具直接读写 L0;L3 高级工具只走 L2 | 写路径无需"先 reindex";读路径接受 eventual |
| **F-3** | L0 是唯一真相源 | 任何 L2 状态都可由 reindex 从 L0 重建,L2 是缓存而非数据库 |
| **F-4** | L4 Action 与 L3 工具是 N:M 编排关系 | Action 不直接碰 L0 / L2 |
| **F-5** | L5 Service 与 Runtime 共享 L0-L4 全栈,**不直接通信** | 只通过 L0 / L2 状态间接耦合;一边崩了不影响另一边的读 |
| **F-6** | L0 与 L2 之间存在 eventual 窗口 | agent 上下文承担近期信息,不依赖 L2 立即可见(详见 §5.5) |
---
## 7. L4 Action 模块映射
第 5 节列了 6 类 Action 及其触发源,本节把这些 Action 落到 **L4 实现模块**(架构角色,不指代源码路径);触发机制(on-demand 路径与 background 路径)见 §5.3。
### 7.1 五个 L4 模块
| 模块 | 实现动作 | 触发 | LLM-driven | 单一职责 |
|---|---|---|---|---|
| **ingester** | ingest | External push / pull | × | 原样落 resource + 抽 frontmatter + 入索引 |
| **notifier** | notify | Reme background(cron + L2 资源自治状态阈值) | × | 从 L2 资源自治状态选候选 → 写 L2 推送队列;Service MCP 拿走 |
| **synchronizer** | sync | Agent on-demand | ✓ | 把当下事件织入 daily 工作叙事 |
| **digester** | digest | Reme background | ✓ | resource + daily 双源合流成 digest 长期条目 |
| **maintainer** | maintain | Reme background | ✓ | digest topic tree 的**密度折叠**(fold-only) |
`retrieve` 不构成独立 L4 模块,理由见 §7.4。
**三个 reme 自治模块**:notifier(机械)、digester(LLM)、maintainer(LLM)。三者都由 scheduler 触发,都消费 L2 自治状态,但只有 notifier 是机械的 —— 候选选择不需要 LLM,LLM 决策在 agent 侧的 sync。
### 7.2 对称结构
```
跨表征层翻译 结构性纪律
(LLM-driven) (机械)
Inbound: ingester
Attention: notifier
Working: synchronizer
Sink: digester
Organization: maintainer (fold-only)
```
五类不同方向的"翻译":
| 模块 | 翻译方向 |
|---|---|
| ingester | 外部异构格式 → vault 统一文件 |
| notifier | L2 资源自治状态 → agent 注意力(`notify` 推送) |
| synchronizer | agent 事件流 → 工作过程叙事(写 hot) |
| digester | 工作过程 + 原始资料 → 长期知识(双源合流,写 cold) |
| maintainer | 散乱叶子 → 有层次的 topic tree(组织 cold) |
ingester 和 notifier 是机械(确定性阈值/流水线);其它三个是 LLM 决策模块,各自跨越一层语义鸿沟。
### 7.3 Maintainer:Topic Tree 密度折叠
`digest/` 整体视为一颗 **topic tree**:文件夹 = 中间节点,文件 = 叶子。Maintainer 唯一职责:随写入持续,某中间节点下叶子过密时,**折叠**为新的子中间节点 + 高密度摘要。
```
触发前:某中间节点叶子过多 / 太碎
digest/infra/
├── logging.md
├── tracing.md
├── metrics.md
├── alerting.md
├── dashboards.md
└── slo.md
折叠后:LLM 判断聚类,引入子中间节点 + 摘要
digest/infra/
├── observability/ ← 新中间节点
│ ├── _index.md ← 新生成的高密度摘要
│ ├── logging.md ← 内容不变,只搬位置
│ ├── tracing.md
│ ├── metrics.md
│ ├── alerting.md
│ ├── dashboards.md
│ └── slo.md
└── ...(未被折叠的叶子原位)
```
**设计承诺**(在 `maintain` 通用不变量之上进一步收紧):
| # | 承诺 | 含义 |
|---|---|---|
| **M-1** | **Fold-only**,无 merge / move / promote / demote / introduce | 树只向下生长,从不反向 |
| **M-2** | 叶子内容 0 修改,只被搬位置 | 与 `maintain` 内容守恒一致 |
| **M-3** | 新中间节点带一个高密度摘要文件,读摘要就能决定要不要深入 | 折叠后可读性不降反升 |
| **M-4** | 每次只处理一个候选节点 | 最小化变更面 |
| **M-5** | 不能聚类时,**不动**(默认保守) | 宁可不折,不要错折 |
LLM 唯一的决策点:
1. 这些叶子能不能聚类(if not → 不动)
2. 新中间节点叫什么、摘要怎么写
其它都机械:阈值判断(L1 file_watcher 派生 L2 自治状态提供信号)、移动文件(crud)、wikilink 重定向(graph/retarget)。
### 7.4 为什么没有 retriever 模块
L4 模块的存在条件 = "有跨原子编排 / 需要 LLM 决策"。Retrieve 不满足:
- 三种问法(state / semantic / topological)各自被 **L3 原子工具**直接覆盖(list+filter / search / traverse)
- 没有跨原子状态、没有 LLM 决策点
- Agent 直接调用 L3 原子即可
---
## 8. 跨切面:Schema
Schema(资料的 frontmatter / wikilink / 章节约定)是横跨三层、各写动作的共同契约。reme4 的核心立场:
| 立场 | 说明 |
|---|---|
| **reme 核心只保留 `name` / `description` 两个字段** | 其它都是 opinionated convention,服务消费层可以替换 |
| **Schema 是"协议"不是"代码"** | 用 markdown 文字描述,LLM agent 自我约束;不内嵌 schema validator |
| **三层共用同一套 wikilink 协议** | 全路径引用,无 short-link / no-ext 解析 |
### 8.1 协议文档(opinionated default)
| 内容 | 谁规定 |
|---|---|
| 目录结构(三层 + folder 单位) | 第 2 节本文档 |
| 动作语义契约(notify / sync / digest / maintain 的输入产出不变量) | 第 3 节本文档 |
| Frontmatter 推荐字段(4 轴等) | `reme4/steps/jobs/protocol.md` (opinionated) |
| 章节约定(Objective/Plan/Progress/...)| sync / digest 各自的 prompt(opinionated) |
### 8.2 重载入口
服务消费层(plugin / 自定义 caller)无需 fork reme,可通过以下方式替换 schema:
| 入口 | 适用场景 |
|---|---|
| 替换 protocol 文档 | 改 frontmatter / wikilink / 章节约定 |
| 替换 prompt 模板 | 改 sync / digest 的决策流程 |
| 替换 toolkit | 改 ReAct agent 可见的工具集 |
---
## 9. 反例:不属于本架构的设计
明确画出**不允许**的设计,免得后续讨论或扩展时滑回去:
| # | 反例 | 违反的不变量 |
|---|---|---|
| ✗-1 | Agent 通过任意 verb 直接写 digest | I-1(digest 写权只属 digester / maintainer) |
| ✗-2 | 多 agent 并发改同一个 daily folder | I-2(daily 单作者) |
| ✗-3 | 任何动作改写 resource 的原文 | I-3(resource immutable) |
| ✗-4 | 跨层引用引入第二套机制(hash-id / external ref / SQL) | I-4(wikilink 是唯一跨层载体) |
| ✗-5 | digester 改写 daily 正文 | `digest` 不变量(§3.5) |
| ✗-6 | maintain 改写 daily / resource 的语义内容 | `maintain` 不变量(§3.6) |
| ✗-7 | `notify` 维护 resource 上的 `referenced_by` 反指 | `notify` 完全单向(§3.3) |
| ✗-8 | 把 state / semantic / topological 合并成单一 read verb | R-1 |
| ✗-9 | Retrieve 自动 eager-expand provenance | R-4 |
| ✗-10 | digest / maintain 同步阻塞 agent 请求 | 5.4 节奏分级 |
| ✗-11 | digest / maintain 强一致(agent 写完 daily 立即可查 digest) | 5.5 eventual consistency |
| ✗-12 | maintainer 做 merge / move / promote / demote 等"通用重组" | M-1(fold-only) |
| ✗-13 | maintainer 改写既有叶子的内容(不只是搬位置) | M-2(叶子内容 0 修改) |
| ✗-14 | L4 模块在 frontmatter 里写 `status` / `pending` 等可派生状态字段 | 状态由 L1 file_watcher 派生到 L2 自治状态,L4 不重复 |
| ✗-15 | 为 retrieve 单设 L4 模块或聚合 verb | §7.4(L3 原子已足够) |
| ✗-16 | notify 写入 vault(在 resource 上加 `notified` frontmatter 或新建 daily 占位) | notify 只写 L2 推送队列,**不落任何文件**;ack 由 L1 watcher 检测 wikilink 派生 |
| ✗-17 | Service 自决推什么 notify 候选 | notify 决策在 Runtime(notifier);Service 只是 MCP transport,从 L2 推送队列读取(F-5) |
| ✗-18 | agent 主动调用 `notify` 想"标记这个 resource 我要看" | notify 是 reme→agent 单向,反向是 agent 用 sync 写 wikilink(自然 ack) |
---
## 附录:术语索引
| 术语 | 定义 |
|---|---|
| **resource/** | 不可变原始资料层 |
| **daily/** | agent 任务工作区层 |
| **digest/** | 沉淀知识层 |
| **State 问** | 在某层做 list + 过滤的状态查询 |
| **Semantic 问** | 跨层全文/向量检索 |
| **Topological 问** | 沿 wikilink 走的拓扑查询 |
| **Provenance** | 下游节点反查到上游来源的能力 |
| **Retarget** | 节点移动时对所有入向 wikilink 的原子重写 |
| **L5 Service** | 服务 agent 请求的进程(HTTP / MCP);执行栈最上层;也是 notify 的 MCP transport |
| **L5 Runtime** | 自治维护 vault 的进程;scheduler 在其中按 L2 自治状态阈值触发 background Action |
| **L4 Action** | 6 类动作语义:ingest / notify / sync / retrieve / digest / maintain |
| **L4 模块** | 实现 Action 的架构角色;五个:ingester / notifier / synchronizer / digester / maintainer(retrieve 不构成独立模块) |
| **ingester** | L4 模块,机械:外部源原样落 resource + 抽 frontmatter + 入索引 |
| **notifier** | L4 模块,机械:从 L2 资源自治状态选 notify 候选 → 写 L2 推送队列;Service MCP 拿走推给 agent |
| **synchronizer** | L4 模块,LLM-driven:agent 事件织入 daily 工作叙事;响应 notify 的也走这里 |
| **digester** | L4 模块,LLM-driven:resource + daily 双源合流为 digest 长期条目 |
| **maintainer** | L4 模块,LLM-driven,**fold-only**:digest topic tree 的密度折叠 |
| **scheduler** | L5 Runtime 内部触发器:按 cron + L2 自治状态阈值拉起 background L4 模块(notifier / digester / maintainer) |
| **Topic tree** | digest/ 的心智模型:文件夹 = 中间节点,文件 = 叶子 |
| **Fold(密度折叠)** | maintainer 唯一操作:把过密叶子归簇到新子中间节点 + 写高密度摘要 |
| **L3 原子工具** | 基础(create/append/edit/read/write/move/delete/list/stat,直 fs)+ 高级(search/traverse/frontmatter,走 L2)两组 |
| **L2 文件状态** | `file_store` + `file_graph` + 自治状态(resource 入流批次/orphan、daily 任务索引、digest 密度水位/断链)+ 推送队列;由 L1 派生(推送队列由 notifier 写) |
| **L2 推送队列** | `notify` 的 L2 状态条目;notifier 写 pending,Service MCP 推送后置 notified,L1 watcher 检测到 daily→resource wikilink 后置 acknowledged |
| **L1 file_watcher** | fs event → L2 state delta 的唯一派生桥;承担索引同步 + 自治状态派生 + `notify` ack 派生 |
| **L0 vault filesystem** | 物理目录:resource/ + daily/ + digest/;唯一真相源 |
| **Eventual consistency** | digest / maintain 异步处理,有可见延迟窗口;L0↔L2 之间 watcher 滞后窗口同理 |
| **Opinionated default** | reme 提供的参考实现,服务层可替换 |

View file

@ -0,0 +1,10 @@
"""Jobs steps — composite ReAct-agent-driven workflows.
Two steps:
digester — cold-write: distill daily notes into digest/ (R-M-W via a ReAct agent).
synchronizer — hot-write: persist in-progress task as a daily note.
"""
from . import digester # noqa: F401 -- @R.register("digester")
from . import synchronizer # noqa: F401 -- @R.register("synchronizer")

View file

@ -0,0 +1,300 @@
"""Smart Digester — knowledge distillation from daily notes to digest/.
The Digester is the **cold-write** counterpart to Synchronizer (hot-write).
It reads completed work in ``daily/<date>/<slug>.md`` note files,
identifies entities / concepts / claims / methods worth preserving
long-term, and sinks them into ``digest/`` as canonical-entry nodes so
the main agent can retrieve them later via search and graph traversal.
``digest/`` is the cold-tier root. Per ``protocol.md``, scope folders
under it may nest arbitrarily; each folder's canonical entry is
``<folder>/<folder>.md``, and slug (folder name) is globally unique across
the whole tree. Pending detection and graph machinery treat nodes at any
depth uniformly.
Drives a ReAct agent with a read/lookup/graph/write toolkit; the
agent follows the protocol in ``protocol.md`` (the opinionated
default schema + R-M-W decision tree). The schema is convention-driven
— reme core only reserves ``name`` / ``description``, so the agent
owns its own discipline rather than relying on a post-write linter.
Distillation state lives in the daily note's ``status`` frontmatter
— a **daily-tier convention owned by this digester** (reme core
reserves only ``name`` / ``description``; ``status`` is just an
extra). After processing each daily, the agent must call
``frontmatter_update`` with ``metadata={"status": "completed"}``
(or ``metadata={"status": "skipped"}`` when intentionally bypassed). Convention: absent
≡ ``pending``, so the next pass finds residual work via
``file_list path=daily recursive=true`` + per-item ``frontmatter_read`` to filter for absent ``status``.
Only the digester writes ``status``; Synchronizer / hand-edits must
leave it alone.
No degraded path — distillation strictly requires an LLM. When ``as_llm``
is unavailable, the step short-circuits with ``skipped=True`` and an
error message.
Override interface — schema is a service-consumption concern, not a
core invariant, so this step ships an **opinionated default** that any
caller can fully replace without touching reme4:
* ``protocol`` / ``protocol_path`` constructor args replace the
``protocol.md`` injected as ``{protocol}`` in the system prompt
(use when keeping the default prompt template but swapping schema).
* ``prompt_dict`` (inherited from ``BaseStep``) replaces the
``system_prompt`` / ``user_message`` templates wholesale (use when
the prompt structure itself needs to change).
* ``toolkit`` replaces the tool surface ``_DIGESTER_TOOLS`` builds.
Service layers (e.g. plugin-side configs) wire these in via component
config; ``digester.py`` / ``protocol.md`` shipped here are just a
reference implementation of one viable convention.
Toolkit. Each entry in ``_DIGESTER_TOOLS`` is a job name registered
in the active config; ``add_as_tool`` wraps ``job(**kwargs)`` into a
``ToolResponse``. The job indirection means the agent sees the same
tool surface (rich descriptions + JSON schema) as the L2 MCP layer.
"""
import datetime
import zoneinfo
from pathlib import Path
from agentscope.agent import ReActAgent
from agentscope.message import Msg
from agentscope.tool import Toolkit
from pydantic import BaseModel, Field
from ..base_step import BaseStep
from ...components import R
_DIGESTER_TOOLS: tuple[str, ...] = (
"file_list",
"file_read",
"file_stat",
"frontmatter_read",
"traverse",
"file_write",
"file_append",
"file_move",
"frontmatter_update",
"frontmatter_delete",
)
def _pack_daily(file_store, daily_path: str) -> str:
"""Render one daily note file's body into a prompt-friendly block.
A daily note is a single self-contained markdown file at
``daily/<date>/<slug>.md`` — everything the originating task wanted
the digester to see is inline (no sibling materials). External
assets land in ``resource/<date>/`` and are linked from the note's
``## References`` section; the LLM opens those on demand via
``file_read``.
"""
try:
absolute = (Path(file_store.vault_path or ".") / daily_path).resolve()
except Exception as e:
return f"### {daily_path}\n(error resolving path: {type(e).__name__}: {e})\n"
if not absolute.is_file():
return f"### {daily_path}\n(note file not found)\n"
parts: list[str] = [f"### {daily_path}"]
try:
parts.append(absolute.read_text(encoding="utf-8"))
except Exception as e:
parts.append(f"(error reading note: {type(e).__name__}: {e})")
return "\n".join(parts) + "\n"
class DistillResult(BaseModel):
"""Outcome of a single distillation call.
Without per-tool audit (the agent's toolkit is the job surface,
which doesn't expose per-call records back to the orchestrator),
the structured outcome is just the inputs the call was asked to
process plus the LLM's free-form summary. Per-file write outcomes
can be verified by re-reading vault_dir afterwards if needed.
Field semantics:
* ``daily_read`` — daily paths actually processed (input
order, deduped)
* ``summary`` — LLM's free-form one-paragraph summary
* ``skipped`` — True when no LLM available, no daily
paths provided, or the LLM reported ``SKIP``
* ``error`` — short error string when skipped due
to misconfiguration (e.g. no LLM)
"""
used_llm: bool = False
skipped: bool = False
daily_read: list[str] = Field(default_factory=list)
summary: str = ""
error: str = ""
@R.register("digester")
class Digester(BaseStep):
"""Knowledge digester: daily/ → digest/ via a ReAct agent.
Inputs (from RuntimeContext):
daily_paths (list[str], required): vault-relative paths to
daily note files (``daily/<date>/<slug>.md``) to distill.
Pass ``[]`` to no-op.
hint (str, optional): caller guidance to the LLM
(e.g. "focus on the auth-related decisions").
Output (written to context.response.answer):
DistillResult JSON — see model docstring.
"""
def __init__(
self,
toolkit: Toolkit | None = None,
console_enabled: bool = False,
timezone: str | None = None,
protocol: str | None = None,
protocol_path: str | None = None,
**kwargs,
):
"""Constructor overrides (service layer customization points):
* ``toolkit`` — replace the agent's tool surface; default builds
one from ``_DIGESTER_TOOLS``.
* ``protocol`` — inline protocol document (highest precedence);
overrides whatever the agent sees under ``{protocol}`` in the
system prompt.
* ``protocol_path`` — path (relative to the vault or absolute) to a protocol
markdown file; used when ``protocol`` is not given.
* ``prompt_dict`` (inherited via ``BaseStep``) — override the
``system_prompt`` / ``user_message`` templates wholesale, e.g.
to swap in a service-layer prompt that hardcodes a different
schema entirely.
With none of the above, falls back to the opinionated default
(the ``protocol.md`` and ``digester.yaml`` shipped alongside
this module).
"""
super().__init__(**kwargs)
self.toolkit = toolkit
self.console_enabled = console_enabled
self.timezone = timezone
self._protocol = self._load_protocol(protocol, protocol_path)
@staticmethod
def _load_protocol(protocol: str | None, protocol_path: str | None) -> str:
"""Resolve the protocol document; explicit string > path > default."""
if protocol is not None:
return protocol
if protocol_path:
path = Path(protocol_path)
if path.exists():
return path.read_text(encoding="utf-8")
default_path = Path(__file__).parent / "protocol.md"
return default_path.read_text(encoding="utf-8") if default_path.exists() else ""
def _now(self) -> datetime.datetime:
if self.timezone:
try:
return datetime.datetime.now(zoneinfo.ZoneInfo(self.timezone))
except Exception as e:
self.logger.error(f"Invalid timezone: {self.timezone}, error={e}")
return datetime.datetime.now()
def _vault_dir(self) -> Path:
vr = getattr(self.file_store, "vault_path", None)
return Path(vr).resolve() if vr else Path.cwd().resolve()
def _llm_available(self) -> bool:
"""Pre-flight check: can ``self.as_llm`` resolve without raising?
``BaseStep.as_llm`` asserts when no model is registered, so we
wrap the access here to avoid hard-failing at the call site."""
try:
return self.as_llm is not None
except Exception:
return False
def _build_toolkit(self) -> Toolkit:
"""Bind every digester-relevant job as a tool function."""
toolkit = self.toolkit or Toolkit()
for job_name in _DIGESTER_TOOLS:
self.add_as_tool(toolkit, job_name)
return toolkit
async def execute(self):
assert self.context is not None
daily_paths: list[str] = list(self.context.get("daily_paths") or [])
hint: str = (self.context.get("hint", "") or "").strip()
# No work to do: no daily paths supplied.
if not daily_paths:
result = DistillResult(used_llm=False, skipped=True)
self.context.response.success = True
self.context.response.answer = "Skipped: no daily paths supplied"
self.context.response.metadata.update(result.model_dump())
return
# No LLM available: distillation strictly requires one.
if not self._llm_available():
result = DistillResult(
used_llm=False,
skipped=True,
error="no as_llm configured; distillation requires an LLM",
)
self.context.response.success = False
self.context.response.answer = f"Error: {result.error}"
self.context.response.metadata.update(result.model_dump())
return
# Dedupe daily_paths while preserving order.
seen: set[str] = set()
deduped: list[str] = []
for p in daily_paths:
if p and p not in seen:
seen.add(p)
deduped.append(p)
daily_paths = deduped
# Build the per-daily blob the agent will see.
daily_blob = "\n\n".join(_pack_daily(self.file_store, p) for p in daily_paths)
vault_dir = self._vault_dir()
toolkit = self._build_toolkit()
agent = ReActAgent(
name="reme_digester",
model=self.as_llm,
sys_prompt=self.prompt_format(
"system_prompt",
vault_dir=str(vault_dir),
protocol=self._protocol,
),
formatter=self.as_llm_formatter,
toolkit=toolkit,
)
agent.set_console_output_enabled(self.console_enabled)
user_message: str = self.prompt_format(
"user_message",
today=self._now().strftime("%Y-%m-%d"),
hint=hint or "(none)",
daily_blob=daily_blob or "(none)",
)
final_msg: Msg = await agent.reply(
Msg(name="reme", role="user", content=user_message),
)
summary = (final_msg.get_text_content() or "").strip()
result = DistillResult(
used_llm=True,
daily_read=list(daily_paths),
summary=summary,
skipped=summary.upper().startswith("SKIP"),
)
self.context.response.success = True
self.context.response.answer = summary or "Distillation completed"
self.context.response.metadata.update(result.model_dump())

View file

@ -0,0 +1,114 @@
system_prompt: |
You are the digester. You read completed work in `daily/<date>/<slug>.md`
note files and lift entities, concepts, claims, and methods into
canonical entries under `digest/<…>/<slug>/<slug>.md`. The main agent
retrieves them later via search and graph traversal.
vault_dir: {vault_dir}
## Five steps per call
### Step 1 — Read each daily
For every daily note passed in, the full body is already packed
in the user message below. Each note is a single self-contained
markdown file — anything the originating task wanted you to see is
inline. If a note's `## References` section points at
`[[resource/<date>/<name>]]` items and you need them, open them via
`file_read` on demand.
### Step 2 — Lookup candidates
Identify each entity / concept / claim / method named in the daily.
Before deciding to CREATE anything, look it up in `digest/`:
- `file_list path=digest/ recursive=true` to scan the tree, or
- `graph_traverse path=<candidate>.md depth=1` for neighborhood.
**Slugs are globally unique under `digest/`** — a prior occurrence
at any nesting depth means the node already exists. Find and reuse
it; do not CREATE a duplicate under a new path. This is the worst
failure mode of the digester.
### Step 3 — R-M-W decision
Per candidate, pick exactly one branch. Every branch writes only
the node being authored in this step — never sideways into other
nodes' bodies. A relation worth recording is captured by a typed
wikilink in the source body; the inbound view is queried later via
`graph_traverse direction=in`.
- **CREATE** — no hit anywhere under `digest/`:
`file_write digest/<…>/<slug>/<slug>.md`. Place it under the
closest existing semantic parent scope; top-level if none applies.
- **UPDATE** — exact match exists:
Merge new facts into the right section. Prefer
`frontmatter_update` for metadata; `file_append` for purely
additive trailing sections; `file_read` + `file_write` for
mid-body edits. Override stale claims rather than stacking
contradictions. Don't restructure unrelated parts.
- **MOVE / promote** — rename or relocate an existing node:
`file_move` (the retarget pass rewrites inbound wikilinks
atomically; never `file_write` to new + `file_delete` old).
When CREATE / UPDATE writes a relation into the body, prefer a
typed wikilink (`predicate:: [[X]]` or `[predicate:: [[X]]]`) when
the relation has clear semantic weight; default to bare `[[X]]` for
plain mention. See protocol.md §4.1 for the recommended predicate
vocabulary.
If a candidate is only a passing mention with no new fact to write,
do nothing — the mention stays in the daily, search will still
find it, and a future digester pass can lift it when it accumulates
substance worth CREATE/UPDATE.
### Step 4 — Flip status per daily (mandatory)
After processing each daily, call:
frontmatter_update path=daily/<date>/<slug>.md metadata={status: completed}
Use `status=skipped` if the daily had nothing worth lifting (chitchat,
dead end). This flip is the daily-tier convention this digester owns
— absent ≡ pending, so the next pass uses
`file_list path=daily recursive=true` + per-item `frontmatter_read`
to find leftover work. Forgetting the flip leaves the daily in the
queue forever.
### Step 5 — Reply
One short paragraph: which dailies you read, what you CREATEd /
UPDATEd / MOVEd (with paths), and which dailies you flipped to
`completed` vs `skipped`. Don't replay every tool call — those are
on the audit trail.
## Wikilinks
Always full path relative to the vault with `.md`:
`[[digest/<…>/<slug>/<slug>.md]]` for digest entries,
`[[daily/<date>/<slug>.md]]` for daily notes,
`[[resource/<date>/<name>]]` for ingested resources. Short forms or
extension-less forms don't resolve.
## Boundaries
- **Never write under `daily/`** except for the Step 4 `status` flip.
- **Only this digester writes `status`** — sync / hand-edits leave
it alone (absent ≡ pending is exactly how this digester finds its
workload).
## Memory protocol
{protocol}
user_message: |
today: {today}
hint: {hint}
# Daily notes to distill
{daily_blob}
Run the five steps from the system prompt. Reply = one paragraph audit.

View file

@ -0,0 +1,112 @@
# Memory Protocol
Opinionated default contract for writing memory into a reme vault.
reme core reserves only `name` / `description` (both optional);
everything below is convention that consumers may replace.
## 1. Directory architecture
```
<vault>/
├── daily/
│ └── <YYYY-MM-DD>/
│ └── <slug>/
│ ├── <slug>.md # hot summary note (markdown)
│ └── <material>.* # sibling materials (any file type)
└── digest/
└── <slug>/
├── <slug>.md # cold canonical entry
├── <material>.md # supporting docs
└── <subslug>/ # nested narrower scope
└── <subslug>.md
```
- **Hot tier** (`daily/…`) — streaming. One upstream writer per
folder; every other consumer treats it as read-only. The
summary note `<slug>.md` is markdown; siblings may be any
file type the writer chooses.
- **Cold tier** (`digest/…`) — curated. Each folder is a
scope and must contain `<folder>/<folder>.md` as its canonical
entry. A scope's other children are sibling material files or
narrower scope subfolders; nesting depth is unconstrained.
Slugs are globally unique — a folder name appears at most once
anywhere under `digest/`.
Facts flow one-way, `daily/` → `digest/`. References use the
full path relative to the vault: `[[digest/<slug>/<subslug>/<subslug>.md]]`.
## 2. Frontmatter
Reserved (typed; all optional):
| key | type |
|---|---|
| `name` | string |
| `description` | string |
Opinionated default axes (closed enums):
| key | values |
|---|---|
| `lifecycle` | `streaming` / `evolving` / `frozen` |
| `scope` | `instance` / `class` |
| `source` | `auto` / `curated` / `derived` |
| `role` | `profile` / `concept` / `claim` / `method` / `reference` / `observation` / `question` / `fundamentals` |
Any other keys consumers want (e.g. a workflow `status` flag) live
as extras — write them, read them with the `where` filter on `list`
tools (`null` matches absent-or-null); the protocol does not name
or enumerate them.
## 3. Body
Section structure is **advisory** — `## Summary`, `## Key Facts`,
`## Decisions`, `## Related` are convenient defaults but the
protocol mandates no specific section.
## 4. Wikilinks
Three recognized forms:
| Form | Example | Meaning |
|---|---|---|
| Bare | `See [[张三.md]]` | weakest layer — "mention" |
| Line-level Dataview | `colleague:: [[李四.md]]` | typed relation, queryable by predicate |
| Inline-bracketed Dataview | `主导 [负责:: [[项目X.md]]] 的重构` | typed relation, embedded inline |
Targets are stored **verbatim** as full paths relative to the vault.
`[[digest/zhang-san/zhang-san.md]]` resolves; short or
extension-less forms do not — no implicit `.md` completion, no
basename search, no folder-note expansion.
Renaming a node requires atomically rewriting every inbound
wikilink.
### 4.1 Typed predicates (half-open)
Recommended core vocabulary — writers may extend beyond this set,
but new predicates should be reused consistently:
| predicate | meaning |
|---|---|
| `is_a` | hierarchical (X is a kind of Y) |
| `part_of` | containment (X is part of Y) |
| `depends_on` | dependency (X requires Y) |
| `manages` | authority / responsibility |
| `alias_of` | equivalence (X and Y are the same thing) |
| `references` | citation / external pointer |
Use typed `predicate:: [[X]]` only when the relation has clear
semantic weight; default to bare `[[X]]` for plain mentions. Typed
edges become queryable via `graph_traverse predicate=<name>`.
### 4.2 One-way write rule
A wikilink lives in the **source** node's body only — the node whose
prose introduces the relation. The **target** is never modified to
record the inbound relation. Backlinks are discovered at query time
via `graph_traverse direction=in`, never written into target bodies.
This keeps every write authoritative: a node's body reflects only
what its own author/writer chose to say, never sideways annotations
from other nodes' writers.

View file

@ -0,0 +1,276 @@
"""Synchronizer — daily-note sync ReAct agent.
Watches the agent's recent conversation and persists in-progress
tasks as a daily note inside reme's vault_dir, so future agent
invocations can pick the work back up. Pure sync mechanism — does
**not** compress the agent's context (compression is the agent's
own concern).
Note layout: a single markdown file ``daily/<YYYY-MM-DD>/<slug>.md``.
Everything worth preserving (verbatim user prompt, key tool output,
intermediate data) goes inline inside this file — there are no
sibling materials. References to it use the full path relative to
the vault (``[[daily/<date>/<slug>.md]]``); short or no-extension
forms do not resolve. Frontmatter carries ``name`` / ``description``
plus an optional ``inherits`` wikilink for cross-day continuation.
Body splits into ``Objective`` / ``Plan`` / ``Progress`` /
``Findings`` / ``Decisions`` / ``Next`` / ``References`` sections
(the last is a list of ``[[resource/<date>/<name>]]`` wikilinks for
inbound assets landed via ``ingest`` — non-markdown
artefacts the task itself produced are summarized inline).
Inputs (from RuntimeContext):
messages (list[Msg], required): conversation slice to inspect.
note (str, optional): caller-supplied note hint (task name
or ``daily/<date>/<slug>.md`` path) to bias slug selection
and disambiguate same-day tasks.
Output (written to context.response.answer):
{
"skipped": True if the agent reported [SKIP],
"actions": one-line action statement from the agent,
"note": note file path relative to the vault, or None,
"summary": full markdown content of the synced note,
for the calling agent to reload into a
compacted context. None when SKIP / failed.
}
The agent's toolkit is assembled by ``add_as_tool`` — each entry in
``_NOTE_TOOLS`` is a job name (registered in the active config);
the wrapper turns ``job(**kwargs)`` into a ``ToolResponse``. The job
indirection means the agent sees exactly the same tool surface
(name / description / parameter schema) as the L2 MCP layer.
Override interface — note shape (single-file layout, frontmatter
fields, section discipline) is opinionated convention, not a core
invariant, so callers can fully replace it without touching reme4:
* ``prompt_dict`` (inherited from ``BaseStep``) overrides the
``system_prompt`` / ``user_message`` templates wholesale — this is
how a service layer swaps in its own note schema (e.g. a different
section list, different frontmatter fields).
* ``toolkit`` replaces the tool surface ``_NOTE_TOOLS`` builds.
What ships here (``synchronizer.yaml``) is just one viable convention;
the service layer (plugin configs, custom callers) is the right place
to pin down the *deployment-specific* shape.
"""
import datetime
import re
import zoneinfo
from pathlib import Path
from agentscope.agent import ReActAgent
from agentscope.message import Msg
from agentscope.tool import Toolkit
from pydantic import BaseModel, Field
from ..base_step import BaseStep
from ...components import R
_NOTE_PATH_RE = re.compile(r"daily/\d{4}-\d{2}-\d{2}/[^/\s]+\.md")
_NOTE_TOOLS: tuple[str, ...] = (
"file_list",
"file_read",
"file_write",
"file_append",
"file_edit",
"file_stat",
"frontmatter_read",
"frontmatter_update",
"frontmatter_delete",
"daily_read",
"daily_write",
"daily_reindex",
)
def _coerce_messages(raw) -> list[Msg]:
"""Normalize incoming messages to ``Msg`` instances.
The Python caller hands in ``list[Msg]`` directly; the MCP layer
delivers ``list[dict]`` (each dict shaped roughly ``{name?, role?,
content?}``). Both shapes land here.
"""
if not raw:
return []
out: list[Msg] = []
for item in raw:
if isinstance(item, Msg):
out.append(item)
continue
if isinstance(item, dict):
out.append(
Msg(
name=item.get("name") or item.get("role") or "user",
role=item.get("role") or "user",
content=item.get("content", ""),
),
)
return out
def _format_history(messages: list[Msg]) -> str:
"""Render the conversation as a speaker-tagged transcript.
Skips messages whose text content is empty (tool-only frames
don't help the LLM judge task state).
"""
if not messages:
return "(empty)"
lines: list[str] = []
for msg in messages:
speaker = msg.name or msg.role or "?"
text = (msg.get_text_content() or "").strip()
if not text:
continue
lines.append(f"[{speaker}]\n{text}")
return "\n\n".join(lines) or "(no text)"
class SynchronizerResult(BaseModel):
"""Outcome of a single note-sync call.
Without per-tool audit (the agent's toolkit is the job surface,
which doesn't expose per-call records back to the orchestrator),
the structured outcome is just what the agent reports plus what
we re-read from disk after it returns.
"""
used_llm: bool = Field(default=False)
skipped: bool = Field(default=False)
actions: str = Field(
default="",
description="One-line action statement from the agent (e.g. "
"'updated daily/2026-05-15/auth-refactor.md' or '[SKIP]').",
)
note: str | None = Field(
default=None,
description="Note file path relative to the vault, "
"e.g. 'daily/2026-05-15/auth-refactor.md'. None when SKIP / failed.",
)
summary: str | None = Field(
default=None,
description="Full markdown content of the note file. Lets the calling "
"agent reload the warm summary into a freshly compacted context without an extra read.",
)
@R.register("synchronizer")
class Synchronizer(BaseStep):
"""Drive daily-note sync via a ReAct agent."""
def __init__(
self,
toolkit: Toolkit | None = None,
console_enabled: bool = False,
timezone: str | None = None,
inherit_window_days: int = 7,
**kwargs,
):
super().__init__(**kwargs)
self.toolkit = toolkit
self.console_enabled = console_enabled
self.timezone = timezone
self.inherit_window_days = inherit_window_days
def _now(self) -> datetime.datetime:
if self.timezone:
try:
return datetime.datetime.now(zoneinfo.ZoneInfo(self.timezone))
except Exception as e:
self.logger.error(
f"Invalid timezone {self.timezone!r}, falling back to local time: {e}",
)
return datetime.datetime.now()
def _vault_dir(self) -> Path:
wd = getattr(self.file_store, "vault_path", None)
return Path(wd).resolve() if wd else Path.cwd().resolve()
def _build_toolkit(self) -> Toolkit:
"""Bind every note-relevant job as a tool function.
Each entry in ``_NOTE_TOOLS`` is a job name registered in the
active config; ``add_as_tool`` wraps ``job(**kwargs)`` into a
``ToolResponse``. The job indirection means the agent sees exactly
the same tool surface (name / description / parameter schema) as
the L2 MCP layer.
"""
toolkit = self.toolkit or Toolkit()
for job_name in _NOTE_TOOLS:
self.add_as_tool(toolkit, job_name)
return toolkit
async def execute(self):
assert self.context is not None
messages: list[Msg] = _coerce_messages(self.context.get("messages"))
note_hint: str = self.context.get("note", "") or ""
if not messages:
result = SynchronizerResult(used_llm=False, skipped=True)
self.context.response.success = True
self.context.response.answer = "Skipped: no messages supplied"
self.context.response.metadata.update(result.model_dump())
return
toolkit = self._build_toolkit()
agent = ReActAgent(
name="reme_synchronizer",
model=self.as_llm,
sys_prompt=self.prompt_format("system_prompt"),
formatter=self.as_llm_formatter,
toolkit=toolkit,
)
agent.set_console_output_enabled(self.console_enabled)
user_message: str = self.prompt_format(
"user_message",
today=self._now().strftime("%Y-%m-%d"),
vault_dir=str(self._vault_dir()),
inherit_window_days=self.inherit_window_days,
note=note_hint or "(none)",
history=_format_history(messages),
)
final_msg: Msg = await agent.reply(
Msg(name="reme", role="user", content=user_message),
)
actions = (final_msg.get_text_content() or "").strip()
result = SynchronizerResult(used_llm=True, actions=actions)
if "[SKIP]" in actions.upper():
result.skipped = True
# Reload the freshly written note so the calling agent can drop
# it back into a compacted context without an extra read trip.
if not result.skipped:
self._reload_note(result, actions)
self.context.response.success = True
self.context.response.answer = actions or "Synchronization completed"
self.context.response.metadata.update(result.model_dump())
def _reload_note(self, result: SynchronizerResult, actions: str) -> None:
"""Parse the agent's action line for the note path and read the
full file back into ``result.summary``. Best-effort: a parse miss
leaves the context-management fields as None but does not fail
the step (persistence already succeeded)."""
match = _NOTE_PATH_RE.search(actions)
if not match:
return
note_path = match.group(0)
try:
absolute = (Path(self.file_store.vault_path or ".") / note_path).resolve()
text = absolute.read_text(encoding="utf-8")
except Exception as e:
self.logger.warning(f"synchronizer: could not reload note {note_path!r}: {e}")
return
result.note = note_path
result.summary = text

View file

@ -0,0 +1,144 @@
system_prompt: |
You persist an in-progress task as a daily note so the next session
can pick it up. One call writes at most one note.
## Note shape
<vault>/daily/<YYYY-MM-DD>/<slug>.md # the whole note, one file
One note per task per day; cross-day continuation uses an
`inherits:` frontmatter wikilink. Anything worth preserving
(verbatim user prompt, key tool output, intermediate data) goes
inline inside this file — there are no sibling materials.
## Five steps per call
### Step 1 — Skip check
Is the conversation a real multi-step task in progress? Casual Q&A, a
single-shot answer, or idle chat → reply `[SKIP]` (alone, literal) and
stop. When genuinely ambiguous, default to writing — losing work is
worse than an extra note.
### Step 2 — Pick the slug
- Note hint already a kebab-case slug → use verbatim.
- Otherwise mint a stable kebab-case from the task topic (≤60 chars).
- **Reuse the same slug across calls for the same logical thread** —
same slug = same file = upsert. Fragmenting one thread across
multiple slugs is the worst failure mode.
### Step 3 — Discover the branch
Call `file_list path=daily/{today}` (and `daily_read` /
`frontmatter_read` on candidates as needed) to determine one of three
branches:
- **UPDATE** — `daily/{today}/<slug>.md` already exists:
`daily_read slug=<slug>` to fetch body + frontmatter, merge per
Step 4, then `daily_write slug=<slug> overwrite=true` with the
full new body + frontmatter.
- **INHERIT** — no file today, but within the last
{inherit_window_days} days an active note with the same name
exists: confirm via `daily_read` on the earlier note, then
`daily_write slug=<slug>` (default `overwrite=false`) with a
fresh body that sets `inherits: [[daily/<earlier-date>/<earlier-slug>.md]]`
in frontmatter and copies the predecessor's `Objective` + `Plan`.
Progress / Findings / Decisions / Next start empty. **Do not
modify the predecessor.**
- **CREATE** — neither: `daily_write slug=<slug>` (default
`overwrite=false`) with the full body + frontmatter. Idempotent —
no-ops if a same-slug note already exists today (caller falls
back to the UPDATE branch).
### Step 4 — Write
Pick the smallest write for each change:
| Change | Tool |
|---|---|
| New / replacement full note | `daily_write` (one shot: body + frontmatter + index refresh) |
| Append to a trailing append-only section | `file_append` — cheaper than R-M-W |
| One frontmatter key | `frontmatter_update` (call `daily_reindex` after if `name`/`description` changed) |
| Mid-body restructure | `daily_read` + `daily_write overwrite=true` |
Suggested body shape (sections are convention — adapt as fits):
```markdown
---
name: <task name>
description: <2-3 sentences: what + why>
inherits: [[daily/<earlier-date>/<earlier-slug>.md]] # INHERIT only
---
## Objective
<set-once long-term goal>
## Plan
<current approach — rewrite wholesale on update>
## Progress
- <YYYY-MM-DD HH:MM> <entry> <!-- append-only -->
## Findings
- <key fact / conclusion> <!-- append-only -->
## Decisions
- <choice + why> <!-- append-only -->
## Next
- [ ] <todo> <!-- rewrite wholesale -->
## References
- [[resource/<date>/<name>]] — <one line> # external assets the task consumed
```
Section discipline: Progress / Findings / Decisions are append-only
(never delete history). Plan / Next are wholesale-rewritten each call.
Objective is set once. `## References` lists `[[resource/<date>/<name>]]`
wikilinks for inbound assets that arrived through an external channel
and were landed by `ingest`; non-markdown content the task
itself produced (raw outputs, screenshots) should be summarized inline
or skipped — there is no per-note sibling folder anymore.
Wikilink form is fixed: note refs use the full vault-relative
path with `.md` (`[[daily/<date>/<slug>.md]]`); resource refs use the
canonical resource path (`[[resource/<date>/<name>]]`). Short forms
don't resolve.
### Step 5 — Emit one line
Your final reply must be exactly one line in this form:
<action> daily/<date>/<slug>.md
- `<action>` ∈ `created` / `inherited` / `updated`
Or `[SKIP]` (literal, alone) if Step 1 said skip. This line is parsed
mechanically — match the format exactly.
## Boundaries
- **Never write under `digest/`** — that tier is downstream.
- **Never write the `status` frontmatter** — `status` is reserved for
the downstream distillation pass, which uses absence to find pending
work. Touching it from here makes the note invisible to the
next distill run.
- **`daily_write` defaults to `overwrite=false`** (idempotent skip-if-exists,
mirroring the old `daily_resolve` probe). Pass `overwrite=true` only when
you've already read the file via `daily_read` and intend to replace it
(the UPDATE branch). On a surprise collision (`created: false`) — fall
back to UPDATE rather than blindly overwriting.
user_message: |
Today: {today}
Vault dir: {vault_dir}
Inherit window: last {inherit_window_days} days
Note hint: {note}
# Recent conversation
{history}
Run the five steps from the system prompt. Final reply = one line.