Compare commits

...
28 Commits
Author SHA1 Message Date
wuyouMaster 99c6e923bc feat: read-only roominfo (owner/announcement/openim) on GroupSnapshot 2026-09-25 00:06:48 +08:00
wuyouMaster afb876250e feat: merge favorites into export and search_messages 2026-09-24 23:55:52 +08:00
wuyouMaster 19bf497af0 feat: wire favorites read path, AMR ffmpeg decode, voice cache/VoiceTemp walk 2026-09-24 21:27:48 +08:00
wuyouMaster 784bed0ee0 feat: align favorite numeric types with real fav_db_item.type 2026-09-24 21:09:13 +08:00
wuyouMaster 52c56bfc98 feat: map MM_FAV types to message cards and voice disk bypass V2 2026-09-24 20:44:19 +08:00
wuyouMaster e69c3626d6 feat: parse streamvideo/gift cards and heuristic local_type 2/8 2026-09-24 20:06:16 +08:00
wuyouMaster 7864cf001f feat: parse music/subscribe/kefu appmsg cards into read-only share/system 2026-09-24 19:54:24 +08:00
wuyouMaster 264d25b539 feat: parse patMsg/findernamecard/product/friend-verify and unpack local_type 2026-09-24 19:46:40 +08:00
wuyouMaster 62c1b0063f Merge origin/main into local payment export work 2026-09-24 19:25:56 +08:00
wuyouMaster 4aef093802 feat: parse trans_id/fee_type/refund_bank_type into transfer pay info 2026-09-24 19:22:09 +08:00
wuyouMaster 88ee3f19bb feat: map transfer/hb status labels into export and chat cards 2026-09-24 19:01:15 +08:00
qingmao 8503c92f3a Add files via upload 2026-09-22 09:49:09 +08:00
wuyouMaster 95b3ba6121 fix: auto retry scheduled reports after send recovery 2026-09-20 20:40:10 +08:00
wuyouMaster b6d287a299 fix: make packaging checks platform safe 2026-09-20 18:16:41 +08:00
wuyouMaster 4c25236070 fix: ad-hoc sign packaged macOS key helpers and app bundle
electron-builder 26 skips macOS signing entirely when no Developer ID identity is configured, so the packaged key helpers can ship unsigned or with a broken signature and macOS kills them even with SIP disabled. afterPack now verifies and ad-hoc re-signs the packaged xkey helpers and the outer app bundle, failing the build when a signature cannot be repaired.
2026-09-20 18:16:41 +08:00
qingmao c1224cb612 Merge pull request #45 from mmhh256/fix/query-agent-system-messages
fix: 合并 Query Agent 的连续 system messages
2026-09-20 16:11:08 +08:00
Wxw-Gu 5036639aa1 test: 测试用例 2026-09-20 15:56:27 +08:00
Wxw-Gu 9593ca0f54 fix: 修复语音消息取错、时长精度与转写不刷新
- 时长解析:改取 <voicemsg voicelength>(毫秒),不再误取 length(SILK 字节数)
- 时长显示:四舍五入对齐微信口径,有原生时长时不被解码时长覆盖
- 转写:重新识别改为强制重算,跳过身份级缓存短路(音频级缓存仍生效)
- 日报:语音累计秒数显示取整,不再出现小数
2026-09-20 15:40:01 +08:00
mmhh256 b441287d54 fix: 合并 Query Agent 的连续 system messages 2026-09-19 14:01:39 +08:00
Wxw-Gu 9b82d29037 fix: 兼容微信新版系统消息模板格式,入群通知不再显示成「撤销」 2026-09-18 14:41:50 +08:00
Wxw-Gu 0f270366aa feat: 图片文字索引按时间分段优先处理最近图片
- 首次索引先处理最近 7 天,再依次回溯 30 天 / 近一年 / 更早历史
 - 完成后新到的图片单独补齐,不受历史回填影响
 - 覆盖度增加时间维度,可区分「最近已完整」与「更早仍在补齐」
2026-09-18 14:39:05 +08:00
Wxw-Gu 12b8fb34c6 fix: 增加sip文案 2026-09-18 09:38:33 +08:00
Wxw-Gu 774346d503 fix: 日报头像补齐下载重试 2026-09-17 19:24:27 +08:00
Wxw-Gu b97e1f4bc6 fix: 问问微信恢复引用编号,日报修复成员头像与 hero 溢出
- 问问微信:Host 分配稳定 citationId 并注入模型上下文,正文 [E#] 与证据按钮同号
- 问问微信:回答返回前校验引用,幻觉编号移除并提示,不再出现不可信的引用按钮
- Agent Hub:收紧「最近会话」快捷路由,内容查询交回 Query Agent
2026-09-17 18:56:32 +08:00
Wxw-Gu 4b2395a8d2 feat: 重写agent hub连接器,新增对话记录和微信输入状态
- 新增 Agent Hub 对话记录面板,按会话回看机器人与微信用户的完整收发内容
- 新增微信原生正在输入状态,长任务维持 typing,异常路径强制收尾
2026-09-17 17:32:37 +08:00
Wxw-Gu 1337bcb7de feat: 新增mac ocr转文字,增加到问一问微信 图片索引优化速度
- 图片文字索引性能与进度诚实化
- 问问微信:证据卡区分「消息类型」与「派生来源」,派生命中内容自报来源
- 问问微信:回答规则禁止未真实执行的多轮承诺
- 本地图片文字识别:支持 macOS 系统 OCR(Apple Vision)
2026-09-17 13:11:20 +08:00
电摇小子 b8f08d54c8 feat: 新增微信图片文字索引与问问微信图片检索能力 2026-09-16 10:42:52 +08:00
电摇小子 24399f1d70 feat: 新增 Windows 本地图片文字识别能力 2026-09-16 00:23:23 +08:00
213 changed files with 29292 additions and 3453 deletions
-8
View File
@@ -25,11 +25,6 @@ jobs:
node-version: 22
cache: pnpm
- uses: actions/setup-go@v5
with:
go-version-file: services/wechat-connector/go.mod
cache-dependency-path: services/wechat-connector/go.sum
- name: Install dependencies
run: pnpm install
@@ -51,9 +46,6 @@ jobs:
- name: Skill installation instruction tests
run: pnpm test:skill-install
- name: WeChat connector tests
run: pnpm test:wechat-connector
- name: Build Electron test application
run: pnpm test:e2e:build
-1
View File
@@ -10,7 +10,6 @@ out
coverage/
playwright-report/
test-results/
resources/connectors/wechat/
resources/connectors/wechat-personal/
.omc
.codex/
+2
View File
@@ -49,6 +49,8 @@ Agent Hub 让微信机器人调用本机 TraceMemo;Reader Skill / Local HTTP A
## 开发文档
- [开发、测试与构建](./development/overview.md)
- [界面开发规范:按钮与主题色](./development/ui-guidelines.md)
- [微信系统消息解析与格式兼容](./development/wechat-system-message-parsing.md)
- [Query Agent POC(开发测试入口)](./development/query-agent-poc.md)
- [本地启动排障](./development/local-startup-troubleshooting.md)
- [macOS 数据访问说明](./platform/macos.md)
@@ -6,7 +6,6 @@
执行 `pnpm dev` 后,以下状态同时满足,说明本地开发环境已经可用:
- 控制台显示连接器已生成,例如 `resources/connectors/wechat/win32-x64/wechat-connector.exe`;
- Electron 窗口已打开,或 `http://localhost:5173/` 返回 HTTP `200`;
- 控制台显示 Local HTTP API 正在监听 `http://127.0.0.1:6131`。
@@ -22,18 +21,6 @@ WCDB_DEBUG_LOGS=1 pnpm dev
开启后会输出 `GETMSG-xxx` 请求耗时和 native `WCDB-EXPLAIN` 执行计划,不记录聊天正文。取消该环境变量或设为 `0` 即可关闭。
## Go 命令找不到
如果 `pnpm dev` 在构建微信连接器时出现 `spawnSync go ENOENT`,先执行:
```bash
go version
```
命令不可用表示当前终端的 `PATH` 没有找到 Go。Windows 默认安装位置是 `C:\Program Files\Go\bin`。确认 Go 已安装并把该目录加入系统 `PATH` 后,关闭并重新打开终端或 IDE,再重新执行 `go version` 和 `pnpm dev`。
如果 Go 刚完成安装,已经打开的终端不会自动继承新的环境变量;重开终端是必要步骤。不要绕过连接器构建直接启动 `electron-vite dev`,否则 Agent Hub 的微信连接器不会生成。
## Electron 二进制缺失或下载失败
`electron-vite dev` 报 `Electron uninstall`,或 Electron 安装器报 `fetch failed`,通常表示 `node_modules/electron/dist` 中的 Electron 二进制缺失或下载未完成。这不是应用业务代码的启动错误。
@@ -68,4 +55,4 @@ Vite 在某些 Windows 环境中只监听 IPv6 本机回环地址 `::1`。这时
## 仍无法启动时
保留首次错误的完整输出,并同时记录操作系统、Node.js、pnpm 和 Go 版本,以及 `pnpm install --frozen-lockfile` 与 `pnpm dev` 的执行结果。不要提交数据库密钥、AI API Key、微信数据路径或聊天内容。
保留首次错误的完整输出,并同时记录操作系统、Node.js 与 pnpm 版本,以及 `pnpm install --frozen-lockfile` 与 `pnpm dev` 的执行结果。不要提交数据库密钥、AI API Key、微信数据路径或聊天内容。
+2 -3
View File
@@ -6,7 +6,6 @@
- Electron + React + TypeScript;
- pnpm 7+;
- Go(构建微信连接器);
- 平台对应的 Electron/native 构建环境。
产品文档的事实来源优先级是:当前源码 → 当前 UI/Renderer → 测试 → package/config → README/docs → 历史资料。功能、API、版本、隐私和兼容性变更时,不要只改 README。
@@ -18,7 +17,7 @@ pnpm install
pnpm dev
```
本地依赖安装、Go 环境和 Electron 二进制下载异常,请查看[本地启动排障](./local-startup-troubleshooting.md)。
本地依赖安装与 Electron 二进制下载异常,请查看[本地启动排障](./local-startup-troubleshooting.md)。
常用检查:
@@ -30,7 +29,7 @@ pnpm test:integration
pnpm test:e2e:build
```
完整测试入口 `pnpm test` 还会运行 Skill 安装指令、微信连接器、构建和 Playwright 测试;需要对应平台环境。
完整测试入口 `pnpm test` 还会运行 Skill 安装指令、构建和 Playwright 测试;需要对应平台环境。
## 代码变更对应文档
+171
View File
@@ -0,0 +1,171 @@
# 界面开发规范:按钮与主题色
这份规范回答一件事:**为什么同一个产品里,有的按钮是主题色,有的还是浏览器默认的黑白方角。**
先看一个真实案例 —— 同一屏里的两组按钮:
```
主界面:「更新图片文字索引」 ← 主题色(正确)
弹窗里:「取消」「开始索引」 ← 浏览器默认样式(错误)
```
两者渲染出来完全不同,用户会以为是两个不同的产品。根因不是"设计没定颜色",
而是**组件在导出时把样式丢了**。下面写清楚怎么避免。
---
## 1. 永远不要写裸 `<button>`
任何可点的按钮都必须来自 `components/ui/button`:
```tsx
import { Button } from '../ui'
<Button variant="outline" onClick={handleCancel}>取消</Button>
```
**唯一的例外**:结构性控件(导航项、Tab、列表行、图标热区)——它们有自己
成套的布局样式,用原生 `<button>` 是合理的,但**必须**带 `className`,
且样式写在对应的 `.scss` 里,不要在 JSX 里临时拼颜色。
```tsx
// 可以:结构性控件,样式来自 .scss
<button type="button" role="tab" className={active ? 'active' : ''} onClick={...}>
今日日报
</button>
```
**判据**:如果这个按钮在别的界面也会以同样形态出现("取消"、"保存"、"删除"),
它就该是 `Button`;如果它只在某一个位置有意义(侧栏导航项),才考虑原生。
---
## 2. 三种角色,只有三个默认变体
`Button` 提供 6 个变体,但**日常只用其中 3 个**:
| 角色 | `variant` | 长什么样 | 用在哪 |
| --- | --- | --- | --- |
| 主要 | `default` | 主题色实底 | 这一步用户唯一该做的事 |
| 次要 | `outline` / `ghost` | 描边 / 无底色 | 取消、返回、并列的辅助操作 |
| 危险 | `destructive` | 红色实底 | 删除、清空、不可恢复的操作 |
另外两个(`secondary` / `link`)按需用;`link` 只用于正文里的行内跳转。
**一条硬约束:同一个界面(或同一个弹窗)里,`default` 最多出现一次。**
两个主题色实底按钮并排,等于没有主次。
---
## 3. 弹窗按钮:组件已经带样式了,不要再包一层
`AlertDialogCancel` 和 `AlertDialogAction` **自带**按钮样式(分别是 `outline`
和 `default`),直接写文字即可:
```tsx
<AlertDialogFooter>
<AlertDialogCancel>取消</AlertDialogCancel>
<AlertDialogAction onClick={handleStart}>开始索引</AlertDialogAction>
</AlertDialogFooter>
```
**不要**再套一层 `Button`:
```tsx
// 反面写法:外层已经有样式了,再包一层只会产生重复类名
<AlertDialogCancel asChild>
<Button variant="outline">取消</Button>
</AlertDialogCancel>
```
需要危险动作时,用 `className` 覆盖(`cn` 走 tailwind-merge,同族类后者生效):
```tsx
<AlertDialogAction className="bg-destructive text-destructive-foreground">
删除
</AlertDialogAction>
```
---
## 4. 颜色只能用语义 token,禁止硬编码
颜色全部走 Tailwind 的语义类,它们背后是 `--tm-*` 变量,换主题时自动跟随:
```
背景 bg-primary / bg-surface / bg-accent / bg-destructive
文字 text-foreground / text-primary-foreground / text-muted-foreground
描边 border-border / border-border-subtle / border-disabled-border
```
```tsx
// 对
<Button className="bg-primary text-primary-foreground">保存</Button>
// 错 —— 换主题时这行不会跟着变
<Button className="bg-[#247a63] text-white">保存</Button>
```
**判据**:JSX 里出现 `#` 开头的颜色、`rgb(...)`、或 Tailwind 的调色板名
(`bg-green-600`、`text-slate-500`)—— 都是漏用 semantic token 的信号。
---
## 5. 「默认样式」的三个常见来源
排查界面里冒出来的黑白方角按钮时,按这个顺序找:
**① 组件导出时把样式丢了。** 最常见。把 Radix 的 primitive 原样导出:
```tsx
// 错:渲染出来就是浏览器默认按钮
const AlertDialogCancel = AlertDialogPrimitive.Cancel
```
正确做法是 `forwardRef` 包一层,挂上 `buttonVariants`:
```tsx
const AlertDialogCancel = React.forwardRef<...>(({ className, ...props }, ref) => (
<AlertDialogPrimitive.Cancel
ref={ref}
className={cn(buttonVariants({ variant: 'outline' }), className)}
{...props}
/>
))
```
**判据**:`components/ui/` 里凡是导出 Radix primitive 的地方,都要确认它是
"样式化的封装"还是"原样透传"。原样透传只对布局容器(`Root` / `Portal` /
`Group`)成立,对**可点元素**(`Close` / `Action` / `Cancel` / `Item`)不成立。
**② `asChild` 里重复包了一层。** 外层已经带样式、子元素又带一次,虽然因为
同族类后生效而不会出错,但会产生冗余类名。**能去掉一层就去掉。**
**③ 原生 `<button>` 忘写 `className`。** 见第 1 节的例外条款 —— 结构性控件也必须
有样式来源。
---
## 6. 提交前检查清单
- [ ] 新增的可点元素来自 `Button`,不是裸 `<button>`
- [ ] 同一界面里 `default` 变体不超过一个
- [ ] 危险操作走 `destructive`,不是红色硬编码
- [ ] 弹窗按钮没有重复包 `Button`
- [ ] JSX 里没有 `#` 开头的颜色、没有 Tailwind 调色板名
- [ ] `components/ui/` 里新导出的可点 primitive 已经挂上 `buttonVariants`
- [ ] 组件测试覆盖到按钮的可见性与点击行为(testid 用 `xxx-yyy` 连字符命名)
---
## 7. 一个反面案例的复盘
弹窗里的「取消 / 开始索引」显示成浏览器默认样式,原因就是第 5 节第 ① 条:
`alert-dialog.tsx` 把 `Cancel` / `Action` 两个 primitive 原样导出了。
修复是给它们各加一个 `forwardRef` 封装,挂上 `buttonVariants`。**组件本身没坏**,
所有调用方一行不用改,样式自动生效 —— 这正是把样式收在 `components/ui/` 里的价值:
**修一处,全产品对齐。**
如果你发现某个地方的按钮"没跟上主题",先别去改那个界面 ——
**先看它用的组件是不是漏了样式。**
@@ -0,0 +1,92 @@
# 微信系统消息(sysmsg)解析与格式兼容
微信的「系统消息」(入群、撤回、成员变动等)以 XML(`<sysmsg>`)存放在消息内容里,
但**同一类提示的 XML 结构会随客户端版本变化**。本文说明 TraceMemo 的解析方式,
以及在遇到新格式时应当怎么扩展。
## 两类格式
### 旧格式:正文直接放在 `<plain>`
```xml
<sysmsg type="delchatroommember">
<delchatroommember>
<plain><![CDATA["成员昵称"通过扫描你分享的二维码加入群聊]]></plain>
<text><![CDATA["成员昵称"通过扫描你分享的二维码加入群聊]]></text>
<link>
<scene>qrcode</scene>
<text><![CDATA[撤销]]></text>
</link>
</delchatroommember>
</sysmsg>
```
解析:命中 `delchatroommember`,直接取 `<plain>`。
### 新格式:正文在 `<template>`,用 `$名称$` 引用 link
```xml
<sysmsg type="sysmsgtemplate">
<sysmsgtemplate>
<content_template type="tmpl_type_profilewithrevokeqrcode">
<plain><![CDATA[]]></plain>
<template><![CDATA["$adder$"通过扫描你分享的二维码加入群聊 $revoke$]]></template>
<link_list>
<link name="adder" type="link_profile">
<memberlist><member>
<username><![CDATA[wxid_xxxxxxxx]]></username>
<nickname><![CDATA[成员昵称]]></nickname>
</member></memberlist>
</link>
<link name="revoke" type="link_revoke_qrcode" hidden="1">
<title><![CDATA[撤销]]></title>
</link>
</link_list>
</content_template>
</sysmsgtemplate>
</sysmsg>
```
三个要点:
- `<plain>` 变成**空 CDATA**,正文挪进 `<template>`;
- 正文里的 `$名称$` 是占位符,按 `<link_list>` 中 `link[name]` 回填;
- `hidden="1"` 的 link 在微信里是**可点击按钮**,纯文本展示时应省略其文案。
## 解析流程
`src/main/message-parser.ts` 的 `parseSystemMessage()` 按以下顺序尝试:
| 顺序 | 分支 | 处理对象 |
| --- | --- | --- |
| 1 | `extractRecallMessage` | `<revokemsg>` 撤回通知 |
| 2 | `extractSysmsgTemplateText` | `<sysmsgtemplate>` 模板消息 |
| 3 | `extractDelChatroomMemberText` | `<delchatroommember>` 成员变动 |
| 4 | 通用提取(`plain` → `text` → `title`),再退回 `fallbackSystemText` | 其余未覆盖类型 |
第 4 步之前会先调用 `stripSysmsgLinkList()` 剥掉 `<link_list>`。
## 为什么必须显式处理新格式
通用提取链只在第 1~3 步全部落空时才执行,而新格式恰好让它落空:
`<plain>` 是空 CDATA,又没有 `<text>`,于是取到 `<title>` ——
那是 `hidden="1"` 按钮的标题。**结果是整条系统消息只剩一个按钮文案**,
例如把「某某通过扫描你分享的二维码加入群聊」显示成「撤销」。
因此三处约束缺一不可:
1. 模板分支必须排在通用提取之前;
2. 占位符回填必须尊重 `hidden="1"`;
3. 通用提取前先剥 `<link_list>`,作为未知类型的防护。
## 新增一类系统消息时
1. 从真实消息中取出 `content`(`<sysmsg>` 原文),确认 `type` 与承载正文的标签;
2. 在 `parseSystemMessage()` 里加一个**早于通用提取**的分支;
3. 补 `tests/unit/message-parser.test.ts` 用例,**新旧两版各一条**,防止回归;
4. 文档与代码注释只写结构,不粘贴真实会话内容、昵称、wxid 或二维码链接。
## 相关位置
- 解析实现:`src/main/message-parser.ts`
- 单元测试:`tests/unit/message-parser.test.ts`
+60 -3
View File
@@ -1,5 +1,62 @@
# macOS 数据访问说明(兼容入口)
# macOS 关闭 SIP 教程
完整内容已移到[macOS 数据访问与系统权限](./platform/macos.md)。
SIP(System Integrity Protection,系统完整性保护)是 macOS 的系统安全机制。关闭 SIP 会降低系统安全性,只建议在确实需要读取或调试本地微信数据时临时关闭;操作完成后,建议重新开启。
保留此文件是为了兼容应用内已经发布的帮助链接。请不要把“关闭 SIP”当作默认安装步骤;只有当当前连接页面明确要求时才处理,并在完成后恢复系统安全设置。
> 只在连接页面明确提示需要关闭 SIP 时才处理。首次连接失败时,先确认微信版本、账号目录和登录时机,再按本文操作。关闭 SIP 不是 TraceMemo 的常规安装步骤,也不应长期保持关闭。
## 准备
- 一台 Mac 电脑,Intel 芯片和 Apple Silicon 芯片均可。
- 需要进入 macOS 恢复模式。
- 请先保存正在编辑的文件,并预留一次重启时间。
## 关闭 SIP
### Intel Mac
1. 关机。
2. 按下开机键后,立刻按住 `Command + R`。
3. 保持按住,直到进入 macOS 恢复模式。
### Apple Silicon Mac(M1/M2/M3/M4/M5)
1. 关机。
2. 长按开机键不放。
3. 直到出现启动选项界面后松开。
4. 选择"选项",进入 macOS 恢复模式。
### 在恢复模式中执行命令
1. 进入恢复模式后,点击顶部菜单栏的 **Utilities(实用工具)**。
2. 选择 **Terminal(终端)**。
3. 在终端中输入:
```bash
csrutil disable
```
4. 按回车执行。
5. 看到关闭成功提示后,重启电脑。
## 确认是否生效
重启回到正常桌面后,打开"终端",执行:
```bash
csrutil status
```
看到 `System Integrity Protection status: disabled.` 才算关闭成功。
若仍显示 `enabled`,说明没有生效。常见原因是没在恢复模式里执行,或系统刚做过大版本更新——
macOS 大版本更新会把 SIP 重置回开启状态,此前关过也会失效,需要重新按上面的步骤操作。
## 重新开启 SIP
拿到数据库密钥后,建议重新进入恢复模式,在终端中执行:
```bash
csrutil enable
```
然后重启电脑,恢复系统安全设置。
+1 -1
View File
@@ -8,7 +8,7 @@ TraceMemo 需要读取微信本地数据。macOS 会根据系统版本、微信
1. 先启动 TraceMemo,阅读连接页面显示的当前前置条件。
2. 确认微信数据目录指向当前账号。
3. 只在页面明确要求时处理系统授权或 SIP;按页面提示完成密钥获取后,恢复你平时使用的安全设置。
3. 只在页面明确要求时处理系统授权或 SIP;关闭 SIP 的具体步骤见[关闭 SIP 教程](../mac-disable-sip.md),按页面提示完成密钥获取后,恢复你平时使用的安全设置。
4. 返回应用重新检测账号、数据库和图片资源状态。
不要直接复制网上针对其他微信版本的命令。系统授权失败时,记录 macOS 版本、微信版本和页面错误,再按[排障文档](../user-guide/troubleshooting.md#连接微信失败)处理。
+2 -1
View File
@@ -19,6 +19,8 @@ asarUnpack:
- node_modules/silk-wasm/**
- node_modules/sherpa-onnx-node/**
- node_modules/sherpa-onnx-*/**
- node_modules/@napi-rs/system-ocr/**
- node_modules/@napi-rs/system-ocr-*/**
extraResources:
# Keep the updater provider in every packaged Windows app. electron-builder also
# regenerates this file during publish, using the same release configuration.
@@ -33,7 +35,6 @@ extraResources:
- mobile_daily_report.html
- mobile_daily_report_v1.html
- mobile_daily_report_v2.html
- connectors/wechat/win32-x64/**
- key/win32/x64/**
- runtime/win32/**
- wcdb/win32/x64/**
+5
View File
@@ -19,6 +19,11 @@ asarUnpack:
- node_modules/silk-wasm/**
- node_modules/sherpa-onnx-node/**
- node_modules/sherpa-onnx-*/**
# System OCR(@napi-rs/system-ocr)的 native binding 必须 unpacked,否则
# macOS 的系统 OCR 会在运行时 MODULE_NOT_FOUND。平台本机的 binding 由
# scripts/after-pack.cjs 校验,外架构的同级包在 afterPack 里被裁掉。
- node_modules/@napi-rs/system-ocr/**
- node_modules/@napi-rs/system-ocr-*/**
extraResources:
# Keep the updater provider in every packaged macOS app. electron-builder also
# regenerates this file during publish, using the same release configuration.
+1 -1
View File
@@ -16,7 +16,7 @@ export default defineConfig({
output: {
entryFileNames: '[name].js'
},
external: ['koffi', 'sherpa-onnx-node']
external: ['koffi', 'sherpa-onnx-node', '@napi-rs/system-ocr']
}
}
},
+12 -14
View File
@@ -25,12 +25,13 @@
},
"main": "./out/main/index.js",
"scripts": {
"test": "pnpm typecheck && pnpm test:unit && pnpm test:component && pnpm test:integration && pnpm test:skill-install && pnpm test:wechat-connector && pnpm test:e2e:build && playwright test",
"test": "pnpm typecheck && pnpm test:unit && pnpm test:component && pnpm test:integration && pnpm test:skill-install && pnpm test:e2e:build && playwright test",
"format": "prettier --write .",
"lint": "eslint --cache .",
"typecheck:node": "tsc --noEmit -p tsconfig.node.json --composite false",
"typecheck:web": "tsc --noEmit -p tsconfig.web.json --composite false",
"typecheck": "npm run typecheck:node && npm run typecheck:web",
"typecheck": "npm run typecheck:node && npm run typecheck:web && npm run typecheck:test",
"typecheck:test": "node scripts/typecheck-tests.cjs",
"test:skill-install": "node scripts/test-skill-install-instruction.cjs",
"cp:env": "node scripts/ensure-env.cjs",
"prepare:env": "node scripts/ensure-env.cjs",
@@ -41,11 +42,10 @@
"prepare:wechat-personal": "node scripts/prepare-wechat-chatter-runtime.cjs",
"start": "electron-vite preview",
"predev": "node scripts/ensure-electron-binary.cjs",
"dev": "node scripts/ensure-env.cjs && node scripts/build-wechat-connector.cjs && electron-vite dev",
"dev": "node scripts/ensure-env.cjs && electron-vite dev",
"dev:update": "cross-env TRACEMEMO_UPDATE_SIMULATION=true pnpm dev",
"poc:query-agent": "electron-vite build && node scripts/run-query-agent-poc.cjs",
"poc:query-agent:run": "node scripts/run-query-agent-poc.cjs",
"test:wechat-connector": "go -C services/wechat-connector test ./... && go -C services/wechat-connector vet ./...",
"test:unit": "vitest run --config vitest.unit.config.ts",
"test:component": "vitest run --config vitest.component.config.ts",
"test:integration": "vitest run --config vitest.integration.config.ts",
@@ -56,21 +56,17 @@
"test:e2e": "pnpm test:e2e:build && playwright test --grep-invert @visual",
"test:visual": "pnpm test:e2e:build && playwright test tests/e2e/visual.spec.ts",
"test:smoke": "node --test tests/smoke/native-environment.test.mjs",
"build:wechat-connector": "node scripts/build-wechat-connector.cjs",
"build:wechat-connector:win": "node scripts/build-wechat-connector.cjs --platform win32 --arch x64",
"build:wechat-connector:mac": "node scripts/build-wechat-connector.cjs --platform darwin --arch arm64",
"build:native-services": "npm run build:wechat-connector",
"build": "npm run typecheck && npm run build:native-services && electron-vite build",
"build": "npm run typecheck && electron-vite build",
"postinstall": "electron-builder install-app-deps && node scripts/prepare-electron-runtime.cjs && node scripts/ensure-electron-binary.cjs",
"build:unpack": "npm run build && electron-builder --config electron-builder.yml --dir",
"build:win": "npm run typecheck && npm run build:wechat-connector:win && npm run prepare:win-runtime && electron-vite build && electron-builder --config electron-builder.win.yml --win --x64",
"build:mac:arm64": "npm run typecheck && node scripts/build-wechat-connector.cjs --platform darwin --arch arm64 && npm run prepare:ffmpeg:mac:arm64 && electron-vite build && electron-builder --config electron-builder.yml --mac --arm64",
"build:mac:x64": "npm run typecheck && node scripts/build-wechat-connector.cjs --platform darwin --arch x64 && npm run prepare:ffmpeg:mac:x64 && electron-vite build && electron-builder --config electron-builder.yml --mac --x64",
"build:win": "npm run typecheck && npm run prepare:win-runtime && electron-vite build && electron-builder --config electron-builder.win.yml --win --x64",
"build:mac:arm64": "npm run typecheck && npm run prepare:ffmpeg:mac:arm64 && electron-vite build && electron-builder --config electron-builder.yml --mac --arm64",
"build:mac:x64": "npm run typecheck && npm run prepare:ffmpeg:mac:x64 && electron-vite build && electron-builder --config electron-builder.yml --mac --x64",
"release": "npm run release:mac && npm run release:win",
"release:mac": "npm run typecheck && node scripts/build-wechat-connector.cjs --platform darwin --arch arm64,x64 && electron-vite build && npm run release:mac:arm64 && npm run release:mac:x64",
"release:mac": "npm run typecheck && electron-vite build && npm run release:mac:arm64 && npm run release:mac:x64",
"release:mac:arm64": "npm run prepare:ffmpeg:mac:arm64 && electron-builder --config electron-builder.yml --mac --arm64 --publish always",
"release:mac:x64": "npm run prepare:ffmpeg:mac:x64 && electron-builder --config electron-builder.yml --mac --x64 --publish always",
"release:win": "npm run typecheck && npm run build:wechat-connector:win && npm run prepare:win-runtime && electron-vite build && electron-builder --config electron-builder.win.yml --win --x64 --publish always",
"release:win": "npm run typecheck && npm run prepare:win-runtime && electron-vite build && electron-builder --config electron-builder.win.yml --win --x64 --publish always",
"release:beta": "cross-env EP_PRE_RELEASE=true npm run release",
"release:stable": "npm run release",
"build:linux": "electron-vite build && electron-builder --config electron-builder.yml --linux"
@@ -79,6 +75,8 @@
"@electron-toolkit/preload": "^3.0.2",
"@electron-toolkit/utils": "^4.0.0",
"@koromix/koffi-win32-x64": "3.1.0",
"@napi-rs/system-ocr": "1.2.0",
"@napi-rs/system-ocr-win32-x64-msvc": "1.2.0",
"@radix-ui/react-alert-dialog": "^1.1.23",
"@radix-ui/react-checkbox": "^1.3.11",
"@radix-ui/react-dialog": "^1.1.23",
+42
View File
@@ -12,6 +12,8 @@ specifiers:
'@electron-toolkit/tsconfig': ^2.0.0
'@electron-toolkit/utils': ^4.0.0
'@koromix/koffi-win32-x64': 3.1.0
'@napi-rs/system-ocr': 1.2.0
'@napi-rs/system-ocr-win32-x64-msvc': 1.2.0
'@playwright/test': ^1.62.1
'@radix-ui/react-alert-dialog': ^1.1.23
'@radix-ui/react-checkbox': ^1.3.11
@@ -86,6 +88,8 @@ dependencies:
'@electron-toolkit/preload': 3.0.2_electron@43.1.0
'@electron-toolkit/utils': 4.0.0_electron@43.1.0
'@koromix/koffi-win32-x64': 3.1.0
'@napi-rs/system-ocr': 1.2.0
'@napi-rs/system-ocr-win32-x64-msvc': 1.2.0
'@radix-ui/react-alert-dialog': 1.1.23_eijghdl4n2x4hz6j4cg7ctgbuu
'@radix-ui/react-checkbox': 1.3.11_eijghdl4n2x4hz6j4cg7ctgbuu
'@radix-ui/react-dialog': 1.1.23_eijghdl4n2x4hz6j4cg7ctgbuu
@@ -1685,6 +1689,44 @@ packages:
- supports-color
dev: true
/@napi-rs/system-ocr/1.2.0:
resolution: {integrity: sha512-r0f2xNH6U+sth44qF+lUP+2WuHSGUBAry5KSCNuaLDGRbgslFqeROr/qJJ/fb6AjBp3Ov+CJP5MdrOWoaoM3cw==}
engines: {node: '>= 10'}
optionalDependencies:
'@napi-rs/system-ocr-darwin-arm64': 1.2.0
'@napi-rs/system-ocr-darwin-x64': 1.2.0
'@napi-rs/system-ocr-win32-arm64-msvc': 1.2.0
'@napi-rs/system-ocr-win32-x64-msvc': 1.2.0
dev: false
/@napi-rs/system-ocr-darwin-arm64/1.2.0:
resolution: {integrity: sha512-cK8dcDBEl3P4A04xmFJSHEJQxfDytaAIFyDCLqavTp92FVU5plESttWzZsqtTkS81/kzKiBfHyPQffSIndfWbQ==}
cpu: [arm64]
os: [darwin]
engines: {node: '>= 10'}
dev: false
/@napi-rs/system-ocr-darwin-x64/1.2.0:
resolution: {integrity: sha512-u3TBvBGrhmT5Os6AfaxbUEg6VHe8lvrFJNPgThJgshJHyRXUx/wCfTyOroJ22KdVCP5AE4GpwS5tFHMb6p6iaQ==}
cpu: [x64]
os: [darwin]
engines: {node: '>= 10'}
dev: false
/@napi-rs/system-ocr-win32-arm64-msvc/1.2.0:
resolution: {integrity: sha512-7ej8uMvmXomw3NXo5gZ5p2Nl6UKsHI+VRU3ELv0mhcxR0sJ6wFifYTu5bJrM1TGcz1/RsaX+TjWMmsDq8vriKQ==}
cpu: [arm64]
os: [win32]
engines: {node: '>= 10'}
dev: false
/@napi-rs/system-ocr-win32-x64-msvc/1.2.0:
resolution: {integrity: sha512-oOoCj3FPWDVctTxx98vMBiMI6m51U+w7SMmMefvmtpcpLelzZ/zYTqdwtWZFAjShaHO+RdaKkcpeVcQuBQiVbA==}
cpu: [x64]
os: [win32]
engines: {node: '>= 10'}
dev: false
/@nodelib/fs.scandir/2.1.5:
resolution: {integrity: sha512-vq24Bq3ym5HEQm2NKCr3yXDwjc7vTsEThRDnkp2DK9p1uqLR+DHurm/NOTo0KG7HYHU7eppKZj3MyqYuMBf62g==}
engines: {node: '>= 8'}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 157 KiB

After

Width:  |  Height:  |  Size: 153 KiB

+17 -9
View File
@@ -68,21 +68,29 @@
.overview {
margin-top: 2px;
}
/*
* hero 头像簇的几何必须由 contract 变量驱动,不能再写死容器尺寸。
*
* 生产导出会额外注入 `report-template-fragment-contract.ts`,其中
* `img.tm-avatar.tm-avatar--hero` 用 !important 把头像钉在
* clamp(28px, var(--tm-avatar-hero-size, 40px), 56px)。
* 旧版这里写死 58x58(单头像 28x28),两个权威打架:头像实际 40px,
* 2 列 x 40px + 3px gap = 83px 塞不进 58px 的盒子,于是头像向右向下溢出容器,
* 视觉上越过卡片内边距、压到卡片边缘之外。
*
* 现在容器尺寸由内容决定(列宽/行高都取同一个变量):头像数 1..4 都不会溢出,
* 主题调整 --tm-avatar-hero-size 时容器与头像也不会分叉。
*/
.avatar-grid {
width: 58px;
height: 58px;
display: grid;
grid-template-columns: 1fr 1fr;
grid-template-columns: repeat(2, var(--tm-avatar-hero-size, 40px));
grid-auto-rows: var(--tm-avatar-hero-size, 40px);
gap: 3px;
flex: 0 0 auto;
}
/* 单头像时不保留空列,簇宽恰好等于一个头像。 */
.avatar-grid.avatar-count-1 {
width: 28px;
height: 28px;
grid-template-columns: 1fr;
}
.avatar-grid.avatar-count-2 {
height: 28px;
grid-template-columns: var(--tm-avatar-hero-size, 40px);
}
.avatar-grid.empty-section {
display: none;
+17 -9
View File
@@ -61,21 +61,29 @@
font-size: 13px;
line-height: 1.55;
}
/*
* hero 头像簇的几何必须由 contract 变量驱动,不能再写死容器尺寸。
*
* 生产导出会额外注入 `report-template-fragment-contract.ts`,其中
* `img.tm-avatar.tm-avatar--hero` 用 !important 把头像钉在
* clamp(28px, var(--tm-avatar-hero-size, 40px), 56px)。
* 旧版这里写死 58x58(单头像 28x28),两个权威打架:头像实际 40px,
* 2 列 x 40px + 3px gap = 83px 塞不进 58px 的盒子,于是头像向右向下溢出容器,
* 视觉上越过卡片内边距、压到卡片边缘之外。
*
* 现在容器尺寸由内容决定(列宽/行高都取同一个变量):头像数 1..4 都不会溢出,
* 主题调整 --tm-avatar-hero-size 时容器与头像也不会分叉。
*/
.avatar-grid {
width: 58px;
height: 58px;
display: grid;
grid-template-columns: 1fr 1fr;
grid-template-columns: repeat(2, var(--tm-avatar-hero-size, 40px));
grid-auto-rows: var(--tm-avatar-hero-size, 40px);
gap: 3px;
flex: 0 0 auto;
}
/* 单头像时不保留空列,簇宽恰好等于一个头像。 */
.avatar-grid.avatar-count-1 {
width: 28px;
height: 28px;
grid-template-columns: 1fr;
}
.avatar-grid.avatar-count-2 {
height: 28px;
grid-template-columns: var(--tm-avatar-hero-size, 40px);
}
.avatar-grid.empty-section {
display: none;
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+1 -1
View File
@@ -1 +1 @@
f1bbd88584075c0e6d3487357a06424949de26d7e6dcad803d210aab83b0cbda xkey_helper_4_1_13
a35d3fa67387e049cc30bc073f9e65b077aa1db9c5c2b183125dec695fa013b9 xkey_helper_4_1_13
+130 -6
View File
@@ -16,6 +16,15 @@ const REQUIRED_RUNTIME_PACKAGES = [
'koffi'
]
// electron-builder 26 skips macOS signing entirely when no Developer ID
// identity is configured, so an unpacked bundle can ship without a usable
// signature. macOS kills a helper whose code or signature is missing or
// modified even when SIP is disabled, which is what customers hit on newer
// macOS releases. Ad-hoc re-sign the runtime helpers and the outer bundle so
// every Mach-O verifies strictly; spctl still rejects ad-hoc code, which is
// acceptable for the SIP-disabled customer workflow.
const MACOS_HELPER_NAMES = ['xkey_helper', 'xkey_helper_4_1_13']
function getRuntimeResources(context) {
const productName = context.packager.appInfo.productFilename
return context.electronPlatformName === 'darwin'
@@ -78,11 +87,106 @@ function validateSherpaRuntime(runtimeResources, platform, arch) {
}
}
/**
* System OCR 用 native package(@napi-rs/system-ocr)。它是 external + asarUnpack,
* 打包后必须以 unpacked 形式存在,否则运行时会 MODULE_NOT_FOUND / native binding missing。
* Windows 与 macOS 都是 supported target,都要做硬校验(Linux 不是)。
*/
function systemOcrTarget(platform, arch) {
return platform === 'win32' ? `${platform}-${arch}-msvc` : `${platform}-${arch}`
}
function validateSystemOcrRuntime(runtimeResources, platform, arch) {
if (platform !== 'win32' && platform !== 'darwin') return
const target = systemOcrTarget(platform, arch)
const basePath = path.join(
runtimeResources,
'app.asar.unpacked',
'node_modules',
'@napi-rs',
'system-ocr'
)
const nativePath = path.join(
runtimeResources,
'app.asar.unpacked',
'node_modules',
'@napi-rs',
`system-ocr-${target}`
)
const requiredFiles = [
path.join(basePath, 'package.json'),
path.join(basePath, 'index.js'),
path.join(nativePath, 'package.json'),
path.join(nativePath, `system-ocr.${target}.node`)
]
const missingFiles = requiredFiles.filter((filePath) => !existsSync(filePath))
if (missingFiles.length > 0) {
throw new Error(`Missing unpacked System OCR runtime: ${missingFiles.join(', ')}`)
}
}
function normalizeBuilderArch(arch) {
if (typeof arch === 'string') return arch
return { 0: 'ia32', 1: 'x64', 2: 'armv7l', 3: 'arm64', 4: 'universal' }[arch] || String(arch)
}
function runCodesign(args) {
execFileSync('/usr/bin/codesign', args, { stdio: 'ignore' })
}
function isMacosCodeValid(targetPath, run = runCodesign) {
try {
run(['--verify', '--strict', targetPath])
return true
} catch {
return false
}
}
function findMacosHelperPaths(runtimeResources) {
return MACOS_HELPER_NAMES.map((name) => path.join(runtimeResources, 'resources', name)).filter(
(helperPath) => existsSync(helperPath)
)
}
function signMacosHelpers(runtimeResources, run = runCodesign) {
const helperPaths = findMacosHelperPaths(runtimeResources)
for (const helperPath of helperPaths) {
chmodSync(helperPath, 0o755)
if (!isMacosCodeValid(helperPath, run)) {
run(['--force', '--sign', '-', helperPath])
}
for (const arch of ['arm64', 'x86_64']) {
try {
run(['--verify', '--strict', '--arch', arch, helperPath])
} catch (error) {
throw new Error(
'macOS helper signature verification failed: ' +
path.basename(helperPath) +
' (' +
arch +
')',
{ cause: error }
)
}
}
}
return helperPaths
}
function signMacosAppBundle(appBundlePath, run = runCodesign) {
if (isMacosCodeValid(appBundlePath, run)) return appBundlePath
run(['--force', '--sign', '-', appBundlePath])
try {
run(['--verify', '--strict', appBundlePath])
} catch (error) {
throw new Error('macOS app bundle signature verification failed: ' + appBundlePath, {
cause: error
})
}
return appBundlePath
}
/**
* A foreign-architecture binary only fails once the user touches the feature
* that needs it, so verify the ones whose filename is shared across
@@ -152,16 +256,24 @@ function validateReaderSkillRuntime(runtimeResources) {
* The loaders pick their package from process.platform/arch, so the siblings
* are dead weight — drop them.
*/
// 每个条目返回 platform package 的**完整后缀**(不含 package 前缀与连字符)。
const NATIVE_RUNTIME_PACKAGES = [
{
modules: [],
prefix: 'sherpa-onnx',
platformName: (platform) => (platform === 'win32' ? 'win' : platform)
platformName: (platform, arch) => `${platform === 'win32' ? 'win' : platform}-${arch}`
},
{
modules: ['@koromix'],
prefix: 'koffi',
platformName: (platform) => platform
platformName: (platform, arch) => `${platform}-${arch}`
},
{
// @napi-rs 的 platform package 目录名带 -msvc 后缀(win32-x64-msvc)。
modules: ['@napi-rs'],
prefix: 'system-ocr',
platformName: (platform, arch) => systemOcrTarget(platform, arch),
foreignPattern: /^system-ocr-[a-z0-9]+-(arm64|x64|ia32|loong64|riscv64)(-msvc)?$/
}
]
@@ -173,12 +285,16 @@ function pruneForeignArchNativeRuntimes(runtimeResources, platform, arch) {
for (const runtime of NATIVE_RUNTIME_PACKAGES) {
const modulesRoot = path.join(unpackedRoot, ...runtime.modules)
if (!existsSync(modulesRoot)) continue
const expected = `${runtime.prefix}-${runtime.platformName(platform)}-${arch}`
const foreign = new RegExp(`^${runtime.prefix}-[a-z0-9]+-(arm64|x64|ia32|loong64|riscv64)$`)
const expected = `${runtime.prefix}-${runtime.platformName(platform, arch)}`
const foreign =
runtime.foreignPattern ||
new RegExp(`^${runtime.prefix}-[a-z0-9]+-(arm64|x64|ia32|loong64|riscv64)$`)
for (const entry of readdirSync(modulesRoot, { withFileTypes: true })) {
if (!entry.isDirectory() || entry.name === expected || !foreign.test(entry.name)) continue
rmSync(path.join(modulesRoot, entry.name), { recursive: true, force: true })
removed.push(runtime.modules.length ? `${runtime.modules.join('/')}/${entry.name}` : entry.name)
removed.push(
runtime.modules.length ? `${runtime.modules.join('/')}/${entry.name}` : entry.name
)
}
}
return removed
@@ -223,6 +339,7 @@ exports.default = async function afterPack(context) {
'Bundled ffmpeg'
)
validateSherpaRuntime(runtimeResources, context.electronPlatformName, arch)
validateSystemOcrRuntime(runtimeResources, context.electronPlatformName, arch)
pruneIntelMacKeyTool(runtimeResources, context.electronPlatformName, arch)
pruneForeignArchConnectors(runtimeResources, context.electronPlatformName, arch)
pruneForeignArchNativeRuntimes(runtimeResources, context.electronPlatformName, arch)
@@ -231,6 +348,9 @@ exports.default = async function afterPack(context) {
execFileSync('/usr/bin/codesign', ['--force', '--sign', '-', ffmpegPath], {
stdio: 'ignore'
})
signMacosHelpers(runtimeResources)
const productName = context.packager.appInfo.productFilename
signMacosAppBundle(path.join(context.appOutDir, productName + '.app'))
}
if (context.electronPlatformName === 'win32') {
@@ -249,7 +369,6 @@ exports.default = async function afterPack(context) {
}
return
}
}
exports.getRuntimeResources = getRuntimeResources
@@ -258,7 +377,12 @@ exports.validateReaderSkillRuntime = validateReaderSkillRuntime
exports.validateFfmpegRuntime = validateFfmpegRuntime
exports.validateSilkWasmRuntime = validateSilkWasmRuntime
exports.validateSherpaRuntime = validateSherpaRuntime
exports.validateSystemOcrRuntime = validateSystemOcrRuntime
exports.pruneIntelMacKeyTool = pruneIntelMacKeyTool
exports.pruneForeignArchConnectors = pruneForeignArchConnectors
exports.pruneForeignArchNativeRuntimes = pruneForeignArchNativeRuntimes
exports.validateRuntimeBinaryArchitecture = validateRuntimeBinaryArchitecture
exports.findMacosHelperPaths = findMacosHelperPaths
exports.isMacosCodeValid = isMacosCodeValid
exports.signMacosHelpers = signMacosHelpers
exports.signMacosAppBundle = signMacosAppBundle
-65
View File
@@ -1,65 +0,0 @@
/* eslint-disable @typescript-eslint/no-require-imports, @typescript-eslint/explicit-function-return-type */
const { execFileSync } = require('node:child_process')
const fs = require('node:fs')
const path = require('node:path')
const projectRoot = path.resolve(__dirname, '..')
const sourceDir = path.join(projectRoot, 'services', 'wechat-connector')
const outputRoot = path.join(projectRoot, 'resources', 'connectors', 'wechat')
function normalizePlatform(value) {
if (value === 'win32' || value === 'windows') return 'windows'
if (value === 'darwin' || value === 'macos') return 'darwin'
if (value === 'linux') return 'linux'
throw new Error(`Unsupported connector platform: ${value}`)
}
function normalizeArch(value) {
if (value === 'x64' || value === 'amd64') return 'amd64'
if (value === 'arm64') return 'arm64'
throw new Error(`Unsupported connector architecture: ${value}`)
}
function detectHostArch() {
if (process.platform !== 'darwin') return process.arch
try {
const arm64Supported = execFileSync('sysctl', ['-n', 'hw.optional.arm64'], {
encoding: 'utf8'
}).trim()
return arm64Supported === '1' ? 'arm64' : process.arch
} catch {
return process.arch
}
}
function parseTargets() {
const platformArg = process.argv.indexOf('--platform')
const archArg = process.argv.indexOf('--arch')
const platforms = platformArg >= 0 ? process.argv[platformArg + 1].split(',') : [process.platform]
const arches = archArg >= 0 ? process.argv[archArg + 1].split(',') : [detectHostArch()]
return platforms.flatMap((platform) =>
arches.map((arch) => ({ goos: normalizePlatform(platform), goarch: normalizeArch(arch) }))
)
}
if (!fs.existsSync(path.join(sourceDir, 'go.mod'))) {
throw new Error(`Repository-local WeChat connector source is missing: ${sourceDir}`)
}
for (const target of parseTargets()) {
const directoryName = `${target.goos === 'windows' ? 'win32' : target.goos}-${target.goarch === 'amd64' ? 'x64' : target.goarch}`
const outputDir = path.join(outputRoot, directoryName)
const outputPath = path.join(
outputDir,
target.goos === 'windows' ? 'wechat-connector.exe' : 'wechat-connector'
)
fs.rmSync(outputDir, { recursive: true, force: true })
fs.mkdirSync(outputDir, { recursive: true })
execFileSync('go', ['build', '-trimpath', '-o', outputPath, '.'], {
cwd: sourceDir,
env: { ...process.env, GOOS: target.goos, GOARCH: target.goarch, CGO_ENABLED: '0' },
stdio: 'inherit'
})
if (target.goos !== 'windows') fs.chmodSync(outputPath, 0o755)
console.log(`[build-wechat-connector] built ${directoryName}: ${outputPath}`)
}
+11 -4
View File
@@ -24,10 +24,17 @@ const avatarSvg = (label, color) =>
`<svg xmlns="http://www.w3.org/2000/svg" width="96" height="96"><rect width="96" height="96" rx="18" fill="${color}"/><text x="48" y="58" text-anchor="middle" font-family="PingFang SC, sans-serif" font-size="36" fill="#0f172a">${label}</text></svg>`
).toString('base64')}`
const localImagePath = '/Users/user/Library/Containers/com.tencent.xinWeChat/Data/Documents/xwechat_files/fixture_account_1a2b/temp/RWTemp/2026-07/fixture-image-hash.png'
const sampleImage = fs.existsSync(localImagePath)
? `data:image/png;base64,${fs.readFileSync(localImagePath).toString('base64')}`
: avatarSvg('图', '#dbeafe')
/**
* 可选的本地样例图(用于人工核对图片区块的排版)。
*
* 走环境变量传入,**不要在源码里写本机路径** —— 微信数据目录会连带暴露
* 系统用户名与账号目录名。不传就退回内置的 SVG 头像占位。
*/
const localImagePath = process.env.REPORT_FIXTURE_IMAGE || ''
const sampleImage =
localImagePath && fs.existsSync(localImagePath)
? `data:image/png;base64,${fs.readFileSync(localImagePath).toString('base64')}`
: avatarSvg('图', '#dbeafe')
const avatars = {
阿宇: avatarSvg('宇', '#dcfce7'),
+140
View File
@@ -0,0 +1,140 @@
/*
* 测试文件的类型检查棘轮(ratchet)。
*
* 背景:`tsconfig.node.json` / `tsconfig.web.json` 的 include 都不含 `tests/`,
* 所以测试里的类型错误对 `pnpm typecheck` 与 CI 完全不可见——已经积累了一批历史债。
*
* 策略:**不阻塞既有债,但禁止新增**。
* - 基线按「文件 → 错误数」记录,而不是只记总数:
* 否则在 A 文件修掉 1 条、同时在 B 文件新增 1 条会互相抵消,棘轮形同虚设。
* - 某个文件的错误数超过基线即失败;新增了带类型错误的文件同样失败。
* - 需要主动下调基线时用 `--update`(只在确实修好了错误之后)。
*
* 用法:
* node scripts/typecheck-tests.cjs # 校验
* node scripts/typecheck-tests.cjs --update # 用当前结果重写基线
*/
/* eslint-disable @typescript-eslint/explicit-function-return-type, @typescript-eslint/no-require-imports */
const { spawnSync } = require('node:child_process')
const fs = require('node:fs')
const path = require('node:path')
const projectRoot = path.resolve(__dirname, '..')
const configPath = path.join(projectRoot, 'tsconfig.test.json')
const baselinePath = path.join(projectRoot, 'tests', 'typecheck-baseline.json')
const ERROR_LINE = /^(.+?)\((\d+),(\d+)\): error (TS\d+): (.*)$/
const MAX_REPORTED = 20
function runTypeScript() {
// 直接用本地 typescript 包,避免依赖 node_modules/.bin 在各平台的差异。
const tscPath = require.resolve('typescript/bin/tsc')
const result = spawnSync(
process.execPath,
[tscPath, '--noEmit', '--pretty', 'false', '-p', configPath],
{ cwd: projectRoot, encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 }
)
if (result.error) throw result.error
return `${result.stdout || ''}${result.stderr || ''}`
}
function relative(file) {
const rel = path.relative(projectRoot, path.resolve(projectRoot, file))
return rel.split(path.sep).join('/')
}
/** 解析出「文件 → 错误数」与「文件 → 错误信息列表」。 */
function collectErrors(output) {
const counts = new Map()
const details = new Map()
for (const line of output.split('\n')) {
const match = ERROR_LINE.exec(line.trim())
if (!match) continue
const file = relative(match[1])
counts.set(file, (counts.get(file) ?? 0) + 1)
if (!details.has(file)) details.set(file, [])
if (details.get(file).length < 3) details.get(file).push(`${match[2]}:${match[3]} ${match[4]}`)
}
return { counts, details }
}
function readBaseline() {
try {
const parsed = JSON.parse(fs.readFileSync(baselinePath, 'utf8'))
if (parsed && typeof parsed.files === 'object' && parsed.files !== null) return parsed
} catch {
// 基线缺失或损坏时按「空基线」处理,会在下面明确报错提示。
}
return null
}
function total(counts) {
let sum = 0
for (const value of counts.values()) sum += value
return sum
}
function main() {
if (!fs.existsSync(configPath)) {
console.error(`[typecheck:test] 缺少 ${path.relative(projectRoot, configPath)}`)
process.exit(1)
}
const { counts, details } = collectErrors(runTypeScript())
if (process.argv.includes('--update')) {
const files = Object.fromEntries([...counts.entries()].sort(([a], [b]) => a.localeCompare(b)))
const payload = {
note: '测试文件类型检查基线:只允许下降,不允许上升。用 node scripts/typecheck-tests.cjs --update 下调。',
total: total(counts),
files
}
fs.mkdirSync(path.dirname(baselinePath), { recursive: true })
fs.writeFileSync(baselinePath, `${JSON.stringify(payload, null, 2)}\n`, 'utf8')
console.log(
`[typecheck:test] 基线已更新:${payload.total} 个错误 / ${Object.keys(files).length} 个文件`
)
return
}
const baseline = readBaseline()
if (!baseline) {
console.error(
`[typecheck:test] 找不到基线 ${path.relative(projectRoot, baselinePath)}。\n` +
' 首次启用请运行:node scripts/typecheck-tests.cjs --update'
)
process.exit(1)
}
const regressions = []
for (const [file, count] of counts) {
const allowed = baseline.files[file] ?? 0
if (count > allowed) regressions.push({ file, count, allowed })
}
const now = total(counts)
const baselineTotal = Number(baseline.total) || 0
const improved = baselineTotal - now
if (regressions.length > 0) {
console.error('[typecheck:test] 测试文件出现新的类型错误 ❌')
console.error(` 基线 ${baselineTotal} → 当前 ${now}(+${now - baselineTotal})\n`)
let printed = 0
for (const item of regressions) {
console.error(` ${item.file} ${item.allowed} → ${item.count}`)
for (const line of details.get(item.file) ?? []) {
if (printed >= MAX_REPORTED) break
console.error(` ${line}`)
printed += 1
}
}
console.error('\n 修好之后用 --update 下调基线(不要为了过检查而放宽它)。')
process.exit(1)
}
console.log(
`[typecheck:test] PASS ✅ 当前 ${now} 个既有类型错误 / ${counts.size} 个文件` +
(improved > 0 ? `(比基线少 ${improved} 个,可运行 --update 下调)` : '')
)
console.log(' 注意:这是棘轮,只保证「不新增」。修完历史债后可改为阻断式检查。')
}
main()
-21
View File
@@ -1,21 +0,0 @@
MIT License
Copyright (c) 2026 fastclaw-ai
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
-25
View File
@@ -1,25 +0,0 @@
# TraceMemo WeChat Connector
This repository-local service provides the minimal WeChat bridge required by TraceMemo:
- QR-code login with a single persisted credential
- account discovery
- inbound long polling and authenticated webhook delivery
- local HTTP health and send endpoints
- text and local/remote media sending
The executable is managed by the Electron main process. It is not a general-purpose agent runtime and does not load external AI command-line tools.
## Commands
```bash
go run . login --json
go run . accounts --json
go run . start --foreground --api-addr 127.0.0.1:18011 --account-id <account-id>
```
Credential and synchronization state is stored under `~/.wechatexplorer/wechat-connector/accounts`. This legacy directory name is intentionally retained so upgrades can reuse existing accounts. A successful login is written before the older credential and synchronization state are removed, so an incomplete login cannot destroy the last working credential.
## Attribution
Low-level protocol and media transport portions are distributed under the MIT license in [LICENSE](LICENSE). TraceMemo-specific process management, webhook contract, product UI, and Agent Hub behavior live in the surrounding TraceMemo project.
-135
View File
@@ -1,135 +0,0 @@
package api
import (
"context"
"encoding/json"
"fmt"
"log"
"net/http"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/messaging"
)
// Server provides an HTTP API for sending messages.
type Server struct {
clients []*ilink.Client
addr string
}
// NewServer creates an API server.
func NewServer(clients []*ilink.Client, addr string) *Server {
if addr == "" {
addr = "127.0.0.1:18011"
}
return &Server{clients: clients, addr: addr}
}
// SendRequest is the JSON body for POST /api/send.
type SendRequest struct {
AccountID string `json:"account_id,omitempty"`
To string `json:"to"`
Text string `json:"text,omitempty"`
MediaURL string `json:"media_url,omitempty"` // image/video/file URL
}
// Run starts the HTTP server. Blocks until ctx is cancelled.
func (s *Server) Run(ctx context.Context) error {
mux := http.NewServeMux()
mux.HandleFunc("/api/send", s.handleSend)
mux.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
fmt.Fprintln(w, "ok")
})
srv := &http.Server{Addr: s.addr, Handler: mux}
go func() {
<-ctx.Done()
srv.Shutdown(context.Background())
}()
log.Printf("[api] listening on %s", s.addr)
if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
return err
}
return nil
}
func (s *Server) handleSend(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost {
http.Error(w, "POST only", http.StatusMethodNotAllowed)
return
}
var req SendRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
http.Error(w, "invalid JSON: "+err.Error(), http.StatusBadRequest)
return
}
if req.To == "" {
http.Error(w, `"to" is required`, http.StatusBadRequest)
return
}
if req.Text == "" && req.MediaURL == "" {
http.Error(w, `"text" or "media_url" is required`, http.StatusBadRequest)
return
}
if len(s.clients) == 0 {
http.Error(w, "no accounts configured", http.StatusServiceUnavailable)
return
}
client := s.clientForAccount(req.AccountID)
if client == nil {
http.Error(w, "requested account is not available", http.StatusNotFound)
return
}
ctx := r.Context()
// Send text if provided
if req.Text != "" {
if err := messaging.SendTextReply(ctx, client, req.To, req.Text, "", ""); err != nil {
log.Printf("[api] send text failed: %v", err)
http.Error(w, "send text failed: "+err.Error(), http.StatusInternalServerError)
return
}
log.Printf("[api] sent text to %s: %q", req.To, req.Text)
// Extract and send any markdown images embedded in text
for _, imgURL := range messaging.ExtractImageURLs(req.Text) {
if err := messaging.SendMediaFromURL(ctx, client, req.To, imgURL, ""); err != nil {
log.Printf("[api] send extracted image failed: %v", err)
} else {
log.Printf("[api] sent extracted image to %s: %s", req.To, imgURL)
}
}
}
// Send media if provided
if req.MediaURL != "" {
if err := messaging.SendMediaFromURL(ctx, client, req.To, req.MediaURL, ""); err != nil {
log.Printf("[api] send media failed: %v", err)
http.Error(w, "send media failed: "+err.Error(), http.StatusInternalServerError)
return
}
log.Printf("[api] sent media to %s: %s", req.To, req.MediaURL)
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]string{"status": "ok"})
}
func (s *Server) clientForAccount(accountID string) *ilink.Client {
if accountID == "" {
return s.clients[0]
}
for _, client := range s.clients {
if client.BotID() == accountID {
return client
}
}
return nil
}
@@ -1,20 +0,0 @@
package api
import (
"testing"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
)
func TestClientForAccountSelectsMatchingBot(t *testing.T) {
oldClient := ilink.NewClient(&ilink.Credentials{ILinkBotID: "bot-old"})
newClient := ilink.NewClient(&ilink.Credentials{ILinkBotID: "bot-new"})
server := NewServer([]*ilink.Client{oldClient, newClient}, "")
if got := server.clientForAccount("bot-new"); got != newClient {
t.Fatal("clientForAccount did not select the requested account")
}
if got := server.clientForAccount("missing"); got != nil {
t.Fatal("clientForAccount should reject an unknown account")
}
}
-8
View File
@@ -1,8 +0,0 @@
module github.com/Wxw-Gu/WechatExplorer/services/wechat-connector
go 1.23.0
require (
github.com/google/uuid v1.6.0
rsc.io/qr v0.2.0
)
-4
View File
@@ -1,4 +0,0 @@
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
rsc.io/qr v0.2.0 h1:6vBLea5/NRMVTz8V66gipeLycZMl/+UlFmk8DvqQ6WY=
rsc.io/qr v0.2.0/go.mod h1:IF+uZjkb9fqyeF/4tlBoynqmQxUoPfWEKh921coOuXs=
-236
View File
@@ -1,236 +0,0 @@
package ilink
import (
"context"
"encoding/json"
"fmt"
"os"
"path/filepath"
"strings"
"time"
)
const (
qrCodeURL = "https://ilinkai.weixin.qq.com/ilink/bot/get_bot_qrcode?bot_type=3"
qrStatusURL = "https://ilinkai.weixin.qq.com/ilink/bot/get_qrcode_status?qrcode="
statusWait = "wait"
statusScanned = "scaned"
statusConfirmed = "confirmed"
statusExpired = "expired"
)
// FetchQRCode retrieves a new QR code for login.
func FetchQRCode(ctx context.Context) (*QRCodeResponse, error) {
c := NewUnauthenticatedClient()
var resp QRCodeResponse
if err := c.doGet(ctx, qrCodeURL, &resp); err != nil {
return nil, fmt.Errorf("fetch QR code: %w", err)
}
return &resp, nil
}
// PollQRStatus polls for QR code scan status until confirmed or expired.
// It calls onStatus for each status change so the caller can display progress.
func PollQRStatus(ctx context.Context, qrcode string, onStatus func(status string)) (*Credentials, error) {
c := NewUnauthenticatedClient()
url := qrStatusURL + qrcode
for {
select {
case <-ctx.Done():
return nil, ctx.Err()
default:
}
pollCtx, cancel := context.WithTimeout(ctx, 40*time.Second)
var resp QRStatusResponse
err := c.doGet(pollCtx, url, &resp)
cancel()
if err != nil {
// Timeout is normal for long-poll, retry
if ctx.Err() != nil {
return nil, ctx.Err()
}
continue
}
if onStatus != nil {
onStatus(resp.Status)
}
switch resp.Status {
case statusConfirmed:
creds := &Credentials{
BotToken: resp.BotToken,
ILinkBotID: resp.ILinkBotID,
BaseURL: resp.BaseURL,
ILinkUserID: resp.ILinkUserID,
}
return creds, nil
case statusExpired:
return nil, fmt.Errorf("QR code expired")
case statusWait, statusScanned:
// Continue polling
default:
// Unknown status, continue
}
}
}
func accountsDir(rootName string) (string, error) {
home, err := os.UserHomeDir()
if err != nil {
return "", err
}
return filepath.Join(home, rootName, "wechat-connector", "accounts"), nil
}
// AccountsDir returns the TraceMemo directory where new credentials are stored.
func AccountsDir() (string, error) {
return accountsDir(".tracememo")
}
// LegacyAccountsDir is read-only compatibility for v2.1.9 and earlier.
func LegacyAccountsDir() (string, error) {
return accountsDir(".wechatexplorer")
}
func accountDirectoryForID(accountID string) (string, error) {
current, err := AccountsDir()
if err != nil {
return "", err
}
if _, err := os.Stat(filepath.Join(current, accountID+".json")); err == nil {
return current, nil
}
legacy, err := LegacyAccountsDir()
if err != nil {
return "", err
}
if _, err := os.Stat(filepath.Join(legacy, accountID+".json")); err == nil {
return legacy, nil
}
return current, nil
}
// NormalizeAccountID converts raw bot ID to filesystem-safe format.
func NormalizeAccountID(raw string) string {
s := raw
for _, ch := range []string{"@", ".", ":"} {
s = filepath.Clean(s)
s = replaceAll(s, ch, "-")
}
return s
}
func replaceAll(s, old, new string) string {
for {
i := indexOf(s, old)
if i < 0 {
return s
}
s = s[:i] + new + s[i+len(old):]
}
}
func indexOf(s, sub string) int {
for i := range s {
if i+len(sub) <= len(s) && s[i:i+len(sub)] == sub {
return i
}
}
return -1
}
// SaveCredentials saves the latest credentials and removes older accounts.
// The new credential is written first so a failed login never destroys the
// previously working credential.
func SaveCredentials(creds *Credentials) error {
dir, err := AccountsDir()
if err != nil {
return err
}
if err := os.MkdirAll(dir, 0o700); err != nil {
return fmt.Errorf("create accounts dir: %w", err)
}
id := NormalizeAccountID(creds.ILinkBotID)
path := filepath.Join(dir, id+".json")
data, err := json.MarshalIndent(creds, "", " ")
if err != nil {
return fmt.Errorf("marshal credentials: %w", err)
}
if err := os.WriteFile(path, data, 0o600); err != nil {
return fmt.Errorf("write credentials: %w", err)
}
entries, err := os.ReadDir(dir)
if err != nil {
return fmt.Errorf("prune old credentials: %w", err)
}
keepPrefix := id + "."
for _, entry := range entries {
if entry.IsDir() || strings.HasPrefix(entry.Name(), keepPrefix) {
continue
}
if filepath.Ext(entry.Name()) != ".json" {
continue
}
if err := os.Remove(filepath.Join(dir, entry.Name())); err != nil && !os.IsNotExist(err) {
return fmt.Errorf("remove old credential %s: %w", entry.Name(), err)
}
}
return nil
}
func loadCredentialsFromDir(dir string) ([]*Credentials, error) {
entries, err := os.ReadDir(dir)
if err != nil {
if os.IsNotExist(err) {
return nil, nil
}
return nil, fmt.Errorf("read accounts dir: %w", err)
}
var result []*Credentials
for _, e := range entries {
if e.IsDir() || filepath.Ext(e.Name()) != ".json" {
continue
}
data, err := os.ReadFile(filepath.Join(dir, e.Name()))
if err != nil {
continue
}
var creds Credentials
if json.Unmarshal(data, &creds) == nil && creds.BotToken != "" {
result = append(result, &creds)
}
}
return result, nil
}
// LoadAllCredentials loads TraceMemo credentials first and falls back to the
// untouched WechatExplorer directory for one-version upgrade compatibility.
func LoadAllCredentials() ([]*Credentials, error) {
current, err := AccountsDir()
if err != nil {
return nil, err
}
credentials, err := loadCredentialsFromDir(current)
if err != nil || len(credentials) > 0 {
return credentials, err
}
legacy, err := LegacyAccountsDir()
if err != nil {
return nil, err
}
return loadCredentialsFromDir(legacy)
}
// CredentialsPath returns the path for display purposes.
func CredentialsPath() (string, error) {
return AccountsDir()
}
@@ -1,86 +0,0 @@
package ilink
import (
"encoding/json"
"os"
"path/filepath"
"testing"
)
func setTestHome(t *testing.T) string {
t.Helper()
home := t.TempDir()
t.Setenv("HOME", home)
t.Setenv("USERPROFILE", home)
return home
}
func TestSaveCredentialsKeepsOnlyLatestAccount(t *testing.T) {
setTestHome(t)
old := &Credentials{ILinkBotID: "bot-old@im.bot", BotToken: "old-token"}
latest := &Credentials{ILinkBotID: "bot-new@im.bot", BotToken: "new-token"}
if err := SaveCredentials(old); err != nil {
t.Fatal(err)
}
dir, err := AccountsDir()
if err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, NormalizeAccountID(old.ILinkBotID)+".sync.json"), []byte(`{}`), 0o600); err != nil {
t.Fatal(err)
}
if err := SaveCredentials(latest); err != nil {
t.Fatal(err)
}
accounts, err := LoadAllCredentials()
if err != nil {
t.Fatal(err)
}
if len(accounts) != 1 || accounts[0].ILinkBotID != latest.ILinkBotID {
t.Fatalf("accounts = %#v", accounts)
}
if _, err := os.Stat(filepath.Join(dir, NormalizeAccountID(old.ILinkBotID)+".sync.json")); !os.IsNotExist(err) {
t.Fatalf("old sync state still exists: %v", err)
}
}
func TestAccountsDirUsesTraceMemoIdentity(t *testing.T) {
home := setTestHome(t)
dir, err := AccountsDir()
if err != nil {
t.Fatal(err)
}
want := filepath.Join(home, ".tracememo", "wechat-connector", "accounts")
if dir != want {
t.Fatalf("AccountsDir() = %q, want %q", dir, want)
}
}
func TestLoadAllCredentialsFallsBackToLegacyDirectory(t *testing.T) {
setTestHome(t)
legacyDir, err := LegacyAccountsDir()
if err != nil {
t.Fatal(err)
}
if err := os.MkdirAll(legacyDir, 0o700); err != nil {
t.Fatal(err)
}
legacy := &Credentials{ILinkBotID: "legacy@im.bot", BotToken: "legacy-token"}
data, err := json.Marshal(legacy)
if err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(legacyDir, NormalizeAccountID(legacy.ILinkBotID)+".json"), data, 0o600); err != nil {
t.Fatal(err)
}
accounts, err := LoadAllCredentials()
if err != nil {
t.Fatal(err)
}
if len(accounts) != 1 || accounts[0].BotToken != legacy.BotToken {
t.Fatalf("accounts = %#v", accounts)
}
if _, err := os.Stat(legacyDir); err != nil {
t.Fatalf("legacy directory changed or removed: %v", err)
}
}
-218
View File
@@ -1,218 +0,0 @@
package ilink
import (
"bytes"
"context"
"crypto/rand"
"encoding/base64"
"encoding/binary"
"encoding/json"
"fmt"
"io"
"net/http"
"time"
)
const (
defaultBaseURL = "https://ilinkai.weixin.qq.com"
longPollTimeout = 35 * time.Second
sendTimeout = 15 * time.Second
)
// Client is an iLink HTTP API client.
type Client struct {
baseURL string
botToken string
botID string
httpClient *http.Client
wechatUIN string
}
// NewClient creates a new iLink API client.
func NewClient(creds *Credentials) *Client {
baseURL := creds.BaseURL
if baseURL == "" {
baseURL = defaultBaseURL
}
return &Client{
baseURL: baseURL,
botToken: creds.BotToken,
botID: creds.ILinkBotID,
httpClient: &http.Client{},
wechatUIN: generateWechatUIN(),
}
}
// NewUnauthenticatedClient creates a client without credentials for login flow.
func NewUnauthenticatedClient() *Client {
return &Client{
baseURL: defaultBaseURL,
httpClient: &http.Client{Timeout: 40 * time.Second},
wechatUIN: generateWechatUIN(),
}
}
// BotID returns the bot's user ID.
func (c *Client) BotID() string {
return c.botID
}
// GetUpdates performs a long-poll for new messages.
func (c *Client) GetUpdates(ctx context.Context, buf string) (*GetUpdatesResponse, error) {
reqBody := GetUpdatesRequest{
GetUpdatesBuf: buf,
BaseInfo: BaseInfo{ChannelVersion: "1.0.0"},
}
ctx, cancel := context.WithTimeout(ctx, longPollTimeout+5*time.Second)
defer cancel()
var resp GetUpdatesResponse
if err := c.doPost(ctx, "/ilink/bot/getupdates", reqBody, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// SendMessage sends a message through iLink.
func (c *Client) SendMessage(ctx context.Context, msg *SendMessageRequest) (*SendMessageResponse, error) {
ctx, cancel := context.WithTimeout(ctx, sendTimeout)
defer cancel()
var resp SendMessageResponse
if err := c.doPost(ctx, "/ilink/bot/sendmessage", msg, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// GetConfig fetches bot config for a user (includes typing_ticket).
func (c *Client) GetConfig(ctx context.Context, userID, contextToken string) (*GetConfigResponse, error) {
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
req := GetConfigRequest{
ILinkUserID: userID,
ContextToken: contextToken,
BaseInfo: BaseInfo{},
}
var resp GetConfigResponse
if err := c.doPost(ctx, "/ilink/bot/getconfig", req, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// SendTyping sends a typing indicator to a user.
func (c *Client) SendTyping(ctx context.Context, userID, typingTicket string, status int) error {
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
req := SendTypingRequest{
ILinkUserID: userID,
TypingTicket: typingTicket,
Status: status,
BaseInfo: BaseInfo{},
}
var resp SendTypingResponse
if err := c.doPost(ctx, "/ilink/bot/sendtyping", req, &resp); err != nil {
return err
}
if resp.Ret != 0 {
return fmt.Errorf("sendtyping failed: ret=%d errmsg=%s", resp.Ret, resp.ErrMsg)
}
return nil
}
// GetUploadURL gets a pre-signed CDN upload URL for media files.
func (c *Client) GetUploadURL(ctx context.Context, req *GetUploadURLRequest) (*GetUploadURLResponse, error) {
ctx, cancel := context.WithTimeout(ctx, sendTimeout)
defer cancel()
var resp GetUploadURLResponse
if err := c.doPost(ctx, "/ilink/bot/getuploadurl", req, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// BaseURL returns the base URL for CDN operations.
func (c *Client) BaseURL() string {
return c.baseURL
}
func (c *Client) doPost(ctx context.Context, path string, body interface{}, result interface{}) error {
data, err := json.Marshal(body)
if err != nil {
return fmt.Errorf("marshal request: %w", err)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL+path, bytes.NewReader(data))
if err != nil {
return fmt.Errorf("create request: %w", err)
}
c.setHeaders(req)
resp, err := c.httpClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
respBody, err := io.ReadAll(resp.Body)
if err != nil {
return fmt.Errorf("read response: %w", err)
}
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("HTTP %d: %s", resp.StatusCode, string(respBody))
}
if err := json.Unmarshal(respBody, result); err != nil {
return fmt.Errorf("unmarshal response: %w", err)
}
return nil
}
func (c *Client) doGet(ctx context.Context, url string, result interface{}) error {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return fmt.Errorf("create request: %w", err)
}
resp, err := c.httpClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
respBody, err := io.ReadAll(resp.Body)
if err != nil {
return fmt.Errorf("read response: %w", err)
}
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("HTTP %d: %s", resp.StatusCode, string(respBody))
}
if err := json.Unmarshal(respBody, result); err != nil {
return fmt.Errorf("unmarshal response: %w", err)
}
return nil
}
func (c *Client) setHeaders(req *http.Request) {
req.Header.Set("Content-Type", "application/json")
req.Header.Set("AuthorizationType", "ilink_bot_token")
req.Header.Set("Authorization", "Bearer "+c.botToken)
req.Header.Set("X-WECHAT-UIN", c.wechatUIN)
}
func generateWechatUIN() string {
var n uint32
_ = binary.Read(rand.Reader, binary.LittleEndian, &n)
s := fmt.Sprintf("%d", n)
return base64.StdEncoding.EncodeToString([]byte(s))
}
-181
View File
@@ -1,181 +0,0 @@
package ilink
import (
"context"
"encoding/json"
"fmt"
"log"
"os"
"path/filepath"
"time"
)
const (
maxConsecutiveFailures = 5
initialBackoff = 3 * time.Second
maxBackoff = 60 * time.Second
sessionExpiredBackoff = 5 * time.Second
errCodeSessionExpired = -14
)
// MessageHandler is called for each received message.
type MessageHandler func(ctx context.Context, client *Client, msg WeixinMessage)
// Monitor manages the long-poll loop for receiving messages.
type Monitor struct {
client *Client
handler MessageHandler
getUpdatesBuf string
bufPath string
failures int
lastActivity time.Time
}
// NewMonitor creates a new long-poll monitor.
func NewMonitor(client *Client, handler MessageHandler) (*Monitor, error) {
accountID := NormalizeAccountID(client.BotID())
accountsRoot, err := accountDirectoryForID(accountID)
if err != nil {
return nil, err
}
bufPath := filepath.Join(accountsRoot, accountID+".sync.json")
m := &Monitor{
client: client,
handler: handler,
bufPath: bufPath,
lastActivity: time.Now(),
}
m.loadBuf()
return m, nil
}
// Run starts the long-poll loop. It blocks until ctx is cancelled.
// Automatically recovers from errors with exponential backoff.
func (m *Monitor) Run(ctx context.Context) error {
log.Println("[monitor] starting long-poll loop")
for {
select {
case <-ctx.Done():
log.Println("[monitor] shutting down")
return ctx.Err()
default:
}
resp, err := m.client.GetUpdates(ctx, m.getUpdatesBuf)
if err != nil {
if ctx.Err() != nil {
return ctx.Err()
}
m.failures++
backoff := m.calcBackoff()
log.Printf("[monitor] GetUpdates error (%d/%d, backoff=%s): %v",
m.failures, maxConsecutiveFailures, backoff, err)
if m.failures == maxConsecutiveFailures {
log.Printf("[monitor] WARNING: %d consecutive failures; reconnect from TraceMemo if this persists.", maxConsecutiveFailures)
}
select {
case <-time.After(backoff):
case <-ctx.Done():
return ctx.Err()
}
continue
}
// Reset failure counter on any successful response
m.failures = 0
m.lastActivity = time.Now()
// Session expired — reset sync buf and reconnect silently
if resp.ErrCode == errCodeSessionExpired {
if m.getUpdatesBuf != "" {
log.Printf("[monitor] session expired, resetting sync buf")
m.getUpdatesBuf = ""
m.saveBuf()
} else {
// Sync buf already empty but still getting session expired:
// the bot token itself has expired. The user needs to re-login.
log.Printf("[monitor] WARNING: WeChat session expired and cannot be auto-recovered; reconnect from TraceMemo.")
}
select {
case <-time.After(sessionExpiredBackoff):
case <-ctx.Done():
return ctx.Err()
}
continue
}
// Other server errors
if resp.Ret != 0 && resp.ErrCode != 0 {
log.Printf("[monitor] server error: ret=%d errcode=%d errmsg=%s", resp.Ret, resp.ErrCode, resp.ErrMsg)
continue
}
// Update buf for next poll
if resp.GetUpdatesBuf != "" {
m.getUpdatesBuf = resp.GetUpdatesBuf
m.saveBuf()
}
// Process messages concurrently — don't block the poll loop
for _, msg := range resp.Msgs {
go m.handler(ctx, m.client, msg)
}
}
}
// calcBackoff returns an exponential backoff duration capped at maxBackoff.
func (m *Monitor) calcBackoff() time.Duration {
d := initialBackoff
for i := 1; i < m.failures; i++ {
d *= 2
if d > maxBackoff {
return maxBackoff
}
}
return d
}
type syncData struct {
GetUpdatesBuf string `json:"get_updates_buf"`
}
func (m *Monitor) loadBuf() {
data, err := os.ReadFile(m.bufPath)
if err != nil {
return
}
var s syncData
if json.Unmarshal(data, &s) == nil && s.GetUpdatesBuf != "" {
m.getUpdatesBuf = s.GetUpdatesBuf
log.Printf("[monitor] loaded sync buf from %s", m.bufPath)
}
}
func (m *Monitor) saveBuf() {
dir := filepath.Dir(m.bufPath)
if err := os.MkdirAll(dir, 0o700); err != nil {
log.Printf("[monitor] failed to create buf dir: %v", err)
return
}
data, _ := json.Marshal(syncData{GetUpdatesBuf: m.getUpdatesBuf})
if err := os.WriteFile(m.bufPath, data, 0o600); err != nil {
log.Printf("[monitor] failed to save buf: %v", err)
}
}
// FormatMessageSummary returns a short description of a message for logging.
func FormatMessageSummary(msg WeixinMessage) string {
text := ""
for _, item := range msg.ItemList {
if item.Type == ItemTypeText && item.TextItem != nil {
text = item.TextItem.Text
break
}
}
if len(text) > 50 {
text = text[:50] + "..."
}
return fmt.Sprintf("from=%s type=%d state=%d text=%q", msg.FromUserID, msg.MessageType, msg.MessageState, text)
}
-219
View File
@@ -1,219 +0,0 @@
package ilink
// Message types
const (
MessageTypeNone = 0
MessageTypeUser = 1
MessageTypeBot = 2
)
// Message states
const (
MessageStateNew = 0
MessageStateGenerating = 1
MessageStateFinish = 2
)
// Item types
const (
ItemTypeNone = 0
ItemTypeText = 1
ItemTypeImage = 2
ItemTypeVoice = 3
ItemTypeFile = 4
ItemTypeVideo = 5
)
// QRCodeResponse is the response from get_bot_qrcode.
type QRCodeResponse struct {
QRCode string `json:"qrcode"`
QRCodeImgContent string `json:"qrcode_img_content"`
}
// QRStatusResponse is the response from get_qrcode_status.
type QRStatusResponse struct {
Status string `json:"status"`
BotToken string `json:"bot_token"`
ILinkBotID string `json:"ilink_bot_id"`
BaseURL string `json:"baseurl"`
ILinkUserID string `json:"ilink_user_id"`
}
// Credentials stores login session data.
type Credentials struct {
BotToken string `json:"bot_token"`
ILinkBotID string `json:"ilink_bot_id"`
BaseURL string `json:"baseurl"`
ILinkUserID string `json:"ilink_user_id"`
}
// BaseInfo is included in request bodies.
type BaseInfo struct {
ChannelVersion string `json:"channel_version,omitempty"`
}
// GetUpdatesRequest is the body for getupdates.
type GetUpdatesRequest struct {
GetUpdatesBuf string `json:"get_updates_buf"`
BaseInfo BaseInfo `json:"base_info"`
}
// GetUpdatesResponse is the response from getupdates.
type GetUpdatesResponse struct {
Ret int `json:"ret"`
ErrCode int `json:"errcode,omitempty"`
ErrMsg string `json:"errmsg,omitempty"`
Msgs []WeixinMessage `json:"msgs"`
GetUpdatesBuf string `json:"get_updates_buf"`
LongPollingTimeoutMs int `json:"longpolling_timeout_ms,omitempty"`
}
// WeixinMessage represents a message from WeChat.
type WeixinMessage struct {
Seq int `json:"seq,omitempty"`
MessageID int64 `json:"message_id,omitempty"`
FromUserID string `json:"from_user_id"`
ToUserID string `json:"to_user_id"`
MessageType int `json:"message_type"`
MessageState int `json:"message_state"`
ItemList []MessageItem `json:"item_list"`
ContextToken string `json:"context_token"`
}
// MessageItem is a single item in a message.
type MessageItem struct {
Type int `json:"type"`
TextItem *TextItem `json:"text_item,omitempty"`
ImageItem *ImageItem `json:"image_item,omitempty"`
VoiceItem *VoiceItem `json:"voice_item,omitempty"`
VideoItem *VideoItem `json:"video_item,omitempty"`
FileItem *FileItem `json:"file_item,omitempty"`
}
// CDN media type constants.
const (
CDNMediaTypeImage = 1
CDNMediaTypeVideo = 2
CDNMediaTypeFile = 3
)
// GetUploadURLRequest is the body for getuploadurl.
type GetUploadURLRequest struct {
FileKey string `json:"filekey"`
MediaType int `json:"media_type"`
ToUserID string `json:"to_user_id"`
RawSize int `json:"rawsize"`
RawFileMD5 string `json:"rawfilemd5"`
FileSize int `json:"filesize"`
NoNeedThumb bool `json:"no_need_thumb"`
AESKey string `json:"aeskey"`
BaseInfo BaseInfo `json:"base_info"`
}
// GetUploadURLResponse is the response from getuploadurl.
type GetUploadURLResponse struct {
Ret int `json:"ret"`
ErrMsg string `json:"errmsg,omitempty"`
UploadParam string `json:"upload_param"`
UploadFullURL string `json:"upload_full_url,omitempty"`
}
// TextItem holds text content.
type TextItem struct {
Text string `json:"text"`
}
// MediaInfo holds CDN media reference for uploaded files.
type MediaInfo struct {
EncryptQueryParam string `json:"encrypt_query_param"`
AESKey string `json:"aes_key"` // base64-encoded
EncryptType int `json:"encrypt_type"` // 1 = AES-128-ECB
}
// VoiceItem holds voice content.
type VoiceItem struct {
Media *MediaInfo `json:"media,omitempty"`
VoiceSize int `json:"voice_size,omitempty"`
EncodeType int `json:"encode_type,omitempty"` // 1=pcm 2=adpcm 3=feature 4=speex 5=amr 6=silk 7=mp3
BitsPerSample int `json:"bits_per_sample,omitempty"`
SampleRate int `json:"sample_rate,omitempty"` // Hz
Playtime int `json:"playtime,omitempty"` // duration in milliseconds
Text string `json:"text,omitempty"` // speech-to-text transcription from WeChat
}
// ImageItem holds image content.
type ImageItem struct {
URL string `json:"url,omitempty"`
Media *MediaInfo `json:"media,omitempty"`
MidSize int `json:"mid_size,omitempty"` // ciphertext size
}
// VideoItem holds video content.
type VideoItem struct {
Media *MediaInfo `json:"media,omitempty"`
VideoSize int `json:"video_size,omitempty"`
}
// FileItem holds file content.
type FileItem struct {
Media *MediaInfo `json:"media,omitempty"`
FileName string `json:"file_name,omitempty"`
Len string `json:"len,omitempty"` // plaintext size as string
}
// SendMessageRequest is the body for sendmessage.
type SendMessageRequest struct {
Msg SendMsg `json:"msg"`
BaseInfo BaseInfo `json:"base_info"`
}
// SendMsg is the message payload for sending.
type SendMsg struct {
FromUserID string `json:"from_user_id"`
ToUserID string `json:"to_user_id"`
ClientID string `json:"client_id"`
MessageType int `json:"message_type"`
MessageState int `json:"message_state"`
ItemList []MessageItem `json:"item_list"`
ContextToken string `json:"context_token"`
}
// SendMessageResponse is the response from sendmessage.
type SendMessageResponse struct {
Ret int `json:"ret"`
ErrMsg string `json:"errmsg,omitempty"`
}
// Typing status constants.
const (
TypingStatusTyping = 1
TypingStatusCancel = 2
)
// GetConfigRequest is the body for getconfig.
type GetConfigRequest struct {
ILinkUserID string `json:"ilink_user_id"`
ContextToken string `json:"context_token,omitempty"`
BaseInfo BaseInfo `json:"base_info"`
}
// GetConfigResponse is the response from getconfig.
type GetConfigResponse struct {
Ret int `json:"ret"`
ErrMsg string `json:"errmsg,omitempty"`
TypingTicket string `json:"typing_ticket,omitempty"`
}
// SendTypingRequest is the body for sendtyping.
type SendTypingRequest struct {
ILinkUserID string `json:"ilink_user_id"`
TypingTicket string `json:"typing_ticket"`
Status int `json:"status"`
BaseInfo BaseInfo `json:"base_info"`
}
// SendTypingResponse is the response from sendtyping.
type SendTypingResponse struct {
Ret int `json:"ret"`
ErrMsg string `json:"errmsg,omitempty"`
}
-200
View File
@@ -1,200 +0,0 @@
package main
import (
"context"
"encoding/base64"
"encoding/json"
"errors"
"flag"
"fmt"
"log"
"os"
"os/signal"
"strings"
"sync"
"syscall"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/api"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/messaging"
"rsc.io/qr"
)
type loginEvent struct {
Status string `json:"status"`
QRCodeDataURL string `json:"qr_code_data_url,omitempty"`
AccountID string `json:"account_id,omitempty"`
WeChatUserID string `json:"wechat_user_id,omitempty"`
}
type accountSummary struct {
AccountID string `json:"account_id"`
WeChatUserID string `json:"wechat_user_id"`
}
func main() {
if len(os.Args) < 2 {
fatal(errors.New("expected one of: login, accounts, start"))
}
var err error
switch os.Args[1] {
case "login":
err = runLogin(os.Args[2:])
case "accounts":
err = runAccounts(os.Args[2:])
case "start":
err = runStart(os.Args[2:])
default:
err = fmt.Errorf("unknown command %q", os.Args[1])
}
if err != nil {
fatal(err)
}
}
func fatal(err error) {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
func signalContext() (context.Context, context.CancelFunc) {
return signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
}
func runLogin(args []string) error {
flags := flag.NewFlagSet("login", flag.ContinueOnError)
jsonOutput := flags.Bool("json", false, "emit JSON Lines events")
if err := flags.Parse(args); err != nil {
return err
}
ctx, cancel := signalContext()
defer cancel()
creds, err := login(ctx, *jsonOutput)
if err != nil {
return err
}
if !*jsonOutput {
fmt.Printf("WeChat account %s connected.\n", creds.ILinkBotID)
}
return nil
}
func login(ctx context.Context, jsonOutput bool) (*ilink.Credentials, error) {
qrResponse, err := ilink.FetchQRCode(ctx)
if err != nil {
return nil, err
}
code, err := qr.Encode(qrResponse.QRCodeImgContent, qr.L)
if err != nil {
return nil, fmt.Errorf("encode QR image: %w", err)
}
emit := func(event loginEvent) {
if jsonOutput {
_ = json.NewEncoder(os.Stdout).Encode(event)
}
}
emit(loginEvent{Status: "qrcode", QRCodeDataURL: "data:image/png;base64," + base64.StdEncoding.EncodeToString(code.PNG())})
lastStatus := ""
creds, err := ilink.PollQRStatus(ctx, qrResponse.QRCode, func(status string) {
if status != lastStatus {
lastStatus = status
emit(loginEvent{Status: status})
}
})
if err != nil {
return nil, err
}
if err := ilink.SaveCredentials(creds); err != nil {
return nil, fmt.Errorf("save credentials: %w", err)
}
emit(loginEvent{Status: "active", AccountID: creds.ILinkBotID, WeChatUserID: creds.ILinkUserID})
return creds, nil
}
func runAccounts(args []string) error {
flags := flag.NewFlagSet("accounts", flag.ContinueOnError)
jsonOutput := flags.Bool("json", false, "print JSON")
if err := flags.Parse(args); err != nil {
return err
}
accounts, err := ilink.LoadAllCredentials()
if err != nil {
return err
}
items := make([]accountSummary, 0, len(accounts))
for _, account := range accounts {
items = append(items, accountSummary{AccountID: account.ILinkBotID, WeChatUserID: account.ILinkUserID})
}
if *jsonOutput {
return json.NewEncoder(os.Stdout).Encode(map[string]any{"accounts": items})
}
for _, item := range items {
fmt.Printf("%s\t%s\n", item.AccountID, item.WeChatUserID)
}
return nil
}
func runStart(args []string) error {
flags := flag.NewFlagSet("start", flag.ContinueOnError)
_ = flags.Bool("foreground", false, "kept for host compatibility")
apiAddr := flags.String("api-addr", "127.0.0.1:18011", "local send API address")
accountID := flags.String("account-id", "", "account to start")
if err := flags.Parse(args); err != nil {
return err
}
accounts, err := ilink.LoadAllCredentials()
if err != nil {
return err
}
if len(accounts) == 0 {
return errors.New("no connected WeChat account; scan a QR code first")
}
selected := accounts[len(accounts)-1]
if *accountID != "" {
selected = nil
for _, account := range accounts {
if account.ILinkBotID == *accountID {
selected = account
break
}
}
if selected == nil {
return fmt.Errorf("account %q not found", *accountID)
}
}
ctx, cancel := signalContext()
defer cancel()
client := ilink.NewClient(selected)
server := api.NewServer([]*ilink.Client{client}, *apiAddr)
webhookURL := strings.TrimSpace(os.Getenv("WECHAT_CONNECTOR_INBOUND_WEBHOOK_URL"))
webhook := messaging.NewInboundWebhook(webhookURL, os.Getenv("WECHAT_CONNECTOR_INBOUND_WEBHOOK_TOKEN"))
monitor, err := ilink.NewMonitor(client, func(messageContext context.Context, source *ilink.Client, message ilink.WeixinMessage) {
if webhookURL != "" {
webhook.Dispatch(messageContext, source, message)
}
})
if err != nil {
return err
}
var wait sync.WaitGroup
wait.Add(2)
go func() {
defer wait.Done()
if err := server.Run(ctx); err != nil && ctx.Err() == nil {
log.Printf("[api] stopped: %v", err)
cancel()
}
}()
go func() {
defer wait.Done()
if err := monitor.Run(ctx); err != nil && ctx.Err() == nil {
log.Printf("[monitor] stopped: %v", err)
cancel()
}
}()
wait.Wait()
return nil
}
-232
View File
@@ -1,232 +0,0 @@
package messaging
import (
"bytes"
"context"
"crypto/aes"
"crypto/md5"
"crypto/rand"
"encoding/base64"
"encoding/hex"
"fmt"
"io"
"net/http"
"net/url"
"strings"
"time"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
)
const cdnBaseURL = "https://novac2c.cdn.weixin.qq.com/c2c"
// UploadedFile holds the result of a CDN upload.
type UploadedFile struct {
DownloadParam string // encrypted query param for download
AESKeyHex string // hex-encoded AES key
FileSize int // plaintext size
CipherSize int // ciphertext size
}
// UploadFileToCDN encrypts and uploads a file to the WeChat CDN.
func UploadFileToCDN(ctx context.Context, client *ilink.Client, data []byte, toUserID string, mediaType int) (*UploadedFile, error) {
// Generate random filekey and AES key
filekey := make([]byte, 16)
aeskey := make([]byte, 16)
if _, err := rand.Read(filekey); err != nil {
return nil, fmt.Errorf("generate filekey: %w", err)
}
if _, err := rand.Read(aeskey); err != nil {
return nil, fmt.Errorf("generate aeskey: %w", err)
}
filekeyHex := hex.EncodeToString(filekey)
aeskeyHex := hex.EncodeToString(aeskey)
// Calculate MD5 of plaintext
hash := md5.Sum(data)
rawMD5 := hex.EncodeToString(hash[:])
// Calculate ciphertext size (PKCS7 padding)
cipherSize := aesECBPaddedSize(len(data))
// Get upload URL from iLink API
uploadReq := &ilink.GetUploadURLRequest{
FileKey: filekeyHex,
MediaType: mediaType,
ToUserID: toUserID,
RawSize: len(data),
RawFileMD5: rawMD5,
FileSize: cipherSize,
NoNeedThumb: true,
AESKey: aeskeyHex,
BaseInfo: ilink.BaseInfo{},
}
uploadResp, err := client.GetUploadURL(ctx, uploadReq)
if err != nil {
return nil, fmt.Errorf("get upload URL: %w", err)
}
if uploadResp.Ret != 0 {
return nil, fmt.Errorf("get upload URL failed: ret=%d errmsg=%s", uploadResp.Ret, uploadResp.ErrMsg)
}
// Encrypt data with AES-128-ECB
encrypted, err := encryptAESECB(data, aeskey)
if err != nil {
return nil, fmt.Errorf("encrypt: %w", err)
}
// Upload to CDN: prefer server-provided full URL, fall back to param-based construction
cdnURL := strings.TrimSpace(uploadResp.UploadFullURL)
if cdnURL == "" {
if uploadResp.UploadParam == "" {
return nil, fmt.Errorf("getuploadurl returned no upload URL (need upload_full_url or upload_param)")
}
cdnURL = fmt.Sprintf("%s/upload?encrypted_query_param=%s&filekey=%s",
cdnBaseURL, url.QueryEscape(uploadResp.UploadParam), url.QueryEscape(filekeyHex))
}
downloadParam, err := uploadToCDN(ctx, encrypted, cdnURL)
if err != nil {
return nil, fmt.Errorf("CDN upload: %w", err)
}
return &UploadedFile{
DownloadParam: downloadParam,
AESKeyHex: aeskeyHex,
FileSize: len(data),
CipherSize: cipherSize,
}, nil
}
// AESKeyToBase64 converts a hex AES key to base64 format for message items.
func AESKeyToBase64(hexKey string) string {
return base64.StdEncoding.EncodeToString([]byte(hexKey))
}
// DownloadFileFromCDN downloads and decrypts a file from the WeChat CDN.
func DownloadFileFromCDN(ctx context.Context, encryptQueryParam, aesKeyBase64 string) ([]byte, error) {
// Decode AES key: base64 -> hex string -> raw bytes
aesKeyHexBytes, err := base64.StdEncoding.DecodeString(aesKeyBase64)
if err != nil {
return nil, fmt.Errorf("decode AES key base64: %w", err)
}
aesKey, err := hex.DecodeString(string(aesKeyHexBytes))
if err != nil {
return nil, fmt.Errorf("decode AES key hex: %w", err)
}
// Download encrypted data from CDN
downloadURL := fmt.Sprintf("%s/download?encrypted_query_param=%s",
cdnBaseURL, url.QueryEscape(encryptQueryParam))
reqCtx, cancel := context.WithTimeout(ctx, 60*time.Second)
defer cancel()
req, err := http.NewRequestWithContext(reqCtx, http.MethodGet, downloadURL, nil)
if err != nil {
return nil, fmt.Errorf("create download request: %w", err)
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, fmt.Errorf("download from CDN: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
body, _ := io.ReadAll(resp.Body)
return nil, fmt.Errorf("CDN download HTTP %d: %s", resp.StatusCode, string(body))
}
encrypted, err := io.ReadAll(resp.Body)
if err != nil {
return nil, fmt.Errorf("read CDN response: %w", err)
}
// Decrypt AES-128-ECB
return decryptAESECB(encrypted, aesKey)
}
// decryptAESECB decrypts data encrypted with AES-128-ECB and removes PKCS7 padding.
func decryptAESECB(ciphertext, key []byte) ([]byte, error) {
block, err := aes.NewCipher(key)
if err != nil {
return nil, err
}
if len(ciphertext)%aes.BlockSize != 0 {
return nil, fmt.Errorf("ciphertext is not a multiple of block size")
}
plaintext := make([]byte, len(ciphertext))
for i := 0; i < len(ciphertext); i += aes.BlockSize {
block.Decrypt(plaintext[i:i+aes.BlockSize], ciphertext[i:i+aes.BlockSize])
}
// Remove PKCS7 padding
if len(plaintext) == 0 {
return plaintext, nil
}
padLen := int(plaintext[len(plaintext)-1])
if padLen > aes.BlockSize || padLen == 0 {
return nil, fmt.Errorf("invalid PKCS7 padding")
}
return plaintext[:len(plaintext)-padLen], nil
}
func uploadToCDN(ctx context.Context, encrypted []byte, cdnURL string) (string, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, cdnURL, bytes.NewReader(encrypted))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/octet-stream")
client := &http.Client{Timeout: 60 * time.Second}
resp, err := client.Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
body, _ := io.ReadAll(resp.Body)
return "", fmt.Errorf("CDN upload HTTP %d: %s", resp.StatusCode, string(body))
}
downloadParam := resp.Header.Get("X-Encrypted-Param")
if downloadParam == "" {
return "", fmt.Errorf("CDN upload: missing X-Encrypted-Param header")
}
return downloadParam, nil
}
// encryptAESECB encrypts data using AES-128-ECB with PKCS7 padding.
func encryptAESECB(plaintext, key []byte) ([]byte, error) {
block, err := aes.NewCipher(key)
if err != nil {
return nil, err
}
// PKCS7 padding
padLen := aes.BlockSize - (len(plaintext) % aes.BlockSize)
padded := make([]byte, len(plaintext)+padLen)
copy(padded, plaintext)
for i := len(plaintext); i < len(padded); i++ {
padded[i] = byte(padLen)
}
// ECB mode: encrypt each block independently
encrypted := make([]byte, len(padded))
for i := 0; i < len(padded); i += aes.BlockSize {
block.Encrypt(encrypted[i:i+aes.BlockSize], padded[i:i+aes.BlockSize])
}
return encrypted, nil
}
func aesECBPaddedSize(plaintextSize int) int {
return (plaintextSize/aes.BlockSize + 1) * aes.BlockSize
}
@@ -1,121 +0,0 @@
package messaging
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"strings"
"time"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
)
const (
webhookAttempts = 3
webhookTimeout = 5 * time.Second
)
type InboundWebhook struct {
url string
token string
client *http.Client
}
type inboundWebhookPayload struct {
AccountID string `json:"account_id"`
FromUserID string `json:"from_user_id"`
MessageID int64 `json:"message_id"`
MessageType int `json:"message_type"`
Items []inboundWebhookItem `json:"items"`
ReceivedAt time.Time `json:"received_at"`
}
type inboundWebhookItem struct {
Type int `json:"type"`
Text string `json:"text,omitempty"`
}
func NewInboundWebhook(url, token string) *InboundWebhook {
return &InboundWebhook{
url: strings.TrimSpace(url),
token: token,
client: &http.Client{Timeout: webhookTimeout},
}
}
// Dispatch is intentionally non-blocking so webhook failures never stall iLink polling.
func (w *InboundWebhook) Dispatch(ctx context.Context, client *ilink.Client, msg ilink.WeixinMessage) {
payload := normalizeInboundMessage(client.BotID(), msg)
go func() {
if err := w.deliver(ctx, payload); err != nil {
log.Printf("[webhook] inbound delivery failed for message %d: %v", msg.MessageID, err)
}
}()
}
func (w *InboundWebhook) deliver(ctx context.Context, payload inboundWebhookPayload) error {
body, err := json.Marshal(payload)
if err != nil {
return fmt.Errorf("encode payload: %w", err)
}
var lastErr error
for attempt := 1; attempt <= webhookAttempts; attempt++ {
if attempt > 1 {
timer := time.NewTimer(time.Duration(attempt-1) * time.Second)
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
}
}
req, reqErr := http.NewRequestWithContext(ctx, http.MethodPost, w.url, bytes.NewReader(body))
if reqErr != nil {
return fmt.Errorf("create request: %w", reqErr)
}
req.Header.Set("Content-Type", "application/json")
if w.token != "" {
req.Header.Set("Authorization", "Bearer "+w.token)
}
resp, doErr := w.client.Do(req)
if doErr != nil {
lastErr = doErr
continue
}
responseBody, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
resp.Body.Close()
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
lastErr = fmt.Errorf("status %s: %s", resp.Status, strings.TrimSpace(string(responseBody)))
if resp.StatusCode >= 400 && resp.StatusCode < 500 {
break
}
}
return lastErr
}
func normalizeInboundMessage(accountID string, msg ilink.WeixinMessage) inboundWebhookPayload {
items := make([]inboundWebhookItem, 0, len(msg.ItemList))
for _, item := range msg.ItemList {
normalized := inboundWebhookItem{Type: item.Type}
if item.TextItem != nil {
normalized.Text = item.TextItem.Text
} else if item.VoiceItem != nil {
normalized.Text = item.VoiceItem.Text
}
items = append(items, normalized)
}
return inboundWebhookPayload{
AccountID: accountID,
FromUserID: msg.FromUserID,
MessageID: msg.MessageID,
MessageType: msg.MessageType,
Items: items,
ReceivedAt: time.Now().UTC(),
}
}
@@ -1,75 +0,0 @@
package messaging
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"sync/atomic"
"testing"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
)
func TestInboundWebhookDeliversNormalizedPayload(t *testing.T) {
var got inboundWebhookPayload
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.Header.Get("Authorization") != "Bearer secret" {
t.Errorf("authorization = %q", r.Header.Get("Authorization"))
}
if err := json.NewDecoder(r.Body).Decode(&got); err != nil {
t.Errorf("decode: %v", err)
}
w.WriteHeader(http.StatusOK)
}))
defer server.Close()
webhook := NewInboundWebhook(server.URL, "secret")
err := webhook.deliver(context.Background(), normalizeInboundMessage("bot-new", ilink.WeixinMessage{
MessageID: 7, FromUserID: "user-1", MessageType: ilink.MessageTypeUser,
ItemList: []ilink.MessageItem{{Type: ilink.ItemTypeText, TextItem: &ilink.TextItem{Text: "最近5条消息"}}},
}))
if err != nil {
t.Fatalf("deliver: %v", err)
}
if got.AccountID != "bot-new" || got.MessageID != 7 || len(got.Items) != 1 || got.Items[0].Text != "最近5条消息" {
t.Fatalf("payload = %#v", got)
}
}
func TestInboundWebhookRetriesServerErrors(t *testing.T) {
var calls atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
if calls.Add(1) < 3 {
http.Error(w, "temporary", http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
}))
defer server.Close()
webhook := NewInboundWebhook(server.URL, "")
if err := webhook.deliver(context.Background(), inboundWebhookPayload{}); err != nil {
t.Fatalf("deliver: %v", err)
}
if calls.Load() != 3 {
t.Fatalf("calls = %d, want 3", calls.Load())
}
}
func TestInboundWebhookDoesNotRetryClientErrors(t *testing.T) {
var calls atomic.Int32
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
calls.Add(1)
http.Error(w, "unauthorized", http.StatusUnauthorized)
}))
defer server.Close()
webhook := NewInboundWebhook(server.URL, "")
if err := webhook.deliver(context.Background(), inboundWebhookPayload{}); err == nil {
t.Fatal("deliver error = nil")
}
if calls.Load() != 1 {
t.Fatalf("calls = %d, want 1", calls.Load())
}
}
@@ -1,103 +0,0 @@
package messaging
import (
"regexp"
"strings"
)
var (
// Code blocks: strip fences, keep code content
reCodeBlock = regexp.MustCompile("(?s)```[^\n]*\n?(.*?)```")
// Inline code: strip backticks, keep content
reInlineCode = regexp.MustCompile("`([^`]+)`")
// Images: remove entirely
reImage = regexp.MustCompile(`!\[[^\]]*\]\([^)]*\)`)
// Links: keep display text only
reLink = regexp.MustCompile(`\[([^\]]+)\]\([^)]*\)`)
// Table separator rows: remove
reTableSep = regexp.MustCompile(`(?m)^\|[\s:|\-]+\|$`)
// Table rows: convert pipe-delimited to space-delimited
reTableRow = regexp.MustCompile(`(?m)^\|(.+)\|$`)
// Headers: remove # prefix
reHeader = regexp.MustCompile(`(?m)^#{1,6}\s+`)
// Bold: **text** or __text__
reBold = regexp.MustCompile(`\*\*(.+?)\*\*|__(.+?)__`)
// Italic: *text* or _text_
reItalic = regexp.MustCompile(`(?:^|[^*])\*([^*]+)\*(?:[^*]|$)|(?:^|[^_])_([^_]+)_(?:[^_]|$)`)
// Strikethrough: ~~text~~
reStrike = regexp.MustCompile(`~~(.+?)~~`)
// Blockquote: > prefix
reBlockquote = regexp.MustCompile(`(?m)^>\s?`)
// Horizontal rule
reHR = regexp.MustCompile(`(?m)^[-*_]{3,}\s*$`)
// Unordered list markers: -, *, +
reUL = regexp.MustCompile(`(?m)^(\s*)[-*+]\s+`)
)
// MarkdownToPlainText converts markdown to readable plain text for WeChat.
func MarkdownToPlainText(text string) string {
result := text
// Code blocks: strip fences, keep code content
result = reCodeBlock.ReplaceAllStringFunc(result, func(match string) string {
parts := reCodeBlock.FindStringSubmatch(match)
if len(parts) > 1 {
return strings.TrimSpace(parts[1])
}
return match
})
// Images: remove entirely
result = reImage.ReplaceAllString(result, "")
// Links: keep display text only
result = reLink.ReplaceAllString(result, "$1")
// Table separator rows: remove
result = reTableSep.ReplaceAllString(result, "")
// Table rows: pipe-delimited to space-delimited
result = reTableRow.ReplaceAllStringFunc(result, func(match string) string {
parts := reTableRow.FindStringSubmatch(match)
if len(parts) > 1 {
cells := strings.Split(parts[1], "|")
for i := range cells {
cells[i] = strings.TrimSpace(cells[i])
}
return strings.Join(cells, " ")
}
return match
})
// Headers: remove # prefix
result = reHeader.ReplaceAllString(result, "")
// Bold
result = reBold.ReplaceAllStringFunc(result, func(match string) string {
parts := reBold.FindStringSubmatch(match)
if parts[1] != "" {
return parts[1]
}
return parts[2]
})
// Strikethrough
result = reStrike.ReplaceAllString(result, "$1")
// Blockquote
result = reBlockquote.ReplaceAllString(result, "")
// Horizontal rule -> empty line
result = reHR.ReplaceAllString(result, "")
// Unordered list: replace markers with "• "
result = reUL.ReplaceAllString(result, "${1}• ")
// Inline code: strip backticks (do after code blocks)
result = reInlineCode.ReplaceAllString(result, "$1")
// Clean up excessive blank lines
result = regexp.MustCompile(`\n{3,}`).ReplaceAllString(result, "\n\n")
return strings.TrimSpace(result)
}
@@ -1,221 +0,0 @@
package messaging
import (
"context"
"fmt"
"io"
"log"
"mime"
"net/http"
"os"
"path/filepath"
"regexp"
"strings"
"time"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
)
// reMarkdownImage matches markdown image syntax: ![alt](url)
var reMarkdownImage = regexp.MustCompile(`!\[[^\]]*\]\(([^)]+)\)`)
// ExtractImageURLs extracts image URLs from markdown text.
func ExtractImageURLs(text string) []string {
matches := reMarkdownImage.FindAllStringSubmatch(text, -1)
var urls []string
for _, m := range matches {
url := strings.TrimSpace(m[1])
if strings.HasPrefix(url, "http://") || strings.HasPrefix(url, "https://") {
urls = append(urls, url)
}
}
return urls
}
// SendMediaFromURL sends a local file or downloads from a URL and sends it as a media message.
func SendMediaFromURL(ctx context.Context, client *ilink.Client, toUserID, mediaURL, contextToken string) error {
// Check if it's a local file
if _, err := os.Stat(mediaURL); err == nil {
return SendMediaFromPath(ctx, client, toUserID, mediaURL, contextToken)
}
// Must be a valid HTTP URL to download
if !strings.HasPrefix(mediaURL, "http://") && !strings.HasPrefix(mediaURL, "https://") {
return fmt.Errorf("unsupported media path (not a local file and not an HTTP URL): %s", mediaURL)
}
data, contentType, err := downloadFile(ctx, mediaURL)
if err != nil {
return fmt.Errorf("download %s: %w", mediaURL, err)
}
return sendMediaData(ctx, client, toUserID, filenameFromURL(mediaURL), mediaURL, data, contentType, contextToken)
}
// SendMediaFromPath reads a local file and sends it as a media message.
func SendMediaFromPath(ctx context.Context, client *ilink.Client, toUserID, path, contextToken string) error {
data, err := os.ReadFile(path)
if err != nil {
return fmt.Errorf("read %s: %w", path, err)
}
return sendMediaData(ctx, client, toUserID, filepath.Base(path), path, data, inferContentType(path), contextToken)
}
func sendMediaData(ctx context.Context, client *ilink.Client, toUserID, fileName, source string, data []byte, contentType, contextToken string) error {
if fileName == "" {
fileName = "file"
}
cdnMediaType, itemType := classifyMedia(contentType, source)
log.Printf("[media] uploading %s (%s, %d bytes) for %s", source, contentType, len(data), toUserID)
uploaded, err := UploadFileToCDN(ctx, client, data, toUserID, cdnMediaType)
if err != nil {
return fmt.Errorf("upload to CDN: %w", err)
}
media := &ilink.MediaInfo{
EncryptQueryParam: uploaded.DownloadParam,
AESKey: AESKeyToBase64(uploaded.AESKeyHex),
EncryptType: 1,
}
var item ilink.MessageItem
switch itemType {
case ilink.ItemTypeImage:
item = ilink.MessageItem{
Type: ilink.ItemTypeImage,
ImageItem: &ilink.ImageItem{
Media: media,
MidSize: uploaded.CipherSize,
},
}
case ilink.ItemTypeVideo:
item = ilink.MessageItem{
Type: ilink.ItemTypeVideo,
VideoItem: &ilink.VideoItem{
Media: media,
VideoSize: uploaded.CipherSize,
},
}
default:
item = ilink.MessageItem{
Type: ilink.ItemTypeFile,
FileItem: &ilink.FileItem{
Media: media,
FileName: fileName,
Len: fmt.Sprintf("%d", uploaded.FileSize),
},
}
}
req := &ilink.SendMessageRequest{
Msg: ilink.SendMsg{
FromUserID: client.BotID(),
ToUserID: toUserID,
ClientID: NewClientID(),
MessageType: ilink.MessageTypeBot,
MessageState: ilink.MessageStateFinish,
ItemList: []ilink.MessageItem{item},
ContextToken: contextToken,
},
BaseInfo: ilink.BaseInfo{},
}
resp, err := client.SendMessage(ctx, req)
if err != nil {
return fmt.Errorf("send media message: %w", err)
}
if resp.Ret != 0 {
return fmt.Errorf("send media failed: ret=%d errmsg=%s", resp.Ret, resp.ErrMsg)
}
log.Printf("[media] sent %s to %s from %s", contentType, toUserID, source)
return nil
}
func downloadFile(ctx context.Context, url string) ([]byte, string, error) {
ctx, cancel := context.WithTimeout(ctx, 60*time.Second)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, "", err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, "", err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, "", fmt.Errorf("HTTP %d", resp.StatusCode)
}
data, err := io.ReadAll(resp.Body)
if err != nil {
return nil, "", err
}
contentType := resp.Header.Get("Content-Type")
if contentType == "" {
contentType = inferContentType(url)
}
return data, contentType, nil
}
func classifyMedia(contentType, url string) (cdnMediaType int, itemType int) {
ct := strings.ToLower(contentType)
if strings.HasPrefix(ct, "image/") || isImageExt(url) {
return ilink.CDNMediaTypeImage, ilink.ItemTypeImage
}
if strings.HasPrefix(ct, "video/") || isVideoExt(url) {
return ilink.CDNMediaTypeVideo, ilink.ItemTypeVideo
}
return ilink.CDNMediaTypeFile, ilink.ItemTypeFile
}
func isImageExt(url string) bool {
ext := strings.ToLower(filepath.Ext(stripQuery(url)))
switch ext {
case ".png", ".jpg", ".jpeg", ".gif", ".webp", ".bmp":
return true
}
return false
}
func isVideoExt(url string) bool {
ext := strings.ToLower(filepath.Ext(stripQuery(url)))
switch ext {
case ".mp4", ".mov", ".webm", ".mkv", ".avi":
return true
}
return false
}
func inferContentType(url string) string {
ext := filepath.Ext(stripQuery(url))
if ct := mime.TypeByExtension(ext); ct != "" {
return ct
}
return "application/octet-stream"
}
func filenameFromURL(rawURL string) string {
u := stripQuery(rawURL)
name := filepath.Base(u)
if name == "" || name == "." || name == "/" {
return "file"
}
return name
}
func stripQuery(rawURL string) string {
if i := strings.IndexByte(rawURL, '?'); i >= 0 {
return rawURL[:i]
}
return rawURL
}
@@ -1,73 +0,0 @@
package messaging
import "testing"
func TestExtractImageURLs(t *testing.T) {
text := "check ![img](https://example.com/a.png) and ![](https://example.com/b.jpg)"
urls := ExtractImageURLs(text)
if len(urls) != 2 {
t.Fatalf("expected 2 urls, got %d", len(urls))
}
if urls[0] != "https://example.com/a.png" {
t.Errorf("urls[0] = %q", urls[0])
}
if urls[1] != "https://example.com/b.jpg" {
t.Errorf("urls[1] = %q", urls[1])
}
}
func TestExtractImageURLs_NoImages(t *testing.T) {
urls := ExtractImageURLs("just plain text")
if len(urls) != 0 {
t.Errorf("expected 0 urls, got %d", len(urls))
}
}
func TestExtractImageURLs_RelativeURL(t *testing.T) {
text := "![img](./local.png)"
urls := ExtractImageURLs(text)
if len(urls) != 0 {
t.Errorf("expected 0 urls for relative path, got %d", len(urls))
}
}
func TestFilenameFromURL(t *testing.T) {
tests := []struct {
url string
want string
}{
{"https://example.com/photo.png", "photo.png"},
{"https://example.com/path/to/report.pdf", "report.pdf"},
{"https://example.com/file", "file"},
}
for _, tt := range tests {
got := filenameFromURL(tt.url)
if got != tt.want {
t.Errorf("filenameFromURL(%q) = %q, want %q", tt.url, got, tt.want)
}
}
}
func TestFilenameFromURL_WithQuery(t *testing.T) {
got := filenameFromURL("https://example.com/photo.png?token=abc")
if got != "photo.png" {
t.Errorf("got %q, want %q", got, "photo.png")
}
}
func TestStripQuery(t *testing.T) {
tests := []struct {
input string
want string
}{
{"https://example.com/a?b=c", "https://example.com/a"},
{"https://example.com/a", "https://example.com/a"},
{"https://example.com/?x=1&y=2", "https://example.com/"},
}
for _, tt := range tests {
got := stripQuery(tt.input)
if got != tt.want {
t.Errorf("stripQuery(%q) = %q, want %q", tt.input, got, tt.want)
}
}
}
@@ -1,86 +0,0 @@
package messaging
import (
"context"
"fmt"
"log"
"github.com/Wxw-Gu/WechatExplorer/services/wechat-connector/ilink"
"github.com/google/uuid"
)
// NewClientID generates a new unique client ID for message correlation.
func NewClientID() string {
return uuid.New().String()
}
// SendTypingState sends a typing indicator to a user via the iLink sendtyping API.
// It first fetches a typing_ticket via getconfig, then sends the typing status.
func SendTypingState(ctx context.Context, client *ilink.Client, userID, contextToken string) error {
// Get typing ticket
configResp, err := client.GetConfig(ctx, userID, contextToken)
if err != nil {
return fmt.Errorf("get config for typing: %w", err)
}
if configResp.TypingTicket == "" {
return fmt.Errorf("no typing_ticket returned from getconfig")
}
// Send typing
if err := client.SendTyping(ctx, userID, configResp.TypingTicket, ilink.TypingStatusTyping); err != nil {
return fmt.Errorf("send typing: %w", err)
}
log.Printf("[sender] sent typing indicator to %s", userID)
return nil
}
// SendTextReply sends a text reply to a user through the iLink API.
// If clientID is empty, a new one is generated.
func SendTextReply(ctx context.Context, client *ilink.Client, toUserID, text, contextToken, clientID string) error {
if clientID == "" {
clientID = NewClientID()
}
// Convert markdown to plain text for WeChat display
plainText := MarkdownToPlainText(text)
req := &ilink.SendMessageRequest{
Msg: ilink.SendMsg{
FromUserID: client.BotID(),
ToUserID: toUserID,
ClientID: clientID,
MessageType: ilink.MessageTypeBot,
MessageState: ilink.MessageStateFinish,
ItemList: []ilink.MessageItem{
{
Type: ilink.ItemTypeText,
TextItem: &ilink.TextItem{
Text: plainText,
},
},
},
ContextToken: contextToken,
},
BaseInfo: ilink.BaseInfo{},
}
resp, err := client.SendMessage(ctx, req)
if err != nil {
return fmt.Errorf("send message: %w", err)
}
if resp.Ret != 0 {
return fmt.Errorf("send message failed: ret=%d errmsg=%s", resp.Ret, resp.ErrMsg)
}
log.Printf("[sender] sent reply to %s: %q", toUserID, truncate(text, 50))
return nil
}
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
return s[:n] + "..."
}
+28 -4
View File
@@ -1172,13 +1172,37 @@ const renderExportScript = (name: string): string => `
}
const renderPaymentContent = (data, kind) => {
const isTransfer = kind === 'transfer'
const pay = (isTransfer ? data.transfer : data.pay) || {}
const amount = pay.amountText
const memo = pay.payMemo
const statusText = isTransfer ? pay.transferStatusText : pay.redPacketStatusText
const meta = [amount, memo ? '备注:' + memo : '', statusText]
.filter(Boolean)
.join(' · ')
const title = isTransfer
? (data.title || '微信转账')
: (pay.sendTitle || pay.receiveTitle || data.title || '微信红包')
const blurb =
data.description ||
data.des ||
(isTransfer ? '转账消息' : (pay.sceneText || '恭喜发财,大吉大利'))
// 状态/金额/备注始终单独露出,避免被消息原文 des 盖住。
const description = meta
? (blurb && blurb !== meta ? blurb + ' · ' + meta : meta)
: blurb
const footerBits = isTransfer
? [pay.paySubtype === '3' ? '收款' : pay.paySubtype === '1' || pay.paySubtype === '4' ? '转账' : '',
pay.transferId ? '单号 ' + pay.transferId : '']
: [pay.hbType ? '类型 ' + pay.hbType : '', pay.sendId ? 'sendid ' + pay.sendId : '']
return '<div class="structured-content payment-content ' + (isTransfer ? 'transfer' : 'red-packet') +
'" data-rich-kind="' + (isTransfer ? 'transfer' : 'redPacket') + '">' +
'<div class="structured-kicker">' + (isTransfer ? '微信转账' : '微信红包') + '</div>' +
'<div class="structured-title">' + displayText(data.title || (isTransfer ? '微信转账' : '微信红包')) + '</div>' +
'<div class="structured-description">' +
displayText(data.description || data.des || (isTransfer ? '转账消息' : '恭喜发财,大吉大利')) +
'</div></div>'
'<div class="structured-title">' + displayText(title) + '</div>' +
'<div class="structured-description">' + displayText(description) + '</div>' +
(footerBits.filter(Boolean).length
? '<div class="structured-footer">' + displayText(footerBits.filter(Boolean).join(' · ')) + '</div>'
: '') +
'</div>'
}
const renderVoipContent = (data) => {
const title = Number(data.roomType) === 1 ? '视频通话' : '语音通话'
+32 -4
View File
@@ -25,6 +25,7 @@ import { mergeCachedSelfInfo, type CachedSelfInfo } from './services/bootstrap-c
import type { VoiceRecognitionUseCase } from './voice-pipeline/voice-recognition-use-case'
import { imageFileQuality } from '../shared/image-quality'
import { resolveMemberName } from '../shared/member-names'
import { FavoritesService } from './favorites-service'
import { filesystemSafeName } from '../shared/contact-name'
const jobs = new Set<string>()
@@ -853,6 +854,28 @@ async function runSingleExport(
percent: Math.max(1, Math.round(((targetOrder + 1) / targets.length) * 10))
})
}
if (request.includeFavorites) {
const wcdb = chat.getChatDb()?.getWcdb4Client()
if (wcdb) {
try {
const favoriteMessages = await new FavoritesService(wcdb).listExportMessages(500)
for (const [messageOrder, message] of favoriteMessages.entries()) {
if (!request.kinds.includes(kindOf(message))) continue
messageEntries.push({
message: {
...message,
exportConversationId: 'favorites',
exportConversationName: '收藏'
},
targetOrder: targets.length,
messageOrder
})
}
} catch (error) {
console.warn('[export] favorites merge skipped:', error)
}
}
}
const messages = messageEntries
.sort((left, right) => {
const byTime = Number(left.message.createTime || 0) - Number(right.message.createTime || 0)
@@ -1246,10 +1269,15 @@ async function runSingleExport(
const wavChannels =
audioBuffer.length >= 44 ? audioBuffer.readUInt16LE(22) : 1
const pcmBytes = Math.max(0, audioBuffer.length - 44)
message.voiceDuration = Math.max(
1,
Math.round(pcmBytes / (wavSampleRate * wavChannels * 2))
)
// 口径统一:消息解析阶段已从 <voicemsg voicelength> 拿到微信的原始秒数(带小数),
// 它是唯一权威来源,不要覆盖。只有拿不到时才退回用 PCM 字节数估算——
// 那份估算是整秒、且下限 1 秒(WAV 缺失头部时的兜底),语义不同。
if (message.voiceDuration == null) {
message.voiceDuration = Math.max(
1,
Math.round(pcmBytes / (wavSampleRate * wavChannels * 2))
)
}
} catch (error) {
keepMediaError(
request,
+71
View File
@@ -0,0 +1,71 @@
import type { Message, ParsedContent } from '../shared/types'
import { favoriteRowToContent } from '../shared/favorites'
import type { Wcdb4Client } from './wcdb4-client'
/** 收藏只读服务:favorite.db → 可读卡片 / 导出消息。 */
export class FavoritesService {
constructor(private readonly wcdb4Client: Wcdb4Client) {}
async listContents(limit = 200): Promise<ParsedContent[]> {
const rows = await this.wcdb4Client.listFavoriteItems(limit)
return rows.map((row) => favoriteRowToContent(row))
}
/** 收藏转成导出/搜索可用的 Message 列表(合成会话「收藏」)。 */
async listExportMessages(limit = 200): Promise<Message[]> {
const rows = await this.wcdb4Client.listFavoriteItems(limit)
return rows.map((row, index) => {
const contentData = favoriteRowToContent(row)
const createTime = Number(row.update_time) || 0
return {
id: `fav-${row.local_id ?? index}`,
from: 'favorite',
type: '收藏',
datetime: createTime ? new Date(createTime * 1000).toISOString() : '',
content: textOf(contentData),
isSender: false,
createTime,
name: '收藏',
contentData
} as Message
})
}
/** 只读关键词搜索收藏文本(title/desc/正文)。 */
async search(query: string, limit = 50): Promise<ParsedContent[]> {
const needle = String(query || '').trim().toLowerCase()
if (!needle) return []
const items = await this.listContents(500)
return items.filter((item) => textOf(item).toLowerCase().includes(needle)).slice(0, limit)
}
/** Local Query API 适配:只要文本与时间戳。 */
async searchHits(
query: string,
limit = 20
): Promise<Array<{ text: string; timestamp?: number }>> {
const rows = await this.wcdb4Client.listFavoriteItems(500)
const needle = String(query || '').trim().toLowerCase()
if (!needle) return []
return rows
.map((row) => {
const content = favoriteRowToContent(row)
return {
text: textOf(content),
timestamp: Number(row.update_time) || undefined
}
})
.filter((hit) => hit.text.toLowerCase().includes(needle))
.slice(0, limit)
}
}
function textOf(content: ParsedContent): string {
if (content.type === 'text') return content.content
if (content.type === 'system') return content.content
if (content.type === 'share') return [content.title, content.des].filter(Boolean).join(' · ')
if (content.type === 'location') return content.poiname || content.label || '[位置]'
if (content.type === 'miniProgram') return content.title || '[小程序]'
if (content.type === 'forwardBundle') return content.title || '[聊天记录]'
return '[收藏]'
}
+180 -76
View File
@@ -13,6 +13,8 @@ import {
GroupReportRenderSnapshotExportRequest,
ReportHeat,
ReportSectionMeta,
buildReportAvatarAliasIndex,
mergeReportAvatars,
selectHeroParticipantNames
} from '../shared/group-report'
import { resolveMd5, getGroupSnapshot } from './services/chat-service'
@@ -99,51 +101,155 @@ const fallbackAvatar = (name: string): RenderedAvatar => {
const hue = hashName(name) % 360
const initial = escapeHtml(Array.from(name.trim())[0] || '?')
const svg = `<svg xmlns="http://www.w3.org/2000/svg" width="96" height="96"><rect width="96" height="96" rx="18" fill="hsl(${hue} 45% 82%)"/><text x="48" y="58" text-anchor="middle" font-family="-apple-system,BlinkMacSystemFont,PingFang SC,sans-serif" font-size="38" fill="hsl(${hue} 35% 28%)">${initial}</text></svg>`
return { source: `data:image/svg+xml;base64,${Buffer.from(svg).toString('base64')}`, fallback: true }
}
const imageMimeType = (contentType: string | null, source: string): string => {
if (contentType?.startsWith('image/')) return contentType.split(';')[0]
const extension = path.extname(source).toLowerCase()
if (extension === '.png') return 'image/png'
if (extension === '.webp') return 'image/webp'
if (extension === '.gif') return 'image/gif'
return 'image/jpeg'
}
const embedAvatar = async (source: string | undefined, name: string): Promise<RenderedAvatar> => {
if (!source) return fallbackAvatar(name)
if (/^data:image\/[a-z0-9.+/-]+;base64,[a-z0-9+/=]+$/i.test(source)) return { source, fallback: false }
try {
if (/^https?:\/\//i.test(source)) {
const response = await fetch(source, {
headers: {
'User-Agent': 'Mozilla/5.0 TraceMemo',
Referer: 'https://weixin.qq.com/'
},
signal: AbortSignal.timeout(8000)
})
if (!response.ok) throw new Error(`HTTP ${response.status}`)
const mime = imageMimeType(response.headers.get('content-type'), source)
return { source: `data:${mime};base64,${Buffer.from(await response.arrayBuffer()).toString('base64')}`, fallback: false }
}
const localPath = source.startsWith('file://') ? new URL(source) : source
const buffer = await fs.readFile(localPath)
return { source: `data:${imageMimeType(null, source)};base64,${buffer.toString('base64')}`, fallback: false }
} catch (error) {
console.warn(`[GroupReport] avatar fallback for ${name}:`, error)
return fallbackAvatar(name)
return {
source: `data:image/svg+xml;base64,${Buffer.from(svg).toString('base64')}`,
fallback: true
}
}
/**
* 只认真实图片字节。
*
* `content-type` 与 URL 扩展名都**不可信**:微信 CDN 在限流 / 反盗链时会返回 200 + HTML 正文。
* 旧实现按扩展名猜 mime 并默认 `image/jpeg`,会把 HTML 内联成"解码失败的 data URL" ——
* 在报告里表现为**空白头像**(比首字占位更糟:用户看不到任何东西,也不知道为什么)。
*/
const IMAGE_MAGIC: Array<{ mime: string; bytes: number[] }> = [
{ mime: 'image/jpeg', bytes: [0xff, 0xd8, 0xff] },
{ mime: 'image/png', bytes: [0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a] },
{ mime: 'image/gif', bytes: [0x47, 0x49, 0x46, 0x38] },
{ mime: 'image/bmp', bytes: [0x42, 0x4d] }
]
export const detectImageMime = (bytes: Buffer): string | undefined => {
for (const signature of IMAGE_MAGIC) {
if (
bytes.length >= signature.bytes.length &&
signature.bytes.every((byte, index) => bytes[index] === byte)
) {
return signature.mime
}
}
if (
bytes.length >= 12 &&
bytes.toString('ascii', 0, 4) === 'RIFF' &&
bytes.toString('ascii', 8, 12) === 'WEBP'
) {
return 'image/webp'
}
const head = bytes.subarray(0, 64).toString('utf8').trimStart()
if (head.startsWith('<svg')) return 'image/svg+xml'
if (head.startsWith('<?xml') && head.includes('<svg')) return 'image/svg+xml'
return undefined
}
const AVATAR_FETCH_TIMEOUT_MS = 8000
/** 首次 + 一次重试:单次瞬时失败(限流 / 连接重置 / 超时)不该让一个人永久退回首字。 */
const AVATAR_FETCH_ATTEMPTS = 2
/** 同一 origin 的并发上限。几十个头像同时打一个 CDN 会显著抬高被限流的概率。 */
const AVATAR_FETCH_CONCURRENCY = 6
/** 进程内头像缓存条目上限(老报告重渲染 / 连续生成同一群时不必重复下载)。 */
const AVATAR_CACHE_LIMIT = 256
const sleep = (ms: number): Promise<void> => new Promise((resolve) => setTimeout(resolve, ms))
const avatarEmbedCache = new Map<string, RenderedAvatar>()
const rememberAvatarEmbed = (source: string, rendered: RenderedAvatar): void => {
if (rendered.fallback) return
avatarEmbedCache.set(source, rendered)
while (avatarEmbedCache.size > AVATAR_CACHE_LIMIT) {
const oldest = avatarEmbedCache.keys().next().value
if (oldest === undefined) break
avatarEmbedCache.delete(oldest)
}
}
/** 有界并发:把 N 个任务压到 limit 个同时在飞,结果顺序与输入一致。 */
export const mapWithConcurrency = async <T, R>(
items: readonly T[],
limit: number,
task: (item: T) => Promise<R>
): Promise<R[]> => {
const results: R[] = new Array(items.length)
let cursor = 0
const workers = Array.from({ length: Math.max(1, Math.min(limit, items.length)) }, async () => {
for (;;) {
const index = cursor
cursor += 1
if (index >= items.length) return
results[index] = await task(items[index])
}
})
await Promise.all(workers)
return results
}
const readAvatarSource = async (source: string): Promise<RenderedAvatar> => {
if (/^https?:\/\//i.test(source)) {
const response = await fetch(source, {
headers: {
'User-Agent': 'Mozilla/5.0 TraceMemo',
Referer: 'https://weixin.qq.com/'
},
signal: AbortSignal.timeout(AVATAR_FETCH_TIMEOUT_MS)
})
if (!response.ok) throw new Error(`HTTP ${response.status}`)
const bytes = Buffer.from(await response.arrayBuffer())
const mime = detectImageMime(bytes)
if (!mime) {
throw new Error(
`not an image (content-type=${response.headers.get('content-type') || 'unknown'}, ${bytes.length} bytes)`
)
}
return { source: `data:${mime};base64,${bytes.toString('base64')}`, fallback: false }
}
const localPath = source.startsWith('file://') ? new URL(source) : source
const bytes = await fs.readFile(localPath)
const mime = detectImageMime(bytes)
if (!mime) throw new Error(`not an image (${bytes.length} bytes)`)
return { source: `data:${mime};base64,${bytes.toString('base64')}`, fallback: false }
}
export const embedAvatar = async (
source: string | undefined,
name: string
): Promise<RenderedAvatar> => {
if (!source) return fallbackAvatar(name)
if (/^data:image\/[a-z0-9.+/-]+;base64,[a-z0-9+/=]+$/i.test(source))
return { source, fallback: false }
const cached = avatarEmbedCache.get(source)
if (cached) return cached
let lastError: unknown
for (let attempt = 1; attempt <= AVATAR_FETCH_ATTEMPTS; attempt += 1) {
try {
const embedded = await readAvatarSource(source)
rememberAvatarEmbed(source, embedded)
return embedded
} catch (error) {
lastError = error
if (attempt < AVATAR_FETCH_ATTEMPTS) await sleep(150 * attempt)
}
}
console.warn(`[GroupReport] avatar fallback for ${name}:`, lastError)
return fallbackAvatar(name)
}
/**
* 从群成员快照反推真头像,填进 metadata.avatars。
* - 没传 talker → 跳过(向后兼容)
* - talker 解析失败 / snapshot 拿不到 → 200 + warn,继续走 fallback
* - 客户端传的 avatars[name](非空)优先;否则从 snapshot 的 m_nsHeadImgUrl 补
* - 同名取首条(P2 风险:群里两人同名)
*
* **必须按多个别名建索引**:报告里的显示名取决于 `memberNameMode`
* (默认 `groupNickname` = 群昵称),而快照的 `nickname` 字段是
* `wechatNickname || groupNickname || username`(见 `normalizeGroupMembers`)。
* 一个成员同时有微信昵称与群昵称且两者不同时,只按 `nickname` 建索引就会**全部对不上**,
* 于是头像 enrichment 静默失效 —— 这正是原先只用单一索引键时的问题。
*/
const enrichAvatarsFromGroup = async (metadata: GroupReportMetadata): Promise<void> => {
if (!metadata.talker) return
@@ -162,22 +268,16 @@ const enrichAvatarsFromGroup = async (metadata: GroupReportMetadata): Promise<vo
return
}
const index = new Map<string, string>()
for (const member of snapshot.members) {
if (member.nickname && member.avatar && !index.has(member.nickname)) {
index.set(member.nickname, member.avatar)
}
}
// 每个成员的所有可用显示名都指向同一个头像 URL;先到先得,避免同名互相覆盖。
// 只按 `member.nickname` 建索引会在"微信昵称 ≠ 群昵称"时全部对不上 —— 见 shared 里的注释。
const index = buildReportAvatarAliasIndex(snapshot.members)
metadata.avatars = metadata.avatars ?? {}
for (const [name, url] of index) {
if (metadata.avatars[name]) continue
metadata.avatars[name] = url
}
const filled = mergeReportAvatars(metadata.avatars, index)
metadata.warnings = metadata.warnings ?? []
metadata.warnings.push(
`enriched ${index.size} member avatars from snapshot (${snapshot.memberCount} members)`
`enriched ${filled}/${index.size} member avatar aliases from snapshot (${snapshot.memberCount} members)`
)
}
@@ -276,12 +376,22 @@ const renderReportHtml = async (request: GroupReportExportRequest): Promise<stri
report.media?.voiceHighlights?.forEach((item) => avatarNames.add(item.sender))
report.media?.funBadges?.forEach((item) => avatarNames.add(item.owner))
const avatars = new Map<string, RenderedAvatar>()
await Promise.all(
Array.from(avatarNames).map(async (name) => {
avatars.set(name, await embedAvatar(metadata.avatars[name], name))
})
// 有界并发 + 单条重试:几十个头像同时打同一个 CDN 会被限流,瞬时失败会让一个人
// 在整份报告里永久退化成首字占位(实测同一天三次生成:0% / 0% / 32% 失败)。
const renderedAvatars = await mapWithConcurrency(
Array.from(avatarNames),
AVATAR_FETCH_CONCURRENCY,
async (name) => [name, await embedAvatar(metadata.avatars[name], name)] as const
)
const avatars = new Map(renderedAvatars)
const fallbackCount = renderedAvatars.filter(([, item]) => item.fallback).length
if (fallbackCount > 0) {
// 只记数量,不记人名:报告本身已经有名字,这里只需要一个可诊断的信号。
metadata.warnings = metadata.warnings ?? []
metadata.warnings.push(
`avatar fallback ${fallbackCount}/${renderedAvatars.length}: 未取到真实头像,已用首字占位`
)
}
const avatar = (name: string): RenderedAvatar => avatars.get(name) || fallbackAvatar(name)
const renderAvatar = (
name: string,
@@ -295,9 +405,7 @@ const renderReportHtml = async (request: GroupReportExportRequest): Promise<stri
}
const heroNames = selectHeroParticipantNames(metadata.heroParticipants)
const heroAvatars = heroNames
.map((name) => renderAvatar(name, 'hero', '', name))
.join('')
const heroAvatars = heroNames.map((name) => renderAvatar(name, 'hero', '', name)).join('')
const heroAvatarClass = heroNames.length ? `avatar-count-${heroNames.length}` : 'empty-section'
const topicCards = report.topics
@@ -480,7 +588,7 @@ const renderReportHtml = async (request: GroupReportExportRequest): Promise<stri
(item, index) => `<div class="rank tm-fragment tm-ranking-item">
${renderAvatar(item.sender, 'ranking')}
<b>${index + 1}. ${escapeHtml(item.sender)}</b>
<span>${item.count} 条 · ${item.durationSec} 秒</span>
<span>${item.count} 条 · ${Math.round(item.durationSec)} 秒</span>
</div>`
)
.join('')
@@ -656,7 +764,10 @@ const renderReportHtml = async (request: GroupReportExportRequest): Promise<stri
for (const [key, value] of Object.entries(values)) html = replacePlaceholder(html, key, value)
// 清空模板中残留的未使用占位符(模板独有但 values 没提供的键)
html = html.replace(/\{\{[A-Z_]+\}\}/g, '')
return addReportCsp(injectReportTemplateFragmentContract(html), resolvedTemplate.source === 'builtin')
return addReportCsp(
injectReportTemplateFragmentContract(html),
resolvedTemplate.source === 'builtin'
)
}
const renderReportSnapshotHtml = async (
@@ -679,7 +790,10 @@ const renderReportSnapshotHtml = async (
REPORT_DATE: request.snapshot.values.REPORT_DATE || escapeHtml(request.snapshot.reportDate)
}
for (const [key, value] of Object.entries(values)) html = replacePlaceholder(html, key, value)
return addReportCsp(html.replace(/\{\{[A-Z0-9_]+\}\}/g, ''), resolvedTemplate.source === 'builtin')
return addReportCsp(
html.replace(/\{\{[A-Z0-9_]+\}\}/g, ''),
resolvedTemplate.source === 'builtin'
)
}
export const extractGroupReportRenderSnapshot = async (
@@ -1021,15 +1135,10 @@ export const exportGroupReport = async (
await fs.writeFile(htmlPath, html, 'utf8')
const htmlEndedAt = new Date()
const pngStartedAt = new Date()
const imageDataUrl = await captureFullPage(
htmlPath,
pngPath,
request.templateId,
{
...resolvedTemplate.definition,
maxCaptureHeight: resolvedTemplate.captureMaxHeight
}
)
const imageDataUrl = await captureFullPage(htmlPath, pngPath, request.templateId, {
...resolvedTemplate.definition,
maxCaptureHeight: resolvedTemplate.captureMaxHeight
})
const pngEndedAt = new Date()
return {
success: true,
@@ -1074,15 +1183,10 @@ export const exportGroupReportSnapshot = async (
await fs.writeFile(htmlPath, html, 'utf8')
const htmlEndedAt = new Date()
const pngStartedAt = new Date()
const imageDataUrl = await captureFullPage(
htmlPath,
pngPath,
request.templateId,
{
...resolvedTemplate.definition,
maxCaptureHeight: resolvedTemplate.captureMaxHeight
}
)
const imageDataUrl = await captureFullPage(htmlPath, pngPath, request.templateId, {
...resolvedTemplate.definition,
maxCaptureHeight: resolvedTemplate.captureMaxHeight
})
const pngEndedAt = new Date()
return {
success: true,
+5 -1
View File
@@ -114,7 +114,11 @@ function getFfmpegCandidates(selectedPath = loadSettings().ffmpegPath): FfmpegCa
)
}
function resolveFfmpegExecutable(): string {
/**
* 解析可用的 ffmpeg 可执行文件。除图片解密自身使用外,也供 System OCR 的
* 图片归一化(GIF/BMP/WebP/TIFF → PNG)复用,避免重复一套路径探测逻辑。
*/
export function resolveFfmpegExecutable(): string {
for (const candidate of getFfmpegCandidates()) {
const pathLike = candidate.executable.includes('/') || candidate.executable.includes('\\')
if (pathLike) {
+278 -17
View File
@@ -29,12 +29,10 @@ import {
ImageDecryptService,
inspectImageDecoderExecutable,
inspectImageDecoderStatus,
resolveFfmpegExecutable,
type DecodedImage
} from './image-decrypt-service'
import {
exportGroupReportSnapshot,
extractGroupReportRenderSnapshot
} from './group-report-service'
import { exportGroupReportSnapshot, extractGroupReportRenderSnapshot } from './group-report-service'
import {
deleteGeneratedReport,
listGeneratedReports,
@@ -44,9 +42,7 @@ import {
} from './report-history-service'
import { reportTemplateService } from './report-template-service'
import { registerReportTemplateIpc } from './report-template-ipc'
import type {
GroupReportRenderSnapshotExportRequest
} from '../shared/group-report'
import type { GroupReportRenderSnapshotExportRequest } from '../shared/group-report'
import type {
SaveGeneratedReportRequest,
PrepareGeneratedReportTemplateSwitchRequest,
@@ -64,6 +60,7 @@ import { apiTokenStore } from './api-token-store'
import { ImageKeyConfigService } from './services/image-key-config-service'
import { AIProviderService } from './services/ai-provider-service'
import { imageInsightService } from './services/image-insight-service'
import { systemOcrService } from './services/system-ocr-service'
import type {
ImageAnalysisRequest,
ImageAnalysisResponse,
@@ -71,6 +68,7 @@ import type {
ImageCandidateQuery,
ImageInsight
} from '../shared/image-insight'
import type { SystemOcrCapability, SystemOcrRequest, SystemOcrResult } from '../shared/system-ocr'
import { KeyServiceMac } from './key-service-mac'
import { KeyService as KeyServiceWin } from './key-service-win'
import * as chat from './services/chat-service'
@@ -111,6 +109,8 @@ import {
} from './services/bootstrap-cache'
import { installSafeConsole } from './safe-log'
import { agentHubService } from './services/agent-hub-service'
import { WechatConnectorService } from './services/wechat-ilink'
import { wechatSendGateway } from './services/wechat-send-gateway'
import { groupExitMonitorService } from './services/group-exit-monitor-service'
import { wechatActionLogService } from './services/wechat-action-log-service'
import { wechatActionGateway } from './services/wechat-action-gateway'
@@ -147,6 +147,11 @@ function nextGetMessagesRequestId(): string {
import type { AppLogEntry } from '../shared/app-log'
import { appUpdateService } from './services/app-update-service'
import { clearCache, getCacheSummary, openKnowledgeDirectory } from './services/cache-service'
import { imageTextIndexService } from './services/image-text-index-service'
import {
IMAGE_TEXT_SEGMENT_MESSAGE_LIMIT,
type ImageTextIndexStartOptions
} from '../shared/image-text-index'
import type { CacheClearScope } from './services/cache-service'
import { configureRecallArchive, RecallArchiveMonitor } from './services/recall-archive-service'
import { VideoAssetService } from './video-asset-service'
@@ -196,6 +201,12 @@ import type {
// handler. Wrap console.* before any other module logs anything.
installSafeConsole()
/**
* 进程内微信 iLink 连接器:inbound 回调与 outbound 发送都在主进程内完成,
* 不依赖子进程,也不开本地 HTTP 端口。
*/
const wechatConnectorService = new WechatConnectorService()
let voiceService: VoiceService | null = null
let voiceRecognition: VoiceRecognitionUseCase | null = null
let voiceBatchService: VoiceBatchService | null = null
@@ -634,8 +645,139 @@ app.whenReady().then(async () => {
voiceRecognition.onTranscriptUpdate((update) =>
knowledgeSearchService?.indexVoiceTranscript(update)
)
/**
* 确保图片解密服务可用(按需创建,与 `db:getImage` 冷路径同一套构造方式)。
*
* 提取成显式入口是因为原来它只存在于 `db:getImage` 的闭包里,
* 别的需要解密的路径(图片文字索引回填)拿不到、只能拿到 `null`。
*/
function ensureImageDecryptService(): ImageDecryptService | null {
if (imageDecryptService) return imageDecryptService
const { xorKey, aesKey } = getConfiguredImageKeys()
if (!aesKey) return null
imageDecryptService = new ImageDecryptService(
xorKey,
aesKey,
chat.getChatDb()?.getWcdb4Client(),
loadSettings().dbRoot
)
return imageDecryptService
}
// 图片文字索引(本地 System OCR 派生文本)。
// 与语音转写完全同构:派生文本在 main 进程解析后贴到消息上,Knowledge 侧只消费结果。
knowledgeSearchService.setImageOcrResolver((conversationId, messageId) =>
imageTextIndexService.getConversationOcr(conversationId).get(messageId)
)
/**
* 图片文字索引需要解密图片。
*
* 解密服务原本只在 `db:getImage`(用户点开某张图)里才懒加载,于是没点开过图片时
* 全量回填会拿到 `null`。这里改成显式"按需确保",凡是需要解密的路径都能自己建起来。
*/
imageTextIndexService.bind({
databaseRoot: join(app.getPath('userData'), 'image-text-index'),
resolveAccountId: () =>
chat.isReady()
? String(chat.getSelfAccountInfo()?.wxid || chat.getCurrentAccountRoot() || '')
: '',
resolveAccountRoot: () => chat.getCurrentAccountRoot() || loadSettings().dbRoot || '',
listContacts: async () => {
const contacts = await chat.listContactsAsync()
return contacts.map((contact) => ({
md5: contact.md5,
m_nsUsrName: contact.m_nsUsrName,
type: contact.type
}))
},
listMessages: (conversationId) =>
chat.listMessagesAsync(
conversationId,
undefined,
undefined,
undefined,
undefined,
'image-text-index'
),
/**
* 图片索引走**专用查询**:只读图片消息,不读整个会话。
*
* 全量读取一个 20 万条消息的会话实测要 15s 以上,而其中 99% 以上的行
* 图片索引根本不看 —— 那是数据边界错了,不是 OCR 慢。
*/
listImageMessages: (conversationId, window) =>
chat.listImageMessagesAsync(
conversationId,
{
...(window?.sinceMs !== undefined ? { sinceMs: window.sinceMs } : {}),
...(window?.beforeMs !== undefined ? { beforeMs: window.beforeMs } : {}),
limit: IMAGE_TEXT_SEGMENT_MESSAGE_LIMIT
},
'image-text-index'
),
countConversationImages: (conversationId, range) =>
chat.countImageMessagesAsync(conversationId, range),
imageWatermark: (conversationId, range) =>
chat.imageConversationWatermarkAsync(conversationId, range),
decryptService: () => ensureImageDecryptService(),
capability: () => systemOcrService.getCapability(),
/**
* OCR 并发度的运行时覆盖;不设置则走 `DEFAULT_IMAGE_TEXT_OCR_CONCURRENCY`。
*
* 同一份二进制、同一批图片只改这一个数,才能把并发度当作对照变量来比较。
* 非法值会被 `resolveImageTextOcrConcurrency` 收敛掉。
*/
...(process.env.TRACEMEMO_OCR_CONCURRENCY
? { ocrConcurrency: Number(process.env.TRACEMEMO_OCR_CONCURRENCY) }
: {}),
/** 低频性能画像:只写性能数字,不含图片内容 / 路径 / 会话标识。 */
logStageProfile: (profile) =>
appLogger.write({
level: 'info',
scope: 'image-text-index',
message: '图片文字索引性能画像',
details: { ...profile }
}),
recognize: async (imageDataUrl) => {
const result = await systemOcrService.recognize({ imageDataUrl })
return {
success: result.success,
text: result.text,
language: result.language,
...(result.errorCode ? { errorCode: result.errorCode } : {})
}
},
// 会话图片全部处理完 → 重建该会话索引,OCR 文本才可被 search_messages 检索。
onConversationIndexed: (conversationId) =>
knowledgeSearchService?.indexImageOcr(conversationId) ?? Promise.resolve()
})
aiSearchPipelineService = new AiSearchPipelineService(knowledgeSearchService, aiProviderService)
localQueryApiService = new LocalQueryApiService(knowledgeSearchService)
// 图片文字索引覆盖度是**独立覆盖维度**:接到 search_messages 的 tool result 上,
// 让 Query Agent 在图片索引没做完时不能凭 0 条证据断言"没有"。
localQueryApiService.setImageTextCoverageProvider(() =>
imageTextIndexService.getCoverageSnapshot()
)
/**
* 单条图片消息的 OCR 派生文本也要接到精确读消息路径上。
*
* 与覆盖度是**两件不同的事**:覆盖度回答"索引建了多少",这里回答
* "这一条图片已经识别出的文字是什么"。只接前者的话,图片索引建好了模型也读不到正文,
* 只能看到一个空的 `attachment`。
*
* 只读派生库,**不触发 OCR / 解密 / 读原图**。
*/
localQueryApiService.setImageOcrEntryProvider((conversationId, messageId) =>
imageTextIndexService.getConversationOcr(conversationId).get(messageId)
)
// 收藏只读检索:search_messages 合并 favorite.db 命中(无库时返回空)。
localQueryApiService.setFavoritesSearchProvider(async (query, limit) => {
const wcdb = chat.getChatDb()?.getWcdb4Client()
if (!wcdb) return []
const { FavoritesService } = await import('./favorites-service')
return new FavoritesService(wcdb).searchHits(query, limit)
})
setLocalQueryApiService(localQueryApiService)
// Query Agent:生产 Runtime 只在这里实例化一次,桌面问问微信与 Agent Hub 共用同一个实例。
queryAgentService = new QueryAgentService(
@@ -649,14 +791,47 @@ app.whenReady().then(async () => {
if (!aiSearchPipelineService) throw new Error('本地搜索服务尚未初始化')
return aiSearchPipelineService.run(request, () => undefined)
},
log: (record) => appLogger.write({ level: record.level, scope: 'query-agent', message: record.message, details: record.details })
log: (record) =>
appLogger.write({
level: record.level,
scope: 'query-agent',
message: record.message,
details: record.details
})
})
agentHubService.setQueryAgentService(queryAgentService)
// 微信 iLink 连接器直接跑在主进程内:inbound 回调与 outbound 发送都不经过本地 HTTP 桥。
agentHubService.setWechatConnector(wechatConnectorService)
wechatSendGateway.configureIlinkSender(async (request) => {
const target = {
to: request.to,
...(request.account_id ? { accountId: request.account_id } : {}),
...(request.context_token ? { contextToken: request.context_token } : {})
}
if (request.type === 'text') {
await wechatConnectorService.sendText({ ...target, text: request.msg })
return
}
if (request.type === 'image' || request.type === 'file') {
if (/^https?:\/\//i.test(request.msg)) {
await wechatConnectorService.sendMediaUrl({ ...target, mediaUrl: request.msg })
} else {
await wechatConnectorService.sendMediaPath({ ...target, filePath: request.msg })
}
return
}
throw new Error('iLink 通道暂不支持发送语音')
})
knowledgeSearchService.onStatusChange((status) => {
for (const window of BrowserWindow.getAllWindows()) {
if (!window.isDestroyed()) window.webContents.send('knowledge:status', status)
}
})
imageTextIndexService.onStatusChange((status) => {
for (const window of BrowserWindow.getAllWindows()) {
if (!window.isDestroyed()) window.webContents.send('image-text-index:status', status)
}
})
voiceRecognition.modelManager.setProgressListener((status) => {
for (const window of BrowserWindow.getAllWindows()) {
if (!window.isDestroyed()) window.webContents.send('voice:modelProgress', status)
@@ -745,12 +920,25 @@ app.whenReady().then(async () => {
ipcMain.handle('cache:getSummary', () => getCacheSummary())
ipcMain.handle('cache:openKnowledgeDirectory', () => openKnowledgeDirectory())
ipcMain.handle('cache:clear', async (_, scope: CacheClearScope) => {
const allowedScopes: CacheClearScope[] = ['bootstrap', 'electron', 'knowledge', 'all']
const allowedScopes: CacheClearScope[] = [
'bootstrap',
'electron',
'knowledge',
'image-text-index',
'all'
]
if (!allowedScopes.includes(scope)) return getCacheSummary()
imageDecryptService = null
// 这里刻意**不再**提前 resetAccount():清理钩子需要先读到派生库里的
// "哪些会话有 OCR 派生文本",才能把这些会话的 Knowledge 索引一起失效。
// 句柄由 beforeClearImageTextIndex 内部的 clear() 自己关闭(删文件前)。
return clearCache(scope, {
beforeClearKnowledge: () =>
knowledgeSearchService?.prepareForCacheClear() || Promise.resolve()
knowledgeSearchService?.prepareForCacheClear() || Promise.resolve(),
beforeClearImageTextIndex: async () => {
await imageTextIndexService.prepareForCacheClear()
imageTextIndexService.resetAccount()
}
})
})
@@ -838,6 +1026,8 @@ app.whenReady().then(async () => {
.catch((error) => console.warn('[WCDB4] message cursor warmup failed:', error))
}
imageDecryptService = null
// 派生库按 accountId 分目录,切账号必须换句柄,否则会串账号。
imageTextIndexService.resetAccount()
console.log(
`[WCDB4] db:init ready sessions=${sessions.length} monitoring=${monitoring} cost=${Date.now() - startedAt}ms`
)
@@ -1017,6 +1207,8 @@ app.whenReady().then(async () => {
aesKey: result.aesKey
})
if (saved.success) imageDecryptService = null
// 派生库按 accountId 分目录,切账号必须换句柄,否则会串账号。
imageTextIndexService.resetAccount()
return {
...result,
success: saved.success,
@@ -1037,6 +1229,8 @@ app.whenReady().then(async () => {
ipcMain.handle('image:saveConfig', (_, request: SaveImageKeyRequest) => {
const result = imageKeyConfigService.save(request)
if (result.success) imageDecryptService = null
// 派生库按 accountId 分目录,切账号必须换句柄,否则会串账号。
imageTextIndexService.resetAccount()
return result
})
@@ -1047,6 +1241,8 @@ app.whenReady().then(async () => {
ipcMain.handle('image:clearConfig', () => {
const result = imageKeyConfigService.clear()
if (result.success) imageDecryptService = null
// 派生库按 accountId 分目录,切账号必须换句柄,否则会串账号。
imageTextIndexService.resetAccount()
return result
})
@@ -1171,7 +1367,7 @@ app.whenReady().then(async () => {
const requestId = nextGetMessagesRequestId()
const startedAt = Date.now()
wcdbDebugLog(
`[${requestId}] IPC db:getMessages start userMd5=${userMd5} start=${startTime || 0} end=${endTime || 0} limit=${options?.limit || 0}`
`[${requestId}] IPC db:getMessages start start=${startTime || 0} end=${endTime || 0} limit=${options?.limit || 0}`
)
try {
const messages = await chat.listMessagesAsync(
@@ -1232,7 +1428,7 @@ app.whenReady().then(async () => {
return { messages: [], found: false, radiusSeconds: 0, truncated: false }
}
wcdbDebugLog(
`[${requestId}] IPC db:getMessagesAround start userMd5=${userMd5} messageId=${target.messageId} anchor=${anchorSeconds || 0}`
`[${requestId}] IPC db:getMessagesAround start messageId=${target.messageId} anchor=${anchorSeconds || 0}`
)
for (const radius of radii) {
const start = Math.max(0, (anchorSeconds as number) - radius)
@@ -1335,6 +1531,30 @@ app.whenReady().then(async () => {
if (!knowledgeSearchService) throw new Error('本地知识库服务尚未初始化')
return knowledgeSearchService.cancelCurrentAccountIndex()
})
// ---- 图片文字索引(本地 System OCR 派生文本,非 AI Provider)----
ipcMain.handle('image-text-index:getStatus', () => imageTextIndexService.getStatus())
/**
* 点击索引前的快速统计:纯 SQL COUNT,**不解密任何图片**。
* 这是「先告诉用户有多少张图片再决定是否开始」能足够快的前提。
*/
ipcMain.handle('image-text-index:count', (_, sinceMs?: number) =>
imageTextIndexService.countImageMessages(sinceMs)
)
ipcMain.handle('image-text-index:start', (_, options?: ImageTextIndexStartOptions) =>
imageTextIndexService.startPass(options ?? {})
)
ipcMain.handle('image-text-index:pause', () => imageTextIndexService.pause())
ipcMain.handle('image-text-index:resume', (_, options?: ImageTextIndexStartOptions) =>
imageTextIndexService.resume(options ?? {})
)
ipcMain.handle('image-text-index:cancel', () => imageTextIndexService.cancel())
ipcMain.handle('image-text-index:clear', () => imageTextIndexService.clear())
// 只重置失败记录(成功记录与其它数据一律不动),供"修好代码后重跑"使用。
ipcMain.handle('image-text-index:resetFailures', () =>
imageTextIndexService.resetRetriableFailures()
)
// 派生索引修复:只重建 Knowledge 里的图片派生条目(L3),**不重新 OCR**(L1 不动)。
ipcMain.handle('image-text-index:repair', () => imageTextIndexService.repairKnowledgeIndex())
ipcMain.handle('ai-search:run', (event, request: AiSearchPipelineRequest) => {
if (!aiSearchPipelineService) throw new Error('本地搜索服务尚未初始化')
return aiSearchPipelineService.run(request, (progress) => {
@@ -1640,12 +1860,15 @@ app.whenReady().then(async () => {
return error ? { success: false, error } : { success: true }
})
ipcMain.handle('voice:recognize', (_, reference: VoiceMessageReference) => {
if (!voiceRecognition) {
return { success: false, code: 'NOT_CONNECTED', error: '语音识别服务尚未初始化' }
ipcMain.handle(
'voice:recognize',
(_, reference: VoiceMessageReference, options?: { force?: boolean }) => {
if (!voiceRecognition) {
return { success: false, code: 'NOT_CONNECTED', error: '语音识别服务尚未初始化' }
}
return voiceRecognition.recognize(reference, options)
}
return voiceRecognition.recognize(reference)
})
)
ipcMain.handle('voice:getTranscriptSnapshot', (_, reference: VoiceMessageReference) => {
return voiceRecognition?.getTranscriptSnapshot(reference) || { state: 'pending' as const }
@@ -1814,6 +2037,13 @@ app.whenReady().then(async () => {
}
})
// System OCR 是独立的本地 Runtime(不是 AI Provider):只注入项目统一的 ffmpeg
// 解析逻辑(GIF/BMP/WebP/TIFF → PNG 归一化)和系统 locale(OCR 语言包探测)。
systemOcrService.bind({
resolveFfmpegExecutable,
locale: () => app.getLocale()
})
/** 日报入口:取会话 Top N 热点图片 + 已缓存的 Insight */
ipcMain.handle(
'image:listCandidates',
@@ -1885,6 +2115,26 @@ app.whenReady().then(async () => {
}
)
// ============================================================
// 本地图片文字识别(System OCR Runtime:Windows 系统 OCR / macOS 系统 OCR)
// ============================================================
// 这是本地 Runtime,不是 AI Vision Provider:
// - 不联网、不上传原图;
// - 不读写 AI Provider / Vision 模型配置;
// - 本 IPC 只返回识别文本、不落库:派生文本的持久化由图片文字索引负责。
ipcMain.handle('system-ocr:getCapability', async (): Promise<SystemOcrCapability> => {
return imageInsightService.getSystemOcrCapability()
})
ipcMain.handle(
'system-ocr:recognize',
async (_, request: SystemOcrRequest): Promise<SystemOcrResult> => {
// 单图识别是用户主动触发的一次操作,结果里已经带了 text / durationMs / errorCode,
// 调用方直接用返回值判断即可,这里不再打日志(尤其不打稳定的图片标识)。
return imageInsightService.extractLocalText(request)
}
)
ipcMain.handle('db:getSticker', async (_, cdnUrl?: string, md5?: string) => {
if (!stickerService) {
stickerService = new StickerService(chat.getChatDb()?.getWcdb4Client())
@@ -1933,6 +2183,8 @@ app.whenReady().then(async () => {
if (aesKey) imageKeyConfigService.save({ resourceRoot, xorKey, aesKey })
else imageKeyConfigService.clear()
imageDecryptService = null
// 派生库按 accountId 分目录,切账号必须换句柄,否则会串账号。
imageTextIndexService.resetAccount()
}
if ('recallProtectionEnabled' in patch && chat.isReady()) {
const currentDb = chat.getChatDb()
@@ -2091,6 +2343,15 @@ app.whenReady().then(async () => {
ipcMain.handle('agent-hub:cancelLogin', () => agentHubService.cancelLogin())
ipcMain.handle('agent-hub:reconnect', () => agentHubService.reconnect())
ipcMain.handle('agent-hub:disconnect', () => agentHubService.disconnect())
// 对话记录:完整收发回看,仅本机,不进日志。
ipcMain.handle('agent-hub:getConversations', () => agentHubService.listConversations())
ipcMain.handle('agent-hub:getConversation', (_, userId: string) =>
agentHubService.getConversation(String(userId || ''))
)
ipcMain.handle('agent-hub:clearConversations', () => {
agentHubService.clearConversations()
return { success: true }
})
ipcMain.handle('wechat-personal:getStatus', () => personalWechatSendService.getStatus())
ipcMain.handle('wechat-personal:getKeepProcess', () =>
personalWechatSendService.getKeepOneBotProcess()
+11 -5
View File
@@ -134,6 +134,15 @@ export function parseXkeyHelperOutput(output: string): DatabaseKeyResult {
return mapXkeyHelperFailure(rawError)
}
export const SIP_ENABLED_ERROR =
'macOS 系统完整性保护(SIP)已开启,无法自动获取数据库密钥。请先关闭 SIP,或改用手动粘贴。'
export function parseSipEnabled(statusOutput: string): boolean {
const status = statusOutput.match(/status:\s*([a-z]+)/i)?.[1]?.toLowerCase()
if (status) return status === 'enabled'
return statusOutput.toLowerCase().includes('enabled')
}
export class KeyServiceMac {
private getMacKeyRuntimeDir(): string {
return path.join(app.getPath('userData'), 'key-runtime')
@@ -306,7 +315,7 @@ export class KeyServiceMac {
private async isSipEnabled(): Promise<boolean> {
try {
const { stdout } = await execFileAsync('/usr/bin/csrutil', ['status'])
return stdout.toLowerCase().includes('enabled')
return parseSipEnabled(stdout)
} catch {
return false
}
@@ -346,10 +355,7 @@ export class KeyServiceMac {
return { success: false, error: '自动获取密钥目前仅支持 macOS' }
}
if (await this.isSipEnabled()) {
return {
success: false,
error: '当前系统还未完成连接环境准备,请按页面提示完成设置。'
}
return { success: false, code: 'SIP_ENABLED', error: SIP_ENABLED_ERROR }
}
try {
+81 -11
View File
@@ -1,6 +1,7 @@
import { monitorEventLoopDelay } from 'perf_hooks'
import * as chat from '../services/chat-service'
import type {
KnowledgeImageOcrState,
KnowledgeAttachmentMetadata,
KnowledgeEvidence,
KnowledgeMessageKind,
@@ -24,6 +25,7 @@ import {
emptyKnowledgeSearchTimings
} from '../../shared/knowledge'
import { KnowledgeService } from './knowledge-service'
import { sourceMessageId } from './message-identity'
import {
voiceAccountIdentity,
voiceMessageIdentity
@@ -118,11 +120,6 @@ function groupMemberDisplayName(member: chat.GroupSnapshot['members'][number]):
)
}
function sourceMessageId(message: chat.FormattedMessage): string {
if (message.localId) return `local:${message.localId}`
if (message.id) return String(message.id)
return `${message.createTime || 0}:${message.serverId || message.content}`
}
function sourceKind(message: chat.FormattedMessage): KnowledgeMessageKind {
if (message.voiceTranscript || message.type === '语音') return 'voice'
@@ -130,6 +127,16 @@ function sourceKind(message: chat.FormattedMessage): KnowledgeMessageKind {
return message.exportMediaType
}
if (message.exportMediaType === 'file') return 'file'
// 索引路径上 `exportMediaType` **不会被赋值**(只有 export-service 会设它),
// 所以图片/视频/表情包必须从 contentData.type 判定,否则图片会静默落成 'other',
// 进而让"图片文字索引"的 Evidence 丢掉真正的来源类型。
if (
message.contentData?.type === 'image' ||
message.contentData?.type === 'video' ||
message.contentData?.type === 'sticker'
) {
return message.contentData.type
}
if (message.contentData?.type === 'share' || message.contentData?.type === 'miniProgram') {
return message.contentData.type === 'share' && message.contentData.typeVal === '6'
? 'file'
@@ -201,12 +208,16 @@ function toSourceMessage(
accountId: string,
conversationId: string,
message: chat.FormattedMessage,
transcriptOverride?: string
transcriptOverride?: string,
imageOcr?: { state: KnowledgeImageOcrState; text: string }
): KnowledgeSourceMessage | null {
if (!message.createTime) return null
const extracted = sourceTextAndAttachment(message)
const voiceTranscript = transcriptOverride?.trim() || message.voiceTranscript?.trim() || undefined
if (!extracted.text && !extracted.attachment && !voiceTranscript) return null
// 图片 OCR 文本走与语音转写完全相同的派生通道:有文本才入库,
// 没有文字的图片(表情包/风景)不会污染索引。
const imageOcrText = imageOcr?.text?.trim() || undefined
if (!extracted.text && !extracted.attachment && !voiceTranscript && !imageOcrText) return null
return {
accountId,
conversationId,
@@ -218,7 +229,9 @@ function toSourceMessage(
kind: sourceKind(message),
text: extracted.text,
attachment: extracted.attachment,
voiceTranscript
voiceTranscript,
...(imageOcrText ? { imageOcrText } : {}),
...(imageOcrText && imageOcr?.state ? { imageOcrState: imageOcr.state } : {})
}
}
@@ -256,6 +269,14 @@ export class KnowledgeSearchService {
private interactiveIdleResolve: (() => void) | null = null
private wcdbQueueMsTotal = 0
private wcdbExecutionMsTotal = 0
/**
* 图片 OCR 文本解析器(由 main 注入)。
*
* 与语音同构:派生文本在**主进程**解析后贴到消息上,派生库不进 worker。
*/
private imageOcrResolver:
| ((conversationId: string, messageId: string) => { state: KnowledgeImageOcrState; text: string } | undefined)
| undefined
private voiceTranscriptResolver:
| ((reference: VoiceMessageReference) => VoiceTranscriptSnapshot)
| undefined
@@ -372,6 +393,15 @@ export class KnowledgeSearchService {
this.voiceTranscriptResolver = resolver
}
/** 注入图片 OCR 文本解析器(本地 System OCR 的派生结果)。 */
setImageOcrResolver(
resolver:
| ((conversationId: string, messageId: string) => { state: KnowledgeImageOcrState; text: string } | undefined)
| undefined
): void {
this.imageOcrResolver = resolver
}
/**
* A successful recognition updates its source conversation. Consecutive
* updates for the same conversation are coalesced because a complete
@@ -1051,7 +1081,7 @@ export class KnowledgeSearchService {
lane: WcdbReadLane = 'interactive'
): ReturnType<typeof chat.listMessagesAsync> {
return this.enqueueWcdbRead(
() => chat.listMessagesAsync(conversationId, startTime, endTime),
() => chat.listMessagesAsync(conversationId, startTime, endTime, undefined, undefined, 'knowledge'),
lane
)
}
@@ -1074,8 +1104,16 @@ export class KnowledgeSearchService {
const reference = this.voiceReferenceFromMessage(message)
const snapshot = reference ? this.voiceTranscriptResolver?.(reference) : undefined
const hydrated = this.withVoiceTranscript(message)
const source = toSourceMessage(accountId, conversationId, hydrated, transcriptOverride)
if (!source || source.kind !== 'voice') return source
const imageOcr = this.imageOcrResolver?.(conversationId, sourceMessageId(message))
const source = toSourceMessage(
accountId,
conversationId,
hydrated,
transcriptOverride,
imageOcr
)
if (!source) return source
if (source.kind !== 'image' && source.kind !== 'voice') return source
return {
...source,
voiceTranscriptState:
@@ -1105,6 +1143,38 @@ export class KnowledgeSearchService {
}
}
/**
* 某个会话的图片 OCR 处理完成 → 重建该会话的索引。
*
* 与"语音转写完成后单会话重索引"完全同构:整会话重读 + completeSnapshot 重建,
* 让 OCR 派生文本进入 chunks/FTS,从而可被 search_messages 检索。
* 原图片消息仍然是 authoritative source —— 这里只是让它多了一段派生文本,
* 不产生任何"OCR 消息"。
*/
async indexImageOcr(conversationId: string): Promise<void> {
if (!chat.isReady()) return
const accountId = this.currentAccountId()
if (!accountId) return
const activeIndex = this.indexing.get(accountId)
if (activeIndex) await activeIndex
const contacts = await this.listContacts()
const contact = contacts.find((item) => item.md5 === conversationId)
if (!contact) return
const messages = await this.listMessages(contact.md5, undefined, undefined, 'background')
const sourceMessages = messages
.map((message) => this.toSourceMessage(accountId, contact.md5, message))
.filter((message): message is KnowledgeSourceMessage => Boolean(message))
await this.service.index({
accountId,
conversations: [
{ conversationId: contact.md5, completeSnapshot: true, messages: sourceMessages }
],
chunker: DEFAULT_KNOWLEDGE_CHUNKER,
fts: DEFAULT_KNOWLEDGE_FTS_CONFIG
})
await this.refreshStatus(accountId)
}
private async indexVoiceTranscriptNow(update: VoiceTranscriptUpdate): Promise<void> {
if (!chat.isReady()) return
if (update.state === 'transcribed' && !update.transcript?.trim()) return
+39 -7
View File
@@ -19,7 +19,11 @@ import type {
KnowledgeSearchTimings,
KnowledgeSearchResult
} from '../../shared/knowledge'
import { emptyKnowledgeSearchTimings, KNOWLEDGE_SCHEMA_VERSION } from '../../shared/knowledge'
import {
emptyKnowledgeSearchTimings,
KNOWLEDGE_SCHEMA_VERSION,
toEvidenceDisplayText
} from '../../shared/knowledge'
import { chunkConversation } from './chunker'
import { normalizeKnowledgeMessage } from './normalizer'
@@ -761,7 +765,8 @@ export class KnowledgeStore {
evidence: asRows(
this.database
.prepare(
`SELECT m.conversation_id, m.message_id, m.create_time, m.searchable_text, m.kind, m.sender_id, m.sender_name
`SELECT m.conversation_id, m.message_id, m.create_time, m.searchable_text, m.kind, m.sender_id, m.sender_name,
m.image_ocr_text, m.voice_transcript
FROM knowledge_messages m
WHERE ${clauses.join(' AND ')}
ORDER BY m.create_time DESC
@@ -789,7 +794,22 @@ export class KnowledgeStore {
timestamp: Number(row.create_time),
messageIds: chunk ? chunk.map((item) => String(item.message_id)) : [messageId],
sourceKind: String(row.kind) as KnowledgeEvidence['sourceKind'],
text: String(row.searchable_text),
// 内部前缀(`图片文字:`)绝不能进 Evidence:面向用户与模型的是可读文本,
// 来源信息由下面的结构化字段表达。
text: toEvidenceDisplayText(String(row.searchable_text)),
...(row.image_ocr_text ? { imageOcrText: String(row.image_ocr_text) } : {}),
/*
* 来源标记按"这条消息带什么派生内容"判定,与 `sourceKind` 正交:
* `image_ocr` = 靠图片里的文字命中,`voice_transcript` = 靠语音转写命中。
*
* 两者都有时以图片 OCR 为先 —— 图片消息不会同时带语音转写,这里只是取确定值,
* 实际不会出现需要二选一的数据。
*/
...(row.image_ocr_text
? { derivedSource: 'image_ocr' as const }
: String(row.voice_transcript || '').trim()
? { derivedSource: 'voice_transcript' as const }
: {}),
score: String(row.kind) === 'system' ? 1 : 0
}
}
@@ -899,6 +919,7 @@ export class KnowledgeStore {
attachment_json TEXT,
voice_transcript TEXT,
voice_transcript_state TEXT,
image_ocr_text TEXT,
PRIMARY KEY (conversation_id, message_id)
) STRICT;
CREATE INDEX IF NOT EXISTS knowledge_messages_conversation_time
@@ -957,6 +978,14 @@ export class KnowledgeStore {
if (!messageColumns.has('voice_transcript_state')) {
this.database.exec('ALTER TABLE knowledge_messages ADD COLUMN voice_transcript_state TEXT')
}
// 图片 OCR 派生文本单独留一列(不只是埋进 searchable_text)。
//
// 为什么必须落列而不是从 searchable_text 里截字符串:Evidence 需要回答
// "这条结果是不是来自图片里的文字",并按此给出来源标记与 OCR 片段。
// 靠解析前缀来判来源,一旦前缀格式调整就会静默失效。
if (!messageColumns.has('image_ocr_text')) {
this.database.exec('ALTER TABLE knowledge_messages ADD COLUMN image_ocr_text TEXT')
}
this.writeMetaIfMissing('schema_version', String(KNOWLEDGE_SCHEMA_VERSION))
const storedAccount = this.readMeta('account_id')
if (storedAccount && storedAccount !== this.accountId) {
@@ -1137,8 +1166,9 @@ export class KnowledgeStore {
const upsert = this.database.prepare(
`INSERT INTO knowledge_messages (
account_id, conversation_id, message_id, create_time, content_hash, searchable_text,
kind, sender_id, sender_name, attachment_json, voice_transcript, voice_transcript_state
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
kind, sender_id, sender_name, attachment_json, voice_transcript, voice_transcript_state,
image_ocr_text
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(conversation_id, message_id) DO UPDATE SET
create_time = excluded.create_time,
content_hash = excluded.content_hash,
@@ -1148,7 +1178,8 @@ export class KnowledgeStore {
sender_name = excluded.sender_name,
attachment_json = excluded.attachment_json,
voice_transcript = excluded.voice_transcript,
voice_transcript_state = excluded.voice_transcript_state`
voice_transcript_state = excluded.voice_transcript_state,
image_ocr_text = excluded.image_ocr_text`
)
for (let index = 0; index < messages.length; index += 1) {
this.assertNotAborted(signal)
@@ -1165,7 +1196,8 @@ export class KnowledgeStore {
message.senderName ?? null,
message.attachment ? encodedJson(message.attachment) : null,
message.voiceTranscript ?? null,
message.voiceTranscriptState ?? null
message.voiceTranscriptState ?? null,
message.imageOcrText ?? null
)
if (index % YIELD_EVERY === 0) {
onProgress(index + 1, 0)
+24
View File
@@ -0,0 +1,24 @@
/**
* 消息身份的**唯一真源**。
*
* 这个规则同时被三处需要:
* - Knowledge 索引写入 `knowledge_messages.message_id`
* - 图片文字索引的 binding(必须与 Knowledge 里的 message_id 完全一致,否则 OCR 文本贴不到消息上)
* - Evidence → 档案跳转的 messageRef
*
* 任何一处各自复制一份,都会在 `local:` 前缀上静默失配(项目里已经有这个坑的历史注释),
* 所以抽成一个模块,谁都不许再抄。
*/
import type * as chat from '../services/chat-service'
/**
* 源消息 → 稳定消息 id。
*
* 降级顺序刻意保守:`localId` 是 WCDB 行内最稳的本地 id;其次用消息自带 id;
* 最后才退化成「时间 + 服务端 id / 内容」的组合(仅在极端缺字段时命中)。
*/
export function sourceMessageId(message: chat.FormattedMessage): string {
if (message.localId) return `local:${message.localId}`
if (message.id) return String(message.id)
return `${message.createTime || 0}:${message.serverId || message.content}`
}
+5
View File
@@ -24,6 +24,10 @@ export function normalizeKnowledgeMessage(
const transcript = compact(source.voiceTranscript)
if (transcript) sections.push(`语音转写:${transcript}`)
// 图片 OCR 文本:与语音同样的"固定前缀"约定,让检索与展示都能识别这是派生内容。
const imageText = compact(source.imageOcrText)
if (imageText) sections.push(`图片文字:${imageText}`)
const attachmentName = compact(source.attachment?.name)
if (attachmentName) {
const label = source.attachment?.kind === 'link' ? '链接' : '附件'
@@ -45,6 +49,7 @@ export function normalizeKnowledgeMessage(
senderId: source.senderId || '',
kind: source.kind,
voiceTranscriptState: source.voiceTranscriptState || '',
imageOcrState: source.imageOcrState || '',
searchableText
})
)
+600 -16
View File
@@ -1,3 +1,5 @@
import { describeRedPacketStatus, describeTransferStatus } from '../shared/payment-status'
type TextContent = { type: 'text'; content: string }
type VoiceContent = { type: 'voice'; duration?: number }
type LocationContent = {
@@ -22,6 +24,44 @@ type ShareContent = {
appname?: string
typeVal?: string
articles?: ShareArticle[]
transfer?: TransferPaymentInfo
}
type TransferPaymentInfo = {
paySubtype?: string
amountText?: string
transcationId?: string
transferId?: string
invalidTime?: string
beginTransferTime?: string
effectiveDate?: string
payMemo?: string
receiverUsername?: string
payerUsername?: string
transferStatus?: string
/** `transfer_status` 展示文案(只读)。 */
transferStatusText?: string
transId?: string
feeType?: string
transferAttach?: string
refundBankType?: string
}
type RedPacketPaymentInfo = {
templateId?: string
receiveTitle?: string
sendTitle?: string
sceneText?: string
senderDes?: string
receiverDes?: string
iconUrl?: string
nativeUrl?: string
sendId?: string
hbType?: string
hbStatus?: string
receiveStatus?: string
/** `hb_status` / `receive_status` 展示文案(只读)。 */
redPacketStatusText?: string
}
type ForwardedMessageItem = {
messageType: number
@@ -51,6 +91,7 @@ type RedPacketContent = {
title: string
description?: string
url?: string
pay?: RedPacketPaymentInfo
}
type VoipContent = { type: 'voip'; duration?: number; status: string; roomType?: number }
type ImageContent = {
@@ -122,21 +163,48 @@ export type ParsedContent =
| SystemContent
| UnknownContent
/**
* `local_type` 可能是打包 u64:低 32 位为类型枚举,高 32 位为标志/子类型。
*/
export function normalizeLocalType(messageType: number | string): number {
try {
const value =
typeof messageType === 'number'
? BigInt(Math.trunc(messageType))
: BigInt(String(messageType).trim())
return Number(value & 0xffffffffn)
} catch {
const fallback = Number(messageType)
return Number.isFinite(fallback) ? Math.trunc(fallback) : 0
}
}
export function parseMessageContent(content: string, messageType: number): ParsedContent {
// Voice rows may keep their binary payload outside msgContent, so an empty
const localType = normalizeLocalType(messageType)
// content string is still a valid voice message.
if (messageType === 34) return { type: 'voice' }
if (localType === 34) {
const duration = parseVoiceDurationSeconds(content)
return duration === undefined ? { type: 'voice' } : { type: 'voice', duration }
}
if (!content || typeof content !== 'string') {
return { type: 'unknown', raw: content || '' }
}
const normalized = content.trim()
switch (messageType) {
switch (localType) {
case 1:
return { type: 'text', content: normalized }
case 2:
return parseLegacyType2(normalized)
case 3:
return parseImageMessage(normalized)
case 8:
return parseLegacyType8(normalized)
case 17:
return parseForwardBundle(normalized)
case 37:
return parseFriendVerifyMessage(normalized)
case 42:
return parseCardMessage(normalized)
case 43:
@@ -153,10 +221,37 @@ export function parseMessageContent(content: string, messageType: number): Parse
case 10002:
return parseSystemMessage(normalized)
default:
return { type: 'unknown', raw: normalized, messageType }
return { type: 'unknown', raw: normalized, messageType: localType }
}
}
/**
* 语音时长藏在解压后的 message_content 里:`<voicemsg ... voicelength="1600" ...>`,单位毫秒。
* Msg_* 表没有 voice_length 列,这是唯一来源。
*
* `<voicemsg>` 上两个极易混淆的属性(真机实测,同一条 1.6 秒语音,2026-09-20):
*
* - `voicelength="1600"` → **毫秒时长**。这条语音微信气泡显示 2"(1.6 秒四舍五入)。
* **要取的是它。**
* - `length="6672"` → **SILK 编码数据的字节数,与时长无关**。
* 已验证:`wcdb_get_voice_data` 取出的 SILK 恰好是 6672 字节,
* 解码后为 51200 字节 PCM(1.6 秒)。误取它会算出 6.672 秒,把 2" 显示成 0:07。
*
* 换算成秒后**刻意保留小数**(1600ms → 1.6):在这里取整会把精度永久丢掉,
* 后面显示层再怎么四舍五入都对不回微信的口径(微信是四舍五入到整秒)。
*
* 注:`<videomsg length="...">` 的 `length` 同理是字节数,不是时长。
*/
function parseVoiceDurationSeconds(content: string): number | undefined {
if (!content || typeof content !== 'string') return undefined
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const rawLength = extractXmlAttribute(decoded, 'voicemsg', 'voicelength')
if (!rawLength) return undefined
const milliseconds = Number(rawLength)
if (!Number.isFinite(milliseconds) || milliseconds <= 0) return undefined
return milliseconds / 1000
}
function parseVideoMessage(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const md5 = normalizeMd5(extractXmlAttribute(decoded, 'videomsg', 'md5'))
@@ -181,6 +276,14 @@ function parseSystemMessage(content: string): ParsedContent {
recall
}
}
const templateText = extractSysmsgTemplateText(decoded)
if (templateText) {
return {
type: 'system',
content: templateText,
raw: content
}
}
const delChatroomMemberText = extractDelChatroomMemberText(decoded)
if (delChatroomMemberText) {
return {
@@ -189,16 +292,19 @@ function parseSystemMessage(content: string): ParsedContent {
raw: content
}
}
// <link_list> 里放的是富文本片段(可能含 hidden="1" 的可点击按钮),
// 不是消息正文;先剥掉再走通用提取,避免把按钮文案当成整条系统消息。
const withoutLinkList = stripSysmsgLinkList(decoded)
const plainText =
extractXmlNodeText(decoded, 'plain') ||
extractXmlNodeText(decoded, 'text') ||
extractXmlNodeText(decoded, 'title') ||
extractXmlValue(decoded, 'plain') ||
extractXmlValue(decoded, 'text') ||
extractXmlValue(decoded, 'title') ||
extractXmlNodeText(withoutLinkList, 'plain') ||
extractXmlNodeText(withoutLinkList, 'text') ||
extractXmlNodeText(withoutLinkList, 'title') ||
extractXmlValue(withoutLinkList, 'plain') ||
extractXmlValue(withoutLinkList, 'text') ||
extractXmlValue(withoutLinkList, 'title') ||
''
const normalized = normalizeSystemText(plainText || fallbackSystemText(decoded))
const normalized = normalizeSystemText(plainText || fallbackSystemText(withoutLinkList))
return {
type: 'system',
content: normalized || '[系统消息]',
@@ -439,6 +545,30 @@ function parseLocationMessage(content: string): ParsedContent {
function parseShareMessage(content: string): ParsedContent {
const appMsgType = extractAppMsgType(content)
const isFileMessage = appMsgType === '6' || appMsgType === '74'
if (/<patMsg\b/i.test(content) || /sysmsg[^>]+type=["']pat["']/i.test(content)) {
return parsePatMessage(content)
}
if (/<findernamecard\b/i.test(content)) {
return parseFinderNameCard(content)
}
if (/<productitem\b/i.test(content) && !extractXmlValue(content, 'title')) {
return parseProductItem(content)
}
if (looksLikeMusicShare(content)) {
return parseMusicShareCard(content)
}
if (looksLikeSubscribeCard(content)) {
return parseSubscribeCard(content)
}
if (looksLikeKefuCard(content)) {
return parseKefuCard(content)
}
if (looksLikeStreamVideo(content)) {
return parseStreamVideoCard(content)
}
if (looksLikeGiftCard(content)) {
return parseGiftCard(content)
}
if (appMsgType === '19') {
return parseForwardBundle(content)
}
@@ -481,12 +611,25 @@ function parseShareMessage(content: string): ParsedContent {
}
}
if (appMsgType === '2001') {
if (appMsgType === '2001' || /mmpayhb|receivehongbao|wxpay:\/\/c2cbizmessagehandler\/hongbao/i.test(content)) {
const pay = parseWcpayInfo(content)
const url = decodeXmlUrl(extractXmlValue(content, 'url')) || undefined
return {
type: 'redPacket',
title: extractXmlValue(content, 'title') || '微信红包',
description: extractXmlValue(content, 'des') || '恭喜发财,大吉大利',
url: decodeXmlUrl(extractXmlValue(content, 'url')) || undefined
title:
pay.sendTitle ||
pay.receiveTitle ||
extractXmlValue(content, 'title') ||
'微信红包',
description:
extractXmlValue(content, 'des') ||
pay.sceneText ||
'恭喜发财,大吉大利',
url,
pay: {
...pay,
sendId: pay.sendId || extractSendIdFromUrl(url)
}
}
}
@@ -506,6 +649,12 @@ function parseShareMessage(content: string): ParsedContent {
const typeVal = extractXmlValue(content, 'type') || ''
if (!title && !url) {
if (/<productitem\b/i.test(content)) return parseProductItem(content)
if (looksLikeMusicShare(content)) return parseMusicShareCard(content)
if (looksLikeSubscribeCard(content)) return parseSubscribeCard(content)
if (looksLikeKefuCard(content)) return parseKefuCard(content)
if (looksLikeStreamVideo(content)) return parseStreamVideoCard(content)
if (looksLikeGiftCard(content)) return parseGiftCard(content)
return { type: 'unknown', raw: content }
}
@@ -516,7 +665,372 @@ function parseShareMessage(content: string): ParsedContent {
url,
appname,
typeVal,
articles: articles.length > 1 ? articles : undefined
articles: articles.length > 1 ? articles : undefined,
transfer: typeVal === '2000' ? parseWcpayInfo(content) : undefined
}
}
/** 只读解析 `<wcpayinfo>`(转账 / 红包展示字段;不做支付)。 */
function parseWcpayInfo(content: string): TransferPaymentInfo & RedPacketPaymentInfo {
const block = /<wcpayinfo>([\s\S]*?)<\/wcpayinfo>/i.exec(content)?.[1] || content
const val = (tag: string): string | undefined => {
const raw = extractXmlValue(block, tag)
return raw ? decodeXmlEntities(raw) || undefined : undefined
}
const transferStatus = val('transfer_status')
const hbStatus = val('hb_status')
const receiveStatus = val('receive_status')
return {
paySubtype: val('paysubtype'),
amountText: val('feedesc'),
transcationId: val('transcationid'),
transferId: val('transferid'),
invalidTime: val('invalidtime'),
beginTransferTime: val('begintransfertime'),
effectiveDate: val('effectivedate'),
payMemo: val('pay_memo') || undefined,
receiverUsername: val('receiver_username'),
payerUsername: val('payer_username'),
transferStatus,
transferStatusText: describeTransferStatus(transferStatus),
transId: val('trans_id'),
feeType: val('fee_type'),
transferAttach: val('transfer_attach'),
refundBankType: val('refund_bank_type'),
templateId: val('templateid'),
receiveTitle: val('receivertitle'),
sendTitle: val('sendertitle'),
sceneText: val('scenetext'),
senderDes: val('senderdes'),
receiverDes: val('receiverdes'),
iconUrl: decodeXmlUrl(val('iconurl') || '') || undefined,
nativeUrl: val('nativeurl'),
hbType: val('hb_type'),
hbStatus,
receiveStatus,
redPacketStatusText: describeRedPacketStatus(hbStatus, receiveStatus)
}
}
function extractSendIdFromUrl(url?: string): string | undefined {
if (!url) return undefined
try {
return new URL(url).searchParams.get('sendid') || undefined
} catch {
return /sendid=(\d+)/i.exec(url)?.[1]
}
}
/** 拍一拍(`HandlePatMsg` / `<patMsg>`)→ 系统文案。 */
function parsePatMessage(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const template =
extractXmlValue(decoded, 'template') ||
extractXmlValue(decoded, 'pattemplate') ||
extractXmlValue(decoded, 'plain') ||
''
const from = extractXmlValue(decoded, 'fromusername') || extractXmlValue(decoded, 'fromuser')
const to = extractXmlValue(decoded, 'tousername') || extractXmlValue(decoded, 'touser')
const pat = extractXmlValue(decoded, 'pat') || extractXmlValue(decoded, 'patted')
const text =
normalizeSystemText(
template.replace(/\$from\$/g, from || '').replace(/\$to\$/g, to || '').replace(/\$pat\$/g, pat || '')
) || normalizeSystemText(decoded.replace(/<[^>]+>/g, ' ')) || '拍了拍'
return { type: 'system', content: text, raw: content }
}
/** 视频号名片(`<findernamecard>`)→ 只读 share。 */
function parseFinderNameCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const username = extractXmlValue(decoded, 'username') || extractXmlValue(decoded, 'finderUsername')
const nickname = extractXmlValue(decoded, 'nickname') || extractXmlValue(decoded, 'finderNickname')
const title = nickname || username || extractXmlValue(decoded, 'title') || '视频号名片'
const url = decodeXmlUrl(extractXmlValue(decoded, 'url') || extractXmlValue(decoded, 'appPageUrl') || '')
return {
type: 'share',
title,
des: username && nickname ? username : undefined,
url: url || '',
appname: '视频号',
typeVal: extractAppMsgType(content) || 'findernamecard'
}
}
/** 商品卡(`<productitem>`)→ 只读 share。 */
function parseProductItem(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'productName') ||
extractXmlValue(decoded, 'product_name') ||
extractXmlValue(decoded, 'title') ||
'商品'
const des =
extractXmlValue(decoded, 'productDesc') ||
extractXmlValue(decoded, 'sellingPrice') ||
extractXmlValue(decoded, 'referdes')
const url = decodeXmlUrl(extractXmlValue(decoded, 'url') || extractXmlValue(decoded, 'productUrl') || '')
return {
type: 'share',
title,
des: des || undefined,
url: url || '',
appname: extractXmlValue(decoded, 'sellername') || '商品',
typeVal: extractAppMsgType(content) || 'productitem'
}
}
/** 好友验证(local_type=37)→ 系统文案,避免 unknown 黑块。 */
function parseFriendVerifyMessage(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const text = extractSysmsgTemplateText(decoded)
if (text) return { type: 'system', content: text, raw: content }
const nickname = extractXmlValue(decoded, 'nickname') || extractXmlValue(decoded, 'NickName')
const username = extractXmlValue(decoded, 'username') || extractXmlValue(decoded, 'UserName')
const contentText =
extractXmlValue(decoded, 'content') ||
extractXmlValue(decoded, 'bighead') ||
extractXmlValue(decoded, 'source')
const label = normalizeSystemText(
[nickname || username ? `${nickname || username}` : '', contentText || '好友验证消息']
.filter(Boolean)
.join(' · ')
)
return { type: 'system', content: label || '好友验证消息', raw: content }
}
function looksLikeMusicShare(content: string): boolean {
return /songalbumurl|songlyric|musicShareItem|tingListenItem|ListenItem|music_share_item|musicurl|<songlyri|<songalbu/i.test(
content
)
}
function looksLikeSubscribeCard(content: string): boolean {
return /subscribeMessage|subscribe_msg|SubscribeMsg|updatablemsg|wadynamicpageinf|OnSubscriptionCustom/i.test(
content
)
}
function looksLikeKefuCard(content: string): boolean {
return /opencustomerservicemsg|wa_app_kefu_message|kefumenu|kf_order|kf_user_|ChatKfTemplate|AppReaderTemplate|template_header/i.test(
content
)
}
/** 流视频 / 长视频卡(`streamvideotitle` / `finderMegaVideo`)。 */
function looksLikeStreamVideo(content: string): boolean {
return /streamvideotitle|streamvideoword|streamvideoweburl|streamvideothumb|finderMegaVideo|finderLiveInvite/i.test(
content
)
}
/**
* 礼物 / 礼品卡(`csgift` / `giftcarditem`)。
* 只读展示;**不**调用 `acceptgiftcard` / `preacceptgiftcard` / `getcardgiftinfo`。
*/
function looksLikeGiftCard(content: string): boolean {
return /<csgift\b|<ecsgift\b|<giftcarditem\b|<giftcard\b|giftcarditem|acceptgiftcard/i.test(content)
}
/** 流视频卡 → 只读 share。 */
function parseStreamVideoCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'streamvideotitle') ||
extractXmlValue(decoded, 'title') ||
extractXmlValue(decoded, 'sourcetitle') ||
'视频'
const des =
extractXmlValue(decoded, 'streamvideoword') ||
extractXmlValue(decoded, 'des') ||
extractXmlValue(decoded, 'contentdescshowtext') ||
undefined
const url =
decodeXmlUrl(
extractXmlValue(decoded, 'streamvideoweburl') ||
extractXmlValue(decoded, 'url') ||
extractXmlValue(decoded, 'weburl')
) || ''
return {
type: 'share',
title,
des,
url,
appname: /finderMegaVideo|finderLiveInvite/i.test(content) ? '视频号' : '视频',
typeVal: extractAppMsgType(content) || 'streamvideo'
}
}
/** 礼物 / 礼品卡 → 只读 share(无 title 时 system 文案)。 */
function parseGiftCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'title') ||
extractXmlValue(decoded, 'gifttitle') ||
extractXmlValue(decoded, 'cardtitle') ||
extractXmlValue(decoded, 'brandname')
const des =
extractXmlValue(decoded, 'des') ||
extractXmlValue(decoded, 'giftwording') ||
extractXmlValue(decoded, 'cardwording') ||
extractXmlValue(decoded, 'description')
const url = decodeXmlUrl(extractXmlValue(decoded, 'url') || extractXmlValue(decoded, 'cardurl') || '')
if (!title && !des) {
return {
type: 'system',
content: normalizeSystemText(decoded.replace(/<[^>]+>/g, ' ')) || '礼物卡',
raw: content
}
}
return {
type: 'share',
title: title || '礼物卡',
des: des || undefined,
url: url || '',
appname: extractXmlValue(decoded, 'brandname') || '礼物',
typeVal: extractAppMsgType(content) || 'giftcard'
}
}
/**
* local_type=2:真机 histogram 未采到;按内容形态尽量落到 system/text,
* 否则保留 unknown(见 wechat-message-type-coverage.md)。
*/
function parseLegacyType2(content: string): ParsedContent {
if (/<sysmsg\b|<patMsg\b|revoke/i.test(content)) {
return parseSystemMessage(content)
}
if (!/<[a-zA-Z!]/.test(content)) {
return { type: 'text', content }
}
return parseSystemMessage(content)
}
/**
* local_type=8:社区多见于 GIF/大表情变体;无真机样本时按 sticker → image → system
* 逐级尝试,避免直接黑块。
*/
function parseLegacyType8(content: string): ParsedContent {
const sticker = parseStickerMessage(content)
if (sticker.type === 'sticker' && sticker.md5) return sticker
const image = parseImageMessage(content)
if (image.type === 'image' && (image.md5 || image.datName)) return image
if (/<sysmsg\b|<emoji\b|<msg\b/i.test(content)) {
const system = parseSystemMessage(content)
if (system.type === 'system' && system.content && system.content !== content) return system
}
return { type: 'unknown', raw: content, messageType: 8 }
}
/** 音乐 / 听歌分享(`musicShareItem` / `ListenItem` / `songalbumurl`)→ 只读 share。 */
function parseMusicShareCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'title') ||
extractXmlValue(decoded, 'songname') ||
extractXmlValue(decoded, 'musicTitle') ||
'音乐分享'
const des =
extractXmlValue(decoded, 'des') ||
extractXmlValue(decoded, 'singername') ||
extractXmlValue(decoded, 'albumname') ||
undefined
const url =
decodeXmlUrl(
extractXmlValue(decoded, 'url') ||
extractXmlValue(decoded, 'musicurl') ||
extractXmlValue(decoded, 'streamweburl') ||
extractXmlValue(decoded, 'weburl')
) || ''
const appname =
extractXmlValue(decoded, 'appname') ||
extractXmlValue(decoded, 'publisher') ||
(looksLikeListenItem(content) ? '听一听' : '音乐')
return {
type: 'share',
title,
des,
url,
appname,
typeVal: extractAppMsgType(content) || 'music'
}
}
function looksLikeListenItem(content: string): boolean {
return /tingListenItem|ListenItem|MMLISTEN_ITEM_TYPE/i.test(content)
}
/**
* 订阅号 / 可更新消息模板(`subscribeMessage` / `updatablemsg`)→ 只读 share。
* 模板头 `template_header` / `template_detail` 作标题与摘要。
*/
function parseSubscribeCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'template_header') ||
extractXmlValue(decoded, 'title') ||
extractXmlValue(decoded, 'templatetitle') ||
'订阅消息'
const des =
extractXmlValue(decoded, 'template_detail') ||
extractXmlValue(decoded, 'des') ||
extractXmlValue(decoded, 'digest') ||
undefined
const url =
decodeXmlUrl(
extractXmlValue(decoded, 'url') ||
extractXmlValue(decoded, 'jumpUrl') ||
extractXmlValue(decoded, 'weburl')
) || ''
return {
type: 'share',
title,
des,
url,
appname: extractXmlValue(decoded, 'appname') || '订阅消息',
typeVal: extractAppMsgType(content) || 'subscribe'
}
}
/**
* 客服 / 门店模板卡(`opencustomerservicemsg` / `wa_app_kefu_message` / `kefumenu`)
* → 只读 share;无 title 时退化为 system 文案,避免 unknown 黑块。
*/
function parseKefuCard(content: string): ParsedContent {
const decoded = decodeXmlEntities(stripChatroomPrefix(content))
const title =
extractXmlValue(decoded, 'template_header') ||
extractXmlValue(decoded, 'title') ||
extractXmlValue(decoded, 'kf_title') ||
extractXmlValue(decoded, 'templatetitle')
const des =
extractXmlValue(decoded, 'template_detail') ||
extractXmlValue(decoded, 'kf_order_text') ||
extractXmlValue(decoded, 'des') ||
extractXmlValue(decoded, 'content')
const url =
decodeXmlUrl(
extractXmlValue(decoded, 'url') ||
extractXmlValue(decoded, 'jumpUrl') ||
extractXmlValue(decoded, 'weburl')
) || ''
const appname =
extractXmlValue(decoded, 'kf_user_name') ||
extractXmlValue(decoded, 'appname') ||
'客服消息'
if (!title && !des) {
return {
type: 'system',
content: normalizeSystemText(decoded.replace(/<[^>]+>/g, ' ')) || '客服消息',
raw: content
}
}
return {
type: 'share',
title: title || '客服消息',
des: des || undefined,
url: url || '',
appname,
typeVal: extractAppMsgType(content) || 'kefu'
}
}
@@ -838,6 +1352,76 @@ function extractDelChatroomMemberText(xml: string): string {
return ''
}
/** 剥掉 <link_list> 区块 —— 其中的文案属于富文本片段,不是消息正文。 */
function stripSysmsgLinkList(xml: string): string {
return String(xml || '').replace(/<link_list\b[\s\S]*?<\/link_list>/gi, ' ')
}
/**
* 微信 4.x 起,部分系统消息改成「模板」格式,正文不再写在 <plain> 里。
*
* 旧格式(正文就在 <plain>,取到即可):
*
* <sysmsg type="delchatroommember"><delchatroommember>
* <plain><![CDATA["成员昵称"通过扫描你分享的二维码加入群聊]]></plain>
* <link><scene>qrcode</scene><text><![CDATA[撤销]]></text>…</link>
* </delchatroommember></sysmsg>
*
* 新格式(<plain> 变空,正文挪进 <template>,用 $名称$ 引用 <link_list> 里的 link):
*
* <sysmsg type="sysmsgtemplate"><sysmsgtemplate>
* <content_template type="tmpl_type_profilewithrevokeqrcode">
* <plain><![CDATA[]]></plain>
* <template><![CDATA["$adder$"通过扫描你分享的二维码加入群聊 $revoke$]]></template>
* <link_list>
* <link name="adder" type="link_profile">
* <memberlist><member><nickname><![CDATA[成员昵称]]></nickname></member></memberlist>
* </link>
* <link name="revoke" type="link_revoke_qrcode" hidden="1">
* <title><![CDATA[撤销]]></title>
* </link>
* </link_list>
* </content_template>
* </sysmsgtemplate></sysmsg>
*
* 两个要点:
* 1. 正文取自 <template>,其中的 $名称$ 占位符按 <link_list> 的 link name 回填;
* 2. hidden="1" 的 link 在微信里是可点击按钮,纯文本展示时省略其文案。
*
* 漏掉这段会让新格式消息落进通用提取链:<plain> 为空、又没有 <text>,
* 于是取到 <title> —— 也就是那个隐藏按钮的标题,整条系统消息只剩一个按钮名。
*/
function extractSysmsgTemplateText(xml: string): string {
if (!/<sysmsgtemplate\b|<content_template\b/i.test(xml)) return ''
const template = extractXmlNodeText(xml, 'template')
if (!template) return ''
const links = new Map<string, { text: string; hidden: boolean }>()
const linkPattern = /<link\b([^>]*)>([\s\S]*?)<\/link>/gi
let linkMatch: RegExpExecArray | null
while ((linkMatch = linkPattern.exec(xml)) !== null) {
const name = extractXmlValue(linkMatch[1], 'name')
if (!name) continue
links.set(name, {
// title 用于按钮文案,nickname 用于成员展示名,text 作最后兜底。
text:
extractXmlNodeText(linkMatch[2], 'title') ||
extractXmlNodeText(linkMatch[2], 'nickname') ||
extractXmlNodeText(linkMatch[2], 'text') ||
'',
hidden: /hidden\s*=\s*["']1["']/i.test(linkMatch[1])
})
}
const rendered = template.replace(/\$([A-Za-z0-9_]+)\$/g, (_raw, name: string) => {
const link = links.get(name)
return link && !link.hidden ? link.text : ''
})
return normalizeSystemText(rendered)
}
function normalizeMd5(value: unknown): string | undefined {
const md5 = String(value || '')
.trim()
+43 -15
View File
@@ -23,6 +23,45 @@ import {
const aiProvider = new AIProviderService()
/**
* 用真实群成员快照补全消息的显示名与头像。
*
* **头像与名称的门槛刻意不同**:
*
* 此前这里先用 `isInternalName(message.name)` 做整体早退 —— 只有当消息里的名字还是内部标识
* (空 / `wxid_*` / `*@chatroom` / 18+ 位字母数字)时才继续。但 `listMessages` 产出的 `name`
* 优先取 `senderNickname`,在真实群里通常是**已可读的昵称**,于是整条记录被跳过、
* `member.avatar` 永远补不上,导出层只能退化成首字头像;软件内日报没有这个门槛,
* 所以它能显示真实头像。
*
* 现在:**只要 senderId 命中真实群成员就允许补头像**;名称只在解析结果确实是可读名时才采用,
* 避免把调用方已有的昵称降级成空串或内部标识。消息自带的 `img` 始终优先。
*/
export function hydrateGroupMemberIdentity(
messages: Message[],
members: ReadonlyArray<NonNullable<ReturnType<typeof getGroupSnapshot>>['members'][number]>,
memberNameMode: ScheduledReportMemberNameMode
): Message[] {
const index = new Map(
members.map((member) => [
member.wxid,
{ name: resolveMemberName(member, memberNameMode), avatar: member.avatar }
])
)
return messages.map((message) => {
const member = index.get(String(message.senderId || message.name || ''))
if (!member) return message
const shouldFillAvatar = !message.img && Boolean(member.avatar)
const resolvedName = member.name && !isInternalName(member.name) ? member.name : message.name
if (!shouldFillAvatar && resolvedName === message.name) return message
return {
...message,
name: resolvedName,
...(shouldFillAvatar ? { img: member.avatar } : {})
}
})
}
export interface AgentGroupReportRequest {
group: string
range?: SummaryDateRange | 'recent24h'
@@ -101,22 +140,11 @@ export async function generateAgentGroupReport(
const snapshot = getGroupSnapshot(contact.md5)
if (snapshot) {
const members = new Map(
snapshot.members.map((member) => [
member.wxid,
{
name: resolveMemberName(member, request.memberNameMode || 'groupNickname'),
avatar: member.avatar
}
])
messages = hydrateGroupMemberIdentity(
messages,
snapshot.members,
request.memberNameMode || 'groupNickname'
)
messages = messages.map((message) => {
if (!isInternalName(message.name)) return message
const member = members.get(String(message.senderId || message.name || ''))
return member?.name
? { ...message, name: member.name, img: message.img || member.avatar }
: message
})
}
const input = await buildGroupReportInput(messages, contact as Contact, true, 'full')
@@ -0,0 +1,193 @@
import { chmodSync, mkdirSync, readFileSync, renameSync, rmSync, writeFileSync } from 'node:fs'
import { dirname } from 'node:path'
import { randomUUID } from 'node:crypto'
import {
agentHubKindPlaceholder,
type AgentHubConversation,
type AgentHubConversationMessage,
type AgentHubConversationSummary,
type AgentHubMessageKind,
type AgentHubMessageStatus
} from '../../shared/agent-hub-conversation'
/**
* Agent Hub 对话记录存储。
*
* 设计取向:
* - 一个 JSON 文件装全部会话,按「最近活跃」排序,避免为几十个会话开目录;
* - 每个会话最多保留 N 条(默认 500),超出丢最旧的,文件不会无限增长;
* - 会话数也有上限(默认 50),只保留最近活跃的;
* - 原子写(tmp + rename)+ 0600,任何一个环节失败都不能影响真实收发。
*
* 与 `WechatInboundInbox` 的区别:收件箱是**待处理的在途消息**(处理完即删),
* 这里是**供人回看的历史**(按上限长期保留)。
*/
const DEFAULT_MAX_MESSAGES_PER_CONVERSATION = 500
const DEFAULT_MAX_CONVERSATIONS = 50
interface ConversationFile {
version: 1
conversations: AgentHubConversation[]
}
export interface AgentHubConversationStoreOptions {
filePath: () => string
maxMessagesPerConversation?: number
maxConversations?: number
now?: () => number
createId?: () => string
}
export interface AppendInput {
userId: string
accountId?: string
direction: 'in' | 'out'
kind: AgentHubMessageKind
text: string
messageId?: string
status?: AgentHubMessageStatus
errorCode?: string
createdAt?: number
}
export class AgentHubConversationStore {
private readonly options: AgentHubConversationStoreOptions
private readonly maxMessages: number
private readonly maxConversations: number
private cache: AgentHubConversation[] | null = null
constructor(options: AgentHubConversationStoreOptions) {
this.options = options
this.maxMessages = options.maxMessagesPerConversation ?? DEFAULT_MAX_MESSAGES_PER_CONVERSATION
this.maxConversations = options.maxConversations ?? DEFAULT_MAX_CONVERSATIONS
}
/** 追加一条消息,返回受影响会话的摘要与这条消息(供 UI 增量刷新)。 */
append(
input: AppendInput
): { summary: AgentHubConversationSummary; message: AgentHubConversationMessage } | null {
const userId = String(input.userId || '').trim()
if (!userId) return null
const createdAt = input.createdAt ?? this.options.now?.() ?? Date.now()
const message: AgentHubConversationMessage = {
id: this.options.createId?.() ?? randomUUID(),
direction: input.direction,
kind: input.kind,
text: String(input.text ?? ''),
createdAt,
...(input.status ? { status: input.status } : {}),
...(input.errorCode ? { errorCode: input.errorCode } : {}),
...(input.messageId ? { messageId: input.messageId } : {})
}
try {
const conversations = this.load()
const index = conversations.findIndex((item) => item.userId === userId)
if (index >= 0) {
const existing = conversations[index]
const messages = [...existing.messages, message].slice(-this.maxMessages)
conversations[index] = {
...existing,
...(input.accountId ? { accountId: input.accountId } : {}),
lastAt: createdAt,
messages
}
} else {
conversations.push({
userId,
...(input.accountId ? { accountId: input.accountId } : {}),
firstAt: createdAt,
lastAt: createdAt,
messages: [message]
})
}
// 只保留最近活跃的 maxConversations 个会话。
const trimmed = conversations
.sort((left, right) => right.lastAt - left.lastAt)
.slice(0, this.maxConversations)
this.persist(trimmed)
this.cache = trimmed
const summary = this.toSummary(trimmed.find((item) => item.userId === userId) ?? null)
return summary ? { summary, message } : null
} catch (error) {
// 记录历史失败绝不能影响真实收发。
console.warn('[AgentHubConversation] 写入对话记录失败:', error)
return null
}
}
listSummaries(): AgentHubConversationSummary[] {
return this.load()
.slice()
.sort((left, right) => right.lastAt - left.lastAt)
.map((conversation) => this.toSummary(conversation))
.filter((summary): summary is AgentHubConversationSummary => summary !== null)
}
get(userId: string): AgentHubConversation | null {
const normalized = String(userId || '').trim()
if (!normalized) return null
const found = this.load().find((item) => item.userId === normalized)
if (!found) return null
return { ...found, messages: found.messages.map((message) => ({ ...message })) }
}
clear(): void {
this.persist([])
this.cache = []
}
private toSummary(conversation: AgentHubConversation | null): AgentHubConversationSummary | null {
if (!conversation) return null
const last = conversation.messages[conversation.messages.length - 1]
const preview = last ? last.text.trim() || agentHubKindPlaceholder(last.kind) : ''
return {
userId: conversation.userId,
...(conversation.accountId ? { accountId: conversation.accountId } : {}),
firstAt: conversation.firstAt,
lastAt: conversation.lastAt,
messageCount: conversation.messages.length,
lastPreview: preview,
lastDirection: last?.direction ?? 'in'
}
}
private load(): AgentHubConversation[] {
if (this.cache) return this.cache
let parsed: ConversationFile | null = null
try {
parsed = JSON.parse(readFileSync(this.options.filePath(), 'utf8')) as ConversationFile
} catch {
parsed = null
}
const conversations = Array.isArray(parsed?.conversations)
? parsed!.conversations.filter((item): item is AgentHubConversation =>
Boolean(item && typeof item === 'object' && String(item.userId || '').trim())
)
: []
this.cache = conversations
return conversations
}
private persist(conversations: AgentHubConversation[]): void {
const path = this.options.filePath()
mkdirSync(dirname(path), { recursive: true, mode: 0o700 })
const tempPath = `${path}.tmp-${process.pid}-${Date.now()}`
try {
writeFileSync(
tempPath,
JSON.stringify({ version: 1, conversations } satisfies ConversationFile),
{ encoding: 'utf8', mode: 0o600 }
)
chmodSync(tempPath, 0o600)
renameSync(tempPath, path)
chmodSync(path, 0o600)
} catch (error) {
rmSync(tempPath, { force: true })
throw error
}
}
}
+36 -1
View File
@@ -87,15 +87,50 @@ export function matchGroupMemberChatIntent(text: string): GroupMemberChatIntent
return null
}
/**
* 「列出会话」的**形态标记**:用户在问"有哪些 / 和谁 / 列表 / 最近 N 条",
* 而不是在问"聊了什么内容"。
*/
const RECENT_LIST_SHAPE = /(哪些|哪个|都有谁|都跟谁|和谁|跟谁|是谁|列表|名单|\d{1,2}(条|个|位))/
/** 名单类问题必须落到会话 / 联系人这个对象上。 */
const RECENT_LIST_TARGET = /(聊天|会话|联系人|好友|人|群|消息|窗口)/
/** 「和谁 / 跟谁」问法本身就在问会话对象,不要求额外载体词。 */
const RECENT_PEER_QUESTION = /(和谁|跟谁)/
/** 内容探针:问的是消息里的内容 / 是否提到某事物 —— 必须交给 Query Agent。 */
const RECENT_CONTENT_PROBE =
/(提到|提过|说过|说啥|说什么|说了什么|聊了啥|聊了什么|都聊什么|都说什么|什么话题|聊到|讨论|内容|讲了什么|哪条|哪一句|有没有|是否)/
/**
* "最近有哪些会话"类请求。
*
* 这是**确定性能力**(列出会话),不是消息内容查询 —— Query Agent 无法表达,
* 因此保留为不经过模型的无 LLM 快捷路径。
*
* 判定必须**正向**:只有用户确实在要一份"会话 / 联系人名单"时才算 recent_list。
* 早先的实现只要求「最近」+「消息|会话|聊天」同时出现,于是
* 「最近群里聊的消息里有没有提到报价?」这类**内容查询**会被截走,
* 直接回一串会话名,用户永远得不到答案。
*/
export function matchRecentChatIntent(text: string): number | null {
const normalized = text.replace(/\s+/g, '')
if (!normalized.includes('最近') || !/(消息|会话|聊天)/.test(normalized)) return null
if (!normalized.includes('最近')) return null
// 内容探针优先排除:问"有没有提到 X / 谁提过 X / 聊了什么"是在查消息内容,不是要名单。
if (RECENT_CONTENT_PROBE.test(normalized)) return null
/**
* 只有两种形态算"要最近会话列表":
* 1) 「和谁 / 跟谁」问法 —— 它本身就在问会话对象,不需要额外的载体词
* (如「最近和谁聊过」);
* 2) 名单形态 + 会话载体 —— 如「最近有哪些聊天」「最近 5 个会话」「最近3条消息」。
*
* 两种都不满足时交给 Query Agent:形状不像"要名单"的,就是在问内容。
*/
const listLike =
RECENT_PEER_QUESTION.test(normalized) ||
(RECENT_LIST_SHAPE.test(normalized) && RECENT_LIST_TARGET.test(normalized))
if (!listLike) return null
const limit = Number(normalized.match(/\d{1,2}/)?.[0] || 5)
return Math.max(1, Math.min(20, limit))
}
File diff suppressed because it is too large Load Diff
+20 -3
View File
@@ -249,16 +249,33 @@ function selectCoverage(
}
}
/**
* Citation sanitize 的允许集合。
*
* 两个入口都只能引用「Host 侧已分配」的编号,但它们的形状不同:
* - Query Agent:`citationId` 字符串集合(`EvidenceCollector` 分配的 `E1`…`En`);
* - Legacy AI Search:Final Evidence 列表(取其中的 `id`)。
*
* 故意不接受裸 `string`:那会被当成字符序列迭代,静默退化成单字符白名单。
*/
export type CitationAllowList =
| ReadonlyArray<string | Pick<AiSearchFinalEvidence, 'id'>>
| ReadonlySet<string>
/** Do not expose citations that cannot resolve to program-owned Final Evidence. */
export function sanitizeAnswerCitations(
answer: string,
evidence: Array<Pick<AiSearchFinalEvidence, 'id'>>
allowed: CitationAllowList
): CitationValidationResult {
const allowed = new Set(evidence.map((item) => item.id))
const allowedIds = new Set<string>()
for (const entry of allowed) {
if (typeof entry === 'string') allowedIds.add(entry)
else if (entry && typeof entry.id === 'string') allowedIds.add(entry.id)
}
const invalidCitationIds = new Set<string>()
const sanitized = answer.replace(/\[E(\d+)\]/g, (citation, number: string) => {
const id = `E${number}`
if (allowed.has(id as AiSearchFinalEvidence['id'])) return citation
if (allowedIds.has(id)) return citation
invalidCitationIds.add(id)
return ''
})
+40 -13
View File
@@ -86,7 +86,8 @@ export class AskWechatService {
const diagnostics = this.diagnostics(
{ provider: '', model: '', modelCallCount: 0, toolCallCount: 0, traces: [] },
startedAt,
'runtime_error'
'runtime_error',
{ question }
)
this.writeLog('error', `Query Agent Runtime 异常(${this.options.entry})`, diagnostics)
return this.fallback(request, 'runtime_error', diagnostics)
@@ -97,13 +98,13 @@ export class AskWechatService {
engine: 'query-agent',
status: 'error',
message: EMPTY_QUESTION_MESSAGE,
diagnostics: this.diagnostics(result, startedAt, 'invalid_question')
diagnostics: this.diagnostics(result, startedAt, 'invalid_question', { question })
}
}
if (result.errorKind === 'provider_unavailable' || result.errorKind === 'provider_failure') {
const outcome: AskWechatOutcome = result.errorKind
const diagnostics = this.diagnostics(result, startedAt, outcome)
const diagnostics = this.diagnostics(result, startedAt, outcome, { question })
this.writeLog('warn', `查询 Provider 不可用(${this.options.entry})`, diagnostics)
return {
engine: 'query-agent',
@@ -114,13 +115,13 @@ export class AskWechatService {
}
if (result.errorKind === 'tool_limit') {
const diagnostics = this.diagnostics(result, startedAt, 'tool_limit')
const diagnostics = this.diagnostics(result, startedAt, 'tool_limit', { question })
this.writeLog('warn', `查询超出工具调用上限(${this.options.entry})`, diagnostics)
return this.fallback(request, 'runtime_error', diagnostics)
}
if (!result.answer?.trim()) {
const diagnostics = this.diagnostics(result, startedAt, 'runtime_error')
const diagnostics = this.diagnostics(result, startedAt, 'runtime_error', { question })
this.writeLog('warn', `Query Agent 未返回回答(${this.options.entry})`, diagnostics)
return this.fallback(request, 'runtime_error', diagnostics)
}
@@ -128,14 +129,19 @@ export class AskWechatService {
const answer = result.answer.trim()
// 澄清回答也记录:下一句("是 BOBO")需要接得上上文。
this.memory.record(conversationKey, question, answer)
const diagnostics = this.diagnostics(result, startedAt, 'answered')
const diagnostics = this.diagnostics(result, startedAt, 'answered', { question, answer })
this.writeLog('info', `Query Agent 回答完成(${this.options.entry})`, diagnostics)
return {
engine: 'query-agent',
status: 'answered',
answer,
// 直接透传 Runtime 收集的真实证据:UI 不允许从 answer 文本反解析。
// citationId 由 Runtime 分配,Adapter 不改写、不重编号。
evidence: (result.evidence || []) as AskWechatEvidenceItem[],
// Host 侧 citation 校验中被移除的非法编号(非空 = 模型引用过不存在的 E#)。
...(result.invalidCitationIds?.length
? { invalidCitationIds: result.invalidCitationIds }
: {}),
stats: buildAskWechatStats(result, request.scope, startedAt),
diagnostics
}
@@ -167,7 +173,12 @@ export class AskWechatService {
requestId: request.requestId,
text: request.text
})
this.writeLog('warn', `Query Agent 失败后回退 Legacy(${this.options.entry})`, diagnostics, reason)
this.writeLog(
'warn',
`Query Agent 失败后回退 Legacy(${this.options.entry})`,
diagnostics,
reason
)
return { engine: 'legacy', status: 'legacy', reason, result: legacyResult }
} catch {
this.writeLog('error', `Legacy fallback 也失败(${this.options.entry})`, diagnostics, reason)
@@ -187,17 +198,35 @@ export class AskWechatService {
> &
Partial<Pick<QueryAgentResult, 'totalMs'>>,
startedAt: number,
outcome: AskWechatOutcome
outcome: AskWechatOutcome,
/**
* 问答原文(可选)。只在本地应用日志里用,不上传、不进遥测。
*
* 排查这类"同一问题时对时错"的故障,光有工具名与次数是不够的 ——
* 必须能对着"问题 + 模型回答"回放,否则无法判断是理解错了、链路断了,还是索引没建。
*/
content?: { question?: string; answer?: string }
): QueryAgentDiagnostics {
const traces = result.traces || []
// 图片 OCR 的两条结构化事实:不回读正文,只统计"取到了几条"与"当时覆盖度是多少"。
const imageOcrTextCount = traces.reduce((sum, trace) => sum + (trace.imageOcrTextCount || 0), 0)
const coverageState = traces
.map((trace) => trace.imageOcrCoverageState)
.filter((value): value is string => typeof value === 'string')
.at(-1)
return {
entry: this.options.entry,
provider: result.provider,
model: result.model,
modelCallCount: result.modelCallCount,
toolCallCount: result.toolCallCount,
tools: (result.traces || []).map((trace) => trace.toolName),
tools: traces.map((trace) => trace.toolName),
totalMs: result.totalMs || Date.now() - startedAt,
outcome
outcome,
...(imageOcrTextCount > 0 ? { imageOcrTextCount } : {}),
...(coverageState ? { imageOcrCoverageState: coverageState } : {}),
...(content?.question ? { question: content.question } : {}),
...(content?.answer ? { answer: content.answer } : {})
}
}
@@ -245,9 +274,7 @@ export function buildAskWechatStats(
const totalMs = result.totalMs || Date.now() - startedAt
// 真实拆解:模型总耗时直接来自每次模型调用的测量;本地查询 = 所有 Tool 的 durationMs 之和。
// 两者不互相推算(用 total - model 反推会把"框架开销"混进"本地查询",那是另一种谎)。
const modelDurationsMs = (result.modelDurationsMs || []).filter((value) =>
Number.isFinite(value)
)
const modelDurationsMs = (result.modelDurationsMs || []).filter((value) => Number.isFinite(value))
const toolDurationsMs = (result.traces || [])
.map((trace) => trace.durationMs)
.filter((value) => Number.isFinite(value))
+16
View File
@@ -8,9 +8,12 @@ export type { CacheClearScope } from '../../shared/cache'
const BOOTSTRAP_CACHE_DIR = path.join(app.getPath('userData'), 'cache', 'bootstrap')
const KNOWLEDGE_CACHE_DIR = path.join(app.getPath('userData'), 'knowledge')
const IMAGE_TEXT_INDEX_CACHE_DIR = path.join(app.getPath('userData'), 'image-text-index')
export interface CacheClearOptions {
beforeClearKnowledge?: () => Promise<void>
/** 清理图片文字索引前调用:停任务 + 关闭派生库句柄。 */
beforeClearImageTextIndex?: () => Promise<void>
}
function inspectDirectory(directory: string): { sizeBytes: number; fileCount: number } {
@@ -46,6 +49,7 @@ export function getCacheSummary(): CacheSummary {
const bootstrap = inspectDirectory(BOOTSTRAP_CACHE_DIR)
const electron = inspectDirectory(path.join(app.getPath('userData'), 'Cache'))
const knowledge = inspectDirectory(KNOWLEDGE_CACHE_DIR)
const imageTextIndex = inspectDirectory(IMAGE_TEXT_INDEX_CACHE_DIR)
const items: CacheSummaryItem[] = [
{
id: 'bootstrap',
@@ -65,6 +69,13 @@ export function getCacheSummary(): CacheSummary {
description:
'为问问微信建立的所有账号本地检索索引。清理后需手动重新建立,不影响微信原始数据。',
...knowledge
},
{
id: 'image-text-index',
label: '图片文字索引',
description:
'本机从微信图片里识别出的文字及其检索索引。清理后无法搜索图片中的文字,可重新建立;不影响微信原始图片与聊天记录。',
...imageTextIndex
}
]
return {
@@ -89,6 +100,11 @@ export async function clearCache(
await options.beforeClearKnowledge?.()
await fs.remove(KNOWLEDGE_CACHE_DIR)
}
if (scope === 'image-text-index' || scope === 'all') {
// 先停下任务再删库,避免"边写边删"。
await options.beforeClearImageTextIndex?.()
await fs.remove(IMAGE_TEXT_INDEX_CACHE_DIR)
}
return getCacheSummary()
}
+346 -68
View File
@@ -16,6 +16,8 @@ import {
} from '../../shared/windows-runtime'
import { mergeRecallArchiveMessages, recordRecallArchiveMessages } from './recall-archive-service'
import type { ExportImageQuality } from '../../shared/image-quality'
import type { ImageMessageCountProbe } from '../../shared/image-text-index'
import { imageTextWindowToSeconds } from '../../shared/image-text-index'
import { wcdbDebugLog } from '../wcdb-debug'
import {
buildContactSearchIndex,
@@ -23,6 +25,87 @@ import {
type ContactSearchIndex
} from '../../shared/contact-search'
/**
* 谁在读消息。
*
* 只允许下面这几个固定标签 —— 日志里**不能**出现会话 md5 / session id / wxid /
* 群名 / 联系人 / 路径,所以调用方身份只能靠标签表达。
* 落在集合外的调用点一律记 `unknown`。
*/
export type ListMessagesCaller =
| 'image-text-index'
| 'group-monitor'
| 'archive'
| 'knowledge'
| 'unknown'
/** 进程生命周期内单调递增的读取序号,用来把"同一段时间的几次调用"关联起来(不是稳定标识)。 */
let listMessagesRequestSeq = 0
export function nextListMessagesRequestId(): string {
listMessagesRequestSeq += 1
return `request-${listMessagesRequestSeq}`
}
/**
* 一次 `listMessages` 的性能拆解。
*
* 存在的意义:大会话的全量读取会把主进程卡住数秒,而原来只有一行 `totalMs`,
* 无法判断时间花在 **WCDB 查询**、**JS 逐条格式化**,还是 **内容解析**上。
*
* 覆盖面:`totalMs` 是外层入口的整段耗时;`formatMs` 包含 `contentParseMs` 与
* `dateFormatMs`(后两者是它的子集,不可与 `formatMs` 相加)。
*/
export interface ListMessagesPerf {
caller: ListMessagesCaller
requestId: string
/** WCDB 返回的原始行数(异步路径里含原生查询时间)。 */
rawRows: number
formattedRows: number
totalMs: number
/** 取原始行:同步 `getUserMessages` 或 `await getUserMessagesAsync`。 */
rawReadMs: number
/** 逐条构造 `FormattedMessage`(整个 `map`)。 */
formatMs: number
/** └ 其中:日期格式化(`toLocaleString`)。 */
dateFormatMs: number
/** └ 其中:内容解析(`parseMessageContent` / `parseStickerMessageFromRow`)。 */
contentParseMs: number
/** 召回归档合并与排序。 */
sortMs: number
/** `totalMs` 减去上面已计部分。 */
otherMs: number
}
function emptyPerf(caller: ListMessagesCaller, requestId: string): ListMessagesPerf {
return {
caller,
requestId,
rawRows: 0,
formattedRows: 0,
totalMs: 0,
rawReadMs: 0,
formatMs: 0,
dateFormatMs: 0,
contentParseMs: 0,
sortMs: 0,
otherMs: 0
}
}
/**
* 只在**值得看**的时候打一行:大会话、或者总耗时已经明显影响交互。
* 单行、可 grep、无任何会话标识。
*/
function logListMessagesPerf(perf: ListMessagesPerf): void {
perf.otherMs = Math.max(0, perf.totalMs - perf.rawReadMs - perf.formatMs - perf.sortMs)
const noteworthy = perf.formattedRows >= 20_000 || perf.totalMs >= 1_000
if (!noteworthy) return
console.log(
`[ChatServicePerf] caller=${perf.caller} request=${perf.requestId} rows=${perf.formattedRows} rawRows=${perf.rawRows} totalMs=${perf.totalMs} rawReadMs=${perf.rawReadMs} formatMs=${perf.formatMs} dateFormatMs=${perf.dateFormatMs} contentParseMs=${perf.contentParseMs} sortMs=${perf.sortMs} otherMs=${perf.otherMs}`
)
}
export function getCurrentKey(): string {
if (!dbRef) return ''
try {
@@ -97,6 +180,8 @@ export interface GroupSnapshot {
roomId: string
memberCount: number
groupName?: string
/** 只读 roominfo(contact.db / chatroom);字段可能随版本缺失。 */
roomInfo?: GroupRoomInfo
members: {
wxid: string
nickname: string
@@ -107,6 +192,18 @@ export interface GroupSnapshot {
}[]
}
/** 群 roominfo 只读字段(列名兼容多版本;缺失则为 undefined)。 */
export interface GroupRoomInfo {
roomId: string
owner?: string
announcement?: string
announcementEditor?: string
maxMemberCount?: number
chatName?: string
openImAccountType?: string
isOpenIm?: boolean
}
export interface GroupMembershipSnapshot {
roomId: string
memberIds: string[]
@@ -198,7 +295,8 @@ export function isReady(): boolean {
/** Session 行的时间字段可能是秒,也可能是毫秒;1e11 以下按秒换算。 */
function sessionTimeToEpochMs(value: unknown): number | null {
const numeric = typeof value === 'number' ? value : typeof value === 'string' ? Number(value) : NaN
const numeric =
typeof value === 'number' ? value : typeof value === 'string' ? Number(value) : NaN
if (!Number.isFinite(numeric) || numeric <= 0) return null
return Math.round(numeric < 1e11 ? numeric * 1000 : numeric)
}
@@ -378,28 +476,52 @@ function listSourceMessages(
endTime?: number,
options?: { limit?: number },
rawMessagesOverride?: WechatMessage[],
requestId = 'NO-REQUEST'
requestId = 'NO-REQUEST',
perf?: ListMessagesPerf
): FormattedMessage[] {
if (!dbRef) return []
const startedAt = Date.now()
const wcdb4Client = dbRef.getWcdb4Client()
const username = wcdb4Client.getUsernameByMd5(userMd5)
const isGroupChat = Boolean(username?.endsWith('@chatroom'))
wcdbDebugLog(
`[${requestId}] ChatService listSourceMessages start md5=${userMd5} username=${username || ''} start=${startTime || 0} end=${endTime || 0} limit=${options?.limit || 0}`
`[${requestId}] ChatService listSourceMessages start start=${startTime || 0} end=${endTime || 0} limit=${options?.limit || 0} hasOverride=${rawMessagesOverride ? 1 : 0}`
)
const rawReadStartedAt = Date.now()
const rawMessages =
rawMessagesOverride ?? dbRef.getUserMessages(userMd5, startTime, endTime, options)
if (perf) {
perf.rawReadMs += Date.now() - rawReadStartedAt
perf.rawRows += rawMessages.length
}
wcdbDebugLog(
`[${requestId}] ChatService raw snapshot ready raw=${rawMessages.length} cost=${Date.now() - startedAt}ms`
`[${requestId}] ChatService raw snapshot ready raw=${rawMessages.length} cost=${Date.now() - rawReadStartedAt}ms`
)
/** 把"内容解析"单独计时,才能区分"消息多"和"每条都在做解析"。 */
const timedParse = <T>(fn: () => T): T => {
if (!perf) return fn()
const startedAt = Date.now()
try {
return fn()
} finally {
perf.contentParseMs += Date.now() - startedAt
}
}
const formatStartedAt = Date.now()
const formatted = rawMessages.map((msg: WechatMessage) => {
const rawMsgType = parseInt(msg.messageType)
const msgType = normalizeMsgType(msg.messageType)
const createTime = parseInt(msg.msgCreateTime)
const date = new Date(createTime * 1000)
/**
* 逐条 `toLocaleString` 每次都会新建一个 ICU formatter —— 大会话里这是主要成本,
* 所以单独计时,避免它被笼统算进"格式化耗时"。
*/
const dateFormatStartedAt = perf ? Date.now() : 0
const datetimeText = date.toLocaleString('zh-CN', { hour12: false })
if (perf) perf.dateFormatMs += Date.now() - dateFormatStartedAt
const isMine = msg.mesDes !== 1
const localId = parseInt(msg.mesLocalID) || 0
@@ -433,7 +555,7 @@ function listSourceMessages(
/<patinfo\b|<type>\s*62\s*<\/type>/i.test(rawContent) ||
([10000, 10002].includes(msgType) && /拍了拍/i.test(rawContent))
if (isPatMessage) {
const system = parseMessageContent(content, 10000)
const system = timedParse(() => parseMessageContent(content, 10000))
const patContent =
system.type === 'system'
? { ...system, pat: true }
@@ -460,9 +582,9 @@ function listSourceMessages(
/<(?:emoji|sticker|emoticon)\b/i.test(content) || /<type>\s*47\s*<\/type>/i.test(content)
const rowSticker =
inferredMsgType === 47 || (inferredMsgType === 49 && !isQuotePayload && hasStickerPayload)
? parseStickerMessageFromRow(msg, content)
? timedParse(() => parseStickerMessageFromRow(msg, content))
: undefined
const parsedContent = parseMessageContent(content, inferredMsgType)
const parsedContent = timedParse(() => parseMessageContent(content, inferredMsgType))
const rowStickerUrl = rowSticker?.type === 'sticker' ? String(rowSticker.url || '') : ''
const parsedShareUrl = parsedContent.type === 'share' ? parsedContent.url : ''
const redPacketUrl = rowStickerUrl || parsedShareUrl
@@ -534,7 +656,7 @@ function listSourceMessages(
}
if (!contentData && typeof content === 'string' && /^[0-9a-fA-F]{64,}$/.test(content.trim())) {
const parsed = parseStickerMessageFromRow(msg, content)
const parsed = timedParse(() => parseStickerMessageFromRow(msg, content))
if (parsed.type === 'sticker') {
if (!parsed.url && parsed.md5) {
parsed.url = wcdb4Client.resolveEmoticonCdnUrl(parsed.md5)
@@ -561,6 +683,9 @@ function listSourceMessages(
: msg.mesLocalID || Math.random().toString()
)
const imageContent = contentData?.type === 'image' ? contentData : undefined
// 语音时长来自 message_content 的 <voicemsg voicelength>(毫秒)——注意不是 length,
// 那是 SILK 数据字节数。已在 parseMessageContent 里换算成秒。
const voiceDuration = contentData?.type === 'voice' ? contentData.duration : undefined
// Local ids repeat across conversations. Scope media handles to this database
// connection and image without changing the message id used by other clients.
const mediaId = imageContent
@@ -615,7 +740,7 @@ function listSourceMessages(
from: contentData?.type === 'system' ? 'system' : isMine ? 'assistant' : 'user',
isSender: isMine,
type: displayType,
datetime: date.toLocaleString('zh-CN', { hour12: false }),
datetime: datetimeText,
content,
img,
name,
@@ -629,13 +754,15 @@ function listSourceMessages(
createTime,
recoveredFromRecallJournal,
contentData,
media
media,
voiceDuration
}
})
console.log(
`[ChatService] listMessages end md5=${userMd5} formatted=${formatted.length} cost=${Date.now() - startedAt}ms`
)
if (perf) {
perf.formatMs += Date.now() - formatStartedAt
perf.formattedRows += formatted.length
}
return formatted
}
@@ -643,13 +770,38 @@ export function listMessages(
userMd5: string,
startTime?: number,
endTime?: number,
options?: { limit?: number }
options?: { limit?: number },
caller: ListMessagesCaller = 'unknown'
): FormattedMessage[] {
const sourceMessages = listSourceMessages(userMd5, startTime, endTime, options)
if (!dbRef) return sourceMessages
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
recordRecallArchiveMessages(userMd5, username, sourceMessages)
return mergeRecallArchiveMessages(userMd5, sourceMessages, startTime, endTime, options?.limit)
const perf = emptyPerf(caller, nextListMessagesRequestId())
const totalStartedAt = Date.now()
try {
const sourceMessages = listSourceMessages(
userMd5,
startTime,
endTime,
options,
undefined,
perf.requestId,
perf
)
if (!dbRef) return sourceMessages
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
const recallStartedAt = Date.now()
recordRecallArchiveMessages(userMd5, username, sourceMessages)
const result = mergeRecallArchiveMessages(
userMd5,
sourceMessages,
startTime,
endTime,
options?.limit
)
perf.sortMs += Date.now() - recallStartedAt
return result
} finally {
perf.totalMs = Date.now() - totalStartedAt
logListMessagesPerf(perf)
}
}
export async function listMessagesAsync(
@@ -657,42 +809,128 @@ export async function listMessagesAsync(
startTime?: number,
endTime?: number,
options?: { limit?: number },
requestId = 'NO-REQUEST'
requestId = '',
caller: ListMessagesCaller = 'unknown'
): Promise<FormattedMessage[]> {
if (!dbRef) return []
const startedAt = Date.now()
wcdbDebugLog(`[${requestId}] ChatService listMessagesAsync start md5=${userMd5}`)
const rawMessages = await dbRef.getUserMessagesAsync(
userMd5,
startTime,
endTime,
options,
requestId
)
wcdbDebugLog(
`[${requestId}] ChatService getUserMessagesAsync end raw=${rawMessages.length} cost=${Date.now() - startedAt}ms`
)
const sourceMessages = listSourceMessages(
userMd5,
startTime,
endTime,
options,
rawMessages,
requestId
)
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
recordRecallArchiveMessages(userMd5, username, sourceMessages)
const result = mergeRecallArchiveMessages(
userMd5,
sourceMessages,
startTime,
endTime,
options?.limit
)
wcdbDebugLog(
`[${requestId}] ChatService listMessagesAsync end formatted=${result.length} cost=${Date.now() - startedAt}ms`
)
return result
const perf = emptyPerf(caller, requestId || nextListMessagesRequestId())
const totalStartedAt = Date.now()
try {
wcdbDebugLog(`[${perf.requestId}] ChatService listMessagesAsync start`)
const rawReadStartedAt = Date.now()
const rawMessages = await dbRef.getUserMessagesAsync(
userMd5,
startTime,
endTime,
options,
perf.requestId
)
perf.rawReadMs += Date.now() - rawReadStartedAt
wcdbDebugLog(
`[${perf.requestId}] ChatService getUserMessagesAsync end raw=${rawMessages.length} cost=${Date.now() - rawReadStartedAt}ms`
)
const sourceMessages = listSourceMessages(
userMd5,
startTime,
endTime,
options,
rawMessages,
perf.requestId,
perf
)
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
const recallStartedAt = Date.now()
recordRecallArchiveMessages(userMd5, username, sourceMessages)
const result = mergeRecallArchiveMessages(
userMd5,
sourceMessages,
startTime,
endTime,
options?.limit
)
perf.sortMs += Date.now() - recallStartedAt
wcdbDebugLog(
`[${perf.requestId}] ChatService listMessagesAsync end formatted=${result.length} cost=${Date.now() - totalStartedAt}ms`
)
return result
} finally {
perf.totalMs = Date.now() - totalStartedAt
logListMessagesPerf(perf)
}
}
/**
* 只取**图片消息**(图片文字索引专用)。
*
* 与 `listMessagesAsync` 的唯一差别是"读哪些行":由 WCDB 在 SQL 层按消息类型过滤,
* 而不是把整个会话读进来再在 JS 里筛。格式化和消息身份走的是**同一套代码**
* (同一个 `listSourceMessages`),所以 `messageId` / `contentData` / 派生键完全不变。
*
* 存在的理由:大会话(十几万到二十几万条消息)全量读一次要 15s 以上,
* 而图片索引只关心图片;这是数据边界错了,不是性能调优问题。
*/
export async function listImageMessagesAsync(
userMd5: string,
window: {
/** 闭下界(epoch ms)。 */
sinceMs?: number
/** 开上界(epoch ms)。 */
beforeMs?: number
limit?: number
} = {},
requestId = '',
caller: ListMessagesCaller = 'unknown'
): Promise<FormattedMessage[]> {
if (!dbRef) return []
const perf = emptyPerf(caller, requestId || nextListMessagesRequestId())
const totalStartedAt = Date.now()
// ms 半开区间 → 秒闭区间。换算只有共享契约里那一处实现。
const { sinceSec, beforeSecInclusive } = imageTextWindowToSeconds(window)
const startTime = sinceSec ?? undefined
const endTime = beforeSecInclusive ?? undefined
try {
const rawReadStartedAt = Date.now()
const rawMessages = await dbRef.getWcdb4Client().listImageMessagesAsync(userMd5, {
...(window.sinceMs !== undefined ? { sinceMs: window.sinceMs } : {}),
...(window.beforeMs !== undefined ? { beforeMs: window.beforeMs } : {}),
...(window.limit !== undefined ? { limit: window.limit } : {}),
// recent-first:同一时间窗内**新的图片先处理**。
order: 'desc',
requestId: perf.requestId
})
perf.rawReadMs += Date.now() - rawReadStartedAt
/**
* 时间边界必须同时交给格式化与召回归档合并。
*
* 少了这一步,归档合并会把**窗口之外**的撤回图片补回来 —— 于是"最近 7 天"
* 这一段会混进十年前的消息,分段窗口形同虚设。
*/
const sourceMessages = listSourceMessages(
userMd5,
startTime,
endTime,
window.limit !== undefined ? { limit: window.limit } : undefined,
rawMessages,
perf.requestId,
perf
)
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
const recallStartedAt = Date.now()
recordRecallArchiveMessages(userMd5, username, sourceMessages)
// 召回归档里可能还留着已被撤回的图片;跳过合并会漏索引,所以照旧合并。
const result = mergeRecallArchiveMessages(
userMd5,
sourceMessages,
startTime,
endTime,
window.limit
)
perf.sortMs += Date.now() - recallStartedAt
return result
} finally {
perf.totalMs = Date.now() - totalStartedAt
logListMessagesPerf(perf)
}
}
export async function listMessagesForExport(
@@ -701,22 +939,61 @@ export async function listMessagesForExport(
endTime?: number
): Promise<FormattedMessage[]> {
if (!dbRef) return []
const rawMessages = await dbRef.getUserMessagesForExport(userMd5, startTime, endTime)
const sourceMessages = listSourceMessages(userMd5, startTime, endTime, undefined, rawMessages)
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
recordRecallArchiveMessages(userMd5, username, sourceMessages)
const mergedMessages = mergeRecallArchiveMessages(userMd5, sourceMessages, startTime, endTime)
console.log(
`[ChatService] listMessagesForExport end md5=${userMd5} source=${sourceMessages.length} merged=${mergedMessages.length}`
)
return mergedMessages
const perf = emptyPerf('archive', nextListMessagesRequestId())
const totalStartedAt = Date.now()
try {
const rawReadStartedAt = Date.now()
const rawMessages = await dbRef.getUserMessagesForExport(userMd5, startTime, endTime)
perf.rawReadMs += Date.now() - rawReadStartedAt
const sourceMessages = listSourceMessages(
userMd5,
startTime,
endTime,
undefined,
rawMessages,
perf.requestId,
perf
)
const username = dbRef.getWcdb4Client().getUsernameByMd5(userMd5) || ''
const recallStartedAt = Date.now()
recordRecallArchiveMessages(userMd5, username, sourceMessages)
const mergedMessages = mergeRecallArchiveMessages(userMd5, sourceMessages, startTime, endTime)
perf.sortMs += Date.now() - recallStartedAt
return mergedMessages
} finally {
perf.totalMs = Date.now() - totalStartedAt
logListMessagesPerf(perf)
}
}
/**
* Count voice rows without hydrating message content. This is used by the
* batch-selection view, where loading every conversation would make opening
* Settings noticeably slow.
* 图片消息计数(SQL 统计,不解密)。
*
* 返回 `count: null` 表示**统计失败**,不是 0 张。调用方必须区分这两件事 ——
* 否则"数不出来"会被显示成"账号里没有图片",用户会因此放弃建立索引。
*/
export async function countImageMessagesAsync(
userMd5: string,
range?: number | { sinceMs?: number; beforeMs?: number }
): Promise<ImageMessageCountProbe> {
if (!dbRef) return { count: null, typeColumn: null, error: '微信数据库尚未就绪' }
return dbRef.getWcdb4Client().countImageMessagesAsync(userMd5, range)
}
/**
* 图片消息的增量水位(条数 + 最大插入序)。
*
* 增量索引**不能只比条数**:召回一张旧图的同时新增一张新图,条数不变但集合变了。
* 返回 null = 当前数据库不支持该统计(调用方须退化成"每轮重扫",宁可慢也不可漏)。
*/
export async function imageConversationWatermarkAsync(
userMd5: string,
range?: number | { sinceMs?: number; beforeMs?: number }
): Promise<{ count: number; maxLocalId: number } | null> {
if (!dbRef) return null
return dbRef.getWcdb4Client().imageConversationWatermarkAsync(userMd5, range)
}
export async function countVoiceMessagesAsync(
userMd5: string,
startTime?: number,
@@ -744,7 +1021,7 @@ export function getGroupSnapshot(userMd5: string): GroupSnapshot | null {
avatar: member.m_nsHeadImgUrl || ''
}))
return { roomId, memberCount: members.length, members }
return { roomId, memberCount: members.length, members, roomInfo: wcdb4Client.getRoomInfo(roomId) || undefined }
}
export async function getGroupSnapshotAsync(userMd5: string): Promise<GroupSnapshot | null> {
@@ -769,7 +1046,8 @@ export async function getGroupSnapshotAsync(userMd5: string): Promise<GroupSnaps
roomId,
groupName: session?.nickname || undefined,
memberCount: members.length,
members
members,
roomInfo: (await wcdb4Client.getRoomInfoAsync(roomId)) || undefined
}
}
@@ -27,6 +27,8 @@ import {
isFreshImageInsight,
isHotImageCandidate
} from '../../shared/image-insight'
import type { SystemOcrCapability, SystemOcrRequest, SystemOcrResult } from '../../shared/system-ocr'
import { systemOcrService } from './system-ocr-service'
/**
* 单张图片的最小信息(由 renderer 从已加载的 messages 中提取并传入 main)。
@@ -312,6 +314,35 @@ class ImageInsightService {
listBySession(sessionId: string, limit?: number): ImageInsight[] {
return imageInsightsStore.listBySession(sessionId, limit)
}
// ============================================================
// 本地图片文字识别(System OCR)
// ============================================================
//
// 与 Vision 路径的关系:
// ImageInsightService 是统一编排入口,下面挂两条互不干扰的运行时——
// - Vision Model Runtime(AIProviderService,走 AI Provider,可能联网)
// - System OCR Runtime(SystemOcrService,纯本地,不联网)
//
// 边界与约束:
// 1. 本地 OCR 结果属于 **派生内容**,原始消息始终是权威来源;
// 本服务只返回识别文本,不做持久化 —— 落库与 Knowledge 回填在图片文字索引侧。
// 2. 本地 OCR 结果 **不会** 写入 image-insights.json——那是 Vision 结果的缓存,
// 两者的缓存键空间也不同(见 buildSystemOcrCacheKey)。
// 3. 这里不读取也绝不修改 AI Vision Provider / 模型配置。
/** 本机是否支持本地图片文字识别(System OCR,引擎按平台决定)。 */
getSystemOcrCapability(): Promise<SystemOcrCapability> {
return systemOcrService.getCapability()
}
/**
* 只做「把图片里的文字读出来」。不发网络请求,不动 AI Provider 配置。
* 失败不抛,返回带 errorCode 的结果。
*/
extractLocalText(request: SystemOcrRequest): Promise<SystemOcrResult> {
return systemOcrService.recognize(request)
}
}
export const imageInsightService = new ImageInsightService()
File diff suppressed because it is too large Load Diff
+671
View File
@@ -0,0 +1,671 @@
/**
* 图片文字索引的**派生存储**(与 Knowledge 派生库物理分离)。
*
* 为什么单独一个库而不是往 knowledge.sqlite 里加表:
* - 清理语义干净:整个能力 = 三个文件(.sqlite/-wal/-shm),删掉即可,不留残渣。
* - 零迁移风险:不动已发布的 knowledge schema(升级不得破坏既有派生库)。
* - 去重语义天然:artifact 按「图片内容 + OCR 运行时指纹」唯一,binding 承担多来源。
*
* 账号隔离与 Knowledge 一致:路径按 accountId 摘要分目录 + 库内 account_id 自证。
*/
import { createHash } from 'node:crypto'
import { mkdirSync, rmSync, statSync, existsSync } from 'node:fs'
import { dirname, join, resolve } from 'node:path'
import { DatabaseSync } from 'node:sqlite'
import {
IMAGE_TEXT_BACKFILL_TIER_ORDER,
IMAGE_TEXT_INDEX_SCHEMA_VERSION,
type ImageOcrArtifact,
type ImageOcrBinding,
type ImageOcrPersistedState,
type ImageTextBackfillTier,
type ImageTextIndexStorageStats,
type ImageTextTierRunState
} from '../../shared/image-text-index'
const MAX_SAFE_ACCOUNT_SEGMENT = /^[a-f0-9]{32}$/
function sha256Hex(value: string): string {
return createHash('sha256').update(value).digest('hex')
}
/** 派生库目录名不直接暴露 accountId。 */
export function imageTextIndexAccountKey(accountId: string): string {
return sha256Hex(`image-text-index-account-v1:${accountId}`).slice(0, 32)
}
export function getImageTextIndexDatabasePath(databaseRoot: string, accountId: string): string {
const accountKey = imageTextIndexAccountKey(accountId)
if (!MAX_SAFE_ACCOUNT_SEGMENT.test(accountKey)) {
throw new Error('Invalid image text index account key')
}
return join(resolve(databaseRoot), accountKey, 'image-text-index.sqlite')
}
/**
* 精确删除三件套(与 Knowledge 的 removeKnowledgeDatabase 同构)。
*
* **删除后必须回验**:Windows 上只要还有句柄(WAL/SHM 未关干净、别的进程打开了库),
* `rmSync` 可能不报错却没真的删掉 —— 那就是"看似清理成功,实际没删"。
* 这里把没删掉的路径返回给调用方,让上层能如实报告失败,而不是假装成功。
*/
export function removeImageTextIndexDatabase(databasePath: string): {
removed: boolean
leftovers: string[]
} {
const leftovers: string[] = []
for (const suffix of ['', '-wal', '-shm']) {
const target = `${databasePath}${suffix}`
if (!existsSync(target)) continue
try {
rmSync(target, { force: true })
} catch {
// 删除失败(典型原因是文件仍被占用)→ 由下面的回验兜住。
}
if (existsSync(target)) leftovers.push(target)
}
return { removed: leftovers.length === 0, leftovers }
}
function asRows(value: unknown): Record<string, unknown>[] {
return Array.isArray(value) ? (value as Record<string, unknown>[]) : []
}
function artifactFromRow(row: Record<string, unknown>): ImageOcrArtifact {
return {
accountId: String(row.account_id),
artifactKey: String(row.artifact_key),
imageIdentity: String(row.image_identity),
state: String(row.state) as ImageOcrPersistedState,
text: String(row.text ?? ''),
charCount: Number(row.char_count ?? 0),
engine: String(row.engine),
platform: String(row.platform),
runtimeVersion: row.runtime_version ? String(row.runtime_version) : null,
language: row.language ? String(row.language) : null,
...(row.error_code ? { errorCode: String(row.error_code) } : {}),
createdAt: Number(row.created_at),
updatedAt: Number(row.updated_at)
}
}
/** 会话内「消息 → OCR 文本」,供 Knowledge 索引时解析(与语音 resolver 同构)。 */
export interface ConversationImageOcrEntry {
state: ImageOcrPersistedState
text: string
}
export class ImageTextIndexStore {
private readonly database: DatabaseSync
constructor(
private readonly databasePath: string,
private readonly accountId: string
) {
mkdirSync(dirname(databasePath), { recursive: true })
this.database = new DatabaseSync(databasePath)
this.initialize()
}
private initialize(): void {
this.database.exec(`
PRAGMA journal_mode = WAL;
PRAGMA synchronous = NORMAL;
PRAGMA busy_timeout = 5000;
CREATE TABLE IF NOT EXISTS image_ocr_meta (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
) STRICT;
CREATE TABLE IF NOT EXISTS image_ocr_artifacts (
artifact_key TEXT PRIMARY KEY,
account_id TEXT NOT NULL,
image_identity TEXT NOT NULL,
state TEXT NOT NULL,
text TEXT NOT NULL,
char_count INTEGER NOT NULL DEFAULT 0,
engine TEXT NOT NULL,
platform TEXT NOT NULL,
runtime_version TEXT,
language TEXT,
error_code TEXT,
created_at INTEGER NOT NULL,
updated_at INTEGER NOT NULL
) STRICT;
CREATE INDEX IF NOT EXISTS image_ocr_artifacts_identity
ON image_ocr_artifacts (image_identity);
CREATE TABLE IF NOT EXISTS image_ocr_bindings (
conversation_id TEXT NOT NULL,
message_id TEXT NOT NULL,
account_id TEXT NOT NULL,
create_time INTEGER NOT NULL,
sender_id TEXT,
sender_name TEXT,
image_identity TEXT NOT NULL,
artifact_key TEXT NOT NULL,
state TEXT NOT NULL,
updated_at INTEGER NOT NULL,
PRIMARY KEY (conversation_id, message_id)
) STRICT;
CREATE INDEX IF NOT EXISTS image_ocr_bindings_identity
ON image_ocr_bindings (image_identity);
CREATE INDEX IF NOT EXISTS image_ocr_bindings_state
ON image_ocr_bindings (state);
-- 每个会话的扫描 checkpoint:重启后据此跳过已完成的会话。
CREATE TABLE IF NOT EXISTS image_ocr_scan_state (
conversation_id TEXT PRIMARY KEY,
account_id TEXT NOT NULL,
state TEXT NOT NULL,
image_total INTEGER NOT NULL DEFAULT 0,
image_processed INTEGER NOT NULL DEFAULT 0,
image_max_local_id INTEGER NOT NULL DEFAULT 0,
updated_at INTEGER NOT NULL
) STRICT;
`)
// 探测式加列(与 Knowledge 一致):旧库缺列时补上,不做版本号比较。
const bindingColumns = new Set(
asRows(this.database.prepare('PRAGMA table_info(image_ocr_bindings)').all()).map((row) =>
String(row.name)
)
)
if (!bindingColumns.has('artifact_key')) {
this.database.exec(
"ALTER TABLE image_ocr_bindings ADD COLUMN artifact_key TEXT NOT NULL DEFAULT ''"
)
}
const scanColumns = new Set(
asRows(this.database.prepare('PRAGMA table_info(image_ocr_scan_state)').all()).map((row) =>
String(row.name)
)
)
if (!scanColumns.has('image_max_local_id')) {
// 旧库补列后默认 0:等于「水位未知」,下一次 pass 会重扫该会话并写入真实水位。
this.database.exec(
'ALTER TABLE image_ocr_scan_state ADD COLUMN image_max_local_id INTEGER NOT NULL DEFAULT 0'
)
}
const storedAccount = this.readMeta('account_id')
if (storedAccount && storedAccount !== this.accountId) {
throw new Error('Image text index account isolation check failed')
}
if (!storedAccount) this.writeMeta('account_id', this.accountId)
if (!this.readMeta('schema_version')) {
this.writeMeta('schema_version', String(IMAGE_TEXT_INDEX_SCHEMA_VERSION))
}
}
private readMeta(key: string): string | null {
const row = this.database
.prepare('SELECT value FROM image_ocr_meta WHERE key = ?')
.get(key) as Record<string, unknown> | undefined
return row ? String(row.value) : null
}
private writeMeta(key: string, value: string): void {
this.database
.prepare(
'INSERT INTO image_ocr_meta (key, value) VALUES (?, ?) ON CONFLICT(key) DO UPDATE SET value = excluded.value'
)
.run(key, value)
}
// ---------------------------------------------------------------- artifacts
getArtifact(artifactKey: string): ImageOcrArtifact | null {
const row = this.database
.prepare('SELECT * FROM image_ocr_artifacts WHERE artifact_key = ?')
.get(artifactKey) as Record<string, unknown> | undefined
return row ? artifactFromRow(row) : null
}
putArtifact(artifact: ImageOcrArtifact): void {
if (artifact.accountId !== this.accountId) {
throw new Error('Image OCR artifact account does not match database')
}
const existing = this.getArtifact(artifact.artifactKey)
this.database
.prepare(
`INSERT INTO image_ocr_artifacts (
artifact_key, account_id, image_identity, state, text, char_count,
engine, platform, runtime_version, language, error_code, created_at, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(artifact_key) DO UPDATE SET
state = excluded.state,
text = excluded.text,
char_count = excluded.char_count,
runtime_version = excluded.runtime_version,
language = excluded.language,
error_code = excluded.error_code,
updated_at = excluded.updated_at`
)
.run(
artifact.artifactKey,
artifact.accountId,
artifact.imageIdentity,
artifact.state,
artifact.text,
artifact.charCount,
artifact.engine,
artifact.platform,
artifact.runtimeVersion,
artifact.language,
artifact.errorCode ?? null,
existing?.createdAt ?? artifact.createdAt,
artifact.updatedAt
)
}
// ----------------------------------------------------------------- bindings
/** 写入绑定;同一 OCR 结果可被多个会话/消息引用。 */
putBinding(binding: ImageOcrBinding): void {
if (binding.accountId !== this.accountId) {
throw new Error('Image OCR binding account does not match database')
}
this.database
.prepare(
`INSERT INTO image_ocr_bindings (
conversation_id, message_id, account_id, create_time, sender_id, sender_name,
image_identity, artifact_key, state, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(conversation_id, message_id) DO UPDATE SET
create_time = excluded.create_time,
sender_id = excluded.sender_id,
sender_name = excluded.sender_name,
image_identity = excluded.image_identity,
artifact_key = excluded.artifact_key,
state = excluded.state,
updated_at = excluded.updated_at`
)
.run(
binding.conversationId,
binding.messageId,
binding.accountId,
binding.createTime,
binding.senderId ?? null,
binding.senderName ?? null,
binding.imageIdentity,
binding.artifactKey,
binding.state,
binding.updatedAt
)
}
/**
* 某会话的「消息 → OCR 文本」映射。
*
* 与语音的 `withVoiceTranscript` 同构:在**主进程**把派生文本贴到消息上,
* 再交给 Knowledge 索引,派生库不需要被 worker 打开。
*/
getConversationOcr(
conversationId: string
): Map<string, ConversationImageOcrEntry> {
const rows = asRows(
this.database
.prepare(
`SELECT b.message_id AS message_id, b.state AS binding_state,
a.state AS artifact_state, a.text AS text
FROM image_ocr_bindings b
LEFT JOIN image_ocr_artifacts a ON a.artifact_key = b.artifact_key
WHERE b.conversation_id = ?`
)
.all(conversationId)
)
const result = new Map<string, ConversationImageOcrEntry>()
for (const row of rows) {
result.set(String(row.message_id), {
state: String(row.artifact_state || row.binding_state) as ImageOcrPersistedState,
text: String(row.text ?? '')
})
}
return result
}
// -------------------------------------------------------------- checkpoint
/**
* 已持久化的图片消息总数统计。
*
* 必须落盘:派生库只知道自己**处理过**什么,不知道源数据里**一共**有多少图片。
* 一旦把这个 total 只放在内存里,应用重启后 coverage 就会退化成
* 「processed / processed」→ 把 30% 的部分索引谎报成 100% 完整覆盖。
*/
readCountedTotal(): { total: number; countedAt: number; complete: boolean } | null {
const total = this.readMeta('total_image_messages')
const countedAt = this.readMeta('total_image_counted_at')
if (total === null || countedAt === null) return null
const parsedTotal = Number(total)
const parsedCountedAt = Number(countedAt)
if (!Number.isFinite(parsedTotal) || !Number.isFinite(parsedCountedAt)) return null
return {
total: parsedTotal,
countedAt: parsedCountedAt,
// 统计时若有会话没数上(数据库不支持该统计),分母就是偏小的 →
// 绝不能据此声称"已覆盖全部",否则少数的那些会话会被静默算进"已覆盖"。
complete: this.readMeta('total_image_messages_complete') === '1'
}
}
/**
* 扫描进度:所有会话累计「应扫多少张图片消息」与「实际扫过多少张」。
*
* 与上面的 `readCountedTotal()` 分工必须分清:
* - `readCountedTotal()` 是 `countImageMessages()` 给出的**预估**分母(遍历消息表数出来的,
* 会随新消息变动,且与「归档合并后流水线真正拿到的消息集合」并不完全一致);
* - 这里是流水线**真实走过**的集合。
*
* 进度必须用后者。拿预估值当分母,进度会永远差最后几个百分点,
* 让已经跑完的索引一直显示成"部分完成"。
*/
readScanProgress(): { total: number; processed: number } {
const row = this.database
.prepare(
'SELECT SUM(image_total) AS total, SUM(image_processed) AS processed FROM image_ocr_scan_state'
)
.get() as Record<string, unknown> | undefined
return {
total: Number(row?.total ?? 0) || 0,
processed: Number(row?.processed ?? 0) || 0
}
}
writeCountedTotal(input: { total: number; countedAt: number; complete: boolean }): void {
this.writeMeta('total_image_messages', String(input.total))
this.writeMeta('total_image_counted_at', String(input.countedAt))
this.writeMeta('total_image_messages_complete', input.complete ? '1' : '0')
}
// ---------------------------------------------------- recent-first 回填状态
//
// 全部走 `image_ocr_meta`(key/value),**不加表、不加列** —— 这是 old checkpoint
// 兼容性的来源:旧库没有这些 key 时读出来就是"没有计划",于是下一次 pass
// 以当前时刻为锚点重新建计划。已经处理过的图片由 terminal binding 兜住,
// 不会因为"计划是新的"而重新 OCR。
/**
* 读取回填计划状态。
*
* `anchorMs === null` = 从未规划过(新用户,或从旧版本升级且没有这些 key)。
* 调用方此时必须**新建**计划,而不是假设"已完成"。
*/
readBackfillState(): {
anchorMs: number | null
tierStates: Partial<Record<ImageTextBackfillTier, ImageTextTierRunState>>
coveredToMs: number | null
} {
const anchorRaw = this.readMeta('backfill_anchor_ms')
const anchorMs = anchorRaw === null ? null : Number(anchorRaw)
let tierStates: Partial<Record<ImageTextBackfillTier, ImageTextTierRunState>> = {}
const statesRaw = this.readMeta('backfill_tier_states')
if (statesRaw) {
try {
const parsed = JSON.parse(statesRaw) as Record<string, unknown>
for (const tier of IMAGE_TEXT_BACKFILL_TIER_ORDER) {
const value = parsed[tier]
if (value === 'pending' || value === 'running' || value === 'complete') {
tierStates[tier] = value
}
}
} catch {
// 状态串损坏时按"没有分段状态"处理:最坏情况是重扫一遍,
// 而 terminal binding 保证不会重复 OCR。
tierStates = {}
}
}
const coveredRaw = this.readMeta('backfill_covered_to_ms')
const coveredToMs = coveredRaw === null ? null : Number(coveredRaw)
return {
anchorMs: anchorMs !== null && Number.isFinite(anchorMs) ? anchorMs : null,
tierStates,
coveredToMs: coveredToMs !== null && Number.isFinite(coveredToMs) ? coveredToMs : null
}
}
writeBackfillAnchor(anchorMs: number): void {
this.writeMeta('backfill_anchor_ms', String(Math.floor(anchorMs)))
}
/**
* 写入单个分段的运行态。
*
* 一个 pass 是单线程串行推进分段,所以"读-改-写"整体落在一条 meta 行里;
* 这里额外做的只有"保留其它分段"。
*/
writeBackfillTierState(tier: ImageTextBackfillTier, state: ImageTextTierRunState): void {
const current = this.readBackfillState().tierStates
const next = { ...current, [tier]: state }
this.writeMeta('backfill_tier_states', JSON.stringify(next))
}
/** 增量补齐水位:处理到哪儿了(`create_time <= ms` 都已就绪)。只增不减。 */
writeBackfillCoveredToMs(coveredToMs: number): void {
const previous = this.readBackfillState().coveredToMs
if (previous !== null && previous >= coveredToMs) return
this.writeMeta('backfill_covered_to_ms', String(Math.floor(coveredToMs)))
}
readScanState(): Map<
string,
{ state: string; imageTotal: number; processed: number; maxLocalId: number }
> {
const rows = asRows(
this.database
.prepare(
'SELECT conversation_id, state, image_total, image_processed, image_max_local_id FROM image_ocr_scan_state'
)
.all()
)
const map = new Map<
string,
{ state: string; imageTotal: number; processed: number; maxLocalId: number }
>()
for (const row of rows) {
map.set(String(row.conversation_id), {
state: String(row.state),
imageTotal: Number(row.image_total ?? 0),
processed: Number(row.image_processed ?? 0),
maxLocalId: Number(row.image_max_local_id ?? 0)
})
}
return map
}
writeScanState(input: {
conversationId: string
state: 'done' | 'partial'
imageTotal: number
imageProcessed: number
/** 本会话图片消息的最大插入序(增量水位)。 */
maxLocalId: number
}): void {
this.database
.prepare(
`INSERT INTO image_ocr_scan_state (
conversation_id, account_id, state, image_total, image_processed,
image_max_local_id, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(conversation_id) DO UPDATE SET
state = excluded.state,
image_total = excluded.image_total,
image_processed = excluded.image_processed,
image_max_local_id = excluded.image_max_local_id,
updated_at = excluded.updated_at`
)
.run(
input.conversationId,
this.accountId,
input.state,
input.imageTotal,
input.imageProcessed,
input.maxLocalId,
Date.now()
)
}
// ------------------------------------------------------------------ 统计
/**
* 有 OCR 派生绑定的会话集合。
*
* 清理时**必须**先拿到它:OCR 文本早已被灌进 Knowledge 的 chunks / FTS,
* 只删派生病不会让那些派生文字失效 —— 用户仍会从旧索引里搜到图片里的文字。
*/
conversationIdsWithOcr(): string[] {
const rows = asRows(
this.database.prepare('SELECT DISTINCT conversation_id FROM image_ocr_bindings').all()
)
return rows.map((row) => String(row.conversation_id))
}
/**
* **有 OCR 派生文本**(artifact 里 char_count > 0)的会话集合。
*
* 派生索引修复只需要重建它们:只有这些会话的 Knowledge 里"应该"存在图片派生文字。
* 全是 `empty` 的会话本来就没有派生文字可修,重建它们只是白读一遍 WCDB。
*/
conversationIdsWithIndexedOcr(): string[] {
const rows = asRows(
this.database
.prepare(
`SELECT DISTINCT b.conversation_id AS conversation_id
FROM image_ocr_bindings b
JOIN image_ocr_artifacts a ON a.artifact_key = b.artifact_key
WHERE a.char_count > 0`
)
.all()
)
return rows.map((row) => String(row.conversation_id))
}
/**
* 重置指定状态的失败记录,让它们可以被下一轮 pass 重新处理。
*
* 用途:**代码修好后**,把上一次 bug 造成的假失败(例如整批 `decrypt_failed`)
* 变成可重试状态,而不是要求用户删掉整个派生库 —— 那会连已经成功的记录一起丢掉。
*
* 三件事一起做,缺一不可:
* 1. 删掉这些失败绑定;
* 2. 删掉它们所在会话的 checkpoint —— 否则 pass 会以"该会话已完成"直接跳过,
* 表现为"点了重试但什么都没发生";
* 3. 删掉因此变成孤儿的 artifact(**没有任何绑定再引用的**才删,成功记录一条不动)。
*/
resetFailures(states: ImageOcrPersistedState[]): number {
if (!states.length) return 0
const placeholders = states.map(() => '?').join(', ')
const affected = asRows(
this.database
.prepare(
`SELECT DISTINCT conversation_id FROM image_ocr_bindings WHERE state IN (${placeholders})`
)
.all(...states)
).map((row) => String(row.conversation_id))
const artifactKeys = asRows(
this.database
.prepare(
`SELECT DISTINCT artifact_key FROM image_ocr_bindings WHERE state IN (${placeholders})`
)
.all(...states)
).map((row) => String(row.artifact_key))
const info = this.database
.prepare(`DELETE FROM image_ocr_bindings WHERE state IN (${placeholders})`)
.run(...states)
const clearScan = this.database.prepare(
'DELETE FROM image_ocr_scan_state WHERE conversation_id = ?'
)
for (const conversationId of affected) clearScan.run(conversationId)
const dropOrphan = this.database.prepare(
`DELETE FROM image_ocr_artifacts
WHERE artifact_key = ?
AND NOT EXISTS (SELECT 1 FROM image_ocr_bindings b WHERE b.artifact_key = ?)`
)
for (const artifactKey of artifactKeys) dropOrphan.run(artifactKey, artifactKey)
return Number(info.changes ?? 0)
}
/** 按状态聚合绑定数 —— 覆盖度与进度都从这里取,保证与库内真实一致。 */
countByState(): Record<string, number> { const rows = asRows(
this.database
.prepare('SELECT state, COUNT(*) AS total FROM image_ocr_bindings GROUP BY state')
.all()
)
const counts: Record<string, number> = {}
for (const row of rows) counts[String(row.state)] = Number(row.total ?? 0)
return counts
}
storageStats(): ImageTextIndexStorageStats {
const counts = this.countByState()
const textRow = this.database
.prepare(
"SELECT COUNT(*) AS total FROM image_ocr_artifacts WHERE state = 'indexed' AND length(text) > 0"
)
.get() as Record<string, unknown> | undefined
const updatedRow = this.database
.prepare('SELECT MAX(updated_at) AS latest FROM image_ocr_bindings')
.get() as Record<string, unknown> | undefined
let totalBytes = 0
for (const suffix of ['', '-wal', '-shm']) {
try {
totalBytes += statSync(`${this.databasePath}${suffix}`).size
} catch {
// 文件可能尚未创建;忽略。
}
}
const latest = updatedRow?.latest
return {
indexedImages: counts['indexed'] ?? 0,
ocrTextCount: Number(textRow?.total ?? 0),
totalBytes,
updatedAt: latest === null || latest === undefined ? null : Number(latest)
}
}
/** 清空全部派生数据(表级清空;文件级删除由 service 负责)。 */
clearAll(): void {
this.database.exec(`
DELETE FROM image_ocr_bindings;
DELETE FROM image_ocr_artifacts;
DELETE FROM image_ocr_scan_state;
DELETE FROM image_ocr_meta WHERE key IN (
'total_image_messages',
'total_image_counted_at',
'total_image_messages_complete',
'backfill_anchor_ms',
'backfill_tier_states',
'backfill_covered_to_ms'
);
`)
}
/**
* 清空并重置检查点。
*
* 注意:必须同时清 `scan_state`,否则清理后再次索引会因为「会话已完成」
* 而直接跳过 —— UI 会停在「未建立」但实际再也不跑。
*/
clearDerivedData(): void {
this.clearAll()
}
close(): void {
try {
// 先折叠 WAL 再关连接:否则 -wal / -shm 可能仍被持有,
// Windows 上会导致后续 rmSync 静默失败("清理成功"但文件还在)。
this.database.exec('PRAGMA wal_checkpoint(TRUNCATE);')
} catch {
// 库可能已经处于不可写状态;关闭仍然要做。
}
try {
this.database.close()
} catch {
// best effort
}
}
}
+181 -7
View File
@@ -1,5 +1,6 @@
import { listContactsAsync, listMessagesAsync, isReady, type FormattedContact, type FormattedMessage } from './chat-service'
import { resolveContact } from './contact-resolution-service'
import { sourceMessageId } from '../knowledge/message-identity'
import type { KnowledgeSearchService } from '../knowledge/knowledge-search-service'
import { inferAiSearchTimeRange } from '../../shared/ai-search'
import { KNOWLEDGE_FRESHNESS_TOLERANCE_MS } from '../../shared/knowledge'
@@ -13,6 +14,7 @@ import type {
ResolvedCorpusScope,
ResolvedTimeRange,
QueryIndexCoverage,
QueryImageTextCoverage,
QuerySearchTimings,
QueryMessagesRequest,
SearchMessagesRequest,
@@ -26,6 +28,14 @@ import {
decodeMessageRef as fromRef,
normalizeMessageIdentity
} from '../../shared/local-query-api'
import {
IMAGE_TEXT_BACKFILL_TIER_LABEL,
describeImageTextCoverage,
describeImageTextCoveredRanges,
imageTextCoverageState,
type ImageTextIndexCoverage
} from '../../shared/image-text-index'
import { toEvidenceDisplayText } from '../../shared/knowledge'
const LIMIT_MAX = 200
const CONTEXT_MAX = 50
@@ -112,6 +122,59 @@ function buildIndexCoverage(
}
}
/**
* 图片文字索引未完成时,必须附加的零结果诚实性约束。
*
* 图片 OCR 是**独立的**覆盖维度:它可能是"未建立"或"只做了 30%"。
* 此时 0 条图片证据只是**索引缺口**,不是**事实空缺**。
*/
const IMAGE_OCR_ZERO_RESULT_CAUTION =
'涉及图片、截图、海报里的文字的问题,当前不能因为没搜到就回答"没有"。'
/**
* 图片文字索引覆盖度 → 可直接引用的结论句。
*
* 与 `buildIndexCoverage` 同思路:只给结构化数字,模型会自己换算、甚至反过来
* 宣称"覆盖完整"。这里由 Engine 给出结论句,模型只需引用。
*/
export function buildImageOcrCoverage(
coverage: ImageTextIndexCoverage | null
): QueryImageTextCoverage | undefined {
if (!coverage) return undefined
const state = imageTextCoverageState(coverage)
const countedNote = coverage.countedAt
? `(图片数量统计于 ${formatLocalMinute(coverage.countedAt)})`
: ''
const base = describeImageTextCoverage(coverage)
/**
* recent-first 之后必须把"哪段时间能下确定性结论"一起给出。
*
* 否则模型只看一个总百分比:30% 时它会以为连最近一周都不可信(过度保守没坏处),
* 但 99% 时它会以为"去年也能放心下结论"(这就把索引缺口说成了事实空缺)。
*/
const tiers = coverage.tiers ?? []
const coveredRanges = tiers
.filter((entry) => entry.state === 'complete')
.map((entry) => IMAGE_TEXT_BACKFILL_TIER_LABEL[entry.tier])
const rangeNote = tiers.length ? describeImageTextCoveredRanges(coverage) : ''
return {
state,
totalImageMessages: coverage.totalImageMessages,
processed: coverage.processed,
indexed: coverage.indexed,
empty: coverage.empty,
missing: coverage.missing,
failed: coverage.failed,
pending: coverage.pending,
...(coverage.countedAt ? { countedAtLabel: formatLocalMinute(coverage.countedAt) } : {}),
...(coveredRanges.length ? { coveredRanges } : {}),
summary:
state === 'complete'
? `${base}${countedNote}`
: `${base}${countedNote}${rangeNote}${IMAGE_OCR_ZERO_RESULT_CAUTION}`
}
}
const KIND_LABELS: Record<QueryMessageType, string> = {
text: '文本',
image: '图片',
@@ -149,7 +212,19 @@ function resolvedTimeRange(input: QueryTimeRange, now = new Date()): ResolvedTim
const range = inferAiSearchTimeRange(phrase[input.kind], map[input.kind], now)
return { kind: input.kind, startTime: range.startTime, endTime: range.endTime, label: range.label }
}
function toQueryMessage(conversationId: string, message: FormattedMessage, target: FormattedContact): QueryMessage {
/**
* 一条消息的展示形态。
*
* `imageOcr` 是可选的**派生文本**(来自本地图片文字索引,只读、不触发 OCR)。
* 图片消息的正文永远是空的 —— 识别出的文字必须走独立字段,
* 否则"图片里的文字"会被伪装成"群友发的文字消息"。
*/
function toQueryMessage(
conversationId: string,
message: FormattedMessage,
target: FormattedContact,
imageOcr?: { state: string; text: string }
): QueryMessage {
const kind = kindOf(message)
const content = message.contentData
const attachment =
@@ -163,7 +238,16 @@ function toQueryMessage(conversationId: string, message: FormattedMessage, targe
? { kind: 'file' as const, name: message.exportMediaName || (content?.type === 'share' ? content.title : undefined), url: content?.type === 'share' ? content.url : undefined }
: undefined
const text = message.content?.trim() || message.voiceTranscript?.trim() || undefined
return { messageRef: toRef(conversationId, message.id), timestamp: (message.createTime || 0) * 1000, datetime: message.datetime, sender: message.isSender ? '我' : (message.name || target.m_nsNickName), direction: message.isSender ? 'to_target' : 'from_target', messageType: kind, sourceKind: kind, ...(attachment ? { attachment } : {}), ...(text ? { text } : {}) }
const derived = kind === 'image' ? imageOcr : undefined
const ocrText = derived && derived.state === 'indexed' ? derived.text.trim() : ''
const imageTextState: QueryMessage['imageTextState'] = !derived
? 'not_indexed'
: ocrText
? 'indexed'
: derived.state === 'empty'
? 'empty'
: 'not_indexed'
return { messageRef: toRef(conversationId, message.id), timestamp: (message.createTime || 0) * 1000, datetime: message.datetime, sender: message.isSender ? '我' : (message.name || target.m_nsNickName), direction: message.isSender ? 'to_target' : 'from_target', messageType: kind, sourceKind: kind, ...(attachment ? { attachment } : {}), ...(text ? { text } : {}), ...(ocrText ? { imageOcrText: ocrText, derivedSource: 'image_ocr' as const } : {}), ...(kind === 'image' ? { imageTextState } : {}) }
}
/**
@@ -216,7 +300,61 @@ function outsideScopeError(scope: ResolvedCorpusScope, actual: string): { status
}
export class LocalQueryApiService {
constructor(private readonly knowledge?: KnowledgeSearchService, private readonly nowProvider: () => Date = () => new Date()) {}
/**
* 图片文字索引覆盖度提供者(同步、只读)。
*
* 刻意不在 Knowledge worker 里算:OCR 派生库(`image-text-index.sqlite`)与
* `knowledge.sqlite` 物理分离,worker 不该为了一个覆盖度数字去开它。
*/
private imageTextCoverage: () => ImageTextIndexCoverage | null = () => null
/**
* 单条图片消息的 OCR 派生文本提供者(同步、只读)。
*
* L4(查询层)**只读** L1(OCR artifact)—— 这里绝不允许触发 OCR、解密或读原图。
* 之前 `query_messages` 缺这一环,导致"图片已经识别出文字"这件事在精确读消息
* 这条路径上完全不可见:模型只拿到一个空的 `attachment`,于是把"索引缺口"
* 说成"图片里没有文字",甚至反过来建议用户去建立已经建好的索引。
*/
private imageOcrEntry:
| ((conversationId: string, messageId: string) => { state: string; text: string } | undefined)
| undefined
/** 收藏关键词检索(favorite.db,只读)。 */
private favoritesSearch:
| ((query: string, limit: number) => Promise<Array<{ text: string; timestamp?: number }>>)
| undefined
constructor(
private readonly knowledge?: KnowledgeSearchService,
private readonly nowProvider: () => Date = () => new Date()
) {}
/** 注入收藏检索(`FavoritesService.searchHits`),供 `search_messages` 合并只读命中。 */
setFavoritesSearchProvider(
provider: (query: string, limit: number) => Promise<Array<{ text: string; timestamp?: number }>>
): void {
this.favoritesSearch = provider
}
/**
* 注入图片文字索引覆盖度提供者。
*
* 用 setter 而不是构造参数:避免给第二个带默认值的参数写 `undefined` 占位,
* 也让测试可以直接注入假的覆盖度。
*/
setImageTextCoverageProvider(provider: () => ImageTextIndexCoverage | null): void {
this.imageTextCoverage = provider
}
/** 注入单条图片消息的 OCR 派生文本解析器(只读;见 `imageOcrEntry` 的约束)。 */
setImageOcrEntryProvider(
provider:
| ((conversationId: string, messageId: string) => { state: string; text: string } | undefined)
| undefined
): void {
this.imageOcrEntry = provider
}
capabilities(): QueryCapabilitiesResponse {
return { version: 1, tools: { query_messages: { operation: '读取指定联系人的确定性消息', directions: ['any', 'from_target', 'to_target'], messageTypes: kinds, timeRanges: ['all', 'today', 'yesterday', 'this_week', 'last_7_days', 'this_month', 'previous_month', 'this_year', 'previous_year', 'absolute'], limitMax: LIMIT_MAX }, search_messages: { operation: '受限 Knowledge 关键词检索', timeRanges: ['all', 'today', 'yesterday', 'this_week', 'last_7_days', 'this_month', 'previous_month', 'this_year', 'previous_year', 'absolute'], limitMax: LIMIT_MAX }, message_context: { operation: '读取消息前后文', timeRanges: ['all'], limitMax: CONTEXT_MAX }, conversation_overview: { operation: '按会话时间片提取概览证据', timeRanges: ['all', 'today', 'yesterday', 'this_week', 'last_7_days', 'this_month', 'previous_month', 'this_year', 'previous_year', 'absolute'], limitMax: LIMIT_MAX } } }
}
@@ -282,7 +420,18 @@ export class LocalQueryApiService {
const raw = await listMessagesAsync(contact.md5, range.startTime, range.endTime)
const direction = request.direction || 'any'; const allowed = new Set(request.messageTypes || kinds)
const filtered = raw.filter((message) => !(request.excludeSystem !== false && kindOf(message) === 'system')).filter((message) => allowed.has(kindOf(message))).filter((message) => direction === 'any' || (direction === 'to_target' ? message.isSender : !message.isSender)).sort((a, b) => ((a.createTime || 0) - (b.createTime || 0)) * ((request.order || 'asc') === 'asc' ? 1 : -1)).slice(0, Math.min(LIMIT_MAX, Math.max(1, request.limit || 20)))
return { status: 'completed' as const, target: contactView(contact), query: { direction, messageTypes: request.messageTypes || [], order: request.order || 'asc', limit: Math.min(LIMIT_MAX, Math.max(1, request.limit || 20)), excludeSystem: request.excludeSystem !== false, resolvedTimeRange: range }, coverage: { state: 'complete' as const }, returnedCount: filtered.length, messages: filtered.map((message) => toQueryMessage(contact.md5, message, contact)), scope: corpus.scope }
// 图片 OCR 派生文本:L4 只读 L1,**不触发 OCR / 解密 / 读原图**。
// 键必须用 `sourceMessageId(message)`(binding 主键就是它),不能用裸 `message.id`,
// 否则 `local:` 前缀会让查表静默失配 —— 与 Knowledge 用同一条身份规则。
const messages = filtered.map((message) => {
const ocr =
kindOf(message) === 'image'
? this.imageOcrEntry?.(contact.md5, sourceMessageId(message))
: undefined
return toQueryMessage(contact.md5, message, contact, ocr)
})
const imageOcrCoverage = buildImageOcrCoverage(this.imageTextCoverage())
return { status: 'completed' as const, target: contactView(contact), query: { direction, messageTypes: request.messageTypes || [], order: request.order || 'asc', limit: Math.min(LIMIT_MAX, Math.max(1, request.limit || 20)), excludeSystem: request.excludeSystem !== false, resolvedTimeRange: range }, coverage: { state: 'complete' as const }, returnedCount: filtered.length, messages, scope: corpus.scope, ...(imageOcrCoverage ? { imageOcrCoverage } : {}) }
}
async search(request: SearchMessagesRequest) {
const requestStartedAt = Date.now()
@@ -359,6 +508,8 @@ export class LocalQueryApiService {
const covered = indexCovers(found.indexLatestAt, requestedEnd, found.sourceLatestAt)
const indexCoverage = buildIndexCoverage(found.indexLatestAt, found.sourceLatestAt, covered)
// 图片文字索引是**独立覆盖维度**:文字索引再完整也不代表图片里的文字搜得到。
const imageOcrCoverage = buildImageOcrCoverage(this.imageTextCoverage())
const timings: QuerySearchTimings = {
totalMs: Date.now() - requestStartedAt,
scopeMs,
@@ -368,19 +519,36 @@ export class LocalQueryApiService {
enrichmentMs,
...(knowledgeTiming ? { knowledge: knowledgeTiming } : {})
}
let favoriteEvidence: QueryEvidenceItem[] = []
if (this.favoritesSearch && !targetView) {
try {
const hits = await this.favoritesSearch(request.query, Math.min(10, limit))
favoriteEvidence = hits.map((hit, index) => ({
messageRef: `fav-${index}`,
conversationName: '收藏',
timestamp: hit.timestamp ?? 0,
sender: '收藏',
sourceKind: 'other' as QueryMessageType,
text: hit.text
}))
} catch {
favoriteEvidence = []
}
}
return {
status: 'completed' as const,
...(targetView ? { target: targetView } : {}),
resolvedTimeRange: range,
coverage: { state: this.searchCoverage(found, requestedEnd) },
probeCount: probes.length,
evidenceCount: found.evidence.size,
evidence: Array.from(found.evidence.values()).slice(0, limit),
evidenceCount: found.evidence.size + favoriteEvidence.length,
evidence: [...Array.from(found.evidence.values()).slice(0, limit), ...favoriteEvidence],
scope: corpus.scope,
indexLatestAt: found.indexLatestAt,
sourceLatestAt: found.sourceLatestAt,
freshness: { catchUp: freshness.catchUp },
...(indexCoverage ? { indexCoverage } : {}),
...(imageOcrCoverage ? { imageOcrCoverage } : {}),
timings
}
}
@@ -454,7 +622,11 @@ export class LocalQueryApiService {
timestamp: item.timestamp,
sender: item.sender,
sourceKind: item.sourceKind,
text: item.text,
// 兜底再剥一次:不管 Knowledge 侧哪条检索路径产出的文本,
// 面向用户与模型的都不允许出现 `图片文字:` 这类引擎内部标签。
text: toEvidenceDisplayText(item.text),
...(item.derivedSource ? { derivedSource: item.derivedSource } : {}),
...(item.imageOcrText ? { imageOcrText: item.imageOcrText } : {}),
conversationName: owner ? contactView(owner).displayName : undefined,
conversationType: owner?.type
} satisfies QueryEvidenceItem
@@ -592,6 +764,7 @@ export class LocalQueryApiService {
const truncated = raw.length > OVERVIEW_SOURCE_CAP
const evidence = selectTemporalCoverageEvidence(contact, messages, OVERVIEW_EVIDENCE_TARGET)
const state: 'complete' | 'partial' = truncated ? 'partial' : 'complete'
const imageOcrCoverage = buildImageOcrCoverage(this.imageTextCoverage())
return {
status: 'completed' as const,
target: contactView(contact),
@@ -603,6 +776,7 @@ export class LocalQueryApiService {
selection: { mode: 'temporal_coverage' as const, selectedEvidenceCount: evidence.length, sampled: truncated || evidence.length < messages.length },
evidence,
scope: corpus.scope,
...(imageOcrCoverage ? { imageOcrCoverage } : {}),
origin: 'wcdb' as const
}
}
+49
View File
@@ -0,0 +1,49 @@
import { createHash } from 'node:crypto'
/**
* 统一日志脱敏。
*
* 微信连接器与发送日志共用同一套规则:任何 secret 都不能以原值进入普通日志,
* 包括 bot_token、context_token、typing_ticket、AES key、二维码凭据与 Authorization 头。
*
* 只用于自由文本日志;结构化字段(如 sha256 摘要)不要经过这里,否则会被误伤。
*/
const REDACTION_RULES: ReadonlyArray<readonly [RegExp, string]> = [
// Authorization 头
[/Bearer\s+[A-Za-z0-9._~+/-]+=*/gi, 'Bearer [已隐藏]'],
// 二维码图片内容(base64 data URL 可能长达数百 KB)
[/data:image\/[^;]+;base64,[A-Za-z0-9+/=]+/gi, 'data:image/[二维码已隐藏]'],
// JSON / query 形态的 secret 字段
[
/("?(?:bot_token|context_token|typing_ticket|aeskey|aes_key|access_token|refresh_token|authorization)"?\s*[:=]\s*)"?[^\s",}&]+"?/gi,
'$1"[已隐藏]"'
],
// 兼容历史写法:token=xxx / token: xxx
[/(\btoken\s*[:=]\s*)[^\s,}]+/gi, '$1[已隐藏]'],
// iLink 上传参数
[/(encrypted_query_param=)[^\s&"]+/gi, '$1[已隐藏]']
]
/** 对任意自由文本日志做脱敏。 */
export function redactSecrets(message: string): string {
let result = String(message ?? '')
for (const [pattern, replacement] of REDACTION_RULES) {
result = result.replace(pattern, replacement)
}
return result
}
/** 日志中只允许出现 secret 是否存在,不允许出现原值。 */
export function describeSecretPresence(value: string | undefined | null): string {
return value ? 'present' : 'absent'
}
/**
* 生成不可逆的短指纹,用于需要关联同一 secret 又不能落盘的诊断场景。
* 只保留 sha256 前 8 位十六进制,无法反推原值。
*/
export function secretFingerprint(value: string | undefined | null): string {
if (!value) return 'absent'
return createHash('sha256').update(String(value)).digest('hex').slice(0, 8)
}
File diff suppressed because it is too large Load Diff
+73 -27
View File
@@ -46,6 +46,9 @@ const EXECUTIONS_FILE = 'executions.json'
const NOTIFICATIONS_FILE = 'notifications.json'
const SETTINGS_FILE = 'settings.json'
const TICK_MS = 15_000
const AUTO_RESEND_MAX_ATTEMPTS = 6
const AUTO_RESEND_BASE_BACKOFF_MS = 5 * 60_000
const AUTO_RESEND_MAX_BACKOFF_MS = 60 * 60_000
const NOTIFICATION_TEST_MESSAGE = `✅ TraceMemo 定时日报通知已开启
以后定时日报生成或发送出现异常时,
@@ -77,28 +80,6 @@ export interface ScheduledReportSendActionInput {
taskId?: string
}
function buildScheduledReportActionRequest(
input: ScheduledReportSendActionInput
): WechatActionRequest {
return {
idempotencyKey:
input.retryCount && input.retryCount > 0
? `scheduled_report:${input.executionId}:retry:${input.retryCount}`
: `scheduled_report:${input.executionId}`,
origin: 'scheduled_report',
purpose: 'scheduled_report',
triggerType: input.triggerType === 'scheduled' ? 'automation' : 'user',
sourceId: input.executionId,
executionId: input.executionId,
recipient: { type: 'group', id: input.target },
content: { type: 'image', path: input.filePath },
metadata: {
taskId: input.taskId,
retryCount: input.retryCount || 0
}
}
}
const defaultDependencies = (): ScheduledReportDependencies => ({
getCapability: () => personalWechatCapabilityService.getPersonalWechatSendCapability(),
generateReport: generateAgentGroupReport,
@@ -161,6 +142,29 @@ interface ScheduledReportNotificationPayload {
suggestedAction?: string
}
export function buildScheduledReportActionRequest(
input: ScheduledReportSendActionInput
): WechatActionRequest {
return {
idempotencyKey:
input.retryCount && input.retryCount > 0
? `scheduled_report:${input.executionId}:retry:${input.retryCount}`
: `scheduled_report:${input.executionId}`,
origin: 'scheduled_report',
purpose: 'scheduled_report',
triggerType: input.triggerType === 'scheduled' ? 'automation' : 'user',
sourceId: input.executionId,
executionId: input.executionId,
// filehelper 是文件传输助手:唯一的真实发送验收对象,走 contact 而不是群。
recipient: { type: input.target === 'filehelper' ? 'contact' : 'group', id: input.target },
content: { type: 'image', path: input.filePath },
metadata: {
taskId: input.taskId,
retryCount: input.retryCount || 0
}
}
}
export function validateScheduleTime(value: string): boolean {
return /^(?:[01]\d|2[0-3]):[0-5]\d$/.test(String(value || '').trim())
}
@@ -498,11 +502,7 @@ export class ScheduledReportService {
await this.load()
const execution = this.executions!.find((item) => item.id === executionId)
if (!execution) return { success: false, error: '未找到定时日报执行记录' }
const existing = this.retrying.get(executionId)
if (existing) return { success: true, data: await existing }
const promise = this.retrySend(execution).finally(() => this.retrying.delete(executionId))
this.retrying.set(executionId, promise)
const result = await promise
const result = await this.startRetry(execution)
return {
success: result.status !== 'failed',
data: result,
@@ -568,6 +568,7 @@ export class ScheduledReportService {
async tick(at = this.deps.now?.() || new Date()): Promise<void> {
await this.load()
await this.flushNotifications()
await this.flushPendingSends(at)
if (!this.deps.isDatabaseReady()) return
const nowMs = at.getTime()
for (const task of [...this.tasks!]) {
@@ -959,6 +960,15 @@ export class ScheduledReportService {
return completed
}
/** 为同一执行记录复用唯一的重发 Promise,合并自动重发和手动重试。 */
private startRetry(execution: ScheduledReportExecution): Promise<ScheduledReportExecution> {
const existing = this.retrying.get(execution.id)
if (existing) return existing
const promise = this.retrySend(execution).finally(() => this.retrying.delete(execution.id))
this.retrying.set(execution.id, promise)
return promise
}
private clearFailureFields(
execution: ScheduledReportExecution,
patch: Partial<ScheduledReportExecution>
@@ -1190,6 +1200,40 @@ export class ScheduledReportService {
await Promise.all([this.saveNotifications(), this.saveExecutions()])
}
/** waiting_to_send 执行的下一次自动重发时间:自 finishedAt 起按指数退避。 */
private autoResendDueAt(execution: ScheduledReportExecution): number {
const base = Date.parse(execution.finishedAt || execution.startedAt)
const attempt = Math.max(execution.retryCount || 0, 0)
return base + Math.min(AUTO_RESEND_BASE_BACKOFF_MS * 2 ** attempt, AUTO_RESEND_MAX_BACKOFF_MS)
}
/** 能力恢复后自动重发挂起的日报;重试有上限,超限后仍可手动重试。 */
private async flushPendingSends(at: Date): Promise<void> {
await this.load()
const nowMs = at.getTime()
const candidates = this.executions!.filter(
(item) =>
item.status === 'waiting_to_send' &&
item.retryable !== false &&
Boolean(item.pngPath) &&
(item.retryCount || 0) < AUTO_RESEND_MAX_ATTEMPTS &&
!this.retrying.has(item.id) &&
this.autoResendDueAt(item) <= nowMs
)
const retries: Promise<ScheduledReportExecution>[] = []
for (const execution of candidates) {
const task = this.tasks!.find((item) => item.id === execution.taskId)
if (!task || !task.enabled) continue
retries.push(this.startRetry(execution))
}
const results = await Promise.allSettled(retries)
for (const result of results) {
if (result.status === 'rejected') {
console.warn('[ScheduledReport] auto resend failed:', result.reason)
}
}
}
private notificationCapabilityReasonForSend(
result: AgentHubNotificationResult
): ScheduledReportNotificationCapabilityReason {
@@ -1219,6 +1263,8 @@ export class ScheduledReportService {
private resolveTarget(task: ScheduledReportTask): string | undefined {
const explicit = String(task.target || '').trim()
if (explicit.endsWith('@chatroom')) return explicit
// 文件传输助手是唯一允许真实发送验收的个人目标。
if (explicit.toLowerCase() === 'filehelper') return 'filehelper'
const contact = resolveMd5(task.group || explicit)
return contact?.m_nsUsrName?.endsWith('@chatroom') ? contact.m_nsUsrName : undefined
}
+637
View File
@@ -0,0 +1,637 @@
// src/main/services/system-ocr-service.ts
//
// System OCR Runtime(本地图片文字识别)。
//
// 职责边界(只做这些事):
// 1. capability detection
// 2. image normalization / preparation
// 3. OCR execution
// 4. result normalization
// 5. runtime metadata
// 6. error mapping
//
// 明确不做:
// - 不伪装成 AI Provider / Vision Model;不读写 AIVisionRuntimeConfig;
// - 不发任何网络请求;不上传原图;
// - 不遍历历史图片、不做 backfill、不写 Knowledge;
// - 不把 OCR 文本写进 image-insights.json(那是 Vision 结果的缓存)。
//
// Windows 后端:Windows.Media.Ocr.OcrEngine(经 @napi-rs/system-ocr)。
// 已实测的引擎行为(@napi-rs/system-ocr 1.2.0 / Electron 43 / Windows x64):
// - Buffer 输入只接受 PNG;JPEG / WEBP / BMP 会被判为不可识别,
// 所以本服务在边界上统一归一化成 PNG 字节再调用(不落盘)。
// - preferredLangs 只使用第一个语言;语言包缺失时引擎创建失败,
// 抛出的错误是 `Windows error 操作成功完成。 (0x00000000)`(HRESULT 为 S_OK)。
// - 空白图不会报错,返回空文本 → 映射成 OCR_EMPTY_RESULT。
// - CJK 字符之间会被引擎插入空格,结果里做归一化。
//
// macOS 后端:Apple Vision(同一 native 包,darwin binding)。
// 已实测的引擎行为(1.2.0 / macOS 15.7.7 / arm64):
// - Buffer 输入 PNG / JPEG / WEBP / GIF / BMP / TIFF **全部直接可用**,
// 所以 macOS 不做任何归一化,原始字节直通(不落盘、不起 ffmpeg 子进程)。
// - preferredLangs 对识别结果没有可观测影响(Vision 自行决定识别语言),
// 因此默认不传语言提示;显式指定 language 时仍然透传。
// - 畸形图片抛普通 Error:`CRImage Reader Detector was given zero-dimensioned image (0 x 0)`;
// 任一边 ≤2px 抛 `The image is too small in at least one dimension ...` → 都映射成 IMAGE_DECODE_FAILED。
// - macOS 没有"语言包缺失"这一失败模式。
import crypto from 'node:crypto'
import { spawn } from 'node:child_process'
import {
SYSTEM_OCR_CACHE_TTL_MS,
SYSTEM_OCR_PROBE_PNG_BASE64,
buildSystemOcrCacheKey,
detectSystemOcrImageFormat,
isSystemOcrPlatform,
mapSystemOcrNativeError,
normalizeSystemOcrText,
parseImageDataUrl,
resolveSystemOcrEngine,
resolveSystemOcrLanguageTag
} from '../../shared/system-ocr'
import type {
SystemOcrCapability,
SystemOcrEngine,
SystemOcrErrorCode,
SystemOcrImageFormat,
SystemOcrLine,
SystemOcrRequest,
SystemOcrResult
} from '../../shared/system-ocr'
const NATIVE_PACKAGE = '@napi-rs/system-ocr'
const MAX_CACHE_ENTRIES = 32
const FFMPEG_TIMEOUT_MS = 10_000
interface NativeLine {
text: string
confidence: number
boundingBox: { x: number; y: number; width: number; height: number }
}
interface NativeResult {
text: string
confidence: number
lines: NativeLine[]
}
interface NativeRuntime {
version: string | null
recognize: (image: Uint8Array, accuracy?: number, languages?: string[]) => Promise<NativeResult>
}
export interface SystemOcrServiceDeps {
/** 加载 native 运行时;不可用时返回 null(不允许抛) */
loadRuntime?: () => NativeRuntime | null
/**
* 把输入图片转成引擎可接受的字节;失败返回 null。
*
* Windows 后端只吃 PNG,必须走这一步;macOS 的 Vision 直接接受
* PNG / JPEG / WEBP / GIF / BMP / TIFF,默认实现直接透传原始字节。
*/
toPngBytes?: (input: {
buffer: Buffer
format: SystemOcrImageFormat
}) => Promise<Buffer | null>
/**
* ffmpeg 可执行文件解析器。只用于 Windows 上 GIF/BMP/WebP/TIFF → PNG 的兜底归一化。
* main/index.ts 会注入项目统一的解析逻辑(与图片解密共用一套候选路径)。
*/
resolveFfmpegExecutable?: () => string
platform?: NodeJS.Platform
arch?: string
/** 系统 locale(如 zh-CN),用于推导 OCR 语言标签 */
locale?: () => string
}
const failure = (
engine: SystemOcrEngine,
errorCode: SystemOcrErrorCode,
error: string,
startedAt: number
): SystemOcrResult => ({
success: false,
text: '',
lines: [],
language: null,
engine,
durationMs: Date.now() - startedAt,
errorCode,
error
})
/** 未被显式注入时的兜底:环境变量 → 打包内 ffmpeg-static → PATH。 */
const defaultResolveFfmpegExecutable = (): string => {
const fromEnvironment = String(process.env['FFMPEG_BIN'] || '').trim()
if (fromEnvironment) return fromEnvironment
try {
const bundled = require('ffmpeg-static') as string | null
if (bundled) return bundled
} catch {
// 忽略:退回到 PATH 上的 ffmpeg
}
return process.platform === 'win32' ? 'ffmpeg.exe' : 'ffmpeg'
}
const toLines = (lines: NativeLine[] | undefined): SystemOcrLine[] =>
Array.isArray(lines)
? lines.map((line) => ({
text: normalizeSystemOcrText(line.text),
confidence: typeof line.confidence === 'number' ? line.confidence : 1,
boundingBox: {
x: Number(line.boundingBox?.x ?? 0),
y: Number(line.boundingBox?.y ?? 0),
width: Number(line.boundingBox?.width ?? 0),
height: Number(line.boundingBox?.height ?? 0)
}
}))
: []
/** 把任意容器(gif/bmp/webp/tiff)用 ffmpeg 走内存管道转成 PNG。不落盘。仅 Windows 归一化路径会用到。 */
const convertWithFfmpeg = (buffer: Buffer, executable: string): Promise<Buffer | null> =>
new Promise((resolve) => {
let settled = false
const finish = (value: Buffer | null): void => {
if (settled) return
settled = true
resolve(value)
}
let child: ReturnType<typeof spawn>
try {
child = spawn(
executable,
[
'-hide_banner',
'-loglevel',
'error',
'-i',
'pipe:0',
'-frames:v',
'1',
'-f',
'image2pipe',
'-vcodec',
'png',
'pipe:1'
],
{ windowsHide: true }
)
} catch {
finish(null)
return
}
const chunks: Buffer[] = []
const timeout = setTimeout(() => {
try {
child.kill()
} catch {
// best-effort
}
finish(null)
}, FFMPEG_TIMEOUT_MS)
child.stdout?.on('data', (chunk: Buffer) => chunks.push(chunk))
child.on('error', () => {
clearTimeout(timeout)
finish(null)
})
child.on('close', (code) => {
clearTimeout(timeout)
finish(code === 0 && chunks.length > 0 ? Buffer.concat(chunks) : null)
})
child.stdin?.on('error', () => undefined)
child.stdin?.end(buffer)
})
class SystemOcrService {
private runtime: NativeRuntime | null = null
private runtimeLoaded = false
private capability: SystemOcrCapability | null = null
private capabilityPromise: Promise<SystemOcrCapability> | null = null
private readonly cache = new Map<string, { value: SystemOcrResult; expireAt: number }>()
private deps: SystemOcrServiceDeps = {}
constructor(deps: SystemOcrServiceDeps = {}) {
this.deps = deps
}
/** 由 main/index.ts 在 app ready 后调用(可选,用于注入 app.getLocale 等)。 */
bind(deps: SystemOcrServiceDeps): void {
this.deps = { ...this.deps, ...deps }
this.runtime = null
this.runtimeLoaded = false
this.capability = null
this.capabilityPromise = null
}
/** 仅测试用:清空探测与缓存状态。 */
reset(): void {
this.runtime = null
this.runtimeLoaded = false
this.capability = null
this.capabilityPromise = null
this.cache.clear()
}
private get platform(): NodeJS.Platform {
return this.deps.platform ?? process.platform
}
private get arch(): string {
return this.deps.arch ?? process.arch
}
/** 该平台对应的引擎标识。进 artifact 指纹与缓存 key,不要硬编码。 */
private get engine(): SystemOcrEngine {
return resolveSystemOcrEngine(this.platform)
}
/** 本平台是否为 macOS 后端(决定是否跳过图片归一化)。 */
private get isMacBackend(): boolean {
return this.platform === 'darwin'
}
private get locale(): string {
if (this.deps.locale) {
try {
return this.deps.locale()
} catch {
return ''
}
}
try {
const { app } = require('electron') as typeof import('electron')
return app?.getLocale?.() ?? ''
} catch {
return ''
}
}
private loadRuntime(): NativeRuntime | null {
if (this.runtimeLoaded) return this.runtime
this.runtimeLoaded = true
if (this.deps.loadRuntime) {
this.runtime = this.deps.loadRuntime()
return this.runtime
}
if (!isSystemOcrPlatform(this.platform)) {
this.runtime = null
return this.runtime
}
try {
// 原生模块必须在打包时 external + asarUnpack,否则这里会 MODULE_NOT_FOUND。
const nativeModule = require(NATIVE_PACKAGE) as {
recognize: NativeRuntime['recognize']
}
let version: string | null = null
try {
version = (require(`${NATIVE_PACKAGE}/package.json`) as { version?: string }).version ?? null
} catch {
version = null
}
this.runtime =
nativeModule && typeof nativeModule.recognize === 'function'
? { version, recognize: nativeModule.recognize.bind(nativeModule) }
: null
} catch (error) {
console.warn(
'[SystemOcrService] native runtime unavailable engine=%s platform=%s reason=%s',
this.engine,
this.platform,
error instanceof Error ? error.message.split('\n')[0] : String(error)
)
this.runtime = null
}
return this.runtime
}
private async toPngBytes(
buffer: Buffer,
format: SystemOcrImageFormat
): Promise<Buffer | null> {
if (this.deps.toPngBytes) return this.deps.toPngBytes({ buffer, format })
/*
* macOS:Vision 后端直接接受 PNG / JPEG / WEBP / GIF / BMP / TIFF(已实测),
* 归一化没有收益,只会白白多一次转码或一个 ffmpeg 子进程 —— 原字节直通。
*/
if (this.isMacBackend) return buffer
if (format === 'png') return buffer
if (format === 'jpeg') {
// 项目内已有的进程内解码能力,优先于 ffmpeg(更快、无子进程)。
try {
const { nativeImage } = require('electron') as typeof import('electron')
const image = nativeImage.createFromBuffer(buffer)
if (!image.isEmpty()) {
const png = image.toPNG()
if (png && png.length > 0) return png
}
} catch {
// 继续走 ffmpeg 兜底
}
}
try {
const resolveFfmpeg = this.deps.resolveFfmpegExecutable ?? defaultResolveFfmpegExecutable
const png = await convertWithFfmpeg(buffer, resolveFfmpeg())
return png
} catch {
return null
}
}
/** capability 探测:平台 → native 运行时 → 至少一个可用 OCR 语言。 */
async getCapability(force = false): Promise<SystemOcrCapability> {
if (!force && this.capability) return this.capability
if (!force && this.capabilityPromise) return this.capabilityPromise
this.capabilityPromise = this.detectCapability()
try {
this.capability = await this.capabilityPromise
} finally {
this.capabilityPromise = null
}
return this.capability
}
private async detectCapability(): Promise<SystemOcrCapability> {
const base: Pick<
SystemOcrCapability,
'engine' | 'platform' | 'arch' | 'runtimeVersion' | 'language'
> = {
engine: this.engine,
platform: this.platform,
arch: this.arch,
runtimeVersion: null,
language: null
}
if (!isSystemOcrPlatform(this.platform)) {
return {
...base,
available: false,
reason: 'UNSUPPORTED_PLATFORM',
message: '本地图片文字识别目前支持 Windows 与 macOS。'
}
}
const runtime = this.loadRuntime()
if (!runtime) {
return {
...base,
available: false,
reason: 'NATIVE_MODULE_MISSING',
message: '本地文字识别组件不可用,请重新安装 TraceMemo。'
}
}
const probed = await this.probeLanguage(runtime)
if (probed.reason) {
return {
...base,
runtimeVersion: runtime.version,
available: false,
reason: probed.reason,
message: probed.message
}
}
const engineLabel = this.isMacBackend ? 'macOS 系统 OCR' : 'Windows 系统 OCR'
return {
...base,
runtimeVersion: runtime.version,
available: true,
language: probed.language,
message: probed.language
? `本地图片文字识别可用(${engineLabel},${probed.language})。`
: `本地图片文字识别可用(${engineLabel},跟随系统语言)。`
}
}
/**
* 用一个 64x32 纯白 PNG 探测引擎是否真的能跑。
*
* Windows:引擎创建依赖语言包,首选「系统 locale 推导出的标签」,
* 失败再退回「系统用户语言配置」,并据此区分 LANGUAGE_UNAVAILABLE。
*
* macOS:Vision 自行决定识别语言,**没有语言包缺失这一失败模式**,
* 所以不传语言提示,探测失败只可能是引擎本身起不来。
*/
private async probeLanguage(
runtime: NativeRuntime
): Promise<
| { language: string | null; reason?: undefined; message?: undefined }
| { language: null; reason: 'LANGUAGE_UNAVAILABLE' | 'NATIVE_MODULE_MISSING'; message: string }
> {
const probeBuffer = Buffer.from(SYSTEM_OCR_PROBE_PNG_BASE64, 'base64')
if (this.isMacBackend) {
try {
await runtime.recognize(probeBuffer, undefined, undefined)
return { language: null }
} catch (error) {
/*
* 探测图是纯白图。Vision 对"图里没有文字"是**抛错**(`No text recognized`),
* 而抛这个错恰恰证明识别器跑通了 —— 不能当成引擎故障。
*/
if (
mapSystemOcrNativeError(error instanceof Error ? error.message : String(error)) ===
'OCR_EMPTY_RESULT'
) {
return { language: null }
}
return {
language: null,
reason: 'NATIVE_MODULE_MISSING',
message: '本地文字识别引擎初始化失败,请重启 TraceMemo 或重新安装。'
}
}
}
const preferred = resolveSystemOcrLanguageTag(this.locale, this.platform)
const candidates: Array<string | null> = preferred ? [preferred, null] : [null]
let lastCode: SystemOcrErrorCode = 'OCR_FAILED'
for (const candidate of candidates) {
try {
await runtime.recognize(
probeBuffer,
undefined,
candidate ? [candidate] : undefined
)
return { language: candidate }
} catch (error) {
lastCode = mapSystemOcrNativeError(
error instanceof Error ? error.message : String(error)
)
}
}
if (lastCode === 'OCR_LANGUAGE_UNAVAILABLE') {
return {
language: null,
reason: 'LANGUAGE_UNAVAILABLE',
message:
'当前 Windows 未安装可用的 OCR 语言支持,请在系统「语言和区域」里安装简体中文或英文的 OCR 语言包后重试。'
}
}
return {
language: null,
reason: 'NATIVE_MODULE_MISSING',
message: '本地文字识别引擎初始化失败,请重启 TraceMemo 或重新安装。'
}
}
private readCache(key: string, startedAt: number): SystemOcrResult | null {
const hit = this.cache.get(key)
if (!hit) return null
if (hit.expireAt <= Date.now()) {
this.cache.delete(key)
return null
}
return { ...hit.value, durationMs: Date.now() - startedAt, fromCache: true }
}
private writeCache(key: string, value: SystemOcrResult): void {
if (this.cache.size >= MAX_CACHE_ENTRIES) {
const oldest = this.cache.keys().next()
if (!oldest.done) this.cache.delete(oldest.value)
}
this.cache.set(key, { value, expireAt: Date.now() + SYSTEM_OCR_CACHE_TTL_MS })
}
/**
* 识别一张图片里的文字。
* 任意失败都不抛,统一返回 success=false + 产品级 errorCode。
*/
async recognize(request: SystemOcrRequest): Promise<SystemOcrResult> {
const startedAt = Date.now()
const engine = this.engine
const parsed = parseImageDataUrl(request.imageDataUrl)
if (!parsed) {
return failure(
engine,
'UNSUPPORTED_IMAGE',
'仅支持 PNG、JPG、JPEG、WebP、GIF、BMP 图片。',
startedAt
)
}
let sourceBuffer: Buffer
try {
sourceBuffer = Buffer.from(parsed.base64, 'base64')
} catch {
return failure(engine, 'IMAGE_DECODE_FAILED', '图片数据无法解码。', startedAt)
}
if (sourceBuffer.length === 0) {
return failure(engine, 'IMAGE_DECODE_FAILED', '图片数据为空。', startedAt)
}
const format = detectSystemOcrImageFormat(sourceBuffer)
if (!format) {
return failure(engine, 'UNSUPPORTED_IMAGE', '无法识别的图片格式。', startedAt)
}
const imageHash =
request.imageHash?.trim() ||
crypto.createHash('sha256').update(sourceBuffer).digest('hex').slice(0, 32)
const requestedLanguage = request.language?.trim() || null
const capability = await this.getCapability()
if (!capability.available) {
const errorCode: SystemOcrErrorCode =
capability.reason === 'UNSUPPORTED_PLATFORM'
? 'UNSUPPORTED_PLATFORM'
: capability.reason === 'LANGUAGE_UNAVAILABLE'
? 'OCR_LANGUAGE_UNAVAILABLE'
: 'SYSTEM_OCR_UNAVAILABLE'
return failure(engine, errorCode, capability.message, startedAt)
}
const languageForCache = requestedLanguage ?? capability.language
const cacheKey = buildSystemOcrCacheKey({
imageHash,
language: languageForCache,
runtimeVersion: capability.runtimeVersion,
platform: capability.platform,
engine: capability.engine
})
if (requestedLanguage === null) {
const cached = this.readCache(cacheKey, startedAt)
if (cached) return cached
}
const runtime = this.loadRuntime()
if (!runtime) {
return failure(engine, 'SYSTEM_OCR_UNAVAILABLE', '本地文字识别组件不可用。', startedAt)
}
const prepared = await this.toPngBytes(sourceBuffer, format)
if (!prepared || prepared.length === 0 || !detectSystemOcrImageFormat(prepared)) {
return failure(engine, 'IMAGE_DECODE_FAILED', '图片解码失败,无法读取这张图片。', startedAt)
}
const candidates: Array<string | null> = requestedLanguage
? [requestedLanguage]
: capability.language
? [capability.language, null]
: [null]
let lastErrorCode: SystemOcrErrorCode = 'OCR_FAILED'
let lastErrorMessage = ''
let usedLanguage: string | null = null
for (const candidate of candidates) {
try {
// accuracy 传 undefined = 用 native 默认值,而该默认是 `Accurate`
// (见 @napi-rs/system-ocr 的 recognize 文档)。**不要改成 Fast**:
// 低精度档在中文上会明显掉字。Windows 忽略该参数。
const result = await runtime.recognize(
prepared,
undefined,
candidate ? [candidate] : undefined
)
const text = normalizeSystemOcrText(result?.text ?? '')
const lines = toLines(result?.lines)
usedLanguage = candidate
if (!text) {
return failure(engine, 'OCR_EMPTY_RESULT', '没有在这张图片里识别到文字。', startedAt)
}
const succeeded: SystemOcrResult = {
success: true,
text,
lines,
language: usedLanguage,
engine,
durationMs: Date.now() - startedAt
}
/**
* 成功路径**刻意不逐张打日志**。
*
* 后台回填会连续识别几万张图片,逐张一条成功日志既是没有信息量的噪声,
* 又会把日志刷爆。逐张耗时由 `ImageTextIndexService` 的阶段画像低频汇总,
* 单张的 `durationMs` / 字数依然在**返回值**里(设置页的单图诊断就是用它)。
* 只有失败才值得在默认输出里留痕 —— 见下面的 `failed`。
*/
this.writeCache(cacheKey, succeeded)
return succeeded
} catch (error) {
lastErrorMessage = error instanceof Error ? error.message : String(error)
lastErrorCode = mapSystemOcrNativeError(lastErrorMessage)
/*
* macOS 的 Vision 在"图里没有文字"时是抛错(`No text recognized`)而不是返回空文本。
* 它必须走 empty 语义:表情包 / 风景 / 头像都是**正常终态**,不是 OCR 失败 ——
* 否则这些图片会落成可重试失败,被反复重算,覆盖率也会说谎。
*/
if (lastErrorCode === 'OCR_EMPTY_RESULT') {
return failure(engine, 'OCR_EMPTY_RESULT', '没有在这张图片里识别到文字。', startedAt)
}
// 语言不可用才值得换下一个候选;其它错误直接结束,避免无意义重试。
if (lastErrorCode !== 'OCR_LANGUAGE_UNAVAILABLE') break
}
}
console.warn(
'[SystemOcrService] failed engine=%s platform=%s errorCode=%s durationMs=%d',
engine,
capability.platform,
lastErrorCode,
Date.now() - startedAt
)
return failure(
engine,
lastErrorCode,
lastErrorCode === 'OCR_LANGUAGE_UNAVAILABLE'
? '当前 Windows 未安装可用的 OCR 语言支持。'
: '本地文字识别失败,请稍后重试。',
startedAt
)
}
}
export { SystemOcrService }
export const systemOcrService = new SystemOcrService()
+3 -2
View File
@@ -17,7 +17,7 @@ import type {
WechatActionResult
} from '../../shared/wechat-action'
import { personalWechatCapabilityService } from './personal-wechat-capability-service'
import { personalWechatSendService } from './personal-wechat-send-service'
import { wechatSendGateway } from './wechat-send-gateway'
const MAX_AUDIT_RECORDS = 500
const MAX_CONTENT_PREVIEW_LENGTH = 240
@@ -54,7 +54,8 @@ const defaultDependencies = (): Required<
>
> => ({
getCapability: () => personalWechatCapabilityService.getPersonalWechatSendCapability(),
send: (request) => personalWechatSendService.send(request),
// 高层业务审计之后仍然统一走 WechatSendGateway,保证每一次真实发送都有 Send Log。
send: (request) => wechatSendGateway.sendPersonal(request),
getUserDataPath: () => app.getPath('userData'),
now: () => new Date(),
wait: (milliseconds) => new Promise((resolve) => setTimeout(resolve, milliseconds))
@@ -0,0 +1,262 @@
import {
chmodSync,
existsSync,
mkdirSync,
readdirSync,
readFileSync,
renameSync,
rmSync,
writeFileSync,
type Dirent
} from 'node:fs'
import { homedir } from 'node:os'
import { join } from 'node:path'
import type { ILinkCredentials } from './types'
/**
* 账号与游标持久化。
*
* 目录布局沿用历史命名(含 `wechat-connector` 这一层),保证升级后**不需要重新扫码**:
* ~/.tracememo/wechat-connector/accounts/<account-id>.json 凭据
* ~/.tracememo/wechat-connector/accounts/<account-id>.sync.json 长轮询游标
* ~/.tracememo/wechat-connector/accounts/<account-id>.context.json 会话上下文令牌
*
* 历史目录 ~/.wechatexplorer/wechat-connector/accounts 只读兼容,不删除、不覆写。
*/
const CONNECTOR_DIR_NAME = 'wechat-connector'
const ACCOUNTS_DIR_NAME = 'accounts'
const CURRENT_ROOT = '.tracememo'
const LEGACY_ROOT = '.wechatexplorer'
const FILE_MODE = 0o600
const DIR_MODE = 0o700
export type HomeDirectoryResolver = () => string
export function normalizeAccountId(raw: string): string {
const value = String(raw ?? '')
if (!value) return ''
// @ . : 统一替换为 -,使账号 ID 可以作为文件名。
return value.replace(/[@.:]/g, '-')
}
export function accountsDirectory(home: HomeDirectoryResolver = homedir): string {
return join(home(), CURRENT_ROOT, CONNECTOR_DIR_NAME, ACCOUNTS_DIR_NAME)
}
export function legacyAccountsDirectory(home: HomeDirectoryResolver = homedir): string {
return join(home(), LEGACY_ROOT, CONNECTOR_DIR_NAME, ACCOUNTS_DIR_NAME)
}
function ensureDirectory(directory: string): void {
mkdirSync(directory, { recursive: true, mode: DIR_MODE })
}
/** 原子写入:先写临时文件再 rename,避免进程退出留下半截 JSON。 */
function writeJsonAtomically(path: string, value: unknown): void {
const tempPath = `${path}.tmp-${process.pid}-${Date.now()}`
try {
writeFileSync(tempPath, JSON.stringify(value, null, 2), { encoding: 'utf8', mode: FILE_MODE })
chmodSync(tempPath, FILE_MODE)
renameSync(tempPath, path)
chmodSync(path, FILE_MODE)
} catch (error) {
try {
rmSync(tempPath, { force: true })
} catch {
// 清理失败不影响主流程:文件本身不会覆盖有效数据。
}
throw error
}
}
function readJsonIfValid<T>(path: string, validate: (value: T) => boolean): T | undefined {
try {
const parsed = JSON.parse(readFileSync(path, 'utf8')) as T
return validate(parsed) ? parsed : undefined
} catch {
return undefined
}
}
function isCredentials(value: unknown): value is ILinkCredentials {
const candidate = value as Partial<ILinkCredentials> | null
return Boolean(candidate && typeof candidate.bot_token === 'string' && candidate.bot_token)
}
/** 找到某个账号凭据实际所在目录:优先当前目录,其次历史目录。 */
export function resolveAccountDirectory(
accountId: string,
home: HomeDirectoryResolver = homedir
): { directory: string; legacy: boolean } {
const current = accountsDirectory(home)
if (existsSync(join(current, `${accountId}.json`))) return { directory: current, legacy: false }
const legacy = legacyAccountsDirectory(home)
if (existsSync(join(legacy, `${accountId}.json`))) return { directory: legacy, legacy: true }
return { directory: current, legacy: false }
}
/**
* 保存新登录的凭据。
* 先写入新凭据再清理旧账号文件,因此一次失败的登录不会摧毁上一个可用账号。
* 同时保留同前缀的 `<id>.sync.json` / `<id>.context.json`。
*/
export function saveCredentials(
credentials: ILinkCredentials,
home: HomeDirectoryResolver = homedir
): string {
const directory = accountsDirectory(home)
ensureDirectory(directory)
const accountId = normalizeAccountId(credentials.ilink_bot_id)
if (!accountId) throw new Error('登录凭据缺少 ilink_bot_id,无法保存')
const path = join(directory, `${accountId}.json`)
writeJsonAtomically(path, credentials)
const keepPrefix = `${accountId}.`
for (const entry of readdirSync(directory, { withFileTypes: true })) {
if (entry.isDirectory()) continue
if (!entry.name.endsWith('.json')) continue
if (entry.name.startsWith(keepPrefix)) continue
try {
rmSync(join(directory, entry.name), { force: true })
} catch {
// 旧凭据清理失败不影响新凭据可用性。
}
}
return path
}
function loadCredentialsFromDirectory(directory: string): ILinkCredentials[] {
let entries: Dirent[]
try {
entries = readdirSync(directory, { withFileTypes: true })
} catch {
return []
}
const result: ILinkCredentials[] = []
for (const entry of entries) {
if (entry.isDirectory() || !entry.name.endsWith('.json')) continue
const credential = readJsonIfValid<ILinkCredentials>(join(directory, entry.name), isCredentials)
if (credential) result.push(credential)
}
return result
}
/** 当前目录优先;当前目录为空时回退到历史目录(一次性升级兼容,不迁移不删除)。 */
export function loadAllCredentials(home: HomeDirectoryResolver = homedir): ILinkCredentials[] {
const current = loadCredentialsFromDirectory(accountsDirectory(home))
if (current.length > 0) return current
return loadCredentialsFromDirectory(legacyAccountsDirectory(home))
}
export function findCredentials(
accountId: string,
home: HomeDirectoryResolver = homedir
): ILinkCredentials | undefined {
const { directory } = resolveAccountDirectory(accountId, home)
const credential = readJsonIfValid<ILinkCredentials>(
join(directory, `${accountId}.json`),
isCredentials
)
if (credential) return credential
return loadAllCredentials(home).find(
(item) => normalizeAccountId(item.ilink_bot_id) === accountId || item.ilink_bot_id === accountId
)
}
/* ------------------------------------------------------------------ */
/* 长轮询游标 */
/* ------------------------------------------------------------------ */
interface SyncRecord {
get_updates_buf: string
}
/**
* 游标按账号隔离保存。切换账号时绝不沿用上一个账号的游标。
*/
export function loadCursor(accountId: string, home: HomeDirectoryResolver = homedir): string {
const { directory } = resolveAccountDirectory(accountId, home)
const record = readJsonIfValid<SyncRecord>(join(directory, `${accountId}.sync.json`), (value) =>
Boolean(value && typeof value.get_updates_buf === 'string')
)
return record?.get_updates_buf ?? ''
}
export function saveCursor(
accountId: string,
getUpdatesBuf: string,
home: HomeDirectoryResolver = homedir
): void {
// 必须写到 loadCursor 实际读取的目录:如果凭据仍在历史目录(尚未迁移),
// 写进当前目录会导致游标永远读不回来,长轮询就会反复重投同一条消息。
const { directory } = resolveAccountDirectory(accountId, home)
ensureDirectory(directory)
writeJsonAtomically(join(directory, `${accountId}.sync.json`), {
get_updates_buf: getUpdatesBuf
} satisfies SyncRecord)
}
export function clearCursor(accountId: string, home: HomeDirectoryResolver = homedir): void {
const { directory } = resolveAccountDirectory(accountId, home)
try {
rmSync(join(directory, `${accountId}.sync.json`), { force: true })
} catch {
// 游标清理失败不阻断重连流程。
}
}
/* ------------------------------------------------------------------ */
/* 会话上下文令牌 */
/* ------------------------------------------------------------------ */
interface ContextRecord {
tokens: Record<string, { context_token: string; updated_at: number }>
}
/**
* context_token 属于**会话上下文**,不是账号长期凭据。
* 按「账号 + 用户」保存最近一次有效值,用于重启后的主动发送(例如定时日报)。
* 永不写入日志,文件权限 0600。
*/
export function saveContextToken(
accountId: string,
toUserId: string,
contextToken: string,
home: HomeDirectoryResolver = homedir,
now: () => number = Date.now
): void {
if (!contextToken) return
const { directory } = resolveAccountDirectory(accountId, home)
ensureDirectory(directory)
const path = join(directory, `${accountId}.context.json`)
const existing = readJsonIfValid<ContextRecord>(path, (value) =>
Boolean(value && typeof value.tokens === 'object' && value.tokens !== null)
) ?? { tokens: {} }
const tokens: ContextRecord['tokens'] = {
...existing.tokens,
[toUserId]: { context_token: contextToken, updated_at: now() }
}
// 只保留最近 500 条,避免文件无限增长。
const entries = Object.entries(tokens).sort(
(left, right) => right[1].updated_at - left[1].updated_at
)
const trimmed = Object.fromEntries(entries.slice(0, 500))
writeJsonAtomically(path, { tokens: trimmed } satisfies ContextRecord)
}
export function loadContextToken(
accountId: string,
toUserId: string,
home: HomeDirectoryResolver = homedir
): string | undefined {
const { directory } = resolveAccountDirectory(accountId, home)
const record = readJsonIfValid<ContextRecord>(
join(directory, `${accountId}.context.json`),
(value) => Boolean(value && typeof value.tokens === 'object' && value.tokens !== null)
)
return record?.tokens?.[toUserId]?.context_token || undefined
}
+293
View File
@@ -0,0 +1,293 @@
import { ILinkError } from './errors'
import { ILinkClient, type FetchLike } from './client'
import { loadAllCredentials, saveCredentials, type HomeDirectoryResolver } from './account-store'
import {
ILINK_DEFAULT_BASE_URL,
ILINK_QR_STATUS_TIMEOUT_MS,
type ILinkCredentials,
type ILinkQrCodeResponse,
type ILinkQrStatusResponse,
type WechatLoginEvent
} from './types'
const QR_CODE_PATH = '/ilink/bot/get_bot_qrcode?bot_type=3'
const QR_STATUS_PATH = '/ilink/bot/get_qrcode_status'
/** 官方客户端策略(非协议常量):本地二维码 TTL 约 5 分钟,整体等待约 480 秒。 */
const QR_SESSION_TTL_MS = 5 * 60_000
const LOGIN_DEADLINE_MS = 480_000
const MAX_QR_REFRESHES = 3
/** 等待用户输入数字配对码的上限。 */
const VERIFY_CODE_WAIT_MS = 60_000
export type QrEncoder = (content: string) => Promise<string>
export interface QrLoginDependencies {
fetchImpl?: FetchLike
home?: HomeDirectoryResolver
baseUrl?: string
signal?: AbortSignal
onEvent?: (event: WechatLoginEvent) => void
/** 把二维码内容渲染成可直接给 <img src> 的 data URL。 */
qrEncoder?: QrEncoder
/** 手机端要求数字配对码时,由宿主提供;返回 undefined 表示放弃这次登录。 */
verifyCodeProvider?: () => Promise<string | undefined>
now?: () => number
sleep?: (milliseconds: number) => Promise<void>
maxQrRefreshes?: number
deadlineMs?: number
}
/**
* 默认二维码渲染器:把服务端返回的内容编码成二维码 PNG。
* 服务端返回的是二维码页面 URL,因此这里必须自己编码,而不是当作图片直传。
*/
const defaultQrEncoder: QrEncoder = async (content) => {
if (/^data:image\//i.test(content)) return content
const { toDataURL } = await import('qrcode')
return toDataURL(content, { errorCorrectionLevel: 'L', margin: 1, width: 320 })
}
interface LoginState {
host: string
verifyCode?: string
pendingVerifyCode?: boolean
}
/** 等待用户输入配对码;超时视为放弃本次尝试,改为刷新二维码而不是无限等待。 */
function awaitVerifyCode(
provider: () => Promise<string | undefined>,
waitMs: number
): Promise<string | undefined> {
return new Promise<string | undefined>((resolve) => {
let settled = false
const finish = (value: string | undefined): void => {
if (settled) return
settled = true
clearTimeout(timer)
resolve(value)
}
const timer = setTimeout(() => finish(undefined), waitMs)
provider().then(
(value) => finish(value),
() => finish(undefined)
)
})
}
function isCancelled(signal: AbortSignal | undefined): boolean {
return signal?.aborted === true
}
interface QrStatusPollResult {
outcome: 'continue' | 'refresh' | 'done'
credentials?: ILinkCredentials
}
async function pollOnce(
client: ILinkClient,
qrcode: string,
state: LoginState,
deps: Required<Pick<QrLoginDependencies, 'onEvent' | 'now'>> & QrLoginDependencies
): Promise<QrStatusPollResult> {
const query = new URLSearchParams({ qrcode })
if (state.verifyCode) query.set('verify_code', state.verifyCode)
const url = `${state.host}${QR_STATUS_PATH}?${query.toString()}`
let response: ILinkQrStatusResponse
try {
response = await client.getAnonymous<ILinkQrStatusResponse>(url, '二维码状态轮询', {
timeoutMs: ILINK_QR_STATUS_TIMEOUT_MS,
...(deps.signal ? { signal: deps.signal } : {})
})
} catch (error) {
if (isCancelled(deps.signal)) throw error
// 网络超时 / 网关 524 属于长轮询正常控制流:保持当前二维码继续轮询。
return { outcome: 'continue' }
}
const status = String(response.status || '').trim()
switch (status) {
case 'wait':
deps.onEvent({ status: 'wait' })
return { outcome: 'continue' }
case 'scaned':
// 已扫码:若上一次轮询携带过配对码,说明已被接受,清除暂存值。
state.verifyCode = undefined
deps.onEvent({ status: 'scaned' })
return { outcome: 'continue' }
case 'need_verifycode': {
deps.onEvent({ status: 'need_verifycode' })
state.verifyCode = undefined
if (state.pendingVerifyCode) return { outcome: 'continue' }
state.pendingVerifyCode = true
const provider = deps.verifyCodeProvider
if (!provider) return { outcome: 'refresh' }
const code = await awaitVerifyCode(provider, VERIFY_CODE_WAIT_MS)
state.pendingVerifyCode = false
const normalized = String(code ?? '').trim()
if (!normalized) return { outcome: 'refresh' }
state.verifyCode = normalized
return { outcome: 'continue' }
}
case 'verify_code_blocked':
// 多次输入错误被限制:清除配对码并刷新二维码,由调用方计数。
deps.onEvent({ status: 'verify_code_blocked' })
state.verifyCode = undefined
state.pendingVerifyCode = false
return { outcome: 'refresh' }
case 'scaned_but_redirect': {
// 状态轮询需要切换 IDC 节点;只切换状态轮询主机,不影响已保存的 baseurl。
const redirectHost = String(response.redirect_host || '').trim()
if (redirectHost) {
state.host = /^https?:\/\//.test(redirectHost)
? redirectHost.replace(/\/+$/, '')
: `https://${redirectHost.replace(/\/+$/, '')}`
}
return { outcome: 'continue' }
}
case 'binded_redirect': {
// 账号已绑定到本客户端:只有本地确实仍有可用凭据时才能视为成功。
const existing = loadAllCredentials(deps.home)
const reusable = existing[existing.length - 1]
if (!reusable) return { outcome: 'refresh' }
deps.onEvent({ status: 'confirmed' })
return { outcome: 'done', credentials: reusable }
}
case 'expired':
deps.onEvent({ status: 'expired' })
return { outcome: 'refresh' }
case 'confirmed': {
const accountId = String(response.ilink_bot_id || '').trim()
if (!accountId) {
throw new ILinkError({
kind: 'protocol',
message: '登录已确认,但服务端未返回 ilink_bot_id'
})
}
const credentials: ILinkCredentials = {
bot_token: String(response.bot_token || ''),
ilink_bot_id: accountId,
baseurl: String(response.baseurl || deps.baseUrl || ILINK_DEFAULT_BASE_URL),
ilink_user_id: String(response.ilink_user_id || '')
}
if (!credentials.bot_token) {
throw new ILinkError({ kind: 'protocol', message: '登录已确认,但服务端未返回 bot_token' })
}
deps.onEvent({ status: 'confirmed' })
return { outcome: 'done', credentials }
}
default:
// 未知状态按 wait 处理,避免因为服务端新增状态直接打断登录。
return { outcome: 'continue' }
}
}
/**
* 扫码登录。
*
* 状态机与官方客户端对齐:wait / scaned / need_verifycode / verify_code_blocked /
* scaned_but_redirect / binded_redirect / expired / confirmed。
* 成功后立即原子落盘,且先写新凭据再清理旧账号,失败登录不会摧毁可用凭据。
*/
export async function runQrLogin(deps: QrLoginDependencies = {}): Promise<ILinkCredentials> {
const baseUrl = (deps.baseUrl || ILINK_DEFAULT_BASE_URL).replace(/\/+$/, '')
const now = deps.now ?? Date.now
const sleep =
deps.sleep ?? ((milliseconds: number) => new Promise<void>((r) => setTimeout(r, milliseconds)))
const onEvent = deps.onEvent ?? ((): void => undefined)
const maxRefreshes = deps.maxQrRefreshes ?? MAX_QR_REFRESHES
const deadline = deps.deadlineMs ?? LOGIN_DEADLINE_MS
const qrEncoder = deps.qrEncoder ?? defaultQrEncoder
const client = new ILinkClient({
baseUrl,
...(deps.fetchImpl ? { fetchImpl: deps.fetchImpl } : {})
})
const startedAt = now()
let refreshes = 0
while (true) {
if (isCancelled(deps.signal)) {
throw new ILinkError({ kind: 'aborted', message: '登录已取消' })
}
if (now() - startedAt > deadline) {
throw new ILinkError({ kind: 'timeout', message: '登录超时,请重新获取二维码' })
}
// local_token_list 只在取二维码时上报,便于服务端判断是否已绑定。
const localTokenList = loadAllCredentials(deps.home)
.map((item) => item.bot_token)
.filter(Boolean)
.slice(-10)
const qrResponse = await client.postAnonymous<ILinkQrCodeResponse>(
QR_CODE_PATH,
{ local_token_list: localTokenList },
'获取登录二维码',
{ timeoutMs: 20_000, ...(deps.signal ? { signal: deps.signal } : {}) }
)
const qrcode = String(qrResponse.qrcode || '').trim()
const qrcodeContent = String(qrResponse.qrcode_img_content || '').trim()
if (!qrcode || !qrcodeContent) {
throw new ILinkError({ kind: 'protocol', message: '服务端未返回有效的登录二维码' })
}
const qrCodeDataUrl = await qrEncoder(qrcodeContent)
onEvent({ status: 'qrcode', qrCodeDataUrl })
const state: LoginState = { host: baseUrl }
const sessionDeadline = now() + QR_SESSION_TTL_MS
let refresh = false
while (!refresh) {
if (isCancelled(deps.signal)) {
throw new ILinkError({ kind: 'aborted', message: '登录已取消' })
}
if (now() > sessionDeadline || now() - startedAt > deadline) {
throw new ILinkError({ kind: 'timeout', message: '二维码已过期,请重新获取' })
}
const result = await pollOnce(client, qrcode, state, {
...deps,
onEvent,
now
})
if (result.outcome === 'done') {
const credentials = result.credentials!
saveCredentials(credentials, deps.home)
onEvent({
status: 'active',
accountId: credentials.ilink_bot_id,
wechatUserId: credentials.ilink_user_id
})
return credentials
}
if (result.outcome === 'refresh') {
refreshes += 1
if (refreshes > maxRefreshes) {
throw new ILinkError({
kind: 'protocol',
message: '二维码多次刷新后仍未完成登录,请稍后重试'
})
}
refresh = true
continue
}
// 避免空转打满接口。
await sleep(300)
}
}
}
+284
View File
@@ -0,0 +1,284 @@
import { classifyNetworkError, describeErrorChain, excerptBody, ILinkError } from './errors'
import {
buildAuthorizedHeaders,
buildBaseInfo,
buildCommonHeaders,
buildQrStatusHeaders,
type ILinkHeaderOptions
} from './headers'
import {
ILINK_CONFIG_TIMEOUT_MS,
ILINK_DEFAULT_BASE_URL,
ILINK_SEND_TIMEOUT_MS,
type ILinkGetConfigResponse,
type ILinkGetUpdatesResponse,
type ILinkGetUploadUrlRequest,
type ILinkGetUploadUrlResponse,
type ILinkSendMessageRequest,
type ILinkSendMessageResponse,
type ILinkSendTypingResponse
} from './types'
export type FetchLike = typeof fetch
export interface ILinkClientOptions {
baseUrl?: string
botToken?: string
fetchImpl?: FetchLike
headers?: ILinkHeaderOptions
}
export interface ILinkRequestOptions {
timeoutMs: number
signal?: AbortSignal
/** 是否在失败时把响应体拼进错误信息(默认只带 HTTP 状态码)。 */
includeBody?: boolean
}
function withTimeout(
signal: AbortSignal | undefined,
timeoutMs: number
): { signal: AbortSignal; cleanup: () => void } {
const controller = new AbortController()
const timer = setTimeout(() => controller.abort(new Error('ilink request timeout')), timeoutMs)
const onAbort = (): void => controller.abort(signal?.reason)
if (signal) {
if (signal.aborted) controller.abort(signal.reason)
else signal.addEventListener('abort', onAbort, { once: true })
}
return {
signal: controller.signal,
cleanup: () => {
clearTimeout(timer)
signal?.removeEventListener('abort', onAbort)
}
}
}
async function parseJson<T>(response: Response, context: string, includeBody: boolean): Promise<T> {
const text = await response.text()
if (!response.ok) {
throw new ILinkError({
kind: 'http',
message: includeBody
? `${context}失败:HTTP ${response.status} ${excerptBody(text)}`
: `${context}失败:HTTP ${response.status}`,
httpStatus: response.status
})
}
if (!text.trim()) return {} as T
try {
return JSON.parse(text) as T
} catch {
throw new ILinkError({
kind: 'protocol',
message: `${context}失败:响应不是合法 JSON`,
httpStatus: response.status
})
}
}
/**
* iLink HTTP 客户端。
*
* 只负责传输层:网络失败 / 超时 / 非 2xx / JSON 解析失败会抛 ILinkError。
* 业务层 ret / errcode 的判断交给调用方(长轮询需要区分 -14)。
*/
export class ILinkClient {
private baseUrlValue: string
private botTokenValue: string
private readonly fetchImpl: FetchLike
private readonly headerOptions: ILinkHeaderOptions
constructor(options: ILinkClientOptions = {}) {
this.baseUrlValue = (options.baseUrl || ILINK_DEFAULT_BASE_URL).replace(/\/+$/, '')
this.botTokenValue = options.botToken || ''
this.fetchImpl = options.fetchImpl || globalThis.fetch
this.headerOptions = options.headers || {}
}
get baseUrl(): string {
return this.baseUrlValue
}
/** 登录返回 baseurl / redirect_host 后更新后续业务 API 节点。 */
setBaseUrl(baseUrl: string): void {
const normalized = String(baseUrl || '').trim()
if (!normalized) return
this.baseUrlValue = /^https?:\/\//.test(normalized)
? normalized.replace(/\/+$/, '')
: `https://${normalized.replace(/\/+$/, '')}`
}
setBotToken(botToken: string): void {
this.botTokenValue = botToken
}
get hasBotToken(): boolean {
return Boolean(this.botTokenValue)
}
private async request(
url: string,
init: RequestInit,
context: string,
options: ILinkRequestOptions
): Promise<Response> {
const { signal, cleanup } = withTimeout(options.signal, options.timeoutMs)
try {
return await this.fetchImpl(url, { ...init, signal })
} catch (error) {
const kind = classifyNetworkError(error)
const aborted = options.signal?.aborted === true
throw new ILinkError({
kind: aborted ? 'aborted' : kind,
// 带上完整错误链:顶层 fetch failed 没有信息量,原因在 cause 上。
message: `${context}失败:${describeErrorChain(error)}`,
cause: error
})
} finally {
cleanup()
}
}
/** 未鉴权 POST(获取二维码)。 */
async postAnonymous<T>(
path: string,
body: unknown,
context: string,
options: ILinkRequestOptions
): Promise<T> {
const response = await this.request(
`${this.baseUrlValue}${path}`,
{
method: 'POST',
headers: buildCommonHeaders(),
body: JSON.stringify(body)
},
context,
options
)
return parseJson<T>(response, context, options.includeBody === true)
}
/** 鉴权 POST(登录后的全部业务接口)。 */
async post<T>(
path: string,
body: unknown,
context: string,
options: ILinkRequestOptions
): Promise<T> {
const response = await this.request(
`${this.baseUrlValue}${path}`,
{
method: 'POST',
headers: buildAuthorizedHeaders(this.botTokenValue),
body: JSON.stringify(body)
},
context,
options
)
return parseJson<T>(response, context, options.includeBody === true)
}
/** 二维码状态轮询使用裸 URL(可能指向 redirect_host)。 */
async getAnonymous<T>(url: string, context: string, options: ILinkRequestOptions): Promise<T> {
const response = await this.request(
url,
{ method: 'GET', headers: buildQrStatusHeaders() },
context,
options
)
return parseJson<T>(response, context, options.includeBody === true)
}
private baseInfo(): ReturnType<typeof buildBaseInfo> {
return buildBaseInfo(this.headerOptions)
}
async getUpdates(
getUpdatesBuf: string,
options: ILinkRequestOptions
): Promise<ILinkGetUpdatesResponse> {
return this.post<ILinkGetUpdatesResponse>(
'/ilink/bot/getupdates',
{ get_updates_buf: getUpdatesBuf, base_info: this.baseInfo() },
'getupdates',
options
)
}
async sendMessage(
msg: ILinkSendMessageRequest['msg'],
signal?: AbortSignal
): Promise<ILinkSendMessageResponse> {
return this.post<ILinkSendMessageResponse>(
'/ilink/bot/sendmessage',
{ msg, base_info: this.baseInfo() },
'sendmessage',
{ timeoutMs: ILINK_SEND_TIMEOUT_MS, signal, includeBody: true }
)
}
async getUploadUrl(
request: Omit<ILinkGetUploadUrlRequest, 'base_info'>,
signal?: AbortSignal
): Promise<ILinkGetUploadUrlResponse> {
return this.post<ILinkGetUploadUrlResponse>(
'/ilink/bot/getuploadurl',
{ ...request, base_info: this.baseInfo() },
'getuploadurl',
{ timeoutMs: ILINK_SEND_TIMEOUT_MS, signal, includeBody: true }
)
}
async getConfig(
ilinkUserId: string,
contextToken: string,
signal?: AbortSignal
): Promise<ILinkGetConfigResponse> {
return this.post<ILinkGetConfigResponse>(
'/ilink/bot/getconfig',
{
ilink_user_id: ilinkUserId,
...(contextToken ? { context_token: contextToken } : {}),
base_info: this.baseInfo()
},
'getconfig',
{ timeoutMs: ILINK_CONFIG_TIMEOUT_MS, signal }
)
}
/**
* sendtyping:status=1 开始输入、status=2 取消。
* 只用于输入状态,**不是** sendmessage 的鉴权凭据。
*/
async sendTyping(
ilinkUserId: string,
typingTicket: string,
status: number,
signal?: AbortSignal
): Promise<ILinkSendTypingResponse> {
return this.post<ILinkSendTypingResponse>(
'/ilink/bot/sendtyping',
{
ilink_user_id: ilinkUserId,
typing_ticket: typingTicket,
status,
base_info: this.baseInfo()
},
'sendtyping',
{ timeoutMs: ILINK_CONFIG_TIMEOUT_MS, signal }
)
}
/** 生命周期通知:notifystart / notifystop。失败只告警,不阻断消息循环。 */
async notifyLifecycle(action: 'start' | 'stop', signal?: AbortSignal): Promise<void> {
await this.post<{ ret?: number; errmsg?: string }>(
`/ilink/bot/msg/notify${action}`,
{ base_info: this.baseInfo() },
`notify${action}`,
{ timeoutMs: ILINK_CONFIG_TIMEOUT_MS, signal }
)
}
}
+166
View File
@@ -0,0 +1,166 @@
import { ILINK_STALE_TOKEN_CODE } from './types'
export type ILinkErrorKind =
/** DNS / TCP / TLS 等网络层失败 */
| 'network'
/** 客户端主动超时(长轮询正常控制流) */
| 'timeout'
/** 非 2xx HTTP 响应 */
| 'http'
/** 响应体不是合法 JSON,或业务 ret/errcode 非 0 */
| 'protocol'
/** bot token 失效(ret/errcode = -14),必须重新登录 */
| 'stale_token'
/** 调用方主动取消 */
| 'aborted'
const MAX_BODY_EXCERPT = 400
export interface ILinkErrorOptions {
kind: ILinkErrorKind
message: string
httpStatus?: number
ret?: number
errcode?: number
errmsg?: string
cause?: unknown
}
/**
* iLink 统一错误类型。
*
* 刻意不携带请求头 / token:错误对象可能被上层原样写入日志,
* 因此只暴露可安全打印的字段。
*/
export class ILinkError extends Error {
readonly kind: ILinkErrorKind
readonly httpStatus?: number
readonly ret?: number
readonly errcode?: number
readonly errmsg?: string
override readonly cause?: unknown
constructor(options: ILinkErrorOptions) {
super(options.message)
this.name = 'ILinkError'
this.kind = options.kind
if (options.httpStatus !== undefined) this.httpStatus = options.httpStatus
if (options.ret !== undefined) this.ret = options.ret
if (options.errcode !== undefined) this.errcode = options.errcode
if (options.errmsg !== undefined) this.errmsg = options.errmsg
if (options.cause !== undefined) this.cause = options.cause
}
get isStaleToken(): boolean {
return (
this.kind === 'stale_token' ||
this.ret === ILINK_STALE_TOKEN_CODE ||
this.errcode === ILINK_STALE_TOKEN_CODE
)
}
}
export function isILinkError(error: unknown): error is ILinkError {
return error instanceof ILinkError
}
/**
* 把 undici / Node 的错误链压平成可诊断文本。
*
* fetch 层失败时顶层只有一句没有信息量的 `fetch failed`,真正的原因
* (ENOTFOUND / ECONNREFUSED / TLS 握手失败 / socket 超时)在 `error.cause` 链上。
* 日志里必须能看到它,否则线上只能看到"连不上"却不知道为什么。
*/
export function describeErrorChain(error: unknown, maxDepth = 4): string {
const parts: string[] = []
let current: unknown = error
for (let depth = 0; depth < maxDepth && current; depth += 1) {
const candidate = current as {
name?: unknown
code?: unknown
message?: unknown
cause?: unknown
}
const fragments = [candidate.name, candidate.code, candidate.message]
.map((value) => (typeof value === 'string' ? value.trim() : ''))
.filter(Boolean)
parts.push(fragments.join(' ') || String(current))
current = candidate.cause
}
return parts.join(' <- ')
}
/** 沿错误链取第一个有值的 errno 风格 code。 */
function firstErrorCode(error: unknown, maxDepth = 4): string {
let current: unknown = error
for (let depth = 0; depth < maxDepth && current; depth += 1) {
const code = (current as { code?: unknown } | null)?.code
if (typeof code === 'string' && code.trim()) return code.trim()
current = (current as { cause?: unknown } | null)?.cause
}
return ''
}
/** 网络层错误分类;用于决定退避策略。 */
export function classifyNetworkError(error: unknown): ILinkErrorKind {
const code = firstErrorCode(error)
const combined = `${code} ${describeErrorChain(error)}`.toUpperCase()
if (/ABORT|CANCEL/.test(combined)) return 'aborted'
if (
/TIMEOUT|ETIMEDOUT|UND_ERR_CONNECT_TIMEOUT|UND_ERR_HEADERS_TIMEOUT|UND_ERR_BODY_TIMEOUT/.test(
combined
)
) {
return 'timeout'
}
return 'network'
}
/** 截断响应体,避免把大段 HTML / 二进制写进日志。 */
export function excerptBody(body: string): string {
const normalized = String(body ?? '')
.replace(/\s+/g, ' ')
.trim()
return normalized.length > MAX_BODY_EXCERPT
? `${normalized.slice(0, MAX_BODY_EXCERPT)}…`
: normalized
}
/** 把任意 error 压成可安全记录的字符串。 */
export function describeError(error: unknown): string {
if (!isILinkError(error) && error instanceof Error) return describeErrorChain(error)
if (!isILinkError(error) && !(error instanceof Error)) return String(error)
if (isILinkError(error)) {
const parts = [error.message]
if (error.httpStatus !== undefined) parts.push(`http=${error.httpStatus}`)
if (error.ret !== undefined) parts.push(`ret=${error.ret}`)
if (error.errcode !== undefined) parts.push(`errcode=${error.errcode}`)
if (error.errmsg) parts.push(`errmsg=${error.errmsg}`)
return parts.join(' | ')
}
return error instanceof Error ? error.message : String(error)
}
/**
* 协议层业务错误断言:ret / errcode 非 0 即失败。
* HTTP 200 不能单独证明调用成功。
*/
export function assertBusinessOk(
response: { ret?: number; errcode?: number; errmsg?: string },
context: string
): void {
const ret = Number(response.ret ?? 0)
const errcode = Number(response.errcode ?? 0)
if (ret === 0 && errcode === 0) return
const stale = ret === ILINK_STALE_TOKEN_CODE || errcode === ILINK_STALE_TOKEN_CODE
throw new ILinkError({
kind: stale ? 'stale_token' : 'protocol',
message: stale
? `${context}失败:微信登录凭证已失效`
: `${context}失败:ret=${ret} errcode=${errcode}${response.errmsg ? ` errmsg=${response.errmsg}` : ''}`,
ret,
errcode,
...(response.errmsg ? { errmsg: response.errmsg } : {})
})
}
+84
View File
@@ -0,0 +1,84 @@
import { randomBytes } from 'node:crypto'
import {
ILINK_APP_CLIENT_VERSION,
ILINK_APP_ID,
ILINK_CHANNEL_VERSION,
ILINK_DEFAULT_BOT_AGENT,
type ILinkBaseInfo
} from './types'
/** bot_agent 的官方约束:仅 ASCII、总分不超过 256 字节、非法 token 丢弃。 */
const BOT_AGENT_MAX_BYTES = 256
const BOT_AGENT_FALLBACK = 'OpenClaw'
const BOT_AGENT_TOKEN = /^[!-~]+(?:\/[!-~]+)?(?:\([!-~ ]*\))?$/
export interface ILinkHeaderOptions {
/** 覆盖 bot_agent,用于多产品共用同一实现时做归因。仅进入 base_info,不额外造头。 */
botAgent?: string
}
/**
* 清洗 bot_agent。
* 该字段只用于服务端观测聚合,不参与鉴权与路由,因此宁可回退也不抛错。
*/
export function sanitizeBotAgent(value: string | undefined): string {
const tokens = String(value ?? '')
.split(/\s+/)
.map((token) => token.trim())
.filter((token) => token.length > 0 && BOT_AGENT_TOKEN.test(token))
let result = tokens.join(' ')
if (!result) result = BOT_AGENT_FALLBACK
while (Buffer.byteLength(result, 'utf8') > BOT_AGENT_MAX_BYTES && result.includes(' ')) {
result = result.slice(0, result.lastIndexOf(' '))
}
if (Buffer.byteLength(result, 'utf8') > BOT_AGENT_MAX_BYTES) result = BOT_AGENT_FALLBACK
return result
}
/**
* X-WECHAT-UIN:随机 uint32 的十进制字符串再做 base64。
* 官方要求**每次请求重新生成**以防重放,一个客户端只生成一次是不合规的。
*/
export function generateWechatUin(): string {
const value = randomBytes(4).readUInt32LE(0)
return Buffer.from(String(value), 'utf8').toString('base64')
}
/** 所有 POST 请求都会带的公共头(含登录前的二维码接口)。 */
export function buildCommonHeaders(): Record<string, string> {
return {
'Content-Type': 'application/json',
'iLink-App-Id': ILINK_APP_ID,
'iLink-App-ClientVersion': ILINK_APP_CLIENT_VERSION
}
}
/**
* 二维码状态轮询使用未鉴权的最小头集合(官方 2.4.6 不再带 Content-Type)。
* 注意:不自造 SKRouteTag,那是官方内部的路由/调试开关。
*/
export function buildQrStatusHeaders(): Record<string, string> {
return {
'iLink-App-Id': ILINK_APP_ID,
'iLink-App-ClientVersion': ILINK_APP_CLIENT_VERSION
}
}
/** 登录后的业务请求头;未登录时不得携带 Authorization。 */
export function buildAuthorizedHeaders(botToken: string): Record<string, string> {
if (!botToken) throw new Error('buildAuthorizedHeaders 需要有效的 bot token')
return {
...buildCommonHeaders(),
AuthorizationType: 'ilink_bot_token',
Authorization: `Bearer ${botToken}`,
'X-WECHAT-UIN': generateWechatUin()
}
}
/** 每个业务请求体都要带 base_info。 */
export function buildBaseInfo(options: ILinkHeaderOptions = {}): ILinkBaseInfo {
return {
channel_version: ILINK_CHANNEL_VERSION,
bot_agent: sanitizeBotAgent(options.botAgent ?? ILINK_DEFAULT_BOT_AGENT)
}
}
+507
View File
@@ -0,0 +1,507 @@
import { ILinkClient, type FetchLike } from './client'
import { ILinkPoller } from './poller'
import { normalizeInboundMessage } from './messages'
import { sendText, type SendTextResult } from './sender'
import {
sendMediaFromPath,
sendMediaFromSource,
sendMediaFromUrl,
type SendMediaOptions
} from './media'
import { runQrLogin, type QrEncoder } from './auth'
import { assertBusinessOk, ILinkError, describeError } from './errors'
import { TypingCoordinator, type TypingLease, type TypingBeginInput } from './typing'
import {
findCredentials,
loadAllCredentials,
loadContextToken,
loadCursor,
normalizeAccountId,
resolveAccountDirectory,
saveContextToken,
saveCursor,
type HomeDirectoryResolver
} from './account-store'
import type {
ILinkCredentials,
WechatConnectorAccount,
WechatConnectorPhase,
WechatInboundMessage,
WechatLoginEvent
} from './types'
export type ConnectorLogLevel = 'info' | 'warn' | 'error'
export interface WechatConnectorHost {
onLog?: (level: ConnectorLogLevel, message: string) => void
/** 一批入站消息;全部成功 resolve 后才推进游标。 */
onMessages?: (messages: WechatInboundMessage[]) => Promise<void>
onPhaseChange?: (phase: WechatConnectorPhase, error?: string) => void
onLoginEvent?: (event: WechatLoginEvent) => void
}
export interface WechatConnectorServiceOptions {
home?: HomeDirectoryResolver
fetchImpl?: FetchLike
qrEncoder?: QrEncoder
now?: () => number
sleep?: (milliseconds: number) => Promise<void>
botAgent?: string
/** 「正在输入」心跳间隔;默认 5 秒(测试可缩短)。 */
typingKeepaliveMs?: number
/** typing_ticket 缓存有效期。 */
typingTicketTtlMs?: number
/** 定时器注入点,便于测试确定性触发 keepalive。 */
typingSchedule?: (tick: () => void, intervalMs: number) => () => void
/** 便于测试注入更短的长轮询超时/退避。 */
pollOverrides?: {
normalFailureDelayMs?: number
repeatedFailureDelayMs?: number
defaultTimeoutMs?: number
}
}
export interface ConnectorSendTextInput {
to: string
text: string
accountId?: string
contextToken?: string
signal?: AbortSignal
}
/**
* TraceMemo 的微信 iLink 连接器,直接跑在 Electron main process 内。
*
* inbound 通过 onMessages 回调交给 Agent Hub,outbound 通过 sendText / sendMedia 直接调用;
* 不依赖子进程,也不开本地 HTTP 端口。
*/
export class WechatConnectorService {
private readonly options: WechatConnectorServiceOptions
private readonly typing: TypingCoordinator
private host: WechatConnectorHost = {}
private phase: WechatConnectorPhase = 'stopped'
private activeAccountId?: string
private client: ILinkClient | null = null
private pollerAbort: AbortController | null = null
private pollerPromise: Promise<void> | null = null
private loginAbort: AbortController | null = null
private loginPromise: Promise<void> | null = null
private pendingVerifyCode: ((code: string | undefined) => void) | null = null
constructor(options: WechatConnectorServiceOptions = {}) {
this.options = options
this.typing = new TypingCoordinator({
// 取票与下发都绑定"当前账号"的 client,切账号后自然走新 client。
fetchTicket: async ({ ilinkUserId, contextToken }) => {
const client = this.client
if (!client) return undefined
const response = await client.getConfig(ilinkUserId, contextToken ?? '')
assertBusinessOk(response, 'getconfig')
return response.typing_ticket
},
sendTyping: async ({ ilinkUserId, ticket, status }) => {
const client = this.client
if (!client) return false
const response = await client.sendTyping(ilinkUserId, ticket, status)
return Number(response.ret ?? 0) === 0 && Number(response.errcode ?? 0) === 0
},
log: (level, message) => this.log(level, message),
...(options.now ? { now: options.now } : {}),
...(options.typingKeepaliveMs !== undefined
? { keepaliveMs: options.typingKeepaliveMs }
: {}),
...(options.typingTicketTtlMs !== undefined
? { ticketTtlMs: options.typingTicketTtlMs }
: {}),
...(options.typingSchedule ? { schedule: options.typingSchedule } : {})
})
}
/**
* 开始「正在输入」。永不抛异常:拿不到 ticket 或服务端失败时返回空实现,
* 调用方照常执行业务,但**仍然要**在 finally 里 stop(幂等且安全)。
*/
beginTyping(input: TypingBeginInput): Promise<TypingLease> {
return this.typing.begin({
...input,
...(input.accountId ? {} : this.activeAccountId ? { accountId: this.activeAccountId } : {})
})
}
setHost(host: WechatConnectorHost): void {
this.host = host
}
getPhase(): WechatConnectorPhase {
return this.phase
}
getActiveAccountId(): string | undefined {
return this.activeAccountId
}
isRunning(): boolean {
return this.phase === 'polling'
}
listAccounts(): WechatConnectorAccount[] {
return loadAllCredentials(this.options.home).map((credentials) => ({
accountId: credentials.ilink_bot_id,
wechatUserId: credentials.ilink_user_id
}))
}
/** 主动发送时取该会话最近一次有效的 context_token。 */
resolveContextToken(accountId: string | undefined, toUserId: string): string | undefined {
const normalized = normalizeAccountId(accountId || this.activeAccountId || '')
if (!normalized) return undefined
return loadContextToken(normalized, toUserId, this.options.home)
}
/** 当前账号凭据所在目录(凭据 / 游标 / 会话令牌同目录)。 */
getAccountDirectory(accountId: string): { directory: string; legacy: boolean } {
return resolveAccountDirectory(normalizeAccountId(accountId), this.options.home)
}
/* ------------------------------ 登录 ------------------------------ */
/**
* 启动扫码登录。立即返回,进度通过 onLoginEvent 上报。
* 重复调用会先取消进行中的登录。
*/
startLogin(): void {
if (this.loginPromise) return
const controller = new AbortController()
this.loginAbort = controller
this.loginPromise = this.runLogin(controller).finally(() => {
if (this.loginAbort === controller) this.loginAbort = null
this.loginPromise = null
})
}
isLoginInProgress(): boolean {
return this.loginPromise !== null
}
cancelLogin(): void {
this.pendingVerifyCode?.(undefined)
this.pendingVerifyCode = null
this.loginAbort?.abort()
this.loginAbort = null
}
/** 手机端要求数字配对码时由宿主回填。 */
submitVerifyCode(code: string): void {
const resolver = this.pendingVerifyCode
this.pendingVerifyCode = null
resolver?.(code)
}
private async runLogin(controller: AbortController): Promise<void> {
this.setPhase('starting')
try {
await runQrLogin({
...(this.options.home ? { home: this.options.home } : {}),
...(this.options.fetchImpl ? { fetchImpl: this.options.fetchImpl } : {}),
...(this.options.qrEncoder ? { qrEncoder: this.options.qrEncoder } : {}),
...(this.options.now ? { now: this.options.now } : {}),
...(this.options.sleep ? { sleep: this.options.sleep } : {}),
signal: controller.signal,
onEvent: (event) => this.host.onLoginEvent?.(event),
verifyCodeProvider: () =>
new Promise<string | undefined>((resolve) => {
this.pendingVerifyCode = resolve
})
})
this.log('info', '扫码登录成功')
// 登录成功后的连接由宿主在收到 active 事件后调用 start()。
this.setPhase('stopped')
} catch (error) {
if (error instanceof ILinkError && error.kind === 'aborted') {
this.log('info', '扫码登录已取消')
this.setPhase('stopped')
return
}
const message = describeError(error)
this.log('error', `扫码登录失败:${message}`)
this.setPhase('error', message)
}
}
/* --------------------------- 连接与轮询 --------------------------- */
/** 启动某个账号的长轮询。重复调用会先停止当前账号。 */
async start(accountId: string): Promise<void> {
await this.stop()
const credentials = findCredentials(normalizeAccountId(accountId), this.options.home)
if (!credentials) {
throw new ILinkError({ kind: 'protocol', message: `未找到账号 ${accountId} 的登录凭据` })
}
const normalizedId = normalizeAccountId(credentials.ilink_bot_id)
const client = new ILinkClient({
baseUrl: credentials.baseurl,
botToken: credentials.bot_token,
...(this.options.fetchImpl ? { fetchImpl: this.options.fetchImpl } : {}),
...(this.options.botAgent ? { headers: { botAgent: this.options.botAgent } } : {})
})
this.client = client
this.activeAccountId = normalizedId
// 账号或 token 变了:旧 ticket 一律作废。
this.typing.invalidateTickets()
this.setPhase('starting')
// 生命周期通知是尽力而为:失败只告警,绝不阻断消息循环。
await this.notifyLifecycle('start')
const controller = new AbortController()
this.pollerAbort = controller
const poller = new ILinkPoller({
fetchUpdates: (getUpdatesBuf, timeoutMs) =>
client.getUpdates(getUpdatesBuf, { timeoutMs, signal: controller.signal }),
onMessages: async (messages) => {
this.rememberContextTokens(normalizedId, messages)
await this.host.onMessages?.(messages)
},
normalize: (raw) => normalizeInboundMessage(normalizedId, raw),
loadCursor: () => loadCursor(normalizedId, this.options.home),
saveCursor: (getUpdatesBuf) => saveCursor(normalizedId, getUpdatesBuf, this.options.home),
signal: controller.signal,
log: (level, message) => this.log(level, message),
onStaleToken: (message) => {
// 凭证已失效:缓存的 typing_ticket 不再可用。
this.typing.invalidateTickets()
this.setPhase('stale_token', message)
},
...(this.options.now ? { now: this.options.now } : {}),
...(this.options.sleep ? { sleep: this.options.sleep } : {}),
...(this.options.pollOverrides?.normalFailureDelayMs !== undefined
? { normalFailureDelayMs: this.options.pollOverrides.normalFailureDelayMs }
: {}),
...(this.options.pollOverrides?.repeatedFailureDelayMs !== undefined
? { repeatedFailureDelayMs: this.options.pollOverrides.repeatedFailureDelayMs }
: {}),
...(this.options.pollOverrides?.defaultTimeoutMs !== undefined
? { defaultTimeoutMs: this.options.pollOverrides.defaultTimeoutMs }
: {})
})
const pollPromise = poller
.run()
.catch((error) => {
if (controller.signal.aborted) return
const message = describeError(error)
this.log('error', `长轮询异常退出:${message}`)
this.setPhase('error', message)
})
.finally(() => {
if (this.pollerAbort === controller) {
this.pollerAbort = null
this.pollerPromise = null
}
})
this.pollerPromise = pollPromise
if (this.phase !== 'stale_token' && this.phase !== 'error') this.setPhase('polling')
this.log('info', `微信连接器已启动(账号 ${normalizedId})`)
}
async stop(): Promise<void> {
// 先收掉输入状态,避免连接器停了微信端还显示"对方正在输入"。
this.typing.clear()
const controller = this.pollerAbort
const running = this.pollerPromise
this.pollerAbort = null
if (controller && !controller.signal.aborted) controller.abort()
if (running) await running.catch(() => undefined)
if (this.client) {
// 用独立短超时发送停止通知,避免被长轮询的取消信号一起取消。
await this.notifyLifecycle('stop')
}
this.client = null
this.activeAccountId = undefined
if (this.phase !== 'error') this.setPhase('stopped')
}
private async notifyLifecycle(action: 'start' | 'stop'): Promise<void> {
const client = this.client
if (!client || !client.hasBotToken) return
try {
await client.notifyLifecycle(action)
} catch (error) {
this.log('warn', `notify${action} 未成功(不影响消息循环):${describeError(error)}`)
}
}
private rememberContextTokens(accountId: string, messages: WechatInboundMessage[]): void {
for (const message of messages) {
if (!message.contextToken) continue
try {
saveContextToken(
accountId,
message.fromUserId,
message.contextToken,
this.options.home,
this.options.now
)
} catch (error) {
this.log('warn', `会话上下文令牌保存失败:${describeError(error)}`)
}
}
}
/* ------------------------------ 发送 ------------------------------ */
/**
* 主动发送时的 context_token 解析:
* 显式传入优先,其次按「账号 + 用户」取最近一次有效值。
*/
resolveOutgoingContextToken(input: {
accountId?: string
to: string
contextToken?: string
}): string | undefined {
const explicit = String(input.contextToken ?? '').trim()
if (explicit) return explicit
return this.resolveContextToken(input.accountId, input.to)
}
private requireClient(accountId?: string): ILinkClient {
if (this.phase === 'stale_token') {
throw new ILinkError({
kind: 'stale_token',
message: '微信登录凭证已失效,请重新扫码登录'
})
}
if (!this.client) {
throw new ILinkError({ kind: 'protocol', message: '微信连接器尚未启动' })
}
if (accountId) {
const normalized = normalizeAccountId(accountId)
if (this.activeAccountId && normalized !== this.activeAccountId) {
throw new ILinkError({
kind: 'protocol',
message: `账号 ${accountId} 当前未连接`
})
}
}
if (!this.client.hasBotToken) {
throw new ILinkError({ kind: 'stale_token', message: '微信登录凭证不可用' })
}
return this.client
}
async sendText(input: ConnectorSendTextInput): Promise<SendTextResult> {
const client = this.requireClient(input.accountId)
const contextToken = this.resolveOutgoingContextToken({
...(input.accountId ? { accountId: input.accountId } : {}),
to: input.to,
...(input.contextToken ? { contextToken: input.contextToken } : {})
})
return sendText(client, {
to: input.to,
text: input.text,
...(contextToken ? { contextToken } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
async sendMedia(
input: Omit<SendMediaOptions, 'fetchImpl'> & { accountId?: string }
): Promise<{ clientId: string; itemType: number }> {
const client = this.requireClient(input.accountId)
const contextToken = this.resolveOutgoingContextToken({
...(input.accountId ? { accountId: input.accountId } : {}),
to: input.to,
...(input.contextToken ? { contextToken: input.contextToken } : {})
})
return sendMediaFromSource(client, {
to: input.to,
source: input.source,
mediaKind: 'file',
...(contextToken ? { contextToken } : {}),
...(this.options.fetchImpl ? { fetchImpl: this.options.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
async sendMediaPath(input: {
accountId?: string
to: string
filePath: string
contextToken?: string
signal?: AbortSignal
}): Promise<{ clientId: string; itemType: number }> {
const client = this.requireClient(input.accountId)
const contextToken = this.resolveOutgoingContextToken({
...(input.accountId ? { accountId: input.accountId } : {}),
to: input.to,
...(input.contextToken ? { contextToken: input.contextToken } : {})
})
return sendMediaFromPath(client, {
to: input.to,
filePath: input.filePath,
...(contextToken ? { contextToken } : {}),
...(this.options.fetchImpl ? { fetchImpl: this.options.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
async sendMediaUrl(input: {
accountId?: string
to: string
mediaUrl: string
contextToken?: string
signal?: AbortSignal
}): Promise<{ clientId: string; itemType: number }> {
const client = this.requireClient(input.accountId)
const contextToken = this.resolveOutgoingContextToken({
...(input.accountId ? { accountId: input.accountId } : {}),
to: input.to,
...(input.contextToken ? { contextToken: input.contextToken } : {})
})
return sendMediaFromUrl(client, {
to: input.to,
mediaUrl: input.mediaUrl,
...(contextToken ? { contextToken } : {}),
...(this.options.fetchImpl ? { fetchImpl: this.options.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
/** 会话凭据(含 bot_token)只读暴露给需要的内部调用方;不用于日志。 */
getCredentials(accountId?: string): ILinkCredentials | undefined {
const id = normalizeAccountId(accountId || this.activeAccountId || '')
if (!id) return undefined
return findCredentials(id, this.options.home)
}
private setPhase(phase: WechatConnectorPhase, error?: string): void {
this.phase = phase
this.host.onPhaseChange?.(phase, error)
}
private log(level: ConnectorLogLevel, message: string): void {
this.host.onLog?.(level, message)
}
}
export { ILinkClient } from './client'
export { ILinkPoller } from './poller'
export * from './types'
export { normalizeInboundMessage, extractInboundText } from './messages'
export { markdownToPlainText, extractMarkdownImageUrls } from './markdown'
export { ILinkError, describeError, isILinkError } from './errors'
export {
loadAllCredentials,
normalizeAccountId,
saveCredentials,
accountsDirectory,
legacyAccountsDirectory,
loadCursor,
saveCursor,
loadContextToken,
saveContextToken
} from './account-store'
export { sendText, createClientId } from './sender'
export { runQrLogin } from './auth'
@@ -0,0 +1,58 @@
/**
* 微信纯文本渲染。
*
* 微信气泡不渲染 Markdown,所以外发前把常见 Markdown 语法降级成可读纯文本。
* 刻意不处理斜体(`*text*`):微信里 `*` 常作为普通字符出现,转换会误伤正文。
*/
const RE_CODE_BLOCK = /```[^\n]*\n?([\s\S]*?)```/g
const RE_INLINE_CODE = /`([^`]+)`/g
const RE_IMAGE = /!\[[^\]]*\]\([^)]*\)/g
const RE_LINK = /\[([^\]]+)\]\([^)]*\)/g
const RE_TABLE_SEPARATOR = /^\|[\s:|-]+\|$/gm
const RE_TABLE_ROW = /^\|(.+)\|$/gm
const RE_HEADER = /^#{1,6}\s+/gm
const RE_BOLD = /\*\*(.+?)\*\*|__(.+?)__/g
const RE_STRIKE = /~~(.+?)~~/g
const RE_BLOCKQUOTE = /^>\s?/gm
const RE_HORIZONTAL_RULE = /^[-*_]{3,}\s*$/gm
const RE_UNORDERED_LIST = /^(\s*)[-*+]\s+/gm
const RE_EXCESS_BLANK_LINES = /\n{3,}/g
const RE_MARKDOWN_IMAGE_URL = /!\[[^\]]*\]\(([^)]+)\)/g
export function markdownToPlainText(text: string): string {
let result = String(text ?? '')
result = result.replace(RE_CODE_BLOCK, (_match, body: string) => String(body ?? '').trim())
result = result.replace(RE_IMAGE, '')
result = result.replace(RE_LINK, '$1')
result = result.replace(RE_TABLE_SEPARATOR, '')
result = result.replace(RE_TABLE_ROW, (_match, body: string) =>
String(body ?? '')
.split('|')
.map((cell) => cell.trim())
.join(' ')
)
result = result.replace(RE_HEADER, '')
result = result.replace(RE_BOLD, (_match, strong: string, alternative: string) =>
strong !== undefined && strong !== '' ? strong : String(alternative ?? '')
)
result = result.replace(RE_STRIKE, '$1')
result = result.replace(RE_BLOCKQUOTE, '')
result = result.replace(RE_HORIZONTAL_RULE, '')
result = result.replace(RE_UNORDERED_LIST, '$1• ')
result = result.replace(RE_INLINE_CODE, '$1')
result = result.replace(RE_EXCESS_BLANK_LINES, '\n\n')
return result.trim()
}
/** 提取 Markdown 中内嵌的 http(s) 图片地址,用于"文字 + 随后补发图片"。 */
export function extractMarkdownImageUrls(text: string): string[] {
const urls: string[] = []
for (const match of String(text ?? '').matchAll(RE_MARKDOWN_IMAGE_URL)) {
const url = String(match[1] ?? '').trim()
if (url.startsWith('http://') || url.startsWith('https://')) urls.push(url)
}
return urls
}
+391
View File
@@ -0,0 +1,391 @@
import { createCipheriv, createDecipheriv, createHash, randomBytes } from 'node:crypto'
import { readFileSync } from 'node:fs'
import { basename, extname } from 'node:path'
import type { FetchLike, ILinkClient } from './client'
import { assertBusinessOk, excerptBody, ILinkError } from './errors'
import { buildSendMessageBody, createClientId } from './sender'
import {
ILINK_CDN_BASE_URL,
ILINK_CDN_MEDIA_TYPE_FILE,
ILINK_CDN_MEDIA_TYPE_IMAGE,
ILINK_CDN_MEDIA_TYPE_VIDEO,
ILINK_ITEM_TYPE_FILE,
ILINK_ITEM_TYPE_IMAGE,
ILINK_ITEM_TYPE_VIDEO,
type ILinkMessageItem
} from './types'
/**
* 媒体消息(图片 / 视频 / 文件)。
*
* 协议要求端到端加密:随机 16 字节 AES-128-ECB 密钥 → 加密后 PUT 到 CDN →
* sendmessage 里带 CDN 引用与 base64(hex key)。
*
* 本项目不生成缩略图:官方文档提到 IMAGE/VIDEO 可携带缩略图字段,
* 这里显式声明 `no_need_thumb: true`,避免服务端等待一个永远不会上传的缩略图。
*/
const CDN_REQUEST_TIMEOUT_MS = 60_000
const AES_BLOCK_SIZE = 16
const MIME_BY_EXTENSION: Readonly<Record<string, string>> = {
'.png': 'image/png',
'.jpg': 'image/jpeg',
'.jpeg': 'image/jpeg',
'.gif': 'image/gif',
'.webp': 'image/webp',
'.bmp': 'image/bmp',
'.mp4': 'video/mp4',
'.mov': 'video/quicktime',
'.webm': 'video/webm',
'.mkv': 'video/x-matroska',
'.avi': 'video/x-msvideo',
'.pdf': 'application/pdf',
'.txt': 'text/plain',
'.zip': 'application/zip'
}
const IMAGE_EXTENSIONS = new Set(['.png', '.jpg', '.jpeg', '.gif', '.webp', '.bmp'])
const VIDEO_EXTENSIONS = new Set(['.mp4', '.mov', '.webm', '.mkv', '.avi'])
export function aesEcbPaddedSize(plaintextSize: number): number {
return (Math.floor(plaintextSize / AES_BLOCK_SIZE) + 1) * AES_BLOCK_SIZE
}
export function encryptAesEcb(plaintext: Buffer, key: Buffer): Buffer {
const cipher = createCipheriv('aes-128-ecb', key, null)
return Buffer.concat([cipher.update(plaintext), cipher.final()])
}
export function decryptAesEcb(ciphertext: Buffer, key: Buffer): Buffer {
if (ciphertext.length === 0 || ciphertext.length % AES_BLOCK_SIZE !== 0) {
throw new ILinkError({ kind: 'protocol', message: '密文长度不是 AES 块大小的整数倍' })
}
const decipher = createDecipheriv('aes-128-ecb', key, null)
return Buffer.concat([decipher.update(ciphertext), decipher.final()])
}
/** 协议要求把 hex 字符串再做 base64 放进 item.media.aes_key。 */
export function aesKeyToBase64(hexKey: string): string {
return Buffer.from(hexKey, 'utf8').toString('base64')
}
/** 反向解析 item.media.aes_key:base64 → hex 字符串 → 原始密钥。 */
export function aesKeyFromBase64(base64Key: string): Buffer {
return Buffer.from(Buffer.from(base64Key, 'base64').toString('utf8'), 'hex')
}
export function stripQuery(rawUrl: string): string {
const index = rawUrl.indexOf('?')
return index >= 0 ? rawUrl.slice(0, index) : rawUrl
}
export function inferContentType(source: string): string {
return MIME_BY_EXTENSION[extname(stripQuery(source)).toLowerCase()] || 'application/octet-stream'
}
export function classifyMedia(
contentType: string,
source: string
): { cdnMediaType: number; itemType: number } {
const normalized = String(contentType || '').toLowerCase()
const extension = extname(stripQuery(source)).toLowerCase()
if (normalized.startsWith('image/') || IMAGE_EXTENSIONS.has(extension)) {
return { cdnMediaType: ILINK_CDN_MEDIA_TYPE_IMAGE, itemType: ILINK_ITEM_TYPE_IMAGE }
}
if (normalized.startsWith('video/') || VIDEO_EXTENSIONS.has(extension)) {
return { cdnMediaType: ILINK_CDN_MEDIA_TYPE_VIDEO, itemType: ILINK_ITEM_TYPE_VIDEO }
}
return { cdnMediaType: ILINK_CDN_MEDIA_TYPE_FILE, itemType: ILINK_ITEM_TYPE_FILE }
}
export interface UploadedMedia {
downloadParam: string
aesKeyHex: string
fileSize: number
cipherSize: number
}
async function fetchWithTimeout(
fetchImpl: FetchLike,
url: string,
init: RequestInit,
context: string
): Promise<Response> {
const controller = new AbortController()
const timer = setTimeout(() => controller.abort(new Error('cdn timeout')), CDN_REQUEST_TIMEOUT_MS)
try {
const response = await fetchImpl(url, { ...init, signal: controller.signal })
if (!response.ok) {
const body = await response.text().catch(() => '')
throw new ILinkError({
kind: 'http',
message: `${context}失败:HTTP ${response.status} ${excerptBody(body)}`,
httpStatus: response.status
})
}
return response
} catch (error) {
if (error instanceof ILinkError) throw error
throw new ILinkError({
kind: 'network',
message: `${context}失败:${error instanceof Error ? error.message : String(error)}`,
cause: error
})
} finally {
clearTimeout(timer)
}
}
/** 加密并上传到微信 CDN,返回 sendmessage 需要的媒体引用。 */
export async function uploadMediaToCdn(
client: ILinkClient,
input: {
data: Buffer
toUserId: string
mediaType: number
fetchImpl?: FetchLike
signal?: AbortSignal
}
): Promise<UploadedMedia> {
const fetchImpl = input.fetchImpl ?? globalThis.fetch
const fileKey = randomBytes(16)
const aesKey = randomBytes(16)
const fileKeyHex = fileKey.toString('hex')
const aesKeyHex = aesKey.toString('hex')
const rawFileMd5 = createHash('md5').update(input.data).digest('hex')
const cipherSize = aesEcbPaddedSize(input.data.length)
const uploadResponse = await client.getUploadUrl(
{
filekey: fileKeyHex,
media_type: input.mediaType,
to_user_id: input.toUserId,
rawsize: input.data.length,
rawfilemd5: rawFileMd5,
filesize: cipherSize,
no_need_thumb: true,
aeskey: aesKeyHex
},
input.signal
)
assertBusinessOk(uploadResponse, '获取 CDN 上传地址')
const encrypted = encryptAesEcb(input.data, aesKey)
const explicitUrl = String(uploadResponse.upload_full_url ?? '').trim()
const uploadParam = String(uploadResponse.upload_param ?? '').trim()
if (!explicitUrl && !uploadParam) {
throw new ILinkError({
kind: 'protocol',
message: '服务端未返回可用的 CDN 上传地址'
})
}
const cdnUrl =
explicitUrl ||
`${ILINK_CDN_BASE_URL}/upload?encrypted_query_param=${encodeURIComponent(uploadParam)}&filekey=${encodeURIComponent(fileKeyHex)}`
const uploadResult = await fetchWithTimeout(
fetchImpl,
cdnUrl,
{
method: 'POST',
headers: { 'Content-Type': 'application/octet-stream' },
body: new Uint8Array(encrypted)
},
'CDN 上传'
)
const downloadParam = uploadResult.headers.get('X-Encrypted-Param') ?? ''
if (!downloadParam) {
throw new ILinkError({
kind: 'protocol',
message: 'CDN 上传成功但缺少 X-Encrypted-Param 响应头'
})
}
return { downloadParam, aesKeyHex, fileSize: input.data.length, cipherSize }
}
async function downloadFromCdn(
fetchImpl: FetchLike,
encryptQueryParam: string,
aesKeyBase64: string
): Promise<Buffer> {
const url = `${ILINK_CDN_BASE_URL}/download?encrypted_query_param=${encodeURIComponent(encryptQueryParam)}`
const response = await fetchWithTimeout(fetchImpl, url, { method: 'GET' }, 'CDN 下载')
const ciphertext = Buffer.from(await response.arrayBuffer())
return decryptAesEcb(ciphertext, aesKeyFromBase64(aesKeyBase64))
}
function buildMediaItem(
itemType: number,
media: { encrypt_query_param: string; aes_key: string; encrypt_type: number },
uploaded: UploadedMedia,
fileName: string
): ILinkMessageItem {
if (itemType === ILINK_ITEM_TYPE_IMAGE) {
return { type: ILINK_ITEM_TYPE_IMAGE, image_item: { media, mid_size: uploaded.cipherSize } }
}
if (itemType === ILINK_ITEM_TYPE_VIDEO) {
return { type: ILINK_ITEM_TYPE_VIDEO, video_item: { media, video_size: uploaded.cipherSize } }
}
return {
type: ILINK_ITEM_TYPE_FILE,
file_item: {
media,
file_name: fileName || 'file',
len: String(uploaded.fileSize)
}
}
}
export interface SendMediaOptions {
to: string
data: Buffer
fileName: string
contentType: string
source: string
contextToken?: string
fetchImpl?: FetchLike
signal?: AbortSignal
}
/** 上传并发送一条媒体消息。 */
export async function sendMedia(
client: ILinkClient,
options: SendMediaOptions
): Promise<{ clientId: string; itemType: number }> {
const to = String(options.to ?? '').trim()
if (!to) throw new ILinkError({ kind: 'protocol', message: '发送媒体需要有效的接收者' })
const { cdnMediaType, itemType } = classifyMedia(options.contentType, options.source)
const uploaded = await uploadMediaToCdn(client, {
data: options.data,
toUserId: to,
mediaType: cdnMediaType,
...(options.fetchImpl ? { fetchImpl: options.fetchImpl } : {}),
...(options.signal ? { signal: options.signal } : {})
})
const clientId = createClientId()
const response = await client.sendMessage(
buildSendMessageBody({
item: buildMediaItem(
itemType,
{
encrypt_query_param: uploaded.downloadParam,
aes_key: aesKeyToBase64(uploaded.aesKeyHex),
encrypt_type: 1
},
uploaded,
options.fileName
),
to,
clientId,
...(options.contextToken ? { contextToken: options.contextToken } : {})
}),
options.signal
)
assertBusinessOk(response, '发送媒体')
return { clientId, itemType }
}
/** 发送本地文件。 */
export async function sendMediaFromPath(
client: ILinkClient,
input: {
to: string
filePath: string
contextToken?: string
fetchImpl?: FetchLike
signal?: AbortSignal
}
): Promise<{ clientId: string; itemType: number }> {
let data: Buffer
try {
data = readFileSync(input.filePath)
} catch (error) {
throw new ILinkError({
kind: 'protocol',
message: `读取媒体文件失败:${error instanceof Error ? error.message : String(error)}`
})
}
return sendMedia(client, {
to: input.to,
data,
fileName: basename(input.filePath),
contentType: inferContentType(input.filePath),
source: input.filePath,
...(input.contextToken ? { contextToken: input.contextToken } : {}),
...(input.fetchImpl ? { fetchImpl: input.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
/** 发送远程媒体:下载后走同一条加密上传链路。 */
export async function sendMediaFromUrl(
client: ILinkClient,
input: {
to: string
mediaUrl: string
contextToken?: string
fetchImpl?: FetchLike
signal?: AbortSignal
}
): Promise<{ clientId: string; itemType: number }> {
if (!/^https?:\/\//i.test(input.mediaUrl)) {
throw new ILinkError({ kind: 'protocol', message: '远程媒体必须是 http(s) 地址' })
}
const fetchImpl = input.fetchImpl ?? globalThis.fetch
const response = await fetchWithTimeout(
fetchImpl,
input.mediaUrl,
{ method: 'GET' },
'下载远程媒体'
)
const data = Buffer.from(await response.arrayBuffer())
const headerType = response.headers.get('content-type') ?? ''
const name = basename(stripQuery(input.mediaUrl)) || 'file'
return sendMedia(client, {
to: input.to,
data,
fileName: name,
contentType: headerType || inferContentType(input.mediaUrl),
source: input.mediaUrl,
...(input.contextToken ? { contextToken: input.contextToken } : {}),
...(input.fetchImpl ? { fetchImpl: input.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
/** 发送媒体来源(本地路径优先,其次 http(s) URL)。供统一发送入口复用。 */
export async function sendMediaFromSource(
client: ILinkClient,
input: {
to: string
source: string
mediaKind: 'image' | 'file'
contextToken?: string
fetchImpl?: FetchLike
signal?: AbortSignal
}
): Promise<{ clientId: string; itemType: number }> {
const isLocal = !/^https?:\/\//i.test(input.source)
if (isLocal) {
return sendMediaFromPath(client, {
to: input.to,
filePath: input.source,
...(input.contextToken ? { contextToken: input.contextToken } : {}),
...(input.fetchImpl ? { fetchImpl: input.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
return sendMediaFromUrl(client, {
to: input.to,
mediaUrl: input.source,
...(input.contextToken ? { contextToken: input.contextToken } : {}),
...(input.fetchImpl ? { fetchImpl: input.fetchImpl } : {}),
...(input.signal ? { signal: input.signal } : {})
})
}
export { downloadFromCdn }
@@ -0,0 +1,61 @@
import {
ILINK_ITEM_TYPE_TEXT,
type ILinkMessageItem,
type ILinkWeixinMessage,
type WechatInboundItem,
type WechatInboundMessage
} from './types'
/**
* 把协议原始消息归一化成 Agent Hub 消费的形状。
*
* 关键点是 `context_token` **必须完整透传**:它属于会话上下文,
* 回复同一个会话时要原样回传,缺失会导致回复落到错误的会话线程。
*/
export function normalizeInboundMessage(
accountId: string,
raw: ILinkWeixinMessage
): WechatInboundMessage | undefined {
const fromUserId = String(raw.from_user_id ?? '').trim()
if (!fromUserId) return undefined
const items: WechatInboundItem[] = (raw.item_list ?? []).map((item) => {
const normalized: WechatInboundItem = { type: Number(item.type ?? 0) }
const text = itemText(item)
if (text) normalized.text = text
return normalized
})
const message: WechatInboundMessage = {
accountId,
fromUserId,
messageId:
raw.message_id !== undefined && raw.message_id !== null ? String(raw.message_id) : '',
messageType: Number(raw.message_type ?? 0),
items,
receivedAt: Date.now()
}
if (raw.seq !== undefined && raw.seq !== null) message.seq = Number(raw.seq)
const sessionId = String(raw.session_id ?? '').trim()
if (sessionId) message.sessionId = sessionId
const groupId = String(raw.group_id ?? '').trim()
if (groupId) message.groupId = groupId
const contextToken = String(raw.context_token ?? '').trim()
if (contextToken) message.contextToken = contextToken
return message
}
function itemText(item: ILinkMessageItem): string {
const direct = String(item.text_item?.text ?? '').trim()
if (direct) return direct
// 语音条目的 text 是微信侧语音转文字结果,保留以便上层按需使用。
return String(item.voice_item?.text ?? '').trim()
}
/** 拼接入站消息中的文本片段:只有类型为「文本」的条目参与拼接。 */
export function extractInboundText(items: WechatInboundItem[] | undefined): string {
return (items ?? [])
.filter((item) => item.type === ILINK_ITEM_TYPE_TEXT && item.text?.trim())
.map((item) => item.text!.trim())
.join(' ')
}
+187
View File
@@ -0,0 +1,187 @@
import { describeError, ILinkError } from './errors'
import {
ILINK_LONG_POLL_MAX_TIMEOUT_MS,
ILINK_LONG_POLL_TIMEOUT_MS,
type ILinkGetUpdatesResponse,
type ILinkWeixinMessage,
type WechatInboundMessage
} from './types'
export type PollLogLevel = 'info' | 'warn' | 'error'
export type PollLog = (level: PollLogLevel, message: string) => void
export interface ILinkPollerOptions {
/** 拉取一批消息;timeoutMs 来自服务端上一次响应里的建议值。 */
fetchUpdates: (getUpdatesBuf: string, timeoutMs: number) => Promise<ILinkGetUpdatesResponse>
/**
* 处理一批入站消息。
* **只有全部消息被成功接收后才会推进游标**,因此这里抛错会让整批重投。
*/
onMessages: (messages: WechatInboundMessage[]) => Promise<void>
normalize: (raw: ILinkWeixinMessage) => WechatInboundMessage | undefined
loadCursor: () => string
saveCursor: (getUpdatesBuf: string) => void
signal: AbortSignal
log: PollLog
/** bot token 失效(-14):停止轮询,交由上层通知用户重新登录。 */
onStaleToken?: (message: string) => void
now?: () => number
sleep?: (milliseconds: number) => Promise<void>
normalFailureDelayMs?: number
repeatedFailureDelayMs?: number
repeatedFailureThreshold?: number
defaultTimeoutMs?: number
maxTimeoutMs?: number
}
const DEFAULT_NORMAL_DELAY_MS = 2_000
const DEFAULT_REPEATED_DELAY_MS = 30_000
const DEFAULT_REPEATED_THRESHOLD = 3
/**
* 长轮询循环。
*
* 1. **先 dispatch 成功,再持久化新游标**。先推进游标再投递的话,进程在两步之间
* 崩溃就会静默丢消息;这里换取 at-least-once(允许重复,不允许丢失)。
* 2. **客户端自身超时不算失败**:保留游标直接进入下一轮,不计退避。
* 3. **ret/errcode = -14 表示 bot token 已失效**,必须停止业务请求并让用户重新登录;
* 重置游标后紧密重试只会打满接口。
* 4. 采用服务端 `longpolling_timeout_ms` 建议值,而不是写死超时。
*/
export class ILinkPoller {
private readonly options: ILinkPollerOptions
private cursor: string
private consecutiveFailures = 0
/** 是否正处于"连不上"的状态;用于在恢复时补一条日志,否则故障窗口在日志里看不出边界。 */
private degraded = false
private nextTimeoutMs: number
constructor(options: ILinkPollerOptions) {
this.options = options
this.cursor = options.loadCursor()
this.nextTimeoutMs = options.defaultTimeoutMs ?? ILINK_LONG_POLL_TIMEOUT_MS
}
get getUpdatesBuf(): string {
return this.cursor
}
async run(): Promise<void> {
const {
signal,
log,
sleep = (milliseconds: number) => new Promise<void>((r) => setTimeout(r, milliseconds))
} = this.options
const normalDelay = this.options.normalFailureDelayMs ?? DEFAULT_NORMAL_DELAY_MS
const repeatedDelay = this.options.repeatedFailureDelayMs ?? DEFAULT_REPEATED_DELAY_MS
const threshold = this.options.repeatedFailureThreshold ?? DEFAULT_REPEATED_THRESHOLD
const maxTimeout = this.options.maxTimeoutMs ?? ILINK_LONG_POLL_MAX_TIMEOUT_MS
log('info', this.cursor ? '长轮询已启动(恢复上次游标)' : '长轮询已启动')
while (!signal.aborted) {
let response: ILinkGetUpdatesResponse
try {
response = await this.options.fetchUpdates(this.cursor, this.nextTimeoutMs)
} catch (error) {
if (signal.aborted) return
const kind = error instanceof ILinkError ? error.kind : 'network'
// 长轮询客户端超时属于正常控制流:保留游标,立即进入下一轮。
if (kind === 'timeout') continue
// 调用方主动取消(例如停止连接器)
if (kind === 'aborted') return
this.consecutiveFailures += 1
this.degraded = true
const delay = this.failureDelay(normalDelay, repeatedDelay, threshold)
log('warn', `获取更新失败,${Math.round(delay / 1000)} 秒后重试:${describeError(error)}`)
await sleep(delay)
continue
}
if (signal.aborted) return
const ret = Number(response.ret ?? 0)
const errcode = Number(response.errcode ?? 0)
if (ret === -14 || errcode === -14) {
const message = '当前微信机器人登录凭证已失效,需要重新扫码登录'
log('error', message)
// 刻意不清理游标:重新登录后若账号相同,仍可续上原有进度。
this.options.onStaleToken?.(message)
return
}
if (ret !== 0 || errcode !== 0) {
this.consecutiveFailures += 1
this.degraded = true
const delay = this.failureDelay(normalDelay, repeatedDelay, threshold)
log(
'warn',
`服务端返回错误(ret=${ret} errcode=${errcode}${
response.errmsg ? ` errmsg=${response.errmsg}` : ''
}),${Math.round(delay / 1000)} 秒后重试`
)
await sleep(delay)
continue
}
this.consecutiveFailures = 0
if (this.degraded) {
this.degraded = false
log('info', '与微信服务器的连接已恢复')
}
const messages = this.collectMessages(response.msgs)
if (messages.length > 0) {
try {
await this.options.onMessages(messages)
} catch (error) {
// 接收失败:绝不推进游标,让服务端重新投递这一批。
this.consecutiveFailures += 1
this.degraded = true
const delay = this.failureDelay(normalDelay, repeatedDelay, threshold)
log('error', `入站消息处理失败,游标保持不变以便重投:${describeError(error)}`)
await sleep(delay)
continue
}
}
const nextBuf = String(response.get_updates_buf ?? '')
if (nextBuf && nextBuf !== this.cursor) {
this.cursor = nextBuf
try {
this.options.saveCursor(this.cursor)
} catch (error) {
log('warn', `游标持久化失败,下次启动可能重复投递:${describeError(error)}`)
}
}
const suggested = Number(response.longpolling_timeout_ms ?? 0)
if (Number.isFinite(suggested) && suggested > 0) {
this.nextTimeoutMs = Math.min(suggested, maxTimeout)
}
}
}
private collectMessages(raw: ILinkWeixinMessage[] | undefined): WechatInboundMessage[] {
if (!Array.isArray(raw) || raw.length === 0) return []
const result: WechatInboundMessage[] = []
for (const item of raw) {
const normalized = this.options.normalize(item)
if (!normalized) {
this.options.log('warn', '收到一条无法识别的入站消息,已跳过')
continue
}
result.push(normalized)
}
return result
}
private failureDelay(normalDelay: number, repeatedDelay: number, threshold: number): number {
if (this.consecutiveFailures >= threshold) {
this.consecutiveFailures = 0
return repeatedDelay
}
return normalDelay
}
}
+85
View File
@@ -0,0 +1,85 @@
import { randomBytes } from 'node:crypto'
import { assertBusinessOk, ILinkError } from './errors'
import { markdownToPlainText } from './markdown'
import type { ILinkClient } from './client'
import {
ILINK_ITEM_TYPE_TEXT,
ILINK_MESSAGE_STATE_FINISH,
ILINK_MESSAGE_TYPE_BOT,
type ILinkMessageItem,
type ILinkSendMessageRequest
} from './types'
/**
* Bot 外发消息的 `from_user_id`。
*
* 官方客户端(2.4.6)对 Bot 主动外发固定传空字符串;传 bot id 也能被服务端接受。
* 这里遵循官方实现。
*/
const BOT_OUTBOUND_FROM_USER_ID = ''
export function createClientId(): string {
return `tracememo-${randomBytes(16).toString('hex')}`
}
/** 组装 sendmessage 请求体;文本与媒体共用,每个请求只放一个 item。 */
export function buildSendMessageBody(input: {
item: ILinkMessageItem
to: string
contextToken?: string
clientId: string
runId?: string
}): ILinkSendMessageRequest['msg'] {
return {
from_user_id: BOT_OUTBOUND_FROM_USER_ID,
to_user_id: input.to,
client_id: input.clientId,
message_type: ILINK_MESSAGE_TYPE_BOT,
message_state: ILINK_MESSAGE_STATE_FINISH,
item_list: [input.item],
// 缺失时传空字符串而不是省略字段:服务端按字段存在性判断会话。
context_token: String(input.contextToken ?? ''),
...(input.runId ? { run_id: input.runId } : {})
}
}
export interface SendTextOptions {
to: string
text: string
/** 会话上下文令牌:回复当前会话时必须原样回传。 */
contextToken?: string
clientId?: string
runId?: string
signal?: AbortSignal
}
export interface SendTextResult {
clientId: string
/** 实际投递的纯文本内容(Markdown 已降级)。 */
plainText: string
}
export async function sendText(
client: ILinkClient,
options: SendTextOptions
): Promise<SendTextResult> {
const to = String(options.to ?? '').trim()
if (!to) throw new ILinkError({ kind: 'protocol', message: '发送文本需要有效的接收者' })
const plainText = markdownToPlainText(options.text)
if (!plainText) throw new ILinkError({ kind: 'protocol', message: '发送文本内容为空' })
const clientId = options.clientId || createClientId()
const response = await client.sendMessage(
buildSendMessageBody({
item: { type: ILINK_ITEM_TYPE_TEXT, text_item: { text: plainText } },
to,
clientId,
...(options.contextToken ? { contextToken: options.contextToken } : {}),
...(options.runId ? { runId: options.runId } : {})
}),
options.signal
)
assertBusinessOk(response, '发送文本')
return { clientId, plainText }
}
+279
View File
@@ -0,0 +1,279 @@
/**
* WeChat iLink 协议类型与常量。
*
* 与官方客户端(channel_version 2.4.6 基线)对齐,字段名以官方协议为准。
* 这里只描述协议本身,不引入任何 Electron 概念,便于将来被 Tauri 或其他宿主复用。
*/
/** 业务入口固定地址;登录与业务 API 都从这里开始。 */
export const ILINK_DEFAULT_BASE_URL = 'https://ilinkai.weixin.qq.com'
/** CDN 上传/下载基地址。 */
export const ILINK_CDN_BASE_URL = 'https://novac2c.cdn.weixin.qq.com/c2c'
/** base_info.channel_version 上报值。 */
export const ILINK_CHANNEL_VERSION = '2.4.6'
/** iLink-App-Id 固定值。 */
export const ILINK_APP_ID = 'bot'
/** iLink-App-ClientVersion 编码为 (major << 16) | (minor << 8) | patch,2.4.6 => 132102。 */
export const ILINK_APP_CLIENT_VERSION = String((2 << 16) | (4 << 8) | 6)
/** base_info.bot_agent 默认值;仅用于服务端观测聚合,不参与鉴权。 */
export const ILINK_DEFAULT_BOT_AGENT = 'TraceMemo/1.0.0'
/** 长轮询默认与上限。 */
export const ILINK_LONG_POLL_TIMEOUT_MS = 35_000
export const ILINK_LONG_POLL_MAX_TIMEOUT_MS = 120_000
export const ILINK_QR_STATUS_TIMEOUT_MS = 35_000
export const ILINK_SEND_TIMEOUT_MS = 15_000
export const ILINK_CONFIG_TIMEOUT_MS = 10_000
/** bot token 失效(官方 2.4.5 起把内部命名从 session expired 改为 stale token)。 */
export const ILINK_STALE_TOKEN_CODE = -14
/** message_type */
export const ILINK_MESSAGE_TYPE_BOT = 2
/** message_state */
export const ILINK_MESSAGE_STATE_FINISH = 2
/** item type */
export const ILINK_ITEM_TYPE_TEXT = 1
export const ILINK_ITEM_TYPE_IMAGE = 2
export const ILINK_ITEM_TYPE_VOICE = 3
export const ILINK_ITEM_TYPE_FILE = 4
export const ILINK_ITEM_TYPE_VIDEO = 5
/** CDN media_type */
export const ILINK_CDN_MEDIA_TYPE_IMAGE = 1
export const ILINK_CDN_MEDIA_TYPE_VIDEO = 2
export const ILINK_CDN_MEDIA_TYPE_FILE = 3
/** sendtyping 的输入状态值。 */
export const ILINK_TYPING_STATUS_TYPING = 1
export const ILINK_TYPING_STATUS_CANCEL = 2
/**
* 长任务期间维持"正在输入"心跳的间隔。
*
* 协议文档只定义了 status=1/2 两个状态、没有规定间隔;但微信客户端的输入指示
* 会自行消失,所以长任务必须周期性重发 status=1 才能一直亮着。
* 官方客户端按约 5 秒维持,这里沿用同一节奏。
*/
export const ILINK_TYPING_KEEPALIVE_MS = 5_000
/**
* typing_ticket 的缓存有效期。
* 官方按「账号 + 对端用户」缓存、首次取一次;这里用 24 小时兜底,
* 遇到服务端报错会立即失效并在下次需要时重取。
*/
export const ILINK_TYPING_TICKET_TTL_MS = 24 * 60 * 60 * 1000
/** 二维码登录状态机。 */
export type ILinkQrStatus =
| 'wait'
| 'scaned'
| 'confirmed'
| 'expired'
| 'need_verifycode'
| 'verify_code_blocked'
| 'scaned_but_redirect'
| 'binded_redirect'
| (string & {})
export interface ILinkBaseInfo {
channel_version: string
bot_agent: string
}
export interface ILinkQrCodeResponse {
qrcode: string
qrcode_img_content: string
ret?: number
errmsg?: string
}
export interface ILinkQrStatusResponse {
status: ILinkQrStatus
bot_token?: string
ilink_bot_id?: string
ilink_user_id?: string
baseurl?: string
redirect_host?: string
ret?: number
errcode?: number
errmsg?: string
}
/** 持久化到 ~/.tracememo/wechat-connector/accounts/<id>.json 的凭据。 */
export interface ILinkCredentials {
bot_token: string
ilink_bot_id: string
baseurl: string
ilink_user_id: string
}
export interface ILinkTextItem {
text: string
}
export interface ILinkMediaInfo {
encrypt_query_param: string
aes_key: string
encrypt_type: number
}
export interface ILinkImageItem {
url?: string
media?: ILinkMediaInfo
mid_size?: number
}
export interface ILinkVideoItem {
media?: ILinkMediaInfo
video_size?: number
}
export interface ILinkFileItem {
media?: ILinkMediaInfo
file_name?: string
len?: string
}
export interface ILinkVoiceItem {
media?: ILinkMediaInfo
voice_size?: number
encode_type?: number
playtime?: number
text?: string
}
export interface ILinkMessageItem {
type: number
text_item?: ILinkTextItem
image_item?: ILinkImageItem
voice_item?: ILinkVoiceItem
video_item?: ILinkVideoItem
file_item?: ILinkFileItem
}
export interface ILinkWeixinMessage {
seq?: number
message_id?: number
from_user_id?: string
to_user_id?: string
message_type?: number
message_state?: number
item_list?: ILinkMessageItem[]
context_token?: string
session_id?: string
group_id?: string
}
export interface ILinkGetUpdatesResponse {
ret?: number
errcode?: number
errmsg?: string
msgs?: ILinkWeixinMessage[]
get_updates_buf?: string
longpolling_timeout_ms?: number
}
export interface ILinkSendMessageRequest {
msg: {
from_user_id: string
to_user_id: string
client_id: string
message_type: number
message_state: number
item_list: ILinkMessageItem[]
context_token: string
run_id?: string
}
base_info: ILinkBaseInfo
}
export interface ILinkSendMessageResponse {
ret?: number
errmsg?: string
}
export interface ILinkGetUploadUrlRequest {
filekey: string
media_type: number
to_user_id: string
rawsize: number
rawfilemd5: string
filesize: number
no_need_thumb: boolean
aeskey: string
base_info: ILinkBaseInfo
}
export interface ILinkGetUploadUrlResponse {
ret?: number
errmsg?: string
upload_param?: string
upload_full_url?: string
}
export interface ILinkGetConfigResponse {
ret?: number
errcode?: number
errmsg?: string
typing_ticket?: string
}
export interface ILinkSendTypingRequest {
ilink_user_id: string
typing_ticket: string
status: number
base_info: ILinkBaseInfo
}
export interface ILinkSendTypingResponse {
ret?: number
errcode?: number
errmsg?: string
}
/**
* 归一化后的入站消息。Agent Hub 只消费这个形状,
* 不直接依赖协议原始字段,方便将来替换 transport。
*/
export interface WechatInboundItem {
type: number
text?: string
}
export interface WechatInboundMessage {
accountId: string
fromUserId: string
messageId: string
seq?: number
sessionId?: string
groupId?: string
messageType: number
/** 会话上下文令牌:回复必须原样回传,不得用于其他会话,也不得写入普通日志。 */
contextToken?: string
items: WechatInboundItem[]
receivedAt: number
}
/** 账号摘要(供 UI 账号列表使用)。 */
export interface WechatConnectorAccount {
accountId: string
wechatUserId: string
}
/** 登录过程中向宿主上报的事件。 */
export type WechatLoginEvent =
| { status: 'qrcode'; qrCodeDataUrl: string }
| { status: 'wait' | 'scaned' | 'need_verifycode' | 'verify_code_blocked' | 'expired' }
| { status: 'confirmed' }
| { status: 'active'; accountId: string; wechatUserId: string }
/** 连接器对外暴露的运行态。 */
export type WechatConnectorPhase = 'stopped' | 'starting' | 'polling' | 'stale_token' | 'error'
+338
View File
@@ -0,0 +1,338 @@
import { secretFingerprint } from '../log-redaction'
import {
ILINK_TYPING_KEEPALIVE_MS,
ILINK_TYPING_STATUS_CANCEL,
ILINK_TYPING_STATUS_TYPING,
ILINK_TYPING_TICKET_TTL_MS
} from './types'
/**
* 微信原生「正在输入」状态。
*
* 协议:`getconfig` 取 `typing_ticket`,再用 `sendtyping` 下发 status=1 / status=2。
* `typing_ticket` 只用于输入状态,**不是** sendmessage 的鉴权凭据。
*
* 这个模块要解决四件事:
* 1. **非致命**:所有失败都只记日志,绝不影响 Query Agent / 日报 / 最终消息发送。
* 2. **必须收尾**:调用方在 finally 里 stop,任何异常路径都不会留下"对方正在输入"。
* 3. **引用计数**:同一用户并发两个任务时,先结束的那个不能把另一个的 typing 一起取消。
* 4. **少打接口**:typing_ticket 按「账号 + 对端用户」缓存,并发取票去重。
*/
export interface TypingLease {
/** 幂等;可安全地在 finally 里调用。永不抛异常。 */
stop(): Promise<void>
}
export interface TypingBeginInput {
accountId?: string
to: string
contextToken?: string
}
export interface TypingCoordinatorDependencies {
/** 取 typing_ticket;失败/无票返回 undefined 即可(上层已 catch)。 */
fetchTicket: (input: {
ilinkUserId: string
contextToken?: string
}) => Promise<string | undefined>
/** 下发 status=1/2;返回是否成功。 */
sendTyping: (input: { ilinkUserId: string; ticket: string; status: number }) => Promise<boolean>
log: (level: 'info' | 'warn' | 'error', message: string) => void
now?: () => number
keepaliveMs?: number
ticketTtlMs?: number
maxPeers?: number
/**
* 定时器注入点:返回取消函数。
* 测试用可控实现,生产用 setInterval。
*/
schedule?: (tick: () => void, intervalMs: number) => () => void
}
interface PeerTypingState {
key: string
accountId: string
to: string
sessionId: string
refCount: number
/** 当前微信端是否处于「正在输入」。 */
active: boolean
startedAt: number
cancelKeepalive: (() => void) | null
ticket?: string
ticketFetchedAt: number
/** 并发取票去重。 */
ticketFetch: Promise<string | undefined> | null
contextToken?: string
/** 串行化 activate / deactivate,避免本次 TYPING 被上一次的 CANCEL 吃掉。 */
queue: Promise<void>
keepaliveFailureLogged: boolean
lastUseAt: number
}
const DEFAULT_MAX_PEERS = 50
const NOOP_LEASE: TypingLease = { stop: async () => undefined }
export class TypingCoordinator {
private readonly deps: TypingCoordinatorDependencies
private readonly keepaliveMs: number
private readonly ticketTtlMs: number
private readonly maxPeers: number
private readonly states = new Map<string, PeerTypingState>()
constructor(dependencies: TypingCoordinatorDependencies) {
this.deps = dependencies
this.keepaliveMs = dependencies.keepaliveMs ?? ILINK_TYPING_KEEPALIVE_MS
this.ticketTtlMs = dependencies.ticketTtlMs ?? ILINK_TYPING_TICKET_TTL_MS
this.maxPeers = dependencies.maxPeers ?? DEFAULT_MAX_PEERS
}
/** 当前处于「正在输入」的对端数量;用于测试与观测。 */
get activePeerCount(): number {
let count = 0
for (const state of this.states.values()) if (state.active) count += 1
return count
}
/**
* 开始一次输入状态。**永不抛异常**:拿不到 ticket 或服务端失败时返回一个空实现,
* 业务侧照常执行。
*/
async begin(input: TypingBeginInput): Promise<TypingLease> {
const to = String(input.to ?? '').trim()
if (!to) return NOOP_LEASE
const accountId = String(input.accountId ?? '').trim()
const key = `${accountId}::${to}`
try {
this.pruneIfNeeded()
const state = this.stateFor(key, accountId, to)
state.refCount += 1
state.lastUseAt = this.now()
if (input.contextToken) state.contextToken = input.contextToken
const shouldActivate = state.refCount === 1
if (shouldActivate) {
// 串行化:等上一次 deactivate 落地,避免 TYPING 紧接着被 CANCEL 掉。
state.queue = state.queue.then(() => this.activate(state))
await state.queue
}
return this.createLease(state)
} catch (error) {
// 兜底:任何意外都退化为空实现,绝不让 typing 影响业务。
this.deps.log('warn', `typing.begin.failed error=${describeTypingError(error)}`)
return NOOP_LEASE
}
}
/** 账号切换 / 重新登录 / bot token 失效后调用:清空 ticket 缓存。 */
invalidateTickets(): void {
for (const state of this.states.values()) {
state.ticket = undefined
state.ticketFetchedAt = 0
state.ticketFetch = null
}
}
/** 连接器停止时调用:清掉全部状态与定时器,避免留下悬挂的 keepalive。 */
clear(): void {
for (const state of this.states.values()) {
if (state.cancelKeepalive) {
state.cancelKeepalive()
state.cancelKeepalive = null
}
state.refCount = 0
state.active = false
}
this.states.clear()
}
private createLease(state: PeerTypingState): TypingLease {
let stopped = false
return {
stop: async (): Promise<void> => {
if (stopped) return
stopped = true
try {
state.refCount = Math.max(0, state.refCount - 1)
if (state.refCount > 0) return
state.queue = state.queue.then(() => this.deactivate(state))
await state.queue
} catch (error) {
this.deps.log(
'warn',
`typing.stop.leaked session=${state.sessionId} error=${describeTypingError(error)}`
)
}
}
}
}
private stateFor(key: string, accountId: string, to: string): PeerTypingState {
const existing = this.states.get(key)
if (existing) return existing
const state: PeerTypingState = {
key,
accountId,
to,
// 日志里只出现不可逆短指纹,不出现 openid 原文。
sessionId: secretFingerprint(key),
refCount: 0,
active: false,
startedAt: 0,
cancelKeepalive: null,
ticketFetchedAt: 0,
ticketFetch: null,
queue: Promise.resolve(),
keepaliveFailureLogged: false,
lastUseAt: this.now()
}
this.states.set(key, state)
return state
}
private async activate(state: PeerTypingState): Promise<void> {
if (state.active) return
const ticket = await this.ensureTicket(state)
if (!ticket) {
// 无票(getconfig 失败 / 服务端没给)时静默降级,业务照常。
return
}
const ok = await this.trySendTyping(state, ticket, ILINK_TYPING_STATUS_TYPING)
if (!ok) {
// ticket 可能已失效:丢掉缓存,下次需要时重取。
state.ticket = undefined
state.ticketFetchedAt = 0
this.deps.log('warn', `typing.start.failed session=${state.sessionId} ticketPresent=true`)
return
}
state.active = true
state.startedAt = this.now()
state.keepaliveFailureLogged = false
const schedule = this.deps.schedule ?? defaultSchedule
state.cancelKeepalive = schedule(() => {
void this.keepalive(state)
}, this.keepaliveMs)
this.deps.log('info', `typing.start session=${state.sessionId} ticketPresent=true`)
}
private async deactivate(state: PeerTypingState): Promise<void> {
if (state.cancelKeepalive) {
state.cancelKeepalive()
state.cancelKeepalive = null
}
if (!state.active) return
state.active = false
const durationMs = Math.max(0, this.now() - state.startedAt)
const ticket = state.ticket
if (!ticket) return
const ok = await this.trySendTyping(state, ticket, ILINK_TYPING_STATUS_CANCEL)
if (ok) {
this.deps.log(
'info',
`typing.stop session=${state.sessionId} duration=${durationMs}ms success=true`
)
} else {
this.deps.log(
'warn',
`typing.stop.failed session=${state.sessionId} duration=${durationMs}ms success=false`
)
}
}
/**
* 长任务的输入状态会自己消失,必须周期性重发 status=1。
* 失败只在同一会话里记一次日志,避免每 5 秒刷屏。
*/
private async keepalive(state: PeerTypingState): Promise<void> {
if (!state.active || !state.ticket) return
const ok = await this.trySendTyping(state, state.ticket, ILINK_TYPING_STATUS_TYPING)
if (ok) return
if (state.keepaliveFailureLogged) return
state.keepaliveFailureLogged = true
this.deps.log('warn', `typing.keepalive.failed session=${state.sessionId}`)
}
private async ensureTicket(state: PeerTypingState): Promise<string | undefined> {
const cached = String(state.ticket ?? '')
if (cached && this.now() - state.ticketFetchedAt < this.ticketTtlMs) return cached
if (state.ticketFetch) return state.ticketFetch
const pending = (async (): Promise<string | undefined> => {
try {
const ticket = await this.deps.fetchTicket({
ilinkUserId: state.to,
...(state.contextToken ? { contextToken: state.contextToken } : {})
})
const normalized = String(ticket ?? '').trim()
if (!normalized) {
this.deps.log(
'warn',
`typing.ticket.missing session=${state.sessionId} ticketPresent=false`
)
return undefined
}
state.ticket = normalized
state.ticketFetchedAt = this.now()
return normalized
} catch (error) {
this.deps.log(
'warn',
`typing.ticket.failed session=${state.sessionId} ticketPresent=false error=${describeTypingError(error)}`
)
return undefined
} finally {
state.ticketFetch = null
}
})()
state.ticketFetch = pending
return pending
}
private async trySendTyping(
state: PeerTypingState,
ticket: string,
status: number
): Promise<boolean> {
try {
return await this.deps.sendTyping({ ilinkUserId: state.to, ticket, status })
} catch (error) {
// 由调用方决定记哪条日志,这里只吞掉异常保证非致命。
void error
return false
}
}
private pruneIfNeeded(): void {
if (this.states.size <= this.maxPeers) return
const idle = [...this.states.values()]
.filter((state) => state.refCount === 0 && !state.active)
.sort((left, right) => left.lastUseAt - right.lastUseAt)
for (const state of idle) {
if (this.states.size <= this.maxPeers) break
this.states.delete(state.key)
}
}
private now(): number {
return this.deps.now?.() ?? Date.now()
}
}
function defaultSchedule(tick: () => void, intervalMs: number): () => void {
const timer = setInterval(tick, intervalMs)
// keepalive 不该阻止进程退出。
if (typeof timer.unref === 'function') timer.unref()
return () => clearInterval(timer)
}
function describeTypingError(error: unknown): string {
return error instanceof Error ? error.message : String(error)
}
+189
View File
@@ -0,0 +1,189 @@
import { chmodSync, mkdirSync, renameSync, rmSync, readFileSync, writeFileSync } from 'node:fs'
import { dirname } from 'node:path'
import type { WechatInboundItem, WechatInboundMessage } from './wechat-ilink/types'
/**
* Agent Hub 入站收件箱。
*
* 存在意义只有一个:**拿到消息就立刻落盘,然后才允许长轮询推进游标**。
* 先推进游标再异步投递的话,进程在两步之间崩溃就会静默丢消息;
* 因此这里的语义是 at-least-once:允许重复,不允许丢失。
*
* 代价:极端情况下(处理完但删除前崩溃)会重复处理同一条消息。
* Agent Hub 有 message_id 去重 + 业务侧幂等,重复可以接受,丢消息不行。
*
* 文件含聊天文本与 context_token,因此固定 0600 权限、处理完即删除,且从不写入日志。
*/
export interface WechatInboundInboxEntry {
key: string
accountId: string
fromUserId: string
messageId: string
contextToken?: string
items: WechatInboundItem[]
receivedAt: number
attempts: number
}
interface InboxFile {
entries: WechatInboundInboxEntry[]
}
export interface WechatInboundInboxOptions {
filePath: () => string
maxAttempts?: number
maxEntries?: number
now?: () => number
}
const DEFAULT_MAX_ATTEMPTS = 3
const DEFAULT_MAX_ENTRIES = 200
export function inboxKeyFor(
message: WechatInboundMessage,
fallbackIndex: number,
now: number
): string {
if (message.messageId) return `${message.accountId}::${message.messageId}`
return `${message.accountId}::${message.fromUserId}::${now}::${fallbackIndex}`
}
export class WechatInboundInbox {
private readonly options: WechatInboundInboxOptions
private readonly maxAttempts: number
private readonly maxEntries: number
private entries: WechatInboundInboxEntry[] | null = null
constructor(options: WechatInboundInboxOptions) {
this.options = options
this.maxAttempts = options.maxAttempts ?? DEFAULT_MAX_ATTEMPTS
this.maxEntries = options.maxEntries ?? DEFAULT_MAX_ENTRIES
}
/**
* 持久化接收一批消息,返回**本次新增**的条目。
* 已存在(重复投递)的条目不会重复返回,也不会重复处理。
*/
accept(messages: WechatInboundMessage[]): WechatInboundInboxEntry[] {
const current = this.load()
const known = new Set(current.map((entry) => entry.key))
const accepted: WechatInboundInboxEntry[] = []
const now = this.options.now?.() ?? Date.now()
messages.forEach((message, index) => {
const key = inboxKeyFor(message, index, now)
if (known.has(key)) return
known.add(key)
accepted.push({
key,
accountId: message.accountId,
fromUserId: message.fromUserId,
messageId: message.messageId,
...(message.contextToken ? { contextToken: message.contextToken } : {}),
items: message.items.map((item) => ({ ...item })),
receivedAt: message.receivedAt || now,
attempts: 0
})
})
if (accepted.length === 0) return []
const next = [...current, ...accepted].slice(-this.maxEntries)
this.persist(next)
this.entries = next
return accepted
}
/** 尚未处理完成的条目,按接收顺序返回。 */
pending(): WechatInboundInboxEntry[] {
return this.load().map((entry) => ({ ...entry }))
}
contains(key: string): boolean {
return this.load().some((entry) => entry.key === key)
}
size(): number {
return this.load().length
}
/** 处理成功,从收件箱移除。 */
complete(key: string): void {
const current = this.load()
const next = current.filter((entry) => entry.key !== key)
if (next.length === current.length) return
this.persist(next)
this.entries = next
}
/**
* 处理失败:累加尝试次数。
* 达到上限后放弃,返回 abandoned=true,由调用方明确记一条 error 日志——
* 静默丢弃是不允许的,但无限重试同一条毒消息同样不允许。
*/
recordFailure(key: string): { attempts: number; abandoned: boolean } {
const current = this.load()
let attempts = 0
let abandoned = false
const next = current
.map((entry) => {
if (entry.key !== key) return entry
attempts = entry.attempts + 1
abandoned = attempts >= this.maxAttempts
return { ...entry, attempts }
})
.filter((entry) => !(entry.key === key && abandoned))
this.persist(next)
this.entries = next
return { attempts, abandoned }
}
clear(): void {
this.persist([])
this.entries = []
}
private load(): WechatInboundInboxEntry[] {
if (this.entries) return this.entries
let parsed: InboxFile | null = null
try {
parsed = JSON.parse(readFileSync(this.options.filePath(), 'utf8')) as InboxFile
} catch {
parsed = null
}
const entries = Array.isArray(parsed?.entries)
? parsed!.entries.filter((entry): entry is WechatInboundInboxEntry => {
if (!entry || typeof entry !== 'object') return false
const candidate = entry as Partial<WechatInboundInboxEntry>
return Boolean(
String(candidate.key || '').trim() && String(candidate.fromUserId || '').trim()
)
})
: []
this.entries = entries
return entries
}
private persist(entries: WechatInboundInboxEntry[]): void {
const path = this.options.filePath()
try {
mkdirSync(dirname(path), { recursive: true, mode: 0o700 })
const tempPath = `${path}.tmp-${process.pid}-${Date.now()}`
writeFileSync(tempPath, JSON.stringify({ entries } satisfies InboxFile, null, 2), {
encoding: 'utf8',
mode: 0o600
})
chmodSync(tempPath, 0o600)
renameSync(tempPath, path)
chmodSync(path, 0o600)
} catch (error) {
// 落盘失败必须向上抛:调用方要放弃推进游标,让服务端重新投递。
try {
rmSync(`${path}.tmp-${process.pid}`, { force: true })
} catch {
// 忽略临时文件清理失败。
}
throw error
}
}
}
+357
View File
@@ -0,0 +1,357 @@
import { randomUUID } from 'node:crypto'
import type {
PersonalWechatSendRequest,
PersonalWechatSendResult
} from '../../shared/personal-wechat'
import {
normalizeWechatSendRequest,
resolveSendTransport,
type WechatSendErrorCode,
type WechatSendLogEntry,
type WechatSendRequest,
type WechatSendResult,
type WechatSendTransport
} from '../../shared/wechat-send'
import { isILinkError } from './wechat-ilink/errors'
import { wechatSendLogService, type WechatSendLogService } from './wechat-send-log-service'
/** iLink 发送适配器的注入点;由主进程在启动时接到 WechatConnectorService 上。 */
export type IlinkSendAdapter = (request: WechatSendRequest) => Promise<void>
export interface WechatSendGatewayDependencies {
now?: () => number
createRequestId?: () => string
/** 个人微信(注入式发送)适配器。 */
sendPersonal?: (request: PersonalWechatSendRequest) => Promise<PersonalWechatSendResult>
/** iLink 适配器;未注入时 iLink 发送记为 TRANSPORT_UNAVAILABLE。 */
sendIlink?: IlinkSendAdapter
log?: WechatSendLogService
}
class UnsupportedSendTypeError extends Error {}
function errorMessage(error: unknown): string {
return error instanceof Error ? error.message : String(error)
}
function classifyIlinkError(error: unknown): WechatSendErrorCode {
if (isILinkError(error) && error.isStaleToken) return 'STALE_TOKEN'
return 'SEND_FAILED'
}
/**
* 统一微信发送入口。
*
* ```text
* 业务层
* │
* ▼
* WechatSendGateway
* ├── Send Log
* ├── iLink adapter → WechatConnectorService
* └── Personal adapter → PersonalWechatSendService(Windows / macOS 注入式)
* ```
*
* 统一的是 TM 上层发送模型,不强行统一底层协议:
* 个人微信仍然是 `{toWxid,type,msg}` 心智模型,iLink 仍然是 sendmessage。
*/
export class WechatSendGateway {
private readonly deps: Required<Pick<WechatSendGatewayDependencies, 'now' | 'createRequestId'>> &
WechatSendGatewayDependencies
private ilinkAdapter: IlinkSendAdapter | null
constructor(dependencies: WechatSendGatewayDependencies = {}) {
this.deps = {
now: dependencies.now ?? (() => Date.now()),
createRequestId: dependencies.createRequestId ?? (() => randomUUID()),
log: dependencies.log ?? wechatSendLogService,
...(dependencies.sendPersonal ? { sendPersonal: dependencies.sendPersonal } : {}),
...(dependencies.sendIlink ? { sendIlink: dependencies.sendIlink } : {})
}
this.ilinkAdapter = dependencies.sendIlink ?? null
}
/** 主进程启动时注入 iLink 通道(WechatConnectorService)。 */
configureIlinkSender(adapter: IlinkSendAdapter): void {
this.ilinkAdapter = adapter
}
hasIlinkSender(): boolean {
return this.ilinkAdapter !== null
}
listSendLog(): WechatSendLogEntry[] {
return this.deps.log?.list() ?? []
}
/**
* 统一发送入口。
* 刻意不抛异常:调用方永远拿到结构化结果,失败也会留下 Send Log。
*/
async send(input: unknown): Promise<WechatSendResult> {
const startedAt = this.deps.now()
const fallbackRequestId =
input &&
typeof input === 'object' &&
typeof (input as { request_id?: unknown }).request_id === 'string'
? String((input as { request_id: string }).request_id).trim()
: ''
const requestId = fallbackRequestId || this.deps.createRequestId()
const request = normalizeWechatSendRequest(input, { createRequestId: () => requestId })
if (!request) {
const raw = (input ?? {}) as Partial<WechatSendRequest>
const transport = resolveSendTransport({
...(raw.transport === 'ilink' || raw.transport === 'personal'
? { transport: raw.transport }
: {}),
...(typeof raw.context_token === 'string' ? { context_token: raw.context_token } : {})
})
return this.finish({
requestId,
transport,
type: typeof raw.type === 'string' ? (raw.type as WechatSendRequest['type']) : 'text',
to: typeof raw.to === 'string' ? raw.to : '',
msg: typeof raw.msg === 'string' ? raw.msg : '',
startedAt,
errorCode: 'INVALID_REQUEST',
error: '发送请求不合法:缺少接收者、类型或内容'
})
}
const transport = resolveSendTransport(request)
try {
if (transport === 'ilink') {
if (!this.ilinkAdapter) {
return this.finish({
requestId: request.request_id,
transport,
type: request.type,
to: request.to,
msg: request.msg,
startedAt,
...(request.account_id ? { accountId: request.account_id } : {}),
errorCode: 'TRANSPORT_UNAVAILABLE',
error: 'iLink 发送通道尚未初始化'
})
}
await this.ilinkAdapter(request)
} else {
const personalResult = await this.sendPersonalRequest(unifiedToPersonalRequest(request))
if (!personalResult.success) {
return this.finish({
requestId: request.request_id,
transport,
type: request.type,
to: request.to,
msg: request.msg,
startedAt,
...(request.account_id ? { accountId: request.account_id } : {}),
errorCode: 'SEND_FAILED',
error: personalResult.error || '个人微信发送失败'
})
}
}
return this.finish({
requestId: request.request_id,
transport,
type: request.type,
to: request.to,
msg: request.msg,
startedAt,
...(request.account_id ? { accountId: request.account_id } : {}),
status: 'sent'
})
} catch (error) {
const errorCode: WechatSendErrorCode =
error instanceof UnsupportedSendTypeError
? 'UNSUPPORTED_TYPE'
: transport === 'ilink'
? classifyIlinkError(error)
: 'SEND_FAILED'
return this.finish({
requestId: request.request_id,
transport,
type: request.type,
to: request.to,
msg: request.msg,
startedAt,
...(request.account_id ? { accountId: request.account_id } : {}),
errorCode,
error: errorMessage(error)
})
}
}
/**
* 兼容入口:既有个人微信调用方直接给 `PersonalWechatSendRequest`。
* 走同一条 Send Log,但保持原有返回类型,避免打断现有业务与测试。
*/
async sendPersonal(request: PersonalWechatSendRequest): Promise<PersonalWechatSendResult> {
const requestId = this.deps.createRequestId()
const startedAt = this.deps.now()
const preview = personalToPreview(request)
try {
const result = await this.sendPersonalRequest(request)
this.record({
requestId,
transport: 'personal',
type: preview.type,
to: preview.to,
msg: preview.msg,
startedAt,
status: result.success ? 'sent' : 'failed',
...(result.success ? {} : { errorCode: 'SEND_FAILED' as WechatSendErrorCode })
})
return result
} catch (error) {
this.record({
requestId,
transport: 'personal',
type: preview.type,
to: preview.to,
msg: preview.msg,
startedAt,
status: 'failed',
errorCode: 'SEND_FAILED'
})
throw error
}
}
private async sendPersonalRequest(
request: PersonalWechatSendRequest
): Promise<PersonalWechatSendResult> {
const sender = this.deps.sendPersonal ?? (await defaultPersonalSender())
return sender(request)
}
private finish(input: {
requestId: string
transport: WechatSendTransport
type: WechatSendRequest['type']
to: string
msg: string
startedAt: number
accountId?: string
status?: 'sent' | 'failed' | 'blocked'
errorCode?: WechatSendErrorCode
error?: string
}): WechatSendResult {
const durationMs = Math.max(0, this.deps.now() - input.startedAt)
const status = input.status ?? 'failed'
this.record({
requestId: input.requestId,
transport: input.transport,
type: input.type,
to: input.to,
msg: input.msg,
startedAt: input.startedAt,
status,
durationMs,
...(input.accountId ? { accountId: input.accountId } : {}),
...(input.errorCode ? { errorCode: input.errorCode } : {})
})
return {
request_id: input.requestId,
success: status === 'sent',
status,
transport: input.transport,
duration_ms: durationMs,
...(input.errorCode ? { error_code: input.errorCode } : {}),
...(input.error ? { error: input.error } : {})
}
}
private record(input: {
requestId: string
transport: WechatSendTransport
type: WechatSendRequest['type']
to: string
msg: string
startedAt: number
status: 'sent' | 'failed' | 'blocked'
durationMs?: number
accountId?: string
errorCode?: WechatSendErrorCode
}): void {
const log = this.deps.log
if (!log) return
try {
log.record(
log.buildEntry({
request_id: input.requestId,
transport: input.transport,
to: input.to,
type: input.type,
msg: input.msg,
status: input.status,
timestamp: this.deps.now(),
...(input.durationMs !== undefined ? { duration_ms: input.durationMs } : {}),
...(input.accountId ? { account_id: input.accountId } : {}),
...(input.errorCode ? { error_code: input.errorCode } : {})
})
)
} catch (error) {
console.warn('[WechatSendGateway] 发送日志记录失败:', error)
}
}
}
/** 统一模型 → 个人微信模型。文件类型个人通道不支持,显式报错而不是静默降级。 */
export function unifiedToPersonalRequest(request: WechatSendRequest): PersonalWechatSendRequest {
const base = { to: request.to, isGroup: request.is_group === true }
if (request.type === 'text') {
return { ...base, type: 'text', text: request.msg }
}
if (request.type === 'image') {
return { ...base, type: 'image', filePath: request.msg }
}
if (request.type === 'voice') {
const metadata = request.metadata ?? {}
const fromId = typeof metadata.fromId === 'string' ? metadata.fromId.trim() : ''
const durationMs =
typeof metadata.durationMs === 'number' && Number.isFinite(metadata.durationMs)
? metadata.durationMs
: undefined
return {
...base,
type: 'voice',
filePath: request.msg,
...(fromId ? { fromId } : {}),
...(durationMs !== undefined ? { durationMs } : {})
}
}
throw new UnsupportedSendTypeError('个人微信通道暂不支持发送文件')
}
function personalToPreview(request: PersonalWechatSendRequest): {
type: WechatSendRequest['type']
to: string
msg: string
} {
if (request.type === 'text') return { type: 'text', to: request.to, msg: request.text }
return {
type: request.type === 'image' ? 'image' : 'voice',
to: request.to,
msg: request.filePath
}
}
let cachedPersonalSender:
| ((request: PersonalWechatSendRequest) => Promise<PersonalWechatSendResult>)
| null = null
async function defaultPersonalSender(): Promise<
(request: PersonalWechatSendRequest) => Promise<PersonalWechatSendResult>
> {
if (!cachedPersonalSender) {
const module = await import('./personal-wechat-send-service')
cachedPersonalSender = (request) => module.personalWechatSendService.send(request)
}
return cachedPersonalSender
}
/** 兼容性名称。 */
export const wechatSendGateway = new WechatSendGateway()
@@ -0,0 +1,147 @@
import { app } from 'electron'
import { createHash } from 'node:crypto'
import fs from 'fs-extra'
import path from 'path'
import {
buildSendPreview,
type WechatSendLogEntry,
type WechatSendStatus,
type WechatSendTransport,
type WechatSendType
} from '../../shared/wechat-send'
const MAX_LOG_ENTRIES = 500
/**
* 只允许白名单字段落盘。
* 即使调用方不小心把 context_token / bot_token 之类的字段混进条目,
* 也绝不会被写进 Send Log——隐私边界放在这里兜底,而不是依赖调用方自觉。
*/
function sanitizeEntry(entry: WechatSendLogEntry): WechatSendLogEntry {
return {
request_id: String(entry.request_id ?? ''),
timestamp: Number(entry.timestamp) || 0,
transport: entry.transport,
to: String(entry.to ?? ''),
type: entry.type,
status: entry.status,
...(entry.msg_preview ? { msg_preview: entry.msg_preview } : {}),
...(entry.msg_hash ? { msg_hash: entry.msg_hash } : {}),
...(entry.account_id ? { account_id: entry.account_id } : {}),
...(entry.duration_ms !== undefined ? { duration_ms: entry.duration_ms } : {}),
...(entry.error_code ? { error_code: entry.error_code } : {})
}
}
export interface WechatSendLogServiceOptions {
getUserDataPath?: () => string
maxEntries?: number
}
/**
* 发送日志(Send Log)。
*
* 与 WechatActionGateway 的 Action 审计刻意分成两层:
* - Send Log:**每一次**经过 WechatSendGateway 的发送都记录,包括普通 Agent 问答回复;
* - Action 审计:只记录需要业务审计的高层动作(定时日报、退群通知、用户主动 TTS)。
*
* 记录内容:request_id / 时间 / transport / 目标 / 类型 / 截断预览 / sha256 / 状态 / 耗时。
* **不记录**:context_token、bot_token、完整聊天文本、完整本地路径、Authorization 头。
*/
export class WechatSendLogService {
private readonly getUserDataPath: () => string
private readonly maxEntries: number
constructor(options: WechatSendLogServiceOptions = {}) {
this.getUserDataPath = options.getUserDataPath || (() => app.getPath('userData'))
this.maxEntries = options.maxEntries ?? MAX_LOG_ENTRIES
}
/** 对消息原文计算稳定哈希;不做任何截断,用于后续比对是否同一条内容。 */
hashMessage(msg: string): string {
return `sha256:${createHash('sha256')
.update(String(msg ?? ''))
.digest('hex')}`
}
/** 组装一条日志条目;预览会自动截断,媒体只保留文件名。 */
buildEntry(input: {
request_id: string
transport: WechatSendTransport
to: string
type: WechatSendType
msg: string
status: WechatSendStatus
timestamp: number
duration_ms?: number
account_id?: string
error_code?: WechatSendLogEntry['error_code']
}): WechatSendLogEntry {
const preview = buildSendPreview(input.msg, input.type)
return {
request_id: input.request_id,
timestamp: input.timestamp,
transport: input.transport,
to: input.to,
type: input.type,
status: input.status,
...(preview ? { msg_preview: preview } : {}),
msg_hash: this.hashMessage(input.msg),
...(input.account_id ? { account_id: input.account_id } : {}),
...(input.duration_ms !== undefined ? { duration_ms: input.duration_ms } : {}),
...(input.error_code ? { error_code: input.error_code } : {})
}
}
list(): WechatSendLogEntry[] {
return this.readAll().map((entry) => ({ ...entry }))
}
record(entry: WechatSendLogEntry): void {
const records = this.readAll()
const sanitized = sanitizeEntry(entry)
const withoutSameRequest = records.filter((item) => item.request_id !== sanitized.request_id)
const next = [sanitized, ...withoutSameRequest].slice(0, this.maxEntries)
try {
fs.ensureDirSync(path.dirname(this.filePath()))
fs.writeJsonSync(this.filePath(), next, { spaces: 2 })
} catch (error) {
// 发送日志写失败不能影响真实发送结果。
console.warn('[WechatSendLog] 发送日志写入失败:', error)
}
}
clear(): void {
try {
fs.writeJsonSync(this.filePath(), [], { spaces: 2 })
} catch {
// 忽略:清空失败不影响后续发送。
}
}
private filePath(): string {
return path.join(this.getUserDataPath(), 'actions', 'wechat-send-log.json')
}
private readAll(): WechatSendLogEntry[] {
try {
const value = fs.readJsonSync(this.filePath()) as unknown
if (!Array.isArray(value)) return []
return value.filter((item): item is WechatSendLogEntry => {
if (!item || typeof item !== 'object') return false
const record = item as Partial<WechatSendLogEntry>
// `to` 允许为空:非法请求同样要留痕,此时没有可用的接收者。
return Boolean(
String(record.request_id || '').trim() &&
String(record.type || '').trim() &&
String(record.status || '').trim() &&
Number.isFinite(Number(record.timestamp))
)
})
} catch {
return []
}
}
}
export const wechatSendLogService = new WechatSendLogService()
+50 -1
View File
@@ -126,5 +126,54 @@ export class AudioDecoderRegistry {
}
export function createDefaultAudioDecoderRegistry(): AudioDecoderRegistry {
return new AudioDecoderRegistry().register(new SilkAudioDecoder())
return new AudioDecoderRegistry()
.register(new SilkAudioDecoder())
.register(new AmrAudioDecoder())
}
/**
* V3 AMR:外部 `ffmpeg` 解码为 PCM(MultiMediaDyn `CAMRDecoder` 未导出)。
*/
export class AmrAudioDecoder implements VoiceAudioDecoder {
readonly codec = 'amr'
async decode(source: EncodedVoiceSource): Promise<DecodedVoiceAudio> {
const { spawn } = await import('child_process')
const { tmpdir } = await import('os')
const { promises: fs } = await import('fs')
const { randomBytes } = await import('crypto')
const stamp = randomBytes(8).toString('hex')
const inPath = join(tmpdir(), `tracememo-amr-in-${stamp}.amr`)
const outPath = join(tmpdir(), `tracememo-amr-out-${stamp}.wav`)
await fs.writeFile(inPath, source.data)
try {
await new Promise<void>((resolve, reject) => {
const child = spawn(
'ffmpeg',
['-y', '-i', inPath, '-f', 'wav', '-acodec', 'pcm_s16le', '-ar', '8000', '-ac', '1', outPath],
{ stdio: ['ignore', 'ignore', 'pipe'] }
)
let err = ''
child.stderr.on('data', (chunk) => {
err += String(chunk)
})
child.on('error', reject)
child.on('close', (code) => {
if (code === 0) resolve()
else reject(new Error(`ffmpeg amr decode failed: ${err.slice(-200)}`))
})
})
const wav = await fs.readFile(outPath)
const pcm = wav.length > 44 && wav.toString('ascii', 0, 4) === 'RIFF' ? wav.subarray(44) : wav
return {
pcm,
sampleRate: 8000,
channels: 1,
sourceHash: source.sourceHash
}
} finally {
await fs.rm(inPath, { force: true })
await fs.rm(outPath, { force: true })
}
}
}
+21 -13
View File
@@ -37,21 +37,29 @@ export class VoicePipeline {
async run(
accountId: string,
reference: VoiceMessageReference,
signal?: AbortSignal
signal?: AbortSignal,
options?: { force?: boolean }
): Promise<{ transcript: string; language?: string; durationMs: number; cached: boolean }> {
const messageIdentity = voiceMessageIdentity(reference)
const compatible = this.transcripts.findCompatible({
accountId,
messageIdentity,
processorVersion: this.audioProcessor.version,
...this.recognizer.metadata
})
if (compatible?.transcript.trim()) {
return {
transcript: compatible.transcript.trim(),
language: compatible.language,
durationMs: compatible.durationMs,
cached: true
/*
* findCompatible 只按消息身份匹配,不含 audio_hash:音频被换掉(例如取音频的
* 逻辑修好后)时它仍会命中旧记录。用户主动触发的识别必须跳过它,重新取一次
* 音频 —— 音频级缓存 find(key) 带 audio_hash,才是正确的失效机制。
*/
if (!options?.force) {
const compatible = this.transcripts.findCompatible({
accountId,
messageIdentity,
processorVersion: this.audioProcessor.version,
...this.recognizer.metadata
})
if (compatible?.transcript.trim()) {
return {
transcript: compatible.transcript.trim(),
language: compatible.language,
durationMs: compatible.durationMs,
cached: true
}
}
}
@@ -23,6 +23,7 @@ type TranscriptUpdateListener = (update: VoiceTranscriptUpdate) => Promise<void>
type RecognitionOptions = {
priority?: VoiceRecognitionPriority
publishTranscriptUpdate?: boolean
force?: boolean
}
export class VoiceRecognitionUseCase {
@@ -108,7 +109,9 @@ export class VoiceRecognitionUseCase {
error: '请先下载语音识别模型'
} as const
}
const result = await pipeline.run(accountId, reference, signal)
const result = await pipeline.run(accountId, reference, signal, {
force: options?.force
})
if (signal.aborted || !this.isCurrentAccount(accountId, generation)) {
throw new DOMException('Recognition cancelled', 'AbortError')
}
+102 -1
View File
@@ -1,4 +1,6 @@
import { createHash } from 'crypto'
import { promises as fs } from 'fs'
import { join } from 'path'
import { Wcdb4Client } from './wcdb4-client'
import {
createDefaultAudioDecoderRegistry,
@@ -73,11 +75,19 @@ export class VoiceService {
pcmResult.audio.sampleRate,
pcmResult.audio.channels
)
// duration 由 PCM 字节数反算(wavData 去掉 44 字节头即 PCM),
// 用来和用户实际听到的长度对齐、排查「时长显示不对」这类问题。
// 注意它**不是权威值**:权威时长在消息 XML 的 <voicemsg voicelength>(毫秒),
// 显示层以那个为准;这里只是解码结果的自证。
const durationSeconds =
pcmData.length / (pcmResult.audio.sampleRate * pcmResult.audio.channels * 2)
console.log(
'[VoiceService] wavData length:',
wavData.length,
'base64 length:',
wavData.toString('base64').length
wavData.toString('base64').length,
'duration:',
`${durationSeconds.toFixed(2)}s`
)
const base64Data = wavData.toString('base64')
@@ -203,6 +213,17 @@ export class VoiceService {
svrId || 0
)
if (!voiceResult.success || !voiceResult.hex) {
const disk = await this.resolveFromDisk(sessionId, localId, createTime, svrId)
if (disk?.length) {
return {
success: true,
source: {
data: disk,
codec: 'silk',
sourceHash: createHash('sha256').update(disk).digest('hex')
}
}
}
return { success: false, error: voiceResult.error || '获取语音数据失败' }
}
@@ -231,6 +252,38 @@ export class VoiceService {
return candidates
}
/**
* V2:旁路 `wcdb_get_voice_data`,在账号目录下找语音文件
* (`GetMsgAudioPath` 同域:Message / MsgAndFiles / VoiceTemp)。
*/
private async resolveFromDisk(
sessionId: string,
localId: number,
createTime: number,
svrId?: string | number
): Promise<Buffer | null> {
const accountRoot = this.wcdb4Client.getAccountRoot?.()
if (!accountRoot) return null
const needles = [
String(localId),
String(createTime),
svrId !== undefined && svrId !== null && String(svrId) !== '0' ? String(svrId) : '',
sessionId.replace(/@chatroom$/, '')
].filter(Boolean)
const files = await listVoiceDiskCandidates(accountRoot)
for (const file of files) {
const name = file.toLowerCase()
if (!needles.some((needle) => needle && name.includes(needle.toLowerCase()))) continue
try {
const data = await fs.readFile(file)
if (data.length) return data
} catch {
// try next candidate
}
}
return null
}
private decodeVoiceBlob(hex: string): Buffer | null {
try {
const hexClean = hex.replace(/\s+/g, '')
@@ -266,3 +319,51 @@ export class VoiceService {
return Buffer.concat([header, pcmData])
}
}
const VOICE_DISK_EXTS = new Set(['.aud', '.silk', '.amr'])
const VOICE_DISK_DIRS = [
'msg/Message',
'msg/MsgAndFiles',
'msg/VoiceTemp',
'msg/History',
'Message',
'MsgAndFiles',
'VoiceTemp',
// 真机布局(2026-09-24):cache/YYYY-MM/Message/<md5>/VoiceTemp/
'cache'
]
/** 语音磁盘候选路径(扩展名/目录与 GetMsgAudioPath / VoiceTemp 对齐)。 */
export async function listVoiceDiskCandidates(accountRoot: string): Promise<string[]> {
const out: string[] = []
const walk = async (dir: string, depth: number): Promise<void> => {
if (out.length > 200 || depth > 6) return
let names: string[]
try {
names = await fs.readdir(dir)
} catch {
return
}
for (const name of names) {
const full = join(dir, name)
let stat: Awaited<ReturnType<typeof fs.stat>>
try {
stat = await fs.stat(full)
} catch {
continue
}
if (stat.isDirectory()) {
await walk(full, depth + 1)
continue
}
const ext = name.slice(name.lastIndexOf('.')).toLowerCase()
if (!VOICE_DISK_EXTS.has(ext)) continue
out.push(full)
}
}
for (const rel of VOICE_DISK_DIRS) {
await walk(join(accountRoot, rel), 0)
if (out.length > 200) break
}
return out
}
+439 -1
View File
@@ -5,7 +5,54 @@ import crypto from 'crypto'
import { createRequire } from 'module'
import { createConnection, Socket } from 'net'
import { getResourceRoots } from './resource-paths'
export type RoomInfoRow = {
roomId: string
owner?: string
announcement?: string
announcementEditor?: string
maxMemberCount?: number
chatName?: string
openImAccountType?: string
isOpenIm?: boolean
}
function pickString(row: Record<string, unknown>, keys: string[]): string | undefined {
for (const key of keys) {
const value = row[key]
if (typeof value === 'string' && value.trim()) return value.trim()
if (typeof value === 'number' && Number.isFinite(value)) return String(value)
}
return undefined
}
export function normalizeRoomInfoRow(chatroomId: string, row: Record<string, unknown>): RoomInfoRow {
const openImAccountType =
pickString(row, ['openim_acct_type', 'openimAcctType', 'openim_type']) || undefined
const maxRaw = pickString(row, ['max_member_count', 'maxMemberCount', 'member_max'])
return {
roomId: chatroomId,
owner: pickString(row, ['roomowner', 'room_owner', 'owner', 'owner_id']),
announcement: pickString(row, [
'textannouncement',
'announcement',
'announcement_',
'chatroomannouncement'
]),
announcementEditor: pickString(row, [
'announcement_editor',
'announcement_editor_',
'announcementeditor'
]),
maxMemberCount: maxRaw && Number.isFinite(Number(maxRaw)) ? Number(maxRaw) : undefined,
chatName: pickString(row, ['chatroomnick', 'chat_name', 'chatroom_name', 'displayname']),
openImAccountType,
isOpenIm: openImAccountType ? openImAccountType !== '0' && openImAccountType !== '' : undefined
}
}
import { wcdbDebugLog } from './wcdb-debug'
import type { ImageMessageCountProbe } from '../shared/image-text-index'
import { imageTextWindowToSeconds } from '../shared/image-text-index'
export interface Wcdb4Session {
username: string
@@ -1232,11 +1279,15 @@ export class Wcdb4Client {
}
async countVoiceMessagesAsync(
username: string,
md5OrUsername: string,
startTime?: number,
endTime?: number
): Promise<number | null> {
if (!this.wcdbGetMessageTableStats || !this.wcdbExecQuery) return null
// 与图片计数同因的修正:同一个 md5/username 混淆在这里也存在,
// 而且它更隐蔽 —— 匹配不到表时循环不执行,函数会**返回 0 而不是报错**。
const username = this.resolveMessageUsername(md5OrUsername)
if (!username) return null
let tables: Wcdb4MessageStore[]
try {
@@ -1276,6 +1327,243 @@ export class Wcdb4Client {
return total
}
/**
* 消息类型列的可能名字(按顺序探测,命中即用)。
*
* 为什么不能直接硬编码 `"local_type"`:`pickValue(row, [...别名])` 那套别名列表只作用于
* **已经读出来的行**;一旦把列名写进 WHERE,列名不同的库会当场抛错,再被 catch 吞成
* `null` —— 表现就是"检测到 0 张图片"。所以必须先探测真实列名。
*/
private readonly messageTypeColumnCandidates = [
'local_type',
'localType',
'msg_type',
'msgType',
'message_type',
'messageType',
'type',
'WCDB_CT_local_type'
]
/** 每个消息分片的真实类型列名;探测一次即缓存,避免每个会话都跑一次 PRAGMA。 */
private readonly messageTypeColumnCache = new Map<string, string | null>()
private resolveMessageTypeColumn(store: Wcdb4MessageStore): string | null {
const cacheKey = `${store.dbPath}\u0000${store.tableName}`
const cached = this.messageTypeColumnCache.get(cacheKey)
if (cached !== undefined) return cached
let resolved: string | null = null
try {
const columns = this.readMessageColumns(store).map((column) => column.name)
for (const candidate of this.messageTypeColumnCandidates) {
const hit = columns.find((name) => name.toLowerCase() === candidate.toLowerCase())
if (hit) {
resolved = hit
break
}
}
} catch {
resolved = null
}
this.messageTypeColumnCache.set(cacheKey, resolved)
return resolved
}
/**
* 把「会话 md5」解析成原生接口真正需要的 username。
*
* `contact.md5` 是 `md5(wxid)` 的**哈希**(见 chat-service 的 `dbRef.md5(user.m_nsUsrName)`),
* 而 `wcdbGetMessageTableStats` / `wcdbGetMessages` 这些原生接口要的是**原始 username**。
* 直接把 md5 当 username 传,原生侧匹配不到任何表 —— 表现为"未找到该会话的消息表",
* 而按表统计的计数会静默变成 0。
*
* 既有读消息路径一直做了这层转换(`listSourceMessages` 里的 `getUsernameByMd5`),
* **统计/水位路径漏了**,所以这里统一补上。
*
* 解析不到时原样返回:调用方本来就传 username 的路径仍然可用。
*/
private resolveMessageUsername(md5OrUsername: string): string {
const value = String(md5OrUsername || '').trim()
if (!value) return value
const bySession = this.getUsernameByMd5(value)
if (bySession) return bySession
// 有些群只以 `Chat_<md5>` 表存在、不在 session 列表里;退回按聊天表映射解析
// (与 wechat-db 的 `chatMd5ToUsername` 同一套依据)。
try {
const byChatTable = this.getChatTables().find((table) => table.name === `Chat_${value}`)
if (byChatTable?.db_number) return byChatTable.db_number
} catch {
// 映射不可用时退回原值。
}
return value
}
/**
* 归一化图片消息的时间范围参数。
*
* 同时接受旧的裸 `sinceMs` 与新的 `{ sinceMs, beforeMs }`:
* 前者有若干既有调用点(统计卡片、增量水位),不该为了新功能去改它们;
* 后者是 recent-first 分段计划要的半开区间。
*/
private normalizeImageRange(
input?: number | { sinceMs?: number; beforeMs?: number }
): { sinceMs?: number; beforeMs?: number } {
if (typeof input === 'number') {
return Number.isFinite(input) && input > 0 ? { sinceMs: input } : {}
}
if (!input) return {}
const range: { sinceMs?: number; beforeMs?: number } = {}
if (typeof input.sinceMs === 'number' && Number.isFinite(input.sinceMs) && input.sinceMs > 0) {
range.sinceMs = input.sinceMs
}
if (
typeof input.beforeMs === 'number' &&
Number.isFinite(input.beforeMs) &&
input.beforeMs > 0
) {
range.beforeMs = input.beforeMs
}
return range
}
/**
* 图片消息的 WHERE 片段;`[sinceMs, beforeMs)` 半开区间。
*
* 边界换算**只走** `imageTextWindowToSeconds`:COUNT 与列表两条路径必须用
* 同一份换算,否则"统计说有 3 张、列表却返回 2 张",而调用方会据此把一段
* 标成"已覆盖"。
*/
private imageMessageWhere(
column: string,
input?: number | { sinceMs?: number; beforeMs?: number }
): string {
const clauses = [`(${this.quoteSqlIdentifier(column)} & 65535) = 3`]
// 微信的 create_time 是**秒**,`[sinceMs, beforeMs)` 转成秒闭区间
// `[sinceSec, beforeSecInclusive]`;两端同一套下取整,保证不重不漏。
const { sinceSec, beforeSecInclusive } = imageTextWindowToSeconds(
this.normalizeImageRange(input)
)
if (sinceSec !== null) clauses.push(`"create_time" >= ${sinceSec}`)
if (beforeSecInclusive !== null) clauses.push(`"create_time" <= ${beforeSecInclusive}`)
return clauses.join(' AND ')
}
/**
* 统计图片消息条数。
*
* 与 `countVoiceMessagesAsync` 同构:纯 SQL COUNT,**不解密任何图片** ——
* 这是「点击索引前先告诉用户有多少张图片」能足够快的前提。
*
* 与语音版本的关键差别:这里**必须区分「0 张」与「统计失败」**。
* `count: null` 表示没数成,调用方绝不能把它当成 0。
*/
async countImageMessagesAsync(
md5OrUsername: string,
input?: number | { sinceMs?: number; beforeMs?: number }
): Promise<ImageMessageCountProbe> {
if (!this.wcdbGetMessageTableStats || !this.wcdbExecQuery) {
return { count: null, typeColumn: null, error: '当前数据服务不支持消息表统计' }
}
const username = this.resolveMessageUsername(md5OrUsername)
if (!username) {
return { count: null, typeColumn: null, error: '无法解析该会话的标识' }
}
let tables: Wcdb4MessageStore[]
try {
tables = await this.listMessageStoresAsync(username)
} catch {
return { count: null, typeColumn: null, error: '读取消息分片失败' }
}
if (!tables.length) {
return { count: null, typeColumn: null, error: '未找到该会话的消息表' }
}
let total = 0
let typeColumn: string | null = null
for (const table of tables) {
const column = this.resolveMessageTypeColumn(table)
if (!column) {
return { count: null, typeColumn: null, error: '消息表缺少可识别的消息类型列' }
}
if (!typeColumn) typeColumn = column
try {
const rows = await this.callJsonAsync<Record<string, unknown>[]>(
this.wcdbExecQuery as unknown as KoffiAsyncFunction,
'message',
table.dbPath,
`SELECT COUNT(*) AS "image_count" FROM ${this.quoteSqlIdentifier(table.tableName)} WHERE ${this.imageMessageWhere(column, input)}`
)
const value = Number(this.pickValue(rows[0] || {}, ['image_count', 'count', 'COUNT(*)']))
if (Number.isFinite(value)) total += value
} catch {
return { count: null, typeColumn: null, error: '图片消息统计查询失败' }
}
}
return { count: total, typeColumn }
}
/**
* 图片消息的增量水位:`count` + `max(local_id)`。
*
* 为什么不能只靠 `countImageMessagesAsync`:
* 图片总数相同**不代表**图片集合没变。撤回一张旧图 + 新增一张新图,count 不变,
* 但新图的 `local_id` 更大。只看 count 会静默跳过该会话,新图片永远搜不到。
*
* `local_id` 是 WCDB 每张消息表内的插入序(自增),所以:
* - 任何 append → `max_local_id` 严格变大;
* - 「删旧 + 增新」且总数不变 → `max_local_id` 也变大,照样被发现;
* - 只有「删掉非最大的那张且不新增」才不变,而此时集合缩小、无需重扫。
*
* 仍然是一条 SQL 聚合,**不解密任何图片**,成本与 count 同量级。
*/
async imageConversationWatermarkAsync(
md5OrUsername: string,
input?: number | { sinceMs?: number; beforeMs?: number }
): Promise<{ count: number; maxLocalId: number } | null> {
if (!this.wcdbGetMessageTableStats || !this.wcdbExecQuery) return null
// 与计数同因:必须先把会话 md5 解析成原生接口要的 username,否则永远匹配不到消息表。
const username = this.resolveMessageUsername(md5OrUsername)
if (!username) return null
let tables: Wcdb4MessageStore[]
try {
tables = await this.listMessageStoresAsync(username)
} catch {
return null
}
let count = 0
let maxLocalId = 0
for (const table of tables) {
// 同样探测真实列名:硬编码列名会让水位查询静默失败,进而退化成"永远重扫"或"永远跳过"。
const column = this.resolveMessageTypeColumn(table)
if (!column) return null
try {
const rows = await this.callJsonAsync<Record<string, unknown>[]>(
this.wcdbExecQuery as unknown as KoffiAsyncFunction,
'message',
table.dbPath,
`SELECT COUNT(*) AS "image_count", MAX("local_id") AS "image_max_local_id" FROM ${this.quoteSqlIdentifier(table.tableName)} WHERE ${this.imageMessageWhere(column, input)}`
)
const row = rows[0] || {}
const tableCount = Number(this.pickValue(row, ['image_count', 'count', 'COUNT(*)']))
const tableMax = Number(
this.pickValue(row, ['image_max_local_id', 'max_local_id', 'MAX("local_id")'])
)
if (Number.isFinite(tableCount)) count += tableCount
if (Number.isFinite(tableMax) && tableMax > maxLocalId) maxLocalId = tableMax
} catch (error) {
console.warn(
`[WCDB4] image watermark failed username=${username} db=${table.dbPath} table=${table.tableName}:`,
error
)
return null
}
}
return { count, maxLocalId }
}
private readSessionRows(): Record<string, unknown>[] {
if (!this.wcdbGetSessions) return []
const rows = this.callJson<Record<string, unknown>[]>((handle, outJson) =>
@@ -1329,6 +1617,99 @@ export class Wcdb4Client {
return this.finalizeMessages(username, allRows, startTime, endTime, limit)
}
/**
* 只读**图片消息**,供图片文字索引使用。
*
* 为什么需要它:`getMessagesAsync` 会把整个会话的消息都读出来,
* 一个 20 万条消息的会话要花十几秒(实测 `rawReadMs≈15s`),
* 而图片索引只关心其中的图片 —— 那是错误的数据边界。
*
* 实现上刻意**复用 `finalizeMessages`**(即 `normalizeMessage` + 群昵称解析 + 排序),
* 这样产出的 `Wcdb4Message` 与全量路径**逐字段同构**,`messageId` / `contentData`
* 语义完全一致 —— 否则 artifact / binding / checkpoint 的键会全变。
* 变化的只有"读哪些行":靠 `imageMessageWhere` 在 SQL 层过滤。
*
* 这里仍是一次性读完该会话的图片(**未分页**):行数由图片数量决定而不是消息数量,
* 已经比全量小一到两个数量级。超过 `limit` 会被截断并告警,调用方应改成分页。
*/
async listImageMessagesAsync(
md5OrUsername: string,
options: {
sinceMs?: number
/**
* 开区间上界。recent-first 的分段窗口靠它把"最近 30 天"和"更早"切开,
* 与 `sinceMs` 一起构成 `[sinceMs, beforeMs)`。
*/
beforeMs?: number
limit?: number
/**
* 行序。默认 `asc` 保持既有行为不变;recent-first 的窗口用 `desc`,
* 让同一窗口内**新的图片先被处理**(用户先受益,且断点续跑更有意义)。
*/
order?: 'asc' | 'desc'
requestId?: string
} = {}
): Promise<Wcdb4Message[]> {
if (!this.wcdbExecQuery) return []
const requestId = options.requestId ?? 'NO-REQUEST'
const username = this.resolveMessageUsername(md5OrUsername)
if (!username) return []
const startedAt = Date.now()
let tables: Wcdb4MessageStore[] = []
try {
tables = await this.listMessageStoresAsync(username)
} catch (error) {
console.warn('[WCDB4] image message table stats failed:', error)
return []
}
if (!tables.length) return []
const limit = Math.max(1, options.limit ?? 200_000)
const allRows: Record<string, unknown>[] = []
let successfulTables = 0
for (const table of tables) {
// 真实类型列名必须逐表探测:硬编码会让过滤静默失效,把全量消息当图片读回来。
const column = this.resolveMessageTypeColumn(table)
if (!column) continue
try {
const where = this.imageMessageWhere(column, options)
// `local_id` 参与排序:`create_time` 同秒的消息需要一个稳定次序,
// 否则多次读取的行序可能不同,调用方无法做稳定游标。
const direction = options.order === 'desc' ? 'DESC' : 'ASC'
const sql = `SELECT * FROM ${this.quoteSqlIdentifier(table.tableName)} WHERE ${where} ORDER BY "create_time" ${direction}, "local_id" ${direction} LIMIT ${limit}`
const queryStartedAt = Date.now()
const rows = await this.callJsonAsync<Record<string, unknown>[]>(
this.wcdbExecQuery as unknown as KoffiAsyncFunction,
'message',
table.dbPath,
sql
)
successfulTables += 1
if (Array.isArray(rows)) allRows.push(...rows)
wcdbDebugLog(
`[${requestId}] WCDB image messages table=${table.tableName} rows=${Array.isArray(rows) ? rows.length : 0} cost=${Date.now() - queryStartedAt}ms`
)
} catch (error) {
console.warn(
`[WCDB4] image message scan failed table=${table.tableName}:`,
error
)
}
}
if (tables.length > 0 && successfulTables === 0) return []
const messages = this.finalizeMessages(username, allRows)
if (allRows.length >= limit) {
// 不静默丢数据:截断会让该会话被标成"处理完了",下一遍靠水位修正。
console.warn(`[WCDB4] image message scan truncated rows=${allRows.length} limit=${limit}`)
}
wcdbDebugLog(
`[${requestId}] WCDB image messages end rows=${messages.length} cost=${Date.now() - startedAt}ms`
)
return messages
}
private async getMessagesByTableScanAsync(
username: string,
startTime?: number,
@@ -2342,6 +2723,63 @@ export class Wcdb4Client {
return this.accountRoot
}
/**
* 只读 roominfo(contact.db / chatroom 或兼容列名)。
*/
getRoomInfo(chatroomId: string): RoomInfoRow | null {
if (!this.wcdbExecQuery || !chatroomId.endsWith('@chatroom')) return null
const escaped = chatroomId.replace(/'/g, "''")
const sql = `SELECT * FROM chatroom WHERE chatroomname = '${escaped}' LIMIT 1`
try {
const rows = this.callJson<Record<string, unknown>[]>((handle, outJson) =>
this.wcdbExecQuery!(handle, 'contact', '', sql, outJson)
)
return rows?.[0] ? normalizeRoomInfoRow(chatroomId, rows[0]) : null
} catch {
return null
}
}
async getRoomInfoAsync(chatroomId: string): Promise<RoomInfoRow | null> {
if (!this.wcdbExecQuery || !chatroomId.endsWith('@chatroom')) return null
const escaped = chatroomId.replace(/'/g, "''")
const sql = `SELECT * FROM chatroom WHERE chatroomname = '${escaped}' LIMIT 1`
try {
const rows = await this.callJsonAsync<Record<string, unknown>[]>(
this.wcdbExecQuery as unknown as KoffiAsyncFunction,
'contact',
'',
sql
)
return rows?.[0] ? normalizeRoomInfoRow(chatroomId, rows[0]) : null
} catch (error) {
console.warn('[WCDB4] getRoomInfoAsync failed:', error)
return null
}
}
/** 只读列收藏(favorite.db / fav_db_item),供导出与搜索。 */
async listFavoriteItems(limit = 200): Promise<Record<string, unknown>[]> {
if (!this.wcdbExecQuery) return []
const sql = `SELECT local_id, server_id, type, update_time, content, fromusr, realchatname FROM fav_db_item ORDER BY update_time DESC LIMIT ${Math.max(
1,
Math.min(1000, limit)
)}`
try {
const rows = await this.callJsonAsync<Record<string, unknown>[]>(
this.wcdbExecQuery as unknown as KoffiAsyncFunction,
'favorite',
path.join(this.accountRoot, 'db_storage/favorite/favorite.db'),
sql
)
return Array.isArray(rows) ? rows : []
} catch (error) {
console.warn('[WCDB4] listFavoriteItems failed:', error)
return []
}
}
getKey(): string {
return this.key
}
+88 -5
View File
@@ -5,7 +5,10 @@ import {
GroupReportExportResult,
GroupReportRenderSnapshotExportRequest
} from '../shared/group-report'
import type { InstalledReportTemplate, ReportTemplateOperationResult } from '../shared/report-template-package'
import type {
InstalledReportTemplate,
ReportTemplateOperationResult
} from '../shared/report-template-package'
import type {
ReportTemplateCatalogInstallResult,
ReportTemplateCatalogResult
@@ -57,7 +60,19 @@ import type {
ImageCandidateQuery,
ImageInsight
} from '../shared/image-insight'
import type { SystemOcrCapability, SystemOcrRequest, SystemOcrResult } from '../shared/system-ocr'
import type {
ImageTextIndexCountResult,
ImageTextIndexRepairResult,
ImageTextIndexStartOptions,
ImageTextIndexStatus
} from '../shared/image-text-index'
import type { AgentHubActionResult, AgentHubLogEntry, AgentHubStatus } from '../shared/agent-hub'
import type {
AgentHubConversation,
AgentHubConversationMessage,
AgentHubConversationSummary
} from '../shared/agent-hub-conversation'
import type {
PersonalWechatGeneratedTtsVoiceRequest,
PersonalWechatGeneratedTtsVoiceResult,
@@ -94,7 +109,7 @@ import type {
AppUpdateOpenDownloadPageResult,
AppUpdateState
} from '../shared/app-update'
import type { CacheSummary } from '../shared/cache'
import type { CacheClearScope, CacheSummary } from '../shared/cache'
import type { ExportRequest, ExportJobProgress, ExportResult } from '../shared/export'
import type {
VoiceBatchPreflight,
@@ -140,6 +155,37 @@ import type {
TextToSpeechSettingsResult
} from '../shared/text-to-speech'
export type TransferPaymentInfo = {
paySubtype?: string
amountText?: string
transcationId?: string
transferId?: string
invalidTime?: string
beginTransferTime?: string
effectiveDate?: string
payMemo?: string
receiverUsername?: string
payerUsername?: string
transferStatus?: string
transferStatusText?: string
}
export type RedPacketPaymentInfo = {
templateId?: string
receiveTitle?: string
sendTitle?: string
sceneText?: string
senderDes?: string
receiverDes?: string
iconUrl?: string
nativeUrl?: string
sendId?: string
hbType?: string
hbStatus?: string
receiveStatus?: string
redPacketStatusText?: string
}
export type ParsedContent =
| { type: 'text'; content: string }
| { type: 'voice'; duration?: number }
@@ -152,6 +198,7 @@ export type ParsedContent =
url: string
appname?: string
typeVal?: string
transfer?: TransferPaymentInfo
}
| {
type: 'miniProgram'
@@ -163,7 +210,13 @@ export type ParsedContent =
thumbDatName?: string
thumbDataUrl?: string
}
| { type: 'redPacket'; title: string; description?: string; url?: string }
| {
type: 'redPacket'
title: string
description?: string
url?: string
pay?: RedPacketPaymentInfo
}
| { type: 'voip'; duration?: number; status: string; roomType?: number }
| { type: 'image'; md5?: string; datName?: string; aeskey?: string; encrypVer?: number }
| {
@@ -210,7 +263,7 @@ declare global {
openAppUpdateDownloadPage: () => Promise<AppUpdateOpenDownloadPageResult>
onAppUpdateState: (callback: (state: AppUpdateState) => void) => () => void
getCacheSummary: () => Promise<CacheSummary>
clearCache: (scope: 'bootstrap' | 'electron' | 'knowledge' | 'all') => Promise<CacheSummary>
clearCache: (scope: CacheClearScope) => Promise<CacheSummary>
openKnowledgeDirectory: () => Promise<{ success: boolean; error?: string }>
initDb: (key: string, accountRoot: string) => Promise<boolean | DatabaseInitResult>
discoverAccounts: (inputPath: string) => Promise<AccountDiscoveryResult>
@@ -363,7 +416,10 @@ declare global {
cancelVoiceModelDownload: () => Promise<{ success: boolean }>
removeVoiceModel: () => Promise<VoiceModelStatus>
openVoiceModelDirectory: () => Promise<{ success: boolean; error?: string }>
recognizeVoice: (reference: VoiceMessageReference) => Promise<VoiceRecognitionResult>
recognizeVoice: (
reference: VoiceMessageReference,
options?: { force?: boolean }
) => Promise<VoiceRecognitionResult>
getVoiceTranscriptSnapshot: (
reference: VoiceMessageReference
) => Promise<VoiceTranscriptSnapshot>
@@ -660,6 +716,23 @@ declare global {
sessionId: string,
limit?: number
) => Promise<{ success: boolean; insights: ImageInsight[] }>
// 本地图片文字识别(System OCR,本地 Runtime,非 AI Provider)
getSystemOcrCapability: () => Promise<SystemOcrCapability>
recognizeLocalImageText: (request: SystemOcrRequest) => Promise<SystemOcrResult>
getImageTextIndexStatus: () => Promise<ImageTextIndexStatus>
countImageMessages: (sinceMs?: number) => Promise<ImageTextIndexCountResult>
startImageTextIndex: (
options?: ImageTextIndexStartOptions
) => Promise<{ started: boolean; state: string }>
pauseImageTextIndex: () => Promise<{ paused: boolean; state: string }>
resumeImageTextIndex: (
options?: ImageTextIndexStartOptions
) => Promise<{ started: boolean; state: string }>
cancelImageTextIndex: () => Promise<{ cancellable: boolean; cancelled: boolean }>
clearImageTextIndex: () => Promise<{ removed: boolean; removedBytes: number }>
resetImageTextIndexFailures: () => Promise<{ reset: number }>
repairImageTextIndex: () => Promise<ImageTextIndexRepairResult>
onImageTextIndexStatus: (callback: (status: ImageTextIndexStatus) => void) => () => void
getPersonalWechatSenderStatus: () => Promise<PersonalWechatSenderStatus>
getPersonalWechatSendCapability: () => Promise<PersonalWechatSendCapability>
getPersonalWechatKeepOneBotProcess: () => Promise<boolean>
@@ -722,6 +795,16 @@ declare global {
reconnectAgentHub: () => Promise<AgentHubActionResult>
disconnectAgentHub: () => Promise<AgentHubActionResult>
selectAgentHubTestImage: () => Promise<{ canceled: boolean; path?: string }>
getAgentHubConversations: () => Promise<AgentHubConversationSummary[]>
getAgentHubConversation: (userId: string) => Promise<AgentHubConversation | null>
clearAgentHubConversations: () => Promise<{ success: boolean }>
onAgentHubConversation: (
callback: (payload: {
summary: AgentHubConversationSummary
message: AgentHubConversationMessage
}) => void
) => () => void
onAgentHubConversationsCleared: (callback: () => void) => () => void
onAgentHubStatus: (callback: (status: AgentHubStatus) => void) => () => void
onAgentHubLog: (callback: (entry: AgentHubLogEntry) => void) => () => void
}
+82 -7
View File
@@ -30,7 +30,18 @@ import type {
ImageCandidateQuery,
ImageInsight
} from '../shared/image-insight'
import type { SystemOcrCapability, SystemOcrRequest, SystemOcrResult } from '../shared/system-ocr'
import type {
ImageTextIndexCountResult,
ImageTextIndexRepairResult,
ImageTextIndexStartOptions,
ImageTextIndexStatus
} from '../shared/image-text-index'
import type { AgentHubLogEntry, AgentHubStatus } from '../shared/agent-hub'
import type {
AgentHubConversationMessage,
AgentHubConversationSummary
} from '../shared/agent-hub-conversation'
import type {
PersonalWechatGeneratedTtsVoiceRequest,
PersonalWechatGeneratedTtsVoiceResult,
@@ -62,7 +73,7 @@ import type { AppLogEntry } from '../shared/app-log'
import type { AppUpdateState } from '../shared/app-update'
import type { GroupExitMonitorState } from '../shared/group-exit-monitor'
import type { ActionLogEntry } from '../shared/action-log'
import type { CacheSummary } from '../shared/cache'
import type { CacheClearScope, CacheSummary } from '../shared/cache'
import type { ExportRequest, ExportJobProgress } from '../shared/export'
import type { ImageDecoderSelectionResult, ImageDecoderStatus } from '../shared/image-decryption'
import type { AccountDiscoveryResult } from '../shared/database-key'
@@ -122,7 +133,7 @@ const api = {
return () => ipcRenderer.removeListener('app-update:state', listener)
},
getCacheSummary: (): Promise<CacheSummary> => ipcRenderer.invoke('cache:getSummary'),
clearCache: (scope: 'bootstrap' | 'electron' | 'knowledge' | 'all'): Promise<CacheSummary> =>
clearCache: (scope: CacheClearScope): Promise<CacheSummary> =>
ipcRenderer.invoke('cache:clear', scope),
openKnowledgeDirectory: (): Promise<{ success: boolean; error?: string }> =>
ipcRenderer.invoke('cache:openKnowledgeDirectory'),
@@ -206,9 +217,7 @@ const api = {
*
* 事件带 requestId:UI 必须只认自己那一次请求,否则用户连问两次时阶段文案会串台。
*/
onAskWechatProgress: (
callback: (requestId: string, event: QueryAgentProgressEvent) => void
) => {
onAskWechatProgress: (callback: (requestId: string, event: QueryAgentProgressEvent) => void) => {
const listener = (
_event: Electron.IpcRendererEvent,
requestId: string,
@@ -276,8 +285,10 @@ const api = {
removeVoiceModel: (): Promise<VoiceModelStatus> => ipcRenderer.invoke('voice:removeModel'),
openVoiceModelDirectory: (): Promise<{ success: boolean; error?: string }> =>
ipcRenderer.invoke('voice:openModelDirectory'),
recognizeVoice: (reference: VoiceMessageReference): Promise<VoiceRecognitionResult> =>
ipcRenderer.invoke('voice:recognize', reference),
recognizeVoice: (
reference: VoiceMessageReference,
options?: { force?: boolean }
): Promise<VoiceRecognitionResult> => ipcRenderer.invoke('voice:recognize', reference, options),
getVoiceTranscriptSnapshot: (
reference: VoiceMessageReference
): Promise<VoiceTranscriptSnapshot> =>
@@ -462,6 +473,48 @@ const api = {
limit?: number
): Promise<{ success: boolean; insights: ImageInsight[] }> =>
ipcRenderer.invoke('image:listInsights', sessionId, limit),
// 本地图片文字识别(System OCR,本地 Runtime,非 AI Provider)
getSystemOcrCapability: (): Promise<SystemOcrCapability> =>
ipcRenderer.invoke('system-ocr:getCapability'),
recognizeLocalImageText: (request: SystemOcrRequest): Promise<SystemOcrResult> =>
ipcRenderer.invoke('system-ocr:recognize', request),
// 图片文字索引(微信图片 → 本地解密 → System OCR → 派生文本 → Knowledge)
getImageTextIndexStatus: (): Promise<ImageTextIndexStatus> =>
ipcRenderer.invoke('image-text-index:getStatus'),
/** 点击索引前的快速统计(SQL COUNT,不解密图片)。 */
countImageMessages: (sinceMs?: number): Promise<ImageTextIndexCountResult> =>
ipcRenderer.invoke('image-text-index:count', sinceMs),
startImageTextIndex: (
options?: ImageTextIndexStartOptions
): Promise<{ started: boolean; state: string }> =>
ipcRenderer.invoke('image-text-index:start', options),
pauseImageTextIndex: (): Promise<{ paused: boolean; state: string }> =>
ipcRenderer.invoke('image-text-index:pause'),
resumeImageTextIndex: (
options?: ImageTextIndexStartOptions
): Promise<{ started: boolean; state: string }> =>
ipcRenderer.invoke('image-text-index:resume', options),
cancelImageTextIndex: (): Promise<{ cancellable: boolean; cancelled: boolean }> =>
ipcRenderer.invoke('image-text-index:cancel'),
clearImageTextIndex: (): Promise<{ removed: boolean; removedBytes: number }> =>
ipcRenderer.invoke('image-text-index:clear'),
/** 只重置失败记录(成功记录与其它数据不动),供"修好代码后重跑"。 */
resetImageTextIndexFailures: (): Promise<{ reset: number }> =>
ipcRenderer.invoke('image-text-index:resetFailures'),
/**
* 派生索引修复:只重建 Knowledge 里的图片派生条目与 FTS。
*
* 已有的 OCR 结果(L1)一条都不动 —— 修复索引问题永远不该让几万张图片重算。
*/
repairImageTextIndex: (): Promise<ImageTextIndexRepairResult> =>
ipcRenderer.invoke('image-text-index:repair'),
onImageTextIndexStatus: (callback: (status: ImageTextIndexStatus) => void) => {
const listener = (_event: Electron.IpcRendererEvent, status: ImageTextIndexStatus): void =>
callback(status)
ipcRenderer.on('image-text-index:status', listener)
return () => ipcRenderer.removeListener('image-text-index:status', listener)
},
getPersonalWechatSenderStatus: (): Promise<PersonalWechatSenderStatus> =>
ipcRenderer.invoke('wechat-personal:getStatus'),
getPersonalWechatSendCapability: (): Promise<PersonalWechatSendCapability> =>
@@ -559,6 +612,28 @@ const api = {
reconnectAgentHub: () => ipcRenderer.invoke('agent-hub:reconnect'),
disconnectAgentHub: () => ipcRenderer.invoke('agent-hub:disconnect'),
selectAgentHubTestImage: () => ipcRenderer.invoke('agent-hub:selectTestImage'),
getAgentHubConversations: () => ipcRenderer.invoke('agent-hub:getConversations'),
getAgentHubConversation: (userId: string) =>
ipcRenderer.invoke('agent-hub:getConversation', userId),
clearAgentHubConversations: () => ipcRenderer.invoke('agent-hub:clearConversations'),
onAgentHubConversation: (
callback: (payload: {
summary: AgentHubConversationSummary
message: AgentHubConversationMessage
}) => void
) => {
const listener = (
_event: Electron.IpcRendererEvent,
payload: { summary: AgentHubConversationSummary; message: AgentHubConversationMessage }
): void => callback(payload)
ipcRenderer.on('agent-hub:conversation', listener)
return () => ipcRenderer.removeListener('agent-hub:conversation', listener)
},
onAgentHubConversationsCleared: (callback: () => void) => {
const listener = (): void => callback()
ipcRenderer.on('agent-hub:conversationsCleared', listener)
return () => ipcRenderer.removeListener('agent-hub:conversationsCleared', listener)
},
onAgentHubStatus: (callback: (status: AgentHubStatus) => void) => {
const listener = (_event: Electron.IpcRendererEvent, status: AgentHubStatus): void =>
callback(status)

Some files were not shown because too many files have changed in this diff Show More