docs: update README and documentation to clarify model configuration and restart requirements

- Enhanced instructions for restarting Cursor after upgrades or initial model configuration to ensure proper functionality.
- Updated references from "Troubleshooting" to "Frequently Asked Questions" for better alignment with user needs.
- Improved clarity in the installation and usage steps across multiple language versions.
This commit is contained in:
leokun
2026-08-28 16:12:06 +08:00
parent d0033a768b
commit e874d68b79
19 changed files with 323 additions and 483 deletions
View File
@@ -1,117 +0,0 @@
---
title: Why We Rebuilt the Codebase
description: From the v0.0.49 Go monolith to a Rust rewrite — a deep dive into the harness, the protocol layer, the agent loop engine, architectural invariants, and the ecosystem.
---
In the [v1.0 Roadmap](https://github.com/leookun/cursor-byok/discussions/32) we wrote:
> We have to admit that vibe coding brought huge advantages, but also a heavy burden. Everything except the code is now perfectly clear, so we decided to rebuild the codebase.
The product direction, what users need, how the protocol works — all clear. The code itself had become the biggest obstacle. This post goes deep: what exactly was wrong with the old version, how each layer of the new architecture is designed, and which invariants hold it together.
## v0.0.49: a monolith at its product ceiling
The pre-rewrite version (preserved on [archive/v0.0.49](https://github.com/leookun/cursor-byok/tree/archive/v0.0.49)) was a Go monolith with a webview frontend: 17 packages under `internal/` — `mitm`, `certs`, `cursoraccount`, `modelchannel`, `backend`, `bridge`… The product model was **full replacement**: once started, it took over all of Cursor's backend traffic and substituted the login state with a generated fake account.
That premise dictated everything: a start/stop switch, a fake account system to maintain, and official plugins and codebase indexing going dark while taken over. Users chose between the "official world" and the "local world". No patch could change the premise — mixing official and local models freely required a new one.
Meanwhile, code grown through rapid iteration had no dedicated protocol layer: every Cursor client update (new task modes, multitask, checkpoints) meant poking holes throughout the business logic. Boundaries by convention, state scattered everywhere, regression by hand. That is what "a heavy burden" meant concretely.
## The architecture after the rebuild
The new version is a Rust workspace where the boundaries are the directories:
```text
cursor-byok
├── server/src/
│ ├── harness/ # Local CA and MITM proxy: first-stage traffic routing
│ ├── cursor/ # Cursor protocol layer: session actors, projection, prompt compiler, tool dispatch
│ ├── run/ # Agent loop engine: model cycles × tool rounds
│ ├── provider/ # Model adapters: three protocols normalized into one event stream
│ ├── store/ # SQLite: append-only revision DAG + content-addressed storage
│ └── web/ # Web retrieval
├── crates/
│ ├── semble-core/ # Code indexing and hybrid retrieval
│ └── semble-mcp/ # Exposed as an MCP server
└── apps/desktop/ # Tauri + React desktop app
```
Let's walk through it along the lifecycle of a request.
## Harness: "coexistence" is a two-stage routing decision
The new version takes over exactly one thing — model calls. In code, that sentence is a precise two-stage decision:
**Stage one, at the MITM proxy.** A local CA (a 3072-bit RSA root that never leaves your machine) lets the proxy inspect Cursor's HTTPS. Only `*.cursor.sh` traffic is ever TLS-intercepted (`is_cursor_host`); everything else is tunneled untouched. Within Cursor traffic, only a fixed path allowlist — agent runs (`AgentService/RunSSE`, `BidiService/BidiAppend`), model catalogs, account display endpoints — is rewritten to the local service. Sign-in, the plugin marketplace, codebase indexing, and the rest go straight to official servers. Tab completion follows the user's setting: "Direct" passes through; the public or self-hosted service routes out.
**Stage two, at the local service.** Even for locally routed requests, the server decodes the first `AgentClientMessage` and looks up its `model_id` in the local model store. Found: handle locally (BYOK). Not found — the user picked an official model — buffer and forward the request back to `api2.cursor.sh`. A Cursor run is a pair of requests (upstream BidiAppend + a long-lived RunSSE stream), and both must take the same path, so a `wait_route` synchronization point blocks the SSE stream until the run is classified `Local` or `Upstream`.
This is exactly how the fake account and the start/stop switch disappeared: Auto and official models keep flowing to official billing; only your configured models enter the local path — one login, one client, routed per request.
## The protocol layer: Cursor's protocol behind one door
`server/src/cursor/` translates Cursor's private protocol into neutral events the engine understands. When the client updates, this is the only layer that moves. A few designs worth unpacking:
**One actor per request.** Each Cursor run gets a `CursorActor` holding a command mailbox, a cancellation token, and a replayable `OutputHub` — late SSE subscribers get all buffered frames replayed, so reconnects lose nothing. Upstream messages pass through an `OrderedInbox` keyed by sequence number: out-of-order arrivals are buffered, duplicates dropped. The protocol layer tolerates network nondeterminism by construction.
**Projection is the heart of the layer.** Internally there is one canonical message type (`CanonicalMessage`); the checkpoints sent to Cursor and the requests sent to model providers are both **projections** of it. The Cursor-side checkpoint has a hard check: stable history may only grow. If re-projection finds that a previously generated root "changed" or "shrank", it errors out — failing loudly beats silently corrupting state.
**Prompts are compiled.** Seven modes — Agent, Ask, Plan, Debug, Multitask, Subagent, Compaction — each built from a prompt template, a runtime template, and a tool manifest. The compiler substitutes placeholders, trims tools by model capability (no image generation → the tool is removed), and appends dynamic MCP tools in **deterministic order**. Determinism is not a style preference; the caching section below shows it is a correctness requirement.
**Tools execute on the client.** Shell commands, file edits, and searches requested by the model are not executed by the server — they are encoded as instructions for the Cursor client to run in the user's environment, with results returned via BidiAppend. Concurrent edits to the same path are serialized by a scheduler, and streaming arguments are projected into live UI previews.
## The loop engine: a minimal agent kernel
`server/src/run/` is the execution kernel, fully independent of the Cursor protocol. A run is a loop:
```text
loop {
project history → call model (consume_model_cycle)
├── no tool calls → append final reply, finish
└── tool calls → execute tool round → commit results → continue
}
```
**The model cycle is a strict state machine.** `consume_model_cycle` consumes the normalized event stream and enforces protocol invariants: `Start` must come first, text/thinking/tool blocks must open and close in pairs, and the finish reason must agree with the presence of tool calls. On cancellation it cleanly closes any open blocks before exiting — otherwise Cursor would merge deltas from two different outputs.
**Commit barriers.** Every state commit (initial messages, settled tool results, compaction, the final reply) waits for the Cursor checkpoint worker to acknowledge the write before the engine proceeds. Checkpoint publication is a hard synchronization point — the engine never runs ahead of published state.
**Auto-compaction is an explicit prefix reset.** The engine calibrates token estimates with the latest real usage; when estimated input exceeds `context window − 10,000` reserved tokens, it compacts: retain **exactly one** latest request-context message, summarize the rest (capped at 4,096 output tokens, reasoning disabled), then rebuild the revision in deterministic order. If compaction is interrupted, a fallback summary of the most recent text keeps the loop going.
**Cancellation and recovery.** Cancellation tokens propagate from the session all the way into the provider stream; activating a new run in the same conversation cancels the previous one. After a restart, an in-flight tool round can be reconstructed from the Cursor checkpoint and resumed (`RunAction::Resume`) — which is only possible because projection is bidirectional.
## The invariant: prefix cache stability
This is where "structure determines correctness" shows most clearly. Providers cache prompts by prefix, and the price difference can be 10×. To hit the cache, history must be an **append-only log**: everything sent to the model last turn must be a byte-level prefix of the next turn.
Convention cannot defend that, so it is built into every layer:
- **Storage**: history is an immutable revision DAG — each revision is a complete ordered message list identified by a SHA-256 content digest; appends create child revisions, identical appends reuse existing nodes, and the only "rewrite" is compaction, which branches from the root. Runtime events have unique identities; retries are idempotent, and reusing an identity with different content is a hard error.
- **Projection**: the system prompt and stable tool prefix are byte-stable when inputs are unchanged. Request context (rules, skills, MCP metadata) is separated from the per-turn runtime message: unchanged content is never re-appended; changed content appends a new message right before the runtime message — old ones are never rewritten. Context going A → B → A appends a third distinct A.
- **Provider**: tool-call IDs are normalized before leaving the server; thinking-block signatures (Anthropic's `signature`, OpenAI Responses' encrypted reasoning items) are preserved as replay state and replayed verbatim next turn.
GPT-family cache hit rates staying consistently high is not luck — it is this entire chain. The repo even carries a dedicated engineering-constraint document (`.agents/skills/cursor-prefix-stability`); every change touching this chain goes through its checklist.
## The provider layer: three protocols, one event stream
Three adapters — Anthropic Messages, OpenAI Chat Completions, OpenAI Responses — emit one normalized event vocabulary (`Start`, open/delta/close for text, thinking and tool blocks, `Usage`, `Done`). The engine and protocol layer never know which upstream they are talking to — which is exactly why "pick OpenAI Chat for everything else" works architecturally.
Every call is observable: sanitized request headers and body, streamed response chunks (buffered with limits), time to first token, token usage — fully replayable in detailed mode. The desktop app's call-details page reads precisely this data.
## Ecosystem: the rebuild is more than code
**The official ecosystem is preserved as-is.** The biggest ecosystem consequence of coexistence is "no subtraction": plugins, Skills, MCP, codebase indexing, and Tab completion all keep working once you sign in with your own account. The old version could not do this because it replaced the entire backend.
**Semantic search is standalone and measurable.** semble implements code indexing and hybrid retrieval as an independent crate, exposed as an MCP server. It ships a reproducible benchmark (real React/Vue commits, queries hand-labeled to implementation lines): natural-language queries reach 95% Recall@5, literal and symbol queries reach 100% — with a public quality-gate file, so any retrieval regression fails CI.
**The Tab service is open at three levels.** Public service (hosted by the author), Direct (your own account's official service), or self-hosted (`cursor-tab-server` lives on the archive branch) — switchable in system settings.
**The engineering system is part of the ecosystem too.** The repo maintains domain-constraint documents (database schema, prefix-cache stability, release process, i18n…) shared by humans and AI working on the code; `make check` runs formatting, linting, the full test suite, and frontend type checks in one command. This documentation site — including the interactive demo on the homepage built from real desktop components and mock data — lives in the same repo and evolves with the product.
**Direction is driven by community discussion.** Tab support, Gemini, image generation, all task modes, and codebase indexing from the roadmap have all landed in the rebuilt version, with more IDEs planned.
## Was it worth it
The rebuild consumed weeks that could have gone into features. In exchange: the product model no longer limits itself; protocol updates have a fixed landing place (only `cursor/` moves); correctness is guarded by types, invariants, and regression tests instead of memory; and every new capability grows on a product that already works.
The problem used to be "everything is clear except the code" — now the code is clear too. Thoughts on the rebuild? Join the [Roadmap discussion](https://github.com/leookun/cursor-byok/discussions/32).
@@ -1,117 +0,0 @@
---
title: 为什么我们重构代码?
description: 从 v0.0.49 的 Go 单体到 Rust 重写——harness、协议层、Agent 循环引擎、架构不变量与生态的深度解析。
---
在 [v1.0 Roadmap](https://github.com/leookun/cursor-byok/discussions/32) 里我们写过这样一段话:
> 不得不承认 vibe coding 带来了巨大的优势,但也带来了沉重的负担。现在我认为除了代码之外的一切内容都十分清晰,所以决定在近期重构代码。
产品方向、用户需要什么、协议怎么工作——这些都清楚,唯独代码本身成了继续前进的最大阻力。这篇文章把这次重构讲透:旧版到底哪里不行,新架构的每一层是怎么设计的,以及哪些不变量在支撑它。
## v0.0.49:一个走到产品上限的单体
重构前的版本(保留在 [archive/v0.0.49](https://github.com/leookun/cursor-byok/tree/archive/v0.0.49) 分支)是一个 Go 单体加 webview 前端:`internal/` 下 17 个包,`mitm`、`certs`、`cursoraccount`、`modelchannel`、`backend`、`bridge`……产品模式是**整体替换**——服务启动后接管 Cursor 的全部后端流量,用内置生成的 fake 账户顶替登录态。
这个前提决定了一切:需要「启动/停止服务」的开关,需要维护假账户体系,官方的插件、代码库索引在接管状态下全部失效。用户在「官方世界」和「本地世界」之间二选一。修补任何一处都动摇不了这个前提——想让官方模型和自己的模型自由混用,必须换前提。
同时,快速迭代长出来的代码没有独立的协议层:Cursor 客户端每次更新(新任务模式、multitask、checkpoint),都要在业务逻辑里到处开洞。模块边界靠约定、状态散落各处、回归靠人肉。这就是「沉重的负担」的具体含义。
## 重构后的架构
新版是一个 Rust workspace,边界即目录:
```text
cursor-byok
├── server/src/
│ ├── harness/ # 本地 CA 与 MITM 代理:流量的第一级路由
│ ├── cursor/ # Cursor 协议层:会话 actor、投影、提示词编译、工具调度
│ ├── run/ # Agent 循环引擎:模型周期 × 工具轮次
│ ├── provider/ # 模型适配:三种协议归一化为统一事件流
│ ├── store/ # SQLite:append-only 修订 DAG + 内容寻址存储
│ └── web/ # 网页检索能力
├── crates/
│ ├── semble-core/ # 代码索引与混合检索
│ └── semble-mcp/ # 以 MCP 服务形态暴露
└── apps/desktop/ # Tauri + React 桌面端
```
下面按请求的生命周期,逐层拆开。
## Harness:「并存」是一个两级路由决策
新版只接管一件事——模型调用。这句话在代码里是一个精确的两级决策:
**第一级在 MITM 代理层。** 本地 CA(3072 位 RSA 根证书,只存在于你的机器)让代理能解析 Cursor 发出的 HTTPS。但只有 `*.cursor.sh` 的流量会被 TLS 解析(`is_cursor_host`),其余一律原样隧道。解析后再看路径:只有一个固定白名单——Agent 运行(`AgentService/RunSSE`、`BidiService/BidiAppend`)、模型目录、账号信息展示等——会被改写到本地服务;登录、插件市场、代码库索引等其余请求原封不动发往官方服务器。Tab 补全按用户设置:选「直连」就直通官方,选公益或自建才路由出去。
**第二级在本地服务层。** 即使请求进了本地,服务端解码首个 `AgentClientMessage` 后按 `model_id` 查本地模型配置:查到,本地处理(BYOK);查不到——说明用户选的是官方模型——整个请求缓冲后转发回 `api2.cursor.sh`。Cursor 的运行由一对请求组成(上行 BidiAppend + 下行 RunSSE 流),两者必须走同一条路,所以有一个 `wait_route` 同步点:SSE 流会阻塞等待路由分类(`Local` 或 `Upstream`)后再跟随。
fake 账户、启停开关就是这样消失的:官方模型选 Auto 或官方型号照常走官方计费,选你配置的模型才进本地——同一个登录态,同一个客户端,逐请求分流。
## 协议层:把 Cursor 协议关进一个模块
`server/src/cursor/` 的职责是把 Cursor 的私有协议翻译成引擎能理解的中立事件,客户端更新时只有这一层需要动。几个值得展开的设计:
**每个请求一个 actor。** 每个 Cursor 运行请求对应一个 `CursorActor`,持有命令信箱、取消令牌和一个可重放的输出中枢(`OutputHub`)——SSE 订阅者迟到时,缓冲的帧会全部重放,客户端断线重连不丢内容。上行消息经过一个按序号重排的 `OrderedInbox`:乱序到达缓冲,重复到达丢弃,协议层天然容忍网络的不确定性。
**投影(projection)是协议层的核心。** 服务端内部只有一种「规范消息」(`CanonicalMessage`),发给 Cursor 的检查点、发给模型服务商的请求,都是从规范消息**投影**出来的两种视图。Cursor 侧的检查点还有一条硬性校验:稳定历史只能追加,如果重投影时发现某个已生成的根「变了」或「变少了」,直接报错——宁可失败也不静默破坏状态。
**提示词是编译出来的。** Agent、Ask、Plan、Debug、Multitask、Subagent、Compaction 七种模式,每种由 prompt 模板 + runtime 模板 + 工具清单构成,编译器负责占位符替换、按模型能力裁剪工具(不支持图片生成就移除对应工具)、以确定性顺序追加动态 MCP 工具。确定性不是风格偏好,后面讲缓存时会看到它是正确性要求。
**工具在客户端执行。** 模型发起的 shell、读写文件、搜索,服务端并不亲自执行,而是编码成指令下发给 Cursor 客户端,由客户端在用户环境里跑,结果再经 BidiAppend 回传。同路径的并发编辑有串行化调度,流式参数还会被投影成 UI 的实时预览。
## Loop 引擎:Agent 循环的最小内核
`server/src/run/` 是与 Cursor 协议无关的执行内核。一次运行就是一个循环:
```text
loop {
投影历史 → 调用模型(consume_model_cycle)
├── 无工具调用 → 追加最终回复,结束
└── 有工具调用 → 执行工具轮次(tool_round) → 提交结果 → 继续循环
}
```
**模型周期是严格的状态机。** `consume_model_cycle` 消费归一化事件流,校验协议不变量:必须先 Start、文本/思考/工具块必须成对开闭、`finish_reason` 与工具调用必须一致。取消发生时它会先干净地闭合未闭合的块再退出——否则 Cursor 端会把两次输出的增量拼在一起。
**提交屏障(commit barrier)。** 引擎每次提交状态(初始消息、工具结果落定、压缩、最终回复),都要等 Cursor 检查点工作线程确认写完才继续。检查点发布是引擎与协议层之间的硬同步点——引擎永远不会跑到已发布状态的前面。
**自动压缩是显式的前缀重置。** 引擎用最近一次真实用量校准 token 估算,当估算输入超过 `上下文窗口 − 10_000` 预留时触发压缩:保留**恰好一条**最新的 request-context 消息,其余历史交给模型摘要(上限 4096 token、关闭思考),然后以确定性顺序重建修订。压缩若被打断,退化为截取最近文本的兜底摘要,循环继续。
**取消与恢复。** 取消令牌从会话一路传到 provider 流;同一会话激活新运行会自动取消旧的。服务重启后,进行中的工具轮次能从 Cursor 检查点反解出来继续执行(`RunAction::Resume`)——这依赖投影是双向的。
## 不变量:前缀缓存稳定性
这是整次重构里最能体现「结构决定正确性」的部分。模型服务商的提示词缓存按前缀命中,收费差可达 10 倍。要吃到缓存,历史就必须是**只追加的日志**:上一轮发给模型的完整内容,必须是下一轮的字节级前缀。
约定守不住这件事,所以它被做进了每一层:
- **存储层**:会话历史是不可变的修订 DAG——每个修订是完整的有序消息列表,由 SHA-256 内容摘要标识;追加产生子修订,内容相同的追加会复用已有节点;唯一能「改写」的操作是压缩,而它是从根分支出新链。运行时事件有唯一标识,重试幂等,内容不同的复用直接报错。
- **投影层**:系统提示词与稳定工具前缀在输入不变时字节级稳定;规则/技能等请求上下文与当轮 runtime 消息分离,内容不变不重复追加,变了就在 runtime 消息前追加新的一条——从不改写旧的。上下文 A→B→A 会追加第三条 A,而不是复用第一条。
- **provider 层**:工具调用 ID 在发给服务商前统一归一化;思考块的签名(Anthropic 的 signature、OpenAI Responses 的加密推理项)被原样保存为回放状态,下一轮逐字回放。
新版 GPT 系列缓存命中率能稳定在高位,靠的不是运气,是这一整条链。仓库里甚至有一份专门的工程约束文档(`.agents/skills/cursor-prefix-stability`),任何触碰这条链的改动都要过它的检查单。
## Provider 层:三种协议,一种事件流
Anthropic Messages、OpenAI Chat Completions、OpenAI Responses 三个适配器,输出统一的归一化事件(`Start`、文本/思考/工具块的开闭与增量、`Usage`、`Done`)。引擎和协议层完全不知道上游是谁——这就是「其他模型无脑选 OpenAI Chat」在架构上成立的原因。
每次调用都有观测:请求头体(脱敏)、逐块流式响应(带缓冲与限额)、首字延迟、token 用量,详细模式下可完整回放。桌面端的调用详情页读的就是这些数据。
## 生态:重构不只是代码
**官方生态原样保留。** 并存设计最大的生态意义是「不减法」:插件、Skills、MCP、代码库索引、Tab 补全,登录自己的账号后全部照常。旧版做不到,是因为它把整个后端都换掉了。
**语义搜索是独立的、可度量的。** semble 作为独立 crate 实现代码索引与混合检索,以 MCP 服务形态接入。它有一套可复现的对比基准(React/Vue 真实提交、人工标注到实现行):自然语言查询 Recall@5 达 95%,字面与符号查询达 100%,并有公开的质量门槛文件——CI 里任何一次检索质量回退都会挡住合并。
**Tab 服务分层开放。** 公益服务(作者部署)、直连(自己账号的官方服务)、自建(`cursor-tab-server` 留在归档分支)三种模式,在系统设置里一键切换。
**工程体系也是生态。** 仓库里维护着一组领域约束文档(数据库 schema、前缀缓存稳定性、发布流程、i18n……),人和 AI 协作时共用同一套检查单;`make check` 一条命令跑完格式、静态检查、全量测试和前端类型检查。这个文档站本身——包括首页那个用真实桌面组件加 mock 数据构建的可交互 demo——也在同一个仓库里,与产品同步演进。
**方向由社区讨论驱动。** Roadmap 里的 Tab 支持、Gemini、图片生成、全部任务模式、代码库索引都已在重构后的版本里落地,更多 IDE 的支持在计划中。
## 值得吗
重构耗掉了数周本可以用来加功能的时间。换来的是:产品模式不再自我设限;协议更新有了固定的落点(只动 `cursor/`);正确性由类型、不变量和回归测试守护,而不是靠记忆;每个新能力都长在一个已经能工作的产品上。
「除了代码之外的一切都清晰」的问题,现在代码也清晰了。对这次重构有想法?欢迎到 [Roadmap 讨论区](https://github.com/leookun/cursor-byok/discussions/32) 聊聊。