--- title: Why We Rebuilt the Codebase description: From the v0.0.49 Go monolith to a Rust rewrite — a deep dive into the harness, the protocol layer, the agent loop engine, architectural invariants, and the ecosystem. --- In the [v1.0 Roadmap](https://github.com/leookun/cursor-byok/discussions/32) we wrote: > We have to admit that vibe coding brought huge advantages, but also a heavy burden. Everything except the code is now perfectly clear, so we decided to rebuild the codebase. The product direction, what users need, how the protocol works — all clear. The code itself had become the biggest obstacle. This post goes deep: what exactly was wrong with the old version, how each layer of the new architecture is designed, and which invariants hold it together. ## v0.0.49: a monolith at its product ceiling The pre-rewrite version (preserved on [archive/v0.0.49](https://github.com/leookun/cursor-byok/tree/archive/v0.0.49)) was a Go monolith with a webview frontend: 17 packages under `internal/` — `mitm`, `certs`, `cursoraccount`, `modelchannel`, `backend`, `bridge`… The product model was **full replacement**: once started, it took over all of Cursor's backend traffic and substituted the login state with a generated fake account. That premise dictated everything: a start/stop switch, a fake account system to maintain, and official plugins and codebase indexing going dark while taken over. Users chose between the "official world" and the "local world". No patch could change the premise — mixing official and local models freely required a new one. Meanwhile, code grown through rapid iteration had no dedicated protocol layer: every Cursor client update (new task modes, multitask, checkpoints) meant poking holes throughout the business logic. Boundaries by convention, state scattered everywhere, regression by hand. That is what "a heavy burden" meant concretely. ## The architecture after the rebuild The new version is a Rust workspace where the boundaries are the directories: ```text cursor-byok ├── server/src/ │ ├── harness/ # Local CA and MITM proxy: first-stage traffic routing │ ├── cursor/ # Cursor protocol layer: session actors, projection, prompt compiler, tool dispatch │ ├── run/ # Agent loop engine: model cycles × tool rounds │ ├── provider/ # Model adapters: three protocols normalized into one event stream │ ├── store/ # SQLite: append-only revision DAG + content-addressed storage │ └── web/ # Web retrieval ├── crates/ │ ├── semble-core/ # Code indexing and hybrid retrieval │ └── semble-mcp/ # Exposed as an MCP server └── apps/desktop/ # Tauri + React desktop app ``` Let's walk through it along the lifecycle of a request. ## Harness: "coexistence" is a two-stage routing decision The new version takes over exactly one thing — model calls. In code, that sentence is a precise two-stage decision: **Stage one, at the MITM proxy.** A local CA (a 3072-bit RSA root that never leaves your machine) lets the proxy inspect Cursor's HTTPS. Only `*.cursor.sh` traffic is ever TLS-intercepted (`is_cursor_host`); everything else is tunneled untouched. Within Cursor traffic, only a fixed path allowlist — agent runs (`AgentService/RunSSE`, `BidiService/BidiAppend`), model catalogs, account display endpoints — is rewritten to the local service. Sign-in, the plugin marketplace, codebase indexing, and the rest go straight to official servers. Tab completion follows the user's setting: "Direct" passes through; the public or self-hosted service routes out. **Stage two, at the local service.** Even for locally routed requests, the server decodes the first `AgentClientMessage` and looks up its `model_id` in the local model store. Found: handle locally (BYOK). Not found — the user picked an official model — buffer and forward the request back to `api2.cursor.sh`. A Cursor run is a pair of requests (upstream BidiAppend + a long-lived RunSSE stream), and both must take the same path, so a `wait_route` synchronization point blocks the SSE stream until the run is classified `Local` or `Upstream`. This is exactly how the fake account and the start/stop switch disappeared: Auto and official models keep flowing to official billing; only your configured models enter the local path — one login, one client, routed per request. ## The protocol layer: Cursor's protocol behind one door `server/src/cursor/` translates Cursor's private protocol into neutral events the engine understands. When the client updates, this is the only layer that moves. A few designs worth unpacking: **One actor per request.** Each Cursor run gets a `CursorActor` holding a command mailbox, a cancellation token, and a replayable `OutputHub` — late SSE subscribers get all buffered frames replayed, so reconnects lose nothing. Upstream messages pass through an `OrderedInbox` keyed by sequence number: out-of-order arrivals are buffered, duplicates dropped. The protocol layer tolerates network nondeterminism by construction. **Projection is the heart of the layer.** Internally there is one canonical message type (`CanonicalMessage`); the checkpoints sent to Cursor and the requests sent to model providers are both **projections** of it. The Cursor-side checkpoint has a hard check: stable history may only grow. If re-projection finds that a previously generated root "changed" or "shrank", it errors out — failing loudly beats silently corrupting state. **Prompts are compiled.** Seven modes — Agent, Ask, Plan, Debug, Multitask, Subagent, Compaction — each built from a prompt template, a runtime template, and a tool manifest. The compiler substitutes placeholders, trims tools by model capability (no image generation → the tool is removed), and appends dynamic MCP tools in **deterministic order**. Determinism is not a style preference; the caching section below shows it is a correctness requirement. **Tools execute on the client.** Shell commands, file edits, and searches requested by the model are not executed by the server — they are encoded as instructions for the Cursor client to run in the user's environment, with results returned via BidiAppend. Concurrent edits to the same path are serialized by a scheduler, and streaming arguments are projected into live UI previews. ## The loop engine: a minimal agent kernel `server/src/run/` is the execution kernel, fully independent of the Cursor protocol. A run is a loop: ```text loop { project history → call model (consume_model_cycle) ├── no tool calls → append final reply, finish └── tool calls → execute tool round → commit results → continue } ``` **The model cycle is a strict state machine.** `consume_model_cycle` consumes the normalized event stream and enforces protocol invariants: `Start` must come first, text/thinking/tool blocks must open and close in pairs, and the finish reason must agree with the presence of tool calls. On cancellation it cleanly closes any open blocks before exiting — otherwise Cursor would merge deltas from two different outputs. **Commit barriers.** Every state commit (initial messages, settled tool results, compaction, the final reply) waits for the Cursor checkpoint worker to acknowledge the write before the engine proceeds. Checkpoint publication is a hard synchronization point — the engine never runs ahead of published state. **Auto-compaction is an explicit prefix reset.** The engine calibrates token estimates with the latest real usage; when estimated input exceeds `context window − 10,000` reserved tokens, it compacts: retain **exactly one** latest request-context message, summarize the rest (capped at 4,096 output tokens, reasoning disabled), then rebuild the revision in deterministic order. If compaction is interrupted, a fallback summary of the most recent text keeps the loop going. **Cancellation and recovery.** Cancellation tokens propagate from the session all the way into the provider stream; activating a new run in the same conversation cancels the previous one. After a restart, an in-flight tool round can be reconstructed from the Cursor checkpoint and resumed (`RunAction::Resume`) — which is only possible because projection is bidirectional. ## The invariant: prefix cache stability This is where "structure determines correctness" shows most clearly. Providers cache prompts by prefix, and the price difference can be 10×. To hit the cache, history must be an **append-only log**: everything sent to the model last turn must be a byte-level prefix of the next turn. Convention cannot defend that, so it is built into every layer: - **Storage**: history is an immutable revision DAG — each revision is a complete ordered message list identified by a SHA-256 content digest; appends create child revisions, identical appends reuse existing nodes, and the only "rewrite" is compaction, which branches from the root. Runtime events have unique identities; retries are idempotent, and reusing an identity with different content is a hard error. - **Projection**: the system prompt and stable tool prefix are byte-stable when inputs are unchanged. Request context (rules, skills, MCP metadata) is separated from the per-turn runtime message: unchanged content is never re-appended; changed content appends a new message right before the runtime message — old ones are never rewritten. Context going A → B → A appends a third distinct A. - **Provider**: tool-call IDs are normalized before leaving the server; thinking-block signatures (Anthropic's `signature`, OpenAI Responses' encrypted reasoning items) are preserved as replay state and replayed verbatim next turn. GPT-family cache hit rates staying consistently high is not luck — it is this entire chain. The repo even carries a dedicated engineering-constraint document (`.agents/skills/cursor-prefix-stability`); every change touching this chain goes through its checklist. ## The provider layer: three protocols, one event stream Three adapters — Anthropic Messages, OpenAI Chat Completions, OpenAI Responses — emit one normalized event vocabulary (`Start`, open/delta/close for text, thinking and tool blocks, `Usage`, `Done`). The engine and protocol layer never know which upstream they are talking to — which is exactly why "pick OpenAI Chat for everything else" works architecturally. Every call is observable: sanitized request headers and body, streamed response chunks (buffered with limits), time to first token, token usage — fully replayable in detailed mode. The desktop app's call-details page reads precisely this data. ## Ecosystem: the rebuild is more than code **The official ecosystem is preserved as-is.** The biggest ecosystem consequence of coexistence is "no subtraction": plugins, Skills, MCP, codebase indexing, and Tab completion all keep working once you sign in with your own account. The old version could not do this because it replaced the entire backend. **Semantic search is standalone and measurable.** semble implements code indexing and hybrid retrieval as an independent crate, exposed as an MCP server. It ships a reproducible benchmark (real React/Vue commits, queries hand-labeled to implementation lines): natural-language queries reach 95% Recall@5, literal and symbol queries reach 100% — with a public quality-gate file, so any retrieval regression fails CI. **The Tab service is open at three levels.** Public service (hosted by the author), Direct (your own account's official service), or self-hosted (`cursor-tab-server` lives on the archive branch) — switchable in system settings. **The engineering system is part of the ecosystem too.** The repo maintains domain-constraint documents (database schema, prefix-cache stability, release process, i18n…) shared by humans and AI working on the code; `make check` runs formatting, linting, the full test suite, and frontend type checks in one command. This documentation site — including the interactive demo on the homepage built from real desktop components and mock data — lives in the same repo and evolves with the product. **Direction is driven by community discussion.** Tab support, Gemini, image generation, all task modes, and codebase indexing from the roadmap have all landed in the rebuilt version, with more IDEs planned. ## Was it worth it The rebuild consumed weeks that could have gone into features. In exchange: the product model no longer limits itself; protocol updates have a fixed landing place (only `cursor/` moves); correctness is guarded by types, invariants, and regression tests instead of memory; and every new capability grows on a product that already works. The problem used to be "everything is clear except the code" — now the code is clear too. Thoughts on the rebuild? Join the [Roadmap discussion](https://github.com/leookun/cursor-byok/discussions/32).