53 Commits
Author SHA1 Message Date
leokun 9a8fde279c feat(byok): add external API for local apps 2026-09-28 19:06:50 +08:00
The Gru fdae9c41c7 fix(run): make automatic compaction recover an over-limit conversation (#426)
Automatic compaction cannot do its job once a conversation crosses the
context window, so the conversation stays there permanently. Observed
against a 1M-token Anthropic window:

1. The summarize call replays the full history. It runs precisely
   because that history is too large, so the request is itself over the
   limit ("prompt is too long"), or it ends with an assistant/tool
   message that Anthropic refuses as a prefill. Either way the run falls
   back to the 12K truncated JSON summary, which discards the context.
   In one trace the summarizer received 771 messages (2.78 MB) and
   returned a single token.

2. The compaction check uses a 10K fixed reserve. The estimate trails
   the provider's own count by the request context and provider-side
   overhead that the message-tail estimate does not model; a 948K
   estimate passed the check and Anthropic counted 1,017,628.

3. When the provider does refuse the prompt, the run retries the same
   prompt eight times at 5s intervals and then fails. Nothing compacts.

Fixes, all in server/src/run:

- compaction_history trims the summarizer input to the context budget
  at user-turn boundaries (never splitting a tool call from its
  results) and guarantees it ends with a user message.
- context_budget keeps 10% of the window free instead of a fixed 10K,
  so the reserve scales with the model and absorbs the drift.
- A provider refusal matching is_context_overflow compacts once and
  retries the turn instead of failing it.
- 4xx responses other than 408/425/429 are terminal. A rejected request
  fails identically every time, so retrying only delays the error.

Separately, Cursor can resume a finished turn whose checkpoint already
ends with the assistant, which Anthropic also rejects as a prefill.
run/history.rs appends a transient user tail to every provider request
that would otherwise end with the assistant. The tail is never
persisted, so committed checkpoints stay an exact prefix of the next
turn and the usage anchor still counts persisted messages only.
2026-09-06 20:33:03 +08:00
leokun 02db171f2c Merge pull request #390 from kevin9327/fix/local-rules-test-blob-ack 2026-09-04 16:06:56 +08:00
leokun cd09c60657 Merge remote-tracking branch 'origin/main' 2026-09-04 15:46:50 +08:00
leokun 52c52abc66 Merge pull request #397 from kevin9327/fix/duplicate-tool-call-id-hang
fix(conversation): scope completed tool call ids to their round
2026-09-04 15:46:17 +08:00
leokun 17342167ac feat: update desktop settings and server compatibility 2026-09-04 15:33:55 +08:00
leokun 924b5e5926 feat(cursor): cli 接入本地模型路由与配置
- feat(服务配置): 本地响应配置并禁用 HTTP/2 传输
- feat(模型目录): 提供默认模型接口及本地路由凭据
- fix(待办状态): 基于检查点合并增量待办并保留未变更项
- fix(模型参数): 忽略未知参数以兼容新版 Cursor 请求
- test(本地路由): 覆盖模型元数据路由与凭据行为
2026-09-03 16:01:25 +08:00
ProtectCookies 42811a27f5 feat: 新增 Commit 设置及本地提交信息生成
- 新增 CommitSettingsCard 组件,支持选择生成模型与编辑提示词
- 新增 /settings/commit GET/PUT 接口及 CommitSettings 持久化
- 实现 WriteGitCommitMessage RPC 本地生成,空 model_id 时直连转发
- 新增 NetworkService/IsConnected 探针响应,防止流式生成被中断
- 添加 commit prompt 模板及 proto 消息定义
- 补充 zh-CN / en-US 国际化词条
2026-09-03 10:31:41 +08:00
kevin9327andClaude Opus 4.8 a42f84cfe7 fix(conversation): scope completed tool call ids to their round
A provider that reuses a tool call id across two rounds of one run
wedges the run permanently. `ToolDispatcher::start_batch` skips any call
whose id is in `ToolBatchState::completed`, so the second call is never
dispatched and never produces a `ToolCompletion`, while
`tool_round::execute` blocks waiting for `calls.len()` results with no
timeout on that path. The client sees the tool call appear and then
nothing: no completion, no further output, no end-stream frame.

The `completed` set is built once per run and never cleared, so it is
run-scoped. A tool call id is only unique within a round, which the
schema already states as `UNIQUE (round_id, call_id)`; the sibling
runtime completed-map is likewise already cleared per round at
output.rs:615.

Clear `completed` when `ExecuteToolRound` begins a new round, tracked
independently of `active_round` so it does not depend on the order in
which ToolRoundStarted and ExecuteToolRound are observed. Replaying a
round still skips the calls that round already committed.

The practical trigger is openai_chat.rs:196, which synthesizes
`call-{index}` from a per-stream index when a provider omits tool call
ids, so `call-0` recurs on every model call. That file is left alone
here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-03 07:35:10 +09:00
leokun 8fdcdd7f84 Merge pull request #395 from kevin9327/chore/restore-make-check-green
chore: restore a green `make check` on main
2026-09-03 00:32:31 +08:00
leookun f22c7b6680 feat: integrate app version into control service and update ads endpoint
- Added `app_version` field to `ControlService` and updated its initialization to include the app version.
- Modified the `ADS_ENDPOINT` to point to a local server for development purposes.
- Refactored conversation command and output handling to utilize a new `RunFinish` enum for better state management.
- Enhanced the conversation runtime to handle queued user messages after a turn has ended, ensuring smooth transitions between turns.
- Added tests to validate the new behavior of queued messages and transport handling.
2026-09-02 10:23:50 +08:00
leookun 5cdf642dd1 feat: enhance bidi request handling and observability tracing
- Updated the `append` function to include a flag for replacing closing requests, improving request handling.
- Refactored the `run_sse_handler` and `bidi_handler` functions to utilize a new tracing mechanism, enhancing observability.
- Introduced a new `trace_outcome` function to standardize tracing outcomes for requests.
- Removed the `CursorTraceRecorder` in favor of a new `CursorTraceService` for better performance and non-blocking behavior.
- Added tests to validate the new tracing functionality and ensure correct behavior during request processing.
2026-09-02 01:04:54 +08:00
leokun 2dad593263 feat: add resource limits management and network client integration
- Introduced a new module for managing process resource limits, specifically for raising the open file limit on Unix systems.
- Added a `NetworkClients` struct to handle reusable outbound HTTP clients, improving network request management.
- Updated various components, including `ControlService` and `CursorProxy`, to utilize the new network client structure for better client handling.
- Enhanced the API router to accept network clients, ensuring consistent client usage across different services.
- Added tests to validate the integration of network clients and resource limits functionality.
2026-09-01 21:03:02 +08:00
leokun 76417e005b feat: add usage snapshot event and enhance compaction logic
- Introduced `UsageSnapshot` event to track token usage during conversation runs.
- Updated `RunEngine` to emit usage snapshots, providing better visibility into token consumption.
- Refactored compaction logic to utilize a new `compaction_estimate` function for improved token budget management.
- Added tests to validate timeout constants for blob synchronization and ensure correct behavior of usage tracking during compaction.
2026-09-01 20:04:35 +08:00
leokun 2c63bd845a feat: track interaction events during automatic compaction
- Added tracking for interaction events in the `Output` struct, including `summary_started` and `token_delta`.
- Updated the `run` function to push relevant interaction events to the `interaction_events` vector.
- Enhanced the automatic compaction test to verify the immediate reset of cursor usage and the correct logging of interaction events.
2026-09-01 16:54:45 +08:00
leookun d004139526 feat: implement context usage anchor for improved token estimation
- Introduced `ContextUsageAnchor` struct to track context input tokens and message count for conversations.
- Updated token estimation functions to utilize the context usage anchor, enhancing accuracy in estimating tokens for projected messages.
- Refactored compaction logic to incorporate context usage anchor, allowing for more efficient management of token budgets during model runs.
- Added tests to validate the behavior of the context usage anchor across different scenarios, including model switching and message additions.
2026-09-01 16:07:51 +08:00
leokun d83e14af9a refactor: remove retry_count from ProviderConfig and enhance error handling in tool execution
- Removed the `retry_count` field from `ProviderConfig` as it is no longer needed.
- Introduced `argument_error` field in `ToolCall` to capture errors related to tool arguments.
- Updated various components to handle argument errors more gracefully, including in the `ToolDispatcher` and `ConversationOutput`.
- Enhanced tests to validate the new error handling and ensure proper functionality of tool calls.
2026-09-01 10:14:53 +08:00
leokun 29fde7d7c7 feat: enhance context token estimation and compaction logic
- Added `estimate_context_tokens` function to calculate provider-visible context size based on prompt specifications and projected messages.
- Updated `CheckpointBuilder` to record estimated context tokens during message processing.
- Refactored compaction logic to utilize the new token estimation, ensuring proper context management during model runs.
- Introduced tests to validate context estimation and compaction behavior under various scenarios.
2026-09-01 10:10:09 +08:00
kevin9327andClaude Opus 4.8 38a4c3f516 chore: restore a green make check on main
`make check` currently fails on main before any change is made: one
`cargo fmt --all -- --check` diff and four `cargo clippy --workspace
--all-targets -- -D warnings` errors. All five are pre-existing and
none of them change behaviour.

- `server/tests/knowledge_rules.rs:125` — rustfmt wants the long
  `assert!` split across lines. Applied `cargo fmt --all` verbatim.
- `server/src/plugin/data.rs:206,215` — `path` is only read under
  `#[cfg(unix)]`, so every other target sees an unused binding. Added
  a `#[cfg(not(unix))] { let _ = path; }` arm, matching the
  `let _ = error;` idiom already used at line 189 of the same file.
  Windows behaviour is unchanged: these helpers stay no-ops there.
- `server/src/provider/openai_responses.rs:158` — `collapsible_match`.
  Applied clippy's own suggestion (move `thinking_open` into a match
  guard). The match ends in `_ => {}`, so a failed guard falls through
  to a no-op exactly as the inner `if` did.
- `server/src/store/models.rs:327` — `items_after_test_module`. Moved
  `optional_u64` and `to_i64` above `mod tests`; the bodies are
  untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-31 19:59:28 +09:00
kevin9327 e87abace8b fix(tests): acknowledge conversation Blob writes in the local rules test
`local_markdown_rules_land_in_the_request_context_message` never finishes:
`cargo test --workspace` fails on `main` with

    panicked at server\tests\local_rules_context.rs:64:14:
    run finishes within timeout: Elapsed(())

The Run publishes the conversation checkpoint by asking the client to write
Blobs, and it does not continue until every `KvServerMessage` is answered
with a `SetBlobResult`. The test drained the output stream without replying,
so the Run stalled after the first frame, the provider was never invoked, and
none of the assertions the test exists for were ever reached.

Answer the Blob writes the way every other transport test already does
(`error_lifecycle.rs`, `conversation_delivery.rs`, `interrupt.rs`). With the
acknowledgement in place the Run reaches `EndStream` in ~0.3s and the original
assertions run and pass, so `merge_local_rules` is now genuinely covered:
exactly one `request-context:` message is projected and it carries
`<user_rule>Always answer in haiku.</user_rule>`.

No production code changes.

Before: `cargo test -p cursor-server --test local_rules_context`
        -> FAILED (0 passed; 1 failed) after a 5s timeout
After:  `cargo test -p cursor-server --test local_rules_context`
        -> ok (1 passed; 0 failed) in 0.28s
2026-08-31 19:20:30 +09:00
leokun 4c3fe230ce Merge remote-tracking branch 'origin/main' into pr-385-merge
# Conflicts:
#	server/tests/interrupt.rs
2026-08-31 16:12:11 +08:00
leokun 45e694fd63 Merge pull request #386 from kevin9327/fix/empty-tool-arguments
fix: handle tool calls with empty arguments
2026-08-31 13:57:45 +08:00
leookun e7a1cca4c6 feat: add group name functionality to models
- Introduced a new `group_name` field in the model configuration to allow for custom provider-group display names.
- Updated the `CursorModelCards`, `CursorModelEditor`, and `CursorSettingsPage` components to support group settings.
- Enhanced the UI to include group settings options, allowing users to modify group names and associated configurations.
- Added localization strings for new group settings features in both English and Chinese.
- Implemented a database migration to add the `group_name` column to the model configurations.
2026-08-30 23:28:05 +08:00
kevin9327andClaude Opus 4.8 e673a034df fix: handle tool calls with empty arguments
A tool call that carries no arguments streams no argument text, so
`arguments_text` is empty and `from_str("")` fails with `EOF while parsing
a value`, aborting the whole run. The model cycle already guards this, but
two other consumers did not:

- `ConversationOutput` re-parses the streamed text on `ToolCallEnd`; and
- `create_tool_round` stored the empty text verbatim in the
  `arguments_json` column, so re-loading the round (`commit_tool_result`
  and the round loader) then failed on `from_str("")`.

Treat empty argument text as an empty object in the output projection, and
persist `{}` for it so the `arguments_json` column always holds valid JSON.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:38:39 +09:00
kevin9327andClaude Opus 4.8 6ac666da4f fix(conversation): clear pending runtime user-message injections
Runtime user messages and context injections both queue into
`pending_injections`, but with different keys: injections use the raw
injection id (committed under `inject-context:{id}`) while user messages
use the full `user-message:{id}` event id. The commit-correlation handler
only stripped the `inject-context:` prefix, so a user message's entry was
never removed.

Consequences:
- the client never received `ContextInjectionDelivered` /
  `UserMessageAppended` for the message; and
- `pending_injections` stayed non-empty, so every later `ExecuteToolRound`
  was detached without dispatching its tools and `tool_round::execute`
  blocked forever -- a hung turn whenever the model made a tool call after
  the interruption.

Derive the lookup key by stripping the injection prefix when present and
otherwise using the event id verbatim, so both kinds are cleared and their
delivered/appended events fire.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:23:57 +09:00
leookun 66ac94d5b1 feat: implement web cache for persisting and serving fetched content
- Introduced a new `WebCache` module to manage web content caching.
- Added functionality to store fetched content and serve it via a dedicated route.
- Integrated web cache into the search module for improved content retrieval.
- Implemented database migration management with detailed diagnostics for better error handling during startup.
- Updated the SQLite store to utilize the new migration system for enhanced database management.
2026-08-30 04:29:39 +08:00
leookun 44e2d8057a feat: initialize server structure and database schema
- Added initial server setup with Cargo.toml defining dependencies and project structure.
- Created build.rs for generating protobuf bindings and validating wire contracts.
- Established database schema with initial migration files for conversations, messages, and runs.
- Introduced tools and prompts for Cursor functionality, enhancing user interaction capabilities.
2026-08-30 01:33:39 +08:00
leookun d200b3791d new plan 2026-08-29 22:05:04 +08:00
leookun 279e6bb07c refactor: align project directory architecture 2026-08-29 20:45:51 +08:00
leokun 6447fd998a Merge pull request #367 from jiah0231/fix/runtime-unsupported-tool-recovery
修复:运行中遇到已移除/未知工具时避免 Agent 中断Fix/runtime unsupported tool recovery
2026-08-28 23:20:32 +08:00
leookun 5cc2401ce2 feat: add first valid response timing and related metrics
- Introduced `first_valid_response_ms` and `ttfr_ms` to track the timing of the first valid response in LLM calls.
- Updated relevant interfaces and components to display and utilize the new metrics, including CallDetails, CallTable, and LatencyChart.
- Enhanced the database schema to accommodate the new timing fields.
- Implemented logic in the service layer to record the first valid response during model interactions.
2026-08-28 22:56:39 +08:00
leookun 59379f1f72 fix: repair stream lifecycle reliability 2026-08-28 13:00:39 +08:00
jiah0231 2c32dad271 test: cover AwaitShell emitted during active run 2026-08-28 12:21:45 +08:00
jiah0231 edc28de86b 修复:兼容已移除工具,避免旧会话 Resume 中断
兼容 v0.1.5-beta.1 删除 AwaitShell 后的旧会话 Resume,并将未知/已移除工具降级为模型可见失败结果,避免整个 Agent Run 被 Protocol Error 直接终止。
2026-08-28 12:12:50 +08:00
leookun b15e149b9d fix:add anthropic cache block 2026-08-28 10:33:00 +08:00
leookun 71ddf71ec1 feat(api): enhance statistics storage management with scope options
- Introduced a new type `StatisticsStorageScope` to specify the scope for clearing statistics.
- Updated the `clearStatisticsStorage` API method to accept a scope parameter, allowing for selective clearing of detailed records or all statistics.
- Modified the demo API to handle the new scope parameter appropriately.
- Updated the SettingsPage component to include a selection for clearing scope, enhancing user control over statistics management.
- Added new translations for the updated messages related to statistics clearing in both English and Chinese.
2026-08-28 10:18:11 +08:00
leokun c5d578c5b1 feat(docs): add documentation site and demo features
- Introduced a new documentation site for Cursor BYOK using Next.js and Fumadocs.
- Added a product demo page with a corresponding Vite configuration.
- Implemented a demo API to simulate LLM calls and responses.
- Enhanced the Makefile to include new build and development commands for the documentation.
- Updated package.json scripts for building and running the documentation site.
- Created various components and layouts for the documentation structure, including blog and user documentation sections.
- Added styling for the new components and layouts to ensure a cohesive design.
- Included a README and other necessary files for local development and deployment.
2026-08-27 21:45:28 +08:00
leokun ee915ee760 fix: preempt root loop on context injection 2026-08-27 17:01:03 +08:00
leokun 5450fc76e2 fix: shell 2026-08-26 20:20:38 +08:00
leokun df053c3720 fix: harden concurrent persistence and task recovery 2026-08-26 16:46:41 +08:00
leookun e1937233ec merge old config 2026-08-26 15:46:11 +08:00
leookun 847e92c7ea refactor: update cursor request handling and improve parent request management
- Changed `run_id` to `request_id` in `CursorParent` struct for clarity.
- Enhanced the `prepare` function to handle parent requests asynchronously, ensuring proper error handling for active runs.
- Updated database interactions to include `cursor_request_id` for better tracking of requests.
- Added tests to verify the behavior of reused cursor request IDs and their mapping to distinct executions.
2026-08-26 01:46:30 +08:00
leokun 081e1f50e2 feat: NormalizedProvider 2026-08-25 20:50:57 +08:00
leookun 5f87357681 fix: restore Windows shell parsing metadata 2026-08-25 13:31:14 +08:00
leokun eb26b17ba0 fix: stabilize todo state and responses streams 2026-08-24 18:43:44 +08:00
leokun 24177fcb6e fix: restore service entrypoint and clean checks 2026-08-24 15:46:14 +08:00
leokun 7fa4953883 fix: harden provider, tool, and desktop behavior 2026-08-24 15:37:42 +08:00
leokun 4ddd3adb3f feat: normalize MCP tool names and enhance OpenAI response handling 2026-08-24 12:22:11 +08:00
leookun 1a0cf89fe1 fix: set OpenAI Responses image detail to auto 2026-08-24 12:00:49 +08:00
leookun 81a1afdae9 feat: add model connectivity tests and harden cursor heartbeats 2026-08-24 09:57:28 +08:00