154 Commits
Author SHA1 Message Date
leookun d004139526 feat: implement context usage anchor for improved token estimation
- Introduced `ContextUsageAnchor` struct to track context input tokens and message count for conversations.
- Updated token estimation functions to utilize the context usage anchor, enhancing accuracy in estimating tokens for projected messages.
- Refactored compaction logic to incorporate context usage anchor, allowing for more efficient management of token budgets during model runs.
- Added tests to validate the behavior of the context usage anchor across different scenarios, including model switching and message additions.
2026-09-01 16:07:51 +08:00
leokun 8c6c415a84 feat: enhance token usage merging and total token calculation
- Updated the `merge_usage` function to include `total_tokens` in the usage merging process.
- Implemented logic to calculate `total_tokens` based on the sum of `context_input_tokens` and `output_tokens`.
- Added a new test to verify that streamed usage correctly includes cached input in the total token count.
2026-09-01 11:13:00 +08:00
leokun d83e14af9a refactor: remove retry_count from ProviderConfig and enhance error handling in tool execution
- Removed the `retry_count` field from `ProviderConfig` as it is no longer needed.
- Introduced `argument_error` field in `ToolCall` to capture errors related to tool arguments.
- Updated various components to handle argument errors more gracefully, including in the `ToolDispatcher` and `ConversationOutput`.
- Enhanced tests to validate the new error handling and ensure proper functionality of tool calls.
2026-09-01 10:14:53 +08:00
leokun 29fde7d7c7 feat: enhance context token estimation and compaction logic
- Added `estimate_context_tokens` function to calculate provider-visible context size based on prompt specifications and projected messages.
- Updated `CheckpointBuilder` to record estimated context tokens during message processing.
- Refactored compaction logic to utilize the new token estimation, ensuring proper context management during model runs.
- Introduced tests to validate context estimation and compaction behavior under various scenarios.
2026-09-01 10:10:09 +08:00
kevin9327andClaude Opus 4.8 1269110615 fix(tools): report the ReadLints paths that were never checked
ReadLints accepts an array of paths, but codec::request encodes only
paths[0] into the DiagnosticsArgs exec, and the exec protocol cannot
carry more than one path. Nothing downstream mentions the drop: for
{"paths": ["a.ts", "b.ts", "c.ts"]} the model is handed
"No diagnostics found in a.ts" with is_error false, so it concludes
b.ts and c.ts are clean when neither was ever opened. The existing
truncation notice cannot catch this, since it compares total_diagnostics
against a per-file diagnostics count.

Name the unread paths in the model-facing result, using the bracketed
notice convention already in this file and the ToolCall that output()
is already given (as task() and render::read() already use it).
Single-path calls are unchanged.

This does not close the capability gap. A real fix fans out one
DiagnosticsArgs exec per path and aggregates the results into the
repeated FileDiagnostics that ReadLintsToolSuccess already defines,
which needs multi-exec reservation and a completion barrier because
take_exec drops the pending entry on the first result. That is a
separate change; this one only stops the silent misinformation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-31 20:13:22 +09:00
kevin9327andClaude Opus 4.8 38a4c3f516 chore: restore a green make check on main
`make check` currently fails on main before any change is made: one
`cargo fmt --all -- --check` diff and four `cargo clippy --workspace
--all-targets -- -D warnings` errors. All five are pre-existing and
none of them change behaviour.

- `server/tests/knowledge_rules.rs:125` — rustfmt wants the long
  `assert!` split across lines. Applied `cargo fmt --all` verbatim.
- `server/src/plugin/data.rs:206,215` — `path` is only read under
  `#[cfg(unix)]`, so every other target sees an unused binding. Added
  a `#[cfg(not(unix))] { let _ = path; }` arm, matching the
  `let _ = error;` idiom already used at line 189 of the same file.
  Windows behaviour is unchanged: these helpers stay no-ops there.
- `server/src/provider/openai_responses.rs:158` — `collapsible_match`.
  Applied clippy's own suggestion (move `thinking_open` into a match
  guard). The match ends in `_ => {}`, so a failed guard falls through
  to a no-op exactly as the inner `if` did.
- `server/src/store/models.rs:327` — `items_after_test_module`. Moved
  `optional_u64` and `to_i64` above `mod tests`; the bodies are
  untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-31 19:59:28 +09:00
kevin9327 ab4d3adee5 fix(model): stop token accumulation from erasing counts a later call omits
`Usage::add_assign` folds every field through

    fn sum(left: Option<u64>, right: Option<u64>) -> Option<u64> {
        left?.checked_add(right?)
    }

so `sum(Some(900), None)` is `None`, not `Some(900)`. Adding a record that does
not report a field therefore does not leave the running total alone - it wipes
it.

Both accumulators are per-turn and run over every provider call in the turn:
`run/engine.rs::accumulate_usage` (the compaction call plus each model cycle)
and `cursor/conversation/output.rs::turn_usage`, which is what `turn_ended`
reports to Cursor.

Providers report these fields inconsistently *across calls of one turn*, which
is all it takes. `openai_usage` reads `cache_read_tokens` from
`prompt_tokens_details`, an object most OpenAI-compatible gateways omit until
the prompt cache warms: call 1 (cold) yields `None`, call 2 yields
`Some(1024)`, and the turn reports `None`. The same happens to
`reasoning_tokens` when only some cycles reason, and to `total_tokens` on
gateways that omit it from the streaming usage chunk. Any tool-using turn makes
more than one call, so this is the normal case rather than an edge case.

Treat an unreported count as zero and keep the field unknown only when neither
side reported it. `checked_add` also becomes `saturating_add`: with the new
rule, silently turning an overflow into `None` would be the same erasure by
another route, and token counts never approach `u64::MAX` anyway.
`context_input_tokens` is untouched.

Before (with `left?.checked_add(right?)` restored):
  cargo test -p cursor-server --lib model::observability
  -> 2 failed: assertion `left == right` failed: left: None, right: Some(900)
After:
  cargo test -p cursor-server --lib model::observability -> 3 passed
2026-08-31 19:39:02 +09:00
kevin9327 24c41d0347 fix(tools): keep result truncation inside its byte budget and terminating
`truncate_text` looks for a fixed point where the kept prefix length equals the
byte count printed in its own truncation notice. Two things go wrong when the
limit is small, and both are reachable because two call sites pass a *remaining*
budget rather than a constant:

* `gate_grep_content` -> `truncate_text("Grep", .., budget.content_bytes)`
* `gate_mcp` -> `truncate_text("MCP text", .., remaining_text)`

1. The loop can spin forever. `notice.len()` grows with the decimal digit count
   of `shown`, so `kept.len()` alternates between two values across a power-of-
   ten boundary and never equals `shown`. With `tool_name = "Grep"` this happens
   at `limit` 78 and 170; with `"MCP text"` at 82 and 174. The conversation task
   then spins at 100% CPU and the turn never completes. `truncate_middle` in
   `model/tool_result_replay.rs` already guards against exactly this.

2. When the notice itself does not fit, `available` saturates to 0 and the
   function returns the ~68 byte notice alone, i.e. *more* than `limit`.
   `gate_grep_content` then evaluates `budget.content_bytes -= <68 bytes>` on a
   budget of at most 68, which panics with "attempt to subtract with overflow"
   in debug/test builds and wraps in release, silently disabling the 32 KiB
   content cap for the rest of the result.

Both are ordinary Grep results away: 16 matches of ~2 KiB leave a two-digit
remainder of the 32 KiB budget, and the next match then hits the small-limit
path.

Fix: return a plain prefix when the notice cannot fit, and stop as soon as the
kept length repeats a previous value. The reported byte count is then off by one
at most in that rare oscillating case, and the result is guaranteed to be at
most `limit` bytes. Every constant-limit call site is unaffected: their notices
always fit and their limits do not oscillate. `budget.content_bytes` now uses
`saturating_sub`, matching every other subtraction in this file.

Verified against the unfixed function:
  cargo test -p cursor-server --lib gate::tests::truncate_text_never_exceeds_its_limit
  -> FAILED: limit 1 produced 67 bytes
  cargo test -p cursor-server --lib gate::tests::grep_content_gate_survives
  -> FAILED: panicked at gate.rs:261: attempt to subtract with overflow
  cargo test -p cursor-server --lib gate::tests::truncate_text_terminates
  -> never returns (killed after 45s)
  cargo test -p cursor-server --lib gate::tests::grep_content_gate_terminates
  -> never returns (killed after 60s)
After: cargo test -p cursor-server --lib tools::tool_call_result::gate -> 4 passed
2026-08-31 19:35:39 +09:00
kevin9327 b6fc7326e7 fix(provider): keep OpenAI Chat tool calls that finish with reason "stop"
`map_finish` matched `"stop" | "content_filter"` before the `has_tools`
fallback, so the fallback only ever applied to finish reasons the adapter did
not recognise. When an OpenAI-compatible server streams tool calls and then
reports `finish_reason: "stop"` the adapter yielded `Done(Stop)` even though
`ToolCallStart` / `ToolCallEnd` had already been emitted.

`run/model_cycle.rs` rejects that combination:

    let has_tool_calls = !calls.is_empty();
    if matches!(finish_reason, FinishReason::ToolUse) != has_tool_calls {
        return Err(failure(RunFailure::Protocol(
            "finish reason and tool calls disagree".into()), ...));
    }

so the whole turn fails with a protocol error and the tool never runs. Servers
that report `"stop"` alongside `tool_calls` are common in BYOK setups
(llama.cpp, Ollama's OpenAI shim, several proxies), which makes those models
unusable for anything agentic.

Move the `has_tools` arm ahead of `"stop" | "content_filter"` so an observed
tool call outranks the label the provider attached to the stop. This matches
the two sibling adapters (`anthropic.rs` checks `_ if saw_tool` before every
reason except `Length`, `openai_responses.rs` derives the reason from
`saw_tool` alone) and this adapter's own `[DONE]` fallback, which already
infers `ToolUse` from `!tools.is_empty()`. `"length"` still wins so a
truncated response is still reported as truncated, and behaviour with no tool
calls is byte-for-byte unchanged.

Before (with the arm restored):
  cargo test -p cursor-server --lib provider::openai_chat
  -> observed_tool_calls_outrank_a_stop_finish_reason FAILED
     assertion `left == right` failed: left: Stop, right: ToolUse
After:
  cargo test -p cursor-server --lib provider::openai_chat -> 3 passed
2026-08-31 19:26:38 +09:00
kevin9327 e87abace8b fix(tests): acknowledge conversation Blob writes in the local rules test
`local_markdown_rules_land_in_the_request_context_message` never finishes:
`cargo test --workspace` fails on `main` with

    panicked at server\tests\local_rules_context.rs:64:14:
    run finishes within timeout: Elapsed(())

The Run publishes the conversation checkpoint by asking the client to write
Blobs, and it does not continue until every `KvServerMessage` is answered
with a `SetBlobResult`. The test drained the output stream without replying,
so the Run stalled after the first frame, the provider was never invoked, and
none of the assertions the test exists for were ever reached.

Answer the Blob writes the way every other transport test already does
(`error_lifecycle.rs`, `conversation_delivery.rs`, `interrupt.rs`). With the
acknowledgement in place the Run reaches `EndStream` in ~0.3s and the original
assertions run and pass, so `merge_local_rules` is now genuinely covered:
exactly one `request-context:` message is projected and it carries
`<user_rule>Always answer in haiku.</user_rule>`.

No production code changes.

Before: `cargo test -p cursor-server --test local_rules_context`
        -> FAILED (0 passed; 1 failed) after a 5s timeout
After:  `cargo test -p cursor-server --test local_rules_context`
        -> ok (1 passed; 0 failed) in 0.28s
2026-08-31 19:20:30 +09:00
leokun ee2592c469 Merge branch 'main' of github.com:leookun/cursor-byok 2026-08-31 16:15:58 +08:00
leokun 49c1fb6378 feat: add Task tool functionality
- Introduced a new `Task` presentation type in the `ToolCallStream` to handle task-related projections.
- Implemented the `TaskProjection` struct with fields for description, prompt, subagent type, model, resume, and environment.
- Added a `project` method to `TaskProjection` to process task-related events and generate interaction updates.
- Created a `task_partial` function to format task updates for the agent server message.
- Included unit tests to verify the correct behavior of task description projections.
2026-08-31 16:15:42 +08:00
leokun 4c3fe230ce Merge remote-tracking branch 'origin/main' into pr-385-merge
# Conflicts:
#	server/tests/interrupt.rs
2026-08-31 16:12:11 +08:00
leokun ac14245d19 Merge pull request #383 from kevin9327/fix/editnotebook-empty-old-string
fix(tools): reject empty old_string in EditNotebook
2026-08-31 16:06:42 +08:00
leokun 84addec26a Merge pull request #384 from kevin9327/fix/bash-shell-alias
fix(tools): complete the bash shell alias in the tool codec
2026-08-31 16:06:18 +08:00
leokun 5de547041c Merge branch 'main' of github.com:leookun/cursor-byok 2026-08-31 15:38:11 +08:00
leokun 8bd0d70add fix: plugin effort compress 2026-08-31 15:38:02 +08:00
leokun 45e694fd63 Merge pull request #386 from kevin9327/fix/empty-tool-arguments
fix: handle tool calls with empty arguments
2026-08-31 13:57:45 +08:00
leookun e7a1cca4c6 feat: add group name functionality to models
- Introduced a new `group_name` field in the model configuration to allow for custom provider-group display names.
- Updated the `CursorModelCards`, `CursorModelEditor`, and `CursorSettingsPage` components to support group settings.
- Enhanced the UI to include group settings options, allowing users to modify group names and associated configurations.
- Added localization strings for new group settings features in both English and Chinese.
- Implemented a database migration to add the `group_name` column to the model configurations.
2026-08-30 23:28:05 +08:00
leookun 21048fb34b feat: add Grok authentication plugin with OAuth support
- Introduced a new Grok authentication plugin, including essential files such as main.ts, provider.ts, and resources.ts.
- Implemented OAuth2 device authorization flow in oauth.ts, allowing users to sign in with xAI.
- Added model discovery and quota management functionalities in models.ts and resources.ts.
- Created a JSON configuration file (plugin.json) for plugin metadata and permissions.
- Included SVG asset for the Grok icon.
- Developed comprehensive tests in grok_test.ts to ensure functionality and reliability of the plugin.
2026-08-30 22:19:21 +08:00
leookun 3184615719 refactor: improve file writing and error handling in PluginDataStore
- Replaced asynchronous file operations with a synchronous approach to ensure file handles are properly closed before replacement.
- Introduced a new `write_once` function for atomic file writing, including directory creation, temporary file writing, and target file replacement.
- Enhanced error handling to retry on transient errors specific to Windows, improving robustness during file operations.
- Updated the `clear` method in `PluginRegistry` to ensure OAuth sessions are only removed after successful persistence of resources.
2026-08-30 21:22:27 +08:00
leookun 5e405b31b2 feat: add copy button functionality to OAuthMethodCard
- Implemented a new button to copy the user code in the OAuthMethodCard, enhancing user experience.
- Added styles for the copy button to match the UI design.
- Updated localization files to include new strings for the copy action in both English and Chinese.
2026-08-30 21:15:03 +08:00
leookun 3a2d47954e refactor: improve error handling and logging in plugin and account services
- Added detailed error messages for plugin data read/write failures, including file paths for better debugging.
- Updated logging levels for upstream request rejections in account services to debug for less critical issues.
- Enhanced error handling in the plugin worker to provide clearer context when starting the plugin worker fails.
- Introduced new functions for merging extra parameters and applying body allowlists in provider services, improving request validation.
2026-08-30 21:10:08 +08:00
leookun 453c120740 fix: plugin in windows 2026-08-30 20:57:02 +08:00
leookun a5bbe67845 refactor: replace Button with TruncatedButton in CursorModelCards and PluginManagementPage
- Updated the UI components in CursorModelCards and PluginManagementPage to use TruncatedButton for better text handling and display.
- Adjusted styles in CursorSettings and PluginManagementPage to ensure proper button layout and responsiveness.
- Added new ActionMenu component for handling additional actions in PluginManagementPage.
- Enhanced localization files to include new strings for the ActionMenu and TruncatedButton components.
2026-08-30 20:45:27 +08:00
leookun 05181f9e8a Merge branch 'main' into feat/plugin 2026-08-30 20:01:41 +08:00
leookun 4373507571 feat: add provider stream idle timeout and error handling
- Introduced a new `provider_stream_idle_timeout` configuration to manage idle timeouts for provider streams.
- Enhanced error handling in the `ProviderRouter` to include specific timeout errors for both request and stream idle scenarios.
- Updated the `AnthropicProvider` and `OpenAiChatProvider` to utilize the new error handling functions for improved SSE error reporting.
- Added tests to verify the correct behavior of timeout handling and error extraction from provider events.
2026-08-30 20:01:08 +08:00
leookun e6130e01a7 feat: plugin system 2026-08-30 20:00:53 +08:00
kevin9327andClaude Opus 4.8 e673a034df fix: handle tool calls with empty arguments
A tool call that carries no arguments streams no argument text, so
`arguments_text` is empty and `from_str("")` fails with `EOF while parsing
a value`, aborting the whole run. The model cycle already guards this, but
two other consumers did not:

- `ConversationOutput` re-parses the streamed text on `ToolCallEnd`; and
- `create_tool_round` stored the empty text verbatim in the
  `arguments_json` column, so re-loading the round (`commit_tool_result`
  and the round loader) then failed on `from_str("")`.

Treat empty argument text as an empty object in the output projection, and
persist `{}` for it so the `arguments_json` column always holds valid JSON.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:38:39 +09:00
kevin9327andClaude Opus 4.8 6ac666da4f fix(conversation): clear pending runtime user-message injections
Runtime user messages and context injections both queue into
`pending_injections`, but with different keys: injections use the raw
injection id (committed under `inject-context:{id}`) while user messages
use the full `user-message:{id}` event id. The commit-correlation handler
only stripped the `inject-context:` prefix, so a user message's entry was
never removed.

Consequences:
- the client never received `ContextInjectionDelivered` /
  `UserMessageAppended` for the message; and
- `pending_injections` stayed non-empty, so every later `ExecuteToolRound`
  was detached without dispatching its tools and `tool_round::execute`
  blocked forever -- a hung turn whenever the model made a tool call after
  the interruption.

Derive the lookup key by stripping the injection prefix when present and
otherwise using the event id verbatim, so both kinds are cleared and their
delivered/appended events fire.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:23:57 +09:00
kevin9327andClaude Opus 4.8 9e2d22418b fix(tools): complete the bash shell alias in the tool codec
The tool dispatcher already treats `bash`/`Bash` as an alias of `Shell`
(routing, `is_shell_tool`, and `block_until_ms` normalization), but the
codec only matched `shell`:

- `tool_placeholder` returned `unsupported tool: bash`, which aborts the
  turn while streaming the tool call, before it ever runs;
- `request` returned `tool bash is not executed through ExecServerMessage`
  (after already reserving an exec slot); and
- `stream_closed` built its shell-specific error result only for `Shell`.

Anthropic models frequently emit `Bash` even when the tool is advertised
as `Shell`, so the alias must hold across the codec. Match `bash` wherever
the codec special-cases `shell`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:15:29 +09:00
kevin9327andClaude Opus 4.8 97ee138de8 fix(tools): reject empty old_string in EditNotebook
StrReplace rejects an empty `old_string`, but the EditNotebook cell-edit
path did not. Because `str::match_indices("")` matches at every byte
boundary, editing a non-empty cell with an empty `old_string` failed with
a misleading "old_string is not unique in the notebook cell; found N
occurrences" error, and editing an empty cell silently prepended
`new_string`.

Add the same guard StrReplace already uses so both edit tools reject an
empty `old_string` consistently.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 19:07:28 +09:00
leookun 1609b57433 plugin runtime 2026-08-30 13:35:03 +08:00
leookun 2ff74bc6a8 feat: webfetch 2026-08-30 11:54:32 +08:00
leookun 66ac94d5b1 feat: implement web cache for persisting and serving fetched content
- Introduced a new `WebCache` module to manage web content caching.
- Added functionality to store fetched content and serve it via a dedicated route.
- Integrated web cache into the search module for improved content retrieval.
- Implemented database migration management with detailed diagnostics for better error handling during startup.
- Updated the SQLite store to utilize the new migration system for enhanced database management.
2026-08-30 04:29:39 +08:00
leookun fd7998e43c feat: enhance desktop application startup diagnostics and logging
- Introduced a new `startup` module to capture and manage startup diagnostics.
- Implemented logging functionality with daily rotation and error reporting for application startup failures.
- Updated the main application run function to return an exit code based on startup success or failure.
- Added new dependencies for logging and diagnostics in `Cargo.toml`.
- Increased request timeout in the control service for improved stability.
- Expanded maximum search results in the federation module for better search capabilities.
2026-08-30 03:19:01 +08:00
leookun 44e2d8057a feat: initialize server structure and database schema
- Added initial server setup with Cargo.toml defining dependencies and project structure.
- Created build.rs for generating protobuf bindings and validating wire contracts.
- Established database schema with initial migration files for conversations, messages, and runs.
- Introduced tools and prompts for Cursor functionality, enhancing user interaction capabilities.
2026-08-30 01:33:39 +08:00
leookun d200b3791d new plan 2026-08-29 22:05:04 +08:00
leookun 279e6bb07c refactor: align project directory architecture 2026-08-29 20:45:51 +08:00
leokun c983124288 Merge pull request #374 from zwenooo/codex/fix-cursor-reasoning-parameter
fix: align Cursor reasoning parameter metadata
2026-08-28 23:22:06 +08:00
leokun 6447fd998a Merge pull request #367 from jiah0231/fix/runtime-unsupported-tool-recovery
修复:运行中遇到已移除/未知工具时避免 Agent 中断Fix/runtime unsupported tool recovery
2026-08-28 23:20:32 +08:00
leookun 5cc2401ce2 feat: add first valid response timing and related metrics
- Introduced `first_valid_response_ms` and `ttfr_ms` to track the timing of the first valid response in LLM calls.
- Updated relevant interfaces and components to display and utilize the new metrics, including CallDetails, CallTable, and LatencyChart.
- Enhanced the database schema to accommodate the new timing fields.
- Implemented logic in the service layer to record the first valid response during model interactions.
2026-08-28 22:56:39 +08:00
leokun 5a0bc2e0e9 fix: persist each provider retry as its own llm_calls row
Retry attempts now finish the failed call, start a new call id, and record the next request so call history shows intermediate failures.
2026-08-28 21:09:45 +08:00
leokun 6a3570a95a refactor: implement retry mechanism for API calls in Anthropic, OpenAI Chat, and OpenAI Responses providers
- Replaced direct API call handling with a retry mechanism using `send_with_retry`.
- Improved error handling and response management for better reliability in network requests.
2026-08-28 20:37:52 +08:00
leokun 921bcfd247 feat: implement automatic call refresh on CallsPage
- Added a refreshCalls function to appStore to fetch updated call data.
- Integrated useEffect in CallsPage to periodically refresh calls every 2 seconds and on visibility change.
2026-08-28 19:46:37 +08:00
leokun e27b989002 Fix compaction projection and MCP image truncation 2026-08-28 18:54:27 +08:00
zwen befb18d9c3 fix: align Cursor reasoning parameter metadata 2026-08-28 15:24:12 +08:00
leookun 59379f1f72 fix: repair stream lifecycle reliability 2026-08-28 13:00:39 +08:00
jiah0231 2c32dad271 test: cover AwaitShell emitted during active run 2026-08-28 12:21:45 +08:00
jiah0231 85241a0413 docs: describe AwaitShell as runtime unsupported tool 2026-08-28 12:21:34 +08:00