Automatic compaction cannot do its job once a conversation crosses the
context window, so the conversation stays there permanently. Observed
against a 1M-token Anthropic window:
1. The summarize call replays the full history. It runs precisely
because that history is too large, so the request is itself over the
limit ("prompt is too long"), or it ends with an assistant/tool
message that Anthropic refuses as a prefill. Either way the run falls
back to the 12K truncated JSON summary, which discards the context.
In one trace the summarizer received 771 messages (2.78 MB) and
returned a single token.
2. The compaction check uses a 10K fixed reserve. The estimate trails
the provider's own count by the request context and provider-side
overhead that the message-tail estimate does not model; a 948K
estimate passed the check and Anthropic counted 1,017,628.
3. When the provider does refuse the prompt, the run retries the same
prompt eight times at 5s intervals and then fails. Nothing compacts.
Fixes, all in server/src/run:
- compaction_history trims the summarizer input to the context budget
at user-turn boundaries (never splitting a tool call from its
results) and guarantees it ends with a user message.
- context_budget keeps 10% of the window free instead of a fixed 10K,
so the reserve scales with the model and absorbs the drift.
- A provider refusal matching is_context_overflow compacts once and
retries the turn instead of failing it.
- 4xx responses other than 408/425/429 are terminal. A rejected request
fails identically every time, so retrying only delays the error.
Separately, Cursor can resume a finished turn whose checkpoint already
ends with the assistant, which Anthropic also rejects as a prefill.
run/history.rs appends a transient user tail to every provider request
that would otherwise end with the assistant. The tail is never
persisted, so committed checkpoints stay an exact prefix of the next
turn and the usage anchor still counts persisted messages only.
A provider that reuses a tool call id across two rounds of one run
wedges the run permanently. `ToolDispatcher::start_batch` skips any call
whose id is in `ToolBatchState::completed`, so the second call is never
dispatched and never produces a `ToolCompletion`, while
`tool_round::execute` blocks waiting for `calls.len()` results with no
timeout on that path. The client sees the tool call appear and then
nothing: no completion, no further output, no end-stream frame.
The `completed` set is built once per run and never cleared, so it is
run-scoped. A tool call id is only unique within a round, which the
schema already states as `UNIQUE (round_id, call_id)`; the sibling
runtime completed-map is likewise already cleared per round at
output.rs:615.
Clear `completed` when `ExecuteToolRound` begins a new round, tracked
independently of `active_round` so it does not depend on the order in
which ToolRoundStarted and ExecuteToolRound are observed. Replaying a
round still skips the calls that round already committed.
The practical trigger is openai_chat.rs:196, which synthesizes
`call-{index}` from a per-stream index when a provider omits tool call
ids, so `call-0` recurs on every model call. That file is left alone
here.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Added `app_version` field to `ControlService` and updated its initialization to include the app version.
- Modified the `ADS_ENDPOINT` to point to a local server for development purposes.
- Refactored conversation command and output handling to utilize a new `RunFinish` enum for better state management.
- Enhanced the conversation runtime to handle queued user messages after a turn has ended, ensuring smooth transitions between turns.
- Added tests to validate the new behavior of queued messages and transport handling.
- Updated the `append` function to include a flag for replacing closing requests, improving request handling.
- Refactored the `run_sse_handler` and `bidi_handler` functions to utilize a new tracing mechanism, enhancing observability.
- Introduced a new `trace_outcome` function to standardize tracing outcomes for requests.
- Removed the `CursorTraceRecorder` in favor of a new `CursorTraceService` for better performance and non-blocking behavior.
- Added tests to validate the new tracing functionality and ensure correct behavior during request processing.
- Introduced a new module for managing process resource limits, specifically for raising the open file limit on Unix systems.
- Added a `NetworkClients` struct to handle reusable outbound HTTP clients, improving network request management.
- Updated various components, including `ControlService` and `CursorProxy`, to utilize the new network client structure for better client handling.
- Enhanced the API router to accept network clients, ensuring consistent client usage across different services.
- Added tests to validate the integration of network clients and resource limits functionality.
- Introduced `UsageSnapshot` event to track token usage during conversation runs.
- Updated `RunEngine` to emit usage snapshots, providing better visibility into token consumption.
- Refactored compaction logic to utilize a new `compaction_estimate` function for improved token budget management.
- Added tests to validate timeout constants for blob synchronization and ensure correct behavior of usage tracking during compaction.
- Added tracking for interaction events in the `Output` struct, including `summary_started` and `token_delta`.
- Updated the `run` function to push relevant interaction events to the `interaction_events` vector.
- Enhanced the automatic compaction test to verify the immediate reset of cursor usage and the correct logging of interaction events.
- Introduced `ContextUsageAnchor` struct to track context input tokens and message count for conversations.
- Updated token estimation functions to utilize the context usage anchor, enhancing accuracy in estimating tokens for projected messages.
- Refactored compaction logic to incorporate context usage anchor, allowing for more efficient management of token budgets during model runs.
- Added tests to validate the behavior of the context usage anchor across different scenarios, including model switching and message additions.
- Removed the `retry_count` field from `ProviderConfig` as it is no longer needed.
- Introduced `argument_error` field in `ToolCall` to capture errors related to tool arguments.
- Updated various components to handle argument errors more gracefully, including in the `ToolDispatcher` and `ConversationOutput`.
- Enhanced tests to validate the new error handling and ensure proper functionality of tool calls.
- Added `estimate_context_tokens` function to calculate provider-visible context size based on prompt specifications and projected messages.
- Updated `CheckpointBuilder` to record estimated context tokens during message processing.
- Refactored compaction logic to utilize the new token estimation, ensuring proper context management during model runs.
- Introduced tests to validate context estimation and compaction behavior under various scenarios.
`make check` currently fails on main before any change is made: one
`cargo fmt --all -- --check` diff and four `cargo clippy --workspace
--all-targets -- -D warnings` errors. All five are pre-existing and
none of them change behaviour.
- `server/tests/knowledge_rules.rs:125` — rustfmt wants the long
`assert!` split across lines. Applied `cargo fmt --all` verbatim.
- `server/src/plugin/data.rs:206,215` — `path` is only read under
`#[cfg(unix)]`, so every other target sees an unused binding. Added
a `#[cfg(not(unix))] { let _ = path; }` arm, matching the
`let _ = error;` idiom already used at line 189 of the same file.
Windows behaviour is unchanged: these helpers stay no-ops there.
- `server/src/provider/openai_responses.rs:158` — `collapsible_match`.
Applied clippy's own suggestion (move `thinking_open` into a match
guard). The match ends in `_ => {}`, so a failed guard falls through
to a no-op exactly as the inner `if` did.
- `server/src/store/models.rs:327` — `items_after_test_module`. Moved
`optional_u64` and `to_i64` above `mod tests`; the bodies are
untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`local_markdown_rules_land_in_the_request_context_message` never finishes:
`cargo test --workspace` fails on `main` with
panicked at server\tests\local_rules_context.rs:64:14:
run finishes within timeout: Elapsed(())
The Run publishes the conversation checkpoint by asking the client to write
Blobs, and it does not continue until every `KvServerMessage` is answered
with a `SetBlobResult`. The test drained the output stream without replying,
so the Run stalled after the first frame, the provider was never invoked, and
none of the assertions the test exists for were ever reached.
Answer the Blob writes the way every other transport test already does
(`error_lifecycle.rs`, `conversation_delivery.rs`, `interrupt.rs`). With the
acknowledgement in place the Run reaches `EndStream` in ~0.3s and the original
assertions run and pass, so `merge_local_rules` is now genuinely covered:
exactly one `request-context:` message is projected and it carries
`<user_rule>Always answer in haiku.</user_rule>`.
No production code changes.
Before: `cargo test -p cursor-server --test local_rules_context`
-> FAILED (0 passed; 1 failed) after a 5s timeout
After: `cargo test -p cursor-server --test local_rules_context`
-> ok (1 passed; 0 failed) in 0.28s
- Introduced a new `group_name` field in the model configuration to allow for custom provider-group display names.
- Updated the `CursorModelCards`, `CursorModelEditor`, and `CursorSettingsPage` components to support group settings.
- Enhanced the UI to include group settings options, allowing users to modify group names and associated configurations.
- Added localization strings for new group settings features in both English and Chinese.
- Implemented a database migration to add the `group_name` column to the model configurations.
A tool call that carries no arguments streams no argument text, so
`arguments_text` is empty and `from_str("")` fails with `EOF while parsing
a value`, aborting the whole run. The model cycle already guards this, but
two other consumers did not:
- `ConversationOutput` re-parses the streamed text on `ToolCallEnd`; and
- `create_tool_round` stored the empty text verbatim in the
`arguments_json` column, so re-loading the round (`commit_tool_result`
and the round loader) then failed on `from_str("")`.
Treat empty argument text as an empty object in the output projection, and
persist `{}` for it so the `arguments_json` column always holds valid JSON.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Runtime user messages and context injections both queue into
`pending_injections`, but with different keys: injections use the raw
injection id (committed under `inject-context:{id}`) while user messages
use the full `user-message:{id}` event id. The commit-correlation handler
only stripped the `inject-context:` prefix, so a user message's entry was
never removed.
Consequences:
- the client never received `ContextInjectionDelivered` /
`UserMessageAppended` for the message; and
- `pending_injections` stayed non-empty, so every later `ExecuteToolRound`
was detached without dispatching its tools and `tool_round::execute`
blocked forever -- a hung turn whenever the model made a tool call after
the interruption.
Derive the lookup key by stripping the injection prefix when present and
otherwise using the event id verbatim, so both kinds are cleared and their
delivered/appended events fire.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Introduced a new `WebCache` module to manage web content caching.
- Added functionality to store fetched content and serve it via a dedicated route.
- Integrated web cache into the search module for improved content retrieval.
- Implemented database migration management with detailed diagnostics for better error handling during startup.
- Updated the SQLite store to utilize the new migration system for enhanced database management.
- Added initial server setup with Cargo.toml defining dependencies and project structure.
- Created build.rs for generating protobuf bindings and validating wire contracts.
- Established database schema with initial migration files for conversations, messages, and runs.
- Introduced tools and prompts for Cursor functionality, enhancing user interaction capabilities.
- Introduced `first_valid_response_ms` and `ttfr_ms` to track the timing of the first valid response in LLM calls.
- Updated relevant interfaces and components to display and utilize the new metrics, including CallDetails, CallTable, and LatencyChart.
- Enhanced the database schema to accommodate the new timing fields.
- Implemented logic in the service layer to record the first valid response during model interactions.
- Introduced a new type `StatisticsStorageScope` to specify the scope for clearing statistics.
- Updated the `clearStatisticsStorage` API method to accept a scope parameter, allowing for selective clearing of detailed records or all statistics.
- Modified the demo API to handle the new scope parameter appropriately.
- Updated the SettingsPage component to include a selection for clearing scope, enhancing user control over statistics management.
- Added new translations for the updated messages related to statistics clearing in both English and Chinese.
- Introduced a new documentation site for Cursor BYOK using Next.js and Fumadocs.
- Added a product demo page with a corresponding Vite configuration.
- Implemented a demo API to simulate LLM calls and responses.
- Enhanced the Makefile to include new build and development commands for the documentation.
- Updated package.json scripts for building and running the documentation site.
- Created various components and layouts for the documentation structure, including blog and user documentation sections.
- Added styling for the new components and layouts to ensure a cohesive design.
- Included a README and other necessary files for local development and deployment.
- Changed `run_id` to `request_id` in `CursorParent` struct for clarity.
- Enhanced the `prepare` function to handle parent requests asynchronously, ensuring proper error handling for active runs.
- Updated database interactions to include `cursor_request_id` for better tracking of requests.
- Added tests to verify the behavior of reused cursor request IDs and their mapping to distinct executions.