Portuguese and other interface languages have no commit prompt catalog; sending pt-BR made settings fail to deserialize.
Co-authored-by: Cursor <cursoragent@cursor.com>
Enable Portuguese-speaking users with full pt-BR translations and let users
set per-million token rates so home cost estimates reflect their actual pricing.
Co-authored-by: Cursor <cursoragent@cursor.com>
Root cause: Google's fetchAvailableModels API returns different model
catalogs depending on the endpoint. The daily endpoint returns 33
models including Gemini 3.8 Flash series, while the production
endpoint only returns 28 models without 3.8.
Additionally, the User-Agent header was using a stale Electron-style
format (Antigravity/4.3.0) that does not match the real Antigravity
client's UA (antigravity/hub/2.12.2), and unnecessary x-client-name /
x-client-version headers were being sent.
Changes:
- Reorder ANTIGRAVITY_ENDPOINTS to try daily endpoint first (matching
real Antigravity client behavior)
- Update ANTIGRAVITY_USER_AGENT to match the real Antigravity hub UA
- Remove unnecessary x-client-name/x-client-version headers
- Add Gemini 3.8 Flash (high/medium/low/tiered) to static model list
- Add gemini-3.8-flash passthrough in resolveAntigravityModel
`WebSearch::search` fuses the engines with Reciprocal Rank Fusion, which
sums one term per ranked list. `merge` added a term for every result
instead, so an engine that listed the same canonical URL twice had both
of its ranks counted:
1/61 + 1/62 = 0.0325
That is what two independent engines agreeing at rank 1 are worth
(2/61 = 0.0328), produced from a single list. Agreement across engines is
the only ranking signal this federation has, and a repeat inside one
engine forges it.
The repeats come from the repository's own canonicalization, not from
exotic input. `canonicalize` in `search/engine.rs` drops the fragment,
the trailing slash and the `utm_*`, `gclid`, `fbclid` and `mc_*`
parameters, and `redirected_target` unwraps the Bing, DuckDuckGo and
Google redirector links, so rows that are visibly distinct on one result
page collapse onto one URL. Nothing dedupes an engine's own list before
`merge` sees it.
`merge` already knew the rule: the engine label was guarded with
`!existing.engines.contains(&engine)`. Put the score behind the same
guard, so each engine contributes its best rank once. Cross-engine
merging and the longest-snippet rule are untouched.
Automatic compaction cannot do its job once a conversation crosses the
context window, so the conversation stays there permanently. Observed
against a 1M-token Anthropic window:
1. The summarize call replays the full history. It runs precisely
because that history is too large, so the request is itself over the
limit ("prompt is too long"), or it ends with an assistant/tool
message that Anthropic refuses as a prefill. Either way the run falls
back to the 12K truncated JSON summary, which discards the context.
In one trace the summarizer received 771 messages (2.78 MB) and
returned a single token.
2. The compaction check uses a 10K fixed reserve. The estimate trails
the provider's own count by the request context and provider-side
overhead that the message-tail estimate does not model; a 948K
estimate passed the check and Anthropic counted 1,017,628.
3. When the provider does refuse the prompt, the run retries the same
prompt eight times at 5s intervals and then fails. Nothing compacts.
Fixes, all in server/src/run:
- compaction_history trims the summarizer input to the context budget
at user-turn boundaries (never splitting a tool call from its
results) and guarantees it ends with a user message.
- context_budget keeps 10% of the window free instead of a fixed 10K,
so the reserve scales with the model and absorbs the drift.
- A provider refusal matching is_context_overflow compacts once and
retries the turn instead of failing it.
- 4xx responses other than 408/425/429 are terminal. A rejected request
fails identically every time, so retrying only delays the error.
Separately, Cursor can resume a finished turn whose checkpoint already
ends with the assistant, which Anthropic also rejects as a prefill.
run/history.rs appends a transient user tail to every provider request
that would otherwise end with the assistant. The tail is never
persisted, so committed checkpoints stay an exact prefix of the next
turn and the usage anchor still counts persisted messages only.
`gate_mcp_resources` caps the list at `MCP_RESOURCE_LIMIT` (200 resources),
then describes that truncation with `truncation_notice`, which is written
for byte budgets and was handed `MCP_TEXT_LIMIT`. A server returning 250
resources produced:
[truncated: ListMcpResources result exceeded 32768 bytes; showing 200 of 250 bytes]
Neither figure describes what happened: 32768 is a text budget this path
never applies, and the counts are resources rather than bytes. The notice
goes into a sentinel resource's description, so it is what the model reads
to learn why the list is short -- and it invites the conclusion that the
list was cut for size and would fit under a smaller byte budget.
State the cap that was actually applied, in its own unit, matching the
wording the sibling item-count truncation in `gate_mcp` already uses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`8942287` renamed `ProxyMode::System` to `ProxyMode::Default`, changing the
persisted wire value from `"system"` to `"default"`. No migration rewrites
the existing `service_settings` row and `ProxyMode` has no alias, so any
install that ever saved proxy settings on an earlier build now holds a row
this build cannot deserialize:
unknown variant `system`, expected `default` or `custom`
`proxy_settings_secret` turned that into a hard error, and it is the first
statement of every outbound client factory, so on upgrade the failure hits
model calls, plugin installs, search, and the Cursor upstream proxy alike.
It is also unrecoverable from the UI. `proxy_settings` reads the same row,
so the settings page cannot render the proxy card, and `set_proxy_settings`
reads the existing row before it writes, so the user cannot overwrite the
row that broke them. Only editing SQLite by hand clears it.
Read the row through a fallback that logs and returns the default instead.
The affected rows are exactly the ones whose mode meant "no outbound proxy",
which is what the default already is, so nothing is silently changed for
them; a genuinely corrupt row costs the user a re-entered address instead of
a dead install.
Deliberately not `#[serde(alias = "system")]`: `8942287` added a test
asserting that value no longer parses, and this keeps that true while making
the persisted row survivable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Updated the default proxy mode in `api.ts`, `ProxySettingsCard.tsx`, and `SettingsPage.tsx` to "default".
- Adjusted related translations in `catalog.json`, `en-US.json`, and `zh-CN.json`.
- Modified the `ProxyMode` enum in `settings.rs` to reflect the change from "system" to "default".
- Enhanced proxy handling in the server code to support the new default mode.
A provider that reuses a tool call id across two rounds of one run
wedges the run permanently. `ToolDispatcher::start_batch` skips any call
whose id is in `ToolBatchState::completed`, so the second call is never
dispatched and never produces a `ToolCompletion`, while
`tool_round::execute` blocks waiting for `calls.len()` results with no
timeout on that path. The client sees the tool call appear and then
nothing: no completion, no further output, no end-stream frame.
The `completed` set is built once per run and never cleared, so it is
run-scoped. A tool call id is only unique within a round, which the
schema already states as `UNIQUE (round_id, call_id)`; the sibling
runtime completed-map is likewise already cleared per round at
output.rs:615.
Clear `completed` when `ExecuteToolRound` begins a new round, tracked
independently of `active_round` so it does not depend on the order in
which ToolRoundStarted and ExecuteToolRound are observed. Replaying a
round still skips the calls that round already committed.
The practical trigger is openai_chat.rs:196, which synthesizes
`call-{index}` from a per-stream index when a provider omits tool call
ids, so `call-0` recurs on every model call. That file is left alone
here.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Changed the `ADS_ENDPOINT` from a local server URL to the production URL for ads.
- Commented out the local server URL for clarity and future reference.
- Added `app_version` field to `ControlService` and updated its initialization to include the app version.
- Modified the `ADS_ENDPOINT` to point to a local server for development purposes.
- Refactored conversation command and output handling to utilize a new `RunFinish` enum for better state management.
- Enhanced the conversation runtime to handle queued user messages after a turn has ended, ensuring smooth transitions between turns.
- Added tests to validate the new behavior of queued messages and transport handling.
- Updated the `append` function to include a flag for replacing closing requests, improving request handling.
- Refactored the `run_sse_handler` and `bidi_handler` functions to utilize a new tracing mechanism, enhancing observability.
- Introduced a new `trace_outcome` function to standardize tracing outcomes for requests.
- Removed the `CursorTraceRecorder` in favor of a new `CursorTraceService` for better performance and non-blocking behavior.
- Added tests to validate the new tracing functionality and ensure correct behavior during request processing.
- Introduced a new module for managing process resource limits, specifically for raising the open file limit on Unix systems.
- Added a `NetworkClients` struct to handle reusable outbound HTTP clients, improving network request management.
- Updated various components, including `ControlService` and `CursorProxy`, to utilize the new network client structure for better client handling.
- Enhanced the API router to accept network clients, ensuring consistent client usage across different services.
- Added tests to validate the integration of network clients and resource limits functionality.
- Introduced `UsageSnapshot` event to track token usage during conversation runs.
- Updated `RunEngine` to emit usage snapshots, providing better visibility into token consumption.
- Refactored compaction logic to utilize a new `compaction_estimate` function for improved token budget management.
- Added tests to validate timeout constants for blob synchronization and ensure correct behavior of usage tracking during compaction.
- Added a new test to validate the projection of reasoning response items to valid input items in the Codex API.
- Introduced `CallRecorder` to track network requests and responses during plugin interactions.
- Updated the `PluginRegistry` and `PluginWorker` to support call recording, ensuring that reasoning items are correctly processed and recorded.
- Refactored the `responses_input` function to handle reasoning items more effectively, improving the overall response handling logic.
- Added tracking for interaction events in the `Output` struct, including `summary_started` and `token_delta`.
- Updated the `run` function to push relevant interaction events to the `interaction_events` vector.
- Enhanced the automatic compaction test to verify the immediate reset of cursor usage and the correct logging of interaction events.