Automatic compaction cannot do its job once a conversation crosses the
context window, so the conversation stays there permanently. Observed
against a 1M-token Anthropic window:
1. The summarize call replays the full history. It runs precisely
because that history is too large, so the request is itself over the
limit ("prompt is too long"), or it ends with an assistant/tool
message that Anthropic refuses as a prefill. Either way the run falls
back to the 12K truncated JSON summary, which discards the context.
In one trace the summarizer received 771 messages (2.78 MB) and
returned a single token.
2. The compaction check uses a 10K fixed reserve. The estimate trails
the provider's own count by the request context and provider-side
overhead that the message-tail estimate does not model; a 948K
estimate passed the check and Anthropic counted 1,017,628.
3. When the provider does refuse the prompt, the run retries the same
prompt eight times at 5s intervals and then fails. Nothing compacts.
Fixes, all in server/src/run:
- compaction_history trims the summarizer input to the context budget
at user-turn boundaries (never splitting a tool call from its
results) and guarantees it ends with a user message.
- context_budget keeps 10% of the window free instead of a fixed 10K,
so the reserve scales with the model and absorbs the drift.
- A provider refusal matching is_context_overflow compacts once and
retries the turn instead of failing it.
- 4xx responses other than 408/425/429 are terminal. A rejected request
fails identically every time, so retrying only delays the error.
Separately, Cursor can resume a finished turn whose checkpoint already
ends with the assistant, which Anthropic also rejects as a prefill.
run/history.rs appends a transient user tail to every provider request
that would otherwise end with the assistant. The tail is never
persisted, so committed checkpoints stay an exact prefix of the next
turn and the usage anchor still counts persisted messages only.
cursor-byok
cursor-byok is a local implementation of Cursor's backend.
About
cursor-byok is an open-source local model gateway for Cursor. It runs a service on your machine that connects Cursor to the model APIs you configure, routes model requests through your own providers, and preserves Cursor Agent capabilities such as tool calling, Skills, and MCP.
You can connect OpenAI- and Anthropic-compatible services, customize endpoints, model IDs, API keys, and request parameters, and use model channels beyond the options built into the platform.
Important
cursor-byok is free and open source, but the model APIs you connect may charge for usage. This is an independent project and is not affiliated with or endorsed by Cursor or its developers.
Features
- Bring your own model channels: Configure your own API endpoint, credentials, and model IDs.
- Multiple API protocols: Use OpenAI- and Anthropic-compatible APIs or a custom endpoint.
- Model management: Add, duplicate, edit, reorder, and batch-test multiple model configurations.
- Connection benchmarks: Measure time to first token, generation speed, and inspect raw provider responses.
- Agent workflows: Keep tool calling, Skills, MCP, and multi-turn conversations available.
- Session metrics: Track token usage, cache hit rate, conversation turns, and estimated value.
- Cross-platform: Run on macOS, Windows, and Linux.
Quick Start
- Download the latest build for your platform from GitHub Releases.
- Launch cursor-byok, open Model Settings, and enter the endpoint, API key, and model ID.
- Test the model configuration. Once it passes, return to the dashboard and start the service.
- Test the model configuration. Once it passes, return to the dashboard and start the service.
- After upgrading Cursor or configuring a model for the first time, quit Cursor completely and restart it, then start a new conversation and select the configured model.
For complete installation steps, system configuration, and Frequently Asked Questions, see the User Guide.
Model Management
Model configurations support both OpenAI and Anthropic API protocols. Each model channel can independently define its context window, maximum output tokens, reasoning effort, custom headers, and additional request parameters.
How It Works
Cursor client
│
│ Agent requests and tool results
▼
cursor-byok local service
│
│ OpenAI- / Anthropic-compatible requests
▼
Your model API
cursor-byok handles protocol adaptation, model request forwarding, tool-call coordination, and conversation state on your machine. API keys and application settings are stored locally; requests are still sent to the model provider you configure.
Why This Project
Many Agent products bundle their tool capabilities with a fixed set of models, subscriptions, and billing options, leaving users limited to the channels offered by the platform.
cursor-byok is built to return model choice to the user. Developers can make full use of the APIs and credits they already have, choose the models and providers that fit their needs, and self-host related services when required.
Roadmap
The project will continue to improve model compatibility, Agent tooling, local runtime stability, and the self-hosting experience while exploring support for more IDE, chat, and Agent workflows.
See the release roadmap for plans and progress.
Community and Support
- User Guide
- GitHub Issues
- Telegram community
- QQ groups:
1095916242,1094411438,1095918002,1094419321
Development and Contributing
Issues and pull requests are welcome. See the Contributing Guide for prerequisites, build commands, project structure, and contribution guidelines.
Contributors
License
This project is open source under the MIT License.


