TokenControlPlane
TokenControlPlane
Stop runaway agents. One gateway in front of all your agent traffic — MCP tools and LLM models — one budget, one kill switch. Sits between AI clients (Cursor, Claude Desktop, Windsurf) and remote MCP servers / OpenAI-compatible model APIs; hard token budgets; a single static Go binary.
What problem does this solve?
AI agents can make many tool and model calls without a single place to control the total amount of traffic they generate.
With MCP, an agent may call multiple tools repeatedly, and with LLM APIs those calls can also consume significant token budgets. Controls implemented independently inside each client or MCP server make it difficult to enforce a shared budget, revoke access immediately, or see usage across tools and models.
TokenControlPlane puts a gateway between AI clients and their MCP servers / OpenAI-compatible LLM APIs.
It provides:
Per-key token budgets
Per-server token caps
A kill switch for individual gateway keys
Rate limiting and runaway-loop detection
Tool allow/deny policies
Usage and activity visibility
One gateway key across MCP and LLM traffic
Support for both remote HTTP MCP servers and local stdio MCP servers
OpenAI-compatible LLM routing through the same gateway
The goal is not to replace MCP authorization or provider-side limits. It provides a centralized control point for agent traffic.
Related MCP server: mcp-guardian
Why does MCP need this?
MCP defines how an AI application communicates with tools and resources, but it does not by itself provide a centralized budget or operational control plane for all of the traffic generated by an agent.
For example, an agent may have access to several MCP servers:
Agent
├── MCP server A
├── MCP server B
├── MCP server C
└── LLM providerWithout a shared gateway, each component can have its own credentials, limits, and monitoring.
TokenControlPlane changes the topology to:
┌── MCP server A
├── MCP server B
Agent ── TCP ────┼── MCP server C
└── LLM provider(s)
│
TokenControlPlaneThe gateway key becomes the control boundary.
This makes it possible to suspend one agent/key without changing the upstream MCP server, apply budgets across multiple services, and inspect MCP and LLM activity from one place.
TokenControlPlane also supports local stdio MCP servers by spawning the configured process and exposing it through the same gateway interface used for remote MCP servers.
Source: github.com/devthinker-ai/TokenControlPlane

Dashboard overview: servers, active keys, requests today, budget used, daily tokens, and top tools — MCP and LLM traffic in one place.
5-minute quickstart
docker run -p 8080:8080 -v mcp-data:/data ghcr.io/devthinker-ai/tokencontrolplane:latestOpen http://localhost:8080 → register
Add a remote MCP server (streamable-HTTP URL + API key, or OAuth sign-in for servers like Higgsfield — see docs/UPSTREAM_AUTH.md), or a local (stdio) server (
npx/uvx/ binary — see docs/STDIO.md). Browse tools and (Pro) control access — see docs/TOOL_POLICY.md.Generate a gateway key (
tcp_*)Paste the Claude Desktop / Cursor snippet into your client config
Call a tool — usage shows up on the dashboard; kill the key anytime → clients get 402
Total: about five minutes, zero other tooling.

Add remote HTTP or local stdio MCP servers; health, tool counts, and protocol in one list.

Issue tcp_* keys, set per-server token caps, kill a runaway key → clients get 402 on the next request.
Or build from source:
make build
./bin/tokencontrolplane -addr :8080make build-go builds without npm if pkg/web/dist already exists.
Updating
Migrations are forward-only (no down-migrations in v1). Before each unapplied migration the gateway writes *.db.backup via VACUUM INTO (one generation).
Docker
docker compose pull && docker compose up -dMigrations run on container start. Tag images as ghcr.io/devthinker-ai/tokencontrolplane:<TAG> and :latest on release (matches github.com/devthinker-ai/TokenControlPlane). Pinning :latest is fine because migrations are forward-only and tested; pin a TAG when you need to stay put.
Self-hosted binary
tokencontrolplane update --check # exit 1 if newer
tokencontrolplane update # download, verify SHA256SUMS, replace binary
tokencontrolplane rollback # restore tokencontrolplane.prev (one generation only)Air-gapped:
tokencontrolplane update --file ./tokencontrolplane_linux_amd64_v1.0.0 --sha256 <hex>
# or place SHA256SUMS in the cwd and omit --sha256Set NO_UPDATE_CHECK=1 or UPDATE_URL for mirrors. Rollback refuses if the DB schema is newer than the binary’s embedded migrations — restore a DB backup in that case.
Dashboard
GET /api/v1/version joins build metadata with the license update window. When a newer release exists and the window is active, the dashboard banner offers Update now (admin can run a verified server-side update). If the window expired, the banner points at renew — the version you own keeps working. Set NO_UPDATE_CHECK=1 to disable outbound checks (air-gap). See docs/RELEASING.md.
Stop runaway agents
Code | Meaning |
401 | Bad gateway key |
403 | Disabled key or server |
402 | Budget exhausted or kill switch (distinct JSON bodies) |
429 | Rate limit ( |
503 | Upstream circuit open |
Kill switch: kill a key from the dashboard or
/admin→ immediate 402 on the next request.Local (stdio) MCP servers: spawn
npx/uvx/ binaries; same client URL as remote — see docs/STDIO.md.One-time license, no subscription — pay once, self-host, keep the version you own.
Team seats — invite colleagues (no SMTP); named-team flat key (docs/TEAM.md).
LLM providers — OpenAI-compatible chat completions through the same keys/budgets (docs/LLM_ROUTING.md).
Auto-kill: >120 requests / 60s per key (configurable via
LOOP_THRESHOLD) → key auto-killed.Budgets: estimated tokens = bytes/4; enforced pre-flight (the oversized call completes; the next dies).

Tool policy: enable or disable what agents may call without changing the upstream.

LLM providers and model routes — same keys and budgets as MCP traffic.

Quick-add presets (Anthropic, Groq, Mistral, Gemini, Ollama, …) or a custom OpenAI-compatible base URL.

Activity: kills, budget hits, circuit open/closed, tool denied by policy, OAuth, plan changes.

Team seats and roles — humans in the dashboard, not API consumers.

Settings: authenticator 2FA for dashboard users; optional SMTP for invites and password reset.
Self-hosted licensing
Offline RS256 license JWTs (TOKENCONTROLPLANE_LICENSE_KEY). Unlicensed installs get the free tier (3 servers, 3 seats, 5M monthly tokens). Invalid/expired keys fail open to free — a bad key never takes down an air-gapped gateway.
Self-hosted licensing is offline and honor-based; keys are bound to plan caps, not machines, in v1. No phone-home.
Buy & install
No subscription. Pay once via Lemon Squeezy (merchant of record — they handle VAT/sales tax/refunds). You keep the version you bought; renew only for the next version. After checkout the license key appears in your dashboard License page and in the LS order email.
TOKENCONTROLPLANE_LICENSE_KEY=*** tokencontrolplane
# or in docker-compose:
# environment:
# TOKENCONTROLPLANE_LICENSE_KEY: "***"Sign keys manually with tokencontrolplane-license (private key never in the repo):
tokencontrolplane-license --private-key ~/.secrets/license.pem --in claims.jsonConfiguration
Everything is env vars (see tokencontrolplane -h):
Variable | Purpose |
| Dashboard session secret |
|
|
| Public URL for client snippets |
|
|
| Self-hosted license JWT |
| RSA private key path (dashboard mint / reissue) |
| Founders offer cap (default 100; |
| Lemon Squeezy API key |
| Webhook HMAC secret |
| LS catalog IDs |
| SQLite file (default |
| Override GitHub releases API (mirrors / air-gap) |
|
|
|
|
Docker Compose
cp .env.example .env # optional
docker compose up --buildSQLite persists under ./data.
Development
go test ./...
cd frontend && npm test && npm run buildCGO is off (CGO_ENABLED=0); SQLite is pure Go (modernc.org/sqlite).
Licensing
Two different things are licensed here:
The code — Apache-2.0 (LICENSE). Free to use, modify, and redistribute, including commercially. You must keep the license and copyright notices. There is no requirement to share changes back, no copyleft, and no fee.
The license key — commercial license. A one-time purchase (see Buy & install) that unlocks Pro/Team plan caps. The key is an offline RS256 JWT signed with the vendor's private key; it grants plan capabilities, not a copy of the code.
This is the standard open-core model: the source is fully open and auditable (self-hosters can verify exactly what runs), while revenue comes from plan-gated capabilities delivered via a signed key.
What this means in practice:
You can… | Under Apache-2.0 |
Run the gateway free (free tier, fail-open) | Yes |
Buy a key for Pro/Team caps | Yes |
Fork, modify, or redistribute the source, even commercially | Yes |
Forge a license key | No — only the vendor's private key can sign one |
Honor-based gating. The license check is plain source code, so a determined fork can change which caps it enforces. That is permitted by Apache-2.0. Enforcement is honor-based, not cryptographic — like other open-core products (GitLab, Keycloak). No phone-home: an unlicensed or expired key fails open to the free tier and never takes a gateway down.
Private key hygiene. Only the public key is committed (pkg/license/license_pub.pem). The signing private key is kept out of the repo and injected at build/sign time via TOKENCONTROLPLANE_LICENSE_PRIVATE or --private-key.
This server cannot be deployed
Maintenance
Related MCP Connectors
AgentGuard — 20-tool AI safety MCP: policy preflight, risk scoring, audit logging, rate limits.
Zero-trust gateway for AI agents: score tool calls, verify agent cards, enforce policy, audit.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Security gateway for AI agents: policy, approval, and audited execution, no secrets shared.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceGuardrail sidecar proxy between AI agents and their MCP/REST/CLI tools. Policy engine, human approval gates, time-limited grants, rate limiting, and OTEL tracing. One Go binary, one YAML config, fail-closed by default.1Apache 2.0
- AlicenseAqualityBmaintenanceSecurity, cost, and health governance proxy for MCP infrastructure. Enforces YAML-configurable security policies (blocklists, rate limits, token budgets), tracks real token costs via tiktoken, monitors server health with live JSON-RPC probes. Features OAuth 2.1/OIDC with RBAC, web dashboard, payload normalization, semantic shell AST analysis, mTLS, and a formal STRIDE threat model.4173 npm3MIT
- AlicenseNot gradedqualityCmaintenanceBudget & cost control for AI agents: hard per-agent spend caps, rate limits, idempotency, and human-in-the-loop approval — enforced before each LLM call, not after the invoice. One hosted MCP endpoint (no proxy or self-hosting), settled via x402 (USDC on Base).MIT
- AlicenseAqualityAmaintenanceRuntime governance and budget guardrails for Claude Code, Cursor, and autonomous AI agents. Enforces per-session spend caps, verifier safety gates, and runaway loop prevention.24673 npm364Apache 2.0