Skip to main content
Glama
KitsuneTech1

Token Guardian MCP

by KitsuneTech1
README.md
# Token Guardian MCP

Token Guardian is a local, read-only MCP server for finding Claude Code and Codex token waste without making either client dumber.

It does two jobs:

- Reports which agents, models, and sessions are processing the most tokens.
- Recommends the cheapest safe model and effort level for a specific task.

It also includes a dry-run `route-task` command for testing prompt-first routing before the MCP is registered anywhere.

## Hard safety boundary

Token Guardian cannot change model settings, effort levels, context limits, compaction, MCP registration, or running sessions. It has no model credentials and makes no paid API calls.

The usage tool calls `ccusage` with `--offline` and `--no-cost`. Codex titles are optional metadata read through `sqlite3 -readonly`. Every MCP tool is marked read-only, non-destructive, idempotent, and closed-world.

## Tools

### `token_guardian_usage_snapshot`

Reads one through thirty days of local usage and returns:

- Totals by agent and model
- Cache-read, output, and frontier-model shares
- Largest sessions
- Exact duplicate Codex work titles
- Evidence-backed quick wins

Processed tokens are a workload diagnostic. They are not the same as a subscription meter, especially when cached input dominates.

### `token_guardian_recommend_route`

Accepts the client, task, risk, and current context size. It returns a model, effort level, reasons, and an optional frontier validator.

The policy fails closed. Security, architecture, production, destructive, high-risk, critical, or ambiguous work stays on a frontier model. Cheap models handle bounded mechanical work. Balanced models handle normal coding and debugging, with a frontier review when needed.

## Test prompt-first routing

Build the project, then run the router from any folder:

```powershell
node ./dist/route-cli.js --client codex --prompt "Implement a bounded TypeScript parser with tests." --cwd .
```

The working folder does not need to be a repository. The router primarily uses the prompt. It checks up to 256 names in the current folder for optional project markers, without opening file contents or scanning subfolders.

The command only prints a recommendation and a session-scoped launch command. It does not run that command, change configuration, register the MCP, or touch an active session. Vague continuation prompts such as `yeah, test that` stay on the frontier route because their real context is missing.

## Local requirements

- Node.js 24 or newer
- `ccusage` on PATH
- `sqlite3` on PATH for optional Codex thread titles

## Build and test

```powershell
npm install
npm test
npm run typecheck
npm run build
```

Run the built stdio server:

```powershell
node dist/index.js
```

Inspect it without registering it in a live client:

```powershell
npx @modelcontextprotocol/inspector node dist/index.js
```

## Registration status

This first build is intentionally not registered with Claude, Codex, or the shared Kitsune gateway. Registration and any client restart are a separate change after the isolated server is accepted.

## License

MIT. See [LICENSE](LICENSE).

TDQS

A3.7/5.0

Scored across 2 tools

Disambiguation5/5

Each tool serves a distinct, non-overlapping purpose: one provides a usage snapshot with improvement suggestions, the other recommends a routing choice for a task. There is no ambiguity between them.

Naming Consistency5/5

Both tool names follow a consistent 'token_guardian_<verb>_<noun>' pattern using snake_case, with descriptive verbs ('usage_snapshot', 'recommend_route') that clearly indicate their function.

Tool Count3/5

With only 2 tools, the server feels minimal for a domain that might benefit from additional diagnostics (e.g., cost breakdown, model listing) or configuration advice. The count is on the low end of acceptable for a focused utility.

Completeness2/5

The server explicitly avoids any action tools, being read-only. It covers snapshot analysis and routing recommendations but lacks tools for detailed queries (e.g., by time period, by model) or applying any changes, leaving notable gaps for hands-on usage optimization.

Maintenance

ActivitySlowing
ResponsivenessNo issues