claude-cost-audit-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@claude-cost-audit-mcpestimate cost for 12k input and 4k output tokens on Claude Sonnet 5"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
claude-cost-audit-mcp
An MCP server for exact Claude API cost calculation — correct prompt-cache economics (write 1.25x at 5-minute
TTL / 2x at 1-hour TTL, read 0.1x), cache break-even analysis, and a scanner for the single most common
Claude-cost mistake: estimating tokens with OpenAI's tiktoken.
The mistake this catches
Per Anthropic's own documentation: "Do not use tiktoken. It's OpenAI's tokenizer. It undercounts Claude tokens by
~15-20% on typical text, and by much more on code or non-English input." Any cost estimate or context-budget
check built on tiktoken for a Claude model is silently wrong — this scans source for that exact pattern
(foreign tokenizer + Claude/Anthropic usage in the same file) and points to the fix
(messages.count_tokens).
Related MCP server: tokentoll
Cache economics people get wrong
Caching doesn't pay off starting from the second request the way most people assume. The real break-even is
2 reuses at 5-minute TTL, but 3 reuses at 1-hour TTL — because the 1-hour write costs more (2x vs 1.25x).
check_cache_breakeven computes this from Anthropic's real multipliers instead of a guessed rule of thumb, and
flags prefixes below the model's own minimum-cacheable-token threshold (which is not monotonic across model
generations — 512 tokens on the newest models, 4096 on some older ones).
Tools
calculate_message_cost
Exact $ cost from token counts, using Anthropic's real current per-model pricing (including Claude Sonnet 5's time-limited introductory rate, resolved automatically from the date you pass).
check_cache_breakeven
Given expected request volume reusing the same prefix, tells you whether caching saves money and at which TTL.
audit_tokenizer_usage
Scans source for tiktoken/gpt-tokenizer/etc. alongside Claude API usage.
Use it
Hosted (recommended): MCPize — free tier, $7/mo Pro.
Self-host:
npm install
node server.jsPart of a small suite
mcp-schema-audit-mcp, cron-schedule-audit-mcp, regex-safety-audit-mcp.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Compare up-to-date pricing for 40+ LLMs (incl. Chinese) & estimate cost from tokens. EN/zh.
Token guard and rate limiter preventing runaway API cost spikes for OpenAI and Anthropic.
Find AI model pricing, estimate token costs and compare offers. No API key required.
Statically audits MCP tool surfaces for token cost, schema quality, and design issues.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceProvides intelligent analysis of token usage patterns and optimization recommendations to improve efficiency and reduce costs in Claude Code sessions. Offers real-time analysis, cost metrics, and actionable insights for better context window and tool usage optimization.3 npm-
- AlicenseAqualityCmaintenanceScan codebases for LLM API calls and estimate monthly costs. Compare costs between git refs to catch cost regressions during code review.24MIT
- AlicenseAqualityDmaintenanceEstimates Claude API token counts, per-model cost, and prompt-caching break-even without an API key or network.5MIT
- AlicenseNot gradedqualityBmaintenanceAnalyzes Claude Code session token usage and cost locally — where spend actually lands across cache-read, cache-write and output, and what is consuming the context window. Read-only and offline: it parses your own session files and exposes analyze_claude_cost, get_cost_benchmark and tokenscope_share_summary.113 npm4MIT