AI Cost Lab MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AI Cost Lab MCP Servercompare the API costs of GPT-4o and Claude 3.5 Sonnet for 10k input tokens"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI Cost Lab
An educational calculator for comparing text-model API costs and the operating cost of an AI workflow. Built by Bhanu Vadlakonda. MIT licensed.
What it does
Compare up to three models using the same input, output, cached-token and call assumptions
Estimate cost per resolved business task with capped retries, tool fees and human escalations
Allocate setup costs across months and shared monthly infrastructure costs across workloads
Explore four synthetic token-budget workflows and three editable embedding/infrastructure examples
Expose the same deterministic math through four read-only MCP tools
It makes no model API calls and needs no model-provider API key. No analytics, database, account registration or request-body logging is included in this code. Host and proxy logging policies remain the deployer's responsibility.
Related MCP server: mcp-gateway
Run locally
Requires Node.js 22 or later. There are no npm dependencies, and no install or lockfile is required.
npm test
npm startOpen http://127.0.0.1:3000. npm test builds the Worker and runs all suites. After changing assets or the MCP server, run npm run build before starting the server. The default listener is loopback only; HOST and PORT configure it.
The website works without a token. MCP tool calls are denied until the server is configured with a private AI_COST_LAB_MCP_TOKEN environment variable. Clients send that value in Authorization: Bearer .... Keep it outside source control and URLs. Do not paste real credentials into examples, issues or browser inputs.
Deploy the website
For a website-only deployment, upload only the eight top-level files inside dist/ to any static host. Do not upload dist/server/. The UI runs entirely in the browser and does not require the MCP endpoint. Serving the generated Worker is another option.
For website plus MCP, npm run build generates dist/server/index.js, a dependency-free Web Standards Worker with a fetch(request, env) entry point. Supply env.AI_COST_LAB_MCP_TOKEN using the host's secret configuration. Alternatively run npm start behind an HTTPS reverse proxy. Configure TLS, rate limits, resource limits and operational logging appropriately before Internet exposure. The adapter explicitly serves an asset allowlist and /mcp; it does not expose source or environment files.
The portable adapter is a small bearer-token reference implementation, not an OAuth authorization server. Some ChatGPT or other MCP clients require OAuth discovery and registration; those clients need a compatible authenticated gateway. This repository does not install a plugin or reproduce any hosted service's account bindings. Neither the Node nor Worker adapter trusts client-supplied identity headers.
MCP tools
POST JSON-RPC to /mcp. Supported protocol versions: 2024-11-05, 2025-03-26 and 2025-06-18. Stateless initialization, tool listing and calls are supported; no SSE, sessions, resources or prompts. Initialization and tool discovery contain public definitions and are available without authentication. Data-bearing tool calls require adapter authentication.
compare_model_api_costs: identical numeric workload across 1–3 curated models; returns token subtotalsestimate_workflow_cost: explicit per-model success assumptions, retry/escalation inputs and optional additional costsget_workflow_examples: synthetic step-by-step token examplesget_infrastructure_scenarios: editable sourced-rate example budgets and their mapping into workflow costs
Schemas live in src/mcp.mjs. All tools are read-only, deterministic and accept numeric assumptions or defined identifiers. Do not submit customer text or secrets. Unknown or invalid fields are rejected. Requests are limited to 64 KiB. User success rates are assumptions, never quality benchmarks inferred from prices.
Understanding the results
API view is a token subtotal. Cached tokens are a subset of total input, never counted twice. Tokens differ between model tokenizers; matching numeric token counts approximates matching workloads. Include billable reasoning in output assumptions.
One workflow attempt may contain multiple model calls. A retry repeats the entire attempt and stops after success. With success probability p and r retries, expected attempts are sum((1-p)^i, i=0..r). Only remaining failures may escalate to a human; AI and human resolutions do not overlap. These assumptions treat attempts as independent with constant costs, which may not match real failure patterns.
Additional costs include:
One-time parsing, embedding and indexing, amortized over chosen months
Query embeddings, retrieval, reranking, transfer and other usage fees per complete attempt
Monthly vector storage, compute, monitoring and maintenance
Setup and fixed costs use the allocated share; per-attempt usage does not. Avoid counting the same provider bill in multiple categories. Blank/omitted costs remain unknown. Explicit zero means reviewed and not applicable or counted elsewhere. A complete result means the entered scope is complete, not a guarantee of total business cost. At zero volume, fixed/setup costs still exist and per-task costs are undefined.
Infrastructure presets use hypothetical sizes and labour budgets with dated provider rates. They do not establish capacity. Preset selection is a preview; click “Use this example in Workflow cost” to apply it. Query usage follows actual expected attempts including retries; fixed budgets do not auto-scale. The static corpus assumption excludes ongoing corpus updates, backups and high availability. Parsing and indexing stay unknown. Scenarios below the provider minimum or outside valid mapped cost bounds cannot be loaded; this prevents misleading linear approximations.
Prices and sources
Presets are a snapshot checked 2026-10-01, in USD per million tokens, with editable rates in the UI. They are not live quotes. Review official pricing before relying on results.
Model records, exact official URLs and caveats:
dist/models.mjsInfrastructure rates, formulas and source URLs:
dist/infrastructure.mjs
Gemini 3.8 Flash rates in this snapshot are promotional through 2026-12-31; update them before relying on later estimates. Cache writes/storage, long-context and regional premiums, taxes, negotiated discounts and unentered expenses may be excluded. See visible warnings and tool metadata. The infrastructure example uses self-managed PostgreSQL/pgvector on Railway, not a managed database capacity guarantee. Labour hours and hourly rates are illustrative, not vendor prices or industry benchmarks.
Source map and tests
dist/calc.mjs: shared pure calculation and validation enginedist/models.mjs,dist/workflows.mjs,dist/infrastructure.mjs: data and illustrative assumptionsdist/app.mjs,dist/index.html,dist/style.css,dist/favicon.svg: browser appsrc/mcp.mjs,src/auth.mjs: protocol, tools and adapter authenticationbuild.mjs: bundles an explicit frontend allowlist and server modulesserver.mjs: Node HTTP adaptertest*.mjs: 391 assertions covering math, bounds, completeness, workflow normalization, UI events, MCP parity and authentication
UI tests use a simulated DOM; they are not a substitute for browser, accessibility or production deployment testing. CI runs the same dependency-free suite on Node 22. No dependency lockfile is needed because no packages are installed. The source is an initial standalone export; hosting history and private configuration are intentionally absent.
Contributing
Run npm test before proposing changes. Include tests for calculation changes and retain unknown-versus-zero semantics. Keep rate changes dated and link primary sources. Do not add fabricated model-quality scores. Review generated Worker output locally, but do not commit it. Do not include credentials or private customer assumptions in reports.
License and attribution
MIT © 2026 Bhanu Vadlakonda. See LICENSE. See THIRD_PARTY_NOTICES.md for dependency and source attribution notes.
Document-size teaching examples
The infrastructure cards translate the existing embedded-token budget into approximate ten-page PDFs or 1,000-word articles. Assumptions are visible: 500 English words/page, roughly four tokens per three words, and 25% extra for repeated text between passages. These alternative equivalents are not measured document sizes, a universal language conversion or a change to the billed-token calculation. Scanned pages, complex tables and images can require extra work that remains unpriced. The whole library is prepared for search initially; each question sends only selected relevant passages to the answering model.
This server cannot be deployed
Maintenance
Related MCP Connectors
Sandbox workspace tools: search, file read, DB queries, integrations. Returns synthetic data.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Discover Frontier inference capabilities and read sanitized usage through read-only tools.
A read-only verified record of agent-operable GTM tools: search, fetch, compare, track changes.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides read-only tools to generate test plans for AI agents, score agent runs, and produce reusable evaluation scorecards.3MIT
- AlicenseNot gradedqualityCmaintenanceEnables a language model to safely query internal services through a closed set of read-only, schema-validated tools, with full auditing and refusal logging.MIT
- FlicenseNot gradedqualityCmaintenanceEnables MCP tool calls with strict schema validation and stdio isolation, while providing a security gateway for tool-level authorization, streaming PII redaction, and model failover routing.-
- AlicenseNot gradedqualityCmaintenanceEnables external MCP clients to invoke governed tools with schema validation, budgets, policy checks, human approval for side-effecting actions, and prompt-injection quarantine.MIT