Skip to main content
Glama

tokentoll

在代码审查中捕获 LLM 成本变化。LLM 支出的 Infracost。

CI PyPI version GitHub Marketplace License: MIT Python 3.10+

一个 CLI 工具和 GitHub Action,用于静态分析代码中的 LLM API 调用,估算其成本,并在终端或 PR 评论中向您展示每次更改的成本影响。零运行时依赖。

问题所在

将模型从 gpt-4o-mini 切换到 gpt-4o 会使成本增加 15 倍。 热路径中的一个新 API 调用可能会让您的账单每月增加 10,000 美元。 这些变化隐藏在正常的代码审查中。

tokentoll 可以查找代码中的 LLM API 调用,估算其成本,并在更改进入生产环境之前向您展示每次更改的成本影响。

Related MCP server: CosTrack MCP

快速入门

pip install tokentoll

# Scan current directory for LLM API calls and their costs
tokentoll scan .

# Show cost impact of your last commit
tokentoll diff HEAD~1

# Compare two branches
tokentoll diff main..feature-branch

GitHub Action

name: LLM Cost Diff
on:
  pull_request:
    paths:
      - "**.py"

permissions:
  pull-requests: write

jobs:
  cost-diff:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: Jwrede/tokentoll@v0.6.1

检测范围

SDK

模式

状态

OpenAI

chat.completions.create, responses.create

已支持

Anthropic

messages.create, messages.stream

已支持

Google GenAI

models.generate_content

已支持

LiteLLM

completion, acompletion

已支持

LangChain

ChatOpenAI, ChatAnthropic, init_chat_model

已支持

Zhipu AI

ZhipuAiClient, ZhipuAI (GLM 模型)

已支持

JS/TS SDKs

计划中

输出示例

tokentoll scan

LLM API Calls Detected
============================================================

File: src/agents/summarizer.py
  Line 42: openai client.chat.completions.create
           Model: gpt-4o | Max tokens: 4096
           Est. cost/call: $0.03 | Monthly (1000 calls/month per call site): $26.50

  Line 78: openai client.chat.completions.create
           Model: gpt-4o-mini | Max tokens: 1000
           Est. cost/call: $0.000301 | Monthly (1000 calls/month per call site): $0.30

--
Total estimated monthly cost: $26.80
  1000 calls/month per call site

tokentoll diff

LLM Cost Diff: main..feature-branch
============================================================

+ ADDED    src/agents/rewriter.py:35
           openai | Model: gpt-4o
           Est. cost/call: $0.03 | Monthly: +$26.50

~ MODIFIED src/agents/summarizer.py:42
           openai | Model: gpt-4o -> gpt-4o-mini
           Est. cost/call: $0.03 -> $0.000301 | Monthly: -$26.20

--
Monthly cost impact: +$0.30
  Added: 1 | Changed: 1 | Removed: 0
  1000 calls/month per call site

工作原理

  Source Code (.py files)
         |
         v
  +-------------+     +------------------+
  | AST Scanner |---->| SDK Detectors    |
  | (ast.parse) |     | OpenAI, Anthropic|
  +-------------+     | Google, LiteLLM  |
                       | LangChain        |
                       +------------------+
                              |
                              v
                       +------------------+
                       | Pricing Engine   |
                       | 2200+ models     |
                       | Auto-cached      |
                       +------------------+
                              |
                  +-----------+-----------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Scan Report|         | Diff Engine  |
           | (costs)    |         | (old vs new) |
           +------------+         +-------------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Table/JSON |         | Table/JSON/  |
           |            |         | PR Comment   |
           +------------+         +-------------+
  1. 使用 ast 模块解析 Python 文件以查找 LLM API 调用

  2. 多遍常量传播通过变量、os.getenv() 回退、类属性、构造函数参数、字典内容和 **kwargs 解包来解析模型名称

  3. 从本地缓存(源自 LiteLLM,2200+ 模型)中查找定价

  4. 对于 diff 模式:比较两个 git 引用之间的调用并计算成本增量

  5. 以表格、JSON 或 GitHub PR 评论的形式输出成本报告

CLI 参考

tokentoll scan [PATH...] [--format table|json|markdown] [--calls-per-month N] [--config PATH]
tokentoll diff [REF] [--base REF] [--head REF] [--format table|json|markdown|github-comment] [--config PATH]
tokentoll update    # Update bundled pricing data

MCP 服务器

tokentoll 包含一个 MCP(模型上下文协议)服务器,允许 Claude Code 和其他 MCP 主机直接从代理对话中检查 LLM 代码更改的成本影响。

安装

pip install tokentoll[mcp]

在 Claude Code 中注册

claude mcp add --transport stdio tokentoll -- tokentoll-mcp

工具

工具

描述

scan

在目录中查找 LLM API 调用并估算每月成本。接受路径和可选的 calls_per_month。

diff

比较两个 git 引用之间的 LLM 成本。接受 base_ref 和可选的 head_ref(默认为 HEAD)。

两个工具均返回 JSON 输出。

使用案例示例

Claude Code 可以在提交前检查其自身更改的成本影响。例如,在将模型从 gpt-4o 切换到 gpt-4o-mini 后,代理可以针对 HEAD 调用 diff 工具,以在创建提交之前验证成本降低情况。

定价数据

定价已捆绑并可离线工作。要更新到最新价格:

tokentoll update

定价数据源自 LiteLLM 的 model_prices_and_context_window.json,涵盖了 OpenAI、Anthropic、Google、AWS Bedrock、Azure 等 300 多个模型。

动态模型默认值

当 tokentoll 遇到模型名称为无法解析的变量的调用时,它会应用合理的每 SDK 默认值,以便您仍然可以获得成本估算:

SDK

默认模型

OpenAI

gpt-4o

Anthropic

claude-sonnet-4-20250514

Google GenAI

gemini-2.0-flash

LiteLLM

gpt-4o

LangChain

gpt-4o

Zhipu AI

zai/glm-4.6

这些默认值在扫描输出中显示为 gpt-4o (default)。您可以使用 .tokentoll.yml 配置文件(见下文)按项目或按路径覆盖它们。

配置

在项目根目录中创建 .tokentoll.yml 以自定义行为。tokentoll 会通过从扫描目录向上遍历来自动查找此文件。

# Default model for all dynamic (unresolved) calls
default_model: gpt-4o

# Per-SDK defaults (override the built-in defaults above)
default_models:
  openai: gpt-4o-mini
  anthropic: claude-haiku-3-20240307

# Assumed calls per month per call site
calls_per_month: 5000

# Skip cost estimation entirely for dynamic (unresolved) models. When true,
# calls whose model name cannot be resolved statically are reported with no
# cost rather than priced against a default. Useful for projects that prefer
# silence over a guess.
skip_dynamic_models: false

# Exclude paths from scanning (prefix match or glob pattern)
exclude:
  - tests/
  - examples/
  - docs/
  - "*_test.py"

# Per-path overrides (longest prefix match)
overrides:
  - path: src/agents/
    default_model: gpt-4o
    calls_per_month: 10000
  - path: src/azure/
    skip_dynamic_models: true

动态模型默认值的解析顺序:每 SDK 配置 (default_models) > 通用配置 (default_model) > 内置 SDK 默认值。

您还可以传递 --config path/to/.tokentoll.yml 来使用特定的配置文件。

Token 估算

默认情况下,tokentoll 使用字符/4 的启发式方法估算 token 数量。为了获得更准确的估算,请安装 tiktoken:

pip install tiktoken

当 tiktoken 可用时,tokentoll 会为每个模型使用正确的 tokenizer 编码。未知模型会回退到 cl100k_base。Tiktoken 是延迟加载的,编码器会被缓存,因此如果您不需要它,则不会有启动惩罚。

智能变量解析

真实代码库很少将模型名称作为字符串字面量传递。tokentoll 的多遍常量传播引擎遵循:

DEFAULT_MODEL = os.getenv("MODEL", "gpt-4o")

class Config:
    model: str = DEFAULT_MODEL

config = Config()
kwargs = {"model": config.model, "max_tokens": 2000}
client.chat.completions.create(**kwargs)
# tokentoll resolves: model="gpt-4o", max_tokens=2000
  • 变量赋值 (MODEL = "gpt-4o")

  • os.getenv() / os.environ.get() 回退值

  • 函数默认参数

  • 类属性默认值

  • 构造函数参数传播

  • 字典字面量和下标内容

  • **kwargs 解包

路线图

  • 上下文感知调用频率(计划中):从周围代码推断每月调用次数(FastAPI 路由处理程序 = 高流量,脚本 = 低,循环 = 乘数),而不是假设所有调用点具有统一的容量。

  • JS/TS 支持(计划中):检测 JavaScript 和 TypeScript 文件中的 LLM 调用。

  • 成本警报:可配置的阈值,当 PR 超过成本增量时使 CI 失败。

限制

  • 无法解析在运行时从外部配置文件或数据库加载的模型。这些调用使用每 SDK 默认值(可通过 .tokentoll.yml 配置)。

  • 除非安装了 tiktoken,否则 token 估算使用字符/4 的启发式方法。

  • 每月估算假设每个调用点的调用量是统一的(可通过 --calls-per-month、.tokentoll.yml 或路径覆盖进行配置)。使用 exclude 选项跳过测试和示例文件。

  • 目前仅支持 Python(计划支持 JS/TS)。

许可证

MIT

Available Tools

2 tools
diffA

Compare LLM costs between two git refs.

Shows which LLM call sites were added, removed, or changed between the base and head refs, along with the cost impact of those changes.

Args: base_ref: The base git ref (branch, tag, or commit) to compare from. head_ref: The head git ref to compare to. Defaults to HEAD.

Returns: JSON string with the diff results including cost changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
base_refYes
head_refNoHEAD

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It indicates a read-like operation (diff) and describes the output, but does not explicitly state side effects or permissions. Adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with purpose, and includes parameter docs and return type. Every sentence adds value without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema (not shown), the description adequately covers purpose, parameters, and output format. It could include examples or edge cases but is sufficiently complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description documents both parameters: base_ref as the base git ref and head_ref as the head ref defaulting to HEAD. This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares LLM costs between two git refs, specifying it shows added, removed, or changed call sites and cost impact. This distinguishes it from the sibling 'scan' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (to compare costs between refs) but does not explicitly state when not to use it or mention alternatives. Usage is well implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scanA

Scan a directory for LLM API calls and estimate monthly costs.

Finds all LLM API call sites (OpenAI, Anthropic, etc.) in the given path and produces a cost estimate based on token counts and pricing.

Args: path: Directory or file path to scan. Defaults to current directory. calls_per_month: Assumed monthly call volume per call site. If not provided, the CLI default (1000) is used.

Returns: JSON string with the scan results including call sites and cost estimates.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
calls_per_monthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It details the scanning action, cost estimation, and return format. While it doesn't cover every edge case (e.g., recursion depth or error handling), it provides sufficient behavioral insight for a read-only analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a lead sentence, then details in Args and Returns sections. Every sentence adds value, and the format is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 optional params, no annotations), the description covers the core behavior and return type adequately. It could mention recursion or failure modes, but it is sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the description fully explains both parameters: 'path' (directory/file, default current dir) and 'calls_per_month' (monthly volume, default null implying CLI default of 1000). This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans a directory for LLM API calls and estimates costs, specifying providers and purpose. This is a specific verb+resource that distinguishes it from the sibling 'diff'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use the tool (scanning directories for LLM calls and cost estimation). However, it does not explicitly mention when not to use it or provide alternatives, which prevents a top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observeddiff
    • First observedscan

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools, diff and scan, have clearly distinct purposes: scan finds LLM call sites and estimates costs, while diff compares costs between git refs. No overlap or ambiguity.

Naming Consistency4/5

Both tool names are single verbs ('diff', 'scan'), which is consistent in style. While not a verb_noun pattern, the naming is uniform and intuitive for the domain.

Tool Count3/5

With only 2 tools, the server is very focused. This can be appropriate for a narrow utility, but it feels thin for a full server. A few more tools (e.g., pricing config) might improve scope.

Completeness3/5

The tools cover two core operations: scanning and diffing. However, there is no tool for managing pricing configurations or listing assumptions, which could be gaps for advanced use.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI cost calculation, comparison, and optimization across major providers like Anthropic, OpenAI, Google, Meta, and Mistral. Supports cost estimation, budget-aware model finding, and token estimation through a simple API and MCP integration.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Exposes boyter/scc code counting and complexity analysis to LLM agents via read-only tools like counting lines, finding top files, and cost estimation.
    7
    BSD 3-Clause