Skip to main content
Glama

tokentoll

コードレビューでLLMのコスト変動をキャッチ。LLM支出のためのInfracost。

CI PyPI version GitHub Marketplace License: MIT Python 3.10+

LLM API呼び出しを静的に解析し、コストを試算して、ターミナルやPRコメントで変更によるコストへの影響を表示するCLIツールおよびGitHub Actionです。ランタイム依存関係はゼロです。

課題

gpt-4o-mini から gpt-4o へのモデル変更だけで、コストは 15倍 に増加します。 ホットパスに新しいAPI呼び出しを追加すると、月額 10,000ドル の請求増につながる可能性があります。 こうした変更は通常のコードレビューでは見落とされがちです。

tokentollはコード内のLLM API呼び出しを見つけ出し、コストを試算して、本番環境に反映される前に変更によるコストへの影響を表示します。

Related MCP server: CosTrack MCP

クイックスタート

pip install tokentoll

# Scan current directory for LLM API calls and their costs
tokentoll scan .

# Show cost impact of your last commit
tokentoll diff HEAD~1

# Compare two branches
tokentoll diff main..feature-branch

GitHub Action

name: LLM Cost Diff
on:
  pull_request:
    paths:
      - "**.py"

permissions:
  pull-requests: write

jobs:
  cost-diff:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: Jwrede/tokentoll@v0.6.1

検出対象

SDK

パターン

ステータス

OpenAI

chat.completions.create, responses.create

サポート済み

Anthropic

messages.create, messages.stream

サポート済み

Google GenAI

models.generate_content

サポート済み

LiteLLM

completion, acompletion

サポート済み

LangChain

ChatOpenAI, ChatAnthropic, init_chat_model

サポート済み

Zhipu AI

ZhipuAiClient, ZhipuAI (GLMモデル)

サポート済み

JS/TS SDKs

計画中

出力例

tokentoll scan

LLM API Calls Detected
============================================================

File: src/agents/summarizer.py
  Line 42: openai client.chat.completions.create
           Model: gpt-4o | Max tokens: 4096
           Est. cost/call: $0.03 | Monthly (1000 calls/month per call site): $26.50

  Line 78: openai client.chat.completions.create
           Model: gpt-4o-mini | Max tokens: 1000
           Est. cost/call: $0.000301 | Monthly (1000 calls/month per call site): $0.30

--
Total estimated monthly cost: $26.80
  1000 calls/month per call site

tokentoll diff

LLM Cost Diff: main..feature-branch
============================================================

+ ADDED    src/agents/rewriter.py:35
           openai | Model: gpt-4o
           Est. cost/call: $0.03 | Monthly: +$26.50

~ MODIFIED src/agents/summarizer.py:42
           openai | Model: gpt-4o -> gpt-4o-mini
           Est. cost/call: $0.03 -> $0.000301 | Monthly: -$26.20

--
Monthly cost impact: +$0.30
  Added: 1 | Changed: 1 | Removed: 0
  1000 calls/month per call site

仕組み

  Source Code (.py files)
         |
         v
  +-------------+     +------------------+
  | AST Scanner |---->| SDK Detectors    |
  | (ast.parse) |     | OpenAI, Anthropic|
  +-------------+     | Google, LiteLLM  |
                       | LangChain        |
                       +------------------+
                              |
                              v
                       +------------------+
                       | Pricing Engine   |
                       | 2200+ models     |
                       | Auto-cached      |
                       +------------------+
                              |
                  +-----------+-----------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Scan Report|         | Diff Engine  |
           | (costs)    |         | (old vs new) |
           +------------+         +-------------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Table/JSON |         | Table/JSON/  |
           |            |         | PR Comment   |
           +------------+         +-------------+
  1. ast モジュールを使用してPythonファイルを解析し、LLM API呼び出しを検出します

  2. マルチパス定数伝播により、変数、os.getenv() のフォールバック、クラス属性、コンストラクタ引数、辞書の内容、**kwargs のアンパックを通じてモデル名を解決します

  3. ローカルキャッシュ(LiteLLMから取得、2200以上のモデル)から価格を検索します

  4. diffモードでは、2つのgitリファレンス間の呼び出しを比較し、コストの差分を計算します

  5. コストレポートをテーブル、JSON、またはGitHub PRコメントとして出力します

CLIリファレンス

tokentoll scan [PATH...] [--format table|json|markdown] [--calls-per-month N] [--config PATH]
tokentoll diff [REF] [--base REF] [--head REF] [--format table|json|markdown|github-comment] [--config PATH]
tokentoll update    # Update bundled pricing data

MCPサーバー

tokentollにはMCP(Model Context Protocol)サーバーが含まれており、Claude Codeやその他のMCPホストがエージェントとの会話から直接LLMコード変更のコスト影響を確認できます。

インストール

pip install tokentoll[mcp]

Claude Codeへの登録

claude mcp add --transport stdio tokentoll -- tokentoll-mcp

ツール

ツール

説明

scan

ディレクトリ内のLLM API呼び出しを検索し、月額コストを試算します。パスとオプションの calls_per_month を受け取ります。

diff

2つのgitリファレンス間のLLMコストを比較します。base_ref とオプションの head_ref(デフォルトはHEAD)を受け取ります。

どちらのツールもJSON形式で出力します。

使用例

Claude Codeは、コミット前に自身の変更によるコストへの影響を確認できます。例えば、モデルを gpt-4o から gpt-4o-mini に変更した後、エージェントは HEAD に対して diff ツールを呼び出し、コミットを作成する前にコスト削減を確認できます。

価格データ

価格データはバンドルされており、オフラインで動作します。最新の価格に更新するには:

tokentoll update

価格データはLiteLLMの model_prices_and_context_window.json から取得されており、OpenAI、Anthropic、Google、AWS Bedrock、Azureなど300以上のモデルをカバーしています。

動的モデルのデフォルト値

tokentollが解決できない変数名がモデル名として使用されている呼び出しに遭遇した場合、SDKごとの適切なデフォルト値を適用してコスト試算を行います:

SDK

デフォルトモデル

OpenAI

gpt-4o

Anthropic

claude-sonnet-4-20250514

Google GenAI

gemini-2.0-flash

LiteLLM

gpt-4o

LangChain

gpt-4o

Zhipu AI

zai/glm-4.6

これらのデフォルト値は、スキャン出力で gpt-4o (default) と表示されます。.tokentoll.yml 設定ファイルを使用して、プロジェクトごとまたはパスごとに上書きできます(下記参照)。

設定

プロジェクトのルートに .tokentoll.yml を作成して動作をカスタマイズします。 tokentollは、スキャン対象ディレクトリから上位ディレクトリを遡ってこのファイルを自動的に検索します。

# Default model for all dynamic (unresolved) calls
default_model: gpt-4o

# Per-SDK defaults (override the built-in defaults above)
default_models:
  openai: gpt-4o-mini
  anthropic: claude-haiku-3-20240307

# Assumed calls per month per call site
calls_per_month: 5000

# Skip cost estimation entirely for dynamic (unresolved) models. When true,
# calls whose model name cannot be resolved statically are reported with no
# cost rather than priced against a default. Useful for projects that prefer
# silence over a guess.
skip_dynamic_models: false

# Exclude paths from scanning (prefix match or glob pattern)
exclude:
  - tests/
  - examples/
  - docs/
  - "*_test.py"

# Per-path overrides (longest prefix match)
overrides:
  - path: src/agents/
    default_model: gpt-4o
    calls_per_month: 10000
  - path: src/azure/
    skip_dynamic_models: true

動的モデルのデフォルト値の解決順序:SDKごとの設定 (default_models) > 一般設定 (default_model) > ビルトインのSDKデフォルト値。

--config path/to/.tokentoll.yml を渡して特定の設定ファイルを使用することも可能です。

トークン試算

デフォルトでは、tokentollは「文字数/4」というヒューリスティックを使用してトークン数を試算します。 より正確な試算を行うには、tiktoken をインストールしてください:

pip install tiktoken

tiktokenが利用可能な場合、tokentollは各モデルに適したトークナイザーエンコーディングを使用します。不明なモデルは cl100k_base にフォールバックします。Tiktokenは遅延読み込みされ、エンコーダーはキャッシュされるため、不要な場合に起動時のペナルティはありません。

スマート変数解決

実際のコードベースでモデル名が文字列リテラルとして渡されることは稀です。tokentollのマルチパス定数伝播エンジンは以下を追跡します:

DEFAULT_MODEL = os.getenv("MODEL", "gpt-4o")

class Config:
    model: str = DEFAULT_MODEL

config = Config()
kwargs = {"model": config.model, "max_tokens": 2000}
client.chat.completions.create(**kwargs)
# tokentoll resolves: model="gpt-4o", max_tokens=2000
  • 変数代入 (MODEL = "gpt-4o")

  • os.getenv() / os.environ.get() のフォールバック値

  • 関数のデフォルト引数

  • クラス属性のデフォルト値

  • コンストラクタ引数の伝播

  • 辞書リテラルおよび添字の内容

  • **kwargs のアンパック

ロードマップ

  • コンテキストを考慮した呼び出し頻度 (計画中): 周囲のコードから月間呼び出し回数を推論します(FastAPIのルートハンドラは高トラフィック、スクリプトは低トラフィック、ループは乗算など)。すべての呼び出し箇所で一律のボリュームを想定するのではなく、これに対応します。

  • JS/TSサポート (計画中): JavaScriptおよびTypeScriptファイル内のLLM呼び出しを検出します。

  • コストアラート: PRがコスト差分を超えた場合にCIを失敗させる設定可能な閾値。

制限事項

  • 実行時に外部設定ファイルやデータベースから読み込まれるモデルは解決できません。 これらの呼び出しにはSDKごとのデフォルト値が使用されます(.tokentoll.yml で設定可能)。

  • tiktoken がインストールされていない限り、トークン試算には「文字数/4」のヒューリスティックが使用されます。

  • 月間試算は、呼び出し箇所ごとに一律のボリュームを想定しています(--calls-per-month、.tokentoll.yml、またはパスごとの上書きで設定可能)。テストファイルやサンプルファイルを除外するには exclude オプションを使用してください。

  • 現在はPythonのみ対応(JS/TSサポートは計画中)。

ライセンス

MIT

Available Tools

2 tools
diffA

Compare LLM costs between two git refs.

Shows which LLM call sites were added, removed, or changed between the base and head refs, along with the cost impact of those changes.

Args: base_ref: The base git ref (branch, tag, or commit) to compare from. head_ref: The head git ref to compare to. Defaults to HEAD.

Returns: JSON string with the diff results including cost changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
base_refYes
head_refNoHEAD

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It indicates a read-like operation (diff) and describes the output, but does not explicitly state side effects or permissions. Adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with purpose, and includes parameter docs and return type. Every sentence adds value without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema (not shown), the description adequately covers purpose, parameters, and output format. It could include examples or edge cases but is sufficiently complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description documents both parameters: base_ref as the base git ref and head_ref as the head ref defaulting to HEAD. This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares LLM costs between two git refs, specifying it shows added, removed, or changed call sites and cost impact. This distinguishes it from the sibling 'scan' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (to compare costs between refs) but does not explicitly state when not to use it or mention alternatives. Usage is well implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scanA

Scan a directory for LLM API calls and estimate monthly costs.

Finds all LLM API call sites (OpenAI, Anthropic, etc.) in the given path and produces a cost estimate based on token counts and pricing.

Args: path: Directory or file path to scan. Defaults to current directory. calls_per_month: Assumed monthly call volume per call site. If not provided, the CLI default (1000) is used.

Returns: JSON string with the scan results including call sites and cost estimates.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
calls_per_monthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It details the scanning action, cost estimation, and return format. While it doesn't cover every edge case (e.g., recursion depth or error handling), it provides sufficient behavioral insight for a read-only analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a lead sentence, then details in Args and Returns sections. Every sentence adds value, and the format is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 optional params, no annotations), the description covers the core behavior and return type adequately. It could mention recursion or failure modes, but it is sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the description fully explains both parameters: 'path' (directory/file, default current dir) and 'calls_per_month' (monthly volume, default null implying CLI default of 1000). This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans a directory for LLM API calls and estimates costs, specifying providers and purpose. This is a specific verb+resource that distinguishes it from the sibling 'diff'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use the tool (scanning directories for LLM calls and cost estimation). However, it does not explicitly mention when not to use it or provide alternatives, which prevents a top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observeddiff
    • First observedscan

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools, diff and scan, have clearly distinct purposes: scan finds LLM call sites and estimates costs, while diff compares costs between git refs. No overlap or ambiguity.

Naming Consistency4/5

Both tool names are single verbs ('diff', 'scan'), which is consistent in style. While not a verb_noun pattern, the naming is uniform and intuitive for the domain.

Tool Count3/5

With only 2 tools, the server is very focused. This can be appropriate for a narrow utility, but it feels thin for a full server. A few more tools (e.g., pricing config) might improve scope.

Completeness3/5

The tools cover two core operations: scanning and diffing. However, there is no tool for managing pricing configurations or listing assumptions, which could be gaps for advanced use.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI cost calculation, comparison, and optimization across major providers like Anthropic, OpenAI, Google, Meta, and Mistral. Supports cost estimation, budget-aware model finding, and token estimation through a simple API and MCP integration.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Exposes boyter/scc code counting and complexity analysis to LLM agents via read-only tools like counting lines, finding top files, and cost estimation.
    7
    BSD 3-Clause