Skip to main content
Glama

tokentoll

Detecte cambios en los costes de LLM durante la revisión de código. Infracost para el gasto en LLM.

CI Versión de PyPI GitHub Marketplace Licencia: MIT Python 3.10+

Una herramienta CLI y una GitHub Action que analiza estáticamente su código en busca de llamadas a la API de LLM, estima su coste y le muestra el impacto en el coste de cada cambio en su terminal o como un comentario en una PR. Sin dependencias en tiempo de ejecución.

El problema

Un simple cambio de modelo de gpt-4o-mini a gpt-4o aumenta los costes 15 veces. Una nueva llamada a la API en una ruta crítica puede añadir 10.000 $/mes a su factura. Estos cambios se ocultan en la revisión de código normal.

tokentoll encuentra llamadas a la API de LLM en su código, estima su coste y le muestra el impacto en el coste de cada cambio antes de que llegue a producción.

Related MCP server: CosTrack MCP

Inicio rápido

pip install tokentoll

# Scan current directory for LLM API calls and their costs
tokentoll scan .

# Show cost impact of your last commit
tokentoll diff HEAD~1

# Compare two branches
tokentoll diff main..feature-branch

GitHub Action

name: LLM Cost Diff
on:
  pull_request:
    paths:
      - "**.py"

permissions:
  pull-requests: write

jobs:
  cost-diff:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: Jwrede/tokentoll@v0.6.1

Qué detecta

SDK

Patrones

Estado

OpenAI

chat.completions.create, responses.create

Soportado

Anthropic

messages.create, messages.stream

Soportado

Google GenAI

models.generate_content

Soportado

LiteLLM

completion, acompletion

Soportado

LangChain

ChatOpenAI, ChatAnthropic, init_chat_model

Soportado

Zhipu AI

ZhipuAiClient, ZhipuAI (modelos GLM)

Soportado

SDKs JS/TS

Planificado

Ejemplo de salida

tokentoll scan

LLM API Calls Detected
============================================================

File: src/agents/summarizer.py
  Line 42: openai client.chat.completions.create
           Model: gpt-4o | Max tokens: 4096
           Est. cost/call: $0.03 | Monthly (1000 calls/month per call site): $26.50

  Line 78: openai client.chat.completions.create
           Model: gpt-4o-mini | Max tokens: 1000
           Est. cost/call: $0.000301 | Monthly (1000 calls/month per call site): $0.30

--
Total estimated monthly cost: $26.80
  1000 calls/month per call site

tokentoll diff

LLM Cost Diff: main..feature-branch
============================================================

+ ADDED    src/agents/rewriter.py:35
           openai | Model: gpt-4o
           Est. cost/call: $0.03 | Monthly: +$26.50

~ MODIFIED src/agents/summarizer.py:42
           openai | Model: gpt-4o -> gpt-4o-mini
           Est. cost/call: $0.03 -> $0.000301 | Monthly: -$26.20

--
Monthly cost impact: +$0.30
  Added: 1 | Changed: 1 | Removed: 0
  1000 calls/month per call site

Cómo funciona

  Source Code (.py files)
         |
         v
  +-------------+     +------------------+
  | AST Scanner |---->| SDK Detectors    |
  | (ast.parse) |     | OpenAI, Anthropic|
  +-------------+     | Google, LiteLLM  |
                       | LangChain        |
                       +------------------+
                              |
                              v
                       +------------------+
                       | Pricing Engine   |
                       | 2200+ models     |
                       | Auto-cached      |
                       +------------------+
                              |
                  +-----------+-----------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Scan Report|         | Diff Engine  |
           | (costs)    |         | (old vs new) |
           +------------+         +-------------+
                  |                       |
                  v                       v
           +------------+         +-------------+
           | Table/JSON |         | Table/JSON/  |
           |            |         | PR Comment   |
           +------------+         +-------------+
  1. Analiza archivos Python usando el módulo ast para encontrar llamadas a la API de LLM

  2. La propagación de constantes de múltiples pasadas resuelve los nombres de los modelos a través de variables, valores de respaldo de os.getenv(), atributos de clase, argumentos de constructor, contenidos de diccionarios y desempaquetado de **kwargs

  3. Busca precios en una caché local (obtenida de LiteLLM, más de 2200 modelos)

  4. Para el modo diff: compara llamadas entre dos referencias de git y calcula el delta de coste

  5. Genera un informe de costes como tabla, JSON o comentario de PR de GitHub

Referencia de CLI

tokentoll scan [PATH...] [--format table|json|markdown] [--calls-per-month N] [--config PATH]
tokentoll diff [REF] [--base REF] [--head REF] [--format table|json|markdown|github-comment] [--config PATH]
tokentoll update    # Update bundled pricing data

Servidor MCP

tokentoll incluye un servidor MCP (Model Context Protocol) que permite a Claude Code y otros hosts MCP comprobar el impacto en el coste de los cambios de código LLM directamente desde una conversación con el agente.

Instalación

pip install tokentoll[mcp]

Registrar con Claude Code

claude mcp add --transport stdio tokentoll -- tokentoll-mcp

Herramientas

Herramienta

Descripción

scan

Encuentra llamadas a la API de LLM en un directorio y estima los costes mensuales. Acepta una ruta y calls_per_month opcional.

diff

Compara los costes de LLM entre dos referencias de git. Acepta base_ref y head_ref opcional (por defecto es HEAD).

Ambas herramientas devuelven una salida JSON.

Caso de uso de ejemplo

Claude Code puede comprobar el impacto en el coste de sus propios cambios antes de realizar el commit. Por ejemplo, después de cambiar un modelo de gpt-4o a gpt-4o-mini, el agente puede llamar a la herramienta diff contra HEAD para verificar la reducción de costes antes de crear el commit.

Datos de precios

Los precios están incluidos y funcionan sin conexión. Para actualizar a los precios más recientes:

tokentoll update

Los datos de precios provienen de model_prices_and_context_window.json de LiteLLM y cubren más de 300 modelos en OpenAI, Anthropic, Google, AWS Bedrock, Azure y más.

Valores predeterminados de modelo dinámicos

Cuando tokentoll encuentra una llamada donde el nombre del modelo es una variable que no puede resolver, aplica un valor predeterminado sensato por SDK para que aún obtenga estimaciones de costes:

SDK

Modelo predeterminado

OpenAI

gpt-4o

Anthropic

claude-sonnet-4-20250514

Google GenAI

gemini-2.0-flash

LiteLLM

gpt-4o

LangChain

gpt-4o

Zhipu AI

zai/glm-4.6

Estos valores predeterminados se muestran como gpt-4o (default) en la salida del escaneo. Puede sobrescribirlos por proyecto o por ruta usando un archivo de configuración .tokentoll.yml (ver más abajo).

Configuración

Cree un .tokentoll.yml en la raíz de su proyecto para personalizar el comportamiento. tokentoll encuentra automáticamente este archivo subiendo desde el directorio escaneado.

# Default model for all dynamic (unresolved) calls
default_model: gpt-4o

# Per-SDK defaults (override the built-in defaults above)
default_models:
  openai: gpt-4o-mini
  anthropic: claude-haiku-3-20240307

# Assumed calls per month per call site
calls_per_month: 5000

# Skip cost estimation entirely for dynamic (unresolved) models. When true,
# calls whose model name cannot be resolved statically are reported with no
# cost rather than priced against a default. Useful for projects that prefer
# silence over a guess.
skip_dynamic_models: false

# Exclude paths from scanning (prefix match or glob pattern)
exclude:
  - tests/
  - examples/
  - docs/
  - "*_test.py"

# Per-path overrides (longest prefix match)
overrides:
  - path: src/agents/
    default_model: gpt-4o
    calls_per_month: 10000
  - path: src/azure/
    skip_dynamic_models: true

El orden de resolución para los modelos predeterminados dinámicos es: configuración por SDK (default_models) > configuración genérica (default_model) > valores predeterminados integrados del SDK.

También puede pasar --config path/to/.tokentoll.yml para usar un archivo de configuración específico.

Estimación de tokens

Por defecto, tokentoll estima los recuentos de tokens usando una heurística de caracteres/4. Para estimaciones más precisas, instale tiktoken:

pip install tiktoken

Cuando tiktoken está disponible, tokentoll utiliza la codificación de tokenizador correcta para cada modelo. Los modelos desconocidos vuelven a cl100k_base. Tiktoken se carga de forma diferida y los codificadores se almacenan en caché, por lo que no hay penalización de inicio si no lo necesita.

Resolución inteligente de variables

Las bases de código reales rara vez pasan nombres de modelos como literales de cadena. El motor de propagación de constantes de múltiples pasadas de tokentoll sigue:

DEFAULT_MODEL = os.getenv("MODEL", "gpt-4o")

class Config:
    model: str = DEFAULT_MODEL

config = Config()
kwargs = {"model": config.model, "max_tokens": 2000}
client.chat.completions.create(**kwargs)
# tokentoll resolves: model="gpt-4o", max_tokens=2000
  • Asignaciones de variables (MODEL = "gpt-4o")

  • Valores de respaldo de os.getenv() / os.environ.get()

  • Parámetros predeterminados de funciones

  • Valores predeterminados de atributos de clase

  • Propagación de argumentos de constructor

  • Contenidos de literales de diccionario y subíndices

  • Desempaquetado de **kwargs

Hoja de ruta

  • Frecuencia de llamadas consciente del contexto (planificado): inferir llamadas/mes a partir del código circundante (manejadores de rutas de FastAPI = tráfico alto, scripts = bajo, bucles = multiplicado) en lugar de asumir un volumen uniforme en todos los sitios de llamada.

  • Soporte JS/TS (planificado): detectar llamadas a LLM en archivos JavaScript y TypeScript.

  • Alertas de costes: umbrales configurables que fallan en la CI cuando una PR excede un delta de coste.

Limitaciones

  • No puede resolver modelos cargados desde archivos de configuración externos o bases de datos en tiempo de ejecución. Estas llamadas utilizan valores predeterminados por SDK (configurables mediante .tokentoll.yml).

  • Las estimaciones de tokens utilizan una heurística de caracteres/4 a menos que esté instalado tiktoken.

  • Las estimaciones mensuales asumen un volumen de llamadas uniforme por sitio de llamada (configurable mediante --calls-per-month, .tokentoll.yml o sobrescrituras por ruta). Use la opción exclude para omitir archivos de prueba y ejemplo.

  • Solo Python por ahora (soporte JS/TS planificado).

Licencia

MIT

Available Tools

2 tools
diffA

Compare LLM costs between two git refs.

Shows which LLM call sites were added, removed, or changed between the base and head refs, along with the cost impact of those changes.

Args: base_ref: The base git ref (branch, tag, or commit) to compare from. head_ref: The head git ref to compare to. Defaults to HEAD.

Returns: JSON string with the diff results including cost changes.

ParametersJSON Schema
NameRequiredDescriptionDefault
base_refYes
head_refNoHEAD

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It indicates a read-like operation (diff) and describes the output, but does not explicitly state side effects or permissions. Adequate but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with purpose, and includes parameter docs and return type. Every sentence adds value without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema (not shown), the description adequately covers purpose, parameters, and output format. It could include examples or edge cases but is sufficiently complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description documents both parameters: base_ref as the base git ref and head_ref as the head ref defaulting to HEAD. This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares LLM costs between two git refs, specifying it shows added, removed, or changed call sites and cost impact. This distinguishes it from the sibling 'scan' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool (to compare costs between refs) but does not explicitly state when not to use it or mention alternatives. Usage is well implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scanA

Scan a directory for LLM API calls and estimate monthly costs.

Finds all LLM API call sites (OpenAI, Anthropic, etc.) in the given path and produces a cost estimate based on token counts and pricing.

Args: path: Directory or file path to scan. Defaults to current directory. calls_per_month: Assumed monthly call volume per call site. If not provided, the CLI default (1000) is used.

Returns: JSON string with the scan results including call sites and cost estimates.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
calls_per_monthNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It details the scanning action, cost estimation, and return format. While it doesn't cover every edge case (e.g., recursion depth or error handling), it provides sufficient behavioral insight for a read-only analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: a lead sentence, then details in Args and Returns sections. Every sentence adds value, and the format is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 optional params, no annotations), the description covers the core behavior and return type adequately. It could mention recursion or failure modes, but it is sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but the description fully explains both parameters: 'path' (directory/file, default current dir) and 'calls_per_month' (monthly volume, default null implying CLI default of 1000). This adds essential meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scans a directory for LLM API calls and estimates costs, specifying providers and purpose. This is a specific verb+resource that distinguishes it from the sibling 'diff'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use the tool (scanning directories for LLM calls and cost estimation). However, it does not explicitly mention when not to use it or provide alternatives, which prevents a top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observeddiff
    • First observedscan

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation5/5

The two tools, diff and scan, have clearly distinct purposes: scan finds LLM call sites and estimates costs, while diff compares costs between git refs. No overlap or ambiguity.

Naming Consistency4/5

Both tool names are single verbs ('diff', 'scan'), which is consistent in style. While not a verb_noun pattern, the naming is uniform and intuitive for the domain.

Tool Count3/5

With only 2 tools, the server is very focused. This can be appropriate for a narrow utility, but it feels thin for a full server. A few more tools (e.g., pricing config) might improve scope.

Completeness3/5

The tools cover two core operations: scanning and diffing. However, there is no tool for managing pricing configurations or listing assumptions, which could be gaps for advanced use.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI cost calculation, comparison, and optimization across major providers like Anthropic, OpenAI, Google, Meta, and Mistral. Supports cost estimation, budget-aware model finding, and token estimation through a simple API and MCP integration.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Exposes boyter/scc code counting and complexity analysis to LLM agents via read-only tools like counting lines, finding top files, and cost estimation.
    7
    BSD 3-Clause