tokentoll
tokentoll
Отслеживайте изменения стоимости LLM при проверке кода. Infracost для расходов на LLM.
Инструмент CLI и GitHub Action, который статически анализирует ваш код на наличие вызовов API LLM, оценивает их стоимость и показывает влияние каждого изменения на расходы в вашем терминале или в виде комментария к PR. Нулевые зависимости во время выполнения.
Проблема
Простая замена модели с gpt-4o-mini на gpt-4o увеличивает расходы в 15 раз.
Новый вызов API в критическом пути может добавить $10 000/мес к вашему счету.
Эти изменения скрыты при обычной проверке кода.
tokentoll находит вызовы API LLM в вашем коде, оценивает их стоимость и показывает влияние каждого изменения на расходы до того, как оно попадет в продакшн.
Related MCP server: CosTrack MCP
Быстрый старт
pip install tokentoll
# Scan current directory for LLM API calls and their costs
tokentoll scan .
# Show cost impact of your last commit
tokentoll diff HEAD~1
# Compare two branches
tokentoll diff main..feature-branchGitHub Action
name: LLM Cost Diff
on:
pull_request:
paths:
- "**.py"
permissions:
pull-requests: write
jobs:
cost-diff:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: Jwrede/tokentoll@v0.6.1Что он обнаруживает
SDK | Шаблоны | Статус |
OpenAI |
| Поддерживается |
Anthropic |
| Поддерживается |
Google GenAI |
| Поддерживается |
LiteLLM |
| Поддерживается |
LangChain |
| Поддерживается |
Zhipu AI |
| Поддерживается |
JS/TS SDKs | Запланировано |
Пример вывода
tokentoll scan
LLM API Calls Detected
============================================================
File: src/agents/summarizer.py
Line 42: openai client.chat.completions.create
Model: gpt-4o | Max tokens: 4096
Est. cost/call: $0.03 | Monthly (1000 calls/month per call site): $26.50
Line 78: openai client.chat.completions.create
Model: gpt-4o-mini | Max tokens: 1000
Est. cost/call: $0.000301 | Monthly (1000 calls/month per call site): $0.30
--
Total estimated monthly cost: $26.80
1000 calls/month per call sitetokentoll diff
LLM Cost Diff: main..feature-branch
============================================================
+ ADDED src/agents/rewriter.py:35
openai | Model: gpt-4o
Est. cost/call: $0.03 | Monthly: +$26.50
~ MODIFIED src/agents/summarizer.py:42
openai | Model: gpt-4o -> gpt-4o-mini
Est. cost/call: $0.03 -> $0.000301 | Monthly: -$26.20
--
Monthly cost impact: +$0.30
Added: 1 | Changed: 1 | Removed: 0
1000 calls/month per call siteКак это работает
Source Code (.py files)
|
v
+-------------+ +------------------+
| AST Scanner |---->| SDK Detectors |
| (ast.parse) | | OpenAI, Anthropic|
+-------------+ | Google, LiteLLM |
| LangChain |
+------------------+
|
v
+------------------+
| Pricing Engine |
| 2200+ models |
| Auto-cached |
+------------------+
|
+-----------+-----------+
| |
v v
+------------+ +-------------+
| Scan Report| | Diff Engine |
| (costs) | | (old vs new) |
+------------+ +-------------+
| |
v v
+------------+ +-------------+
| Table/JSON | | Table/JSON/ |
| | | PR Comment |
+------------+ +-------------+Анализирует файлы Python с помощью модуля
astдля поиска вызовов API LLMМногопроходное распространение констант разрешает имена моделей через переменные, резервные значения
os.getenv(), атрибуты классов, аргументы конструктора, содержимое словарей и распаковку**kwargsИщет цены в локальном кэше (полученном из LiteLLM, 2200+ моделей)
Для режима diff: сравнивает вызовы между двумя ссылками git и вычисляет дельту стоимости
Выводит отчет о стоимости в виде таблицы, JSON или комментария к PR на GitHub
Справочник CLI
tokentoll scan [PATH...] [--format table|json|markdown] [--calls-per-month N] [--config PATH]
tokentoll diff [REF] [--base REF] [--head REF] [--format table|json|markdown|github-comment] [--config PATH]
tokentoll update # Update bundled pricing dataСервер MCP
tokentoll включает сервер MCP (Model Context Protocol), который позволяет Claude Code и другим хостам MCP проверять влияние изменений кода LLM на стоимость непосредственно из диалога с агентом.
Установка
pip install tokentoll[mcp]Регистрация в Claude Code
claude mcp add --transport stdio tokentoll -- tokentoll-mcpИнструменты
Инструмент | Описание |
| Поиск вызовов API LLM в каталоге и оценка ежемесячных расходов. Принимает путь и необязательный параметр |
| Сравнение затрат на LLM между двумя ссылками git. Принимает |
Оба инструмента возвращают вывод в формате JSON.
Пример использования
Claude Code может проверить влияние своих собственных изменений на стоимость перед фиксацией. Например, после замены модели с gpt-4o на gpt-4o-mini агент может вызвать инструмент diff для HEAD, чтобы подтвердить снижение затрат перед созданием коммита.
Данные о ценах
Данные о ценах включены в комплект и работают в автономном режиме. Чтобы обновить цены до последних:
tokentoll updateДанные о ценах получены из файла LiteLLM model_prices_and_context_window.json и охватывают более 300 моделей OpenAI, Anthropic, Google, AWS Bedrock, Azure и других.
Динамические значения моделей по умолчанию
Когда tokentoll встречает вызов, где имя модели является переменной, которую он не может разрешить, он применяет разумное значение по умолчанию для каждого SDK, чтобы вы все равно получили оценки стоимости:
SDK | Модель по умолчанию |
OpenAI |
|
Anthropic |
|
Google GenAI |
|
LiteLLM |
|
LangChain |
|
Zhipu AI |
|
Эти значения по умолчанию отображаются как gpt-4o (default) в выводе сканирования. Вы можете переопределить их для каждого проекта или пути с помощью файла конфигурации .tokentoll.yml (см. ниже).
Конфигурация
Создайте .tokentoll.yml в корне вашего проекта, чтобы настроить поведение. tokentoll автоматически находит этот файл, поднимаясь вверх от сканируемого каталога.
# Default model for all dynamic (unresolved) calls
default_model: gpt-4o
# Per-SDK defaults (override the built-in defaults above)
default_models:
openai: gpt-4o-mini
anthropic: claude-haiku-3-20240307
# Assumed calls per month per call site
calls_per_month: 5000
# Skip cost estimation entirely for dynamic (unresolved) models. When true,
# calls whose model name cannot be resolved statically are reported with no
# cost rather than priced against a default. Useful for projects that prefer
# silence over a guess.
skip_dynamic_models: false
# Exclude paths from scanning (prefix match or glob pattern)
exclude:
- tests/
- examples/
- docs/
- "*_test.py"
# Per-path overrides (longest prefix match)
overrides:
- path: src/agents/
default_model: gpt-4o
calls_per_month: 10000
- path: src/azure/
skip_dynamic_models: trueПорядок разрешения для динамических моделей по умолчанию: конфигурация для конкретного SDK (default_models) > общая конфигурация (default_model) > встроенные значения по умолчанию SDK.
Вы также можете передать --config path/to/.tokentoll.yml, чтобы использовать определенный файл конфигурации.
Оценка токенов
По умолчанию tokentoll оценивает количество токенов, используя эвристику символы/4. Для более точных оценок установите tiktoken:
pip install tiktokenКогда tiktoken доступен, tokentoll использует правильную кодировку токенизатора для каждой модели. Неизвестные модели возвращаются к cl100k_base. Tiktoken загружается лениво, а кодировщики кэшируются, поэтому нет штрафа за запуск, если он вам не нужен.
Умное разрешение переменных
Реальные кодовые базы редко передают имена моделей в виде строковых литералов. Многопроходный движок распространения констант tokentoll отслеживает:
DEFAULT_MODEL = os.getenv("MODEL", "gpt-4o")
class Config:
model: str = DEFAULT_MODEL
config = Config()
kwargs = {"model": config.model, "max_tokens": 2000}
client.chat.completions.create(**kwargs)
# tokentoll resolves: model="gpt-4o", max_tokens=2000Присваивания переменных (
MODEL = "gpt-4o")Резервные значения
os.getenv()/os.environ.get()Параметры функций по умолчанию
Значения атрибутов классов по умолчанию
Распространение аргументов конструктора
Содержимое литералов словарей и индексов
Распаковку
**kwargs
Дорожная карта
Частота вызовов с учетом контекста (запланировано): определение вызовов/мес из окружающего кода (обработчики маршрутов FastAPI = высокий трафик, скрипты = низкий, циклы = умножено) вместо предположения об одинаковом объеме для всех мест вызова.
Поддержка JS/TS (запланировано): обнаружение вызовов LLM в файлах JavaScript и TypeScript.
Оповещения о стоимости: настраиваемые пороги, которые вызывают сбой CI, когда PR превышает дельту стоимости.
Ограничения
Невозможно разрешить модели, загружаемые из внешних файлов конфигурации или баз данных во время выполнения. Эти вызовы используют значения по умолчанию для каждого SDK (настраиваются через
.tokentoll.yml).Оценки токенов используют эвристику символы/4, если не установлен tiktoken.
Ежемесячные оценки предполагают одинаковый объем вызовов для каждого места вызова (настраивается через
--calls-per-month,.tokentoll.ymlили переопределения для каждого пути). Используйте опциюexclude, чтобы пропустить тестовые и примерные файлы.Пока только Python (поддержка JS/TS запланирована).
Лицензия
MIT
Available Tools
2 toolsdiffA
Compare LLM costs between two git refs.
Shows which LLM call sites were added, removed, or changed between the base and head refs, along with the cost impact of those changes.
Args: base_ref: The base git ref (branch, tag, or commit) to compare from. head_ref: The head git ref to compare to. Defaults to HEAD.
Returns: JSON string with the diff results including cost changes.
| Name | Required | Description | Default |
|---|---|---|---|
| base_ref | Yes | ||
| head_ref | No | HEAD |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It indicates a read-like operation (diff) and describes the output, but does not explicitly state side effects or permissions. Adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with purpose, and includes parameter docs and return type. Every sentence adds value without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of an output schema (not shown), the description adequately covers purpose, parameters, and output format. It could include examples or edge cases but is sufficiently complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description documents both parameters: base_ref as the base git ref and head_ref as the head ref defaulting to HEAD. This adds essential meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares LLM costs between two git refs, specifying it shows added, removed, or changed call sites and cost impact. This distinguishes it from the sibling 'scan' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (to compare costs between refs) but does not explicitly state when not to use it or mention alternatives. Usage is well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scanA
Scan a directory for LLM API calls and estimate monthly costs.
Finds all LLM API call sites (OpenAI, Anthropic, etc.) in the given path and produces a cost estimate based on token counts and pricing.
Args: path: Directory or file path to scan. Defaults to current directory. calls_per_month: Assumed monthly call volume per call site. If not provided, the CLI default (1000) is used.
Returns: JSON string with the scan results including call sites and cost estimates.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . | |
| calls_per_month | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It details the scanning action, cost estimation, and return format. While it doesn't cover every edge case (e.g., recursion depth or error handling), it provides sufficient behavioral insight for a read-only analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a lead sentence, then details in Args and Returns sections. Every sentence adds value, and the format is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 optional params, no annotations), the description covers the core behavior and return type adequately. It could mention recursion or failure modes, but it is sufficient for most use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description fully explains both parameters: 'path' (directory/file, default current dir) and 'calls_per_month' (monthly volume, default null implying CLI default of 1000). This adds essential meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans a directory for LLM API calls and estimates costs, specifying providers and purpose. This is a specific verb+resource that distinguishes it from the sibling 'diff'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use the tool (scanning directories for LLM calls and cost estimation). However, it does not explicitly mention when not to use it or provide alternatives, which prevents a top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
diff - First observed
scan
TDQS
Scored across 2 tools
The two tools, diff and scan, have clearly distinct purposes: scan finds LLM call sites and estimates costs, while diff compares costs between git refs. No overlap or ambiguity.
Both tool names are single verbs ('diff', 'scan'), which is consistent in style. While not a verb_noun pattern, the naming is uniform and intuitive for the domain.
With only 2 tools, the server is very focused. This can be appropriate for a narrow utility, but it feels thin for a full server. A few more tools (e.g., pricing config) might improve scope.
The tools cover two core operations: scanning and diffing. However, there is no tool for managing pricing configurations or listing assumptions, which could be gaps for advanced use.
Maintenance
Related MCP Connectors
Exact Claude API cost calc with real cache economics, plus a tiktoken-misuse scanner.
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
AI-powered codebase analysis — call graphs, security, dead code, complexity. 150+ tools.
Codebase intelligence for AI agents — dead code, blast radius, ownership.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI cost calculation, comparison, and optimization across major providers like Anthropic, OpenAI, Google, Meta, and Mistral. Supports cost estimation, budget-aware model finding, and token estimation through a simple API and MCP integration.-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to track LLM costs, enforce budgets, compare models, and estimate expenses through simple tool calls.-
- AlicenseBqualityDmaintenancePredict the cost of an LLM call before you make it, and pick the cheapest model that still does the job, offline, from your editor.732 npmApache 2.0
- AlicenseAqualityDmaintenanceExposes boyter/scc code counting and complexity analysis to LLM agents via read-only tools like counting lines, finding top files, and cost estimation.7BSD 3-Clause