tokentoll
tokentoll
코드 리뷰에서 LLM 비용 변경을 포착하세요. LLM 지출을 위한 Infracost입니다.
LLM API 호출을 위해 코드를 정적으로 분석하고, 비용을 추정하며, 터미널이나 PR 코멘트를 통해 모든 변경 사항의 비용 영향을 보여주는 CLI 도구이자 GitHub Action입니다. 런타임 의존성이 없습니다.
문제점
gpt-4o-mini에서 gpt-4o로 모델을 한 번만 교체해도 비용이 15배 증가합니다. 핫 패스(hot path)에 새로운 API 호출이 추가되면 청구서에 월 $10,000이 추가될 수 있습니다. 이러한 변경 사항은 일반적인 코드 리뷰에서는 숨겨져 있습니다.
tokentoll은 코드에서 LLM API 호출을 찾아 비용을 추정하고, 프로덕션에 반영되기 전에 모든 변경 사항의 비용 영향을 보여줍니다.
Related MCP server: CosTrack MCP
빠른 시작
pip install tokentoll
# Scan current directory for LLM API calls and their costs
tokentoll scan .
# Show cost impact of your last commit
tokentoll diff HEAD~1
# Compare two branches
tokentoll diff main..feature-branchGitHub Action
name: LLM Cost Diff
on:
pull_request:
paths:
- "**.py"
permissions:
pull-requests: write
jobs:
cost-diff:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: Jwrede/tokentoll@v0.6.1감지 대상
SDK | 패턴 | 상태 |
OpenAI |
| 지원됨 |
Anthropic |
| 지원됨 |
Google GenAI |
| 지원됨 |
LiteLLM |
| 지원됨 |
LangChain |
| 지원됨 |
Zhipu AI |
| 지원됨 |
JS/TS SDKs | 계획됨 |
예시 출력
tokentoll scan
LLM API Calls Detected
============================================================
File: src/agents/summarizer.py
Line 42: openai client.chat.completions.create
Model: gpt-4o | Max tokens: 4096
Est. cost/call: $0.03 | Monthly (1000 calls/month per call site): $26.50
Line 78: openai client.chat.completions.create
Model: gpt-4o-mini | Max tokens: 1000
Est. cost/call: $0.000301 | Monthly (1000 calls/month per call site): $0.30
--
Total estimated monthly cost: $26.80
1000 calls/month per call sitetokentoll diff
LLM Cost Diff: main..feature-branch
============================================================
+ ADDED src/agents/rewriter.py:35
openai | Model: gpt-4o
Est. cost/call: $0.03 | Monthly: +$26.50
~ MODIFIED src/agents/summarizer.py:42
openai | Model: gpt-4o -> gpt-4o-mini
Est. cost/call: $0.03 -> $0.000301 | Monthly: -$26.20
--
Monthly cost impact: +$0.30
Added: 1 | Changed: 1 | Removed: 0
1000 calls/month per call site작동 원리
Source Code (.py files)
|
v
+-------------+ +------------------+
| AST Scanner |---->| SDK Detectors |
| (ast.parse) | | OpenAI, Anthropic|
+-------------+ | Google, LiteLLM |
| LangChain |
+------------------+
|
v
+------------------+
| Pricing Engine |
| 2200+ models |
| Auto-cached |
+------------------+
|
+-----------+-----------+
| |
v v
+------------+ +-------------+
| Scan Report| | Diff Engine |
| (costs) | | (old vs new) |
+------------+ +-------------+
| |
v v
+------------+ +-------------+
| Table/JSON | | Table/JSON/ |
| | | PR Comment |
+------------+ +-------------+ast모듈을 사용하여 Python 파일을 파싱하고 LLM API 호출을 찾습니다.다중 패스 상수 전파(Multi-pass constant propagation)를 통해 변수,
os.getenv()폴백, 클래스 속성, 생성자 인수, 딕셔너리 내용 및**kwargs언패킹을 통해 모델 이름을 확인합니다.로컬 캐시(LiteLLM에서 제공, 2200개 이상의 모델)에서 가격을 조회합니다.
diff 모드: 두 git 참조 간의 호출을 비교하고 비용 차이를 계산합니다.
비용 보고서를 표, JSON 또는 GitHub PR 코멘트로 출력합니다.
CLI 참조
tokentoll scan [PATH...] [--format table|json|markdown] [--calls-per-month N] [--config PATH]
tokentoll diff [REF] [--base REF] [--head REF] [--format table|json|markdown|github-comment] [--config PATH]
tokentoll update # Update bundled pricing dataMCP 서버
tokentoll에는 Claude Code 및 기타 MCP 호스트가 에이전트 대화에서 직접 LLM 코드 변경의 비용 영향을 확인할 수 있도록 하는 MCP(Model Context Protocol) 서버가 포함되어 있습니다.
설치
pip install tokentoll[mcp]Claude Code에 등록
claude mcp add --transport stdio tokentoll -- tokentoll-mcp도구
도구 | 설명 |
| 디렉토리에서 LLM API 호출을 찾고 월별 비용을 추정합니다. 경로와 선택적 |
| 두 git 참조 간의 LLM 비용을 비교합니다. |
두 도구 모두 JSON 출력을 반환합니다.
사용 사례 예시
Claude Code는 커밋하기 전에 자체 변경 사항의 비용 영향을 확인할 수 있습니다. 예를 들어, 모델을 gpt-4o에서 gpt-4o-mini로 교체한 후, 에이전트는 HEAD를 대상으로 diff 도구를 호출하여 커밋을 생성하기 전에 비용 절감을 확인할 수 있습니다.
가격 데이터
가격은 번들로 제공되며 오프라인에서 작동합니다. 최신 가격으로 업데이트하려면:
tokentoll update가격 데이터는 LiteLLM의 model_prices_and_context_window.json에서 가져오며 OpenAI, Anthropic, Google, AWS Bedrock, Azure 등 300개 이상의 모델을 다룹니다.
동적 모델 기본값
tokentoll은 모델 이름이 확인할 수 없는 변수인 호출을 발견하면, SDK별 기본값을 적용하여 비용 추정치를 제공합니다:
SDK | 기본 모델 |
OpenAI |
|
Anthropic |
|
Google GenAI |
|
LiteLLM |
|
LangChain |
|
Zhipu AI |
|
이러한 기본값은 스캔 출력에서 gpt-4o (default)로 표시됩니다. .tokentoll.yml 구성 파일(아래 참조)을 사용하여 프로젝트별 또는 경로별로 재정의할 수 있습니다.
구성
프로젝트 루트에 .tokentoll.yml을 생성하여 동작을 사용자 정의하세요. tokentoll은 스캔된 디렉토리에서 상위로 이동하며 이 파일을 자동으로 찾습니다.
# Default model for all dynamic (unresolved) calls
default_model: gpt-4o
# Per-SDK defaults (override the built-in defaults above)
default_models:
openai: gpt-4o-mini
anthropic: claude-haiku-3-20240307
# Assumed calls per month per call site
calls_per_month: 5000
# Skip cost estimation entirely for dynamic (unresolved) models. When true,
# calls whose model name cannot be resolved statically are reported with no
# cost rather than priced against a default. Useful for projects that prefer
# silence over a guess.
skip_dynamic_models: false
# Exclude paths from scanning (prefix match or glob pattern)
exclude:
- tests/
- examples/
- docs/
- "*_test.py"
# Per-path overrides (longest prefix match)
overrides:
- path: src/agents/
default_model: gpt-4o
calls_per_month: 10000
- path: src/azure/
skip_dynamic_models: true동적 모델 기본값에 대한 해결 순서: SDK별 구성(default_models) > 일반 구성(default_model) > 내장 SDK 기본값.
--config path/to/.tokentoll.yml을 전달하여 특정 구성 파일을 사용할 수도 있습니다.
토큰 추정
기본적으로 tokentoll은 문자 수/4 휴리스틱을 사용하여 토큰 수를 추정합니다. 더 정확한 추정을 원하시면 tiktoken을 설치하세요:
pip install tiktokentiktoken을 사용할 수 있는 경우, tokentoll은 각 모델에 맞는 올바른 토크나이저 인코딩을 사용합니다. 알 수 없는 모델은 cl100k_base로 폴백됩니다. Tiktoken은 지연 로딩되며 인코더는 캐시되므로 필요하지 않은 경우 시작 성능 저하가 없습니다.
스마트 변수 확인
실제 코드베이스에서 모델 이름을 문자열 리터럴로 전달하는 경우는 드뭅니다. tokentoll의 다중 패스 상수 전파 엔진은 다음을 따릅니다:
DEFAULT_MODEL = os.getenv("MODEL", "gpt-4o")
class Config:
model: str = DEFAULT_MODEL
config = Config()
kwargs = {"model": config.model, "max_tokens": 2000}
client.chat.completions.create(**kwargs)
# tokentoll resolves: model="gpt-4o", max_tokens=2000변수 할당 (
MODEL = "gpt-4o")os.getenv()/os.environ.get()폴백 값함수 기본 매개변수
클래스 속성 기본값
생성자 인수 전파
딕셔너리 리터럴 및 첨자 내용
**kwargs언패킹
로드맵
컨텍스트 인식 호출 빈도(계획됨): 모든 호출 사이트에서 균일한 볼륨을 가정하는 대신 주변 코드에서 호출/월을 추론합니다 (FastAPI 라우트 핸들러 = 높은 트래픽, 스크립트 = 낮음, 루프 = 곱셈).
JS/TS 지원(계획됨): JavaScript 및 TypeScript 파일에서 LLM 호출을 감지합니다.
비용 알림: PR이 비용 차이를 초과할 때 CI를 실패하게 만드는 구성 가능한 임계값.
제한 사항
런타임에 외부 구성 파일이나 데이터베이스에서 로드된 모델은 확인할 수 없습니다. 이러한 호출은 SDK별 기본값을 사용합니다 (.tokentoll.yml을 통해 구성 가능).
tiktoken이 설치되지 않은 경우 토큰 추정치는 문자 수/4 휴리스틱을 사용합니다.
월별 추정치는 호출 사이트당 균일한 호출 볼륨을 가정합니다 (
--calls-per-month,.tokentoll.yml또는 경로별 재정의를 통해 구성 가능). 테스트 및 예제 파일을 건너뛰려면exclude옵션을 사용하세요.현재는 Python만 지원합니다 (JS/TS 지원 계획됨).
라이선스
MIT
Available Tools
2 toolsdiffA
Compare LLM costs between two git refs.
Shows which LLM call sites were added, removed, or changed between the base and head refs, along with the cost impact of those changes.
Args: base_ref: The base git ref (branch, tag, or commit) to compare from. head_ref: The head git ref to compare to. Defaults to HEAD.
Returns: JSON string with the diff results including cost changes.
| Name | Required | Description | Default |
|---|---|---|---|
| base_ref | Yes | ||
| head_ref | No | HEAD |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It indicates a read-like operation (diff) and describes the output, but does not explicitly state side effects or permissions. Adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with purpose, and includes parameter docs and return type. Every sentence adds value without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of an output schema (not shown), the description adequately covers purpose, parameters, and output format. It could include examples or edge cases but is sufficiently complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description documents both parameters: base_ref as the base git ref and head_ref as the head ref defaulting to HEAD. This adds essential meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares LLM costs between two git refs, specifying it shows added, removed, or changed call sites and cost impact. This distinguishes it from the sibling 'scan' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (to compare costs between refs) but does not explicitly state when not to use it or mention alternatives. Usage is well implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scanA
Scan a directory for LLM API calls and estimate monthly costs.
Finds all LLM API call sites (OpenAI, Anthropic, etc.) in the given path and produces a cost estimate based on token counts and pricing.
Args: path: Directory or file path to scan. Defaults to current directory. calls_per_month: Assumed monthly call volume per call site. If not provided, the CLI default (1000) is used.
Returns: JSON string with the scan results including call sites and cost estimates.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . | |
| calls_per_month | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It details the scanning action, cost estimation, and return format. While it doesn't cover every edge case (e.g., recursion depth or error handling), it provides sufficient behavioral insight for a read-only analysis tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a lead sentence, then details in Args and Returns sections. Every sentence adds value, and the format is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 optional params, no annotations), the description covers the core behavior and return type adequately. It could mention recursion or failure modes, but it is sufficient for most use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description fully explains both parameters: 'path' (directory/file, default current dir) and 'calls_per_month' (monthly volume, default null implying CLI default of 1000). This adds essential meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scans a directory for LLM API calls and estimates costs, specifying providers and purpose. This is a specific verb+resource that distinguishes it from the sibling 'diff'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use the tool (scanning directories for LLM calls and cost estimation). However, it does not explicitly mention when not to use it or provide alternatives, which prevents a top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
diff - First observed
scan
TDQS
Scored across 2 tools
The two tools, diff and scan, have clearly distinct purposes: scan finds LLM call sites and estimates costs, while diff compares costs between git refs. No overlap or ambiguity.
Both tool names are single verbs ('diff', 'scan'), which is consistent in style. While not a verb_noun pattern, the naming is uniform and intuitive for the domain.
With only 2 tools, the server is very focused. This can be appropriate for a narrow utility, but it feels thin for a full server. A few more tools (e.g., pricing config) might improve scope.
The tools cover two core operations: scanning and diffing. However, there is no tool for managing pricing configurations or listing assumptions, which could be gaps for advanced use.
Maintenance
Related MCP Connectors
Exact Claude API cost calc with real cache economics, plus a tiktoken-misuse scanner.
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
AI-powered codebase analysis — call graphs, security, dead code, complexity. 150+ tools.
Codebase intelligence for AI agents — dead code, blast radius, ownership.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI cost calculation, comparison, and optimization across major providers like Anthropic, OpenAI, Google, Meta, and Mistral. Supports cost estimation, budget-aware model finding, and token estimation through a simple API and MCP integration.-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to track LLM costs, enforce budgets, compare models, and estimate expenses through simple tool calls.-
- AlicenseBqualityDmaintenancePredict the cost of an LLM call before you make it, and pick the cheapest model that still does the job, offline, from your editor.732 npmApache 2.0
- AlicenseAqualityDmaintenanceExposes boyter/scc code counting and complexity analysis to LLM agents via read-only tools like counting lines, finding top files, and cost estimation.7BSD 3-Clause