Inferrail
A read-only, local MCP server that answers spend and health questions about your Inferrail receipt ledger without running inference or writing files.
Query attributed spend (
get_spend) — aggregate the local receipt ledger by a dimension:provider,model,route, or any business attribute name (defaults tocustomer, also things likeagentorwork_id).Filter by time window — restrict results with optional
sinceanduntilarguments.Get trustworthy cost numbers — returns known cost in USD as a decimal string (never a float) per group, plus a count of requests whose pricing was unresolvable, so unknown-cost calls are never silently folded in as
$0.Check gateway health (
get_health) — verify an Inferrail gateway is reachable at abase_url(defaulthttp://127.0.0.1:8000) via its/healthendpoint, and separately report the most recent local receipt as evidence of recent attributed activity.Spend nothing to ask — both tools read a local file only (
receipts_path), make no network inference call, and never consume provider budget.Safe by design — read-only: no inference, no file writes.
Tracks token usage and estimated LLM cost for OpenAI-compatible chat completion endpoints, including streaming and tool calls, by routing requests through the Inferrail gateway.
Developer preview. Everything below is implemented and tested. CLI flags, config, and receipt fields may still change before 1.0.
Quickstart
Requires Python 3.11+.
Protect one run, inside your Python app. No config file and no second terminal:
pip install inferrail
export OPENAI_API_KEY=sk-...import inferrail
from openai import OpenAI
base_url = inferrail.start() # the gateway, on a background thread in this process
client = OpenAI(base_url=base_url, api_key="unused")
client.chat.completions.create(
model="gpt-4o-mini", # example: any model your account can use (`inferrail models` lists them)
max_tokens=200,
messages=[{"role": "user", "content": "Summarize this contract."}],
extra_headers={
"X-Inferrail-Attribute-Work-Id": "contract-review-42", # the run
"X-Inferrail-Budget-Usd": "0.50", # its dollar ceiling
},
)Every call that carries the same run id shares that budget, including parallel calls. A call that would push the run past it gets HTTP 402 before it reaches OpenAI. Then:
inferrail work contract-review-42shows what the run cost.
Inferrail doesn't choose a model: whatever model id you send is passed to
the provider. A dollar budget needs a price for that model; inferrail models shows which models have one, and inferrail.start(pricing=...)
adds a price for a new model. Framework snippets (LangChain, LangGraph,
OpenAI Agents SDK, CrewAI, Haystack, LlamaIndex, Microsoft Agent Framework):
recipe.
Try it offline. No API key, no network calls, no provider charges:
pip install inferrail
inferrail demo
inferrail report --by customer --receipts ./inferrail-demo-receipts.jsonlThe demo sends scripted requests through the real engine to a fake
provider with made-up prices labeled DEMO, and prints what they cost by
customer (recording).
Run it as its own process, with the local dashboard. For a gateway shared by several apps, or to watch calls live, set your key in the gateway's terminal:
export OPENAI_API_KEY=sk-... # and/or ANTHROPIC_API_KEY
inferrail serve --quickstart --app-modeOpen the Dashboard: URL it prints, then point your app at the gateway
and tag the work:
from openai import OpenAI
# api_key is a placeholder: your provider key stays in the gateway
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
client.chat.completions.create(
model="gpt-4o-mini", # example model id; Inferrail passes yours through
messages=[{"role": "user", "content": "Summarize this contract."}],
extra_headers={
"X-Inferrail-Attribute-Customer": "acme",
"X-Inferrail-Attribute-Work-Id": "contract-review-42",
"X-Inferrail-Budget-Usd": "0.50", # optional: a dollar ceiling for this work
},
)The call appears in the Live Feed, and the Work screen totals what
contract-review-42 cost. The Anthropic SDK works the same way with
base_url="http://127.0.0.1:8000" (no /v1). More clients and
frameworks: docs/integrations.md.
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install inferrailIf inferrail is still not found, the environment isn't active or pip
installed into a different Python. More in
docs/self-hosting.md.
Related MCP server: AgentCost MCP Server
What you get
Cost per unit of work. Tag calls with any attribute (
customer,workflow,work_id, …). The dashboard andinferrail report --by <tag>total known cost by those tags, andinferrail work <id>shows one job.A dollar budget for one agent run. Declare it in a header. Parallel calls share it safely, and calls that would exceed it get HTTP 402 before they reach the provider. Global, project, daily, and monthly budgets too. Recipe.
Numbers you can trust. A cost is recorded only when the provider reports usage and a price is on file. Otherwise it's
unknown, never a guessed$0.Local records. Receipts are SQLite or JSONL files on your machine. No Inferrail account or hosted service is involved.
Ask your agent. A read-only MCP server answers questions like "How much did work contract-review-42 cost?"
Privacy boundary
Where | What it sees or keeps |
Your provider | The full request, exactly as it would without Inferrail, under the provider's own policies. |
The gateway, in memory | Your provider key (from its own environment) and the prompts and responses it forwards. |
Receipts, on your disk | Provider, model, token usage, the price used and the cost, status, timing, and the attribution tags you send. |
Not in receipts | Prompt and response bodies. |
Attribution tags are stored exactly as sent, so use identifiers and keep
secrets and message content out of them. inferrail verify-payload-free
prints the live receipt schema, and canary tests check that message
bodies never reach receipts. Neither is a security audit. The optional
usage beacon sends nothing unless an endpoint is configured
(details). How to check all of this
yourself: docs/PRODUCT.md.
How it works
The gateway routes each request by model to a configured provider,
checks any budget in scope, forwards the call, and writes one receipt per
request, including requests a budget refused. Reports, the dashboard, and MCP tools all read those
receipts. Details: docs/ARCHITECTURE.md.
MCP
inferrail mcp is a stdio MCP server with two read-only tools over your
local receipts: get_spend (known cost and tokens grouped by provider,
model, route, or any tag such as work_id) and get_health. They don't
run inference or write files.
claude mcp add inferrail \
-e INFERRAIL_RECEIPTS_PATH=/absolute/path/to/inferrail-receipts.jsonl \
-- uvx --with "mcp>=2.0" inferrail mcpOther clients, and where --app-mode keeps its receipts:
docs/integrations.md.
Current support
Endpoints:
POST /v1/chat/completions(OpenAI-compatible) andPOST /v1/messages(Anthropic-compatible), with streaming and tool calls. Any client that lets you set a base URL and sends these request shapes can use them.Providers: OpenAI, Anthropic, and endpoints compatible with either. Built-in prices cover OpenAI and Anthropic models. Other endpoints need a price declared in your config, or their cost stays unknown.
Images in chat messages (
image_urlparts, e.g. a browser agent's screenshots) are forwarded to OpenAI-compatible providers. Under a budget each image reserves a fixed 3,000-token estimate; the actual cost comes from the provider's reported usage.Not supported: the OpenAI Responses API, embeddings, image generation, audio and the Realtime API, batch, and native Gemini or Bedrock APIs. Request fields the gateway can't account for are rejected with a clear error, never silently dropped.
Scope: only calls that go through a running gateway are counted, not everything on your provider account.
Exact contract and non-goals: docs/PRODUCT.md.
These are separate from the gateway above and not needed to use it.
AP invoice-exception recovery (experimental): decide and record a retry or human review for one invoice-extraction exception. Try
inferrail ap demo. Docs.Hosted trial (preview): a short-lived hosted gateway at tryinferrail.com/try. If you add a real key there, the hosted process holds it (key handling).
Work Economics and Economic Authority (experimental, hosted, Base Sepolia testnet only): Work Economics docs · Economic Authority docs · example.
Documentation
Integrations: SDKs, frameworks, attribution, MCP, voice
Self-hosting: install, configuration, storage, budgets, dashboard
Contributing, security, license
Questions and bugs: GitHub issues.
Security or privacy vulnerabilities: report privately, as described in
SECURITY.md. Contributing: CONTRIBUTING.md
(pytest needs no API key or network). License: Apache-2.0.
Available Tools
2 toolsget_healthCheck gateway health and recent attributionA
Check whether an Inferrail gateway is reachable at base_url via its /health endpoint, and separately report the most recent local receipt (if any) as evidence of recent attributed activity. Does NOT make a new inference call -- a health check must never silently spend provider budget as a side effect.
| Name | Required | Description | Default |
|---|---|---|---|
| base_url | No | http://127.0.0.1:8000 | |
| receipts_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it explicitly states the tool does NOT make a new inference call and must never silently spend provider budget. It could add more detail about failure modes or network behavior, but the key side-effect guarantee is clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The primary action is front-loaded, and the critical side-effect warning is placed prominently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are covered elsewhere. The description adequately explains the two parameters and the no-side-effect behavior. It is slightly incomplete in not addressing the sibling tool relationship, but for a simple health-check tool this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning by identifying base_url as the gateway address and mentioning local receipts, which maps to receipts_path. However, it does not explicitly explain the parameter names, defaults, or the nullability of receipts_path.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks gateway reachability via the /health endpoint and reports the most recent local receipt. It names a specific verb and resource, but does not explicitly differentiate from the sibling tool get_spend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this when you need a health check and recent attribution evidence without making an inference call. However, it does not explicitly state when to prefer this over get_spend or provide exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_spendQuery attributed spendA
Query Inferrail's local receipt ledger, aggregated by a dimension (provider, model, route, or any business attribute name such as 'customer' or 'agent'), optionally restricted to a time window. Returns known cost in USD (as a decimal string, never a float) per group, plus a count of requests whose pricing was unresolvable -- those are never silently folded into the cost total as zero. Reads a local file only; makes no network call and cannot incur provider cost.
| Name | Required | Description | Default |
|---|---|---|---|
| by | No | customer | |
| since | No | ||
| until | No | ||
| receipts_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the behavioral burden. It explicitly states it reads a local file, makes no network call, cannot incur provider cost, returns cost as a decimal string (never a float), and does not silently fold unresolvable pricing into zero. This is highly transparent about side effects and return format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose and then adding essential behavioral details. There is no redundancy or filler; every sentence contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, 0% schema coverage, and no annotations, the description does a good job of covering purpose, behavior, and parameter meanings. It lacks explicit parameter-level details like date formats or file path defaults, but the presence of an output schema and the clarity of the description make it sufficient for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the 'by' parameter as an aggregation dimension with examples (provider, model, route, customer, agent) and the time window via 'since' and 'until'. However, it does not explicitly explain 'receipts_path' (though it implies a local file), nor does it mention defaults or date formats. The description covers the core semantics but leaves some details to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it queries the local receipt ledger, aggregates by a dimension (provider, model, route, or business attribute), and optionally restricts to a time window. This is a specific verb and resource, and it is distinct from the sibling get_health, which presumably checks health rather than spend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context about when to use the tool (e.g., it reads a local file only, makes no network call, cannot incur provider cost), but it does not explicitly mention alternatives or when not to use it. The sibling get_health implies a different purpose, but no direct exclusion is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.5- First observed
get_health - First observed
get_spend
TDQS
Scored across 2 tools
get_health and get_spend are clearly distinct: one checks gateway reachability and recent receipt, the other aggregates spend from the local ledger. There is no overlap in purpose or side effects.
Both tools follow the same get_ verb_noun convention (get_health, get_spend), creating a predictable and consistent pattern for the server's small surface.
With only 2 tools, the server feels thin, though the narrow scope of health and spend monitoring makes the count understandable. It is at the low end of acceptable for a focused read-only utility set.
The server covers the two core monitoring concerns for an Inferrail gateway: health and spend. Minor gaps like a tool to list available routes or models exist, but the current surface is sufficient for basic observability.
Maintenance
Related MCP Connectors
Enforce AI budgets before the model call and track cost per customer across 10 providers.
Budget & cost control for AI agents — per-agent spend caps + rate limits before each call.
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Enterprise AI Control Plane: governance, guardrails, spend tracking, compliance & smart routing.
Related MCP Servers
AlicenseNot gradedqualityDmaintenanceProvides cost intelligence and a reputation scoring system to help AI agents optimize spending through smart model selection and local-to-cloud routing. It enables real-time cost tracking and rewards agents for making efficient, high-credibility decisions across various LLM providers.Apache 2.0- FlicenseNot gradedqualityDmaintenanceProvides real-time AI model pricing, cost estimation, and budget management tools to help agents understand and optimize their spending. It enables agents to compare costs across multiple providers and select the most cost-effective models for specific tasks.1-
- AlicenseNot gradedqualityDmaintenanceCompare AI inference pricing across 9 providers in real time. Routing recommendations, spend tracking, and budget alerts for AI agents.102 npmMIT
- AlicenseAqualityDmaintenanceTracks AI agent token usage and spending in real time, with budget alerts, per-task cost breakdown, and a visual dashboard.61MIT