Skip to main content
Glama

Developer preview. Everything below is implemented and tested. CLI flags, config, and receipt fields may still change before 1.0.

Quickstart

Requires Python 3.11+.

Protect one run, inside your Python app. No config file and no second terminal:

pip install inferrail
export OPENAI_API_KEY=sk-...
import inferrail
from openai import OpenAI

base_url = inferrail.start()   # the gateway, on a background thread in this process
client = OpenAI(base_url=base_url, api_key="unused")
client.chat.completions.create(
    model="gpt-4o-mini",   # example: any model your account can use (`inferrail models` lists them)
    max_tokens=200,
    messages=[{"role": "user", "content": "Summarize this contract."}],
    extra_headers={
        "X-Inferrail-Attribute-Work-Id": "contract-review-42",   # the run
        "X-Inferrail-Budget-Usd": "0.50",                        # its dollar ceiling
    },
)

Every call that carries the same run id shares that budget, including parallel calls. A call that would push the run past it gets HTTP 402 before it reaches OpenAI. Then:

inferrail work contract-review-42

shows what the run cost.

Inferrail doesn't choose a model: whatever model id you send is passed to the provider. A dollar budget needs a price for that model; inferrail models shows which models have one, and inferrail.start(pricing=...) adds a price for a new model. Framework snippets (LangChain, LangGraph, OpenAI Agents SDK, CrewAI, Haystack, LlamaIndex, Microsoft Agent Framework): recipe.

Try it offline. No API key, no network calls, no provider charges:

pip install inferrail
inferrail demo
inferrail report --by customer --receipts ./inferrail-demo-receipts.jsonl

The demo sends scripted requests through the real engine to a fake provider with made-up prices labeled DEMO, and prints what they cost by customer (recording).

Run it as its own process, with the local dashboard. For a gateway shared by several apps, or to watch calls live, set your key in the gateway's terminal:

export OPENAI_API_KEY=sk-...          # and/or ANTHROPIC_API_KEY
inferrail serve --quickstart --app-mode

Open the Dashboard: URL it prints, then point your app at the gateway and tag the work:

from openai import OpenAI

# api_key is a placeholder: your provider key stays in the gateway
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")
client.chat.completions.create(
    model="gpt-4o-mini",   # example model id; Inferrail passes yours through
    messages=[{"role": "user", "content": "Summarize this contract."}],
    extra_headers={
        "X-Inferrail-Attribute-Customer": "acme",
        "X-Inferrail-Attribute-Work-Id": "contract-review-42",
        "X-Inferrail-Budget-Usd": "0.50",   # optional: a dollar ceiling for this work
    },
)

The call appears in the Live Feed, and the Work screen totals what contract-review-42 cost. The Anthropic SDK works the same way with base_url="http://127.0.0.1:8000" (no /v1). More clients and frameworks: docs/integrations.md.

python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
python -m pip install inferrail

If inferrail is still not found, the environment isn't active or pip installed into a different Python. More in docs/self-hosting.md.

Related MCP server: AgentCost MCP Server

What you get

  • Cost per unit of work. Tag calls with any attribute (customer, workflow, work_id, …). The dashboard and inferrail report --by <tag> total known cost by those tags, and inferrail work <id> shows one job.

  • A dollar budget for one agent run. Declare it in a header. Parallel calls share it safely, and calls that would exceed it get HTTP 402 before they reach the provider. Global, project, daily, and monthly budgets too. Recipe.

  • Numbers you can trust. A cost is recorded only when the provider reports usage and a price is on file. Otherwise it's unknown, never a guessed $0.

  • Local records. Receipts are SQLite or JSONL files on your machine. No Inferrail account or hosted service is involved.

  • Ask your agent. A read-only MCP server answers questions like "How much did work contract-review-42 cost?"

Privacy boundary

Where

What it sees or keeps

Your provider

The full request, exactly as it would without Inferrail, under the provider's own policies.

The gateway, in memory

Your provider key (from its own environment) and the prompts and responses it forwards.

Receipts, on your disk

Provider, model, token usage, the price used and the cost, status, timing, and the attribution tags you send.

Not in receipts

Prompt and response bodies.

Attribution tags are stored exactly as sent, so use identifiers and keep secrets and message content out of them. inferrail verify-payload-free prints the live receipt schema, and canary tests check that message bodies never reach receipts. Neither is a security audit. The optional usage beacon sends nothing unless an endpoint is configured (details). How to check all of this yourself: docs/PRODUCT.md.

How it works

The gateway routes each request by model to a configured provider, checks any budget in scope, forwards the call, and writes one receipt per request, including requests a budget refused. Reports, the dashboard, and MCP tools all read those receipts. Details: docs/ARCHITECTURE.md.

MCP

inferrail mcp is a stdio MCP server with two read-only tools over your local receipts: get_spend (known cost and tokens grouped by provider, model, route, or any tag such as work_id) and get_health. They don't run inference or write files.

claude mcp add inferrail \
  -e INFERRAIL_RECEIPTS_PATH=/absolute/path/to/inferrail-receipts.jsonl \
  -- uvx --with "mcp>=2.0" inferrail mcp

Other clients, and where --app-mode keeps its receipts: docs/integrations.md.

Current support

  • Endpoints: POST /v1/chat/completions (OpenAI-compatible) and POST /v1/messages (Anthropic-compatible), with streaming and tool calls. Any client that lets you set a base URL and sends these request shapes can use them.

  • Providers: OpenAI, Anthropic, and endpoints compatible with either. Built-in prices cover OpenAI and Anthropic models. Other endpoints need a price declared in your config, or their cost stays unknown.

  • Images in chat messages (image_url parts, e.g. a browser agent's screenshots) are forwarded to OpenAI-compatible providers. Under a budget each image reserves a fixed 3,000-token estimate; the actual cost comes from the provider's reported usage.

  • Not supported: the OpenAI Responses API, embeddings, image generation, audio and the Realtime API, batch, and native Gemini or Bedrock APIs. Request fields the gateway can't account for are rejected with a clear error, never silently dropped.

  • Scope: only calls that go through a running gateway are counted, not everything on your provider account.

Exact contract and non-goals: docs/PRODUCT.md.

These are separate from the gateway above and not needed to use it.

  • AP invoice-exception recovery (experimental): decide and record a retry or human review for one invoice-extraction exception. Try inferrail ap demo. Docs.

  • Hosted trial (preview): a short-lived hosted gateway at tryinferrail.com/try. If you add a real key there, the hosted process holds it (key handling).

  • Work Economics and Economic Authority (experimental, hosted, Base Sepolia testnet only): Work Economics docs · Economic Authority docs · example.

Documentation

Contributing, security, license

Questions and bugs: GitHub issues. Security or privacy vulnerabilities: report privately, as described in SECURITY.md. Contributing: CONTRIBUTING.md (pytest needs no API key or network). License: Apache-2.0.

Available Tools

2 tools
get_healthCheck gateway health and recent attributionA

Check whether an Inferrail gateway is reachable at base_url via its /health endpoint, and separately report the most recent local receipt (if any) as evidence of recent attributed activity. Does NOT make a new inference call -- a health check must never silently spend provider budget as a side effect.

ParametersJSON Schema
NameRequiredDescriptionDefault
base_urlNohttp://127.0.0.1:8000
receipts_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well: it explicitly states the tool does NOT make a new inference call and must never silently spend provider budget. It could add more detail about failure modes or network behavior, but the key side-effect guarantee is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The primary action is front-loaded, and the critical side-effect warning is placed prominently. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are covered elsewhere. The description adequately explains the two parameters and the no-side-effect behavior. It is slightly incomplete in not addressing the sibling tool relationship, but for a simple health-check tool this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning by identifying base_url as the gateway address and mentioning local receipts, which maps to receipts_path. However, it does not explicitly explain the parameter names, defaults, or the nullability of receipts_path.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks gateway reachability via the /health endpoint and reports the most recent local receipt. It names a specific verb and resource, but does not explicitly differentiate from the sibling tool get_spend.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this when you need a health check and recent attribution evidence without making an inference call. However, it does not explicitly state when to prefer this over get_spend or provide exclusion criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_spendQuery attributed spendA

Query Inferrail's local receipt ledger, aggregated by a dimension (provider, model, route, or any business attribute name such as 'customer' or 'agent'), optionally restricted to a time window. Returns known cost in USD (as a decimal string, never a float) per group, plus a count of requests whose pricing was unresolvable -- those are never silently folded into the cost total as zero. Reads a local file only; makes no network call and cannot incur provider cost.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNocustomer
sinceNo
untilNo
receipts_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the behavioral burden. It explicitly states it reads a local file, makes no network call, cannot incur provider cost, returns cost as a decimal string (never a float), and does not silently fold unresolvable pricing into zero. This is highly transparent about side effects and return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary purpose and then adding essential behavioral details. There is no redundancy or filler; every sentence contributes value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters, 0% schema coverage, and no annotations, the description does a good job of covering purpose, behavior, and parameter meanings. It lacks explicit parameter-level details like date formats or file path defaults, but the presence of an output schema and the clarity of the description make it sufficient for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the 'by' parameter as an aggregation dimension with examples (provider, model, route, customer, agent) and the time window via 'since' and 'until'. However, it does not explicitly explain 'receipts_path' (though it implies a local file), nor does it mention defaults or date formats. The description covers the core semantics but leaves some details to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it queries the local receipt ledger, aggregates by a dimension (provider, model, route, or business attribute), and optionally restricts to a time window. This is a specific verb and resource, and it is distinct from the sibling get_health, which presumably checks health rather than spend.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context about when to use the tool (e.g., it reads a local file only, makes no network call, cannot incur provider cost), but it does not explicitly mention alternatives or when not to use it. The sibling get_health implies a different purpose, but no direct exclusion is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.4.5
    • First observedget_health
    • First observedget_spend

TDQS

A4.1/5.0

Scored across 2 tools

Disambiguation5/5

get_health and get_spend are clearly distinct: one checks gateway reachability and recent receipt, the other aggregates spend from the local ledger. There is no overlap in purpose or side effects.

Naming Consistency5/5

Both tools follow the same get_ verb_noun convention (get_health, get_spend), creating a predictable and consistent pattern for the server's small surface.

Tool Count3/5

With only 2 tools, the server feels thin, though the narrow scope of health and spend monitoring makes the count understandable. It is at the low end of acceptable for a focused read-only utility set.

Completeness4/5

The server covers the two core monitoring concerns for an Inferrail gateway: health and spend. Minor gaps like a tool to list available routes or models exist, but the current surface is sufficient for basic observability.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides cost intelligence and a reputation scoring system to help AI agents optimize spending through smart model selection and local-to-cloud routing. It enables real-time cost tracking and rewards agents for making efficient, high-credibility decisions across various LLM providers.
    Apache 2.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides real-time AI model pricing, cost estimation, and budget management tools to help agents understand and optimize their spending. It enables agents to compare costs across multiple providers and select the most cost-effective models for specific tasks.
    1
    -