agentic-ledger
Query and inspect captured LLM agent traces, sessions, and loop runs via MCP tools.
List recent sessions with aggregated call count, cost, tokens, and start time.
Retrieve the full trace for any single LLM call: prompt, response, tools, tokens, cost, latency.
Get all calls in a session chronologically, with optional full message bodies.
Full-text search across captured prompts, outputs, system prompts, agent names, and user IDs.
List loop runs with iterations, sessions, cost, flagged-call counts, and status.
Check a run's status (iterations, cost, tokens, completion promise) to decide whether to continue iterating.
Captures and monitors LLM calls made to OpenAI's API (including GPT models) via a transparent proxy, providing full trace visibility and cost analysis.
Ingests GenAI spans from Pydantic AI (and other OTel-native frameworks) via OpenTelemetry, enabling observability without a proxy.
Ingests GenAI spans from Vercel AI SDK (and other OTel-native frameworks) via OpenTelemetry, enabling observability without a proxy.
Agentic Ledger
Runtime observability for AI agents - see exactly what your agent did, why it did it, and what it cost.
Website: agentic-ledger.dev
The numbers are meant to match your provider bill. If they don't, that's a bug we want.
Works with any agent framework, any LLM provider, any model gateway. Zero code changes required. Point your agent at the proxy and everything is captured automatically.
How it works
Agentic Ledger runs as a transparent proxy between your agent and the LLM provider. It intercepts every request and response, assigns it an action_id, stores it, and returns the upstream response unmodified. Your agent never knows the proxy is there. The full picture, with
diagrams and a module map for contributors, lives in
ARCHITECTURE.md.
Your Agent → Agentic Ledger Proxy → OpenAI / Anthropic / LiteLLM / any LLM
↓
SQLite or Postgres
↓
Live Dashboard + APIRelated MCP server: swarm-at
Quick Start
Step 1 - Start the proxy
Coming from Helicone or LangSmith? The migration page does the translation in two lines. Running a context compressor like Headroom? They chain.
Two commands, zero config, no terminal held hostage:
uv tool install agentic-ledger # or: pipx install agentic-ledger, or pip install -U agentic-ledger
agenticledger start # runs in the background; terminal freedA tool-managed install (uv tool / pipx) gets its own isolated
environment and one unambiguous shim on PATH, so shadowing by another
Python's copy becomes rare and doctor-detectable, and agenticledger upgrade always means exactly one thing. Plain pip works too; if a machine ever grows
competing installs, agenticledger doctor --fix untangles them.
agenticledger start prints the dashboard URL and gives your terminal
back - closing the window doesn't stop it. agenticledger status tells
you it's up and healthy, agenticledger logs shows what it's doing,
agenticledger stop shuts it down. Want a config file anyway?
agenticledger init writes a commented one; see
Configuration for what goes in it.
Or with Docker (no Python required):
docker run -p 8000:8000 \
-e AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \ # optional: omit to route by call format
-v $(pwd)/data:/data \
ghcr.io/shekharbhardwaj/agentic-ledger:latestThe image is multi-arch (amd64/arm64), runs as a non-root user, and every release is signed with Sigstore and ships an SBOM. Hardening a shared deployment (TLS, auth keys, redaction, verification)? See the deployment guide.
Using Anthropic / Claude? Nothing to configure: with no upstream set, the proxy routes each call by its wire format, so Anthropic-style calls go to Anthropic and OpenAI-style calls go to OpenAI, side by side through one proxy. Setting an explicit
upstream_url(a gateway like LiteLLM or OpenRouter, LM Studio, or a pinned provider) switches to the classic one-proxy-one-provider behavior, mismatch hints included.
Or with docker compose (SQLite by default - see docker-compose.yml):
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com docker compose upFor Postgres, layer the override; it swaps the DSN and adds the database service, and the image ships the driver:
docker compose -f docker-compose.yml -f docker-compose.postgres.yml upWith uv:
uv add agentic-ledger
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com uv run python -m agenticledger.proxyWith pip:
python -m venv venv && source venv/bin/activate
pip install -U agentic-ledger
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com ./venv/bin/python -m agenticledger.proxyPostgres? Install the extra and set
AGENTICLEDGER_DSN:pip install "agentic-ledger[postgres]" AGENTICLEDGER_DSN=postgresql://user:password@localhost/agenticledgerNote: the Docker image uses SQLite only. For Postgres with Docker, install via
pipinstead.
OpenTelemetry? Install the extra and set
AGENTICLEDGER_OTEL_ENDPOINT:pip install "agentic-ledger[otel]" AGENTICLEDGER_OTEL_ENDPOINT=http://localhost:4318
Proxy starts on http://localhost:8000. Traces are saved to ~/.agenticledger/agenticledger.db when started with agenticledger start (one home for the background service, wherever you launched it from), to agenticledger.db in the current folder when run in the foreground (agenticledger serve / python -m agenticledger.proxy), or to /data/agenticledger.db in Docker.
Step 2 - Point your agent at the proxy
For Claude Code, BMAD, or OpenClaw, one command writes the config for you (backed up, merged, Docker-aware):
agenticledger connect claude-code # or: bmad, openclawFor everything else, two changes: set base_url to the proxy and add a session ID header to group calls into a run. Everything else - your API key, model, messages - stays exactly the same.
OpenAI:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1", # ← proxy
api_key="your-openai-key",
default_headers={"x-agenticledger-session-id": "run-1"},
)
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Research the top 3 AI trends in 2026"}],
)Anthropic (no upstream config needed: /v1/messages calls route to Anthropic automatically):
import anthropic
client = anthropic.Anthropic(
base_url="http://localhost:8000", # ← proxy
api_key="your-anthropic-key",
default_headers={"x-agenticledger-session-id": "run-1"},
)Azure OpenAI: point AzureOpenAI(azure_endpoint="http://localhost:8000") at the ledger with your resource set as the upstream; deployments are priced from the model the response names. See the Azure guide.
AWS Bedrock: install agentic-ledger[bedrock], give the ledger AWS credentials through the standard chain, and point boto3 (endpoint_url) or Claude Code (ANTHROPIC_BEDROCK_BASE_URL) at it; the ledger re-signs each call itself. Both wires are covered: InvokeModel and the modern Converse/ConverseStream APIs. See the Bedrock guide.
LiteLLM / OpenRouter / any gateway:
# Point Agentic Ledger at your gateway
AGENTICLEDGER_UPSTREAM_URL=http://localhost:4000 uv run python -m agenticledger.proxy
# Then point your agent at Agentic Ledger
client = OpenAI(base_url="http://localhost:8000/v1", ...)Step 3 - Open the dashboard
http://localhost:8000The web app updates live via WebSocket as calls come in. No refresh needed.
Website-matched design - the dashboard shares the charcoal surfaces, ice-blue accents, geometric raccoon mark, and typography of agentic-ledger.dev. Session Overview adds recorded cost, calls, token totals, and a per-call cost chart. Loop Lens pairs iteration costs with a "While you were away" timeline. Unknown call prices remain explicit; all charts use captured ledger data. Calls, Flow, Trace, replay, budget controls, and light/dark appearance remain available.
Loop Lens - every loop run with its observed status (Running / Flagged / Completion declared / Ended / Calls blocked), one open metric strip (recorded spend, the run ceiling with an honest accounting track, model calls), Overview / Activity / Cache views, a recorded-event timeline that jumps straight to the evidence, a Block calls action that refuses a running loop's further calls at the wall (and Allow calls again to lift it; the agent being blocked cannot), per-iteration breakdowns, and plain-English explanations of every flag. Pick any two runs with ⇆ to diff them side by side - cost, iterations, calls, flags, duration with signed deltas, plus a prompt drift diff showing exactly what changed in the system prompt and opening instruction between the runs.
Sessions - flat, scannable rows and three views: call rows (time, model, one status, latency, cost) that expand into a four-tab inspector (Response, Tools, Prompt, Raw), a Flow DAG of agent handoffs, and a Trace waterfall with real parent links from the loop engine. Rows say whose they are at a glance: team badge, red for real failures, amber for deliberate refusals, purple for replays, and a run chip linking each session to its loop.
Replay the whole run - the question that decides a model switch isn't "how did it handle one call?" but "would my loop have survived?" Pick a run or session, pick a destination (a local model is free), and every step re-runs with its original inputs. You get a report card, not homework: "34 / 40 moments matched", the fumbles named ("dropped the tools"), and the cost both ways. Each step is a real captured moment replayed honestly - after step one a different model would have steered a different conversation, so the ledger compares moments, not fairy tales.
In your pocket -
agenticledger shareopens an https tunnel you own (via cloudflared, no account), prints the pairing link, and draws a QR in the terminal: point your phone's camera and the dashboard is in your hand, kill switch, stop all calls, and ceilings included.--wififor a same-network link,--rotateto un-pair every device, or press Pair a device in the dashboard's ⚿ panel. Local machines never need a key; everyone else meets the auto-generated pairing key. The dashboard fits a phone: one pane at a time, a back button, prev/next arrows to flip between runs.The wall, fleet-wide - Stop all calls is the emergency stop for every agent at once: one button in Loop Lens (with a confirm), a banner on every page while it is on, and it survives a restart until someone lifts it. Allow and deny lists for models and providers (
AGENTICLEDGER_ALLOW_MODELS,AGENTICLEDGER_DENY_PROVIDERS, and friends, globs welcome) turn the wrong model away with the rule named; a team card can carry its own lists and can only narrow the fleet's. Every refusal is on the record: rate limits, loop guards, budgets, ceilings, the kill switch and the stop all land as amberblocked:rows with the reason, counted in Reports and/metrics. A loop block is lifted from the session in the dashboard, no restart. The ledger only ever refuses or records; it never rewrites, reroutes, or substitutes a call.The cache audit - every run answers "was I paying full price for repeated text?" Received discount is exact from the provider's own cache reports; the missed amount is a labeled estimate with its method shown; every verdict carries the reason and a one-line fix, including "nothing missed, you're fine". Also at
GET /api/runs/{id}/cache-audit.Yours to keep - dark, light, or system appearance (a browser-local choice), and URLs that hold the investigation: deep links to runs, sessions and single calls (
#/sessions/<id>/calls/<action_id>opens the session with that call expanded), working Back/Forward, no credentials ever in ordinary links. Lists load the newest 50 and keep loading older with one press, filters reach the whole history, and the count on each sidebar is the real total.Named instances -
agenticledger start --name demo --port 8003runs a second ledger beside your everyday one: own state, own database, its dashboard wears an amber name chip so it can never pass for the real thing.stop,status,logs,share, andrunall take--name.The spend meter - a run's detail reads its money live: spent so far, burning $/h, "at this pace $Y by 8:00 AM". Give any run a cost ceiling and the proxy refuses further calls the moment spend reaches it (amber, costing nothing) until you raise or clear it; the ceiling survives restarts and guards auto-detected loops too. A webhook alert fires at 80%.
Names, pins, projects, icons - call it "the overnight auth fix" instead of
cc-73a26366, ★ pin what matters to the top, file work under a project and the Sessions view reads as sections: a heading per project, its sessions beneath, the unfiled pile last. A run filed under a project files its sessions with it. Give a loop or a session an icon and a color from the ✎ editor's picker and it stands out at a glance.Settings - the ⚙ shows what the proxy is actually running with: config file in effect, upstream, budgets, replay targets, each row labeled file / env / default, secrets hidden. Every row the config file can set has a Change link: pick or type the new value and it is written to the config file, then one Restart now applies all of them from the dashboard (the ledger re-executes itself in place). A row set by the environment says so, since the environment always wins.
Replay & what-if - open any call and ↻ Replay it: pick a destination (the panel lists what your local server actually has loaded), and the exact captured prompt re-executes there - same provider, the other one, or a free local model via LM Studio; tool calls, schemas, and system prompts are translated between the Anthropic and OpenAI wire formats automatically. Works even on calls your own budget blocked - the wall can say no and you can still see what would have happened, for $0. Replays tie back to their original with ↩ Open original. The what-if box answers the cheaper question first: reprice any run or session on another model with pure math, no API calls. (Configure
AGENTICLEDGER_REPLAY_API_KEYand/or the per-providerAGENTICLEDGER_REPLAY_*_KEYtargets.)Reports - where the money goes: spend per day, model mix with latency p50/p95/p99, per-agent totals, a by-team table with each team's spend against its card's daily allowance ("who ran dry?" in one glance), and cache savings - what your prompt-cache traffic would have cost at full input rates versus what it actually cost. Errors and blocks are counted apart everywhere: red = something broke, amber = the ledger refused on purpose - a healthy wall never makes a healthy agent look sick
Search - full-text search across all sessions by prompt, output, agent name, or user ID
Configuration
agenticledger init writes agenticledger.toml with every option
commented. Uncomment what you need - a working setup looks like this:
[proxy]
port = 8000
upstream_url = "https://api.anthropic.com"
db = "sqlite:///agenticledger.db"
[keys]
# Prefer *_file: the file's contents are the key, so no secret lives in
# this file or your shell history (chmod 600 the key file).
api_key_file = "~/.agenticledger/api.key" # dashboard/admin access
ingest_key_file = "~/.agenticledger/ingest.key" # closes the open relay
[budgets]
daily = 25.0 # whole-ledger daily ceiling, USD
session = 5.0 # per-session ceiling
[replay]
# Free local replay via LM Studio (any key works there):
openai_url = "http://localhost:1234"
openai_key = "lm-studio"Three rules:
The file is found in this order:
AGENTICLEDGER_CONFIG, then./agenticledger.toml(the folder you start from), then~/.agenticledger/config.toml. First match wins; the startup banner names the file in effect.Anything typed in the command beats the file. Env vars override per-setting (
AGENTICLEDGER_PORT=9000 agenticledger startuses 9000 for that run without touching the file) - which is also why Docker and CI setups configured by env vars are unaffected.Changes apply on restart (
agenticledger stopthenstart).
Every setting in the environment-variable reference
below has a config-file home; an [env] section passes any other
AGENTICLEDGER_* variable through verbatim.
Providers, step by step
Every provider below rides the same proxy; the only thing that changes is
which base URL you point at it. Each recipe assumes the proxy is up
(agenticledger start) and ends with the same check: run one call, open
http://localhost:8000, and see it in Sessions.
OpenAI (and any OpenAI-compatible API)
Point the client at the proxy:
export OPENAI_BASE_URL=http://localhost:8000/v1Keep your
OPENAI_API_KEYexactly as it was - the proxy passes your auth header through untouched.Make a call; it appears in Sessions with an O mark.
Anthropic
Point the client at the proxy:
export ANTHROPIC_BASE_URL=http://localhost:8000Keep your
ANTHROPIC_API_KEYas it was.Make a call; it appears with an A mark. No upstream config needed - the proxy routes Anthropic-shaped calls to Anthropic by wire format.
AWS Bedrock (direct capture)
Give the ledger AWS credentials of its own through the standard chain (env vars,
~/.awsprofile, or an instance role) scoped tobedrock:InvokeModelandbedrock:InvokeModelWithResponseStream, then install the extra and restart:pip install "agentic-ledger[bedrock]" agenticledger stop && agenticledger startCheck the ⚙ Settings panel: the Bedrock row should read "signing as the ledger in ".
Point the client at the proxy - Claude Code:
export CLAUDE_CODE_USE_BEDROCK=1 export ANTHROPIC_BEDROCK_BASE_URL=http://localhost:8000boto3:
boto3.client("bedrock-runtime", endpoint_url="http://localhost:8000").Make a call; it appears with an orange B mark. The ledger strips the caller's identity and re-signs with its own credentials. Full guide: docs/integrations/bedrock.md.
AWS Bedrock through a company gateway
Many companies front Bedrock with a gateway that does the authentication itself; Claude Code is then configured with CLAUDE_CODE_SKIP_BEDROCK_AUTH=1 and ANTHROPIC_BEDROCK_BASE_URL pointing at the gateway, usually from a managed settings file. The ledger sits in front of that gateway and forwards each call exactly as the agent sent it, headers included: no signing, no AWS credentials on your side.
Tell the ledger where the gateway is (the host
ANTHROPIC_BEDROCK_BASE_URLpointed at before):agenticledger config set proxy.bedrock_gateway_url https://bedrock-gateway.company.example agenticledger stop && agenticledger startPoint Claude Code at the ledger instead:
ANTHROPIC_BEDROCK_BASE_URL=http://localhost:8000. If that value lives in the IT-managed settings file (/Library/Application Support/ClaudeCode/managed-settings.json), it overrides everything else and IT has to change it; if it is in~/.claude/settings.json, edit it there./healthand the ⚙ Settings panel read "forwarding as sent to the gateway at ...". Calls appear with the Bedrock mark and the region-prefixed model id.
Azure OpenAI
Set the upstream to your resource:
agenticledger config set proxy.upstream_url https://<resource>.openai.azure.com agenticledger stop && agenticledger startPoint the client's Azure endpoint at
http://localhost:8000; keep yourapi-keyheader as it was.Calls are tagged
azure-openaiand priced by the model the RESPONSE names, so deployment aliases can't hide the real model. Full guide: docs/integrations/azure-openai.md.
Local models (LM Studio, Ollama with the OpenAI API)
Set the upstream to the local server:
agenticledger config set proxy.upstream_url http://localhost:1234 agenticledger stop && agenticledger startexport OPENAI_BASE_URL=http://localhost:8000/v1in the agent.Calls appear with a purple mark and $0 cost. Full guide: docs/integrations/lm-studio.md.
Gateways (OpenRouter, LiteLLM)
Set the upstream to the gateway:
agenticledger config set proxy.upstream_url https://openrouter.ai/api agenticledger stop && agenticledger startexport OPENAI_BASE_URL=http://localhost:8000/v1; keep the gateway key as it was.Gateway-prefixed model ids ("anthropic/claude-...") price correctly via substring matching. Guides: openrouter.md, litellm.md.
Framework-specific recipes (CrewAI, LangGraph, AutoGen, Vercel AI SDK, pydantic-ai, and more) live in docs/integrations/.
Coding agents - Claude Code, Ralph loops & friends
Claude Code (and most coding agents) can be pointed at the proxy with a single environment variable - no headers, no code changes:
agenticledger startexport ANTHROPIC_BASE_URL=http://localhost:8000
claudeNo upstream config needed: calls route to the provider matching their wire format.
Agentic Ledger fingerprints Claude Code traffic automatically: every call is
tagged framework=claude-code, and instead of one undifferentiated bucket,
each Claude Code session appears under its real session UUID (the same id
claude --resume shows), with prompt-cache reads/writes captured and priced
correctly - cache traffic is where most of a coding agent's real spend lives.
Want a loop filed under a name you chose? Put one word in front of the command you already run:
agenticledger run nightly-digest -- python agent.pyYour command runs exactly as before; its LLM calls land on the run tile
named nightly-digest, and each launch counts as the next iteration, so
tomorrow's run joins the same tile. Nothing in your agent's code changes.
Add --project acme to file the run under a dashboard project as it starts.
Running an overnight loop (Ralph-style while :; do cat PROMPT.md | claude -p; done)?
The same command with loop flags re-executes your command each iteration,
attributes every call to the run (via the base URL, no headers needed), and
stops on a completion promise, a budget ceiling, or the iteration cap:
AGENTICLEDGER_UPSTREAM_URL=https://api.anthropic.com \
AGENTICLEDGER_COMPLETION_PROMISE="ALL TASKS COMPLETE" \
uv run python -m agenticledger.proxyagenticledger run overnight --max-iterations 50 --budget 25 -- \
claude -p "$(cat PROMPT.md)" --dangerously-skip-permissionsEach iteration shows up as iteration N of the run in /api/runs; when the
agent prints the completion promise in a response, run status flips to
complete and the loop exits with a cost/token summary. The word after
run is the run's name; without one the run is named after the folder and
the minute (myproject-0819-1936). Rerunning the same name continues its
iteration count instead of restarting at 1. Any existing loop
script works too - poll GET /api/runs/{run_id} yourself, or let the proxy's
budgets (AGENTICLEDGER_BUDGET_DAILY=25.00) hard-stop a runaway loop.
Iterating on the prompt? Rerun and use ⇆ compare in the Loop Lens to diff the two runs - cost, iterations, calls, and flags side by side - so "did the new prompt actually help" gets a number instead of a feeling.
The same recipe works for any client with a base-URL override (Codex CLI,
opencode, OpenClaw, LiteLLM-based stacks) - set the OpenAI/Anthropic base URL
to the proxy and traffic is captured; add x-agenticledger-* headers when you
want explicit attribution.
OTel-native tools (Gemini CLI, Codex [otel], AutoGen/AG2, Pydantic AI,
Vercel AI SDK) don't need the proxy at all - point their OTLP exporter at the
ledger and GenAI spans are ingested directly:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:8000Both OTLP/HTTP encodings are accepted: JSON always, protobuf when the
[otel] extra is installed (the Docker image includes it). gRPC exporters
should switch to HTTP: OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf.
Framework guides - one per integration in docs/integrations: Claude Code, Codex CLI, opencode, OpenClaw, BMAD-METHOD, LangGraph/LangChain, CrewAI, OpenAI Agents SDK, Gemini CLI, AutoGen/AG2, Pydantic AI, Vercel AI SDK, LiteLLM, OpenRouter, and LM Studio (fully offline: local model, local ledger).
Production deployment - TLS termination, auth keys, redaction, image signature/SBOM verification, enterprise mirrors, and scaling guidance in docs/deployment.md.
The numbers
Measured, not promised. Reproduce them with
python scripts/loadtest.py --calls 2000 --seed 1000000 (Apple M-series
MacBook, SQLite backend; re-measured on 0.10 with the provider adapter
architecture in place - same numbers, 2,580 to 2,800 calls/s on both):
What | Result |
Sustained capture throughput | 2,886 proxied calls/sec |
Added latency per call | 10ms p50 · 12ms p95 |
Direct store writes | ~38,000 saves/sec |
One million calls on disk | 271 MB |
Open one session at 1M calls | 2 ms |
Session list at 1M calls | 335 ms |
30-day report at 1M calls | 719 ms |
The honest caveats: the session list aggregates every session on every
load, so it grows with total history; the report window uses a timestamp
index, so it grows with the window's traffic, not the table. Your agent's
provider latency (hundreds of ms per call) dwarfs the proxy's overhead by
an order of magnitude. Postgres numbers vary with your server; the same
script measures them with --dsn. Cost math has its own guardrails and
a five-minute parity check against your provider console: see
docs/accuracy.md.
What gets captured
Every LLM call is stored with:
Field | What it contains |
| UUID assigned at interception time |
| Run grouping (from header) |
| When the call was made |
| Model used |
|
|
| Full message history sent to the model |
| Extracted system prompt |
| Tool definitions available to the model |
| Tools the model decided to call |
| What the tools returned (from next call's messages) |
| Model's text output |
| Why the model stopped |
| Token usage |
| Prompt-cache usage - reads and writes are priced correctly per provider |
| Extended-thinking output (Anthropic), captured separately from |
| Estimated cost based on model pricing |
| End-to-end response time |
| HTTP status from upstream - errors are captured too |
| Upstream error message for non-200 responses |
| From |
| From |
| From |
| From |
| From |
| Parent call in a nested agent graph |
| Agent handoff tracking for the Flow DAG |
API reference
Method | Endpoint | Description |
|
| Liveness - |
|
| Readiness - pings the store; |
|
| Prometheus metrics (captures persisted/dropped, queue depth). |
|
| Audit trail of sensitive actions, newest first, hash-chained; filter by |
|
| Walk the audit hash chain and name the first break, if any (admin). |
|
| Right-to-erasure: delete all of a user's captured calls (admin). |
|
| Live dashboard |
|
| WebSocket stream - powers live dashboard updates |
|
| Sessions with aggregated stats, newest first. Page with |
|
| Loop runs (explicit or auto-inferred) with iterations, cost, status, and flagged-call counts, newest first. Same paging and headers; filters |
|
| One run's status ( |
|
| Stop all calls: the fleet-wide emergency stop. Every LLM call, replays included, is refused at the wall and recorded until lifted; survives a restart (editor). |
|
| Derived tool executions - each tool call paired with its result, latency, and error status |
|
| Whether the loop circuit breaker is holding a session, and why. |
|
| Delete a session and all its calls |
|
| Spend insights: daily trend, model mix with signed cache savings, latency percentiles, per-agent and per-team totals |
|
| Reprice a run/session/call's captured tokens on another model - pure math, zero API calls |
|
| Mint scoped API tokens - including |
|
| One call by id - follow a replay's parent back to its original |
|
| Configured replay destinations (feeds the dashboard's dropdown) |
|
| Models a replay target actually serves ( |
|
| Whether sign-in is configured here, the provider's name, and where it starts. No auth. |
|
| Start the identity-provider sign-in ( |
|
| End this sign-in and clear the cookie |
|
| Everyone who has signed in, with the role their groups grant and live sign-in count (admin) |
|
| End every sign-in of one person, now (admin) |
|
| A one-minute, single-use ticket for the live |
|
| What is the key I'm holding? Name, role, and team (for team cards) - the dashboard's ⚿ panel uses this |
|
| Replay a whole run or session on another model - returns a job id |
|
| Batch progress and the report card |
|
| Name, pin, or file a session/run under a project; mark it with an |
|
| Project names in use |
|
| What the proxy is running with (admin; secrets masked), each row with the config key that sets it and its choices |
|
| The config file in effect and its values (secrets masked); |
|
| Restart the ledger in place so the config file is read again (admin) |
|
| The repeat-discount this run was eligible for and received: verdict, reason, fix, exact discount, labeled estimate |
|
| Re-run framework detection over unattributed history; returns examined, updated, and per-framework counts |
|
| Re-execute a captured call - same provider or translated to the other one ( |
|
| Full-text search across all captured calls |
|
| All calls in a session, ordered by time |
|
| Single call by action ID |
|
| JSON compliance export with SHA-256 integrity hash |
|
| Printable HTML audit report |
|
| MCP tool server - |
|
| OTLP/HTTP JSON ingest - GenAI spans from OTel-native tools become ledger calls ( |
Examples:
# All calls in a session
curl http://localhost:8000/session/run-1
# Search across all sessions
curl "http://localhost:8000/api/search?q=failed+to+connect"
# Download JSON audit trail (includes an integrity tag; keyed HMAC when configured)
curl http://localhost:8000/export/run-1 -o audit-run-1.json
# Printable HTML report - open in browser, print to PDF
open http://localhost:8000/export/run-1/reportMCP server
Agentic Ledger exposes its captured data as an MCP (Model Context Protocol) tool server at POST /mcp. Point Claude Desktop, Cursor, or any MCP-compatible client at it to query traces directly from your AI assistant.
Tools available:
Tool | Description |
| List recent sessions with cost, token, and call count summaries |
| Full trace for a single LLM call - prompt, tool calls, output, tokens, cost |
| All calls in a session in chronological order |
| Full-text search across all captured calls |
| Loop runs with iterations, cost, and status |
| One run's status - lets an agent inspect its own loop and decide whether to continue |
Configure in claude_desktop_config.json (HTTP, against a running proxy):
{
"mcpServers": {
"agenticledger": {
"url": "http://localhost:8000/mcp"
}
}
}Or as a stdio subprocess - for clients that launch servers as commands (no running proxy required; reads the same database):
{
"mcpServers": {
"agenticledger": {
"command": "agenticledger",
"args": ["mcp"],
"env": { "AGENTICLEDGER_DSN": "sqlite:////absolute/path/to/agenticledger.db" }
}
}
}If AGENTICLEDGER_API_KEY is set, pass it as a header:
{
"mcpServers": {
"agenticledger": {
"url": "http://localhost:8000/mcp",
"headers": { "x-agenticledger-api-key": "your-key" }
}
}
}Once connected, you can ask your assistant things like:
"What did the SearchAgent do in the last session?"
"Show me all calls that mentioned rate limit errors"
"What was the total cost of session run-abc123?"
Configuration reference
Every variable below can also live in agenticledger.toml - see
Configuration for the file, the search order, and the
env-always-wins rule.
Environment variables
Core:
Variable | Required | Default | Description |
| No | (off) | Record Claude Code's own |
| No | (none) | Your company's Bedrock gateway. Bedrock-shaped calls are forwarded there exactly as the agent sent them, headers included; the ledger signs nothing and needs no AWS credentials. |
| No | (unset: route by call format) | LLM endpoint to forward requests to. Accepts OpenAI, Anthropic, LiteLLM, OpenRouter, or any OpenAI-compatible URL. Omit it and the proxy routes each call by its wire format: Anthropic-shaped calls to Anthropic, Bedrock paths to Bedrock, everything else to OpenAI. |
| No | (off) |
|
| No |
| Port for the https dashboard listener. |
| No |
| Database. SQLite for local dev, Postgres URL for production. |
| No |
| Host to bind to. Use |
| No |
| Port to run on. |
| No | (none) | Master admin key. When set, the dashboard, read, and management endpoints require authentication; the key grants the |
| No | (none) | Sign in with an identity provider (OpenID Connect, code flow with PKCE). |
| No | (none) |
|
| No |
| How long a sign-in lives: idle limit, and the absolute limit. |
| No | (none) | When set, the proxy forwards a request only if it carries a matching |
| No | (none) | Key for same-provider replay through the proxy's own upstream - the proxy never stores agent credentials, so re-execution needs its own. |
| No | (none) / provider API | Cross-provider replay target: replay any capture on OpenAI-format models. Point |
| No | (none) / provider API | Cross-provider replay target for Claude models. |
| No | (none) | Every key above also reads from a file named by its |
| No | (none) | When set, compliance exports carry a tamper-evident keyed |
| No | (none) | Comma-separated additional request paths to capture, e.g. |
| No |
| Persist captures on a background worker so storage never adds latency to the agent's call. Trade-off: reads become eventually consistent (a just-captured call may not be queryable for a brief moment). Recommended for high throughput. |
| No |
| Max captures buffered in async mode before load is shed (drops are counted in |
| No |
|
|
| No | (off) | Redact PII/secrets in stored data: |
| No | (none) | Extra redaction regexes as JSON: |
| No | (keep forever) | Delete captured calls older than N days via a background purge worker. |
| No |
| Record an audit trail of who viewed/exported/deleted what plus token/erasure actions, failed logins, rejected ingest credentials and MCP reads. Set |
| No |
| Refuse (HTTP 503) any audited action the log cannot record. Off, the failed write is counted ( |
| No | (none) | Key the audit hash chain with HMAC-SHA256. Without it the chain is plain SHA-256: it catches edits, but a writer with database access can re-chain. |
| No |
| Also print each audit row as one JSON line on stdout, for log scrapers and SIEM agents. With OTel export configured, rows are also sent as OTLP log records. |
Cost budgets - block calls that exceed a spend limit (returns HTTP 429):
Variable | Default | Description |
| (none) | Max USD per |
| (none) | Max USD per |
| (none) | Max USD total across all calls per calendar day (UTC). |
| (none) | Max USD per |
|
| HTTP status for budget blocks. |
|
| What happens when a budget is exceeded: |
|
| A model with no price in the packs cannot be counted. |
Budgets and run ceilings hold under concurrency. Each admitted call reserves an estimate (its text at four chars per token plus its max_tokens, priced like any call) until its real cost is recorded, so a burst of parallel calls cannot each pass the same remaining room. The single call that crosses the line still goes through, as one caller always did, so overshoot is bounded to one call's cost. A reservation is released the moment the call is recorded, fails, is dropped, or is refused.
Allow and deny lists - refuse a model or provider before any quota is spent (returns HTTP 403 with the rule named, so agents stop rather than retry). Patterns are shell globs, matched case-insensitively. Deny wins over allow; an allow list that exists admits only what it names. Team cards can carry the same four lists (see Team cards); the fleet lists always apply and a card can only narrow them.
Variable | Default | Description |
| (none) | Comma-separated model patterns, e.g. |
| (none) | Model patterns refused outright, e.g. |
| (none) | Provider names or patterns ( |
| (none) | Providers refused outright. |
Rate limits - block calls that exceed request frequency (returns HTTP 429, sliding 60-second window). Every refusal is recorded as an amber blocked: row with the reason, counted per session and team in Reports and in /metrics (agenticledger_refusals_total{reason=...}), so a retry storm is visible instead of vanishing:
Variable | Default | Description |
| (none) | Max requests per minute globally. |
| (none) | Max requests per minute per |
| (none) | Max requests per minute per |
| (none) | Max requests per minute per |
Loop engine - every call is stitched into ReAct threads (thread_id, step_index, prev_action_id) and fresh-context loop iterations are grouped into runs, with stuck-loop detection:
Variable | Default | Description |
|
|
|
|
| Consecutive identical tool calls (same tool, same arguments) before a thread is flagged stuck. |
| (none) | Flag (and in block mode, stop) threads that exceed this many ReAct steps. |
|
| Max gap between fresh-context spawns (same system prompt) that still count as iterations of one run. |
| (none) | Regex matched against response text. On match the call is flagged |
Alerts - POST to your webhook when a threshold is breached (does not block calls - see Alerts):
Variable | Default | Description |
| (none) | URL to POST alert payloads to. Required for any alerts to fire. Slack, Discord and PagerDuty URLs get their native shape. |
|
| Payload shape: |
| (none) | PagerDuty Events v2 integration key; needed when the webhook is |
| (none) | Where the dashboard is reachable ( |
| (off) | UTC hour (0-23) to POST a daily spend digest - last 24h totals, cache savings, top models/agents - to the alert webhook. Slack-incoming-webhook friendly ( |
| (none) | Alert when a single call costs more than |
| (none) | Alert when a single call takes longer than |
| (none) | Alert when session error rate exceeds |
| (none) | Alert when daily spend crosses |
OpenTelemetry - emit spans to any OTLP-compatible collector (requires pip install "agentic-ledger[otel]" - see OpenTelemetry export):
Variable | Default | Description |
| (none) | OTLP/HTTP base URL, e.g. |
|
| Value of |
| (none) | Comma-separated |
Pricing overrides - override or extend the built-in per-token pricing table (merged at startup):
Variable | Default | Description |
| Fetch the current price packs from the repository into | |
| (none) | Inline JSON map of model → |
| (none) | Path to a JSON file with the same format. Applied after |
Common startup examples
# Local dev - OpenAI (default)
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com uv run python -m agenticledger.proxy
# Local dev - Anthropic
AGENTICLEDGER_UPSTREAM_URL=https://api.anthropic.com uv run python -m agenticledger.proxy
# Local dev - LiteLLM gateway (any model)
AGENTICLEDGER_UPSTREAM_URL=http://localhost:4000 uv run python -m agenticledger.proxy
# Production - Postgres + auth + budgets + rate limits + alerts
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \
AGENTICLEDGER_DSN=postgresql://user:password@localhost/agenticledger \
AGENTICLEDGER_API_KEY=my-secret \
AGENTICLEDGER_BUDGET_DAILY=20.00 \
AGENTICLEDGER_BUDGET_SESSION=2.00 \
AGENTICLEDGER_RATE_LIMIT_SESSION_RPM=20 \
AGENTICLEDGER_RATE_LIMIT_USER_RPM=60 \
AGENTICLEDGER_ALERT_WEBHOOK_URL=https://hooks.slack.com/services/xxx/yyy/zzz \
AGENTICLEDGER_ALERT_COST_PER_CALL=0.50 \
AGENTICLEDGER_ALERT_DAILY_SPEND=15.00 \
uv run python -m agenticledger.proxyWhen AGENTICLEDGER_API_KEY is set, pass it in a header to access protected endpoints:
curl -H "x-agenticledger-api-key: my-secret" http://localhost:8000/session/run-1Keys travel in headers only. A key in a query string (?api_key=, ?token=) is refused with a 401 that says why: URLs end up in access logs, proxy logs, browser history and Referer headers. In a browser, paste the key into the dashboard's ⚿ panel, or open the pairing link from agenticledger share, which carries the key after the # (the URL fragment, which a browser never sends to any server).
Sign in with your identity provider (OpenID Connect)
For people, not scripts: point the ledger at your identity provider and the ⚿ panel gains a Sign in with Okta button (or whatever you name it). The code flow with PKCE, ID tokens verified against the provider's keys (RS256), and your groups decide the role: a person whose groups map to nothing is refused, told why, and recorded. Keys keep working beside it for agents and scripts.
AGENTICLEDGER_OIDC_ISSUER=https://your-org.okta.com
AGENTICLEDGER_OIDC_CLIENT_ID=0oa...
AGENTICLEDGER_OIDC_CLIENT_SECRET_FILE=/run/secrets/oidc # omit for a public client
AGENTICLEDGER_OIDC_ROLE_MAP=ledger-admins=admin,ledger-editors=editor,ledger-viewers=viewer
AGENTICLEDGER_PUBLIC_URL=https://ledger.example.com # the redirect back lands hereRegister https://ledger.example.com/auth/callback as the redirect URI with the provider. A sign-in is a server-side row the browser holds a cookie for (httponly, SameSite=Lax, Secure over https): it ends after 12 idle hours or 7 days (AGENTICLEDGER_SESSION_IDLE_HOURS, AGENTICLEDGER_SESSION_MAX_HOURS), on Sign out, or when an admin ends it (POST /api/people/{id}/signout). Mutating requests that ride a cookie must come from the dashboard's own origin. Every audit row names the person by email. GET /api/people lists who has signed in, with their role and groups.
Scoped access. AGENTICLEDGER_OIDC_SCOPE_MAP=team-alpha=alpha,team-alpha=alpha-infra makes anyone in team-alpha see exactly those projects: the lists, single sessions and runs, search, reports, exports, what-if, replay and the MCP tools all answer inside the scope, and anything outside it reads as not found. Work filed under no project is invisible to a scoped person until someone files it. A person in no mapped group is unscoped and sees everything their role allows, as before. Scoped editors can file work only under their own projects.
To try it without a provider: agenticledger idp runs a test provider on loopback with four fake people (alice is an admin, dave has no mapped group), prints the four lines to set, and says on every page that it is not for production.
Scoped API tokens (RBAC)
The master key is convenient but coarse. For team access, mint scoped, revocable tokens with roles instead of sharing the master secret. Tokens are random secrets shown once at creation; only their SHA-256 hash is stored.
Roles are hierarchical, with one exception: ingest sits outside the hierarchy and opens only the proxy path.
Role | Can |
| send calls through the proxy only (this is what a team card is): attributes each call to its team and carries the team's daily budget; cannot read the dashboard, API, export or MCP |
| read captured data - dashboard, API, export, MCP |
| viewer + delete sessions |
| editor + manage API tokens |
# Mint a viewer token (admin only - use the master key to bootstrap)
curl -X POST http://localhost:8000/api/tokens \
-H "x-agenticledger-api-key: my-secret" \
-H "content-type: application/json" \
-d '{"name": "grafana-readonly", "role": "viewer", "expires_in_days": 90}'
# → {"token_id": "...", "token": "agl_…", "role": "viewer", ...} (token shown once)
# Use it (Authorization: Bearer or the x-agenticledger-token header)
curl -H "Authorization: Bearer agl_…" http://localhost:8000/api/sessions
# List and revoke
curl -H "x-agenticledger-api-key: my-secret" http://localhost:8000/api/tokens
curl -X DELETE -H "x-agenticledger-api-key: my-secret" http://localhost:8000/api/tokens/<token_id>Auth is enforced only when
AGENTICLEDGER_API_KEYis set; the master key is the admin bootstrap for minting tokens. The live/wsfeed accepts a header (Authorization: Bearerorx-agenticledger-token) or a ticket: a browser cannot put a header on a websocket handshake, so the dashboard first callsPOST /api/ws/ticketwith its key in a header and connects with/ws?ticket=..., a random one-minute, single-use ticket that is worthless once used. Unauthenticated connects, and any connect with a key in its URL, are rejected with close code 1008.
Request headers
Pass these from your agent on each LLM call. All optional. They enrich captured data, power the Flow tab, and enable per-dimension budgets and rate limits.
Header | Default | Description |
| (none) | Groups all calls in a run. Use a consistent ID per agent execution (e.g. a UUID or |
| (none) | End user who triggered this run. Enables per-user rate limiting and auditing. |
| (none) | Name of the agent making this call (e.g. |
| (none) | Application name or ID. Useful when multiple apps share one proxy. |
| (none) | The |
|
|
|
| (none) | Agent handing off control (e.g. |
| (none) | Agent receiving control (e.g. |
| (auto-detected) | Framework/tool making the call (e.g. |
| (auto-inferred) | Groups sessions into a loop run (e.g. a Ralph overnight run). When absent, fresh-context sessions sharing a system prompt within |
| (auto-inferred) | Iteration number within the run. |
Single agent - fully annotated:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="your-openai-key",
default_headers={
"x-agenticledger-session-id": "run-abc123",
"x-agenticledger-user-id": "user-42",
"x-agenticledger-agent-name": "researcher",
"x-agenticledger-app-id": "my-app",
"x-agenticledger-environment": "production",
},
)Multi-agent system - tracking handoffs:
from openai import OpenAI
# Orchestrator
orchestrator_client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="your-openai-key",
default_headers={
"x-agenticledger-session-id": "run-abc123",
"x-agenticledger-agent-name": "orchestrator",
},
)
# Researcher (receives handoff from orchestrator)
researcher_client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="your-openai-key",
default_headers={
"x-agenticledger-session-id": "run-abc123",
"x-agenticledger-agent-name": "researcher",
"x-agenticledger-handoff-from": "orchestrator",
"x-agenticledger-handoff-to": "researcher",
},
)The Flow tab renders orchestrator → researcher as a DAG with cost and latency on each node.
OpenAI Agents SDK (openai-agents) - per-agent clients:
The openai-agents SDK uses its own internal OpenAI client. To pass Agentic Ledger headers you need to create a client per agent using OpenAIResponsesModel and set it as the agent's model.
import uuid
import os
from openai import AsyncOpenAI
from agents import Agent
from agents.models.openai_responses import OpenAIResponsesModel
SESSION_ID = f"run-{uuid.uuid4().hex[:8]}" # one per execution
BASE_URL = os.getenv("OPENAI_BASE_URL") # e.g. http://localhost:8000/v1
def al_model(agent_name: str, model: str = "gpt-4o-mini",
handoff_from: str | None = None, handoff_to: str | None = None):
"""Create a model instance that sends Agentic Ledger metadata headers."""
if not BASE_URL:
return model # proxy not configured - use default client
headers = {
"x-agenticledger-session-id": SESSION_ID,
"x-agenticledger-agent-name": agent_name,
}
if handoff_from:
headers["x-agenticledger-handoff-from"] = handoff_from
if handoff_to:
headers["x-agenticledger-handoff-to"] = handoff_to
client = AsyncOpenAI(base_url=BASE_URL, api_key=os.getenv("OPENAI_API_KEY", ""),
default_headers=headers)
return OpenAIResponsesModel(model=model, openai_client=client)
planner = Agent(name="PlannerAgent", model=al_model("PlannerAgent", handoff_to="SearchAgent"), ...)
searcher = Agent(name="SearchAgent", model=al_model("SearchAgent", handoff_from="PlannerAgent", handoff_to="WriterAgent"), ...)
writer = Agent(name="WriterAgent", model=al_model("WriterAgent", handoff_from="SearchAgent", handoff_to="EmailAgent"), ...)
emailer = Agent(name="EmailAgent", model=al_model("EmailAgent", handoff_from="WriterAgent"), ...)Each agent's calls are tagged with its name and pipeline position. The Flow tab renders the full PlannerAgent → SearchAgent → WriterAgent → EmailAgent DAG automatically.
Why per-agent clients?
set_default_openai_client()sets a single global client - fine for single-agent apps, but it can't carry differentagent_nameorhandoff_*headers per agent in a multi-agent system. Per-agentOpenAIResponsesModelinstances are the correct approach.
Alerts
Agentic Ledger posts to your webhook URL when a threshold is breached, a loop is flagged, a run hits a wall, or a run ends. Slack incoming webhooks, Discord webhooks and PagerDuty Events v2 are recognised from the URL and get their native shape (override with AGENTICLEDGER_ALERT_FORMAT); anything else gets the plain JSON below.
Reliable by construction. Every notification is tried three times with backoff (1s, 3s, 9s) off the request path, said once per window (a crossed daily budget once a day, a flagged loop once per ten minutes, a run hitting a wall once an hour), and recorded: the Settings page lists what was sent, whether it landed, after how many tries, and why not, with a Send a test notification button so wiring Slack takes one click. GET /api/notifications returns the same history. Set AGENTICLEDGER_PUBLIC_URL and every notification about a run or session carries a link straight to it.
Payload format (plain JSON; the native shapes carry the same facts):
{
"type": "high_cost",
"message": "Single call cost $0.1842 exceeded threshold $0.10",
"value": 0.1842,
"threshold": 0.10,
"action_id": "a1b2c3d4-...",
"session_id": "run-1",
"agent_name": "researcher",
"timestamp": "2026-04-03T12:00:00+00:00"
}Alert types:
Type | Triggered when |
| A single call exceeds |
| A single call takes longer than |
| Session error rate exceeds |
| Daily total spend crosses |
| A budget limit is hit and |
| A run's spend reaches 80% of its cost ceiling (fired once per run) |
| The loop engine raised flags on a call ( |
| A run's completion promise was seen - the payload carries the full run summary (iterations, cost, tokens, flagged calls) |
| A run went quiet (no calls for the run gap) - the same summary, so an overnight loop's end is in your channel by morning |
| A run went quiet and its last iteration ended in an error - the same summary, flagged as a failure |
| A run's call was refused at the wall (kill switch, cost ceiling, budget, stop all calls, a model or provider list) - once per run and reason per hour, with the reason |
| You pressed Send a test notification |
Team cards - one proxy, many teams
Think allowance cards: you keep the one real provider key, and hand each
team a card of its own. Each card opens the proxy, stamps every call with
the team's name, and can carry its own daily budget - when marketing hits
$10, only marketing gets blocked (with an honest Retry-After).
curl -X POST http://localhost:8000/api/tokens \
-H "x-agenticledger-api-key: $ADMIN_KEY" -H 'content-type: application/json' \
-d '{"name": "marketing", "role": "ingest", "budget_daily": 10.00}'The response shows the card once - the ledger stores only its hash. The
team puts it in x-agenticledger-ingest-key instead of the shared key;
Reports gains a by-team table with errors, blocks, and spend-today against
each card's allowance. Revoke a card with DELETE /api/tokens/{token_id}
and only that team is affected - from that instant the card gets a final
403 ("the answer is no"), which agents accept without retry storms.
Paste a card into the dashboard's ⚿ panel by mistake and it tells you, in
plain words, that cards open the relay, not the dashboard.
A card can also carry its own allow and deny lists for models and
providers, on top of the fleet-wide AGENTICLEDGER_ALLOW_MODELS and
friends. The fleet lists always apply; a card can only narrow them, never
grant a model the fleet denies. Refusals name the rule and the team, and
show up in the by-team table like any other block.
curl -X POST http://localhost:8000/api/tokens \
-H "x-agenticledger-api-key: $ADMIN_KEY" -H 'content-type: application/json' \
-d '{"name": "marketing", "role": "ingest", "budget_daily": 10.00,
"allow_models": ["gpt-4o", "claude-sonnet-*"], "deny_providers": ["bedrock"]}'Budgets vs alerts:
Budgets (
AGENTICLEDGER_BUDGET_*) - block the call before it reaches the LLM. Agent gets HTTP 429.Alerts (
AGENTICLEDGER_ALERT_*) - the call goes through, you get notified after.
What the webhook receives. Every alert is one JSON POST with our own field names: type (see the table below), message, value, threshold, action_id, session_id, agent_name, timestamp. The daily digest (AGENTICLEDGER_DIGEST_HOUR=8) is a separate POST with type: daily_digest, a ready-to-read text block (last-24h spend, cache savings, top models and agents), and totals.
Slack - paste an Incoming Webhook URL (hooks.slack.com): each notification arrives as a titled message with the detail and an "Open in Agentic Ledger" link.
PagerDuty - use the Events API v2 URL (https://events.pagerduty.com/v2/enqueue) and set AGENTICLEDGER_ALERT_PAGERDUTY_KEY to the integration key: notifications trigger incidents with a severity per type (critical for a failed or blocked run, warning for thresholds, info for summaries and digests), a dedup key, and a link to the run.
Discord - a channel webhook URL (discord.com/api/webhooks/...): a titled embed with the detail and the link.
Custom - any HTTP endpoint that accepts a JSON POST.
OpenTelemetry export
Agentic Ledger can emit every intercepted LLM call as an OTel span to any OTLP-compatible collector: Grafana Tempo, Jaeger, Honeycomb, Datadog, Dynatrace, or any vendor that supports OTLP/HTTP.
Install the extra (Docker image includes OTel - no extra step needed when using Docker):
pip install "agentic-ledger[otel]"
# or
uv add "agentic-ledger[otel]"Configure:
Variable | Default | Description |
| (none) | OTLP/HTTP base URL, e.g. |
|
| Value of |
| (none) | Comma-separated |
Example - Grafana Tempo:
AGENTICLEDGER_UPSTREAM_URL=https://api.openai.com \
AGENTICLEDGER_OTEL_ENDPOINT=http://localhost:4318 \
AGENTICLEDGER_OTEL_SERVICE_NAME=my-agent \
uv run python -m agenticledger.proxyExample - Honeycomb:
AGENTICLEDGER_OTEL_ENDPOINT=https://api.honeycomb.io \
AGENTICLEDGER_OTEL_HEADERS=x-honeycomb-team=YOUR_API_KEY,x-honeycomb-dataset=llm-traces \
uv run python -m agenticledger.proxySpan attributes emitted (GenAI semantic conventions):
Attribute | Source |
| Provider ( |
| Always |
| Model ID |
| If set |
| If set |
| Tokens in |
| Tokens out |
| Stop reason |
| Unique call ID |
| Run grouping |
| From header |
| From header |
| Estimated cost |
| End-to-end latency |
| From header |
| Agent handoffs |
| HTTP status from upstream |
Spans are grouped into traces by session_id - all calls in a session appear as one trace in your backend. Parent-child relationships follow x-agenticledger-parent-action-id. Error spans (status_code != 200) are marked with StatusCode.ERROR.
Compliance documents
For a security or privacy review: docs/compliance holds the data-flow diagram, a data-processing description with a DPA annex, the subprocessor statement (none: the software runs where you install it and sends the project nothing), the HIPAA posture, a SOC 2 and ISO 27001 control mapping with evidence for every row, and the support window (the latest minor gets every fix, the previous minor gets security fixes for 90 days). Written to be attached as they are, and honest about what the project does not hold: no SOC 2 report, no ISO certificate, no BAA.
Compliance export
Every session can be exported as an integrity-tagged audit trail - useful for regulated industries, internal audits, or passing traces to external tools.
# Machine-readable JSON with an integrity tag over the calls array
curl http://localhost:8000/export/run-1 -o audit-run-1.json
# Printable HTML - open in browser and print to PDF
open http://localhost:8000/export/run-1/reportThe JSON export carries an integrity tag over the calls array. By default this is a sha256 checksum - it catches accidental corruption but is not a signature (anyone who edits the calls can recompute it). Set AGENTICLEDGER_EXPORT_HMAC_KEY to switch to a keyed hmac-sha256 tag, which is tamper-evident: a recipient holding the key can detect any modification, and the tag cannot be forged without the key.
Releasing
Tagging a version triggers the full release pipeline automatically:
git tag v0.2.0
git push origin v0.2.0This runs three jobs:
Docker - builds
ghcr.io/shekharbhardwaj/agentic-ledger:{version}and:latestfor linux/amd64 + linux/arm64, pushes with SBOM + provenance attestations, signs the digest with Sigstore cosign (keyless), and mirrors to Docker Hub when theDOCKERHUB_USERNAME/DOCKERHUB_TOKENsecrets are configuredPyPI - builds and publishes
agentic-ledger=={version}to PyPI using trusted publishing (no API token needed), with PEP 740 attestationsGitHub Release - creates a release with auto-generated changelog and attaches the image SBOM (SPDX)
First-time PyPI setup (one time only):
Add a new pending publisher:
PyPI project name: agentic-ledger Owner: ShekharBhardwaj Repository: AgenticLedger Workflow name: release.yml Environment name: pypiCreate a
pypienvironment in GitHub: repo → Settings → Environments → New environment → name itpypiThat's it - no secrets needed
Troubleshooting
Start here: agenticledger doctor. One command prints the whole truth
of your machine: every install on PATH and who shadows whom, which Python
owns each one and whether it can actually run (wrong-architecture wheels
and missing dependencies caught by a real import probe), what the
background service is serving, and a fix-it command per finding. Most of
the problems below diagnose themselves with it. Add --fix and it applies the fixes it names: shadow installs evicted with their own interpreter's pip, the PATH prepend offered for cleanup, then a second diagnostic pass.
Old version / commands or env vars named agentledger (no "ic") -
you're running a pre-0.4 release, most likely from a venv that already had
the package installed: plain pip install agentic-ledger says "requirement
already satisfied" and does NOT upgrade. Run agenticledger upgrade - it
uses the Python environment that owns the install, so there's no guessing
which pip is the right one - then restart. The proxy prints its version on the first line at startup, and
curl localhost:8000/health reports it too. Since 0.4.0 everything is named
agenticledger - see the migration notes in the CHANGELOG.
Replay fails with 401 "invalid x-api-key" - AGENTICLEDGER_REPLAY_API_KEY
needs a real provider API key from console.anthropic.com
(or platform.openai.com). A Claude Code subscription login is not an API
key and cannot be used. No key? Replay for free against a local model - see
the LM Studio guide.
incompatible architecture (have 'arm64', need 'x86_64') on macOS - your
terminal is running under Rosetta, so Python picks its x86_64 slice while pip
installed arm64 native wheels. Check with arch (should print arm64 on
Apple Silicon). Quick fix: prefix the command with arch -arm64. Permanent
fix: uncheck "Open using Rosetta" on your terminal app, use an Apple Silicon
build of your editor, and restart any long-lived tmux server.
module 'httpx' has no attribute 'AsyncClient' - fixed in
0.3.0-alpha.2; upgrade with pip install --upgrade agentic-ledger.
Port 8000 already in use - another proxy instance (or app) is running;
stop it or set AGENTICLEDGER_PORT.
401 OAuth access token has expired from Claude Code - the proxy passed
Anthropic's answer through unmodified; re-authenticate with claude →
/login. Errored calls are still captured, so you'll see the 401 in the
dashboard.
/ answers 404 "Web app not built" - you're running from a source
checkout without the web-app build. cd dashboard-app && npm ci && npm run build and restart. PyPI and Docker installs always include the app.
License
MIT
mcp-name: io.github.ShekharBhardwaj/agentic-ledger
Available Tools
6 toolsexplainA
Retrieve the full captured trace for a single LLM call. Returns the prompt, system prompt, tool calls, model response, token usage, cost, and latency.
| Name | Required | Description | Default |
|---|---|---|---|
| action_id | Yes | The action ID from the x-agenticledger-action-id response header. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Though no annotations exist, the description correctly implies a read-only retrieval operation. It could be improved by explicitly stating idempotency or lack of side effects, but the current text is clear and not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: first states the action, second lists returned fields. No fluff, front-loaded, and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 1-parameter tool with no output schema, the description adequately lists return fields (prompt, cost, latency, etc.). Could add format or limit details, but overall sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for 'action_id'. The tool description adds no additional context to the parameter beyond what the schema already provides, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Retrieve the full captured trace') and the resource ('for a single LLM call'), and enumerates specific returned fields (prompt, system prompt, tool calls, etc.), making it distinct from sibling tools that handle sessions, runs, or search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus siblings like 'get_session' or 'search'. The context that it requires an 'action_id' from a header is implied but not compared to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_run_statusA
Status of one loop run: iterations so far, total cost and tokens, flagged calls, and whether the completion promise was seen (status=complete). Loop runners and agents can use this to decide whether to continue iterating.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run ID (from x-agenticledger-run-id or /api/runs). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It lists the exact data returned (iterations, cost, tokens, flagged calls, completion promise) and implies a read-only operation. No side effects or authentication needs are mentioned, but the simplicity of a status check makes this acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first lists the output content concisely, the second provides the use case. Every word adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains return values (iterations, cost, tokens, flagged calls, completion promise) and the tool's purpose. It could mention error behavior or format, but as a simple status tool it is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%—the only parameter (run_id) is well-described. The description adds no further parameter details beyond identifying the run, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns 'Status of one loop run' with specific data fields (iterations, cost, tokens, flagged calls, completion status). It distinguishes from siblings like list_runs and get_session by focusing on a single run's progress.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states that 'loop runners and agents can use this to decide whether to continue iterating', providing a clear use case. However, it does not mention when not to use it or direct alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sessionA
Retrieve the LLM calls of an agent session in chronological order. By default each call is a compact summary (index, model, status, error reason, tokens, cost, latency, tool names, sizes) — sessions can be megabytes, so full prompt/response bodies are returned only with include_messages=true, and single calls are better fetched via the explain tool using the action_id from a summary row.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | The session ID passed via x-agenticledger-session-id. | |
| include_messages | No | Return full message bodies for every call. Default false; can be very large. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the default summary format (fields like model, status, tokens, cost), the size implication ('sessions can be megabytes'), and the effect of include_messages. It doesn't explicitly state absence of side effects, but 'Retrieve' strongly implies a read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core purpose, then packs valuable caveats and alternatives without fluff. Every phrase earns its place, and the em-dash structure improves readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of an output schema, the description compensates by enumerating summary fields and the condition for full bodies. It also addresses size concerns and directs to explain for individual calls, making the tool's behavior clear enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by explaining why include_messages defaults to false (session size) and what the summary contains, which complements the schema's boolean description. It reinforces the session_id's source but doesn't introduce conflicting semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb-resource pair: 'Retrieve the LLM calls of an agent session' and adds 'in chronological order,' which clearly scopes the operation. It distinguishes from siblings by contrasting with the explain tool for single calls and implying list_sessions for session listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use a different tool: 'single calls are better fetched via the explain tool using the action_id from a summary row.' It also provides usage context for the include_messages parameter, warning about large payloads and explaining the default compact summary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsA
List loop runs (explicit x-agenticledger-run-id or auto-inferred fresh-context loops, e.g. Ralph overnight runs) with iterations, sessions, cost, flagged-call counts, and status (running / flagged / complete).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of runs to return (default 20, max 100). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns runs with specific fields (iterations, sessions, cost, flagged-call counts, status) and mentions possible statuses. However, it does not explicitly state that the operation is read-only, nor does it describe ordering, pagination behavior, or data freshness. The behavioral context is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the action ('List loop runs'). It includes necessary detail (explicit vs auto-inferred, example, returned fields) without excessive verbiage. The parenthetical adds context but could be slightly more streamlined. Overall, it is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one optional parameter and no output schema, the description explains the returned fields (iterations, sessions, cost, flagged-call counts, status) but does not specify the return structure (e.g., array of objects), pagination details (how 'limit' interacts with results, whether there is a next page), or how to interpret 'auto-inferred' runs. It is moderately complete but leaves gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the 'limit' parameter is fully described in the schema). The description does not add any additional meaning to the parameter beyond what the schema provides. It does not mention the parameter at all, so the parameter semantics rely entirely on the schema, which is sufficient for a basic optional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as listing loop runs, specifying the resource ('loop runs') and the action ('List'). It distinguishes from sibling tools like 'list_sessions' by focusing on runs and including run-specific details (iterations, sessions, cost, flagged-call counts, status). The parenthetical notes explicit vs auto-inferred runs, further clarifying scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context on when to use the tool (explicit run ID or auto-inferred fresh-context loops) and an example ('Ralph overnight runs'). However, it does not explicitly state when not to use it or mention alternatives among siblings, such as 'get_run_status' for detailed status of a single run. The usage context is clear but lacks exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsA
List recent agent sessions with aggregated stats — call count, total cost, token usage, and start time. Use this to find a session_id before calling get_session or to get a cost overview.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of sessions to return (default 20, max 100). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that it returns aggregated stats and lists recent sessions, but does not elaborate on behavioral traits like read-only nature, order, or pagination beyond the limit parameter. This is adequate for a simple list tool but could be more explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the action, and contains no unnecessary words. Every part adds value, earning a top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one optional parameter and no output schema. The description adequately explains the return content and common use case. It does not detail ordering or filtering beyond 'recent', but given simplicity, it is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with one parameter 'limit' having a clear description. The description adds no extra meaning beyond what the schema already provides, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'List recent agent sessions with aggregated stats' which is a specific verb and resource. It distinguishes from the sibling 'get_session' by noting its use for finding a session_id before calling that tool, thus providing differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use this to find a session_id before calling get_session or to get a cost overview,' giving clear guidance on when to use. However, it does not discuss when not to use or contrast with other siblings like list_runs or get_run_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Full-text search across all captured LLM calls. Searches prompts, outputs, system prompts, agent names, and user IDs. Use this to find calls related to a topic, error, or agent.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (default 20, max 100). | |
| query | Yes | Search term to look for across all captured calls. | |
| include_messages | No | Return full message bodies for each hit. Default false — hits are compact summaries with an action_id to drill in via explain. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses the search scope (all captured calls, specific fields), the return mode (compact summaries with action_id vs full bodies when include_messages is true), and references to 'explain' for deeper inspection. This is useful behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first states what the tool does, the second gives usage guidance. It is front-loaded with the core purpose and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema and annotations, the description provides enough to understand what the tool does, what it searches, and what results look like (compact summaries with action_id). It could mention sorting/pagination, but the limit parameter covers one aspect. Overall, it is complete for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of the parameters with descriptions, so the baseline is 3. The tool description itself does not add much parameter-specific meaning—it only reinforces the purpose. The 'include_messages' behavior is described in the schema, not the main description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the verb and resource: 'Full-text search across all captured LLM calls.' It enumerates the searched fields (prompts, outputs, system prompts, agent names, user IDs), which differentiates it from sibling tools that handle sessions or runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it: 'Use this to find calls related to a topic, error, or agent.' It also hints at an alternative for drill-in ('to drill in via explain') without naming the tool directly. This gives clear context for selection, though it doesn't explicitly mention when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.8.0- Changed
get_session1 field changed- added
Input schema / properties / include_messagesAdded value: +{ + "description": "Return full message bodies for every call. Default false; can be very large.", + "type": "boolean" +}
- Changed
search1 field changed- added
Input schema / properties / include_messagesAdded value: +{ + "description": "Return full message bodies for each hit. Default false — hits are compact summaries with an action_id to drill in via explain.", + "type": "boolean" +}
TDQS
Scored across 6 tools
Each tool targets a distinct resource or action: session lists, session details, single-call traces, full-text search, run lists, and run status. Descriptions explicitly cross-reference when to use which, eliminating ambiguity.
Most tools follow a clear verb_noun pattern (list_sessions, get_session, list_runs, get_run_status), but 'explain' and 'search' deviate as bare verbs, breaking the pattern slightly.
6 tools is well-scoped for an inspection-oriented server, covering session-level and call-level views without unnecessary bloat or gaps.
The set provides a complete inspection workflow: discover sessions/runs, drill into session summaries, zoom into individual calls, and search across all captured data. No obvious missing operations for a read-only ledger.
Maintenance
Related MCP Connectors
See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.
Personal finance ledger for AI agents — query spending, track bills, forecast cash flow.
Log, query, and edit expenses, budgets, and accounts in Ledgy from any MCP-compatible AI assistant.
A read-only verified record of agent-operable GTM tools: search, fetch, compare, track changes.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables AI assistants to query and analyze AI agent sessions from observability providers like Shepherd (AIOBS) and Langfuse, allowing users to debug agent runs, compare sessions, track performance, and analyze LLM usage patterns.18MIT
- AlicenseNot gradedqualityCmaintenanceEnables agent settlement, trust verification, and ledger operations for multi-agent workflows, with tools for blueprint management, credit tracking, and provenance recording.1MIT
- FlicenseNot gradedqualityDmaintenanceRecords all MCP tool interactions in a centralized ledger, enabling developers to trace, replay, inspect, and audit AI agent workflows.-
- AlicenseAqualityBmaintenanceMCP server for the nagi-ledger audit ledger and guardrail toolkit, exposing tools to record and annotate AI agent actions, subagent dispatches, and known dead ends, plus query session reports and statistics.8MIT