Skip to main content
Glama
reliai-in

io.github.loopg/pramana-mcp

by reliai-in

Pramana MCP server

mcp-name: io.github.loopg/pramana-mcp

Decide whether a prompt or model change is safe to ship, by running it against your own production runs.

Pramana is a flight recorder for AI agents. It records every non-deterministic decision your agent makes in production, replays any past run exactly against your changed code, and reports which decisions moved — not which sentences got reworded.

This MCP server lets an assistant read what Pramana recorded, and check a signed evidence bundle offline. It is read-only. There is deliberately no tool that starts a replay or a comparison: a sandboxed batch calls the model for real on every trace, so it spends money, and an MCP tool is a button any assistant can press without a person deciding. Producing a comparison stays a CLI verb behind its own confirmation gate.

Install

Nothing to install — uvx fetches it on demand.

Claude Desktop

~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "pramana": {
      "command": "uvx",
      "args": ["pramana-mcp"],
      "env": {
        "PRAMANA_API_KEY": "<your-engineer-or-admin-api-key>"
      }
    }
  }
}

Claude Code

$ claude mcp add pramana --env PRAMANA_API_KEY=<your-engineer-or-admin-api-key> -- uvx pramana-mcp

Create a key in Settings → API keys at reliai.in. Use an engineer or admin key — an auditor key deliberately never receives recorded prompts or responses, so get_trace would come back empty. Self-hosting? Set PRAMANA_API_URL to your own API.

Related MCP server: Kryve Agent Evaluation MCP

Tools

list_traces

List recorded AI agent runs (traces) captured by Pramana, optionally filtered by agent, status or time range.

"What agent runs did we record last Tuesday?"

{ "agent": "refund-agent", "since": "2026-09-22T00:00:00Z", "limit": 20 }

get_trace

Get one recorded agent run: its steps, model calls, tool calls with arguments, and outcome.

"Walk me through what the agent did in run loan-07 — which tools did it call, and with what?"

{ "trace_id": "loan-07-5a6eef" }

Long payloads are truncated, and the response says so rather than looking complete.

list_model_diffs

List comparison runs that checked a prompt or model change against recorded production runs.

"Have we compared anything against production since the model upgrade?"

{ "since": "2026-09-20T00:00:00Z", "limit": 20 }

get_model_diff

Get the findings of one comparison run: which decisions changed behaviourally, which were cosmetic, and which runs halted.

"What changed when we moved to the new model last Tuesday?"

{ "diff_id": "mdr_9f2c1a", "trace_id": "loan-07-5a6eef" }

Returns bucket counts first, then root findings, then consequent ones — never merged. One root cause that produced forty downstream differences is one thing to investigate, not forty. Passing trace_id is optional but much faster (see What the API could not do below).

verify_bundle

Verify a Pramana signed evidence bundle offline, with no account and no network access.

"Here's the evidence bundle the vendor sent. Is it intact?"

{ "bundle_path": "./bundle.json", "public_key_path": "./pramana-public-key.txt" }

public_key_path is required. The public key shipped inside a bundle is never used to check that same bundle — that would defeat the point.

A pass proves the bundle has not been altered since it was signed. It does not prove that what was captured was everything that happened. This is a proof of record, not a judgement of conduct.

What the API could not do

Two tools are worse than they should be, and it is the API's fault rather than a design choice:

  • There is no account-wide endpoint for comparison runs. They are only reachable per trace (/v1/traces/{id}/model-diffs), so list_model_diffs and a get_model_diff without trace_id walk the 20 most recent traces. A comparison against an older trace will not be found. Every response says so in a notes field rather than quietly returning a short list. The bound is 20 rather than something larger because the walk is that many sequential round trips: measured against the production API, 50 took 6.0s and 20 takes 2.4s.

  • /v1/traces returns agent_count, not agent ids, so filtering by agent means opening each trace. That filter is applied to the 20 most recent traces only, and says so.

Limitations

These are the same limitations as the SDK. They are not softened for a registry listing.

  • You cannot import existing conversation logs. Replay needs the execution trace, and a transcript does not carry one. Your corpus starts the day you instrument.

  • It reports that a decision changed, never whether the change is good. Judging a changed decision is your call.

  • Sandboxing only covers tool calls you wrapped. In a sandboxed batch the tool calls you have wrapped are served from the recording and never executed, and the model is called for real because the new model is the thing you are testing. In a plain replay, nothing leaves the process at all: outbound network access is blocked at the process level, below whatever HTTP client you use. A tool call you did not wrap is invisible to Pramana and will execute normally, once per trace — the CLI prints how many wrapped tool call sites it found before a batch runs, and what that means multiplied across the batch.

  • Python only. No JavaScript or TypeScript SDK.

  • The OpenAI and Anthropic clients are supported. Azure OpenAI and AnthropicBedrock are routed to those adapters and covered by tests; neither has been exercised against live cloud credentials. Raw boto3 is not supported.

  • LangChain is tested. LangGraph, CrewAI and LlamaIndex are not, and are not supported.

  • No SOC 2, no penetration test, no uptime SLA, no on-premise deployment.

  • The evidence public key is not yet published at a stable URL. verify_bundle runs offline and needs no account, but today you still obtain the key from us — which is not the same as independent verification, and is not claimed as such.

Sample evidence bundles

evidence-samples/ holds two signed bundles, a tampered copy of each, and the key that checks them. No account, no network:

$ pip install pramana-verify
$ pramana-verify evidence-samples/production-run.json \
    --pubkey evidence-samples/sample-public-key.txt
OK — 2 event(s), merkle_root=4fd442ccbf515c6018c0d7908bd1b1f9853df86894e67e8c40b5d1590f48418d

$ pramana-verify evidence-samples/production-run.TAMPERED.json \
    --pubkey evidence-samples/sample-public-key.txt
TAMPERED / INVALID:
  - payload …: content does not hash to its reference — this recorded prompt or response was tampered with

evidence-samples/FORMAT.md is the field-by-field specification. Those files are signed with a sample key, not the key that signs real bundles.


Documentation · reliai.in · pramana-sdk · pramana-verify

Available Tools

5 tools
get_model_diffB

Get the findings of one comparison run: which decisions changed behaviourally, which were cosmetic, and which runs halted.

ParametersJSON Schema
NameRequiredDescriptionDefault
diff_idYes
trace_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It does disclose the substantive content of the result (behavioural vs cosmetic changes, halted runs), which is real signal beyond the schema. However, it says nothing about read-only nature, error behaviour for an unknown diff_id, or what the optional trace_id scopes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a compact colon list; no filler. It is slightly dense at the tail but every clause adds information about the return content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return structure need not be restated, and the description adequately previews the result categories. The substantial gap is parameter handling: both diff_id and trace_id are undocumented in schema and description alike, which is not adequate for a 2-param tool with 0% coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never references either parameter. It does not explain what a diff_id is or how to obtain one, nor what the optional trace_id does (filter? scope the run?), leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (findings of one comparison run), then enumerates the three categories of findings returned. This distinguishes it from list_model_diffs, which enumerates runs rather than reporting on one, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of prerequisites, and no routing to alternatives such as list_model_diffs. The agent must infer that a diff_id is obtained by first calling list_model_diffs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_traceA

Get one recorded agent run: its steps, model calls, tool calls with arguments, and outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
trace_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the shape of the returned data, but says nothing about permissions/auth, whether the operation is read-only, or behavior for missing traces. 'Get' implies a read, but that must be inferred.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler: verb, resource, and the returned payload are stated in one pass. Nothing restates the name or wastes tokens.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description need not explain return values, though it does list them helpfully. For a one-parameter read tool this is nearly sufficient; the only gap is the absence of any error/edge-case or auth note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter trace_id has no schema-level documentation, so the description must compensate. It implies trace_id identifies a recorded agent run but adds no format, source, or lookup guidance beyond that implication.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (one recorded agent run) and enumerates the payload an agent receives: steps, model calls, tool calls with arguments, and outcome. It differentiates itself from the plural sibling list_traces only by the word 'one', so differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'one recorded agent run' combined with a required trace_id suggests this is the fetch-by-identifier counterpart to list_traces. There is no explicit when-to-use, when-not-to-use, or named alternative, and no note on what happens when the trace_id is unknown.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_model_diffsC

List comparison runs that checked a prompt or model change against recorded production runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sinceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'List' implies a safe read, but nothing is said about pagination, ordering, default limits, or result volume. It adds domain framing of what a comparison run is, but no operational behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource front-loaded and no filler. The only minor cost is that 'comparison runs' does not echo the tool's own 'model diffs' terminology, requiring the agent to map the two.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format need not be explained, and this is a simple two-parameter list tool. Still, with zero annotation coverage and zero parameter coverage, an agent lacks the 'since' format and limit behavior needed to call it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions 'limit' or 'since'. Whether 'since' is a timestamp, ISO date, or run id is left entirely to guesswork, so the description does not compensate for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('comparison runs') plus the domain meaning of those runs (prompt or model change checked against recorded production runs). It implicitly contrasts with the singular sibling get_model_diff, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this list versus the sibling get_model_diff or list_traces, no prerequisites, and no exclusion conditions. The agent must infer routing purely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tracesB

List recorded AI agent runs (traces) captured by Pramana, optionally filtered by agent, status or time range.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNo
limitNo
sinceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden; 'List' implies a non-mutating read, which is helpful. However, it discloses nothing about pagination, result ordering, default limit behavior, or any authorization requirements, leaving meaningful gaps for a list-style tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the verb and resource come first and the filtering note follows. Efficient, though it could afford one more clause to cover the limit parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and there are zero required parameters. Still, the description is incomplete: it references a nonexistent 'status' filter and never mentions the limit/pagination control, so an agent cannot fully parameterize the call from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does name the agent and time-range (since) filters, but omits the 'limit' parameter entirely and cites a 'status' filter that does not exist in the schema — a mismatch that could mislead parameterization.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb plus resource ('List recorded AI agent runs (traces)') and clarifies that these are Pramana-captured traces. The plural 'traces' implicitly distinguishes it from the sibling get_trace, though it never names that alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It indicates the tool filters by 'agent, status or time range,' which implies the retrieval use case, but gives no explicit when-to-use versus get_trace or any other sibling. There is also no statement of prerequisites or when-not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_bundleB

Verify a Pramana signed evidence bundle offline, with no account and no network access.

ParametersJSON Schema
NameRequiredDescriptionDefault
bundle_pathYes
public_key_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two meaningful behavioral traits: the operation is fully offline and requires no account or network. It stops short of explaining what a failed verification does (error vs. negative result), whether it validates signature only or the full chain, or any resource constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and qualifies it with the two operationally important constraints. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required, and the offline/no-account context is a useful addition. The gap is on the input side: with zero schema coverage, the description leaves the caller without any notion of what the two paths should point to.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both required parameters (bundle_path, public_key_path) are bare strings with no titles beyond the property name. The description mentions a 'signed evidence bundle' but never explains that these are filesystem paths, what format the bundle or key should be in, or how they relate to each other.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (verify) and resource (a Pramana signed evidence bundle), so there is no ambiguity about what the tool does. It is clearly distinct from the trace/diff listing siblings, though it does not explicitly contrast itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'offline, with no account and no network access' implies the conditions under which this tool is appropriate (air-gapped or credential-free verification). However, it never states when to prefer this over other verification paths or what preconditions the caller must satisfy beyond having the two files.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.3
    • First observedget_model_diff
    • First observedget_trace
    • First observedlist_model_diffs
    • First observedlist_traces
    • First observedverify_bundle

TDQS

A3.5/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: two list/get pairs for traces and model diffs, plus an independent offline verification tool. There is no overlap between listing, retrieving, and verifying operations, so misselection is unlikely.

Naming Consistency5/5

All five names follow a consistent snake_case verb_noun pattern (list_traces, get_trace, list_model_diffs, get_model_diff, verify_bundle). The list/get pairing is symmetric and predictable across both resources.

Tool Count5/5

Five tools is well-scoped for an observability/inspection server with two resources plus a verification utility. Each tool earns its place with no redundant entries.

Completeness4/5

The list/get pairs cover discovery and detail for traces and diffs, and verify_bundle handles the evidence-checking lifecycle, so read workflows are complete. Minor gaps exist (e.g. no trace export or diff-creation trigger), but these are workable given the inspection-focused scope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers