Skip to main content
Glama

Verified Burst

Pay-per-correct inference for agents. At a hard, irreversible, or low-confidence decision, an agent escalates to fast silicon (Cerebras), samples best-of-N, runs the answer through a verifier, and settles over x402 only if it passes. A non-verified result costs nothing.

Live on Base mainnet with a self-hosted facilitator (no Coinbase dependency). The buyer brings their own provider key (BYOK) — we sell routing, verification, and settlement, never marked-up tokens.

Why

An agent can already sample itself more (its own best-of-N is correlated — it shares its own blind spots). The one thing it can't self-supply is an independent, zero-downside, keepable verification: a different model family checks the answer, you pay only if it passes, and you keep the receipt. Because the downside is $0 by construction, an agent's budget policy can auto-approve the spend without a human in the loop.

Related MCP server: coinopai-mcp

Quickstart (MCP, one line)

pip install verified-burst

Add the server to your MCP client — the agent gains one tool, buy_verified_burst:

{
  "mcpServers": {
    "verified-burst": {
      "command": "verified-burst",
      "env": {
        "BURST_BUYER_KEY": "0x<wallet-private-key-that-pays-per-call>",
        "BURST_PROVIDER_KEY": "csk-<your-cerebras-key>"
      }
    }
  }
}
  • BURST_BUYER_KEY — the wallet that pays per verified burst (USDC on Base). Required for live settlement.

  • BURST_PROVIDER_KEY — your Cerebras key; bursts run on your tokens (BYOK). Optional.

  • BURST_ENDPOINT — defaults to the hosted broker; override to self-host.

The tool:

buy_verified_burst(request, strategy="best_of_n", n=3,
                   verifier="self_consistency", answer_key=None)
  -> { answer, verified, charged, receipt, settle_tx }

How it works

request → escalate to fast silicon → best-of-N → verify → settle ONLY if passed
                                                              ↑ pay-only-if-verified
  1. 402 challengeGET /v1/burst returns the x402 payment requirements (Base mainnet, USDC).

  2. Authorize — the client signs an x402 (EIP-3009) authorization for the quoted price.

  3. Burst — the answer is generated on the buyer's BYOK key and gated through the chosen verifier.

  4. Settle — USDC is captured only if the verifier passes. A miss settles nothing.

Every response includes a keepable receipt (verified, corrected, independent, generator/verifier model, settle_tx) so verified decisions compound into memory.

Verifiers

verifier

what it does

cost to you

self_consistency

best-of-N must agree

free (BYOK)

judge

an adversarial judge checks the answer

free (BYOK)

independent_judge

a different model family judges (decorrelated errors)

small

independent_quorum

k-of-M independent judges across vendors

small

Independence is the moat: independent_judge/independent_quorum are judged on models in a different family (and, where configured, a different vendor) than the generator, so the check doesn't share the generator's blind spots.

Pricing

A small service fee per burst (routing + verification + the pay-only-if-verified guarantee), paid in USDC on Base — only on a pass. Per-burst $0.002 (fast) to $0.0045 (independent-judge). Generation tokens are billed to the buyer's own key; we never mark up tokens. Hard per-wallet spend cap; abuse breakers on the judge path.

Proof

This is live and settling real money. End-to-end mainnet settlement proof on BaseScan: 0x76921c33…d359c4.

A reproducible catch-rate harness (proof_harness.pyPROOF.md) generates code-checkable items and runs the real product path. At N=120 the base model was 86.7% accurate; the independent judge caught 16/16 mistakes with 0 false-confirms (never charged for a wrong answer) and a 3.8% false-alarm rate (free redo).

Honest caveat: that result is the checkable-answer regime (arithmetic, counting, labels) — the verifier's strongest suit and the product's stated sweet spot. It does not prove catch rate on fuzzy/subjective decisions, and 16 mistakes is a modest denominator. Claims are kept to what's measured.

HTTP API

route

purpose

GET /v1/info

discovery manifest (capabilities, pricing, verifier enum)

GET /v1/burst

x402 challenge for the paid resource (also advertised in WWW-Authenticate / PAYMENT-REQUIRED headers)

POST /v1/burst

buy a verified burst (send X-PAYMENT; BYOK via X-Provider-Key)

GET /v1/quote

price a burst without buying

GET /healthz

liveness

Self-hosting

The broker is stdlib-Python plus the live-payment deps in requirements-live.txt. Run PORT=8402 python3 server.py behind a TLS proxy; set X402_MODE=live with a relayer wallet and X402_PAY_TO. The MCP stdio server is mcp_server.py (sim mode needs no secrets — it starts and introspects out of the box).


Built honestly: only what's measured is claimed. Questions or integrations welcome.

Available Tools

1 tool
buy_verified_burstA

Buy a verified inference burst at a hard/irreversible/low-confidence decision. Escalates to fast silicon, samples best-of-N, gates the answer through a verifier, and charges (x402) ONLY if it passes. Returns the verified answer + a receipt. Budget-capped per agent. Use when getting it wrong is costly.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNobest-of-N sample count.
requestYesThe decision/question to resolve.
strategyNobest_of_n
verifierNoself_consistency
answer_keyNoOptional ["json","<field>"] or ["regex","<pat>"] to normalize answers.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries the burden. It reveals key behaviors: escalation to fast silicon, best-of-N sampling, verification gating, conditional charging 'ONLY if it passes', returns verified answer + receipt, and budget cap. No contradictions or missing critical traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is highly concise—four sentences, each earning its place. Front-loaded with the main purpose, it efficiently covers purpose, process, output, and usage context without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, no output schema, no annotations), the description provides a solid mental model: it explains the flow and output. However, it omits details about the answer_key parameter's role in normalization and does not explain the trade-offs between different strategy/verifier options. Still, it is largely complete for the core functionality.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 60% of parameters (request, n, answer_key have descriptions; strategy and verifier have enums but no descriptions). The description adds minimal parameter-specific meaning beyond the schema (e.g., 'samples best-of-N' relates to n, 'gates through verifier' relates to verifier). It does not significantly enhance understanding of the parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Buy a verified inference burst at a hard/irreversible/low-confidence decision.' It uses specific verbs ('buy', 'escalates', 'samples') and resource ('verified inference burst'), making it unambiguous. Although no siblings are provided, the description is sufficiently unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises when to use: 'Use when getting it wrong is costly.' This gives clear context. However, it does not mention when not to use or list alternatives, which would be ideal but is not required given the lack of sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedbuy_verified_burst

TDQS

A4.4/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion or ambiguity between tools. The tool's purpose is singular and clearly defined.

Naming Consistency5/5

The single tool 'buy_verified_burst' follows a consistent verb_noun pattern that aligns with the server name and clearly conveys its action.

Tool Count3/5

The server has only one tool, which feels minimal; however, for a very narrow, specific service, this count may be acceptable but is at the lower boundary of reasonable scope.

Completeness5/5

The tool 'buy_verified_burst' encapsulates the entire intended workflow (buy, escalate, verify, charge) for verified inference bursts, leaving no obvious gaps for its stated purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers