Skip to main content
Glama
5uper0

parley-mcp

by 5uper0

parley.

The trust layer for the agent economy

Provable consensus for AI agents of rival owners, deterministic red lines, max-min consensus over masked verdicts, and a verifiable non-betrayal transcript. Multilateral (N>2) · non-crypto · self-hosted · Apache-2.0.

ci license python core PRs welcome

Try it live · See it in 30 seconds · Roadmap · Changelog · Security · Contributing

Rival parties reach a max-min decision · a shortcut that would cheat someone is BLOCKED by a red line · every party verifies the tamper-evident receipt.


Install

pip install parley-consensus              # zero-dependency core
pip install "parley-consensus[crypto]"    # + Ed25519-signed verdicts (parley.net.identity, parley.net.bot)

The distribution is parley-consensus (PyPI's parley is an unrelated project); the import name is parley.

from parley.agent import Agent
from parley.consensus import run_consensus
from parley.preferences import HardConstraint, PreferenceSheet

options = [{"venue": "rooftop", "cost": 900}, {"venue": "garden", "cost": 600}, {"venue": "diner", "cost": 300}]
ana = PreferenceSheet("Ana", utility=lambda o: 1.0 if o["venue"] == "rooftop" else 0.4)
bob = PreferenceSheet("Bob", hard=[HardConstraint("budget", lambda o: o["cost"] <= 700)],  # a red line: code, not a preference
                      utility=lambda o: 1 - o["cost"] / 1000)
cara = PreferenceSheet("Cara", utility=lambda o: 0.9 if o["venue"] == "garden" else 0.5)

result = run_consensus([Agent(s.owner, s) for s in (ana, bob, cara)], options)
print(result.status, result.decision)
print("receipt sha256:", result.transcript.hash())
for sheet in (ana, bob, cara):  # each owner replays their own private sheet, locally
    print(sheet.owner, "not betrayed:", result.transcript.verify_non_betrayal(sheet, result.decision))
agreed {'venue': 'garden', 'cost': 600}
receipt sha256: a9cb75c0a15b41360fb2b134e084ca3b577daaf50f5edb657c5c4da43cc2fc72
Ana not betrayed: True
Bob not betrayed: True
Cara not betrayed: True

The rooftop is Ana's favourite and loses anyway: it crosses Bob's budget red line, so it is rejected before any score is weighed. The coordinator only ever saw red-line, never the budget or the sheet.

Contributors install from a clone instead: pip install -e ".[dev]" (see Status).

Use it from an MCP host (Claude, Cursor)

The base install ships a stdlib-only MCP server over stdio. Point your host at it:

{"mcpServers": {"parley": {"command": "parley-mcp"}}}

If the console script is not on the host's PATH, launch it through the interpreter that has the package: {"command": "python", "args": ["-m", "parley.mcp"]}.

Three tools appear: parley_decide (options + each party's red lines and preferences in, status, decision, transcript and transcript_sha256 out), parley_verify_receipt (recompute the hash over a transcript, and recompute the max-min decision from its recorded verdicts) and parley_check_party (replay one party's red lines against a decision). The input is the same JSON as the recipes in examples/demo/. Example prompt:

Three of us are picking a venue: rooftop (900), garden (600), diner (300). Bob will not go over 700; Ana prefers the rooftop, Cara the garden, Bob the cheapest. Use parley_decide, then verify the receipt hash and check that Bob's red line held.

What this mode does and does not give you. The host holds every party's spec and passes them all in one call, so there is no privacy between parties or from the host. What still holds is what the engine enforces in code: a red line rejects an option deterministically, the max-min rule picks among options feasible for everyone, and the receipt hash makes any later edit to the transcript visible, provided the hash is kept by someone other than whoever might edit the transcript: it is unsigned and the server does not store it, so a hash and a transcript from the same hand prove nothing about each other. max_min_verified does not depend on the hash: it recomputes the winner from the recorded verdicts. For private sheets, run one process per owner over HTTP (examples/run_env.py); the coordinator then sees only masked verdicts.


Most agent tooling in 2026 solves either transport/identity (A2A, MCP) or 1:1 agentic commerce (an agent buys/books for you), or cooperative debate between agents of the same owner. Parley targets the gap nobody productised: a group of delegates, each representing a different principal with conflicting and private interests, reaching a decision everyone can trust, and prove their agent didn't betray them.

Who it's for. Teams and platforms where several parties must reach a decision none of them can rig, and prove it afterward. Built first for regulated, multi-party workflows: compliance & onboarding decisions (KYC/AML, sanctions red lines), marketplace & P2P disputes, and delegated governance (committees, panels).

The wedge is three properties, enforced in code rather than left to an LLM's discretion:

  1. Deterministic red lines. Each owner's hard constraints are predicates checked in code. A proposal that crosses one is rejected, never negotiated away. "My agent won't betray me" becomes a provable property, not a hope.

  2. Conflicting-interest consensus. The coordinator sees only masked verdicts, a feasibility flag, a soft score, and a masked reason (ok/red-line), never the private sheets or which constraint was at stake. A decision must be feasible for every agent; among those it picks by an egalitarian max-min rule (lift the least-happy participant), tie-broken by total welfare, social choice, not majority vote. No feasible option ⇒ honest deadlock.

  3. Verifiable transcript. A tamper-evident SHA-256 record of every masked verdict. Each owner can replay their own private sheet locally to prove no red line was crossed, without revealing the sheet to anyone.

Related MCP server: mcp-witness

Why this exists (and what we learned building it)

Agents are starting to act on our behalf: booking, negotiating, onboarding, allocating. The moment two agents serve different owners, "did my agent sell me out?" stops being paranoia and becomes a real question. The crypto answer to agent-trust (put it on a chain, add a token) collapsed 89–99.7%; the durable need is the trust primitive itself, decoupled from tokens.

AI that can't lie, and you can check. The 2026 agent-hype cycle proved virality can be fully decoupled from truth. Parley is the provable version: every verdict in a parley is a re-hashable receipt, and each owner replays their own private sheet to prove no red line was crossed, no trust required, just the check.

Building the v0 taught us the wedge is narrower and sharper than "multi-agent consensus": the value isn't the vote, voting is free (Snapshot, polls). The value is provable non-betrayal under conflicting private interests, that no party had to reveal their private sheet or which red line was at stake, no red line could be traded away, and everyone can check the receipt themselves. That's what doesn't compose from off-the-shelf parts, and it's the one thing we made enforceable in code rather than left to an LLM's goodwill. The worked proof: docs/dogfood-01-p2p-escrow.md.

Status: v0 (working core)

Pure-stdlib, zero-dependency core. In-process transport with a clean seam where A2A (distributed discovery + signed Agent Cards) drops in next. LLM-elicited preference sheets (parley/elicit.py) and Ed25519-signed transcripts (parley/net/identity.py, optional crypto extra) already ship. This v0 deliberately isolates the novel part, the consensus + non-betrayal, and reuses nothing that's already a commodity.

python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q                     # the full suite, all green
.venv/bin/python examples/demo/server.py   # the money-shot: open http://127.0.0.1:8080 → Run
.venv/bin/python examples/meeting.py    # three delegates pick a meeting slot
.venv/bin/python examples/run_env.py    # bots as separate processes, consensus over HTTP
.venv/bin/python examples/real_decision.py --options examples/demo/recipe_committee.json
                                        # run a real decision with real people, each one ratifies

See it in 30 seconds

examples/demo/server.py runs the real engine behind one web page: rival parties, a max-min decision, one option BLOCKED by a red line, and a re-hashable receipt each party verifies privately. A worked example, a P2P escrow dispute end-to-end, with the per-party "this is better because…", is in docs/dogfood-01-p2p-escrow.md (synthetic). A real, anonymized dispute, a retrospective early-lease-exit replay where the tenant ratified a "this would have been better" outcome, is in docs/dogfood-02-early-lease-exit.md (honest boundary: the landlord side is inferred, not ratified). The screenshot-native proof card is examples/demo/proofcard_p2p.html.

No Python? Run the demo with Docker:

docker build -t parley . && docker run --rm -p 8080:8080 parley   # open http://127.0.0.1:8080

Try it live: parleyprotocol.com/demo runs the real engine behind one web page, no install and nothing to sign up for. Prefer to self-host? The repo ships a Dockerfile (and a render.yaml Blueprint) so you can run the same demo anywhere in one command.

Built by an agent fleet

The product is trust between agents of different owners. It was made by a fleet of one owner's agents: they plan, write the code, review each other, and ship behind a hard gate, a human directs and holds the irreversible gates (going public, outreach), the fleet does the engineering. Every change clears the same gate you can run yourself: scripts/ship-gate.sh, tests green, zero-dependency core, examples run, under conventional-commit discipline.

Security (v0.1)

The net layer is hardened against the obvious attacks (see tests/test_redteam.py):

  • Signed verdicts (Ed25519). Each bot signs its verdict over the specific option, binding the content to a key. verify_transcript(require_signed=True) then rejects any verdict that was altered, or that arrives unsigned when a signature was expected. Scope (read this): the check proves each signature is internally consistent with the pubkey carried in that record. It does not yet prove authenticity, that the pubkey is the owner's real key, because there is no trusted owner → key roster: a coordinator that assembles the transcript could substitute its own keypair. Treat signatures today as tamper-evidence, not third-party-provable identity.

  • Auth + rate limiting + input validation. Without a bearer token /consider returns 401; brute-force enumeration is throttled (429); oversized/malformed bodies are rejected (413/400). This closes the preference-extraction hole (an unauthenticated attacker previously reconstructed a bot's private red line by probing).

  • Outcome verification. verify_outcome(transcript) (in parley.consensus) recomputes the max-min winner from the recorded verdicts and checks it matches the announced decision, so a coordinator that finalizes a feasible-but-not-max-min (or infeasible) option is caught. Anyone can replay it over the public record — no private sheet needed. Pass expected_owners= the roster you expect, or a coordinator that drops one owner from every entry still passes. (verify_non_betrayal still only proves your own red lines held.)

Not yet (v0.1 honest limits, do not treat as production-secure for adversarial principals):

  • Authenticity pinning, signatures verify against a self-asserted key, not a trusted roster (above); the roster-pinned check (verify against the key from each bot's discovery Agent Card) is v0.2. Today only the coordinator's own side pins it: the client checks that a signature is present under the card's key; whether the signature is valid is checked by verify_transcript(require_signed=True), which examples/run_env.py now runs and any other caller must run.

  • Replay binding, verdicts carry no session/nonce, so a signed verdict is replayable into another parley that reuses the same option.

  • Range-masked scores, the soft cardinal score is public in the transcript, so an untrusted coordinator can infer preference ordering and each party's feasible region (the reason and the private sheet stay hidden; MPC/range-masking is future work).

  • TLS/mTLS, and game-theoretic collusion / strategic-misreport resistance (the research track).

The hosted demo at parleyprotocol.com uses GA4 for basic traffic analytics — see what's collected. The protocol and the self-hosted docker run path collect nothing.

Where's the money (open-core thesis)

The consensus mechanism itself is a commodity (voting/consensus is free, Snapshot, polls). Value, and willingness to pay, scales with decision stakes × number of parties × need for privacy / neutrality / audit. The paid layer is the trusted neutral broker: hosted identity/trust-registry, audit-grade signed transcripts, and managed self-hosting, not the algorithm. First candidate segments: private multi-party B2B negotiation (procurement/SOW) and auditable delegated governance (committees, panels). Differentiator vs Fetch.ai / Olas (real prior art): they do bilateral commerce on a blockchain; Parley does multilateral group consensus, provable non-betrayal, no crypto, self-hosted.

The demo shows three people whose agents hold private, conflicting constraints reach a slot everyone's red lines allow, then each owner proves non-betrayal, and a second run that hits an honest deadlock instead of forcing a bad decision.

Architecture

Layer

v0

Next

Transport / discovery / identity

in-process

A2A (signed Agent Cards, mDNS/registry), reuse, don't rebuild

Agent brain

pure code, or parley/elicit.py (LLM drafts, participant confirms)

wider elicitation UX

Red-line enforcement

parley/preferences.py (code predicates)

the core; stays deterministic

Consensus

parley/consensus.py (max-min)

Nash bargaining, weighted rules

Verifiability

parley/transcript.py (hash + local replay), optional Ed25519 signing

range-masked scores (MPC)

Adversarial

N/A

Byzantine/collusion resistance, the research-grade contribution

Roadmap → open-core

  • Now: consensus core + red-line guard + verifiable transcript (this repo, Apache-2.0).

  • Next: A2A transport so agents on different machines discover and parley over a LAN. (LLM-elicited sheets and Ed25519-signed transcripts already ship.)

  • Research: inject a lying/colluding agent and show max-min + Byzantine-robust aggregation resists it. This is the part potentially interesting to frontier R&D.

  • Paid (self-host): hosted relay/trust-registry so agents meet safely beyond the LAN, plus managed hosting of agents on your own hardware. Open-source core, paid federation.

First ICP: a closed group (a team/department/family) where one operator deploys all the agents, sidesteps the cold-start network effect. "Delegate the position to your agent, agents find the common decision" is exactly this.

Contributing

The core is small on purpose; the best contributions right now are adversarial tests and a second opinion on the consensus protocol. Start with CONTRIBUTING.md, file a bug with a reproducing recipe, or open a discussion. Report security issues privately via SECURITY.md.

Star history

If Parley's approach to provable non-betrayal is interesting, a star helps others find it.

Star History Chart

License

Apache-2.0, permissive, with an explicit patent grant (clears enterprise legal review). See LICENSE + NOTICE.

Available Tools

3 tools
parley_check_partyA

Replay one party's red lines against a decision: did the outcome cross any of them? Pass the party's spec (the same object you would put in parley_decide's parties) and the decision object returned by parley_decide. Returns JSON: owner and holds (true when every red line of that party holds on the decision). A null decision (deadlock) forces nothing on anyone, so holds is true. Soft preferences are not judged here, only red lines.

ParametersJSON Schema
NameRequiredDescriptionDefault
partyYesOne party: a distinct owner name, hard red lines, soft utility terms.
decisionYesThe decision from parley_decide, or null for a deadlock.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does well: it discloses the null-decision edge case (deadlock forces nothing, so holds is true), states that only red lines are evaluated, and describes the return shape. It does not explicitly say the tool is a side-effect-free computation, but 'replay' plus the read-only framing implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the question the tool answers, then the call contract, then return semantics and the edge case. Every sentence carries information; nothing is redundant with the schema or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, nested-object tool with no output schema, the description supplies the missing pieces: the relationship to sibling tools, the null/deadlock behavior, and the returned fields (owner, holds). An agent has everything needed to invoke and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine cross-tool meaning: `party` is the same object used in parley_decide's parties, and `decision` is parley_decide's output or null for a deadlock. That interop framing is not visible in the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: replay one party's red lines against a decision and report whether any were crossed. The 'did the outcome cross any of them?' framing makes the operation concrete and separable from parley_decide (which produces the decision) and parley_verify_receipt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context: pass the party spec that would go into parley_decide's parties plus the decision parley_decide returned. It also bounds scope by excluding soft preferences ('not judged here, only red lines'), giving a de-facto when-not condition. No explicit comparison to parley_verify_receipt, but the use case is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parley_decideA

Reach a decision among parties with conflicting interests. Give the options and each party's red lines (hard constraints, checked in code) and soft preferences. An option that crosses any party's red line is rejected outright; among the options acceptable to everyone the max-min rule picks the one whose least-satisfied party is best off, tie-broken by total score. No option acceptable to all parties yields an honest status 'deadlock' with decision null, never a forced choice. Returns JSON: status ('agreed'|'deadlock'), decision (the winning option object or null), transcript (every option with each party's masked verdict), and transcript_sha256 (the receipt hash; keep it to verify the transcript later with parley_verify_receipt). HONESTY NOTE: in this mode you, the caller, hold every party's spec and pass them all in one call, so there is no privacy between parties or from you. What still holds: red lines reject an option deterministically (a violating option cannot win), the max-min rule picks among options feasible for everyone, and the returned transcript is tamper-evident via transcript_sha256, provided the hash is kept by someone other than whoever might edit the transcript (it is unsigned and this server does not store it). For private sheets run one process per owner: examples/run_env.py.

ParametersJSON Schema
NameRequiredDescriptionDefault
ruleNoSelection rule. Only 'egalitarian' (max-min) exists.egalitarian
specYesA DecisionSpec: the shared options plus every party's position, the same JSON as examples/demo/recipe_*.json. Extra keys are ignored.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses critical behavior: deterministic red-line rejection, max-min selection among feasible options, honest 'deadlock' status with null decision (no forced choice), and that the transcript is tamper-evident via hash. It also states limitations: no privacy in this mode, hash is unsigned and not stored by server. Missing: whether red-line rejection is one violating option per party or any violating option (it says 'any party's red line is rejected outright', which is clear), and what happens with missing attributes (schema says fails closed). It doesn't mention performance or limits (e.g., max 500 options, 50 parties) but these are in schema. Overall very strong, but not quite 5 because it doesn't explain what 'masked verdict' means in the transcript (only says 'masked verdict' without defining).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, with a detailed honesty note that is valuable but could be more front-loaded. It mixes purpose, mechanics, return values, and caveats in one block. While every sentence is informative, the structure could be improved by separating the return-value details and honesty note into clearer sections. It's not egregiously verbose, but it's not concise either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (nested spec with options, parties, hard/soft constraints), no output schema, and no annotations, the description covers everything an agent needs: what the tool does, when to use it, the algorithm, return format, verification path, and honesty limitations. It even points to an example file. Complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters in detail (including enums, nested structures, defaults). The description adds context about the overall spec structure ('Give the options and each party's red lines... and soft preferences') but no additional syntax or semantics beyond what's in the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states the exact purpose: 'Reach a decision among parties with conflicting interests.' It specifies the mechanism (red lines + max-min) and outcome types, making it clearly distinguishable from siblings parley_verify_receipt (checks hashes) and parley_check_party (validates a single party).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this mode ('you, the caller, hold every party's spec and pass them all in one call') and when not to ('For private sheets run one process per owner: examples/run_env.py'). It also directs to parley_verify_receipt for later verification. This is comprehensive alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parley_verify_receiptA

Two independent checks on a transcript from parley_decide. Returns JSON: match, max_min_verified, expected_sha256, recomputed_sha256. match=true means only that this transcript hashes to the sha256 you passed. The hash is unsigned and this server does not store it, so if the transcript and the hash come from the same party, match proves nothing: whoever edits the record can re-hash it. It detects tampering only when the hash was kept by someone other than whoever could edit the transcript. max_min_verified recomputes the decision from the recorded verdicts: true when the announced result is exactly the max-min option over them (or an honest deadlock), false when the announced winner does not follow from the recorded verdicts. Verdicts are unsigned, so whoever can rewrite the record can rewrite them to fit a swapped winner; only a hash kept by another party, or signed verdicts, rules that out. It needs every entry to carry exactly one verdict per owner. Pass expected_owners (the roster you know took part) so a record missing a whole owner also fails; that roster is only as trustworthy as where you got it. A malformed transcript is a tool error, not a mismatch.

ParametersJSON Schema
NameRequiredDescriptionDefault
sha256YesThe transcript_sha256 you were given, hex.
transcriptYesThe `transcript` object returned by parley_decide, unmodified.
expected_ownersNoOptional. The owners you know took part; every entry must carry exactly this set for max_min_verified to be true.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it explains the trust model for match, states that the hash and verdicts are unsigned and stored nowhere, spells out exactly what max_min_verified recomputes and when it is false, and distinguishes a malformed transcript (tool error) from a genuine mismatch. This is unusually rich behavioral disclosure for a verification tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and return shape are front-loaded, and every subsequent sentence adds distinct semantics (trust caveats, recomputation rules, error behavior) rather than padding. The one drawback is that it is a single dense paragraph of caveats, which makes scanning for a specific fact harder than a lightly structured version would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly enumerates the return fields (match, max_min_verified, expected_sha256, recomputed_sha256), interprets them, and covers the error case. For a nested-input verification tool with no annotations, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains that sha256 is only trustworthy if held by another party, and that expected_owners exists so a record missing a whole owner also fails. These are semantic stakes the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific operation (two independent checks) on a specific resource (a transcript from parley_decide), and it explicitly identifies the upstream tool that produces the input. An agent can tell this apart from parley_decide and parley_check_party without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives strong conditional guidance on when the result is meaningful (only when the hash was kept by a different party than whoever could edit the transcript) and recommends passing expected_owners to catch missing owners. It does not explicitly contrast the use case against the sibling parley_check_party, so it stops short of full when/when-not/alternatives coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.2.0
    • First observedparley_check_party
    • First observedparley_decide
    • First observedparley_verify_receipt

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation4/5

The three tools have distinct targets: parley_decide produces a decision, parley_verify_receipt re-checks the full transcript/hash and max-min, and parley_check_party replays one party's red lines. There is mild overlap between the two verification tools, but the descriptions clearly delineate what each checks.

Naming Consistency5/5

All three names follow a single, predictable pattern: parley_ prefix plus a consistent verb_noun form (decide, verify_receipt, check_party). No mixed conventions or vague verbs.

Tool Count4/5

Three tools is on the thin side, but the domain is a narrow, self-contained decision engine, and decide/verify/check each earns a clear place. It is a touch under-provisioned rather than bloated.

Completeness4/5

The surface covers the core lifecycle: reach a decision, verify the recorded transcript, and re-check a single party. Privacy-mode execution is delegated to examples/run_env.py rather than a tool, and there is no batch check, but nothing leaves an obvious dead end for the stated purpose.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    An MCP server that enforces fail-closed deterministic checks, independent refute-first review, and tamper-evident hash-chained receipts for AI agent outputs before claiming completion.
    4
    3
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Cryptographically anchored, tamper-evident evidence receipts for AI agents — verified run receipts, existence-at-time proofs, and cited answers from an anchored public record. Remote MCP with proof-gated settlement; attests existence and integrity, never truth.
    -