Skip to main content
Glama

Derbent

One guarded pass for all your coding agents.

Derbent MCP server – quality and maintenance score on Glama

Every tool call that reaches Derbent, MCP or a CLI's built-in tools, is decided by one policy and written to a tamper-evident log, and the calls you care about wait for you.

A terminal session: derbent explain shows that git push origin main matches an ask rule; Claude Code
hook calls fed to derbent gate let go test run and deny rm -rf by a rule; derbent receipts lists both
calls and derbent verify finds the chain intact; after one receipt is edited with sqlite3, derbent verify
names receipt 2 as where the chain breaks

Derbent guards against mistakes and prompt injection, and keeps an audit trail. It is not a sandbox: an agent that can already run shell commands as you can get around it (see Limits).

In Turkish history, a derbent was a guarded post on a mountain pass: its keepers decided who went through and kept a record of everyone who did. Derbent does the same for your coding agents' tool calls.

Claude Code, Codex, GitHub Copilot CLI and Antigravity CLI connect to Derbent as one MCP server, and your other MCP servers sit behind it. Each CLI's pre-tool hook sends its built-in tools, such as the shell and file edits, through the same gate. Every call is decided by your rules and written down.

Support for Codex and Antigravity CLI is experimental (configured from their docs, not yet checked in a real session).

Four coding agent CLIs send MCP calls to derbent mcp and built-in tool calls to the derbent gate hook; both write a receipt to derbent.db, which the derbent terminal UI reads to show and approve calls, allowed MCP calls go on to your MCP servers, and a built-in call that is not denied runs in the CLI, which still applies its own permission check to a rule allow

It is one binary for Windows, macOS and Linux. There is no daemon and no network listener: a SQLite file is the only shared state (ADR 0001).

What you get

  • Rules that allow, deny or ask, by agent, tool and argument. The first match wins. The same rules apply to MCP tools and to built-in tools such as the shell, in all four CLIs, and a repository can add project rules that only make them stricter.

  • Approvals. A call your rules ask about waits until you press a in the terminal UI, or is denied after 50 seconds by default.

  • Receipts. Every call gets a hash-chained receipt, and derbent verify names the first receipt where the chain breaks after one was edited, moved, inserted or removed. Keep the head hash it prints to catch the newest ones being deleted too.

Also:

  • Shared memory and handoffs. Agents keep notes per repository and can leave tasks for each other.

  • Pins hold back a downstream server's tool when its definition changes.

  • Budgets stop an agent stuck in a loop.

Related MCP server: AgentGate

Install

Download the binary for your system and SHA256SUMS from the latest release, check the hash, and put it on your PATH as derbent (derbent.exe on Windows):

# Linux on amd64; the other files differ only in the end of the name
curl -LO https://github.com/tunahanaliozturk/derbent/releases/download/v1.0.0/derbent-v1.0.0-linux-amd64
curl -LO https://github.com/tunahanaliozturk/derbent/releases/download/v1.0.0/SHA256SUMS
sha256sum --ignore-missing -c SHA256SUMS
install -m 755 derbent-v1.0.0-linux-amd64 ~/.local/bin/derbent

Or build it with Go 1.27 or later:

go install github.com/tunahanaliozturk/derbent/cmd/derbent@latest

Homebrew (brew install tunahanaliozturk/tap/derbent) and Scoop install it too; see Install and set up.

Every release can be rebuilt byte for byte from its tag. Install and set up has the other systems, how to check a release, and each CLI's entries by hand.

Quick start

derbent init --preset balanced   # add Derbent to each CLI it finds, and write a first config
derbent doctor                   # check what init wrote
derbent                          # open the terminal UI in a terminal of its own

derbent init shows every change and asks once before it makes any. It copies each file it changes first and never replaces an entry you already have. With --dry-run it only shows the changes. Two of them, from a real run with the paths shortened:

claude: ~/.claude.json
  copied first to ~/.claude.json.derbent-backup-20260930T205424Z
  runs:
    claude mcp add --scope user derbent -- ~/bin/derbent mcp --agent claude
derbent: ~/.config/derbent/config.toml
  a new file
  adds the balanced preset (derbent init --preset balanced --print shows it)

It also adds the derbent gate hook to each CLI's settings; Install and set up shows every entry. Start your agents as usual. Their calls now appear in the UI, and calls the rules ask about wait there for you. Built-in tools says what each CLI's hook covers.

To put your other MCP servers behind the gate, so each agent needs only the one derbent entry, see Downstream servers.

How a call is decided

A tool call is checked against your rules, where the first match wins, then against the project's rules, which can only make it stricter, then against budgets; the result is allow, ask, which waits for you, or deny, and every call gets one receipt

Rules live in config.toml in your user config directory (%AppData%\derbent\ on Windows, ~/.config/derbent/ on Linux, ~/Library/Application Support/derbent/ on macOS). A downstream MCP server's tools are named <server>__<tool>, and a CLI's built-in tools native__<tool>:

[[rule]]
tool   = "github__issue_read"
action = "allow"

[[rule]]
tool   = "github__*"          # every other GitHub tool waits for you
action = "ask"

[[rule]]
agent  = "claude"
tool   = "native__Bash"       # Claude Code's shell, through the hook
args   = { command = "git push*" }
action = "ask"

[[rule]]
action = "allow"              # the last rule has no condition

To see how a call would be decided before it happens:

derbent explain --agent claude --tool native__Bash --args '{"command":"git push origin main"}'

It prints each rule and why it matches or not, down to the first match, then the project's rules, budgets, pin and grant, and ends with verdict: ask (rule:3) for this call under the rules above. Rules covers globs, the three presets (watch, balanced, strict), project rules, budgets and rule suggestions.

Receipts

Each receipt stores who called which tool, the decision and a hash of the result, plus the hash of the receipt before it; derbent verify recomputes the chain and names the first broken receipt, and the head hash can be kept elsewhere to check an export

derbent verify                              # count, head hash, and whether the chain is intact
derbent receipts --agent codex --since 1h   # also --tool, --project, --limit, --json

A receipt keeps the arguments after masking secrets, and only the size and hash of what a tool returned (ADR 0004). Someone who can write the database could still rewrite the whole chain or delete the newest receipts, so keep the head hash somewhere else if you want to be able to tell later. Receipts and verify covers exports that can be checked without the database.

Commands

Command

What it does

More

derbent

Terminal UI: waiting calls, agents, live receipts, memory, grants

Approvals

derbent init, derbent doctor

Add Derbent to each CLI, then check the setup

Install

derbent mcp --agent <name>

The MCP server each CLI starts

Install

derbent gate --agent <name>

The pre-tool hook for built-in tools

Built-in tools

derbent pending, approve, deny

Answer waiting calls from any shell

Approvals

derbent grants, revoke

List and take back session grants

Approvals

derbent explain

Show how a call would be decided

Rules

derbent suggest

Print the rules your answers point to

Rules

derbent config check

Validate the config and list each server's tools

Servers

derbent pins

See and accept changed downstream tools

Servers

derbent receipts, verify

Read, export and verify receipts

Receipts

derbent handoffs

List the tasks agents left for each other

Memory and handoffs

derbent version

Print the release

Overhead

Measured on GitHub's hosted runners on 2026-10-01 (Linux on an Intel Xeon Platinum 8370C, Windows on an AMD EPYC 9V74, 4 vCPUs each), median of ten runs at p50, from docs/benchmark-results:

Linux

Windows

MCP tool call, direct to the server

256.5 µs

293.0 µs

MCP tool call, through the gate

756.0 µs

909.0 µs

Starting the binary and exiting

4.292 ms

40.40 ms

Hook call, derbent gate, allowed

6.421 ms

64.51 ms

The gate adds 499.5 µs to an MCP call on Linux and 616.0 µs on Windows, for the extra stdio hop, the rule decision, the project rules check and the receipt written to SQLite. A hook call costs about 2.1 ms more than starting the binary on Linux and 24 ms more on Windows, where most of its cost is the process start. The numbers come from one run on shared runners, and runs differ by more than one run's intervals: the day before, the same code on the same Windows CPU model put the gate 459.5 µs over a direct MCP call, against 616.0 µs here. The results page has p99, calls per second and the caveats.

Limits

The whole list, with the reasons, is under Known limits and risks. The ones to know first:

  • Only calls that pass through the gate are seen. Tools a CLI never shows its hook, such as Codex's hosted web search, are outside it.

  • The CLIs marked experimental at the top have entries and hook adapters that follow each CLI's documentation and have not been checked in a real session. Claude Code and Copilot CLI (1.0.88, on 2026-10-02) have.

  • Argument globs match strings, not meaning: git push* does not match cd repo && git push.

  • Approvals guard against mistakes and prompt injection inside MCP. They do not stop an agent that can already run shell commands as you: it can run derbent approve itself.

  • A receipt export shows its last run unchanged only against a head you kept, and never that nothing was left out of it.

  • A suggestion refuses the shells, launchers and operators it knows, but a text rule can be fooled and those lists cannot be complete, and an allow for a command prefix lets any options through.

  • A handoff's address is a label, not an identity: any agent that can call handoff_take can claim every open handoff addressed to *.

  • Approvals depend on you watching. Unattended, ask means denied after the timeout.

  • Pins trust the first definition they see, including a new tool that an update adds to a pinned server, so name the tools you allow for a server whose updates you do not review.

  • A budget can be passed by the calls in flight at the same moment.

  • CI runs the tests on Windows and Linux and only builds on macOS.

For teams

A team audit trail, with receipts synced off each machine and an export for a SIEM, is an idea, not a plan. If you would use it, say so in this discussion.

Docs

Changes are listed in the changelog. To build, test or send a change, see CONTRIBUTING.md. To report a vulnerability, see SECURITY.md.

Licence

Apache 2.0. See LICENSE.

Available Tools

7 tools
handoff_createB

Leave a task for another agent working on this project, addressed to its agent label, such as reviewer, or to * for any agent. Returns the handoff's id.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesthe agent label to hand the task to, such as reviewer, or * for any agent
bodyYesthe task itself, up to 16 KiB
tagsNooptional labels made of letters, digits, dashes and underscores
titleYesa short summary of the task

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It only mentions the return value (handoff id), omitting side effects, persistence/visibility, permissions, error behavior, or limits for this write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the core action and target agent concept. No wasted words, though the return-value sentence is somewhat redundant given the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Full schema coverage and an output schema mean parameters and return values are covered elsewhere. For a write tool with no annotations, however, the description leaves gaps around behavioral side effects and sibling routing, making it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already documented. The description repeats the 'to' semantics ('agent label ... * for any agent') but adds no additional meaning beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Leave a task') and resource ('handoff') with the target agent concept, so the action is clear. It distinguishes the tool from siblings like handoff_take or handoff_done only implicitly, without naming alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for delegating work to another agent by label, which is helpful context. However, it gives no explicit when-to-use versus handoff_take, handoff_list, or handoff_done, and no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_doneB

Mark a handoff you took as done, with an optional note on what you did.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesid of the handoff you took
noteNooptional: what you did, up to 4 KiB

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
stateYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only restates what the schema already says ('id of the handoff you took', optional note). It does not disclose whether the state change is reversible, whether it requires ownership/permission, whether it is idempotent, or what effect it has on handoff_list results. For a mutation tool with zero annotation coverage this is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the action, the target, and the optional input with no wasted words. Nothing is buried and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema means return values need not be described, and both parameters are covered. However, for a no-annotation mutation tool the description omits prerequisites (must have been taken), failure modes, and post-completion effects, leaving it only minimally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters ('id', 'note' with its 4 KiB limit) are documented in the schema. The description adds no syntax, format, or constraint detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('mark a handoff ... as done') and scopes it to 'a handoff you took', which cleanly separates it from handoff_take, handoff_create, and handoff_list. It stops short of explicitly naming those siblings, so it lands at a clear-but-undifferentiated 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'a handoff you took' implies the precondition that handoff_take must have run first, and 'with an optional note on what you did' clarifies the typical follow-up action. There is no explicit when-not guidance, no mention of the alternative of leaving a handoff open, and no statement of what happens if the handoff was never taken.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_listA
Read-only

List this project's handoffs addressed to you or to any agent, open ones unless state says otherwise. mine lists the ones you created instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
mineNolist the handoffs you created instead of those addressed to you
stateNoopen (the default), taken, done or all
all_projectsNolist every project's handoffs instead of only this one's

Output Schema

ParametersJSON Schema
NameRequiredDescription
noticeYes
handoffsYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful behavioral context the annotations do not: the default is open handoffs, the scope is this project unless all_projects, and it covers handoffs addressed to any agent. It does not describe result shape or ordering, which is partly excused by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the primary scope constraint front-loaded before the mine alternative. Every clause carries information the agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with a full output schema and 100% parameter coverage, the description supplies the defaults (open state, current project) that matter for correct invocation. Only result ordering/pagination behavior is unstated, a minor gap given the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema, including the meaning of mine and the allowed state values. The description reinforces the mine/state defaults but adds no syntax or format detail beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (handoffs) with explicit scope: this project, addressed to you or any agent. It also clarifies the default filter (open) and the alternate mode (mine), so an agent can tell what it returns. It stops short of explicitly naming the mutation siblings (handoff_create/take/done) it should not be confused with, but the read verb makes that distinction self-evident.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"open ones unless state says otherwise" and "mine lists the ones you created instead" imply when the tool applies and how to flip behavior, but there is no explicit when-to-use/when-not framing or reference to the sibling tools (e.g. use handoff_take to claim one). Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handoff_takeA

Take an open handoff addressed to you or to any agent, by its id, and read it in full. Only one agent can take a handoff.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesid of the handoff

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes
toYes
bodyYes
fromYes
tagsYes
stateYes
takenYes
titleYes
noticeYes
createdYes
projectYes
taken_byYes
from_sessionYes
taken_sessionYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the exclusive-take semantics ('Only one agent can take a handoff') and the 'open' state precondition, plus that the result is the full content. It omits what happens on conflict (error vs. silent failure), whether the handoff's state changes, and any authorization requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the action and the exclusivity constraint front-loaded. Nothing needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be spelled out, and the single parameter is fully documented. The description covers purpose, precondition, and the key concurrency rule, leaving only edge-case behavior (already-taken, permissions) unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 100% schema description coverage, so the baseline is 3. The phrase 'by its id' merely restates the schema field and adds no format, sourcing, or validity detail beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (take) and resource (handoff), scoped by id, and clarifies the ownership scope ('addressed to you or to any agent'). It does not name or contrast with siblings such as handoff_list or handoff_done, which would have made it a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: you need an id, presumably from handoff_list, and the handoff must be open. There is no explicit when-to-use vs. when-not guidance, no mention of alternatives, and no statement of what to do if the handoff is already taken.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_readA
Read-only

Read one saved note in full by its id.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesid of the note

Output Schema

ParametersJSON Schema
NameRequiredDescription
atYes
idYes
bodyYes
tagsYes
titleYes
authorYes
noticeYes
projectYes
superseded_byNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=false, so the safety profile is covered. The description adds one useful piece of behavioral context — 'in full' implies untruncated retrieval, unlike a search summary — but says nothing about what happens when the id is absent or invalid, which is the main failure mode for an id lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, front-loaded with the verb and resource, with no filler or redundancy. Every word contributes to identifying the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the single required parameter is fully documented. The only real gap is the absence of routing guidance against memory_search, which is minor for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single 'id' parameter, so the schema already carries the parameter documentation. The description's 'by its id' merely restates it without adding format, range, or sourcing guidance, which is the baseline-3 case for a fully documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Read'), resource ('one saved note'), and scope ('in full by its id'), which distinguishes it implicitly from the sibling memory_search (query-based retrieval vs single-item retrieval). It stops short of naming the alternative, so it is clear but not fully sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the phrase 'by its id' signals you must already have an id, which hints that this is the follow-up to a memory_search result. There is no explicit statement of when to use this versus memory_search, nor any exclusion or prerequisite.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_writeA

Save a note for the other agents working on this project: a decision, a finding, a convention. Returns the note's id. Pass supersedes to replace an older note.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesthe note itself, up to 16 KiB
tagsNooptional labels made of letters, digits, dashes and underscores
titleYesa short summary of the note
supersedesNoid of an earlier note this one replaces

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It does disclose two useful behaviors: the return value (the note's id) and the supersession semantics ('pass supersedes to replace an older note'). It says nothing about permissions, durability, deduplication, or whether superseding makes the old note unreadable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and purpose, then the return value, then the one non-obvious parameter behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists and the description still confirms the id return, and all four parameters are schema-documented. Adequate for a write tool, though lifecycle details (what superseding does to the old note, whether notes persist) remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented (body size cap, tag charset, supersedes meaning). The description only restates the supersedes intent, adding no format or constraint detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Save a note') plus the kinds of content the note holds (decision, finding, convention) and the destination (other agents on this project). It reads clearly as the write counterpart to memory_read/memory_search, but it never names a sibling to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives practical when-to-use context by enumerating the note categories (decision, finding, convention) and implying a shared, cross-agent audience. It stops short of explicit alternatives or exclusions, e.g. when to prefer handoff_create or how it relates to memory_read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedhandoff_create
    • First observedhandoff_done
    • First observedhandoff_list
    • First observedhandoff_take
    • First observedmemory_read
    • First observedmemory_search
    • First observedmemory_write

TDQS

A3.7/5.0

Scored across 7 tools

Disambiguation5/5

The seven tools split cleanly into two non-overlapping domains (handoffs and memory notes), and within each domain the operations target distinct lifecycle stages. handoff_create/take/list/done are unambiguous, and memory_read (by id) versus memory_search (by query) are clearly differentiated.

Naming Consistency4/5

All names use a predictable snake_case namespace_verb pattern (handoff_*, memory_*), which is easy to scan. The only minor wrinkle is handoff_done, which is a state adjective rather than a verb like the others, but consistency overall is strong.

Tool Count5/5

Seven tools is a well-scoped set for a coordination server: four handoff operations and three memory operations, each earning its place. Nothing is redundant or missing at the count level.

Completeness4/5

Coverage is solid: handoffs support create/list/take/done, and memory supports write/read/search plus supersedes for replacement. Minor gaps remain (e.g., no delete or list-all for memory notes, no cancel/release for a taken handoff), but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Local zero-trust permission gateway for AI agents. Enforces policy-based tool authorization, human approvals, scoped permissions, and cryptographically verifiable audit logs.
    4
    56 PyPI
    5
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Human-in-the-loop approval gateway for agent tool calls: agents request, policies decide, humans approve via Slack/Discord/web — with an OWASP-LLM-Top-10-tagged audit trail. Self-hostable.
    10
    18 npm
    33
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Safe, reversible tool execution for AI agents. It sits between an agent and its tool servers, adding contracts, dry-run planning, policy, approvals, saga execution, and rewind.
    MIT